Information processing apparatus and program

The information processing device and program identify speaking users in online conferences by displaying their images or videos in a separate area, overcoming the need for individual user microphones.

JP2025120200APending Publication Date: 2025-08-15FUJIFILM BUSINESS INNOVATION CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025088917
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing technologies fail to identify a user speaking in an online conference without using a microphone for each individual user.

Method used

An information processing device and program that indicate when a user is speaking by creating a second display area on the screen to show the user's image or video, even if the user is not initially logged in to the conference.

Benefits of technology

Enables identification of speaking users without relying on microphones, facilitating participation and communication in online conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025120200000001_ABST
    Figure 2025120200000001_ABST
Patent Text Reader

Abstract

To specify a user who is making a statement in an on-line conference without using a microphone for each of individual users attending the on-line conference.SOLUTION: A processor represents that when a user other than users logging in an on-line conference makes a statement in the on-line conference, the other user is making the statement in the on-line conference, and forms a second display region other than a first display region on a screen and also displays an image or a moving image showing other users in the second display region even when the image or moving image showing the other user is displayed in the first display region before the other user makes the statement.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device and a program. [Background technology]

[0002] Patent Document 1 describes a device that identifies the next speaker in a remote conference via a network.

[0003] Patent document 2 describes a device that inputs audio within a base where a communication device is located, captures footage of the base, and if a speech is made within the base, records the point of speech indicating the speaker's location along with the time, and if multiple speech points within the base are recorded within a specified period of time, determines a shooting range that includes the multiple recorded speech points, and transmits video of the determined shooting range to another communication device located at another base.

[0004] Patent Document 3 describes a device for identifying who is speaking to whom when three or more people in different locations hold an audio conference using a telephone line.

[0005] Patent document 4 describes a device that detects whether a conversation is taking place between participants in a meeting, records the voices emitted by the participants, extracts specific voices from the recorded voices based on the detection result of the conversation status, and creates minutes of the meeting using the specific voices.

[0006] Patent Document 5 describes a system that relaxes the restrictions on speakers who can speak simultaneously while switching speakers as the conference progresses.

[0007] Patent document 6 describes a system that creates a shared conference room that each terminal normally uses and individual conference rooms for individual use by specific groups of terminals, and provides audio conferences for each conference room to which each terminal belongs. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-146072 [Patent Document 2] Japanese Patent Application Laid-Open No. 2017-34312 [Patent Document 3] Japanese Patent Application Laid-Open No. 2001-274912 [Patent Document 4] Japanese Patent Application Laid-Open No. 2013-105374 [Patent Document 5] Japanese Patent Application Laid-Open No. 2009-33594 [Patent Document 6] Japanese Patent Publication No. 2020-141208 Summary of the Invention [Problem to be solved by the invention]

[0009] An object of the present invention is to identify a user who is speaking in an online conference without using a microphone for each individual user participating in the online conference. [Means for solving the problem]

[0010] The invention of claim 1 is an information processing device having a processor, which, when a user other than the user logged in to the online conference speaks during an online conference, indicates in the online conference that the other user is speaking, and even if an image or video representing the other user is displayed in a first display area before the other user speaks, forms a second display area other than the first display area on the screen and displays an image or video representing the other user in the second display area.

[0011] The invention of claim 2 is a program for causing a computer to operate in an online conference such that, when a user other than the user logged in to the online conference speaks, the fact that the other user is speaking is indicated in the online conference, and even if an image or video representing the other user is displayed in a first display area before the other user speaks, a second display area other than the first display area is formed on the screen and an image or video representing the other user is displayed in the second display area. [Effects of the Invention]

[0012] According to the inventions of claims 1 and 2, it is possible to identify users who are speaking in an online conference without using a microphone for each individual user participating in the online conference. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a block diagram showing a configuration of an information processing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of a server according to the present embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of a terminal device according to the present embodiment. [Figure 4] FIG. 1 is a diagram showing a user at a location α. [Figure 5] FIG. 10 is a diagram showing a user at location β. [Figure 6] FIG. 1 illustrates a user at location γ. [Figure 7] FIG. [Figure 8] FIG. [Figure 9] FIG. [Figure 10] FIG. [Figure 11] FIG. [Figure 12] FIG. [Figure 13] FIG. [Figure 14] FIG. [Figure 15] FIG. DETAILED DESCRIPTION OF THE INVENTION

[0014] An information processing system according to this embodiment will be described with reference to Fig. 1. Fig. 1 shows an example of the configuration of an information processing system according to this embodiment.

[0015] As an example, the information processing system according to this embodiment includes a server 10 and N terminal devices (N is an integer equal to or greater than 1). In the example shown in FIG. 1, the information processing system according to this embodiment includes terminal devices 12A, 12B, 12C, . . . , 12N. The number of terminal devices shown in FIG. 1 is merely an example, and one or more terminal devices may be included in the information processing system according to this embodiment. Hereinafter, when there is no need to distinguish between the terminal devices 12A, 12B, 12C, . . . , 12N, they will be referred to as "terminal devices 12." The information processing system according to this embodiment may include devices other than the server 10 and the terminal devices 12.

[0016] The server 10 and each terminal device 12 have a function to communicate with other devices. The communication may be wired communication using a cable, or may be wireless communication. That is, each device may be physically connected to another device via a cable to transmit and receive information, or may transmit and receive information via wireless communication. Examples of wireless communication include short-range wireless communication and Wi-Fi (registered trademark). Examples of short-range wireless communication include Bluetooth (registered trademark), RFID (Radio Frequency Identifier), and NFC. For example, each device may communicate with other devices via a communication path N such as a LAN (Local Area Network) or the Internet.

[0017] The server 10 provides an online service via a communication path N. A user can use the online service using a terminal device 12. For example, by using the online service, the user can transmit information such as sound, an image, a video, a character string, and a vibration to the other party.

[0018] Examples of online services include online meetings, services that provide content online, online games, online shopping, social networking services (SNS), or combinations of these. Online meetings are sometimes called web meetings, remote meetings, or video meetings. Examples of content include entertainment (e.g., concerts, plays, movies, videos, music, etc.), sports, and e-sports. For example, video distribution services and music distribution services are examples of services that provide content online. Users can enjoy entertainment online, or watch sports and e-sports.

[0019] An online service may be a service that uses a virtual space, or may not use a virtual space. Virtual space is a concept that contrasts with real space, and includes, for example, a virtual space realized by a computer, a virtual space formed on a network such as the Internet, a virtual space realized by virtual reality (VR) technology, or cyberspace. For example, a virtual three-dimensional space or a two-dimensional space is an example of a virtual space.

[0020] The server 10 also stores and manages account information for users who use online services. The account information is information for logging in to an online service and using the online service, and includes, for example, a user ID and a password. For example, by transmitting the account information to the server 10 and logging in to the online service, a user associated with the account information is permitted to participate in and use the online service. Of course, a user may be able to use an online service without registering his or her account information with the online service. Also, a user may be able to use an online service without logging in to the online service.

[0021] The terminal device 12 is, for example, a personal computer (hereinafter referred to as "PC"), a tablet PC, a smartphone, a wearable device (for example, AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, a hearable device, or a mobile phone.

[0022] An automatic response partner such as a chatbot may participate in the online service. For example, the automatic response partner functions as a response assistant that responds to user inquiries, receives user comments, analyzes the content of the comments, creates a response to the comments, and notifies the user. The automatic response partner is realized, for example, by executing a program. The program is stored, for example, in the server 10 or another device (for example, another server or the terminal device 12). The automatic response partner may be realized by artificial intelligence (AI). Any algorithm may be used for the artificial intelligence.

[0023] In the following, as an example, it is assumed that an online conference is used by a user, and that sound, images, video, text, vibrations, etc. are transmitted to the other party through the online conference.

[0024] The hardware configuration of the server 10 will be described below with reference to Fig. 2. Fig. 2 shows an example of the hardware configuration of the server 10.

[0025] The server 10 includes, for example, a communication device 14 , a UI 16 , a memory 18 , and a processor 20 .

[0026] The communication device 14 is a communication interface having a communication chip, a communication circuit, etc., and has a function of transmitting information to other devices and a function of receiving information from other devices. The communication device 14 may have a wireless communication function or a wired communication function. The communication device 14 may communicate with other devices by using, for example, short-range wireless communication, or may communicate with other devices via a communication path N.

[0027] The UI 16 is a user interface and includes at least one of a display and an input device. The display is a liquid crystal display, an EL display, or the like. The input device is a keyboard, a mouse, input keys, an operation panel, or the like. The UI 16 may also be a UI such as a touch panel that combines a display and an input device.

[0028] The memory 18 is a device that configures one or more storage areas for storing various types of information. The memory 18 is, for example, a hard disk drive, various types of memory (e.g., RAM, DRAM, ROM, etc.), other storage devices (e.g., optical disks, etc.), or a combination thereof. One or more memories 18 are included in the server 10.

[0029] The processor 20 is configured to control the operation of each part of the server 10. The processor 20 may include a memory. For example, the processor 20 provides online services to users.

[0030] The hardware configuration of the terminal device 12 will be described below with reference to Fig. 3. Fig. 3 shows an example of the hardware configuration of the terminal device 12.

[0031] The terminal device 12 includes, for example, a communication device 22, a UI 24, a memory 26, and a processor 28.

[0032] The communication device 22 is a communication interface having a communication chip, a communication circuit, etc., and has a function of transmitting information to other devices and a function of receiving information transmitted from other devices. The communication device 22 may have a wireless communication function or a wired communication function. The communication device 22 may communicate with other devices by using, for example, short-range wireless communication, or may communicate with other devices via a communication path N.

[0033] The UI 24 is a user interface and includes at least one of a display and an input device. The display is a liquid crystal display, an EL display, or the like. The input device is a keyboard, a mouse, input keys, an operation panel, or the like. The UI 24 may be a UI such as a touch panel that combines a display and an input device.

[0034] The terminal device 12 may also include an imaging device such as a camera, a microphone, and a speaker, and all or some of these may be connected to the terminal device 12. Also, earphones or headphones may be connected to the terminal device 12.

[0035] The memory 26 is a device that configures one or more storage areas for storing various types of information. The memory 26 is, for example, a hard disk drive, various types of memory (e.g., RAM, DRAM, ROM, etc.), other storage devices (e.g., optical disks, etc.), or a combination thereof. One or more memories 26 are included in the terminal device 12.

[0036] The processor 28 is configured to control the operation of each part of the terminal device 12. The processor 28 may include a memory.

[0037] For example, the processor 28 displays images, videos, character strings, etc. sent during the online conference on the display of the terminal device 12, produces sounds sent during the online conference from a speaker, sends images, videos, etc. generated by taking pictures with a camera to the other party during the online conference, and sends sounds picked up by a microphone to the other party during the online conference.

[0038] The terminal device 12 may include at least one of various sensors such as a sensor (e.g., a GPS (Global Positioning System) sensor) that acquires location information of the terminal device 12, a gyro sensor that detects orientation and attitude, and an acceleration sensor.

[0039] Each example of the present embodiment will be described below. The processor 20 of the server 10 or the processor 28 of the terminal device 12 may execute the processing according to each example, or the processor 20 and the processor 28 may cooperate to execute the processing according to each example. A part of a certain process may be executed by the processor 20, and another part of the process may be executed by the processor 28. The server 10, the terminal device 12, or a combination thereof corresponds to an example of an information processing device according to the present embodiment.

[0040] In this embodiment, multiple users are in the same place and an online conference is held at that place. The place is not particularly limited, and may be a closed space (for example, a room or a conference room) or an open space (for example, outdoors).

[0041] Example 1 A first embodiment will be described below. As an example, the same online conference is used at locations α, β, and γ. For example, the same online conference is used at locations α, β, and γ by using terminal devices 12 provided at locations α, β, and γ, respectively. A user at location α, a user at location β, and a user at location γ can exchange information with each other by using the same online conference. Note that the number of locations is merely an example.

[0042] Figure 4 shows a user at location α, Figure 5 shows a user at location β, and Figure 6 shows a user at location γ, where locations α, β, and γ are different locations.

[0043] As shown in FIG. 4, there are four users (e.g., users A, B, C, and D) in location α. As shown in FIG. 5, there are two users (e.g., users E and F) in location β. As shown in FIG. 6, there is one user (e.g., user G) in location γ. In this way, there are multiple users in location α and location β. Note that the number of users in each location is just an example.

[0044] User A uses terminal device 12A, user B uses terminal device 12B, user C uses terminal device 12C, user D uses terminal device 12D, user E uses terminal device 12E, user F uses terminal device 12F, and user G uses terminal device 12G. Each terminal device 12 may be provided with a camera, a microphone, and a speaker.

[0045] A display 30, a microphone 32, and a camera 34 are provided at location α. A speaker may also be provided at location α. The display 30, the microphone 32, the camera 34, and the speaker are shared by users A, B, C, and D at location α and used for online conferences. For example, a screen for online conferences is displayed on the display 30, and images of users participating in the online conference are displayed on the screen.

[0046] A display 36 is provided at location β. A microphone, a camera, and a speaker may also be provided at location β, and these may be shared by users E and F. The display 36, the microphone, the camera, and the speaker are used for an online conference. For example, a screen for an online conference is displayed on the display 36.

[0047] A display, a microphone, a camera and a speaker may also be provided at location γ, and these may be used for the online conference.

[0048] The same location may be determined, for example, based on the IP address of each terminal device 12, based on the physical location of each user or each terminal device 12, based on location information obtained by GPS, using a microphone or speaker, or by each user declaring their own location.

[0049] For example, the processor 20 of the server 10 groups multiple terminal devices 12 whose IP addresses are close to each other into one group, estimates that the multiple terminal devices 12 are installed in the same location, and estimates that multiple users using the multiple terminal devices 12 are in the same location. For example, if user identification information (e.g., a user ID or account information) for identifying a user using the terminal device 12 is registered in the terminal device 12, the processor 20 of the server 10 identifies the user using the terminal device 12 based on the user identification information. For example, if the IP addresses assigned to the terminal devices 12A, 12B, 12C, and 12D are closer to each other than the IP addresses assigned to the terminal devices 12E, 12F, and 12G, the processor 20 of the server 10 estimates that the terminal devices 12A, 12B, 12C, and 12D are installed in the same location α, and estimates that users A, B, C, and D are in the same location α.

[0050] As another example, the physical location of each user may be specified by a character string, a diagram, or the like. For example, an image showing the seating layout of each location is displayed on a display, and the user specifies their own seat and the seats of other users on the image. The processor 20 of the server 10 recognizes the location of each user based on the specification. For example, if user A specifies the seats of users A, B, C, and D on an image showing the seating layout of location α, the processor 20 of the server 10 recognizes that users A, B, C, and D are in the same location α. Furthermore, if user A specifies the seats of users A, B, C, and D and assigns the user identification information of users A, B, C, and D to each user's seat, the processor 20 of the server 10 associates the user's seat location with the user identification information of the user and manages them. This allows management of who is in which seat.

[0051] As another example, the processor 20 of the server 10 may detect the location of each user using each terminal device 12 based on the location information of each terminal device 12 (e.g., location information acquired by GPS), and determine whether each user is in the same location based on the location of each user. For example, if the location information of each of users A, B, C, and D indicates a location within location α, the processor 20 of the server 10 may infer that users A, B, C, and D are in the same location α. The processor 20 of the server 10 may infer that multiple users whose locations are close to each other compared to the locations of other users are in the same location. For example, if the locations of users A, B, C, and D are close to each other compared to the locations of users E, F, and G, the processor 20 of the server 10 may infer that users A, B, C, and D are in the same location.

[0052] As another example, the processor 20 of the server 10 may determine whether users are in the same location based on the on / off status of a microphone or speaker and location information of each user. For example, each user may wear a microphone, or each user's terminal device 12 may be equipped with a microphone. The processor 20 of the server 10 may detect the location of each user using GPS or the like, and may also detect whether each user's microphone is on or off. If only one of multiple users who are close to each other has their microphone on, the processor 20 of the server 10 may estimate the multiple users as a single group and estimate that the multiple users are in the same location. The processor 20 may also determine whether multiple users are in the same location based on the on / off status of a speaker instead of a microphone. Note that if a user is wearing earphones or headphones as a speaker, it is usually assumed that the speaker is on, making it difficult to estimate a group based on the speaker. In this case, the group is estimated based on the on / off status of the microphone.

[0053] As another example, a user may voluntarily report their location. For example, a user may input their location using their terminal device 12, or may say "I'm at ____" at the start of an online conference. The processor 20 of the server 10 may receive the input and detect the location of each user, or may receive the utterance and detect the location of each user.

[0054] The processor 28 of the terminal device 12 may execute the processing by the processor 20 of the server 10 described above to determine who is in which location.

[0055] In the following example, users D, F, and G log in to and participate in the same online conference. For example, user D logs in to and participates in the online conference using terminal device 12D, user F logs in to and participates in the online conference using terminal device 12F, and user G logs in to and participates in the online conference using terminal device 12G. Note that multiple users may each log in to and participate in the online conference using the same terminal device 12.

[0056] Users A, B, and C are participating in the online conference at the same location α as user D without logging in to the online conference. User D is participating in the online conference at the same location β as user E without logging in to the online conference. For example, a user who has obtained permission to participate from a logged-in user may be able to participate in the online conference without logging in.

[0057] A user who logs in to an online conference is assigned a display area formed on the screen for the online conference, and the display area displays images or videos captured by a camera linked to that display area, or images or videos (e.g., icons or avatars) that schematically represent the user assigned to that display area. Along with the image or video being displayed, or without the image or video being displayed, a character string for identifying the user (e.g., name, user ID, account, or nickname) may be displayed. A user who is not logged in to the online conference is not assigned that display area.

[0058] For example, a screen for the online conference is displayed on the display of each of the terminal devices 12D, 12F, and 12G that are logged in to the online conference. A display area assigned to user D, a display area assigned to user F, and a display area assigned to user G are displayed on the screen for the online conference.

[0059] Furthermore, at location α, display 30 is used for the online conference, and at location β, display 36 is used for the online conference, and a screen for the online conference is displayed on displays 30 and 36. For example, display 30 is connected to terminal device 12D and used for the online conference, and display 36 is connected to terminal device 12F and used for the online conference.

[0060] Furthermore, a screen for the online conference may also be displayed on the display of the terminal device 12 of a user who is not logged in to the online conference. For example, a screen for the online conference is displayed on the display of the terminal device 12 of a user who joined the online conference in which users D, F, and G are participating without logging in, and the users who joined the online conference without logging in can share the screen for the online conference.

[0061] In the following, as an example, a screen for an online conference is displayed on the display of the terminal device 12 of each user and on the displays 30 and 36.

[0062] For example, a camera 34 installed at location α is linked to user D, who has logged in to and is participating in an online conference from location α, and images and videos captured by the camera 34 are displayed in a display area assigned to user D on the online conference screen. For example, the camera 34 is connected to terminal device 12D, and image and video data captured by the camera 34 is transmitted to each terminal device 12 via the terminal device 12D and the server 10, and displayed on the online conference screen on the display of each terminal device 12. Images and videos captured by a camera (i.e., a built-in camera) of terminal device 12D instead of the camera 34 may be displayed on the online conference screen. Note that instead of images and videos captured by the camera 34, a schematic image or video representing user D may be displayed, or a character string for identifying user D may be displayed.

[0063] Similarly, the camera of terminal device 12F (i.e., the built-in camera) or the camera installed at location β is linked to user F who has logged in to and is participating in the online conference from location β, and images and videos captured by the camera are displayed in a display area on the online conference screen assigned to user F. Instead of images and videos captured by the camera, a schematic image or video representing user F may be displayed, or a character string for identifying user F may be displayed.

[0064] Similarly, the camera of terminal device 12G (i.e., the built-in camera) or the camera installed at location γ is linked to user G who has logged in to and is participating in the online conference from location γ, and images and videos captured by the camera are displayed in a display area on the online conference screen assigned to user G. Instead of images and videos captured by the camera, a schematic image or video representing user G may be displayed, or a character string for identifying user G may be displayed.

[0065] Furthermore, the microphone 32 installed at the location a is connected to the terminal device 12D, and sound data collected by the microphone 32 is transmitted to the terminal devices 12F and 12G via the terminal device 12D and the server 10, and the sound is emitted from the speakers (i.e., built-in speakers) of the terminal devices 12F and 12G, respectively, or from speakers (i.e., external speakers) connected to the terminal devices 12F and 12G, respectively. Instead of the microphone 32, the microphone of the terminal device 12D may be used, or the microphones of the terminal devices 12A, 12B, and 12C may be used.

[0066] Similarly, sound data picked up by the microphone of terminal device 12F (i.e., the built-in microphone) or the microphone installed at location β is transmitted to terminal devices 12D and 12G via terminal device 12F and server 10, and the sound is emitted from the speakers of terminal devices 12D and 12G (i.e., the built-in speakers) or the speakers connected to terminal devices 12D and 12G (i.e., the external speakers).

[0067] Similarly, sound data picked up by the microphone of terminal device 12G (i.e., the built-in microphone) or the microphone installed at location γ is transmitted to terminal devices 12D and 12F via terminal device 12G and server 10, and the sound is emitted from the speakers of terminal devices 12D and 12F (i.e., the built-in speakers) or the speakers connected to terminal devices 12D and 12F (i.e., the external speakers).

[0068] The microphone and speaker may be worn by the user. For example, if the terminal device 12 is a wearable device such as a hearable device, the user may wear the terminal device 12. In this case, the speaker (for example, earphone or headphone) included in the terminal device 12 is worn on the user's ear, and the microphone included in the terminal device 12 is placed near the user's mouth.

[0069] When a user other than the user logged in to the online conference speaks during the online conference, the processor 20 of the server 10 indicates in the online conference that the other user is speaking. The processor 20 of the server 10 may generate a visual change indicating that the other user is speaking (e.g., displaying an image, video, or text indicating that the other user is speaking), may generate a sound indicating that the other user is speaking (e.g., a sound indicating the name, user ID, account, etc. of the other user), or may indicate that the other user is speaking through vibration. For example, if each user is wearing a hearable device, the processor 20 of the server 10 may notify each user that the other user is speaking via bone conduction.

[0070] For example, if an image or video representing another user who has spoken is not displayed on the online conference screen, the processor 20 of the server 10 displays an image or video representing another user who has spoken on the online conference screen. In this case, the processor 20 of the server 10 may display images or videos representing another user who has spoken and images or videos representing users who have not spoken separately. For example, the processor 20 of the server 10 may display images or videos representing another user who has spoken larger on the online conference screen than images or videos representing users who have not spoken, may decorate the images or videos representing another user who has spoken (for example, by surrounding the images or videos with a frame of a specific color or shape), may blink the images or videos representing the other user who has spoken, or may use other methods to make the images or videos representing the other user who has spoken stand out.

[0071] For example, when user A, who is not logged in to the online conference, speaks at location α, that is, when user A, who is at the same location α as user D, who is logged in to the online conference, speaks, processor 20 of server 10 displays an image or video representing user A on the screen for the online conference. For example, the image or video representing user A is displayed on display 30 as shown in FIG. 4, on display 36 as shown in FIG. 5, and on the display of terminal device 12G as shown in FIG. 6. Furthermore, when the screen for the online conference is displayed on each of the displays of terminal devices 12A to 12F, an image or video representing user A is also displayed on each of the displays of terminal devices 12A to 12F. As described above, the image or video representing user A may be displayed prominently.

[0072] An image or video representing user A may be displayed, or alternatively, a sound or vibration may be generated to indicate that user A has spoken, without displaying an image or video representing user A.

[0073] For example, if the microphone 32 is a directional microphone, the processor 20 of the server 10 can detect the direction from which the sound is coming from at location α based on the sound picked up by the microphone 32. The processor 20 of the server 10 can also detect the location of each user based on the pre-registered location of each user (for example, the location of each user's seat) and the location of each terminal device 12 acquired by GPS. If user A is in the direction from which the sound is coming from, the processor 20 of the server 10 infers that user A has spoken, and generates an image, sound, vibration, or the like indicating that user A is speaking.

[0074] As another example, the processor 20 of the server 10 may identify a user who is speaking based on information about the user's face. For example, images representing the faces of each user participating in an online conference are registered in advance in the server 10. The images representing the users' faces are linked to information for identifying the users. For example, an image representing the face of user A is linked to information for identifying user A and registered in advance in the server 10. The faces of each user are photographed by a camera, and the processor 20 of the server 10 estimates the user who is speaking based on images or videos generated by the photographing. For example, the processor 20 of the server 10 estimates that a user whose mouth is moving is the user who is speaking. Furthermore, the processor 20 of the server 10 identifies the user who is estimated to be speaking by comparing images representing each user registered in advance with images or videos generated by photographing and representing the user who is estimated to be speaking.

[0075] For example, the inside of location α is photographed by camera 34. When user A is speaking, processor 20 of server 10 infers that user A is speaking based on the images and videos captured by camera 34, and recognizes that the user making the speech is user A by comparing the captured images and videos of user A with an image of user A pre-registered in server 10.

[0076] As another example, the processor 20 of the server 10 may identify a user who is speaking based on the user's voice. For example, the voice of each user participating in an online conference is registered in advance in the server 10. The user's voice is linked to information for identifying the user. For example, the voice of user A is linked to information for identifying user A and registered in advance in the server 10. When the voice of the user who has spoken is picked up by a microphone, the processor 20 of the server 10 compares the picked up voice with the voices of each user registered in the server 10 to identify the user who is speaking.

[0077] Furthermore, the processor 20 of the server 10 may point a camera at the user who is speaking and capture an image of the user who is speaking with the camera. For example, when user A is speaking, the processor 20 of the server 10 points the camera 34 at user A to capture an image of user A, and displays the captured image or video on the online conference screen. Note that when the camera 34 is connected to the terminal device 12 (for example, the terminal device 12D), the processor 28 of the terminal device 12 may point the camera 34 at user A to capture an image of user A.

[0078] In the above example, the processor 20 of the server 10 identifies the user who is speaking, but the processor 28 of the terminal device 12 may also identify the user who is speaking. For example, when user A is speaking, the processor 28 of the terminal device 12 (for example, terminal device 12D) provided at location α may identify user A who is speaking.

[0079] The image or video representing the user making the speech (for example, an image or video representing user A) may be an image or video registered in advance in the server 10, or may be an image or video generated by capturing a picture with a camera while the user is making the speech. An image or video (for example, an icon or avatar) that schematically represents the user making the speech may be displayed.

[0080] As described above, when a user other than the user logged in to the online conference (e.g., user A) speaks, the fact that the other user is speaking is indicated and communicated to each user participating in the online conference. This makes it possible to identify a user who is speaking in an online conference without using a microphone for each individual user. In other words, it is possible to identify a user who is speaking without identifying the user based on the sound picked up by the microphone used by each individual user. For example, if at least one microphone (e.g., microphone 32 at location α) installed in the same location (e.g., location α) is turned on, it is possible to identify the user who is speaking.

[0081] Note that user D, who is logged in to an online conference, and users A, B, and C, who are not logged in to the online conference, can be considered users who share at least one device used to participate in the online conference. For example, a display 30 used for the online conference is provided at location α, and users A, B, C, and D share the display 30 to participate in the online conference. Furthermore, a microphone 32, a camera 34, and a speaker used for the online conference are provided at location α, and users A, B, C, and D share the microphone 32, camera 34, and speaker to participate in the online conference. In this way, users A, B, C, and D who are in the same location α share the same display 30, microphone 32, camera 34, and speaker provided at location α, while users who are in different locations β and γ do not share the display 30, microphone 32, camera 34, and speaker provided at location α with users A, B, C, and D. The same is true for locations β and γ.

[0082] The processor 20 of the server 10 may log in a user who is speaking without logging in to the online conference to the online conference. For example, if user A speaks when not logged in to the online conference, the processor 20 of the server 10 logs user A into the online conference. If user A's account information is pre-registered in the server 10, the processor 20 of the server 10 changes user A's login status from not logged in to logged in. As another example, the processor 20 of the server 10 may prompt user A to log in by displaying a login screen on the display of the terminal device 12A. User A can log in to the online conference by entering account information on the login screen. The processor 28 of the terminal device 12 may log in a user who is speaking without logging in to the online conference to the online conference. For example, if user A is speaking, the processor 28 of the terminal device 12A may log user A into the online conference.

[0083] When a microphone picks up sound when a user is not speaking, the processor 20 of the server 10 may infer that another user in the same location as the user is speaking. The processor 20 of the server 10 determines whether each user is speaking based on images or videos captured by a camera. For example, a camera (e.g., an in-camera) of the terminal device 12 captures the face of the user using the terminal device 12, and the processor 20 of the server 10 determines whether the user using the terminal device 12 is speaking based on the images or videos captured by the camera. For example, the number of users in the same location is registered in the server 10. The processor 20 of the server 10 determines whether each user is speaking, subtracts the number of users who are not speaking from the registered number, and infers that the remaining user is the user who is speaking. Note that the processor 28 of the terminal device 12 may also infer the user who is speaking.

[0084] A specific example will be given below. As shown in Fig. 5, there are two users (i.e., users E and F) at location β. The presence of two users at location β is registered in server 10. User F is logged in to the online conference, and user E is not logged in to the online conference.

[0085] Furthermore, the microphone of terminal device 12F that is logged in to the online conference is on, and the microphone of terminal device 12E that is not logged in to the online conference is off. In this case, when audio is picked up by the microphone of terminal device 12F and it is determined that user F is not speaking based on images or videos captured by a camera (for example, an in-camera) of terminal device 12F, processor 20 of server 10 estimates that user E, the remaining user, is speaking. Note that processor 28 of terminal device 12F may estimate that the user speaking is user E.

[0086] In addition, when only one user is present in the same location participating in an online conference, if sound is picked up when that user is not speaking, the processor 20 of the server 10 may stop picking up the sound.

[0087] A specific example will be described. As shown in FIG. 6, only one user G is present at location γ. The microphone of the terminal device 12G is turned on. In this case, if sound is picked up by the microphone of the terminal device 12G and it is determined that user G is not speaking based on images or videos captured by a camera (e.g., an in-camera) of the terminal device 12G, the processor 20 of the server 10 stops picking up sound by the microphone of the terminal device 12G. Stopping picking up sound means turning off the microphone, muting the microphone, or not outputting data of the picked up sound. If sound is picked up even though user G is not speaking, it is presumed that the picked up sound is sound that should not be transmitted to other users via the online conference. In this case, stopping picking up sound can prevent sound that should not be transmitted to other users from being transmitted to other users. Note that the processor 28 of the terminal device 12G may stop picking up sound by the microphone.

[0088] An example of a screen for an online conference will be described below with reference to Fig. 7. Fig. 7 is a diagram showing a screen 38 for an online conference. Fig. 7 shows a display 30 provided at location α, and the screen 38 is displayed on the display 30. Note that the same screen as the screen 38 is also displayed on the displays of the terminal devices 12A to 12G and the display 36 provided at location β.

[0089] Display areas allocated to users logged in to the online conference are formed on the screen 38. For example, users D, F, and G are logged in to the online conference. Display area 38A is allocated to user D, display area 38B is allocated to user F, and display area 38C is allocated to user G, and display areas 38A, 38B, and 38C are formed on the screen 38. Display area 38A displays images and videos captured by a camera associated with user D (for example, camera 34 installed at location α or the camera of terminal device 12D). Display area 38B displays images and videos captured by a camera associated with user F (for example, a camera installed at location β or the camera of terminal device 12F). Display area 38C displays images and videos captured by a camera associated with user G (for example, a camera installed at location γ or the camera of terminal device 12G). A character string for identifying the logged-in user may be displayed together with or without displaying an image or video. The displayed image or video may not be an image or video generated by capturing a photo with a camera, but may be an image or video that schematically represents the user.

[0090] Furthermore, information for identifying users who are logged in to the online conference (e.g., account information) may be displayed on the screen 38. In this example, users D, F, and G are logged in to the online conference, so information for identifying users D, F, and G is displayed on the screen 38.

[0091] For example, images and videos captured by the camera 34 installed at location α are displayed in the display area 38A. When user A, who is not logged in to the online conference, speaks at location α, the processor 20 of the server 10 displays in the display area 38A that user A has spoken. For example, if an image or video representing user A is not displayed in the display area 38A before user A speaks (for example, if user A has not been captured by the camera 34 and an image or video representing user A is not displayed in the display area 38A), the processor 20 of the server 10 displays an image or video representing the speaking user A in the display area 38A. The processor 20 of the server 10 may point the camera 34 at user A to capture user A and display the image or video representing user A generated by the capture in the display area 38A, or may display pre-registered images or videos representing user A in the display area 38A. In the example shown in FIG. 7, an image or video representing user A is displayed in the display area 38A. Display area 38B displays images and videos representing user F who is logged in to the online conference, and display area 38C displays images and videos representing user G who is logged in to the online conference.

[0092] If an image or video representing user A is displayed in display area 38A before user A speaks (for example, if user A is photographed by camera 34 and an image or video representing user A is displayed in display area 38A), processor 20 of server 10 may enlarge and display the image or video representing user A on screen 38, may decorate the image or video representing user A, may make the image or video representing user A blink, or may form another display area on screen 38 other than display areas 38A, 38B, and 38C and display the image or video representing user A in that other display area.

[0093] The processor 20 of the server 10 may display an image or video representing the user A making the comment on the screen 38, or may display a character string indicating that the user A is making the comment without displaying an image or video representing the user A.

[0094] Example 2 Hereinafter, a description will be given of Example 2. In Example 2, as in Example 1, users A, B, C, and D participate in an online conference at location α, users E and F participate in an online conference at location β, and user G participates in an online conference at location γ.

[0095] In the second embodiment, users A to G log in to an online conference and participate in the online conference. Each of the terminal devices 12A to 12G is provided with a camera (for example, an in-camera), and images and videos captured by the camera provided in each terminal device 12 are displayed on a screen for the online conference.

[0096] Fig. 8 shows a screen 38 for an online conference. The screen 38 shown in Fig. 8 is a screen displayed on the display 30 provided at location α. The same screen as the screen 38 is also displayed on the display 36 provided at location β and on the displays of each terminal device 12.

[0097] Because users A to G are logged in to the online conference, a display area is assigned to each of users A to G, and a display area for each user is formed on screen 38. As shown in FIG. 8, display areas 38A to 38G are formed on screen 38. Note that display areas for all logged-in users may be formed on screen 38, or display areas for some of the users may be formed on screen 38. For example, display areas for a predetermined number of users may be formed on screen 38.

[0098] Display area 38A is assigned to user A, and displays images and videos captured by the camera of terminal device 12A. Display area 38B is assigned to user B, and displays images and videos captured by the camera of terminal device 12B. Display area 38C is assigned to user C, and displays images and videos captured by the camera of terminal device 12C. Display area 38D is assigned to user D, and displays images and videos captured by the camera of terminal device 12D. Display area 38E is assigned to user E, and displays images and videos captured by the camera of terminal device 12E. Display area 38F is assigned to user F, and displays images and videos captured by the camera of terminal device 12F. Display area 38G is assigned to user G, and displays images and videos captured by the camera of terminal device 12G in display area 38G. Information for identifying the user may be displayed in each display area together with the images and videos, or without displaying the images and videos.

[0099] 8, an image or video representing a user is displayed in each display area. For example, an image or video representing user A is displayed in display area 38A. The image or video representing user A may be an image or video generated by capturing an image with the camera of terminal device 12A, or may be an image or video that schematically represents user A. The same applies to display areas 38B to 38G.

[0100] Additionally, a list of information (e.g., account information) for identifying users who are logged in to the online conference is displayed on screen 38. Here, as an example, users A to G are logged in to the online conference, so a list of the account information for users A to G is displayed.

[0101] When a user speaks after being designated, the processor 20 of the server 10 indicates in the online conference that the user is speaking. For example, the processor 20 of the server 10 displays an image or video displayed in a display area associated with the designated user, or changes the display format of the display area, to indicate that the designated user is speaking. Specifically, the processor 20 of the server 10 may enlarge the display area associated with the designated user to a size corresponding to the user's speech, enlarge the image or video displayed in the display area to a size corresponding to the user's speech, apply decoration to the display area corresponding to the user's speech (for example, surround the display area with a frame of a specific color or shape), blink the display area, image, or video, or make the display area, image, or video stand out by other methods.

[0102] For example, when a user speaks and the voice is picked up by a microphone, the display area associated with that user may be enlarged, decorations may be applied to that display area, or images or videos displayed in that display area may be enlarged.

[0103] In addition, the processor 28 of the terminal device 12 used by the speaking user may perform processing to highlight the display area, image, or video associated with the speaking user, or the processor 28 of the terminal device 12 that receives the audio data may perform that processing.

[0104] 8, user D is designated and making a statement, and display area 38D associated with user D blinks, display area 38D is decorated, and images and videos displayed in display area 38D blink. For example, the frame of display area 38D is displayed in a color (e.g., red) that corresponds to the statement made by user D. As another example, an image or video of user D may be displayed in an enlarged form.

[0105] The processor 20 of the server 10 may notify other users by sound or vibration that user D has been designated to speak. For example, the processor 20 of the server 10 may generate a sound indicating that user D has been designated to speak from the speaker of each terminal device 12, or may notify other users by bone conduction using a hearable device that a user has been designated to speak.

[0106] The next user to speak is designated by a user who spoke before the current user, or by an authority who has the authority to designate a speaker. The previous user may be the user who spoke immediately before the next user to speak, or may be a user who spoke even earlier. The authority may be, for example, the moderator or organizer of the online conference.

[0107] The user who will speak next may be designated, for example, on the screen 38, by a sound such as a voice, by a gesture such as pointing, or by a gaze.

[0108] When specifying the user who will speak next on screen 38, a display area associated with the user who will speak next may be specified, an image or video displayed in that display area may be specified, or account information of the user who will speak next may be specified from a list of account information. Processor 20 of server 10 accepts the specification and recognizes the user who will speak next. For example, if display area 38D associated with user D is specified, an image or video displayed in display area 38D is specified, or account information of user D is specified from a list of account information, processor 20 of server 10 accepts the specification and recognizes that user D is the user who will speak next.

[0109] When the next user to speak is designated by voice, the user who spoke first or an authorized person, etc., calls out the name, account information, nickname, etc. of the user who will speak next, and the voice is picked up by the microphone, and the processor 20 of the server 10 identifies the next user to speak based on the voice. For example, if the name of user D is called out by voice, the processor 20 of the server 10 identifies user D as the user who will speak next.

[0110] When the next user to speak is designated by a gesture such as pointing, when the user who spoke first or an authority points to the next user with a finger or arm, the scene is captured by a camera, and the processor 20 of the server 10 analyzes the captured image or video to identify the pointed user as the next user to speak. For example, when user D is pointed to, the processor 20 of the server 10 identifies user D as the next user to speak.

[0111] When the next user to speak is designated by gaze, when the user who spoke first or an authority turns their gaze toward the next user to speak, the scene is captured by a camera, and the processor 20 of the server 10 analyzes the image or video generated by the capture to identify the user in the direction of the gaze as the user who will speak next. For example, if the user in the direction of the gaze is user D, the processor 20 of the server 10 identifies user D as the user who will speak next.

[0112] Alternatively, the processor 28 of the terminal device 12 may identify the user who will speak next.

[0113] The processor 20 of the server 10 may set the length of time for a designated user to speak, and may forcibly end the user's speech when that time has elapsed. An end button may be displayed on the screen 38, and when the end button is pressed, the processor 20 of the server 10 may forcibly end the designated user's speech. The end of the designated user's speech may be instructed by voice. If the length of time that the designated user is silent exceeds a threshold, the processor 20 of the server 10 may forcibly end the user's speech. When the user's speech is forcibly ended, the processor 20 of the server 10 stops processing to indicate that the user is the speaker. Furthermore, if a next user is designated, the processor 20 of the server 10 indicates in the online conference that the next user is the user who will speak next.

[0114] The processor 20 of the server 10 may display information indicating that the user has been designated as the next user to speak in the online conference. This display method may be on the screen 38, as in the above-described method, or may be represented by a sound such as a voice, or may be represented by a vibration. This process will be described below with reference to Figures 9 to 12. Figures 9 to 12 show the screen 38 for the online conference.

[0115] As an example, users A, B, C, and D are participating in an online conference at location α, and users E and F are participating in an online conference at location β. Users A to F are logging in to the online conference and participating in the online conference.

[0116] 9 is a screen displayed on the display 30 provided at the location α. The same screen as the screen 38 is also displayed on the display 36 provided at the location β and on the displays of the terminal devices 12A to 12F.

[0117] Since users A to F are logged in to the online conference, display areas 38A to 38F are formed on screen 38.

[0118] In the example shown in FIG. 9, user A is speaking, and display area 38A is enlarged to a size corresponding to the fact that user A is speaking, and accordingly, the images and videos displayed in display area 38A are also enlarged. For example, the images and videos representing user A are enlarged and displayed. Furthermore, display area 38A may be decorated or flashing to indicate that user A is speaking. For example, the frame of display area 38A is displayed in a color (e.g., red) corresponding to the fact that user A is speaking.

[0119] User F is a user designated as the next user to speak. The processor 20 of the server 10 indicates in the online conference that user F is designated as the next user to speak. In other words, user F is a user reserved as the next user to speak, and the processor 20 of the server 10 indicates this reservation in the online conference.

[0120] For example, the processor 20 of the server 10 may display an image or video displayed in the display area 38F associated with user F, or change the display format of the display area 38F, to indicate that user F has been designated as the next user to speak. Specifically, the processor 20 of the server 10 may display the display area 38F in a size or color corresponding to the fact that user F has been designated as the next user to speak (e.g., a size or color corresponding to a reservation), or may display the image or video displayed in the display area 38F in a size or color corresponding to the fact that user F has been designated as the next user to speak, or may apply decoration to the display area 38F corresponding to the fact that user F has been designated as the next user to speak (e.g., decoration corresponding to a reservation), or may cause the display area 38F or the image or video to flash in response to a reservation. In this way, the processor 20 of the server 10 indicates that user F has been reserved as the next user to speak. Note that the fact that user F has been reserved as the next user to speak may be communicated to other users by sound, vibration, or the like.

[0121] In the example shown in Fig. 9, display area 38F is displayed in a color according to the reservation. For example, display area 38A associated with user A who is speaking is enlarged, and the frame of display area 38A is displayed in a color (e.g., red) that indicates that user A is speaking. The frame of display area 38F associated with user F who will speak next (i.e., reserved user F) is displayed in a color (e.g., blue) that corresponds to the fact that user F is reserved as the next user to speak. In this way, the user who is speaking and the user who will speak next are distinguished by color, size, decoration, etc.

[0122] Note that users who will speak third or later may be designated. In this case, images or videos may be displayed in colors corresponding to the order, or decorations corresponding to the order may be applied to the display area.

[0123] The processor 20 of the server 10 may gradually change the display mode of the display area 38F of user F, who is designated as the next user to speak, over time. The processor 20 of the server 10 may gradually increase the size of the display area 38F, or may gradually change the color of the frame of the display area 38F to red (i.e., the color that indicates that the user is speaking). For example, if the length of time for which user A speaks is set, the processor 20 of the server 10 may increase the size of the display area 38F or change the color of the frame of the display area 38F to red as the end time of user A's speech approaches.

[0124] For example, when time has passed since user F was designated as the next user to speak, display area 38F is enlarged from the size at the time of designation, as shown in FIG. 10, and accordingly, the images and videos displayed in display area 38F are also enlarged.

[0125] When user A's speaking time ends and user F's speaking time begins, as shown in Fig. 11, display area 38F is enlarged to a size corresponding to user F being the speaker, and accordingly, the images and videos displayed in display area 38F are also enlarged to a size corresponding to user F being the speaker. In the example shown in Fig. 11, because user A's speaking time has ended, the size of display area 38A is reduced to the size when no user is speaking, and accordingly, the images and videos displayed in display area 38A are also reduced to the size when no user is speaking.

[0126] If a user speaks without making a reservation, the processor 20 of the server 10 may indicate in the online conference that the user who made the statement is speaking. For example, in the situation shown in Fig. 11 (i.e., while user F is speaking according to a reservation), if user A speaks without making a reservation, the processor 20 of the server 10 enlarges the size of the display area 38A to a size that indicates that user A is the speaker, as shown in Fig. 12. The processor 20 of the server 10 may not output the statement of user A, who has not made a reservation, from the speakers of each terminal device 12, and may not change the display format of the display area 38A.

[0127] Example 3 Hereinafter, a description will be given of Example 3. In Example 3, similar to Example 1, users A, B, C, and D participate in an online conference at location α, users E and F participate in an online conference at location β, and user G participates in an online conference at location γ.

[0128] In the third embodiment, the order in which each user speaks is specified (for example, the order is reserved), and the processor 20 of the server 10 switches the user who will speak according to the order. In this case, the processor 20 of the server 10 may log in the user who will speak to the online conference.

[0129] For example, in the example shown in FIG. 4, it is determined that users A, B, C, and D will speak in this order, and this order is registered in the server 10. The processor 20 of the server 10 switches the user who will speak according to this order. For example, if the length of time each user will speak is determined, the processor 20 of the server 10 switches the user who will speak according to the length of time each user will speak. In this case, the processor 20 of the server 10 switches the user's image (which may be a video or a schematic image) displayed on the online conference screen according to this order. For example, the processor 20 of the server 10 displays the image of the user whose turn it is to speak on the online conference screen, and changes the displayed user image according to the order of the users' speech. In the above example, the images of users A, B, C, and D are switched in this order. Alternatively, the processor 20 of the server 10 may log in the user whose turn it is to speak to the online conference and log out the other users from the online conference. In the above example, the logged-in users are switched in the order of users A, B, C, and D. In that order, the accounts of users logged in to the online meeting will be switched.

[0130] The order of speech may be any order. For example, the same user may make multiple consecutive speeches, or a specific order may be set for different locations. For example, user A may make two consecutive speeches, followed by user B making three consecutive speeches. Alternatively, users A and B at location α may make speeches in this order, followed by user F at location β and user G at location γ.

[0131] Example 4 Hereinafter, a fourth embodiment will be described. In the fourth embodiment, users A, B, C, and D participate in an online conference at a location α, and users E and F participate in the online conference at a location β. Users A to F log in to the online conference and participate in the online conference.

[0132] In the fourth embodiment, when the order of speech of each user is specified (for example, when the order is reserved), the processor 20 of the server 10 may display an image of each user (which may be a video or a schematic image) in an online conference in a manner according to the order. For example, the processor 20 of the server 10 displays the image of each user by changing the color, size, arrangement, or a combination thereof according to the order.

[0133] For example, it is decided that users A, B, C, D, E, and F will speak in this order (for example, the order is reserved), and the order is registered in the server 10. The processor 20 of the server 10 displays the images of each user in a manner according to the order.

[0134] Display examples of images of each user are shown in Figures 13 to 15. Figures 13 to 15 show a screen 38 for an online conference displayed on the display 30. A screen similar to screen 38 is also displayed on the display 36 and the displays of the terminal devices 12A to 12F.

[0135] 13, display areas 38A, 38B, 38C, and 38D are formed on screen 38, and processor 20 of server 10 arranges each display area according to the order of speech. Processor 20 of server 10 also enlarges the display area of the user currently speaking (i.e., the display area of the user whose turn it is to speak) more than the display areas of the other users.

[0136] Because user A is the first user to speak and is the currently speaking user, display area 38A is enlarged more than the other display areas, and accordingly, the images and videos (for example, images and videos representing users) displayed in display area 38A are enlarged. Furthermore, processor 20 of server 10 may apply decoration to display area 38A corresponding to the fact that user A is the currently speaking user, or may express the image and video of user A with a color, light, or the like corresponding to the fact that user A is the currently speaking user.

[0137] Since user B is the second user to speak, user C is the third user to speak, and user D is the fourth user to speak, display areas 38B, 38C, and 38D are arranged in that order. Note that, due to space limitations on screen 38, images of users five and beyond are not displayed on screen 38.

[0138] Note that each display area may display a character string or the like indicating the order. For example, the number "1" is displayed in display area 38A, and the number "2" is displayed in display area 38B. The same applies to the other display areas.

[0139] A list of the account information of each user is also displayed on the screen 38, and in this list, the account information of each user is arranged in the order of their comments.

[0140] The length of time for each user's speech is set, and as the time for user B, who is to speak next, approaches, as shown in FIG. 14, processor 20 of server 10 enlarges display area 38B to a size corresponding to the fact that user B is the next user to speak, and enlarges and displays the images and videos displayed in display area 38B. In response to this, processor 20 of server 10 may change the arrangement of each display area. If space is secured on screen 38 as a result of this change in arrangement, display areas for users that were not previously displayed may be displayed on screen 38. In the example shown in FIG. 14, display area 38E associated with fifth user E is displayed on screen 38, and images, videos, etc. of user E are displayed in display area 38E.

[0141] When user A's speaking time ends and it is user B's turn to speak, processor 20 of server 10 enlarges display area 38B to a size that corresponds to user B being the currently speaking user, and accordingly enlarges and displays the images and videos displayed in display area 38B, as shown in Fig. 15. Furthermore, processor 20 of server 10 may apply decoration to display area 38B that corresponds to user B being the currently speaking user, or may represent the images and videos of user B with colors, lights, or the like that correspond to the fact that user B is the currently speaking user.

[0142] 15, the processor 20 of the server 10 may not display the display area 38A associated with the user A who has finished speaking on the screen 38. Of course, when the user A has finished speaking, the display area 38A may be reduced and displayed on the screen 38. (Other embodiments)

[0143] When a user introduces himself / herself at the start of an online conference, the processor 20 of the server 10 may identify the user based on the self-introduction and register the identified user as a user participating in the online conference. For example, if the user introduces himself / herself aloud, the processor 20 of the server 10 may identify the user by the voice. Furthermore, if the self-introduction includes information for identifying the user (e.g., name, user ID, account information, etc.), the processor 20 of the server 10 may identify the user based on that information.

[0144] If the beginning and end of a speech by a user who is speaking are specified by a user (e.g., a moderator, organizer, or authorized person), the processor 20 of the server 10 may switch the image of the speaking user in accordance with the specification.

[0145] The processor 20 of the server 10 may exclude a user who is manually inputting characters using an input device (e.g., a keyboard) of the terminal device 12 from candidates for users who will speak. For example, since a user who is typing on a keyboard is likely to be creating minutes or notes, etc., and is likely not to be speaking, the processor 20 of the server 10 excludes that user from candidates for users who will speak, and identifies a user who will speak from among users other than that user. For example, when identifying a user who will speak based on voice, image, etc., a user who is typing on a keyboard is excluded from candidates for users who will speak, and a user who will speak is identified from among users other than that user.

[0146] The processor 20 of the server 10 may exclude users who use application software other than the application software for using the online conference from candidates for users who will speak. For example, each user participates in an online conference by using application software for the online conference installed on their own terminal device 12. Application software other than the application software for the online conference is installed on the terminal device 12. A user who has launched and is operating application software other than the application software for the online conference is presumed to have no intention of participating in the online conference or to have a weak intention to do so, so the processor 20 of the server 10 excludes the user from candidates for users who will speak and identifies a user who will speak from among the other users.

[0147] In addition, since a user searching for information related to an online conference using a web browser is assumed to have an intention to participate in the online conference, the processor 20 of the server 10 does not need to exclude that user from the list of candidates for users who may speak.

[0148] The functions of each unit of the server 10 and the terminal device 12 are realized, for example, by a combination of hardware and software. For example, the processor of each device reads and executes a program stored in the memory of the device, thereby realizing the function of each device. The program is stored in the memory via a recording medium such as a CD or DVD, or via a communication path such as a network.

[0149] In the above embodiments, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Furthermore, the operations of the processor in the above embodiments may not only be performed by a single processor, but may also be performed by multiple processors located in physically separate locations working together. Furthermore, the order of the operations of the processor is not limited to the order described in the above embodiments, and may be changed as appropriate. [Explanation of symbols]

[0150] 10 servers, 12 terminal devices, 20,28 processors.

Claims

1. a processor; The processor: In an online conference, when a user other than the user who is logged in to the online conference makes a statement, the fact that the other user is making a statement is indicated in the online conference; Even if an image or video representing the other user is displayed in a first display area before the other user makes a comment, a second display area other than the first display area is formed on the screen, and the image or video representing the other user is displayed in the second display area.

1. An information processing device comprising:

2. The computer In an online conference, when a user other than the user who is logged in to the online conference makes a statement, the fact that the other user is making a statement is indicated in the online conference; Even if an image or video representing the other user is displayed in a first display area before the other user makes a comment, a second display area other than the first display area is formed on the screen, and the image or video representing the other user is displayed in the second display area. A program to make it work like this.

Citation Information

Patent Citations

  • Communication device, communication system, and program

    JP2017034312A

  • Public collaboration system

    US20130070045A1

  • Remote place conversation control method, remote place conversation system and recording medium wherein remote place conversation control program is recorded

    JP2001274912A

  • Remote conference system

    JP2009033594A

  • Next speaker guidance system, next speaker guidance method and next speaker guidance program

    JP2012146072A