Terminals, methods, and programs
The system addresses stress and engagement issues in remote meetings by using avatars that react naturally to participant engagement, enhancing participation and meeting dynamics.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- POPOPO INC
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-30
AI Technical Summary
Conventional remote meeting systems cause stress and discomfort due to direct participant image display, leading to poor engagement and reaction from others, and systems using avatars risk compulsion to maintain active attitudes.
A terminal-based system that uses avatars in a virtual space, collecting participant voices and determining their engagement levels to dynamically control the meeting display, including avatar movements and screen configurations to enhance participation and engagement.
Reduces stress and enhances participation in remote meetings by using avatars that react naturally to participant engagement, ensuring smooth and engaging meeting experiences.
Smart Images

Figure 2026123822000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a terminal, a method, and a program.
Background Art
[0002] In recent years, remote meetings using individual terminals have been actively held. In a remote meeting, a camera and a microphone are connected to a personal computer, and the images and voices of the participants are transmitted via a network. A mobile terminal such as a smartphone equipped with an in-camera may also be used.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a conventional remote meeting system that displays images of participants photographed by a camera side by side, there is a problem that many participants feel a sense of oppression because they are facing their own directions. Also, participating in the meeting while showing one's own figure seems to be stressful.
[0005] By turning off the camera and displaying icons representing the participants instead of the photographed images, the stress of being watched is reduced, but there is a problem that the reaction from other participants is poor and the speaker hardly feels a response.
[0006] The conference system described in Patent Document 1 represents conference participants with virtual avatars. In Patent Document 1, the system determines the level of engagement, which is an indicator of a participant's active attitude towards the conference, based on the participant's behavior acquired through a camera, and reflects this level of engagement in each participant's avatar. In Patent Document 1, since an avatar is displayed instead of the participant's actual appearance, the stress of being watched is reduced. However, since the level of engagement is determined for each participant and reflected in their avatar, there is a risk that stress may arise as participants feel compelled to maintain an active attitude in front of the camera.
[0007] This invention has been made in view of the above, and aims to provide a meeting system that reduces the stress of remote meetings, allows for easy participation, and enables smooth proceedings. [Means for solving the problem]
[0008] A terminal according to one aspect of the present invention is a terminal for participating in a meeting held in a virtual space where participants' avatars are placed, and comprises: a collection unit for collecting the voice of the participant; a transmission unit for transmitting the voice data of the participant; a reception unit for receiving the voice data of other participants; a determination unit for determining whether the participant is speaking based on the voice data; a display control unit for determining the screen configuration of the meeting based on the determination results of the participant and the other participants; and a display unit for playing the voice data and displaying the screen of the meeting according to the screen configuration. [Effects of the Invention]
[0009] According to the present invention, it is possible to provide a meeting system that reduces the stress of remote meetings, allows for easy participation, and enables smooth proceedings. [Brief explanation of the drawing]
[0010] [Figure 1] Figure 1 shows an example of the overall configuration of the conference system according to this embodiment. [Figure 2] Figure 2 is a functional block diagram showing an example of the configuration of a terminal in the conference system of this embodiment. [Figure 3] Figure 3 is a flowchart showing an example of the process flow when a terminal sends data. [Figure 4] Figure 4 is a flowchart showing an example of the process flow when a terminal displays the meeting screen. [Figure 5] Figure 5 shows an example of a meeting display screen. [Figure 6] Figure 6 is a flowchart showing an example of the process flow when a terminal displays the meeting screen. [Figure 7] Figure 7 shows an example of how an avatar is displayed during a conversation. [Figure 8] Figure 8 shows an example of how an avatar is displayed during a conversation. [Figure 9] Figure 9 shows an example of how an avatar is displayed during a conversation. [Figure 10] Figure 10 is a flowchart showing an example of the process for bringing avatars closer together during a conversation. [Figure 11] Figure 11 shows an example of how avatars are brought closer together during a conversation. [Figure 12] Figure 12 shows an example of a screen with icons placed on it. [Figure 13] Figure 13 shows an example of the screen displayed when a participant selects an icon. [Modes for carrying out the invention]
[0011] [Example 1] Embodiments of the present invention will be described below with reference to the drawings.
[0012] The conferencing system shown in Figure 1 is a system in which participants join a remote conference held in a virtual space using terminals 10. This conferencing system comprises multiple terminals 10 and a server 30 that are connected via a network for communication. Although only five terminals 10 are shown in Figure 1, the number of terminals 10 that can participate in a remote conference is arbitrary and not limited to this.
[0013] In the virtual space, avatars corresponding to each participant are arranged. An avatar is a computer graphics character representing a participant who participates in a remote meeting. The participant uses the terminal 10 to participate in the meeting in the virtual space using the avatar. Note that the meeting includes chats such as a sidebar meeting.
[0014] The terminal 10 collects the voice of the participant with a microphone, photographs the participant with a camera, and generates control data for controlling the movement and posture of the participant's avatar. The terminal 10 transmits the voice data and control data of the participant. The terminal 10 receives the voice data and control data of other participants, outputs the voice data, controls the corresponding avatar according to the control data, and displays the video obtained by rendering the virtual space. In addition, the terminal 10 determines the state of the participant and transmits the determination result, receives the determination result of the state of other participants from other terminals 10, and determines the display mode of the meeting based on the determination result of the participant and the determination result of other participants.
[0015] As the terminal 10, a personal computer connected with a camera and a microphone may be used, a mobile terminal such as a smartphone equipped with an in-camera may be used, or a virtual reality (VR) device equipped with a controller and a head-mounted display (HMD) may be used.
[0016] The server 30 receives the control data, voice data, and determination result from each terminal 10 and distributes them to each terminal 10.
[0017] Referring to FIG. 2, an example of the configuration of the terminal 10 will be described. The terminal 10 shown in FIG. 2 includes a collection unit 11, a photographing unit 12, a control unit 13, a determination unit 14, a transmission unit 15, a reception unit 16, a display control unit 17, and a display unit 18. Each unit included in the terminal 10 may be configured by a computer including an arithmetic processing device, a storage device, etc., and the processing of each unit may be executed by a program. This program is stored in the storage device included in the terminal 10, and can be recorded on a recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be provided through a network.
[0018] The collection unit 11 collects the participant's voice using the microphone provided by the terminal 10 or a microphone connected to the terminal 10. The collection unit 11 may also receive participant voice data recorded by other devices.
[0019] The shooting unit 12 takes pictures of the participants using the camera provided by terminal 10 or a camera connected to terminal 10. The captured video should preferably show the participants' faces, but it may also show the participants' whole bodies, or the participants may not be visible at all. The shooting unit 12 may also receive captured images taken by other devices.
[0020] The control unit 13 generates control data for controlling the participant's avatar. The control unit 13 may generate control data based on at least one of the participant's voice or a captured image. As a simple example, the control unit 13 generates control data to close the avatar's mouth when the participant is not speaking, and generates control data to move the avatar's mouth in response to the participant's speech when the participant is speaking. The control unit 13 may also determine the avatar's actions based on the participant's facial expression in the captured image.
[0021] Alternatively, the control unit 13 may generate control data without reflecting the participant's state. For example, if a participant is looking away from the meeting screen or has moved out of front of the camera, the control unit 13 may generate control data that causes the avatar to perform natural actions in a meeting, such as nodding or turning to face the speaker, without faithfully reflecting the participant's movements in the avatar. If a participant is actively participating in the meeting, such as looking at the screen and nodding, the control unit 13 may generate control data that reflects the participant's movements in the avatar. This ensures that the participant's avatar responds to the participant's actions in the meeting regardless of their state, allowing the speaker to speak comfortably.
[0022] The control unit 13 may use a machine learning model that has learned voice and avatar movements, and input the voice into the machine learning model to generate avatar control data.
[0023] When a VR device is used as terminal 10, the control unit 13 generates control data to control the avatar based on input from the controller and HMD. The participant's hand gestures, head movements, etc., are reflected in the avatar.
[0024] The determination unit 14 determines the participant's status from the captured image. Specifically, the determination unit 14 determines from the captured image whether the participant is looking at the meeting screen and whether the participant is present. The determination by the determination unit 14 does not need to be strict; for example, if the participant is using a smartphone as terminal 10, the determination unit 14 will determine that the participant is looking at the screen if the captured image shows the participant's face from the front. The determination unit 14 may also determine from the captured image or audio data whether the participant is speaking or not.
[0025] The transmission unit 15 transmits audio data, control data, and a judgment result. The judgment result is information indicating the status of the participant as determined by the judgment unit 14. For example, the judgment result may include states such as looking at the screen, not looking at the screen, being in front of the camera, not being in front of the camera, or speaking. The judgment result may also include time information such as the time spent looking at the screen, the time spent not in front of the camera, or the time spent speaking. The transmitted data is distributed to each terminal 10 via the server 30.
[0026] The receiving unit 16 receives voice data, control data, and judgment results from other terminals 10 via the server 30.
[0027] The display control unit 17 aggregates the determination results received from the determination unit 14 and other terminals 10, and determines the display mode of the meeting based on the aggregated results. The display mode includes, for example, the viewpoint when rendering the virtual space, the screen layout, the placement of objects, the movement and posture of the avatar, and various effects. Examples of aggregated results and display modes are given below.
[0028] If the percentage of participants not looking at the screen exceeds a predetermined threshold, the display control unit 17 will change the viewpoint when rendering the virtual space to a viewpoint that shows a close-up of the speaker in order to attract the participants' attention. At this time, the display control unit 17 may have the speaker's avatar perform a large action, such as banging on a table, or increase the volume of the speaker's voice. If the speaker's avatar is to perform a large action, the display control unit 17 will replace the control data of the speaker's avatar with the control data of the large action.
[0029] If the percentage of participants not looking at the screen exceeds a predetermined threshold and there are no speakers, the display control unit 17 will render the virtual space with a viewpoint that shows a close-up of the meeting organizer's (facilitator's) avatar in order to prompt a transition to the next topic or the end of the meeting.
[0030] If the majority of participants are looking at the screen, the display control unit 17 may render the virtual space from a viewpoint that overlooks the entire conference room, creating the illusion that the participants are listening attentively. The display control unit 17 may randomly select several avatars and have them nod. If an avatar is to nod, the display control unit 17 replaces the control data of the target avatar with control data for nodding.
[0031] In this way, by aggregating the status of participants and determining how the meeting will be displayed based on the aggregated results, the meeting can be conducted smoothly.
[0032] The display unit 18 plays back the received audio data and, in accordance with the instructions of the display control unit 17, places objects, including avatars, in the virtual space, controls the movement and posture of the avatars based on control data, and renders the virtual space to generate video of the meeting. For example, the display unit 18 places objects such as the floor, walls, ceiling, and tables that make up the conference room in the virtual space and places the participants' avatars in predetermined positions. The model data and placement positions of the objects are stored in the storage device provided by the terminal 10. Information necessary to construct the virtual space may be received from the server 30 or other devices when joining the meeting. If the instructions of the display control unit 17 include changing the position of objects, or changing the position and posture of avatars, the display unit 18 changes the position of objects, the position and posture of avatars according to those instructions. If the instructions of the display control unit 17 specify a viewpoint, the display unit 18 renders the virtual space from the specified viewpoint.
[0033] The display unit 18 may have operation buttons on the screen to accept input from participants. For example, when an operation button is pressed, control data is sent that causes the participant's avatar to perform the action corresponding to the operation button.
[0034] Furthermore, the server 30 may perform some of the functions of terminal 10. For example, the server 30 may have the functions of a display control unit 17, aggregate the judgment results from each terminal 10 to determine the display mode, and distribute the display mode to each terminal 10. Alternatively, the server 30 may have the functions of a control unit 13, a judgment unit 14, and a display control unit 17, receive captured images and audio data from each terminal 10, generate control data for each avatar, determine the state of each participant, aggregate the judgment results to determine the display mode, and distribute the control data and display mode to each terminal. The server 30 may also have the functions of a display unit 18 and distribute the rendered video of the virtual space to terminal 10.
[0035] Next, the processing flow of terminal 10 will be explained with reference to the flowcharts in Figures 3 and 4. The processes shown in Figures 3 and 4 are executed at each terminal 10 as needed.
[0036] Figure 3 is a flowchart showing an example of the process flow when terminal 10 sends data.
[0037] In step S11, the collection unit 11 collects the participant's voice, and the shooting unit 12 photographs the participant.
[0038] In step S12, the control unit 13 generates control data for controlling the participant's avatar.
[0039] In step S13, the determination unit 14 determines the participant's condition from the captured image or audio.
[0040] In step S14, the transmission unit 15 transmits voice data, control data, and a judgment result. The transmitted data is distributed to each terminal 10 via the server 30.
[0041] Figure 4 is a flowchart showing an example of the process flow when terminal 10 displays the meeting screen.
[0042] In step S21, the receiving unit 16 receives data transmitted by other terminals 10 from the server 30. The data received may include, for example, voice data, control data, and judgment results.
[0043] In step S22, the display control unit 17 aggregates the received judgment results.
[0044] In step S23, the display control unit 17 determines the display format of the meeting based on the aggregated results.
[0045] In step S24, the display unit 18 plays the audio data, controls the avatar according to the control data, and displays the conference screen according to the display mode.
[0046] Figure 5 shows an example of a meeting display screen. Figure 5(a) is an example of a screen showing the speaker's avatar. Figure 5(b) is an example of a screen showing an overview of the entire meeting room. Figure 5(c) is an example of a screen where the screen is divided into frames, with each frame displaying the avatar of each participant. The display mode of the screen may be determined by the aggregated results of the participant status determinations by terminal 10, or it may be determined randomly by terminal 10. All terminals 10 may display the screen in the same display mode, or they may not. In other words, each terminal 10 may individually determine the display mode, or the display mode determined by any terminal 10 may be distributed to all terminals 10 so that the display mode of all terminals 10 is the same.
[0047] [Example 2] In Example 2, the display mode of the meeting is determined by referring to the participant status determination result and past cuts. The overall configuration of the meeting system and the configuration of the terminal 10 in Example 2 are basically the same as in Example 1. In Example 2, the determination unit 14 determines whether a participant is in conversation, the display control unit 17 identifies the participant in conversation based on the determination result, and determines the cut of the avatar of the participant in conversation based on past cuts. In Example 2, the terminal 10 does not need to have a shooting unit 12.
[0048] Referring to the flowchart in Figure 6, the process by which terminal 10 in Example 2 displays the meeting screen will be explained. Note that the process by which terminal 10 sends data is the same as in Example 1.
[0049] In step S31, the receiving unit 16 receives data transmitted by other terminals 10 from the server 30.
[0050] In step S32, the display control unit 17 identifies the participants in the conversation based on the received determination result. For example, if participant A finishes speaking and another participant B starts speaking within a predetermined time, the display control unit 17 determines that participants A and B are in a conversation.
[0051] In step S33, the display control unit 17 determines the display format of the meeting based on past cuts. A specific example of the processing based on past cuts will be described later.
[0052] In step S34, the display unit 18 plays the audio data, controls the avatar according to the control data, and displays the conference screen according to the display mode.
[0053] Here is an example of processing based on past cuts. As shown in Figure 7, suppose that in the past, participant A's avatar A was displayed with a cut where it was facing to the right of the screen. The display control unit 17 stores the cuts used to display the avatars of participants in the past during a conversation. If participant A is the speaker in the conversation, the display control unit 17 sets the display mode to a cut where avatar A is facing to the right of the screen, similar to past cuts. If the other party in the conversation is participant B, when the display control unit 17 displays participant B's avatar B, it sets the cut so that avatar B is facing to the left of the screen, as shown in Figure 8, so that avatar A and avatar B are facing each other. Subsequently, when participant B speaks, the display control unit 17 sets avatar B to face to the left of the screen. The display control unit 17 may also control the avatar's posture.
[0054] If, in the past, both avatar A and avatar B were displayed with a rightward-facing shot, the display control unit 17 will display a screen showing both avatar A and avatar B, with avatar A facing right and avatar B facing left, as shown in Figure 9, for example. Subsequently, when participant A and participant B converse, the display control unit 17 will adjust the shot so that avatar A is facing right and avatar B is facing left. This allows participants to naturally understand who is talking to whom. Based on past shots, the display control unit 17 determines a display mode that allows participants to naturally understand who is talking.
[0055] If several participants are having a conversation, the display control unit 17 may identify the avatars involved in the conversation and determine the viewpoint so that the avatars involved in the conversation fit within a single screen. The display control unit 17 may also move the positions of the avatars in the virtual space so that they are closer to the other participants. Alternatively, the display control unit 17 may divide the screen into multiple regions and display the avatars involved in the conversation in each region.
[0056] The display control unit 17 may configure the screen differently for each participant using the terminal 10, depending on their role (speaker, facilitator, etc.). For example, the facilitator's screen may be divided into panels, displaying the speaker and the participant who is intently watching the screen. The facilitator can then view the screen and give the participant who is intently watching the screen an opportunity to speak.
[0057] [Differentiation] Next, I will explain the process of bringing avatars closer together during a conversation.
[0058] Referring to the flowchart in Figure 10, the process of bringing avatars closer together during a conversation will be explained. The process in Figure 10 is executed intermittently on each participant's terminal 10 during a conversation involving two or more people.
[0059] In step S41, terminal 10 determines whether the avatar of the participant operating terminal 10 and the avatar of the person they are talking to are in different locations. For example, it is determined that they are in different locations if the avatars in the conversation are at a predetermined distance apart in the virtual space. Alternatively, it may be determined that they are in different locations if another avatar exists between the avatars in the conversation. If the avatars in the conversation are not in different locations, the process ends.
[0060] If the avatars are far apart during the conversation, in step S42, terminal 10 determines, based on the type of terminal 10 itself, whether the participant can freely move their avatar. For example, a participant using a VR device as terminal 10 can freely move their avatar, but a participant using a smartphone as terminal 10 will find it difficult to move their avatar freely. Terminals 10 that can freely move their avatars will terminate the process. Alternatively, the types of terminals 10 of the participants in the conversation may be compared to determine whether terminal 10 will find it difficult to freely move its avatar. For example, if a participant using a personal computer as terminal 10 and a participant using a smartphone as terminal 10 are having a conversation, the personal computer has a keyboard and mouse connected, making it easier to move than a smartphone. Therefore, it may be determined that the avatar of the participant using a smartphone will be difficult to move freely.
[0061] If it is difficult to move the avatar freely, in step S43, terminal 10 moves the position of the participant's avatar closer to the person they are talking to.
[0062] In the example shown in Figure 11, avatar A of a participant using a VR device (hereinafter referred to as terminal 10A) as terminal 10 is conversing with avatar B of a participant using a smartphone (hereinafter referred to as terminal 10B) as terminal 10. In this case, terminal 10A determines in step S32 that avatar A can move freely, and terminal 10B determines in step S32 that avatar B cannot move freely easily. In step S33, terminal 10B moves avatar B's position closer to avatar A. When avatar B teleports, terminal 10B displays a warp effect (e.g., sparkles) at avatar B's position before and after the teleport to indicate that avatar B has teleported, and terminal 10A briefly darkens the screen to switch the cut.
[0063] Next, we will explain how participants can control their avatars via device 10.
[0064] As shown in Figure 12, terminal 10 may place icons 110 on screen 100 and accept input from participants. Each icon 110 has a graphic representing the action the participant wants the avatar to perform. When a participant touches an icon 110, terminal 10 generates and transmits control data for the action corresponding to the icon 110. The control data may include not only the avatar's actions but also the background, effects, and viewpoint.
[0065] Upon receiving the control data, terminal 10 controls the corresponding avatar according to the control data. If the control data includes a background, effects, and viewpoint, the terminal places the background and effects and sets the viewpoint in the virtual space according to the instructions in the control data. For example, Figure 9 is an example of screen 100 when a participant with an opinion selects an icon indicating that the avatar should raise its hand. In the example in Figure 13, the avatar raises its hand, a viewpoint is set that shows the avatar from the front, and a "!" effect is displayed above the avatar's head.
[0066] As described above, the terminal 10 of this embodiment is a terminal for participating in a meeting held in a virtual space where participants' avatars are placed, and includes a collection unit 11 for collecting participants' voices, a control unit 13 for generating control data for controlling participants' avatars, a determination unit 14 for determining the status of participants, a transmission unit 15 for transmitting participants' voice data, control data, and determination results, a reception unit 16 for receiving other participants' voice data, control data, and determination results, a display control unit 17 for determining the display mode of the meeting based on the determination results of the participant and other participants, and a display unit 18 for playing back voice data, controlling avatars based on control data, and displaying the meeting screen according to the display mode. As a result, participants can participate in meetings in the virtual space using avatars, reducing the stress of feeling watched, and the overall atmosphere of the meeting can be reflected in the meeting display by aggregating the status of participants and determining the display mode of the meeting. [Explanation of Symbols]
[0067] 10 devices 11 Collection Department 12 Photography Department 13 Control Unit 14 Judgment section 15 Transmitter 16 Receiving unit 17 Display Control Unit 18 Display 30 servers
Claims
1. A terminal for participating in a meeting held in a virtual space where participants' avatars are placed, A collection unit for collecting the voices of the aforementioned participants, A transmission unit that transmits the voice data of the aforementioned participant, A receiving unit that receives audio data from other participants, A determination unit that determines whether or not the participant is speaking based on the aforementioned audio data, A display control unit that determines the screen configuration of the meeting based on the determination results of the aforementioned participant and the aforementioned other participants, The system includes a display unit that plays the aforementioned audio data and displays the meeting screen according to the aforementioned screen configuration. Terminal.
2. The terminal according to claim 1, The display control unit stores past screen configurations, identifies the participant in the conversation based on the determination result, and determines the cuts of the participant's avatar based on the past screen configuration. Terminal.
3. The terminal according to claim 2, The display control unit determines the orientation of the avatar based on the avatar's past orientation. Terminal.
4. The terminal according to claim 1, The display control unit determines the screen configuration according to the role of the participant. Terminal.
5. A method for participating in a meeting held in a virtual space where participants' avatars are placed, Computers The audio of the aforementioned participants was collected, The audio data of the aforementioned participant is transmitted, Receive audio data from other participants, Based on the aforementioned audio data, it is determined whether or not the participant is speaking. The screen layout of the meeting is determined based on the judgment results of the aforementioned participant and the aforementioned other participants. The audio data is played back, and the screen of the meeting is displayed according to the screen configuration. method.
6. A program for operating a computer as a terminal according to any one of claims 1 to 4.