Information processing device

The information processing apparatus addresses the challenge of initiating conversations in virtual spaces by emphasizing audio and visual cues based on user themes and spatial relationships, facilitating more engaging interactions.

JP2026091363AActive Publication Date: 2026-06-04KDDI CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KDDI CORP
Filing Date
2024-11-21
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

In virtual spaces with multiple potential conversation partners, users find it difficult to initiate suitable conversations due to the overwhelming number of options, making it challenging to engage in meaningful interactions.

Method used

An information processing apparatus that identifies and emphasizes audio and visual cues based on user themes, preferences, and spatial relationships within the virtual space to facilitate targeted conversations.

Benefits of technology

Enhances user engagement by highlighting relevant voices and avatars, thereby simplifying and enriching interactions in virtual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026091363000001_ABST
    Figure 2026091363000001_ABST
Patent Text Reader

Abstract

To make it easier for users to converse in a virtual space. [Solution] The information processing device 1 includes an acquisition unit 131 that acquires audio information indicating the voice of a user using the virtual space, an identification unit 132 that identifies an audio to be emphasized from the audio indicated by the audio information to be output to the user terminal 2 used by the other user, based on user information indicating at least one of the following: a theme set by the other user using the virtual space, the other user's hobbies or preferences, and the relationship between the position and orientation of each other's avatars in the virtual space. An output control unit 133 that outputs the audio to the user terminal 2 in an emphasized state, with the audio identified by the identification unit 132 emphasized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus.

Background Art

[0002] Conventionally, a system is known that arranges avatars of a plurality of users in a virtual space and enables conversations between the plurality of users via the avatars (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When there are a large number of other users with whom a user can have a conversation in a virtual space, it is difficult for the user to find other users with whom they can have a suitable conversation, and there is a problem that it is difficult for the user to start a conversation.

[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to make it easier for a user to have a conversation in a virtual space.

Means for Solving the Problems

[0006] A first aspect of the present invention is an information processing apparatus. This information processing apparatus includes: an acquisition unit that acquires audio information indicating the voice of a user using a virtual space; an identification unit that identifies an audio to be emphasized from the audio indicated by the audio information to be output to a terminal used by the other user, based on user information indicating at least one of the following: a theme set by another user using the virtual space, the other user's hobbies or preferences, and the relationship between the position and orientation of each other's avatars in the virtual space, corresponding to the other user and the user; and an output control unit that outputs the audio identified by the identification unit to the terminal used by the other user in an emphasized state.

[0007] The acquisition unit may acquire the audio information indicating the voices of multiple users participating in an event held in the virtual space, and the identification unit may identify the voice indicated by the audio information based on the user information of the other users using the event.

[0008] The identifying unit may identify a voice to be emphasized from among the voices of the user corresponding to an avatar located within a predetermined range from the position of the avatar corresponding to the other user in the virtual space.

[0009] The output control unit may output to the terminal, at a relatively high volume, the voice identified by the identification unit from among the multiple user voices acquired by the acquisition unit.

[0010] The output control unit may cause the terminal used by the other user to output an image that highlights the avatar corresponding to the user who emitted the voice identified by the identification unit.

[0011] The output control unit may output to a terminal used by another user an image indicating that the user is emitting the sound identified by the identification unit, associated with an avatar corresponding to the user.

[0012] The output control unit may output the video, which has an image showing text corresponding to the voice spoken by the user, to a terminal used by another user.

[0013] The user information may also be information indicating a theme set by the other user, and the identification unit may identify the audio corresponding to the theme from among the audio indicated by the audio information.

[0014] The identifying unit may identify a theme corresponding to each of the multiple voices indicated by the audio information, and may identify the voice among the multiple voices for which a theme corresponding to a theme set by the other user has been identified as the voice to be emphasized.

[0015] The acquisition unit may acquire user voices corresponding to each of the multiple users participating in the event held in the virtual space, and information indicating the theme set by that user, and the identification unit may identify, among the multiple voices indicated by the voice information, the voice corresponding to the user who set the theme corresponding to the theme set by the other user as the voice to be emphasized.

[0016] The user information may be information indicating the hobbies or preferences of the other user, and the identification unit may identify the voices indicated by the voice information that correspond to the hobbies or preferences of the other user as the voices to be emphasized.

[0017] The user information may be information indicating the hobbies or preferences of the other user, and the identification unit may identify the hobbies or preferences of the other user based on conversation history information indicating the conversation history of the other user or behavior history information indicating the behavior history of the other user, and may identify the voices from among the voice information that correspond to the identified hobbies or preferences of the other user as the voice to be emphasized.

[0018] The user information may be the relative relationship of the positions and orientations of the avatars corresponding to the other user and the user in the virtual space, and the specifying unit may specify the user corresponding to the avatar that is a candidate for speaking to the avatar corresponding to the other user based on the relative relationship, and may specify the voice corresponding to the specified user among the voices indicated by the voice information.

[0019] The specifying unit may specify, as the user corresponding to the avatar that is a candidate for speaking to the avatar corresponding to the other user, the user corresponding to the avatar approaching the position of the avatar corresponding to the other user or the user corresponding to the avatar whose line of sight is directed at the avatar corresponding to the other user.

[0020] A second aspect of the present invention is an information processing apparatus. This information processing apparatus includes an acquisition unit that acquires voice information indicating the voice of a user using a virtual space, and a specifying unit that specifies a voice to be emphasized among the voices indicated by the voice information to be output to a terminal used by another user using the virtual space, based on at least any one of user information indicating a theme set by the user using the virtual space, the hobbies or preferences of the user, and the relative relationship of the positions and orientations of the avatars corresponding to the user and another user using the virtual space, and an output control unit that causes the voice to be output to the terminal used by the other user in a state where the voice specified by the specifying unit is emphasized.

Advantages of the Invention

[0021] According to the present invention, there is an effect that it is possible to facilitate conversation among users in a virtual space.

Brief Description of the Drawings

[0022] [Figure 1] It is a diagram for explaining the outline of the information processing apparatus. [Figure 2] It is a diagram showing the functional configuration of the information processing apparatus. [Figure 3]This is a diagram showing an example in which an avatar is emphasized in the video displayed on the user terminal. [Figure 4] This is a diagram showing an example in which an image showing text is added to the video displayed on the user terminal. [Figure 5] This is a flowchart showing the process of causing the information processing apparatus to output audio information and video information to the user terminal 2. [Embodiments for Carrying Out the Invention]

[0023] [Overview of Information Processing Apparatus 1] FIG. 1 is a diagram for explaining the overview of the information processing apparatus 1. The information processing apparatus 1 is a computer that acquires audio information indicating the voices of users participating in an event held in a virtual space, and outputs the audio information to terminals used by other users participating in the event.

[0024] The information processing apparatus 1 is communicably connected to a plurality of user terminals 2 used by a plurality of users participating in the event via a communication network such as the Internet or a mobile phone line. The user terminal 2 is, for example, a VR (Virtual Reality) headset worn by the user. Note that the user terminal 2 is not limited to a VR headset, and may be composed of one or more devices used by the user, such as a smartphone, a display, and a speaker.

[0025] In the example shown in FIG. 1, there are four user terminals 2A, 2B, 2C, 2D used by users UA, UB, UC, and UD as the user terminal 2. Also, in the event venue provided in the virtual space, there are avatars AA, AB, AC, and AD corresponding to these users. In the following description, when the user terminals 2A, 2B, 2C, and 2D are not distinguished, these user terminals are referred to as the user terminal 2. Also, the event is, for example, a symposium or a party held on the virtual space, and the user can talk about a theme set by himself / herself or a conversation related to his / her hobby or preference.

[0026] The information processing device 1 acquires audio information from multiple user terminals 2 indicating the voices emitted by users during an event (Figure 1 (1)). Based on user information indicating at least one of the following: a theme set by other users using the virtual space, the hobbies or preferences of other users, and the relationship between the positions and orientations of each user's avatar in the virtual space: the information processing device 1 identifies the voices to be emphasized among the voices to be output to the user terminal 2 used by those other users (Figure 1 (2)).

[0027] In this embodiment, the other user is the user of user terminal 2, which is the destination of the audio output, and the user is the user of user terminal 2, which is the source of the audio acquisition. For example, in the example shown in Figure 1, if the other user is user UD using user terminal 2D, then the user is three users UA, UB, and UC using user terminals 2A, 2B, and 2C. In this case, the information processing device 1 identifies the audio to be emphasized from among the audio corresponding to user UA, UB, and UC to be output to user terminal 2D used by user UD, based on the user information of user UD as the other user.

[0028] The information processing device 1 transmits audio information emphasizing the identified voice to the user terminal 2 used by other users, thereby causing the user terminal 2 to output the audio with the identified voice emphasized (Figure 1, (3)). For example, in the example shown in Figure 1, if the information processing device 1 identifies an audio that emphasizes the voice of user UA among the voices corresponding to user UA, UB, and UC to be output to user terminal 2D, it transmits audio information emphasizing the voice of user UA among the voices corresponding to user UA, UB, and UC to user terminal 2D, causing the voices of these users to be output to user terminal 2D with user UA's voice emphasized. Similarly, the information processing device 1 identifies the audio to be emphasized for user terminals 2A, 2B, and 2C, and causes them to output the audio with the identified audio emphasized. In this way, the information processing device 1 makes it easier for users to converse in the virtual space.

[0029] [Functional configuration of the information processing device 1] Next, the functional configuration of the information processing device 1 will be described. Figure 2 is a diagram showing the functional configuration of the information processing device 1. The information processing device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.

[0030] The communication unit 11 is a communication interface for sending and receiving data with the user terminal 2. The memory unit 12 is a storage medium for storing various types of data, and includes ROM (Read Only Memory), RAM (Random Access Memory), and hard disks. The memory unit 12 stores programs to be executed by the control unit 13. The memory unit 12 also stores programs that cause the control unit 13 to function as an acquisition unit 131, a specification unit 132, and an output control unit 133.

[0031] The control unit 13 is, for example, a CPU (Central Processing Unit). The control unit 13 functions as an acquisition unit 131, a identification unit 132, and an output control unit 133 by executing a program stored in the storage unit 12.

[0032] The acquisition unit 131 acquires audio information indicating the voice of a user using the virtual space. For example, the acquisition unit 131 acquires audio information corresponding to each of the multiple users participating in an event held in the virtual space from the user terminal 2 used by each of the multiple users. For example, the acquisition unit 131 associates the audio information indicating the voice of the user picked up by a microphone installed on the user terminal 2 with the user ID (Identification) as user identification information for identifying the user, and acquires it from the user terminal 2. The user ID is, for example, the user ID issued to the user when the user participated in the event.

[0033] Furthermore, the acquisition unit 131 acquires user information, including themes set by each of the multiple users, and at least one of the user's hobbies or preferences, and associates this information with the user's voice. For example, when a user participates in an event, the acquisition unit 131 receives at least one of the theme and the user's hobbies or preferences from the user terminal 2. The theme is, for example, a topic the user wants to talk about at the event or a topic they are interested in. The acquisition unit 131 acquires user information, including the theme received from the user terminal 2, user information indicating at least one of the user's hobbies or preferences, and the user ID, and associates this information with the user's voice.

[0034] Furthermore, the information processing device 1 manages the virtual space, and the acquisition unit 131 acquires avatar location information, which indicates the location of each avatar in the event venue set up in the virtual space, and avatar direction information, which indicates the direction the avatar is facing, as user information, associated with the user ID.

[0035] Furthermore, in the virtual space, virtual cameras are provided at the location of each user's avatar, facing the same direction as the avatar, and are associated with the user ID of the user corresponding to the avatar. The acquisition unit 131 acquires video information showing the event venue captured by the virtual camera as user information indicating the relationship between the positions and orientations of the users. The acquisition unit 131 acquires the video information and the user ID associated with the virtual camera. Since the user ID is associated with the audio information, the acquisition unit 131 can acquire the video information in association with the user's audio by acquiring the video information and the user ID.

[0036] The identification unit 132 identifies a voice to be emphasized from among multiple voices indicated by the voice information to be output to the user terminal 2 used by another user, based on the user information of other users using the virtual space. The user information of another user is information indicating at least one of the following: the theme set by the other user, the other user's hobbies or preferences, and the relationship between the position and orientation of each other's avatars in the virtual space.

[0037] For example, the user information of other users is at least one of the following, acquired by the acquisition unit 131: the theme set by the other user, information indicating the other user's hobbies or preferences, avatar position information, avatar direction information, and video information.

[0038] The identification unit 132 identifies a voice to be emphasized from among the voices of users corresponding to avatars located within a predetermined range from the virtual space location of the avatar corresponding to the other user, based on the user information of other users.

[0039] Specifically, first, the identification unit 132 identifies the audio to be output to the user terminal 2 used by other users, based on the avatar location information of multiple users acquired by the acquisition unit 131. The identification unit 132 identifies the user ID associated with the avatar location information that indicates a location within a predetermined range from the location indicated by the avatar location information associated with the user ID of the other user. Then, the identification unit 132 identifies the audio indicated by the audio information acquired in association with the identified user ID as the output target audio, which is the audio to be output to the user terminal 2 used by other users.

[0040] The identification unit 132 then identifies the audio to be emphasized from among the audio to be output, based on the user information of the user corresponding to the identified audio to be output and the user information of other users. For example, the identification unit 132 identifies the audio to be emphasized from among the audio to be output, based on themes, hobbies, or preferences set by other users, which are obtained as user information by the acquisition unit 131.

[0041] For example, the identification unit 132 identifies at least one of the themes, hobbies, or preferences of the other user, as indicated by user information obtained from the user terminal 2 of another user. The identification unit 132 also identifies the user who set themes, hobbies, or preferences that correspond to at least one of the themes, hobbies, or preferences set by the other user, based on the themes, hobbies, or preferences indicated by user information associated with the user ID corresponding to the output audio. For example, the identification unit 132 identifies the user who has set the same theme as the theme set by the other user, the user who has set a theme related to the theme set by the other user, or the user who has the same hobbies or preferences as the other user. The identification unit 132 then identifies the audio corresponding to the identified user as the audio to be emphasized.

[0042] Furthermore, the identification unit 132 identifies the audio to be emphasized from among the audio to be output, based on the virtual space position and orientation of the avatar corresponding to the other user, as indicated by the avatar position information, avatar orientation information, and video information, which are user information obtained from the user terminal 2 of the other user, and the virtual space position and orientation of the avatar corresponding to the audio to be output.

[0043] For example, the identification unit 132 identifies the voice of a user corresponding to an avatar that is within a predetermined range from the virtual space position of the avatar corresponding to the other user and is located at a position corresponding to the orientation of the avatar, based on the avatar position information and orientation information of the avatar corresponding to the other user and the avatar position information and orientation information of a user different from the other user, as the voice to be emphasized. Here, the identification unit 132 may also identify the voice of a user as the voice to be emphasized when the other user's avatar and the user's avatar are facing each other, based on the orientation of the other user's avatar and the orientation of the user's avatar. In this way, the information processing device 1 can identify the voice of a user that another user is about to speak to as the voice to be emphasized.

[0044] Furthermore, the identification unit 132 may identify a user corresponding to an avatar that is a candidate to speak to the avatar corresponding to another user, based on the relative positions and orientations in the virtual space of the avatars of that other user and a different user, as indicated by the user information. The identification unit 132 may then identify the audio corresponding to the identified user among the output audio as the audio to be emphasized.

[0045] For example, the identification unit 132 identifies a user whose avatar is within a predetermined range from the virtual space location of the avatar corresponding to another user, and whose avatar is facing the direction of the other user's avatar, as a user whose avatar is a candidate to speak to the avatar corresponding to the other user.

[0046] Furthermore, the identification unit 132 may identify avatars that are approaching the position of the avatar corresponding to another user, or avatars that are looking at the avatar corresponding to another user, based on the video information acquired by the acquisition unit 131 and at least one of the avatar position information and direction information, and then identify the user corresponding to the identified avatar. For example, the identification unit 132 identifies avatars that are approaching the avatar corresponding to another user based on the video at multiple times corresponding to the position of the other user's avatar, as indicated by the video information acquired by the acquisition unit 131. The identification unit 132 also analyzes the video at each of these multiple times and identifies avatars that are facing the direction of the other avatar for a predetermined percentage or more as avatars that are looking at the avatar corresponding to another user. The identification unit 132 identifies the user corresponding to the identified avatar as the user corresponding to an avatar that is a candidate to speak to the avatar corresponding to the other user.

[0047] The identification unit 132 then identifies the voice of the user corresponding to the identified avatar as the voice to be emphasized. In this way, the information processing device 1 can identify the voice of a user who is about to speak to another user as the voice to be emphasized.

[0048] The output control unit 133 controls the output of the audio indicated by the audio information and the video indicated by the video information to the user terminal 2 used by other users, with the audio identified by the identification unit 132 emphasized. The output control unit 133 controls the output of the audio identified by the identification unit 132 from among the audio to be output that is indicated by the audio information of multiple users acquired by the acquisition unit 131, to the user terminal 2 at a relatively high volume.

[0049] For example, the output control unit 133 adjusts the volume of the audio identified by the identification unit 132 to a relatively higher volume by increasing the volume of the audio identified by the identification unit 132 among the output target audio included in the audio of multiple users corresponding to the multiple audio information acquired by the acquisition unit 131, and decreasing the volume of the audio that was not identified. Then, the output control unit 133 transmits audio information indicating the adjusted output target audio to the user terminal 2 of another user, and causes the user terminal 2 of the other user to output the adjusted output target audio indicated by the audio information.

[0050] The output control unit 133 emphasizes the audio identified by the identification unit 132 by adjusting the audio to be output, but is not limited to this. The output control unit 133 may also process video information acquired by the acquisition unit 131 from a virtual camera corresponding to the position of another user's avatar, thereby emphasizing the audio identified by the identification unit 132.

[0051] For example, the output control unit 133 controls the output of video to the user terminal 2 used by other users, highlighting the avatar corresponding to the user who emitted the voice identified by the identification unit 132. Based on the position and orientation of the virtual camera corresponding to the position of the other users' avatars, and avatar position information indicating the position of the avatars corresponding to each of the multiple users, the output control unit 133 identifies the users corresponding to the multiple avatars shown in the video information acquired from the virtual camera. Then, the output control unit 133 processes the video information to highlight the avatar corresponding to the identified user among the multiple avatars shown in the video information.

[0052] Figure 3 shows an example where an avatar is highlighted in a video displayed on user terminal 2. In the example shown in Figure 3, among the multiple avatars shown in the video, a frame is added around the avatar corresponding to the user identified as having a matching theme with another user, and it can be seen that the theme of that user is displayed in conjunction with that avatar.

[0053] Furthermore, the output control unit 133 may control the output to a user terminal 2 used by other users, associated with an avatar corresponding to the user, to show that the user is emitting the sound identified by the identification unit 132. In this case, the identification unit 132 generates text information indicating the output target sound by analyzing the output target sound. The output control unit 133 may then control the output to a user terminal 2 used by other users, to show a video with an image attached that shows the text indicated by the text information identified by the identification unit 132, corresponding to the sound emitted by the user.

[0054] Figure 4 shows an example where images displaying text are added to a video displayed on user terminal 2. In the example shown in Figure 4, it can be seen that each of the multiple avatars shown in the video has a speech bubble displayed containing text indicating the voice spoken by the user corresponding to that avatar. In addition, in the example shown in Figure 4, it can be seen that the speech bubbles corresponding to avatars of users identified as having the same theme as other users are highlighted compared to the other speech bubbles. In this way, other users who view the video can identify avatars of users who have themes and other topics that match their own and with whom they can easily converse, and engage in conversations on common topics.

[0055] Furthermore, if the output control unit 133 wants to emphasize the voice of a user whose avatar is within a predetermined range from the position of another user's avatar but is not visible in the video information acquired from the virtual camera, it may display a message such as "There is a user with the same theme behind you" to emphasize the voice of the identified user.

[0056] [Operation Flow] Next, we will explain the processing flow in the information processing device 1. Figure 5 is a flowchart showing the processing flow in which the information processing device 1 outputs audio and video information to the user terminal 2. The flowchart shown in Figure 5 is assumed to be executed repeatedly while the user is viewing the event.

[0057] First, the acquisition unit 131 acquires at least one of the following from each of the user terminals 2 of multiple users: theme information indicating a theme set by the user, and hobby information indicating a hobby or preference (S1).

[0058] Next, the acquisition unit 131 acquires audio information, video information, avatar position information, and avatar orientation information corresponding to each of the multiple users who participated in the event held in the predetermined space (S2).

[0059] Next, the identification unit 132 identifies the output target audio from among the audio indicated by the acquired audio information, based on the avatar location information of each of the multiple users, which is the audio to be output to the user terminal 2 of another user (S3).

[0060] Next, the identification unit 132 identifies the audio to be emphasized from the audio to be output, based on the theme information and hobby information acquired in S1, and the audio information, video information, avatar position information, and avatar orientation information acquired in S2 (S4).

[0061] Next, the identification unit 132 processes the audio information indicating the target audio to be output identified in S3 and the video information acquired in S2 to generate audio information and video information that emphasizes the audio identified in S4 (S5).

[0062] The output control unit 133 transmits the audio and video information generated in S5 to the user terminal 2 of another user, thereby causing the audio indicated by the audio information and the video indicated by the video information to be output to the user terminal 2 (S6).

[0063] Furthermore, the identification unit 132 and the output control unit 133 execute processes S3 to S6, treating each of the multiple users participating in the event as another user, i.e., a user to whom audio information will be output. As a result, audio and video are output to each of the multiple users' user terminals 2 with the audio identified by the identification unit 132 emphasized, corresponding to the user of that user terminal 2.

[0064] [Differentiation] In the above-described embodiment, the identification unit 132 identified the voice to be emphasized based on the theme set by the user, as indicated by the user information acquired by the acquisition unit 131, but is not limited to this. The identification unit 132 may also identify the theme corresponding to each of the multiple voices, based on each of the multiple voices indicated by the multiple voice information acquired. For example, the identification unit 132 may identify the theme corresponding to each of the multiple voices by analyzing each of the multiple voices indicated by the multiple voice information acquired. Furthermore, the identification unit 132 may identify the voice of a user speaking about a theme corresponding to a theme set by another user as the voice to be emphasized, from among the voices output to the user terminal 2 used by other users.

[0065] Furthermore, the identification unit 132 identifies, but is not limited to, the voice to be emphasized based on the user's hobbies or preferences indicated by the user information acquired by the acquisition unit 131. The identification unit 132 may also generate conversation history information showing the conversation history of each of the multiple users based on the acquired voice information, or acquire behavior history information showing the actions of each of the multiple users from each of the user terminals 2. The behavior history may be, for example, the location history of the user terminal 2, or the access history of websites on the internet that the user terminal 2 has accessed.

[0066] In this case, the identification unit 132 may identify the hobbies or preferences of each of the multiple users based on the conversation history information or behavioral history information of each of the multiple users. The identification unit 132 may then identify the voices indicated by the voice information output by the multiple users that correspond to the hobbies or preferences of other users as voices to be emphasized.

[0067] [Effects of Information Processing Device 1] As described above, the information processing device 1 according to this embodiment acquires audio information indicating the voice of a user using the virtual space, and based on user information indicating at least one of the following: a theme set by another user using the virtual space, the other user's hobbies or preferences, and the relationship between the position and orientation of each other's avatars in the virtual space: the information processing device 1 identifies the audio to be emphasized from the audio information to be output to the user terminal 2 used by the other user, and outputs the audio indicated by the acquired audio information to the user terminal 2 used by the other user with the identified audio emphasized. In this way, the information processing device 1 makes it easier for users to converse in the virtual space.

[0068] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."

[0069] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of its gist. For example, all or part of the apparatus can be configured by functionally or physically distributing and integrating in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combinations are combined with the effects of the original embodiments. [Explanation of symbols]

[0070] 1. Information Processing Device 11 Communications Department 12 Storage section 13 Control Unit 131 Acquisition Department 132 Specific part 133 Output Control Unit

Claims

1. An acquisition unit that acquires audio information indicating the voice of a user using the virtual space, A selection unit identifies the audio to be emphasized from the audio information output to the terminal used by the other user, based on user information indicating at least one of the following: a theme set by another user using the virtual space, the other user's hobbies or preferences, and the relationship between the position and orientation of each other's avatars in the virtual space. An output control unit that causes the audio identified by the identification unit to be output to a terminal used by the other user in an emphasized state, An information processing device having

2. The acquisition unit acquires the audio information indicating the voices of multiple users participating in an event held in the virtual space, The identifying unit identifies the voice indicated by the voice information based on the user information of the other user who is using the event. The information processing apparatus according to claim 1.

3. The identifying unit identifies a voice to be emphasized from among the voices of the user corresponding to an avatar located within a predetermined range from the position of the avatar corresponding to the other user in the virtual space. The information processing apparatus according to claim 1.

4. The output control unit causes the terminal to output, at a relatively high volume, the voice identified by the identification unit from among the multiple voices of the user acquired by the acquisition unit. The information processing apparatus according to claim 1.

5. The output control unit causes the terminal used by the other user to output an image that highlights the avatar corresponding to the user who emitted the voice identified by the identification unit. The information processing apparatus according to claim 1.

6. The output control unit outputs to a terminal used by another user an image indicating that the user is emitting the sound identified by the identification unit, associated with an avatar corresponding to the user. The information processing apparatus according to claim 5.

7. The output control unit causes the video, which has an image showing text corresponding to the voice spoken by the user, to be output to a terminal used by the other user. The information processing apparatus according to claim 6.

8. The user information mentioned above is information indicating the theme set by the other user. The identifying unit identifies the audio corresponding to the theme from among the audio indicated by the audio information. The information processing apparatus according to claim 1.

9. The identification unit identifies a theme corresponding to each of the multiple voices indicated by the audio information, and identifies the voice among the multiple voices whose theme corresponds to the theme set by the other user as the voice to be emphasized. The information processing apparatus according to claim 8.

10. The acquisition unit acquires the voice of each user corresponding to a user participating in an event held in the virtual space, and associates it with information indicating the theme set by that user. The identification unit identifies, among the multiple voices indicated by the voice information, the voice corresponding to the user who set the theme corresponding to the theme set by the other user, as the voice to be emphasized. The information processing apparatus according to claim 8.

11. The user information is information that indicates the hobbies or preferences of the other user. The identifying unit identifies, among the voices indicated by the voice information, the voices corresponding to the other user's hobbies or preferences as voices to be emphasized. The information processing apparatus according to claim 1.

12. The user information is information that indicates the hobbies or preferences of the other user. The identification unit identifies the other user's hobbies or preferences based on conversation history information showing the other user's conversation history or behavior history information showing the other user's behavior history, and identifies the voices from the voice information that correspond to the identified other user's hobbies or preferences as the voices to be emphasized. The information processing apparatus according to claim 11.

13. The user information is the relative position and orientation of the avatars of the other user and the user in the virtual space. The identifying unit identifies the user corresponding to an avatar that is a candidate to speak to the avatar corresponding to the other user based on the relative relationship, and identifies the voice corresponding to the identified user from among the voices indicated by the voice information. The information processing apparatus according to claim 1.

14. The identifying unit identifies a user corresponding to an avatar that is approaching the position of the avatar corresponding to the other user, or a user corresponding to an avatar that is looking at the avatar corresponding to the other user, as the user corresponding to an avatar that is a candidate to speak to the avatar corresponding to the other user. The information processing apparatus according to claim 13.

15. An acquisition unit that acquires audio information indicating the voice of a user using the virtual space, A selection unit identifies the audio to be emphasized from the audio information output to a terminal used by other users of the virtual space, based on user information indicating at least one of the following: a theme set by the user using the virtual space, the user's hobbies or preferences, and the relative position and orientation of each other's avatars in the virtual space corresponding to the user and other users using the virtual space. An output control unit that causes the audio identified by the identification unit to be output to a terminal used by the other user in an emphasized state, An information processing device having