Information processing device, information processing method, and computer program

The information processing device addresses cumbersome voice transmission in virtual spaces by detecting user voices and altering avatar display modes to indicate transmission ranges, enhancing clarity in virtual communication.

JP2025126304APending Publication Date: 2025-08-28NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025108706
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing technologies for virtual communication spaces require users to manually specify call targets, leading to cumbersome operations and unclear voice transmission directions.

Method used

An information processing device that detects user voices in a virtual space, outputs voices to avatars satisfying predetermined conditions, and changes the display mode of listening avatars to indicate voice transmission ranges.

Benefits of technology

Enables users to recognize the range of sound transmission, reducing the need for manual target specification and clarifying voice recipients in virtual communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126304000001_ABST
    Figure 2025126304000001_ABST
Patent Text Reader

Abstract

To provide an information processing device and the like whereby a user can be made aware of the audio transmission range in a situation where a virtual space is used, and users communicate with each other.SOLUTION: An information processing device includes: detection means that detects audio generated by a user operating an avatar in a virtual space; audio control means that outputs the audio to a user of an avatar that fulfills prescribed conditions in a relation with a speaking avatar that is the avatar operated by the user that provided the audio; and display control means that changes a display mode for a listener avatar that is the avatar fulfilling the prescribed conditions.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technology for controlling a virtual space. [Background technology]

[0002] There are technologies that allow multiple users to communicate using virtual space. For example, Patent Document 1 discloses a technology in which objects with captured images of each user are placed in a three-dimensional space, and calls are made through the three-dimensional space.

[0003] Furthermore, in relation to a technology for enabling multiple users to communicate, Patent Document 2 discloses generating an image that corresponds to real space, in which a person object is placed at a position on the image that corresponds to the position of the person in real space. In the technology of Patent Document 2, when there is a person talking, a link is generated that connects the objects corresponding to the person talking. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2018 / 020766 [Patent Document 2] JP 2018-36871 A Summary of the Invention [Problem to be solved by the invention]

[0005] In both of the techniques disclosed in Patent Documents 1 and 2, it is possible to make a call with a pre-designated user.

[0006] Consider a case where a user uses a virtual space where multiple avatars exist. For example, the user operates an avatar representing the user to move within the virtual space or to make calls with other users operating other avatars. In such a case, if the user needs to specify the target of the call for every call, the user's operation becomes cumbersome. On the other hand, even if the user's voice can be transmitted without specifying the target of the call, the user may not know to which user the voice is being transmitted.

[0007] The present disclosure has been made in consideration of the above-mentioned problems, and one of its objectives is to provide an information processing device, etc. that can allow users to recognize the range of sound transmission when users communicate with each other using a virtual space. [Means for solving the problem]

[0008] An information processing device according to one aspect of the present disclosure includes a detection means for detecting a voice uttered by a user operating an avatar in a virtual space, a voice control means for outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice, and a display control means for changing the display mode of a listening avatar, which is an avatar that satisfies the predetermined condition.

[0009] An information processing method according to one aspect of the present disclosure detects a voice uttered by a user operating an avatar in a virtual space, outputs the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice, and changes the display mode of a listening avatar, which is an avatar that satisfies the predetermined condition.

[0010] A computer-readable storage medium according to one aspect of the present disclosure stores a program that causes a computer to execute the following processes: detecting a voice uttered by a user operating an avatar in a virtual space; outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice; and changing the display mode of a listening avatar, which is an avatar that satisfies the predetermined condition. [Effects of the Invention]

[0011] According to the present disclosure, in a situation where users communicate with each other using a virtual space, it is possible to allow users to recognize the range of sound transmission. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a diagram schematically illustrating an example of a configuration including an information processing apparatus according to a first embodiment of the present disclosure. [Figure 2] FIG. 2 is a diagram schematically illustrating an example of a virtual space displayed on a user terminal according to the first embodiment of the present disclosure. [Figure 3] 1 is a block diagram illustrating an example of a functional configuration of an information processing apparatus according to a first embodiment of the present disclosure. [Figure 4] 5 is a flowchart illustrating an example of an operation of the information processing device according to the first embodiment of the present disclosure. [Figure 5] FIG. 10 is a block diagram illustrating an example of a functional configuration of an information processing device according to a second embodiment of the present disclosure. [Figure 6A] FIG. 10 is a diagram illustrating an example of a listening area according to a second embodiment of the present disclosure. [Figure 6B] FIG. 10 is a diagram illustrating a second example of a listening area according to the second embodiment of the present disclosure. [Figure 6C] FIG. 10 is a diagram illustrating a third example of a listening area according to the second embodiment of the present disclosure. [Figure 7] FIG. 10 is a diagram illustrating an example of a display mode of a listening avatar according to the second embodiment of the present disclosure. [Figure 8]10 is a flowchart illustrating an example of an operation of the information processing device according to the second embodiment of the present disclosure. [Figure 9] FIG. 10 is a block diagram illustrating an example of a functional configuration of an information processing device according to a third embodiment of the present disclosure. [Figure 10A] FIG. 11 is a diagram illustrating an example of a volume control area according to the third embodiment of the present disclosure. [Figure 10B] FIG. 10 is a diagram showing another example of a volume control area according to the third embodiment of the present disclosure. [Figure 11] FIG. 11 is a diagram illustrating an example of a display mode of a listening avatar according to the third embodiment of the present disclosure. [Figure 12] 10 is a flowchart illustrating an example of an operation of the information processing device according to the third embodiment of the present disclosure. [Figure 13] FIG. 10 is a block diagram illustrating an example of a functional configuration of an information processing device according to a modified example of the present disclosure. [Figure 14] FIG. 2 is a block diagram showing an example of a hardware configuration of a computer device that realizes an information processing device according to the first, second, and third embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0014] First Embodiment An overview of the information processing device of the present disclosure will be described.

[0015] FIG. 1 is a diagram schematically illustrating an example of a configuration including an information processing device 100. As shown in FIG. 1, the information processing device 100 is communicably connected to user terminals 200-1, 200-2, . . . , 200-n (n is a natural number equal to or greater than 1) via a wireless or wired network. Here, when there is no need to distinguish between the user terminals 200-1, 200-2, . . . , 200-n, they are simply referred to as user terminals 200. The user terminal 200 is a device operated by a user. The user terminal 200 is, for example, a personal computer, but is not limited to this example. The user terminal 200 may be a smartphone or a tablet terminal, or may be a device including a goggle-type wearable terminal (also referred to as a head-mounted display) having a display. The user terminal 200 also includes input devices such as a keyboard, a mouse, a microphone, and a wearable device that is operated based on the user's actions, and output devices such as a display and a speaker. Furthermore, the user terminal 200 includes at least one of a photographing device and a device capable of reading a voiceprint, a fingerprint, a palmprint, an iris, a vein, and the like.

[0016] First, a virtual space in the present disclosure will be described. The virtual space is a virtual space shared by multiple users, and is a space in which user operations are reflected. The virtual space is also called a VR (Virtual Reality) space. For example, the virtual space is provided by an information processing device 100. A user terminal 200 displays an image representing the virtual space. FIG. 2 is a diagram schematically illustrating an example of a virtual space displayed on the user terminal 200. In the example of FIG. 2, the virtual space is displayed on the display of the user terminal 200. As shown in FIG. 2, the virtual space includes an avatar. The avatar is an object operated by a user. The user uses the virtual space by operating the avatar. For example, the user terminal 200 may display an image of the virtual space from the perspective of an avatar operated by the user. In this case, the image displayed on the user terminal 200 may be updated according to the movement of the avatar. Furthermore, for example, a user may be able to communicate with other users by performing actions on avatars operated by other users. Note that the device providing the virtual space does not have to be the information processing device 100. For example, the virtual space may be provided by an external device (not shown).

[0017] 3 is a block diagram showing an example of the functional configuration of the information processing device 100 according to the first embodiment. As shown in FIG. 3, the information processing device 100 includes a detection unit 110, an audio control unit 120, and a display control unit .

[0018] The detection unit 110 detects a voice uttered by a user operating an avatar in a virtual space. The detection unit 110 is an example of a detection means.

[0019] The voice control unit 120 controls the voice. Here, the user who uttered the voice is also referred to as the speaking user. Furthermore, the avatar operated by the user who uttered the voice is also referred to as the speaking avatar. The voice control unit 120, for example, identifies an avatar that satisfies a predetermined condition in relation to the speaking avatar. The avatar that satisfies the predetermined condition may be, for example, an avatar that exists within a predetermined distance from the speaking avatar, or an avatar that exists in a predetermined area including the speaking avatar. Note that the predetermined condition is not limited to this example. Furthermore, the avatar that satisfies the predetermined condition is also referred to as a listening avatar. The voice control unit 120, for example, outputs the voice from the speaking user to the user of the identified avatar. In this way, the voice control unit 120 outputs the voice to the user of the avatar that satisfies the predetermined condition in relation to the speaking avatar, which is the avatar operated by the user who uttered the voice. The voice control unit 120 is an example of a voice control means.

[0020] The display control unit 130 controls the display of the virtual space. For example, when a hearing avatar that satisfies a predetermined condition exists, the display control unit 130 controls the display mode of the hearing avatar. For example, the display control unit 130 assigns a predetermined symbol or a predetermined color to the hearing avatar. The display control unit 130 changes the display mode of the hearing avatar that satisfies the predetermined condition. The display control unit 130 is an example of a display control means.

[0021] Next, an example of the operation of the information processing device 100 will be described with reference to Fig. 4. In this disclosure, each step in a flowchart will be represented by a number assigned to the step, such as "S1".

[0022] 4 is a flowchart illustrating an example of the operation of the information processing device 100. The detection unit 110 detects a voice uttered by a user operating an avatar in a virtual space (S1). The voice control unit 120 outputs the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice (S2). The display control unit 130 changes the display mode of a listening avatar, which is an avatar that satisfies the predetermined condition (S3).

[0023] In this way, the information processing device 100 of the first embodiment detects a voice uttered by a user operating an avatar in a virtual space, and outputs the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice. The information processing device 100 then changes the display mode of a listening avatar, which is an avatar that satisfies the predetermined condition. By controlling the display mode of the avatar of the user to whom the voice is to be output, the information processing device 100 can inform the user who uttered the voice which users the voice will be transmitted to. In other words, the information processing device 100 of the present disclosure allows users to recognize the transmission range of the voice in a situation where users communicate with each other using a virtual space.

[0024] <Second embodiment> Next, an information processing apparatus according to a second embodiment will be described. In the second embodiment, the information processing apparatus 100 described in the first embodiment will be described in more detail.

[0025] [Details of the information processing device 100] 5 is a block diagram showing an example of a functional configuration of the information processing device 100 according to the second embodiment. As shown in FIG. 5, the information processing device 100 includes a detection unit 110, an audio control unit 120, and a display control unit 130.

[0026] The detection unit 110 detects a voice uttered by a user. For example, when a user utters a voice, the voice is picked up by a microphone or the like provided in the user terminal 200. Voice data relating to the picked-up voice is transmitted to the information processing device 100. The detection unit 110 detects the voice uttered by the user, for example, by receiving the voice data.

[0027] The voice control unit 120 includes a listening area setting unit 121 and a voice output unit 122. The listening area setting unit 121 sets a listening area. The listening area is an area that includes a speaking avatar and indicates a range in which the voice of the speaking user is transmitted. The listening area may be, for example, a range indicating a certain distance from the speaking avatar. FIG. 6A is a diagram illustrating an example of a listening area. For example, as shown in FIG. 6A, the listening area setting unit 121 may set a listening area in a circular shape centered on the speaking avatar. In this case, other avatars present within the listening area are listening avatars. That is, the voice of the speaking user is output from the user terminal 200 used by a user who operates another avatar present within the listening area. Here, the size of the listening area may be set to a constant size for each avatar or may vary depending on the voice volume. Specifically, the listening area setting unit 121 acquires the volume of the speaking user's voice, i.e., volume information, from the voice data of the speaking user. If the volume is higher than a predetermined threshold, the listening area setting unit 121 sets the size of the listening area to be larger than the reference size. If the volume is lower than the predetermined threshold, the listening area setting unit 121 sets the size of the listening area to be smaller than the reference size. Note that this example is not limiting, and the sizes of the listening areas may be set in advance for multiple volume ranges, and the listening area setting unit 121 may determine the listening area depending on which range the acquired volume belongs to. In this way, the listening area setting unit 121 may set the listening area depending on the volume of the voice uttered by the user.

[0028] The listening area is not limited to the above example. For example, the listening area may vary depending on the facial direction of the speaking avatar. FIG. 6B is a diagram illustrating a second example of the listening area. As illustrated in FIG. 6B, the listening area may be set to be wider in the direction in which the speaking avatar's face is facing and narrower in the direction in which the speaking avatar's face is not facing. In the example of FIG. 6B, the speaking avatar at point X is facing in the direction of point Q. Point P indicates a position behind the speaking avatar. Points P, X, and Q are points on a straight line, and points P and Q are points on the circumference of the listening area. In this case, the distance from point X to point Q is longer than the distance from point X to point P. In this way, the listening area setting unit 121 may set a wider listening area in the direction in which the speaking avatar's face is facing. That is, the listening area setting unit 121 may determine the listening area depending on the facial direction of the speaking avatar.

[0029] Furthermore, the listening area may be set in a different shape. FIG. 6C is a diagram showing a third example of a listening area. In the example of FIG. 6C, the listening area is set in a fan shape. In this example, the speaking avatar faces in the direction of point R. Point R is a point on the arc of the fan shape. That is, in the example of FIG. 6C, the listening area is set wide in the direction in which the speaking avatar's face is facing. In this way, the shape of the listening area may be a circle or a fan shape, or may be another shape such as an ellipse or a polygon.

[0030] The audio output unit 122 outputs audio. Specifically, if an avatar other than the speaking avatar is present in the listening area, the audio output unit 122 identifies the avatar as a listening avatar. Then, the audio output unit 122 outputs audio to the user terminal 200 used by the user operating the listening avatar. Here, in the present disclosure, "outputting audio from the user terminal 200 used by the user operating the avatar" may also be expressed as "outputting audio to the user of the avatar." The audio output unit 122 may identify the other avatar as a listening avatar by detecting that the other avatar has entered the listening area. Furthermore, the audio output unit 122 may acquire position information of the other avatars around the speaking avatar, and identify the avatar with position information indicating a position within the listening area from the acquired position information as the listening avatar.

[0031] The display control unit 130 controls the display of the virtual space. Specifically, the display control unit 130 changes the display mode of the listening avatar. FIG. 7 is a diagram showing an example of the display mode of the listening avatar. In the example of FIG. 7, an exclamation mark is added to the listening avatar. The display control unit 130 controls, for example, the user terminal 200 of the speaking user to display the listening avatar with the exclamation mark added. This allows the speaking user to visually recognize which avatar is the listening avatar, i.e., to which user the voice is being transmitted. Note that the display mode is not limited to this example. For example, the display control unit 130 may add other symbols or characters to the listening avatar, or may change the color of part or all of the listening avatar.

[0032] The display control unit 130 may also display a listening area. For example, the display control unit 130 displays a listening area set according to the voice of the speaking user on the user terminal 200 of the speaking user. This allows the speaking user to recognize the range to which the voice is being transmitted. The display control unit 130 may also display a listening area set according to the voice of the speaking user on the user terminal 200 of another user. This allows the other users to recognize who is making the voice.

[0033] [Example of operation of information processing device 100] Next, an example of the operation of the information processing device 100 will be described with reference to FIG. 8. FIG. 8 is a flowchart illustrating an example of the operation of the information processing device 100. When the detection unit 110 detects a voice ("Yes" in S101), the listening area setting unit 121 sets a listening area based on the volume of the voice and the facial direction of the speaking avatar (S102). At this time, the listening area setting unit 121 may set the listening area taking into consideration at least one of the volume of the voice and the facial direction of the speaking avatar. When the detection unit 110 does not detect a voice ("No" in S101), the information processing device 100 ends the processing.

[0034] The display control unit 130 displays the listening area (S103). If the audio output unit 122 identifies a listening avatar ("Yes" in S104), the audio output unit 122 outputs audio to the user of the listening avatar (S105). The display control unit 130 changes the display mode of the listening avatar (S106). If the audio output unit 122 does not identify a listening avatar ("No" in S104), the information processing device 100 ends the process.

[0035] In the above example of operation, the process of S103 does not have to be performed, and the order of the processes of S105 and S106 may be reversed.

[0036] In this way, the information processing device 100 of the second embodiment detects a voice uttered by a user operating an avatar in a virtual space and outputs the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice. The information processing device 100 then changes the display mode of a listening avatar, which is an avatar that satisfies the predetermined condition. By controlling the display mode of the avatar of the user to whom the voice is to be output, the information processing device 100 can inform the user who has uttered the voice which other users the voice will be transmitted to. That is, the information processing device 100 of the second embodiment can allow the user to recognize the voice transmission range in a situation where users communicate with each other using a virtual space. The information processing device 100 may also display a listening area. This allows the information processing device 100 to inform the speaking user which other users the voice will be transmitted to. Furthermore, the information processing device can also inform the user who is speaking.

[0037] Furthermore, the information processing device 100 of the second embodiment sets a listening area that is an area including a speaking avatar, sets an avatar present in the listening area as a listening avatar, and outputs a voice to the user of the listening avatar. This allows the information processing device 100 to transmit a voice to other users without the speaking user having to specify a target user.

[0038] Furthermore, the information processing device 100 of the second embodiment may set the listening area according to the volume of the voice uttered by the user. This allows the information processing device 100 to, for example, set a larger listening area if the volume is louder and a smaller listening area if the volume is quieter. Therefore, the user can freely determine the range to which the voice is to be transmitted by controlling the volume of the voice. Furthermore, the information processing device 100 may set the listening area according to the facial orientation of the listening avatar. This allows the information processing device 100 to, for example, set a larger listening area in the direction in which the listening avatar is facing than in the direction in which the listening avatar is not facing. Therefore, the user will turn their avatar toward the other avatar to whom they want to transmit their voice. In this case, the user of the other avatar can easily determine whether the other avatar is speaking to them. As described above, the information processing device 100 can provide the user with a method of transmitting voice similar to that in real space.

[0039] <Third embodiment> Next, an information processing apparatus according to a third embodiment will be described, with some of the description overlapping with the first and second embodiments being omitted.

[0040] [Details of the information processing device 101] 9 is a block diagram showing an example of the functional configuration of an information processing device 101 according to the third embodiment. The information processing device 101 is a device having a configuration that is partially different from that of the information processing device 100 according to the first and second embodiments. Like the information processing device 100, the information processing device 101 is communicably connected to a plurality of user terminals 200 via a wireless or wired network. As shown in FIG. 9, the information processing device 101 includes a detection unit 110, an audio control unit 123, and a display control unit 131.

[0041] The audio control unit 123 includes a listening area setting unit 124 and an audio output unit 125. In addition to the functions of the listening area setting unit 121, the listening area setting unit 124 may have the following functions: The listening area setting unit 124 sets a volume control area within the listening area. The volume control area is an area in which the volume output to the user of an avatar (i.e., a listening avatar) present within the volume control area is set. Here, the volume of the audio output to the user of the listening avatar is referred to as the output volume.

[0042] FIG. 10A is a diagram showing an example of a volume control area. As shown in FIG. 10A, the volume control area is included in the listening area. In this example, the listening area includes volume control area X, volume control area Y, and volume control area Z. An output volume is set for each volume control area. For example, the output volume of a user of a listening avatar in volume control area Y is greater than the output volume of a user of a listening avatar in volume control area X. Also, for example, the output volume of a user of a listening avatar in volume control area Z is greater than the output volume of a user of a listening avatar in volume control area Y. In this way, the output volume is set to be greater as the distance between the speaking avatar and the listening avatar becomes closer.

[0043] The example of the volume control area is not limited to this example. FIG. 10B is a diagram showing another example of the volume control area. For example, as shown in FIG. 10B, the listening area setting unit 124 may set the volume control area according to the facial direction of the speaking avatar, similar to the listening area. In this case, the listening area setting unit 124 sets a large volume control area in the direction in which the avatar's face is facing. Note that the number of volume control areas may be one or more and is not limited to this example. The listening area setting unit 124 may also determine the size of the volume control area according to the volume of the voice of the speaking user.

[0044] The audio output unit 125 may have the following functions in addition to the functions of the audio output unit 122. The audio output unit 125 outputs audio at different volumes depending on the position of the listening avatar. For example, in the example of FIG. 10A , one listening avatar exists in each of the volume control areas X and Y. In this case, the audio output unit 125 outputs audio at a louder output volume to the user of the listening avatar existing in the volume control area Y than to the user of the listening avatar existing in the volume control area X. In other words, the audio output unit 125 outputs audio at a louder output volume to the user of the listening avatar located closer to the position of the speaking avatar. In other words, the audio output unit 125 attenuates the output volume as the distance from the position of the speaking avatar increases.

[0045] In the above example, the listening area setting unit 124 sets the volume control area, but the method for controlling the output volume is not limited to this example. For example, the volume control area may not be set, and the audio output unit 125 may acquire the distance between the speaking avatar and the listening avatar. Then, the audio output unit 125 may control the output volume so that the output volume increases as the distance decreases.

[0046] In this way, the audio control unit 123 controls the output volume, which is the volume of the audio output to the user of the listening avatar, according to the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the audio.

[0047] The display control unit 131 may have the following functions in addition to the functions of the display control unit 130. The display control unit 131 may display a listening avatar in a different display mode depending on the output volume. FIG. 11 is a diagram showing an example of the display mode of a listening avatar. In the example of FIG. 11, listening avatar A and listening avatar B exist within the listening area. At this time, listening avatar A exists closer than listening avatar B. Then, it is assumed that the audio output unit 125 outputs audio to the user of listening avatar A at a louder output volume than to the user of listening avatar B. In such a case, the display control unit 131 displays listening avatar A and listening avatar B in different display modes. In the example of FIG. 11, two exclamation marks are assigned to listening avatar A, and one exclamation mark is assigned to listening avatar B. This allows the speaking user to visually recognize to which user and at what output volume the audio is being transmitted. Note that the display mode is not limited to this example. For example, the display control unit 131 may assign different symbols or characters to listening avatars with different output volumes, or may change the color of part or all of the listening avatar for each listening avatar.

[0048] The display control unit 131 may also display the volume control area. Furthermore, the display control unit 131 may display the volume control area in a different display mode for each volume control area, as in the examples of Figures 10A and 10B. That is, the display control unit 131 may display the listening area in a different display mode depending on the output volume.

[0049] [Example of operation of information processing device 101] Next, an example of the operation of the information processing device 101 will be described with reference to FIG. 12. FIG. 12 is a flowchart illustrating an example of the operation of the information processing device 101. When the detection unit 110 detects a voice ("Yes" in S201), the listening area setting unit 124 sets a listening area based on the volume of the voice and the facial direction of the speaking avatar (S202). At this time, the listening area setting unit 124 sets the listening area taking into consideration at least one of the volume of the voice and the facial direction of the speaking avatar. At this time, the listening area setting unit 124 may also set a volume control area within the listening area. When the detection unit 110 does not detect a voice ("No" in S201), the information processing device 101 ends the processing.

[0050] The display control unit 131 displays the listening area (S203). At this time, the display control unit 131 may also display a volume control area. If the audio output unit 125 has identified a listening avatar ("Yes" in S204), the audio output unit 125 outputs audio at an output volume corresponding to each user of the listening avatar (S205). At this time, the audio output unit 125 outputs audio at an output volume according to the positions of the speaking avatar and the listening avatar. The display control unit 131 changes the display mode of the listening avatar (S206). At this time, the display control unit 131 may display the listening avatar in a different display mode depending on the position of the listening avatar. If the audio output unit 125 has not identified a listening avatar ("No" in S204), the information processing device 101 ends the processing. In the above example of operation, the process of S203 does not have to be performed, and the order of the processes of S205 and S206 may be reversed.

[0051] In this way, the information processing device 101 of the third embodiment may control the output volume, which is the volume of sound output to the user of the listening avatar, depending on the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the sound. As a result, for example, when the voice of the speaking user is louder, the information processing device 101 can output sound at a higher output volume to the user of the listening avatar. Furthermore, the information processing device 101 can output sound at a higher output volume to the user of the listening avatar who is closer to the speaking avatar. In this way, the information processing device 101 can attenuate the output volume as the distance from the speaking avatar increases. Therefore, the information processing device 101 can provide the user with a method of transmitting sound similar to that in real space.

[0052] Furthermore, the information processing device 101 of the third embodiment may display a listening avatar in a different display mode depending on the output volume. Furthermore, the information processing device 101 may display a listening area in a different display mode depending on the output volume. This allows the information processing device 101 to inform the user who has spoken what volume the voice is being transmitted to and to which user.

[0053] [Variations] In the above embodiment, an example has been described in which the range in which the voice is transmitted and the output volume are set based on the voice uttered by the speaking user. The range in which the voice is transmitted and the output volume may be changed by a user operation.

[0054] 13 is a block diagram showing an example of the functional configuration of a modified information processing device 102. As shown in FIG. 13, the information processing device 102 is obtained by adding a mode change unit 140 to the information processing device 101.

[0055] The mode change unit 140 changes the mode of the audio output method. The modes include, for example, an automatic control mode and a user-specified mode. The automatic control mode is a mode in which the range in which audio is transmitted, the output volume, etc. are automatically set based on the audio uttered by the speaking user, as described in the above embodiment. The user-specified mode is a mode in which the range in which audio is transmitted, the output volume, etc. are set based on the user's specification. The user selects a mode by, for example, operating the user terminal 200. The mode change unit 140 acquires information indicating the selected mode from the user terminal 200, and changes the mode to the selected mode.

[0056] Assume that the user has selected the user-specified mode. At this time, the mode change unit 140 accepts designation of the transmission range and output volume. At this time, the mode change unit 140 may accept designation of a predetermined area in the virtual space as the listening area as the transmission range, or may accept designation of a specific user as the transmission range.

[0057] When the mode change unit 140 receives a designation of an area and an output volume, the audio output unit 125 may output audio at a designated volume, which is the designated output volume, to the user of the avatar in the designated area. Also, when the mode change unit 140 receives a designation of a specific user and an output volume, the audio may be output at the designated volume to the specific user.

[0058] The display control unit 131 may change the display mode of the avatar of the user to whom the sound is to be output. The display control unit 131 may also display the designated area as the listening area.

[0059] <Examples of application scenarios> Next, an example of a scene to which the information processing device of the present disclosure is applied will be described. Note that the following description is merely an example, and the scene to which the information processing device of the present disclosure is applied is not limited to the following scene.

[0060] For example, in the event of a disaster, a disaster response room is set up to enable collaboration and information sharing with people in remote locations. In such a case, a user communicates with members of the disaster response room through a virtual space. For example, suppose that members of the disaster response room are divided into multiple groups and holding a meeting in the virtual space. When a user is holding a meeting in one group, that user can hear the voices of other groups at reduced volume. This allows the user to understand the progress of the other groups. Furthermore, the user can sense the commotion around them and grasp changes in the situation.

[0061] <Example of hardware configuration of information processing device> The hardware constituting the information processing devices of the first, second, and third embodiments described above will now be described. Fig. 14 is a block diagram showing an example of the hardware configuration of a computer device that realizes the information processing device in each embodiment. The information processing device and information processing method described in each embodiment and each modified example are realized in a computer device 90.

[0062] 14, a computer device 90 includes a processor 91, a RAM (Random Access Memory) 92, a ROM (Read Only Memory) 93, a storage device 94, an input / output interface 95, a bus 96, and a drive device 97. Note that the information processing device may be realized by a plurality of electric circuits.

[0063] The storage device 94 stores a program (computer program) 98. The processor 91 executes the program 98 of the information processing device using the RAM 92. Specifically, for example, the program 98 includes a program that causes a computer to execute the processes shown in FIGS. 4, 8, and 12. The processor 91 executes the program 98 to realize the functions of the components of the information processing device. The program 98 may be stored in the ROM 93. The program 98 may also be recorded on the storage medium 80 and read out using the drive device 97, or may be transmitted to the computer device 90 from an external device (not shown) via a network (not shown).

[0064] The input / output interface 95 exchanges data with peripheral devices (such as a keyboard, a mouse, and a display device) 99. The input / output interface 95 functions as a means for acquiring or outputting data. The bus 96 connects the components.

[0065] There are various variations in the method for realizing the information processing device. For example, the information processing device can be realized as a dedicated device. Furthermore, the information processing device can be realized based on a combination of multiple devices.

[0066] The scope of each embodiment also includes a processing method for recording a program for realizing each component of the function of each embodiment on a storage medium, reading the program recorded on the storage medium as code, and executing it on a computer. That is, a computer-readable storage medium is also included in the scope of each embodiment. Furthermore, the storage medium on which the above-mentioned program is recorded and the program itself are also included in each embodiment.

[0067] The storage medium may be, but is not limited to, a floppy disk, a hard disk, an optical disk, a magneto-optical disk, a CD (Compact Disc)-ROM, a magnetic tape, a non-volatile memory card, or a ROM. The programs recorded on the storage medium are not limited to standalone programs that execute processes, but also include programs that run on an OS (Operating System) in cooperation with other software and functions of an expansion board.

[0068] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0069] The above-described embodiments and modifications can be combined as appropriate.

[0070] A part or all of the above-described embodiments may be described as, but not limited to, the following supplementary notes: <Additional Notes> [Appendix 1] a detection means for detecting a voice uttered by a user operating an avatar in a virtual space; a voice control means for outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice; and a display control means for changing the display mode of a listening avatar that satisfies the predetermined condition. Information processing device.

[0071] [Appendix 2] the voice control means sets a listening area that is an area including the speaking avatar, designates an avatar present in the listening area as the listening avatar, and outputs the voice to a user of the listening avatar. 10. The information processing device according to claim 1.

[0072] [Appendix 3] the display control means displays the listening area. 3. The information processing device according to claim 2.

[0073] [Appendix 4] the audio control means sets the listening area in accordance with the volume of the audio uttered by the user; 4. The information processing device according to claim 2 or 3.

[0074] [Appendix 5] the audio control means determines the listening area according to a facial orientation of the listening avatar; 5. The information processing device according to any one of Supplementary Notes 2 to 4.

[0075] [Appendix 6] the voice control means controls an output volume, which is a volume of the voice output to the user of the listening avatar, in accordance with the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the voice; 6. An information processing device according to any one of Supplementary Notes 2 to 5.

[0076] [Appendix 7] the display control means displays the listening avatar in a different display mode depending on the output volume. 7. The information processing device according to claim 6.

[0077] [Appendix 8] the display control means displays the listening area in different display modes depending on the output volume. 8. The information processing device according to claim 6 or 7.

[0078] [Appendix 9] further comprising a mode change means for changing the mode of the audio output method, When the mode change means receives a selection of a user-specified mode, the mode change means further receives a transmission range indicating an object to which the audio is to be transmitted and a volume at which the audio is to be output; the audio control means outputs the audio at a designated volume that is a designated volume to users of avatars within the communication range; 9. An information processing device according to any one of appendices 1 to 8.

[0079] [Appendix 10] Detects the voices emitted by users operating avatars in virtual space, outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice; changing the display mode of a listening avatar that satisfies the predetermined condition; Information processing methods.

[0080] [Appendix 11] In the step of outputting the voice, a listening area is set as an area including the speaking avatar, an avatar existing in the listening area is set as the listening avatar, and the voice is output to a user of the listening avatar. 11. The information processing method according to claim 10.

[0081] [Appendix 12] In the step of changing, the listening area is displayed. 12. The information processing method according to claim 11.

[0082] [Appendix 13] In the step of outputting the sound, the listening area is set according to the volume of the sound uttered by the user. 13. The information processing method according to claim 11 or 12.

[0083] [Appendix 14] In the step of outputting the sound, the listening area is determined according to a facial orientation of the listening avatar. 14. An information processing method according to any one of appendices 11 to 13.

[0084] [Appendix 15] In the step of outputting the sound, an output volume, which is a volume of the sound output to the user of the listening avatar, is controlled according to the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the sound. An information processing method according to any one of appendices 11 to 14.

[0085] [Appendix 16] In the step of outputting the sound, the listening avatar is displayed in a different display mode depending on the output volume. 16. The information processing method according to claim 15.

[0086] [Appendix 17] In the changing step, the listening area is displayed in a different display manner depending on the output volume. 17. The information processing method according to claim 15 or 16.

[0087] [Appendix 18] When the selection of the user-specified mode is accepted, the method further accepts a transmission range indicating a target to which the audio is to be transmitted and a volume at which the audio is to be output; In the step of outputting the sound, the sound is output at a designated volume that is a designated volume to users of avatars within the transmission range. An information processing method according to any one of appendices 10 to 17.

[0088] [Appendix 19] A process of detecting a voice uttered by a user operating an avatar in a virtual space; outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar, which is an avatar operated by the user who uttered the voice; and a process of changing the display mode of a listening avatar that satisfies the predetermined condition. A computer-readable storage medium.

[0089] [Appendix 20] In the process of outputting the voice, a listening area is set as an area including the speaking avatar, an avatar existing in the listening area is set as the listening avatar, and the voice is output to a user of the listening avatar. 20. The computer-readable storage medium of claim 19.

[0090] [Appendix 21] In the process of changing, the listening area is displayed. 21. The computer-readable storage medium of claim 20.

[0091] [Appendix 22] In the process of outputting the sound, the listening area is set according to the volume of the sound uttered by the user. 22. The computer-readable storage medium of claim 20 or 21.

[0092] [Appendix 23] In the process of outputting the sound, the listening area is determined according to a facial orientation of the listening avatar. 23. A computer-readable storage medium according to any one of claims 20 to 22.

[0093] [Appendix 24] In the process of outputting the sound, an output volume, which is a volume of the sound output to the user of the listening avatar, is controlled according to the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the sound. 24. A computer-readable storage medium according to any one of appendices 20 to 23.

[0094] [Appendix 25] In the process of outputting the audio, the listening avatar is displayed in a different display mode depending on the output volume. 25. The computer-readable storage medium of claim 24.

[0095] [Appendix 26] In the changing process, the listening area is displayed in a different display mode depending on the output volume. 26. The computer-readable storage medium of claim 24 or 25.

[0096] [Appendix 27] When the selection of the user-specified mode is accepted, the method further accepts a transmission range indicating a target to which the audio is to be transmitted and a volume at which the audio is to be output; In the process of outputting the sound, the sound is output at a designated volume that is a designated volume to users of avatars within the transmission range. 27. The computer-readable storage medium of any one of claims 19 to 26. [Explanation of symbols]

[0097] 100, 101 Information processing device 110 Detection unit 120, 123 Audio control unit 121, 124 Listening area setting unit 122, 125 Audio output section 130, 131 Display control unit 140 Mode change section 200 user terminals

Claims

1. a distance calculation unit that calculates a distance between a first avatar operated in a virtual space and a second avatar different from the first avatar; a display control unit that displays a predetermined display on the second avatar when the distance is within a predetermined distance, Information processing device.

2. the display control unit displays an exclamation mark on the second avatar when the distance is within the predetermined distance. The information processing device according to claim 1 .

3. an acquisition unit that acquires voice data of a user of the first avatar; a setting unit that sets the predetermined distance in accordance with the volume of the sound based on the audio data, The information processing device according to claim 1 .

4. When the volume level is greater than a predetermined threshold, the setting unit sets the size of the predetermined distance to be greater than a reference size. The information processing device according to claim 3 .

5. When the volume level is lower than a predetermined threshold, the setting unit sets the size of the predetermined distance to be smaller than a reference size.

5. The information processing device according to claim 3.

6. an acquisition unit that acquires voice data of a user of the first avatar; a setting unit that sets a size of a listening area, which is a range indicating the predetermined distance from the first avatar, in accordance with a volume level based on the audio data; the display control unit performs the predetermined display on the second avatar when the second avatar is present within the listening area; The information processing device according to claim 1 .

7. an output unit that outputs a sound based on the audio data to a user of the second avatar at a volume corresponding to a position of the second avatar within the listening area; The volume is determined according to the position within the listening area. The information processing device according to claim 6 .

8. calculating a distance between a first avatar operated in a virtual space and a second avatar different from the first avatar; If the distance is within a predetermined distance, a predetermined display is performed on the second avatar. Information processing methods.

9. A process of calculating a distance between a first avatar operated in a virtual space and a second avatar different from the first avatar; and if the distance is within a predetermined distance, displaying a predetermined image on the second avatar. Computer program.

Citation Information

Patent Citations

  • Avatar display device in virtual space communication system, avatar displaying method and storage medium

    JP2001160154A

  • Information processing method, information processing program, information processing system, and information processing device

    JP2018163420A

  • Information processing device, information processing method, and program

    WO2018135304A1

  • JP36871A

  • Information processing device, information processing method, and program

    WO2018020766A1