Information Processing Apparatus, Information Processing Method, and Computer Program

The information processing apparatus addresses the challenge of unspecified voice transmission in virtual spaces by detecting user voice and adjusting display modes to indicate transmission targets, enhancing user awareness and simplifying operations.

JP7708190B2Active Publication Date: 2025-07-15NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023544951
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-03
Publication Date
2025-07-15
Estimated Expiration
2041-09-03

Smart Images

  • Figure 0007708190000001
    Figure 0007708190000001
  • Figure 0007708190000002
    Figure 0007708190000002
  • Figure 0007708190000003
    Figure 0007708190000003
Patent Text Reader

Abstract

One aim of the present invention is to provide an information processing device whereby a user can be made aware of the audio transmission range in a situation where a virtual space is used and users communicate with each other. This information processing device comprises: a detection means that detects audio generated by a user operating an avatar inside a virtual space; an audio control means that outputs the audio to a user of an avatar that fulfills prescribed conditions in a relationship with a speaking avatar, being an avatar operated by the user that provided the audio; and a display control means that changes the display mode for a listener avatar, being an avatar that fulfills the prescribed conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for controlling a virtual space.

Background Art

[0002] There are techniques for multiple users to communicate using a virtual space. For example, Patent Document 1 discloses a technique of placing an object in which images of each user are inserted in a three-dimensional space and conducting a call through the three-dimensional space.

[0003] Also, in relation to techniques for multiple users to communicate, Patent Document 2 discloses generating an image corresponding to a real space, in which an object of a person is placed at a position on the image corresponding to the position where the person is in the real space. In the technique of Patent Document 2, when there are people on a call, a link connecting the objects corresponding to the people on the call is generated.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the techniques of Patent Documents 1 and 2, it is possible to make a call with a pre-specified user in either case.

[0006] Here, consider the case where a user uses a virtual space in which there are a plurality of avatars. For example, the user operates an avatar representing himself / herself to move within the virtual space or communicate with other users who operate other avatars. In such a case, if it is necessary to specify the communication target in all the communications made by the user, it will be troublesome for the user's operation. On the other hand, even if it is possible to transmit the voice uttered by the user without specifying the communication target, there is a risk that the user does not know to which user the voice is transmitted.

[0007] The present disclosure has been made in view of the above problems, and one of the objectives is to provide an information processing apparatus or the like that can cause a user to recognize the voice transmission range in a situation where users communicate with each other using a virtual space.

Means for Solving the Problems

[0008] An information processing apparatus according to an aspect of the present disclosure includes a detection unit that detects a voice uttered by a user who operates an avatar in a virtual space, and a voice control unit that outputs the voice to a user of an avatar that satisfies a predetermined condition in relation to the speaking avatar, which is the avatar operated by the user who uttered the voice, and a display control unit that changes a display mode of a listening avatar, which is the avatar that satisfies the predetermined condition.

[0009] An information processing method according to an aspect of the present disclosure detects a voice uttered by a user who operates an avatar in a virtual space, outputs the voice to a user of an avatar that satisfies a predetermined condition in relation to the speaking avatar, which is the avatar operated by the user who uttered the voice, and changes a display mode of a listening avatar, which is the avatar that satisfies the predetermined condition.

[0010] A computer-readable storage medium according to one aspect of the present disclosure stores a program that causes a computer to execute a process of detecting a voice uttered by a user who operates an avatar in a virtual space, a process of outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to a speaking avatar that is the avatar operated by the user who uttered the voice, and a process of changing a display mode of a listening avatar that is the avatar that satisfies the predetermined condition.

Advantages of the Invention

[0011] According to the present disclosure, in a situation where users communicate with each other using a virtual space, the voice transmission range can be made recognizable to the users.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 11

Figure 12

Figure 13

Figure 14

Modes for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0014] <First Embodiment> The outline of the information processing apparatus of the present disclosure will be described.

[0015] FIG. 1 is a diagram schematically showing an example of a configuration including an information processing apparatus 100. As shown in FIG. 1, the information processing apparatus 100 is communicably connected via a wireless or wired network to user terminals 200-1, 200-2, ···, 200-n (n is a natural number of 1 or more). Here, when not distinguishing each of the user terminals 200-1, 200-2, ···, 200-n, it is simply referred to as the user terminal 200. The user terminal 200 is a device operated by a user. The user terminal 200 is, for example, a personal computer, but is not limited to this example. The user terminal 200 may be a smartphone or a tablet terminal, or may be a device including a goggle-type wearable terminal (also referred to as a head-mounted display) having a display. Further, the user terminal 200 includes an input device such as a keyboard, a mouse, a microphone, and a wearable device that performs operations based on the user's actions, and an output device such as a display and a speaker. Furthermore, the user terminal 200 includes at least one of a photographing device or a device capable of reading voiceprints, fingerprints, palm prints, irises, veins, and the like.

[0016] First, the virtual space in the present disclosure will be described. The virtual space is a virtual space shared by a plurality of users, and is a space in which the operations of the users are reflected. The virtual space is also called a VR (Virtual Reality) space. For example, the virtual space is provided by the information processing apparatus 100. The user terminal 200 displays an image showing the virtual space. FIG. 2 is a diagram schematically showing an example of the virtual space displayed on the user terminal 200. In the example of FIG. 2, the virtual space is displayed on the display of the user terminal 200. As shown in FIG. 2, the virtual space includes an avatar. The avatar is an object that the user operates. The user uses the virtual space by operating the avatar. For example, an image of the virtual space from the avatar's perspective operated by the user may be displayed on the user terminal 200. In this case, the image displayed on the user terminal 200 may be updated according to the actions of the avatar. Also, for example, the user may be able to communicate with other users by performing an action on the avatar operated by another user. Note that the apparatus that provides the virtual space does not have to be the information processing apparatus 100. For example, an external apparatus (not shown) may provide the virtual space.

[0017] FIG. 3 is a block diagram showing an example of the functional configuration of the information processing apparatus 100 according to the first embodiment. As shown in FIG. 3, the information processing apparatus 100 includes a detection unit 110, an audio control unit 120, and a display control unit 130.

[0018] The detection unit 110 detects the voice uttered by the user who operates the avatar in the virtual space. The detection unit 110 is an example of a detection means.

[0019] The voice control unit 120 controls voice. Here, the user who uttered the voice is also referred to as the speaking user. Also, the avatar operated by the user who uttered the voice is also referred to as the speaking avatar. The voice control unit 120, for example, in relation to the speaking avatar, identifies an avatar that satisfies a predetermined condition. The avatar that satisfies the predetermined condition may be, for example, an avatar that exists within a predetermined distance from the speaking avatar, or an avatar that exists in a predetermined area including the speaking avatar. Note that the predetermined condition is not limited to this example. Also, the avatar that satisfies the predetermined condition is also referred to as the listening avatar. The voice control unit 120 outputs, for example, the voice from the speaking user to the user of the identified avatar. In this way, the voice control unit 120 outputs voice to the user of the avatar that satisfies the predetermined condition in relation to the speaking avatar, which is an avatar operated by the user who uttered the voice. The voice control unit 120 is an example of voice control means.

[0020] The display control unit 130 controls the display of the virtual space. For example, when there is a listening avatar, which is an avatar that satisfies a predetermined condition, the display control unit 130 controls the display mode of the listening avatar. For example, the display control unit 130 assigns a predetermined symbol or a predetermined color to the listening avatar. The display control unit 130 changes the display mode of the listening avatar, which is an avatar that satisfies the predetermined condition. The display control unit 130 is an example of display control means.

[0021] Next, an example of the operation of the information processing apparatus 100 will be described with reference to FIG. 4. In the present disclosure, each step of the flowchart is expressed using the number assigned to each step, such as "S1".

[0022] FIG. 4 is a flowchart for explaining an example of the operation of the information processing apparatus 100. The detection unit 110 detects the voice uttered by the user who operates the avatar in the virtual space (S1). The voice control unit 120 outputs the voice to the user of the avatar that satisfies a predetermined condition in relation to the speaking avatar which is the avatar operated by the user who uttered the voice (S2). The display control unit 130 changes the display mode of the listening avatar which is the avatar that satisfies the predetermined condition (S3).

[0023] As described above, the information processing apparatus 100 according to the first embodiment detects the voice uttered by the user who operates the avatar in the virtual space, and outputs the voice to the user of the avatar that satisfies a predetermined condition in relation to the speaking avatar which is the avatar operated by the user who uttered the voice. Then, the information processing apparatus 100 changes the display mode of the listening avatar which is the avatar that satisfies the predetermined condition. Thereby, since the information processing apparatus 100 controls the display mode of the avatar of the user who is the target of voice output, it can notify the speaking user of which user the voice is transmitted to. That is, the information processing apparatus 100 according to the present disclosure can make the user recognize the voice transmission range in the situation where users communicate with each other using the virtual space.

[0024] <Second Embodiment> Next, the information processing apparatus according to the second embodiment will be described. In the second embodiment, the information processing apparatus 100 described in the first embodiment will be described in more detail.

[0025] [Details of Information Processing Apparatus 100] FIG. 5 is a block diagram showing an example of the functional configuration of the information processing apparatus 100 according to the second embodiment. As shown in FIG. 5, the information processing apparatus 100 includes a detection unit 110, a voice control unit 120, and a display control unit 130.

[0026] The detection unit 110 detects the voice uttered by the user. For example, when the user utters a voice, it is picked up by a microphone or the like provided in the user terminal 200. The voice data, which is data regarding the picked-up voice, is transmitted to the information processing apparatus 100. The detection unit 110 detects the voice uttered by the user, for example, by receiving the voice data.

[0027] The voice control unit 120 includes a listening area setting unit 121 and a voice output unit 122. The listening area setting unit 121 sets a listening area. The listening area is an area that includes the speaking avatar and indicates the range within which the voice of the speaking user is transmitted. The listening area may be, for example, a range indicating a certain distance from the speaking avatar. FIG. 6A is a diagram showing an example of the listening area. For example, as shown in FIG. 6A, the listening area setting unit 121 may set the listening area in a circular shape centered on the speaking avatar. In this case, another avatar existing inside the listening area is the listening avatar. That is, in the user terminal 200 used by the user who operates another avatar existing inside the listening area, the voice of the speaking user is output. Here, the size of the listening area may be determined to be a constant size for each avatar, or may vary depending on the volume of the voice. Specifically, the listening area setting unit 121 acquires information on the loudness of the voice of the speaking user, that is, the volume, from the voice data of the speaking user. When the volume is greater than a predetermined threshold value, the listening area setting unit 121 sets the size of the listening area to be larger than the reference size. When the volume is less than the predetermined threshold value, the listening area setting unit 121 sets the size of the listening area to be smaller than the reference size. Note that this is not limited to this example, and the size of the listening area may be preset for each of a plurality of volume ranges, and the listening area setting unit 121 may determine the listening area according to which range the acquired volume belongs to. In this way, the listening area setting unit 121 may set the listening area according to the volume of the voice uttered by the user.

[0028] The listening area is not limited to the above example. For example, the listening area may vary depending on the face direction of the speaking avatar. FIG. 6B is a diagram showing a second example of the listening area. As shown in FIG. 6B, the listening area may be set wider in the direction in which the face of the speaking avatar is facing and narrower in the direction in which the face of the speaking avatar is not facing. In the example of FIG. 6B, the speaking avatar at point X is facing in the direction of point Q. Point P indicates a position behind the speaking avatar. Points P, X, and Q are points on a straight line, and points P and Q are points on the circumference of the listening area. In this case, the distance from point X to point Q is longer than the distance from point X to point P. Thus, the listening area setting unit 121 may set the listening area wider in the direction in which the face of the speaking avatar is facing. That is, the listening area setting unit 121 may determine the listening area according to the face direction of the speaking avatar.

[0029] Also, the listening area may be set in different forms. FIG. 6C is a diagram showing a third example of the listening area. In the example of FIG. 6C, the listening area is set in a fan shape. In this example, the speaking avatar is facing in the direction of point R. Point R is a point on the arc of the fan. That is, also in the example of FIG. 6C, the listening area is set wider in the direction in which the face of the speaking avatar is facing. Thus, the form of the listening area may be a circle or a fan shape, or may be other forms such as an ellipse or a polygon.

[0030] The voice output unit 122 outputs voice. Specifically, when another avatar different from the speaking avatar exists in the listening area, the voice output unit 122 identifies the other avatar as the listening avatar. Then, the voice output unit 122 outputs voice to the user terminal 200 used by the user who operates the listening avatar. Here, in the present disclosure, "outputting voice on the user terminal 200 used by the user who operates the avatar" may also be expressed as "outputting voice to the user of the avatar". The voice output unit 122 may identify another avatar as the listening avatar by detecting that another avatar has entered the listening area. Further, the voice output unit 122 may acquire the position information of other avatars around the speaking avatar, and among the acquired position information, identify the avatar with the position information indicating a position within the listening area as the listening avatar.

[0031] The display control unit 130 controls the display of the virtual space. Specifically, the display control unit 130 changes the display mode of the listening avatar. FIG. 7 is a diagram showing an example of the display mode of the listening avatar. In the example of FIG. 7, an exclamation mark is given to the listening avatar. The display control unit 130 controls, for example, the user terminal 200 of the speaking user so that the listening avatar with the exclamation mark is displayed. Thereby, the speaking user can visually recognize which avatar is the listening avatar, that is, to which user the voice is being transmitted. Note that the display mode is not limited to this example. For example, the display control unit 130 may give other symbols or characters to the listening avatar, or may change the color of a part or the whole of the listening avatar.

[0032] Further, the display control unit 130 may display the listening area. For example, the display control unit 130 displays the listening area set according to the voice of the speaking user on the user terminal 200 of the speaking user. Thereby, the speaking user can recognize in which range the voice is transmitted. Also, the display control unit 130 may display the listening area set according to the voice of the speaking user on the user terminals 200 of other users as well. Thereby, other users can recognize who is speaking.

[0033] [Operation Example of Information Processing Device 100] Next, an example of the operation of the information processing device 100 will be described with reference to FIG. 8. FIG. 8 is a flowchart for explaining an example of the operation of the information processing device 100. When the detection unit 110 detects voice ("Yes" in S101), the listening area setting unit 121 sets a listening area based on the volume of the voice and the face direction of the speaking avatar (S102). At this time, the listening area setting unit 121 may set the listening area in consideration of at least one of the volume of the voice and the face direction of the speaking avatar. When the detection unit 110 does not detect voice ("No" in S101), the information processing device 100 ends the process.

[0034] The display control unit 130 displays the listening area (S103). When the listening avatar is identified by the voice output unit 122 ("Yes" in S104), the voice output unit 122 outputs voice to the user of the listening avatar (S105). The display control unit 130 changes the display mode of the listening avatar (S106). When the listening avatar is not identified by the voice output unit 122 ("No" in S104), the information processing device 100 ends the process.

[0035] Note that in the above operation example, the process of S103 may not be performed. Also, the order of the processes of S105 and S106 may be reversed.

[0036] In this way, the information processing apparatus 100 of the second embodiment detects the voice uttered by the user who operates the avatar in the virtual space, and outputs the voice to the user of the avatar that satisfies a predetermined condition in relation to the speaking avatar which is the avatar operated by the user who uttered the voice. Then, the information processing apparatus 100 changes the display mode of the listening avatar which is the avatar that satisfies the predetermined condition. Thereby, since the information processing apparatus 100 controls the display mode of the avatar of the user who is the target of voice output, it can notify the speaking user of which user the voice is transmitted to. That is, the information processing apparatus 100 of the second embodiment can cause the user to recognize the voice transmission range in a situation where users communicate with each other using the virtual space. Further, the information processing apparatus 100 may display the listening area. Thereby, the information processing apparatus 100 can notify the speaking user of which user the voice is transmitted to. Furthermore, the information processing apparatus can also notify the user of who is speaking.

[0037] Also, the information processing apparatus 100 of the second embodiment sets a listening area which is an area including the speaking avatar, sets the avatars existing in the listening area as listening avatars, and outputs voice to the users of the listening avatars. Thereby, the information processing apparatus 100 can transmit voice to other users without the speaking user designating the target user.

[0038] In addition, the information processing apparatus 100 according to the second embodiment may set the listening area according to the volume of the voice uttered by the user. Thereby, the information processing apparatus 100 can, for example, set a larger listening area when the volume is large and a smaller listening area when the volume is small. Therefore, the user can determine the range in which they want to freely convey voice by controlling the loudness of their voice. Furthermore, the information processing apparatus 100 may determine the listening area according to the face orientation of the listening avatar. Thereby, the information processing apparatus 100 can, for example, set a larger listening area in the direction in which the face of the listening avatar is facing than in the direction in which the face is not facing. Therefore, the user will face their avatar in the direction of the other avatar to whom they want to convey voice. In this case, it becomes easier for the user of the other avatar to determine whether they are speaking towards themselves. As described above, the information processing apparatus 100 can provide the user with a voice transmission method similar to that in the real space.

[0039] <Third Embodiment> Next, the information processing apparatus according to the third embodiment will be described. Note that descriptions overlapping with those of the first and second embodiments will be partially omitted.

[0040] [Details of Information Processing Apparatus 101] FIG. 9 is a block diagram showing an example of the functional configuration of the information processing apparatus 101 according to the third embodiment. The information processing apparatus 101 is a device with a partially different configuration from the information processing apparatus 100 of the first and second embodiments. Similar to the information processing apparatus 100, the information processing apparatus 101 is communicably connected to a plurality of user terminals 200 via a wireless or wired network. As shown in FIG. 9, the information processing apparatus 101 includes a detection unit 110, an audio control unit 123, and a display control unit 131.

[0041] The voice control unit 123 includes a listening area setting unit 124 and a voice output unit 125. The listening area setting unit 124 may have the following functions in addition to the functions of the listening area setting unit 121. The listening area setting unit 124 sets a volume control area within the listening area. The volume control area is an area where the volume output to the user of the avatar (i.e., the listening avatar) existing within the volume control area is set. Here, the volume of the voice output to the user of the listening avatar is referred to as the output volume.

[0042] Figure 10A is a diagram showing an example of the volume control area. As shown in Figure 10A, the volume control area is included in the listening area. In this example, the listening area includes a volume control area X, a volume control area Y, and a volume control area Z. An output volume is determined for each of the volume control areas. For example, the output volume of the user of the listening avatar existing in the volume control area Y is greater than the output volume of the user of the listening avatar existing in the volume control area X. Also, for example, the output volume of the user of the listening avatar existing in the volume control area Z is greater than the output volume of the user of the listening avatar existing in the volume control area Y. Thus, the output volume is set larger as the distance between the speaking avatar and the listening avatar becomes closer.

[0043] Also, the example of the volume control area is not limited to this example. Figure 10B is a diagram showing another example of the volume control area. For example, as shown in Figure 10B, the listening area setting unit 124 may set the volume control area according to the face direction of the speaking avatar, similar to the listening area. In this case, the listening area setting unit 124 sets the volume control area larger in the direction in which the face of the avatar is facing. Note that the number of volume control areas may be one or more and is not limited to this example. Also, the listening area setting unit 124 may determine the size of the volume control area according to the volume of the voice of the speaking user.

[0044] In addition to the functions of the voice output unit 122, the voice output unit 125 may have the following functions. The voice output unit 125 outputs voice at different volumes according to the position of the listening avatar. For example, in the example of FIG. 10A, assume that there is one listening avatar in each of the volume control region X and the volume control region Y. In this case, the voice output unit 125 outputs voice at a larger output volume to the user of the listening avatar present in the volume control region Y than to the user of the listening avatar present in the volume control region X. That is, the voice output unit 125 outputs voice at a larger output volume to the user of the listening avatar located at a position closer to the position of the speaking avatar. In other words, the voice output unit 125 attenuates the output volume as the distance from the position of the speaking avatar increases.

[0045] Note that in the above example, the listening area setting unit 124 set the volume control area, but the method of controlling the output volume is not limited to this example. For example, the volume control area may not be set, and the voice output unit 125 may acquire the distance between the speaking avatar and the listening avatar. Then, the voice output unit 125 may control the output volume so that the output volume increases as the distance decreases.

[0046] In this way, the voice control unit 123 controls the output volume, which is the volume output to the user of the listening avatar, according to the distance between the position of the speaking avatar and the position of the listening avatar, and the volume of the voice.

[0047] The display control unit 131 may have the following functions in addition to the functions of the display control unit 130. The display control unit 131 may display the listening avatar in different display modes according to the output volume. FIG. 11 is a diagram showing an example of the display mode of the listening avatar. In the example of FIG. 11, there are a listening avatar A and a listening avatar B in the listening area. At this time, the listening avatar A exists at a position closer to the user than the listening avatar B. And it is assumed that the voice output unit 125 outputs voice to the user of the listening avatar A at a larger output volume than the user of the listening avatar B. In such a case, the display control unit 131 displays the listening avatar A and the listening avatar B in different display modes. In the example of FIG. 11, two exclamation marks are assigned to the listening avatar A, and one exclamation mark is assigned to the listening avatar B. Thereby, the speaking user can visually recognize to which user the voice is transmitted and at what output volume. Note that the display mode is not limited to this example. For example, the display control unit 131 may assign different other symbols or characters to the listening avatars with different output volumes, or may change the color of part or all of the listening avatars for each listening avatar.

[0048] Also, the display control unit 131 may display a volume control area. Further, the display control unit 131 may display the volume control area in different display modes for each volume control area as in the examples of FIGS. 10A and 10B. That is, the display control unit 131 may display the listening area in different display modes according to the output volume.

[0049] [Operation Example of Information Processing Apparatus 101] Next, an example of the operation of the information processing apparatus 101 will be described with reference to FIG. 12. FIG. 12 is a flowchart for explaining an example of the operation of the information processing apparatus 101. When the detection unit 110 detects voice ("Yes" in S201), the listening area setting unit 124 sets a listening area based on the volume of the voice and the face direction of the speaking avatar (S202). At this time, the listening area setting unit 124 sets the listening area in consideration of at least one of the volume of the voice and the face direction of the speaking avatar. Also, at this time, the listening area setting unit 124 may set a volume control area within the listening area. When the detection unit 110 does not detect voice ("No" in S201), the information processing apparatus 101 ends the process.

[0050] The display control unit 131 displays the listening area (S203). At this time, the display control unit 131 may also display the volume control area. When the voice output unit 125 identifies the listening avatar ("Yes" in S204), the voice output unit 125 outputs voice at the output volume corresponding to each user of the listening avatar (S205). At this time, the voice output unit 125 outputs voice at the output volume according to the positions of the speaking avatar and the listening avatar. The display control unit 131 changes the display mode of the listening avatar (S206). At this time, the display control unit 131 may display the listening avatar in different display modes according to the position of the listening avatar. When the voice output unit 125 does not identify the listening avatar ("No" in S204), the information processing apparatus 101 ends the process. Note that in the above operation example, the process of S203 may not be performed. Also, the order of the processes of S205 and S206 may be reversed.

[0051] In this way, the information processing apparatus 101 of the third embodiment may control the output volume, which is the volume output to the user of the listening avatar, according to the distance between the position of the speaking avatar and the position of the listening avatar, and the volume of the voice. Thereby, for example, when the voice of the speaking user is louder, the information processing apparatus 101 can output the voice at a larger output volume to the user of the listening avatar. Further, the information processing apparatus 101 can output the voice at a larger output volume to the user of the listening avatar closer to the speaking avatar. In this way, the information processing apparatus 101 can attenuate the output volume as the distance from the speaking avatar increases. Therefore, the information processing apparatus 101 can provide a voice transmission method similar to that in the real space to the user.

[0052] Further, the information processing apparatus 101 of the third embodiment may display the listening avatar in different display modes according to the output volume. Also, the information processing apparatus 101 may display the listening area in different display modes according to the output volume. Thereby, the information processing apparatus 101 can notify the speaking user of which user the voice is transmitted to and at what volume.

[0053] [Modification Example] In the above embodiment, an example in which the range in which the voice is transmitted, the output volume, etc. are set by the voice uttered by the speaking user has been described. The range in which the voice is transmitted and the output volume may be changed by the user's operation.

[0054] FIG. 13 is a block diagram showing an example of the functional configuration of the information processing apparatus 102 of the modification example. As shown in FIG. 13, a mode change unit 140 is added to the information processing apparatus 101 in the information processing apparatus 102.

[0055] The mode change unit 140 changes the mode of the voice output method. There are, for example, an automatic control mode and a user-specified mode. The automatic control mode is a mode in which the range where the voice is transmitted, the output volume, etc. are set automatically according to the voice uttered by the speaking user as described in the above embodiment. The user-specified mode is a mode in which the range where the voice is transmitted, the output volume, etc. are set according to the user's specification. The user selects the mode, for example, by operating the user terminal 200. The mode change unit 140 acquires information indicating the selected mode from the user terminal 200 and changes the mode to the selected mode.

[0056] Suppose the user-specified mode is selected by the user. At this time, the mode change unit 140 accepts the specification of the transmission range and the output volume. At this time, the mode change unit 140 may accept the specification of a predetermined area in the virtual space as the listening area as the transmission range, or may accept the specification of a specific user as the transmission range.

[0057] When the mode change unit 140 accepts the specification of the area and the output volume, the voice output unit 125 may output the voice at the specified volume, which is the specified output volume, to the avatar user in the specified area, for example. Also, when the mode change unit 140 accepts the specification of a specific user and the output volume, the voice may be output at the specified volume to the specified specific user.

[0058] The display control unit 131 may change the display mode of the avatar of the user to whom the voice is output. Also, the display control unit 131 may display the specified area as the listening area.

[0059] <Examples of application scenarios> Next, examples of scenes to which the information processing apparatus of the present disclosure is applied will be described. Note that the following description is also merely an example, and the scenes to which the information processing apparatus of the present disclosure is applied are not limited to the following scenes.

[0060] For example, in the event of a disaster, a disaster control room is set up to cooperate with and share information with people in a distant location. In such a case, the user communicates with the members of the disaster control room via a virtual space. For example, assume that in the virtual space, the members of the disaster control room are divided into multiple groups and are having a meeting. When the user is having a meeting in one group, the user can hear the voices of other groups in a state where the volume is attenuated. Therefore, the user can grasp the progress of other groups. Furthermore, the user can detect the commotion around and perceive changes in the situation.

[0061] <Example of the hardware configuration of the information processing apparatus> The hardware constituting the information processing apparatus according to the first, second, and third embodiments described above will be described. FIG. 14 is a block diagram showing an example of the hardware configuration of a computer apparatus that realizes the information processing apparatus in each embodiment. In the computer apparatus 90, the information processing apparatus and the information processing method described in each embodiment and each modification are realized.

[0062] As shown in FIG. 14, the computer apparatus 90 includes a processor 91, a RAM (Random Access Memory) 92, a ROM (Read Only Memory) 93, a storage device 94, an input / output interface 95, a bus 96, and a drive device 97. Note that the information processing apparatus may be realized by a plurality of electric circuits.

[0063] The memory device 94 stores a program (computer program) 98. The processor 91 executes the program 98 of this information processing apparatus using the RAM 92. Specifically, for example, the program 98 includes programs that cause a computer to execute the processes shown in FIGS. 4, 8, and 12. When the processor 91 executes the program 98, the functions of each component of this information processing apparatus are realized. The program 98 may be stored in the ROM 93. Also, the program 98 may be recorded on the storage medium 80 and read using the drive device 97, or may be transmitted from an external device (not shown) to the computer device 90 via a network (not shown).

[0064] The input / output interface 95 exchanges data with peripheral devices (such as a keyboard, a mouse, and a display device) 99. The input / output interface 95 functions as a means for acquiring or outputting data. The bus 96 connects each component.

[0065] Note that there are various modifications to the method for realizing the information processing apparatus. For example, the information processing apparatus can be realized as a dedicated device. Also, the information processing apparatus can be realized based on a combination of a plurality of devices.

[0066] A processing method in which a program for realizing each component in the functions of each embodiment is recorded on a storage medium, the program recorded on the storage medium is read as code, and executed on a computer is also included in the scope of each embodiment. That is, a computer-readable storage medium is also included in the scope of each embodiment. Also, the storage medium on which the above-described program is recorded, and the program itself are included in each embodiment.

[0067] The storage medium is, for example, a floppy (registered trademark) disk, a hard disk, an optical disk, a magneto-optical disk, a CD (Compact Disc)-ROM, a magnetic tape, a non-volatile memory card, or a ROM, but is not limited to this example. Further, the program recorded on the storage medium is not limited to a program that executes processing alone, and also includes a program that operates on an OS (Operating System) and executes processing in cooperation with other software and the functions of an expansion board within the scope of each embodiment.

[0068] As described above, the present invention has been described with reference to the embodiments, but the present invention is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0069] The above embodiments and modified examples can be combined as appropriate.

[0070] Some or all of the above embodiments can be described as follows in the appended claims, but are not limited thereto <Appended Claim> [Appended Claim 1] Detection means for detecting the voice uttered by a user who operates an avatar in a virtual space, Voice control means for outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to the speaking avatar that is the avatar operated by the user who uttered the voice, Display control means for changing the display mode of a listening avatar that is the avatar that satisfies the predetermined condition, and an information processing apparatus comprising the same. Information processing apparatus.

[0071] [Appended Claim 2] The voice control means sets a listening area that is an area including the speaking avatar, designates an avatar existing in the listening area as the listening avatar, and outputs the voice to the user of the listening avatar. The information processing apparatus according to Appended Claim 1.

[0072] [Appended Claim 3] The display control means displays the listening area. The information processing apparatus according to Supplementary Note 2.

[0073] [Supplementary Note 4] The voice control means sets the listening area according to the volume of the voice uttered by the user. The information processing apparatus according to Supplementary Note 2 or 3.

[0074] [Supplementary Note 5] The voice control means determines the listening area according to the face orientation of the listening avatar. The information processing apparatus according to any one of Supplementary Notes 2 to 4.

[0075] [Supplementary Note 6] The voice control means controls the output volume, which is the volume output to the user of the listening avatar, according to the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the voice. The information processing apparatus according to any one of Supplementary Notes 2 to 5.

[0076] [Supplementary Note 7] The display control means displays the listening avatar in different display modes according to the output volume. The information processing apparatus according to Supplementary Note 6.

[0077] [Supplementary Note 8] The display control means displays the listening area in different display modes according to the output volume. The information processing apparatus according to Supplementary Note 6 or 7.

[0078] [Supplementary Note 9] It further includes mode change means for changing the mode of the voice output method. When the mode change means receives the selection of the user-specified mode, it further receives the specification of the transmission range indicating the target to which the voice is transmitted and the volume of the voice output. The voice control means outputs the voice at the specified volume to the user of the avatar within the transmission range. The information processing apparatus according to any one of Supplementary Notes 1 to 8.

[0079] [Supplementary Note 10] Detecting the voice uttered by a user who operates an avatar in a virtual space, Outputting the voice to a user of an avatar that satisfies a predetermined condition in relation to the speaking avatar, which is the avatar operated by the user who uttered the voice, changing the display mode of the listening avatar, which is the avatar that satisfies the predetermined condition. An information processing method.

[0080] [Supplementary Note 11] In the step of outputting the voice, setting a listening area, which is an area including the speaking avatar, designating the avatar existing in the listening area as the listening avatar, and outputting the voice to the user of the listening avatar. The information processing method according to Supplementary Note 10.

[0081] [Supplementary Note 12] In the step of changing, displaying the listening area. The information processing method according to Supplementary Note 11.

[0082] [Supplementary Note 13] In the step of outputting the voice, setting the listening area according to the volume of the voice uttered by the user. The information processing method according to Supplementary Note 11 or 12.

[0083] [Supplementary Note 14] In the step of outputting the voice, determining the listening area according to the face orientation of the listening avatar. The information processing method according to any one of Supplementary Notes 11 to 13.

[0084] [Supplementary Note 15] In the step of outputting the voice, the output volume, which is the volume output to the user of the listening avatar, is controlled according to the distance between the position of the speaking avatar and the position of the listening avatar, and the volume of the voice. The information processing method according to any one of Appendices 11 to 14.

[0085] [Appendix 16] In the step of outputting the voice, the listening avatar is displayed in different display modes according to the output volume. The information processing method according to Appendix 15.

[0086] [Appendix 17] In the step of changing, the listening area is displayed in different display modes according to the output volume. The information processing method according to Appendix 15 or 16.

[0087] [Appendix 18] When receiving the selection of the user-specified mode, further receive the specification of the transmission range indicating the target to which the voice is transmitted and the volume of the voice output. In the step of outputting the voice, the voice is output at the specified volume, which is the specified volume, to the user of the avatar within the transmission range. The information processing method according to any one of Appendices 10 to 17.

[0088] [Appendix 19] The process of detecting the voice uttered by the user operating the avatar in the virtual space, In the relationship with the speaking avatar, which is the avatar operated by the user who uttered the voice, the process of outputting the voice to the user of the avatar satisfying a predetermined condition, A program that causes a computer to execute the process of changing the display mode of the listening avatar, which is the avatar satisfying the predetermined condition, is stored. A computer-readable storage medium.

[0089] [Appendix 20] In the process of outputting the voice, a listening area, which is an area including the speaking avatar, is set, an avatar existing in the listening area is defined as the listening avatar, and the voice is output to the user of the listening avatar. A computer-readable storage medium according to Supplementary Note 19.

[0090] [Supplementary Note 21] In the process of making the change, the listening area is displayed. A computer-readable storage medium according to Supplementary Note 20.

[0091] [Supplementary Note 22] In the process of outputting the voice, the listening area is set according to the volume of the voice uttered by the user. A computer-readable storage medium according to Supplementary Note 20 or 21.

[0092] [Supplementary Note 23] In the process of outputting the voice, the listening area is determined according to the face direction of the listening avatar. A computer-readable storage medium according to any one of Supplementary Notes 20 to 22.

[0093] [Supplementary Note 24] In the process of outputting the voice, the output volume, which is the volume output to the user of the listening avatar, is controlled according to the distance between the position of the speaking avatar and the position of the listening avatar and the volume of the voice. A computer-readable storage medium according to any one of Supplementary Notes 20 to 23.

[0094] [Supplementary Note 25] In the process of outputting the voice, the listening avatar is displayed in different display modes according to the output volume. A computer-readable storage medium according to Supplementary Note 24.

[0095] [Supplementary Note 26] In the process of the above-mentioned change process, the listening area is displayed in different display modes according to the output volume. A computer-readable storage medium according to appended note 24 or 25.

[0096] [Appended note 27] When receiving a selection of the user-specified mode, further receive a specification of a transmission range indicating a target to which the voice is transmitted and a volume at which the voice is output. In the process of outputting the voice, output the voice at a specified volume, which is the specified volume, to the user of the avatar in the transmission range. A computer-readable storage medium according to any one of appended notes 19 to 26.

Explanation of reference signs

[0097] 100, 101 Information processing apparatus 110 Detection unit 120, 123 Voice control unit 121, 124 Listening area setting unit 122, 125 Voice output unit 130, 131 Display control unit 140 Mode change unit 200 User terminal

Claims

1. An acquisition unit that acquires voice information of a first avatar operated in a virtual space; A setting unit that sets the size of a listening area indicating the range in which the voice information is transmitted according to the volume of the voice information; An output unit that outputs the voice information to a second avatar different from the first avatar existing within the listening area, and an information processing apparatus. Information processing apparatus.

2. The output unit outputs the voice information at a set volume in the listening area, and the information processing apparatus according to Claim 1. The information processing apparatus according to Claim 1.

3. When the volume of the voice information is greater than a predetermined threshold, the setting unit sets the size of the listening area to be larger than the reference size, and the information processing apparatus according to Claim 1 or 2. The information processing apparatus according to Claim 1 or 2.

4. When the volume of the voice information is less than a predetermined threshold, the setting unit sets the size of the listening area to be smaller than the reference size, and the information processing apparatus according to Claim 1 or 2. The information processing apparatus according to Claim 1 or 2.

5. The setting unit sets the size of the listening area according to which range of a plurality of preset volumes the volume of the voice information belongs to, and the information processing apparatus according to Claim 1 or 2. The information processing apparatus according to Claim 1 or 2.

6. The output unit increases the volume of the voice information output to the second avatar as the volume of the voice information of the first avatar increases, and the information processing apparatus according to Claim 1. The information processing apparatus according to Claim 1.

7. The output unit increases the volume of the voice information output to the second avatar as the distance from the first avatar decreases, and the information processing apparatus according to Claim 1. The information processing apparatus according to Claim 1.

8. Further comprising a display control unit that displays the second avatar in different display modes according to the volume, and the information processing apparatus according to any one of Claims 1 to 7. The information processing apparatus according to any one of Claims 1 to 7.

9. A computer, Acquires voice information of a first avatar operated in a virtual space, Sets the size of a listening area indicating the range in which the voice information is transmitted according to the volume of the voice information, Outputs the voice information to a second avatar different from the first avatar existing within the listening area, Information processing method.

10. A process of acquiring voice information of a first avatar operated in a virtual space, A process of setting the size of a listening area indicating the range in which the voice information is transmitted according to the volume of the voice information, A computer program that causes a computer to execute a process of outputting the voice information to a second avatar different from the first avatar that exists within the listening area.

Citation Information

Patent Citations

  • Server and voice signal collection / distribution method

    JP2008159034A

  • Information processing method, information processing program, information processing system, and information processing device

    JP2018163420A

  • Voice processing device, voice processing method, voice processing system, and terminal

    JP2022076189A

  • JP36871A

  • Information processing device, information processing method, and program

    WO2018020766A1