Information processing apparatus, information processing method, and program
Patent Information
- Application Number
- JP2023022712
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2026-01-27
AI Technical Summary
In virtual reality environments, overlapping voices from multiple participants can make it difficult for users to intuitively grasp the content of statements, especially when the manner of speaking and tone of voice are crucial for understanding, leading to mixed and indistinguishable audio that may be hard to hear.
A system determines priority sound data from overlapping audio streams directed to a user's avatar, playing the priority sound data in real time while converting or recording non-priority sound data for later playback or display, ensuring the user can grasp the content of priority sounds immediately and non-priority sounds separately.
Users can appropriately understand the content of priority sounds in real time and review non-priority sounds at a later time, enhancing clarity and comprehension in virtual reality communications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] Users can wear a head mounted display (HMD) and communicate with each other through their avatar in a virtual space using virtual reality (VR) technology. Voice chat is used for communication while wearing the HMD.
[0003] Patent Document 1 discloses a technology for converting audio data into character data in voice communication and displaying the converted character data in chronological order on an HMD. With this technology, when multiple participants make statements to a specific avatar at the same time, even if a participant corresponding to a specific avatar misses the contents of a certain participant's statement, the participant can understand the contents of the statement. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2018-186366 A Summary of the Invention [Problem to be solved by the invention]
[0005] On the other hand, there are cases where participants can intuitively understand the intention of a statement if the content of the statement is notified by a voice including elements such as the way of saying it and the tone of voice, rather than by a display of text. However, when a participant is participating in a presentation in a virtual space and is spoken to by a participant next to them, the voices of both the presenter and the participant next to them mix together, making it difficult for the participant to hear either of them clearly.
[0006] The present invention aims to provide a technique that allows a user controlling an avatar to properly understand the content of sounds when sounds from multiple sound sources are directed at the avatar in a virtual space. [Means for solving the problem]
[0007] One aspect of the present invention is a method for producing a composition comprising the steps of: A determination means for determining priority sound data from among a plurality of pieces of sound data that have overlapping sound generation timings and are directed toward an avatar of a first user in a virtual space; a first control means for controlling the first user to be notified of the contents of the priority sound data by playing the priority sound data at a first timing; a second control means for controlling the content of non-priority sound data that has not been determined as the priority sound data by the determination means to be notified to the first user without playing the non-priority sound data at the first timing; The information processing device is characterized by having:
[0008] One aspect of the present invention is a method for producing a composition comprising the steps of: a determining step of determining priority sound data from among a plurality of pieces of sound data whose sound generation timings overlap each other and are directed to an avatar of a first user in a virtual space; a first control step of controlling so as to notify the first user of the contents of the priority sound data by playing the priority sound data at a first timing; a second control step of controlling to notify the first user of the contents of non-priority sound data that has not been determined as the priority sound data in the determination step without playing the non-priority sound data at the first timing; The information processing method is characterized by having the following features. Effect of the Invention
[0009] According to the present invention, when sounds from multiple sound sources are directed toward an avatar in a virtual space, the user controlling the avatar can properly understand the content of the sounds. [Brief description of the drawings]
[0010] [Figure 1] FIG. 2 is a diagram illustrating a user terminal and a server. [Diagram 2] 11 is a flowchart illustrating acquisition of sound data. [Diagram 3] 13 is a flowchart illustrating a process of a server. [Figure 4] 11 is a flowchart illustrating an output of data. [Diagram 5] 11 is a flowchart illustrating the determination of a top-priority sound source. [Figure 6] FIG. 2 is a diagram illustrating the operation of a user terminal and a server. [Figure 7] FIG. 13 is a diagram illustrating terminal settings. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, an embodiment will be described with reference to the attached drawings. In addition, numerical values, processing timing, processing order, processing subject, and data (information) destination / source / storage location used in the embodiment described below are given as examples for the purpose of concrete explanation, and are not intended to be limited to such examples.
[0012] <Embodiment 1> 1A to 1C are block diagrams showing an example of the configuration of a user terminal 100 and a server 211 according to the first embodiment. Fig. 1A shows an example of the hardware configuration of the user terminal 100. The user terminal 100 has a CPU (Central Processing Unit) 101 and a ROM (Read-Only Memory) 102. The user terminal 100 has a RAM (Random Access Memory) 103 and an HDD (Hard Disk Drive) 104. These components are connected via a bus 105 so as to be able to communicate with each other.
[0013] The CPU 101 executes various processes using programs and data stored in the RAM 103 or the ROM 102. As a result, the CPU 101 controls the overall operation of the user terminal 100, and executes or controls various processes that will be described as processes performed by the user terminal 100.
[0014] The ROM 102 stores setting data, programs, data, etc. The RAM 103 has an area for storing computer programs or data (computer programs or data loaded from the ROM 102 or the HDD 104).
[0015] The RAM 103 also has a work area that is used when the CPU 101 executes various processes. The RAM 103 can provide various areas as needed.
[0016] The HDD 104 is an example of a device capable of storing a large amount of information. For example, an OS (Operating System) is stored in the HDD 104. The HDD 104 stores information (including computer programs) for causing the CPU 101 to execute and control various processes. When the computer programs or data stored in the HDD 104 are loaded into the RAM 103 under the control of the CPU 101, they become the subject of processing by the CPU 101.
[0017] Note that a medium (recording medium) and a drive device (a device for reading and writing computer programs and data from and to a medium) may be provided in addition to (or instead of) the HDD 104. Known examples of such media include a flexible disk (FD), a CD-ROM, a DVD, a Universal Serial Bus (USB) memory, a magneto-optical disk (MO), and a flash memory.
[0018] The hardware configuration of the user terminal 100 is not limited to the configuration shown in FIG. 1A, and can be appropriately modified or changed from the configuration shown in FIG. 1A.
[0019] In the first embodiment, the user terminal 100 is communicably connected to an external device including a plurality of components (HMD 106, microphone 107, speaker 108, and controller 109). However, the user terminal 100 may have some or all of the functions of the external device. For example, in FIG. 1A, the HMD 106 and the user terminal 100 are separate devices, but the HMD 106 and the user terminal 100 may be integrated to configure a single user terminal 100.
[0020] The HMD 106 displays a virtual space (VR space). The HMD 106 incorporates various sensors for detecting the inclination of the HMD 106 and the user's line of sight. The microphone 107 receives sound (such as the user's voice). The speaker 108 outputs sound. The HMD 106 and the speaker 108 can provide various notifications to the user, and therefore can be collectively referred to as a notification device.
[0021] The controller 109 receives input from a user. The input from the user is used to move an avatar in the virtual space or to control a UI (User Interface) displayed in the virtual space.
[0022] 1B shows an example of a hardware configuration of the server 211. The server 211 is an information processing device for controlling the operation of the user terminal 100. The server 211 has a CPU 201, a ROM 202, a RAM 203, and a HDD 204, similar to the user terminal 100. These components are connected to each other via a bus 205 so as to be able to communicate with each other. Furthermore, the HMD 106, the microphone 107, the speaker 108, the controller 109, and the like are not connected to the server 211.
[0023] 1C shows an example of the functional configuration of the user terminal 100 and the server 211. The user terminal 100 includes an input / output unit 112 and a transmission / reception unit 113. The server 211 includes a data control unit 213, a transmission / reception unit 214, a determination unit 215, a control unit 218, and a notification unit 219. The data control unit 213 includes a conversion unit 216 and a recording unit 217.
[0024] Each functional unit of the user terminal 100 and the server 211 shown in Fig. 1C may be implemented by hardware or software (computer program). When each functional unit is implemented by software, the above-mentioned user terminal 100 and the server 211 can be used as a computer device capable of executing a computer program.
[0025] The input / output unit 112 receives data from components (HMD 106, microphone 107, speaker 108, and controller 109) connected to the user terminal 100 and outputs the data to an external device. Outputs the data.
[0026] The transmitting / receiving unit 113 transmits data held by the user terminal 100 to another user terminal 100 via the network 110. In addition, the transmitting / receiving unit 113 receives data transmitted from another user terminal 100 via the network 110.
[0027] The transmitting / receiving unit 214 transmits data held by the server 211 to the user terminal 100 via the network 110. In addition, the transmitting / receiving unit 214 receives data transmitted from the user terminal 100 via the network 110.
[0028] The determination unit 215 determines a "top priority sound source" from among the sound sources (avatars) of multiple sound data (multiple sound data addressed to an avatar) directed substantially simultaneously to an avatar linked to the user terminal 100. Here, the "top priority sound source" is a sound source that is given priority over other sounds (voices) to be transmitted to the user. The multiple sound data have mutually overlapping sound (voice) generation timings (generation periods).
[0029] The conversion unit 216 converts the sound data of a sound source that is determined by the determination unit 215 not to be the "top priority sound source" into character data.
[0030] The recording unit 217 records the sound data of a sound source that is determined by the determination unit 215 not to be the "top priority sound source."
[0031] The control unit 218 controls the user terminal 100 to display in the virtual space the character data converted by the conversion unit 216. Alternatively, the control unit 218 controls the user terminal 100 to start playing the sound data recorded by the recording unit 217.
[0032] The notification unit 219 notifies the user terminal 100 that has output the sound data that the sound data has been converted or recorded by the data control unit 213.
[0033] In the first embodiment, n user terminals 100, from a first user terminal 100-1 used by a first user to an n-th user terminal 100-n used by an n-th user (n>2), are connected to the network 110. Each component of the n-th user terminal 100-n and the component connected to the input / output unit 112-1 of the n-th user terminal 100-n will be described with "-n" added to the end. For example, the input / output unit 112 of the n-th user terminal 100-n will be called "input / output unit 112-n", and the transmission / reception unit 113 of the n-th user terminal 100-n will be called "transmission / reception unit 113-n". For example, the HMD 106 connected to the input / output unit 112-n will be called "HMD 106-n", and the speaker 108 connected to the input / output unit 112-n will be called "speaker 108-n".
[0034] 2A to 4B and 5, a process will be described below in which a first user of a first user terminal 100-1 and a second user of a second user terminal 100-2 talk to a third avatar linked to a third user terminal 100-3. In this case, the first avatar and the second avatar talk to the third avatar in the virtual space, so that the first avatar and the second avatar are the sound sources of the sound data.
[0035] First, the process up to when two pieces of sound data are transmitted to server 211 will be described with reference to FIGS. 2A and 2B.
[0036] (Processing of the first information processing device) The flowchart in FIG. 2A shows the process of the first user terminal 100-1. When -1 accepts a voice, the process of step S1001 starts.
[0037] In step S1001, input / output unit 112-1 acquires, as first sound data, data of a voice uttered by a first user toward a third avatar from microphone 107-1.
[0038] In step S1002, the transmitting / receiving unit 113-1 transmits (provides) the first sound data to the server 211.
[0039] (Processing of the second information processing device) 2B shows the process of second user terminal 100-2. When microphone 107-2 receives a voice, the process of step S1011 starts.
[0040] In step S1011, input / output unit 112-2 acquires data of a voice uttered by the second user toward the third avatar as second sound data from microphone 107-2. Note that the first sound data and the second sound data are generated at the same time.
[0041] In step S1012, the transmitting / receiving unit 113-2 transmits (provides) the second sound data to the server 211.
[0042] (Server processing) Next, the processing of the server 211 will be described with reference to the flowchart of Fig. 3. When multiple pieces of sound data are transmitted to the server 211, the processing of step S2005 starts. When only one piece of sound data is transmitted to the server 211, the server 211 transmits the one piece of sound data to the third user terminal 100-3 in real time without performing the processing of the flowchart. Alternatively, when only one piece of sound data is transmitted to the server 211, the server 211 may treat the one piece of sound data as sound data of the "highest priority sound source" and execute the processing of the flowchart of Fig. 3.
[0043] In step S2005, the transmitting / receiving unit 214 receives the first sound data and the second sound data.
[0044] In step S2006, the determination unit 215 determines a "top priority sound source" from among the sound sources (the first avatar and the second avatar) of the multiple sound data (the first sound data and the second sound data) directed to the third avatar linked to the third user terminal 100-3. Details of the process in step S2006 will be described later with reference to the flowchart in FIG. 5.
[0045] Hereinafter, the processes of steps S2007 and S2008 are performed for each piece of sound data directed to the third avatar (i.e., for each of the first sound data and the second sound data). The sound data that is the target of the processes of steps S2007 and S2008 is referred to as "target sound data."
[0046] In step S2007, the data control unit 213 judges whether the target sound data is sound data of a "highest priority sound source" (hereinafter referred to as "highest priority sound data"). If it is judged that the target sound data is "highest priority sound data", the process proceeds to step S2009 without performing the process of step S2008. If it is judged that the target sound data is sound data of a "non-highest priority sound source" (hereinafter referred to as "non-priority sound data"), the process proceeds to step S2008.
[0047] In step S2008, the data control unit 213 determines whether the control mode for the third avatar is set to the conversion mode or the recording mode. If it is determined that the control mode is set to the recording mode, the conversion unit 216 converts the target sound data into character data. If it is determined that the control mode is set to the recording mode, the recording unit 217 starts recording the target sound data. The control mode may be arbitrarily set by the third user. The conversion or recording of the target sound data by the conversion unit 206 continues as long as the target sound data continues to be transmitted to the server 211.
[0048] In step S2009, data control unit 213 determines whether or not all sound data (all sound data directed to the third avatar) has been processed by steps S2007 and S2008. If it is determined that all sound data has been processed, the process proceeds to step S2010. If it is determined that all sound data has not been processed, the process proceeds to step S2007, where the processes of steps S2007 and S2008 are performed on the unprocessed sound data.
[0049] In step S2010, control unit 218 transmits the "highest priority sound data" to third user terminal 100-3 via transmission / reception unit 214. Furthermore, if one or more pieces of sound data have been converted into character data in step S2008, control unit 218 transmits one or more pieces of character data to third user terminal 100-3.
[0050] In step S2011, the notification unit 219 transmits an exception notification indicating that "sound data has been converted or recorded" to the user terminal 100 associated with the sound source of the data processed in step S2008 via the transmission / reception unit 214. In the following, it is assumed that the second sound data has been converted or recorded. That is, an exception notification indicating that "sound data has been converted or recorded" is transmitted to the second user terminal 100-2.
[0051] In step S2012, the data control unit 213 determines whether a certain amount of time has passed since the time (utterance end time) when the utterance (talking: sound generation) corresponding to the "highest priority sound data" directed to the third avatar ended. If it is determined that the certain amount of time has passed since the utterance end time, the process proceeds to step S2013. If it is determined that the certain amount of time has not passed since the utterance end time, the process of step S2012 is repeated.
[0052] In step S2013, the control unit 218 determines whether the control mode is set to the conversion mode or the recording mode. If it is determined that the control mode is set to the conversion mode, the process proceeds to step S2014. If it is determined that the control mode is set to the recording mode, the process proceeds to step S2015.
[0053] In step S2014, conversion unit 216 stops converting the sound data into character data.
[0054] In step S2015, the control unit 218 starts transmitting the sound data recorded in the recording unit 217 to the third user terminal 100-3 via the transmission / reception unit 214. As a result, the control unit 218 controls the third user terminal 100-3 to start playing the sound data (non-priority sound data) recorded in the recording unit 217.
[0055] In step S2016, the transmitting / receiving unit 214 (control unit 218) continues to transmit the recorded sound data to the third user terminal 100-3 until the reproduction of the sound data in the third user terminal 100-3 is completed.
[0056] In step S2017, the recording unit 217 stops recording the sound data.
[0057] Referring to the flowchart of FIG. 4A, the third user terminal 100-3 receives audio data. The process of notifying the third user of the utterance content shown in FIG. 11 will be described below. When data (audio data or character data) is transmitted to third user terminal 100-3, the process of step S2101 starts.
[0058] In step S2101, transmission / reception section 113-3 receives data transmitted from server 211.
[0059] In step S2102, the input / output unit 112-3 outputs the sound data, of the data received in step S2101, to the speaker 108-3, and outputs the text data to the HMD 106-3. This causes the sound data to be reproduced from the speaker 108-3. Text indicated by the text data is displayed on the HMD 106-3. That is, the content indicated by the sound data is notified to the third user by voice, and the content indicated by the text data is notified to the third user by displaying text.
[0060] Therefore, the process of this flowchart starts at the timing when data is transmitted to the third user terminal 100-3 in step S2010 and at the timing when data is transmitted to the third user terminal 100-3 in step S2016. Therefore, the third user is notified of the utterance content of the data transmitted in step S2010 in real time (without delay from the timing of the sound directed to the third avatar). On the other hand, the third user is notified of the utterance content of the data transmitted in step S2016 with a delay from real time (after the playback of the highest priority sound data has ended).
[0061] 4B, a process (processing of second user terminal 100-2) for notifying the second user that "audio data of the second user's utterance has been recorded or converted" will be described. When an exception notification is transmitted from server 211 to second user terminal 100-2, the process of step S2111 starts.
[0062] In step S2111, transmission / reception section 113-2 receives the exception notification sent from server 211 in step S2011.
[0063] In step S2112, the input / output unit 112-2 outputs the exception notification received in step S2111 to the HMD 106-2. As a result, the HMD 106-2 performs display according to the notification. The input / output unit 112-2 may output the exception notification received in step S2111 to the speaker 108-2. As a result, the speaker 108-2 plays sound data according to the notification.
[0064] For example, when the HMD 106-2 receives an exception notification indicating that sound data has been converted, it displays text indicating that the sound data has been converted to text data (notified by text). This allows the second user to understand that the content of his / her voice has been notified to the third user by text, not by voice.
[0065] Furthermore, for example, when the HMD 106-2 receives an exception notification indicating that sound data has been recorded, it displays characters or the like indicating that the sound data has been recorded (notified with a delay). This allows the second user to understand that the content of his / her voice has been notified to the third user by voice with a delay from real time, not in real time.
[0066] When the HMD 106-2 receives the exception notification, it may notify the second user that "the sound data was not notified to the third user by voice in real time" regardless of the type of the exception notification.
[0067] (Regarding step S2006) Next, the determination process performed in step S2006 will be described with reference to the flowchart in Fig. 5. Here, each user terminal 100 can have a terminal setting for determining a "top priority sound source" from among a plurality of sound sources. The terminal setting is, for example, one of a time priority setting, a relationship priority setting, a volume priority setting, and a distance priority setting, as follows.
[0068] Fig. 7A shows a setting screen displayed on the user terminal 100. A drop-down list 501 shows a method for determining the "highest priority sound source" that is referred to in the determination process performed in step S2006. In the example of Fig. 7A, the current terminal setting is a time-priority setting.
[0069] Fig. 7B shows a state in which the drop-down list 501 is opened by a user operation in the state of Fig. 7A. The user can determine any one of the settings of time priority, relationship priority, volume priority, and distance priority as the terminal setting. The user can also set the "top priority sound source" to be determined only by the method selected by the user himself (only manually).
[0070] In step S3001, the determination unit 215 determines whether or not a "top priority sound source" has been selected by the third user (manually) in the third user terminal 100-3. If it is determined that a "top priority sound source" has been selected by the third user, the process proceeds to step S3002. If it is determined that a "top priority sound source" has not been selected by the third user, the process proceeds to step S3003.
[0071] In step S3002, the determination unit 215 determines the sound source selected in step S3001 as the "top priority sound source."
[0072] In step S3003, the decision unit 215 judges whether the setting (terminal setting) of the third user terminal 100-3 is a time-priority setting. If it is judged that the terminal setting is a time-priority setting, the process proceeds to step S3004. If it is judged that the terminal setting is not a time-priority setting, the process proceeds to step S3005.
[0073] In step S3004, the determination unit 215 determines the "highest priority sound source" based on the order of utterances (sound generation) corresponding to the sound data. Specifically, the determination unit 215 determines the sound source of the sound data that started to be emitted at the earliest time among the sound data directed to the third avatar as the "highest priority sound source". For example, in a state in which a first person is talking to the third avatar and a second person starts to talk to the third avatar, the determination unit 215 determines the avatar of the first person as the "highest priority sound source". Alternatively, the determination unit 215 may determine the sound source of the sound data that started to be emitted at the latest time among the sound data directed to the third avatar as the "highest priority sound source".
[0074] In step S3005, the decision unit 215 judges whether the setting (terminal setting) of the third user terminal 100-3 is a relationship-prioritized setting. If it is judged that the terminal setting is a relationship-prioritized setting, the process proceeds to step S3006. If it is judged that the terminal setting is not a relationship-prioritized setting, the process proceeds to step S3007.
[0075] In step S3006, the determination unit 215 determines a "top priority sound source" from among the sound sources of the multiple sound data directed to the third avatar, based on the relationship between the sound sources of the multiple sound data and the third avatar. For example, if a lecturer's avatar is speaking at a lecture and is spoken to by an avatar of a student next to him, there is a high possibility that the words of the lecturer's avatar are more important to the third avatar than the words of the avatar of the student next to him during the lecture. For this reason, the determination unit 215 determines the lecturer's avatar as the "top priority sound source."
[0076] In step S3007, the decision unit 215 determines whether the setting (terminal setting) of the third user terminal 100-3 is a volume priority setting. If it is determined that the terminal setting is a volume priority setting, the process proceeds to step S3008. If it is determined that the terminal setting is not a volume priority setting, the process proceeds to step S3009.
[0077] In step S3008, the determination unit 215 determines a "highest priority sound source" from among the sound sources of the plurality of sound data based on the volume of the plurality of sound data directed to the third avatar. For example, the determination unit 215 determines the sound source of the sound data with the highest volume as the "highest priority sound source" from among the sound sources of the plurality of sound data. For example, a loud voice is often used when conveying emergency content. Therefore, by determining the sound source of the sound data with the highest volume as the "highest priority sound source", it becomes possible to convey the voice including emergency content to the third user with priority.
[0078] In step S3009, the decision unit 215 determines whether the setting (terminal setting) of the third user terminal 100-3 is a distance-priority setting. If it is determined that the terminal setting is a distance-priority setting, the process proceeds to step S3010. If it is determined that the terminal setting is not a distance-priority setting, the process proceeds to step S3011.
[0079] In step S3010, the determination unit 215 determines a "top priority sound source" from among the sound sources of the multiple sound data directed to the third avatar based on the distance between the sound source of the sound data directed to the third avatar and the third avatar. The determination unit 215, for example, determines the sound source that is the furthest from the multiple sound data sources to the third avatar as the "top priority sound source". For example, when conveying urgent content, a voice may be generated to convey the urgent content even from a position far from the third avatar. Therefore, by determining the sound source that is the furthest from the third avatar as the "top priority sound source", it becomes possible to convey the voice including the urgent content to the third user with priority. Alternatively, the determination unit 215 may determine that the avatar (sound source) that is closest to the third avatar is a person who is closer to the third avatar, and determine the sound source that is closest to the third avatar as the "top priority sound source".
[0080] In step S3011, the determination unit 215 determines that there is no "top priority sound source." In step S3011, the determination unit 215 may calculate the priority of each of the multiple sound sources based on at least two of, for example, the order of remarks, the relationship with the third avatar, the distance from the third avatar, and the volume of the multiple sound data. Then, the determination unit 215 may determine the sound source with the highest priority as the "top priority sound source."
[0081] Next, the operations of the user terminal 100 and the server 211 will be described with reference to FIGS. 6A to 6C.
[0082] FIG. 6A represents a virtual space in which a presentation is taking place. First avatar 401 is an avatar controlled by a first user of first user terminal 100-1. Similarly, second avatar 402 is an avatar controlled by a second user of second user terminal 100-2. Third avatar 403 is an avatar controlled by a third user of third user terminal 100-3. Fourth avatar 404 is an avatar controlled by a fourth user of fourth user terminal 100-4. Fifth avatar 405 is an avatar controlled by a fifth user of fifth user terminal 100-5.
[0083] The letters on the torso of each avatar represent the name of the avatar. The first avatar 401 is the presenter of the presentation. The first avatar 401 is displayed in the virtual space. The first avatar 401 is explaining presentation material 406. The second avatar 402 to the fifth avatar 405 are participants in the presentation and are listening to the explanation by the first avatar 401. The second avatar 402 faces the direction in which the third avatar 403 is located, and is about to talk to the third avatar 403.
[0084] Fig. 6B shows an image (image of the virtual space) displayed on HMD 106-3 of third user terminal 100-3 in the state of Fig. 6A. In the state shown in Fig. 6B, only one of the first users is speaking to third avatar 403, and therefore the audio of the first user's speech is being heard from speaker 108-3.
[0085] FIG. 6C shows an image displayed on HMD 106-3 when second avatar 402 speaks to third avatar 403 from the state of FIG. 6A. That is, the state shown in FIG. 6C is a state in which the second user, together with the first user, is speaking to third avatar 403. Therefore, the voice of the first avatar 401 (first user), which is the "highest priority sound source", is played from speaker 108-3. Meanwhile, the content of the statement of second avatar 402 to third avatar 403, which is the "lowest priority sound source", is displayed in dashed line area 407. The content of the statement of second avatar 402 is displayed as text on HMD 106-3.
[0086] The flow of processing in a specific example when the control mode is the change mode will be described below.
[0087] (1) A first user speaks, "What is XR?" into microphone 107-1. Then, input / output unit 112-1 acquires data of the voice saying, "What is XR?" as first sound data (step S1001). After that, transmission / reception unit 113-1 transmits the first sound data to server 211 (step S1002).
[0088] (2) Meanwhile, a second user speaks into microphone 107-2, saying, "Mr. C, what is MR?". Then, input / output unit 112-1 acquires the voice data of "Mr. C, what is MR?" as second sound data (step S1011). Thereafter, transmission / reception unit 113-2 transmits the second sound data to server 211 (step S1012). It is assumed that the processes of (1) and (2) are performed substantially simultaneously.
[0089] (3) The transmitting / receiving unit 214 receives the first sound data and the second sound data (step S2005).
[0090] (4) The determination unit 215 determines a "highest priority sound source" from among the sound source of the first sound data (first avatar 401) and the sound source of the second sound data (second avatar 402) (step S2006).
[0091] Specifically, first, the determination unit 215 determines whether or not the third user has selected the "highest priority sound source" (step S3001). Here, it is assumed that the third user has not selected the "highest priority sound source" and the terminal setting of the third user terminal 100-3 is the time priority setting (the state of FIG. 7A). Then, after determining that the third user has not selected the "highest priority sound source" (step S3001 No), the determination unit 215 determines that the terminal setting is the time priority setting (step S3003 Yes).
[0092] After that, the determination unit 215 determines the sound source of the sound data that was first emitted by the third avatar 403 linked to the third user terminal 100-3 as the "top priority sound source" (step S3004). In this example, the first sound data of "What is XR?" and the second sound data of "Mr. C, what is MR?" Among the sound data, the first sound data is assumed to be the sound data that was first emitted to the third avatar 403. That is, the determination unit 215 determines the first avatar 401, which is the sound source of the first sound data, as the "highest priority sound source."
[0093] (5) The data control unit 213 determines whether the first sound data is the "highest priority sound data" (step S2007). Then, it is determined that the first sound data is the "highest priority sound data" (step S2007: Yes).
[0094] (6) Converter 216 determines whether or not all of the sound data (the first sound data and the second sound data) directed to third avatar 403 has been processed (step S2009).
[0095] (7) Since the second sound data has not been processed (step S2009 No), the conversion unit 216 determines whether the second sound data is the "highest priority sound data" (step S2007).
[0096] (8) Since the second sound data is not the "highest priority sound data" (step S2007 No), the conversion unit 216 converts the second sound data of the "non-highest priority sound source" (the sound data of "Mr. C, what is MR?") into character data (step S2008). The conversion of sound data into character data can be realized by using existing voice recognition technology.
[0097] (9) Converter 216 determines whether or not all sound data directed to third avatar 403 has been processed (step S2009). Here, it is determined that all sound data has been processed (step S2009: Yes).
[0098] (10) The control unit 218 transmits the first sound data of the "highest priority sound source" and the text data of the "non-highest priority sound source" (text data of "Mr. C, what is MR?") to the third user terminal 100-3 via the transmission / reception unit 214 (step S2010). As a result, the control unit 218 controls the playback of the first sound data of the "highest priority sound source" and the display of the text data of the "non-highest priority sound source".
[0099] (11) The transmitting / receiving unit 113-3 receives the sound data (highest priority sound data) and the character data transmitted from the server 211 (step S2101). Here, the transmitting / receiving unit 113-3 receives the first sound data and the character data "Mr. C, what is MR?"
[0100] (12) The input / output unit 112-3 outputs the received data to the speaker 108-3 or the HMD 106-3 (step S2102). Here, the input / output unit 112-3 outputs the first sound data to the speaker 108-3 and outputs the text data to the HMD 106-3. At this time, the image shown in FIG. 6B changes to the image shown in FIG. 6C in the HMD 106-3.
[0101] (13) The transmitter / receiver 214 transmits an exception notification indicating that the sound data is being converted to the second user terminal 100-2 associated with the second avatar 402 that is the sound source of "Mr. C, what is MR?" (step S2011).
[0102] (14) Transmitting / receiving unit 113-2 receives an exception notification transmitted from server 211 (step S2111). Here, transmitting / receiving unit 113-2 receives the exception notification indicating that the sound data of second avatar 402 is being converted into character data.
[0103] (15) The input / output unit 112-2 outputs the received exception notification to the speaker 108-2 or the HMD 106-2 (step S2112). An exception notification indicating that the audio data of the avatar 402 is being converted into character data is output to the HMD 106-2.
[0104] (16) The conversion unit 216 determines whether a certain amount of time has passed since the end time of the statement corresponding to the "highest priority sound data" (statement end time) (step S2012).
[0105] (17) After a certain time has elapsed from the speech end time (Yes in step S2012), the control unit 218 determines whether the control mode is the conversion mode (step S2013).
[0106] (18) Since the control mode is the conversion mode (Yes in step S2013), the conversion unit 216 stops converting the sound data (step S2014).
[0107] In this way, according to the above (1) to (18), the server 211 converts the sound data of the "non-highest priority sound source" (non-priority sound data) into character data. Then, the server 211 outputs the sound data of the "highest priority sound source" (highest priority sound data) as it is to the speaker 108-3 as sound data. On the other hand, the sound data of the "non-highest priority sound source" is output to the HMD 106-3 as character data. This allows the user to check the content of one person's speech in real time by sound (sound that makes it easy to understand the speaker's intention by the way of speaking and tone of voice, etc.) even in a situation where multiple people's speech timing overlaps in the virtual space. Then, the user can check the content of other people's speech by text.
[0108] In the above description, the control mode is the conversion mode. In the following, the control mode is the recording mode. The above steps (1) to (7) are the same as those described above, so the description will be omitted.
[0109] (8') The recording unit 217 starts recording the second sound data of the "non-highest priority sound source" (step S2008).
[0110] (9') Recording unit 217 determines whether or not all sound data directed to third avatar 403 has been processed (step S2009).
[0111] (10') Since all sound data has been processed (step S2009 Yes), the control unit 218 transmits the “highest priority sound data” (the first sound data of “What is XR”) to the third user terminal 100-3 via the transceiver unit 214 (step S2010).
[0112] (11') The transmitting / receiving unit 113-3 receives sound data transmitted from the server 211 (step S2101). Here, the first sound data of "What is XR?" is received.
[0113] (12') The input / output unit 112-3 outputs the received data to the speaker 108-3 (step S2102). That is, the input / output unit 112-3 outputs the first sound data of "What is XR?" to the speaker 108-3.
[0114] (13') The transmitter / receiver 214 transmits an exception notification indicating that the sound data has been recorded to the second user terminal 100-2 associated with the second avatar 402 that is the sound source of the processed second sound data (step S2011).
[0115] (14') The transmitting / receiving unit 113-2 receives the exception notification sent from the server 211 (step S2111).
[0116] (15') The input / output unit 112-2 transmits the exception notification received in step S2111 to the HMD1 In step S2112, an exception notification indicating that the sound data of second avatar 402 is being recorded is output to HMD 106-2.
[0117] (16') The recording unit 217 judges whether or not a certain period of time has elapsed since the time when the utterance corresponding to the "highest priority sound data" ended (utterance end time) (step S2012).
[0118] (17') After a certain time has elapsed from the speech end time (Yes in step S2012), the control unit 218 judges whether the control mode is set to the conversion mode (step S2013).
[0119] (18') The control mode is set to the recording mode (step S2013 No). Therefore, the control unit 218 transmits sound data recorded in the recording unit 217 to the third user terminal 100-3 so that the third user terminal 100-3 starts playing the sound data (step S2015). Here, the control unit 218 transmits the second sound data (the sound data of "Mr. C, what is MR?") to the third user terminal 100-3 so that the third user terminal 100-3 starts playing the second sound data.
[0120] (19') The transmitting / receiving unit 214 continues to transmit the recorded second sound data (the sound data of "Mr. C, what is MR?") to the third user terminal 100-3 until the reproduction of the second sound data is completed (step S2016).
[0121] (20') The transmitting / receiving unit 113-3 receives data transmitted from the server 211 (step S2101). Here, the transmitting / receiving unit 113-3 receives the recorded second sound data (sound data of "Mr. C, what is MR?").
[0122] (21') The input / output unit 112-3 outputs the received data to the speaker 108-3 (step S2102). Here, the input / output unit 112-3 outputs the recorded second sound data (the sound data of "Mr. C, what is MR?") to the speaker 108-3. As a result, the second sound data is played back after the playback of the first sound data is completed, so that the two sounds are not overlapped in the sound heard by the third user. Therefore, the third user can grasp the contents of both the first sound data and the second sound data through the sound.
[0123] (22') The recording unit 217 stops recording the sound data (step S2017).
[0124] In this way, according to the above (1) to (7) and (8') to (22'), the server 211 records the sound data of the "non-highest priority sound source". Then, the sound data of the "highest priority sound source" is played back as audio in real time, and the sound data of the "non-highest priority sound source" is played back by the speaker 108-3 at a timing after the playback of the sound data of the "highest priority sound source" has finished. This allows the user to check the content of one person's speech in real time by audio, and to check the content of the other people's speech separately, even in a situation where multiple people's speech timing overlaps in the virtual space.
[0125] In the above description, the conversion of sound data by conversion unit 216 is stopped when a certain time has elapsed from the end time of the speech. The conversion of sound data may be stopped by a manual instruction from the third user of third user terminal 100-3. Similarly, the start of playback of recorded data by control unit 218 and the stop of recording of sound data by recording unit 217 may also be stopped by a manual instruction from the third user of third user terminal 100-3.
[0126] Regarding the playback of sound data (recorded sound data) by the control unit 218, a UI for playing the sound data may be displayed on the HMD 106-3 so that the third user can control the behavior of the speaker 108-3 during the playback of the sound data.
[0127] In the above description, it has been explained that the recording unit 217 records the sound data, but the recording method is not limited to audio recording of the sound data. Specifically, the recording unit 217 may record (video record) an image in a virtual space that includes the sound data. In addition, after the conversion unit 216 converts the sound data into character data, the recording unit 217 may record the character data.
[0128] In the above description, the avatar associated with the user terminal 100 is described as the sound source, but the sound source is not limited to this. The sound source may be a virtual object in the virtual space that is not associated with the user terminal 100 (for example, a virtual speaker that emits sound).
[0129] In the above description, a client-server system in which the server 211 exists has been described, but a peer-to-peer system may also be used. In this case, the server 211 does not exist, and the user terminal 100 realizes the functional configuration and processing of the server 211 instead of the server 211.
[0130] Although the present invention has been described in detail based on the preferred embodiments, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate.
[0131] Also, in the above, "If A is equal to or greater than B, proceed to step S1, and if A is smaller (lower) than B, proceed to step S2" may be read as "If A is greater (higher) than B, proceed to step S1, and if A is equal to or less than B, proceed to step S2." Conversely, "If A is greater (higher) than B, proceed to step S1, and if A is equal to or less than B, proceed to step S2" may be read as "If A is greater (higher) than B, proceed to step S1, and if A is smaller (lower) than B, proceed to step S2." Therefore, unless a contradiction occurs, "equal to or greater than A" may be read as "equal to or greater than A (high; long; many)," and "equal to or less than A" may be read as "equal to or less than A (low; short; few)." And, "equal to or greater than A" may be read as "equal to or greater than A," and "equal to or less than A" may be read as "equal to or less than A."
[0132] Each functional unit in each of the above embodiments (variations) may or may not be individual hardware. The functions of two or more functional units may be realized by common hardware. Each of a plurality of functions of one functional unit may be realized by individual hardware. Two or more functions of one functional unit may be realized by common hardware. Furthermore, each functional unit may or may not be realized by hardware such as an ASIC, FPGA, or DSP. For example, the device may have a processor and a memory (storage medium) in which a control program is stored. Then, the functions of at least some of the functional units of the device may be realized by the processor reading and executing the control program from the memory.
[0133] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions.
[0134] The disclosure of the above embodiments includes the following configurations, methods, and programs. (Configuration 1) A determination means for determining priority sound data from among a plurality of pieces of sound data that have overlapping sound generation timings and are directed toward an avatar of a first user in a virtual space; a first control means for controlling the first user to be notified of the contents of the priority sound data by playing the priority sound data at a first timing; a second control means for controlling the content of non-priority sound data that has not been determined as the priority sound data by the determination means to be notified to the first user without playing the non-priority sound data at the first timing; 13. An information processing device comprising: (Configuration 2) the second control means controls to notify the first user of the content of the non-priority sound data by displaying characters indicating the content of the non-priority sound data. 2. The information processing device according to configuration 1. (Configuration 3) the second control means controls the reproduction of the non-priority sound data at a second timing after the reproduction of the priority sound data is completed, thereby notifying the first user of the content of the non-priority sound data. 2. The information processing device according to configuration 1. (Configuration 4) a notification means for notifying the terminal that the non-priority sound data will not be reproduced at the first timing when a terminal of a second user provides the non-priority sound data to the information processing device; The terminal that has been notified that the non-priority sound data will not be reproduced at the first timing notifies the second user that the non-priority sound data will not be reproduced at the first timing. 4. The information processing device according to any one of configurations 1 to 3. (Configuration 5) The determining means determines, from among the plurality of sound data, the sound data which has been emitted to the avatar first as the priority sound data. 5. The information processing device according to any one of configurations 1 to 4. (Configuration 6) The determining means determines the priority sound data based on a relationship between each of the sound sources of the plurality of sound data and the avatar in the virtual space. 5. The information processing device according to any one of configurations 1 to 4. (Configuration 7) The determining means determines the priority sound data based on the volume of the plurality of sound data. 5. The information processing device according to any one of configurations 1 to 4. (Configuration 8) The determining means determines the priority sound data based on a distance between the avatar and each of the sound sources of the plurality of sound data in the virtual space. 5. The information processing device according to any one of configurations 1 to 4. (Configuration 9) The determining means determines, as the priority sound data, sound data of a sound source selected by the first user from among the sound sources of the plurality of sound data in the virtual space. 5. The information processing device according to any one of configurations 1 to 4. (method) a determining step of determining priority sound data from among a plurality of pieces of sound data whose sound generation timings overlap each other and are directed to an avatar of a first user in a virtual space; The priority sound data is reproduced at a first timing, thereby a first control step of controlling to notify the first user; a second control step of controlling to notify the first user of the contents of non-priority sound data that has not been determined as the priority sound data in the determination step without playing the non-priority sound data at the first timing; 13. An information processing method comprising: (program) A program for causing a computer to function as each of the means of the information processing device according to any one of configurations 1 to 9. [Explanation of symbols]
[0135] 100: user terminal, 211: server, 215: decision unit, 218: control unit
Claims
1. a determination means for determining priority sound data from among a plurality of pieces of sound data each having overlapping sound generation timings and directed toward an avatar of a first user in a virtual space; a first control means for controlling the reproduction of the priority sound data at a first timing to notify the first user of the contents of the priority sound data; a second control means for controlling the content of non-priority sound data that has not been determined as the priority sound data by the determination means to be notified to the first user without playing the non-priority sound data at the first timing; 13. An information processing device comprising:
2. the second control means controls to notify the first user of the content of the non-priority sound data by displaying characters indicating the content of the non-priority sound data.
2. The information processing apparatus according to claim 1,
3. the second control means controls the reproduction of the non-priority sound data at a second timing after the reproduction of the priority sound data is completed, thereby notifying the first user of the content of the non-priority sound data.
2. The information processing apparatus according to claim 1,
4. a notification means for notifying the terminal that the non-priority sound data will not be reproduced at the first timing when a terminal of a second user provides the non-priority sound data to the information processing device; The terminal that has been notified that the non-priority sound data will not be reproduced at the first timing notifies the second user that the non-priority sound data will not be reproduced at the first timing.
2. The information processing apparatus according to claim 1,
5. The determining means determines, from among the plurality of sound data, the sound data which has been emitted to the avatar first as the priority sound data.
5. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
6. The determining means determines the priority sound data based on a relationship between each of the sound sources of the plurality of sound data and the avatar in the virtual space.
5. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
7. The determining means determines the priority sound data based on the volume of the plurality of sound data.
5. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
8. The determining means determines the priority sound data based on a distance between the avatar and each of the sound sources of the plurality of sound data in the virtual space.
5. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
9. The determining means determines, as the priority sound data, sound data of a sound source selected by the first user from among the plurality of sound sources of the sound data in the virtual space.
5. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
10. determining priority sound data from among a plurality of sound data whose sound generation timings overlap each other and are directed to an avatar of a first user in a virtual space; a decision step of: a first control step of controlling so as to notify the first user of the contents of the priority sound data by playing the priority sound data at a first timing; a second control step of controlling to notify the first user of the contents of non-priority sound data that has not been determined as the priority sound data in the determination step, without playing the non-priority sound data at the first timing; 13. An information processing method comprising:
11. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 4.