Information processing apparatus, information processing terminal, information processing method, and program

By employing HRTF data for audio-visual localization in remote meetings, the system addresses the challenge of distinguishing between multiple speakers and identifying virtual action performers, resulting in improved clarity and presence in immersive communication.

JP7687339B2Active Publication Date: 2025-06-03SONY GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022547498
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-10
Filing Date
2021-08-27
Publication Date
2025-06-03
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

In remote meetings, participants cannot individually specify another participant to talk to and focus on specific voices, leading to difficulties in distinguishing between multiple speakers and identifying who is performing virtual actions.

Method used

An information processing apparatus and method that utilize HRTF (Head-Related Transfer Function) data to perform audio-visual localization processing, allowing voice content to be output in a realistic manner based on actions by conversation participants, thereby localizing audio-visual content at specific positions.

Benefits of technology

Enables participants to clearly distinguish and focus on specific voices and actions, enhancing the sense of presence and clarity in immersive communication environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687339000001
    Figure 0007687339000001
  • Figure 0007687339000002
    Figure 0007687339000002
  • Figure 0007687339000003
    Figure 0007687339000003
Patent Text Reader

Abstract

The present technology relates to an information processing device, an information processing terminal, an information processing method, and a program with which it is possible for speech content corresponding to actions by participants in a conversation to be outputted in a state of realistic sensations. The information processing device according to one aspect of the present technology comprises: a storage unit for storing HRTF data that corresponds to a plurality of positions based on a listening position; and a sound image localization unit for performing a sound image localization process in which is used HRTF data selected in accordance with an action by a specific participant among the participants in the conversation who are participating via a network, thereby providing speech content selected in accordance with the action so that a sound image is localized at a prescribed position. The present technology can be applied to computers that perform remote conferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology particularly relates to an information processing apparatus, an information processing terminal, an information processing method, and a program that can output voice content according to actions by participants in a conversation in an immersive state.

Background Art

[0002] A so-called remote meeting in which a plurality of participants at a distance use devices such as a PC to hold a meeting has become widespread. By starting a web browser or a dedicated application installed on the PC and accessing an access destination specified by a URL assigned for each meeting, a user who knows the URL can participate in the meeting as a participant.

[0003] The voice of the participant collected by the microphone is transmitted to the devices used by other participants via a server and output from headphones or speakers. Also, the video showing the participant photographed by the camera is transmitted to the devices used by other participants via a server and displayed on the display of the device.

[0004] As a result, each participant can talk while seeing the faces of other participants.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Since one's own speech is shared with all other participants, a participant cannot, for example, individually specify a particular participant and talk only with the specified participant.

[0007] On the contrary, participants cannot focus only on the speech of a specific participant and listen to the content of the speech.

[0008] When a virtual action function such as a raising hand function is used, it may be visually presented on the screen display that a specific participant is performing an action, but it is difficult to tell which participant is performing the action.

[0009] The present technology has been made in view of such a situation, and enables voice content corresponding to an action by a conversation participant to be output in a realistic state.

Means for Solving the Problem

[0010] An information processing apparatus according to one aspect of the present technology includes a storage unit that stores HRTF data corresponding to a plurality of positions based on a listening position, and the HRTF data selected according to an action by a specific participant among the participants of a conversation participating via a network. By performing audio-visual localization processing using the data, an audio-visual localization processing unit that provides voice content selected according to the action so that the audio-visual is localized at a predetermined position.

[0011] An information processing terminal according to another aspect of the present technology stores HRTF data corresponding to a plurality of positions based on a listening position, and is selected according to an action by a specific participant among the participants of a conversation participating via a network. The voice content obtained by performing the audio-visual localization processing is received from an information processing apparatus that performs the audio-visual localization processing to provide the voice content selected according to the action so that the audio-visual is localized at a predetermined position, and a voice receiving unit that outputs the voice is provided.

[0012] In one aspect of the present technology, HRTF data corresponding to a plurality of positions based on a listening position is stored, and audio-visual localization processing using the HRTF data selected according to an action by a specific participant among the participants in a conversation participating via a network is performed, so that the audio-visual content is provided according to the action so that the audio-visual is localized at a predetermined position.

[0013] In another aspect of the present technology, HRTF data corresponding to a plurality of positions based on a listening position is stored, and audio-visual localization processing using the HRTF data selected according to an action by a specific participant among the participants in a conversation participating via a network is performed, so that the audio-visual content selected according to the action is transmitted from an information processing device that provides the audio-visual content so that the audio-visual is localized at a predetermined position, and the audio-visual content obtained by performing the audio-visual localization processing is received and the audio is output.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments for carrying out the present technology will be described. The description will be made in the following order. 1. Configuration of the Tele-communication System 2. Basic Operations 3. Configuration of Each Device 4. Use Cases of Audio-Visual Localization 5. Variations

[0016] <<Configuration of the Tele-communication System>> FIG. 1 is a diagram showing a configuration example of a Tele-communication system according to an embodiment of the present technology.

[0017] The Tele-communication system in FIG. 1 is configured by connecting a plurality of client terminals used by conference participants to a communication management server 1 via a network 11 such as the Internet. In the example of FIG. 1, client terminals 2A to 2D, which are PCs, are shown as client terminals used by users A to D who are conference participants.

[0018] Other devices such as smartphones and tablet terminals having a voice input device such as a microphone (mic) and a voice output device such as headphones and speakers may be used as client terminals. When there is no need to distinguish between client terminals 2A to 2D, they are appropriately referred to as client terminal 2.

[0019] Users A to D are users who participate in the same conference. Note that the number of users participating in the conference is not limited to four.

[0020] The communication management server 1 manages a meeting that can be advanced by multiple users having a conversation online. The communication management server 1 is an information processing device that controls the transmission and reception of voice between client terminals 2 and manages a so-called remote meeting.

[0021] For example, as shown by arrow A1 in the upper part of FIG. 2, the communication management server 1 receives the voice data of user A transmitted from client terminal 2A in response to user A's speech. From client terminal 2A, the voice data of user A collected by a microphone provided in client terminal 2A is transmitted.

[0022] The communication management server 1 transmits the voice data of user A to each of client terminals 2B to 2D as shown by arrows A11 to A13 in the lower part of FIG. 2, and outputs the voice of user A. When user A speaks as a speaker, users B to D become listeners. Hereinafter, the user who becomes the speaker will be appropriately referred to as the speaking user, and the user who becomes the listener will be referred to as the listening user.

[0023] Similarly, when another user speaks, the voice data transmitted from the client terminal 2 used by the speaking user is transmitted to the client terminal 2 used by the listening user via the communication management server 1.

[0024] The communication management server 1 manages the positions of each user in the virtual space. The virtual space is, for example, a three-dimensional space virtually set as the place where the meeting is held. The position in the virtual space is represented by three-dimensional coordinates.

[0025] FIG. 3 is a plan view showing an example of the positions of users in the virtual space.

[0026] In the example of FIG. 3, a vertically long rectangular table T is arranged approximately at the center of the virtual space indicated by the rectangular frame F, and positions P1 to P4, which are positions around the table T, are set as the positions of users A to D, respectively. The front direction of each user is the direction from the position of each user to the table T.

[0027] During the meeting, on the screen of the client terminal 2 used by each user, as shown in FIG. 4, participant icons, which are information visually representing the users, are displayed over the background image representing the place where the meeting is held. The position of the participant icon on the screen is a position corresponding to the position of each user in the virtual space.

[0028] In the example of FIG. 4, the participant icon is configured as a circular image including the face of the user. The participant icon is displayed in a size corresponding to the distance from the reference position set in the virtual space to the position of each user. Participant icons I1 to I4 represent users A to D, respectively.

[0029] For example, the position of each user is automatically set by the communication management server 1 when participating in the meeting. The position in the virtual space may be set by the user himself / herself, such as by moving the participant icon on the screen of FIG. 4.

[0030] The communication management server 1 has HRTF (Head-Related Transfer Function) data, which is data representing the sound transmission characteristics from a plurality of positions to the listening position when each position in the virtual space is the listening position. HRTF data corresponding to a plurality of positions is prepared in the communication management server 1 based on each listening position in the virtual space.

[0031] For each listening user, the communication management server 1 performs audio localization processing using HRTF data on the voice data so that the voice of the speaking user can be heard from the position of the speaking user in the virtual space, and transmits the voice data obtained by performing the audio localization processing.

[0032] As described above, the voice data transmitted to the client terminal 2 is the voice data obtained by performing audio localization processing in the communication management server 1. The audio localization processing includes rendering such as VBAP (Vector Based Amplitude Panning) based on position information and binaural processing using HRTF data.

[0033] That is, the voice of each speaking user is processed in the communication management server 1 as object audio voice data. The channel-based audio data of, for example, 2 channels of L / R generated by the audio localization processing in the communication management server 1 is transmitted from the communication management server 1 to each client terminal 2, and the voice of the speaking user is output from headphones or the like provided in the client terminal 2.

[0034] By performing audio localization processing using HRTF data according to the relative positional relationship between the position of the listening user himself and the position of the speaking user, each listening user will feel that the voice of the speaking user can be heard from the position of the speaking user.

[0035] Figure 5 is a diagram showing an example of how a voice is heard.

[0036] When focusing on User A, whose position P1 is set as a position in the virtual space, as the listening user, the voice of User B is heard from the right side adjacent, as shown by the arrow in Fig. 5, through performing sound image localization processing based on the HRTF data between position P2 (with position P2 as the sound source position) and position P1. The front of User A, who is having a conversation facing the client terminal 2A, is in the direction of the client terminal 2A.

[0037] Also, the voice of User C is heard from the front through performing sound image localization processing based on the HRTF data between position P3 (with position P3 as the sound source position) and position P1. The voice of User D is heard from the far right through performing sound image localization processing based on the HRTF data between position P4 (with position P4 as the sound source position) and position P1.

[0038] The same applies when other users are the listening users. For example, as shown in Fig. 6, the voice of User A is heard from the left side adjacent for User B, who is having a conversation facing the client terminal 2B, and is heard from the front for User C, who is having a conversation facing the client terminal 2C. Also, the voice of User A is heard from the far right for User D, who is having a conversation facing the client terminal 2D.

[0039] In this way, in the communication management server 1, the voice data for each listening user is generated according to the positional relationship between the position of each listening user and the position of the speaking user, and is used for the output of the voice of the speaking user. The voice data transmitted to each listening user becomes voice data with different ways of being heard according to the positional relationship between the position of each listening user and the position of the speaking user.

[0040] Fig. 7 is a diagram showing the state of users participating in a meeting.

[0041] For example, user A, who is wearing headphones and participating in a meeting, will hear the voices of users B to D whose sound images are localized at the positions of the right neighbor, the front, and the back right respectively, and have a conversation. As described with reference to FIG. 5 and the like, based on the position of user A, the positions of users B to D are the positions of the right neighbor, the front, and the back right respectively. Note that in FIG. 7, the coloring of users B to D indicates that users B to D do not actually exist in the same space as the space where user A is having the meeting.

[0042] Note that, as will be described later, background sounds such as bird chirping and BGM are also output based on the voice data obtained by the sound image localization process so that the sound images are localized at predetermined positions.

[0043] The sounds that the communication management server 1 processes include not only the uttered voice but also sounds such as environmental sounds and background sounds. Hereinafter, when it is not necessary to distinguish the types of each sound as appropriate, the sounds that the communication management server 1 processes will be simply described as voices. In reality, the sounds that the communication management server 1 processes include sounds of types other than voices.

[0044] Since the voice of the speaking user can be heard from a position corresponding to the position in the virtual space, the listening user can easily distinguish the voices of each user even when there are a plurality of participants. For example, even when a plurality of users speak simultaneously, the listening user can distinguish each voice.

[0045] In addition, since the voice of the speaking user is felt three-dimensionally, the listening user can obtain from the voice the feeling that the speaking user actually exists at the position of the sound image. The listening user can have a conversation with other users with a sense of presence.

[0046] <<Basic Operations>> Here, the basic operation flow of the communication management server 1 and the client terminal 2 will be described.

[0047] <Operation of Communication Management Server 1> Referring to the flowchart of FIG. 8, the basic processing of the communication management server 1 will be described.

[0048] In step S1, the communication management server 1 determines whether voice data has been transmitted from the client terminal 2, and waits until it is determined that voice data has been transmitted.

[0049] If it is determined in step S1 that voice data has been transmitted from the client terminal 2, then in step S2, the communication management server 1 receives the voice data transmitted from the client terminal 2.

[0050] In step S3, the communication management server 1 performs sound image localization processing based on the location information of each user, and generates voice data for each listening user.

[0051] For example, the voice data for user A is generated such that the sound image of the voice of the speaking user is localized at a position corresponding to the position of the speaking user when the position of user A is used as a reference.

[0052] Also, the voice data for user B is generated such that the sound image of the voice of the speaking user is localized at a position corresponding to the position of the speaking user when the position of user B is used as a reference.

[0053] Similarly, for the voice data for other listening users, HRTF data corresponding to the relative positional relationship with the speaking user is used based on the position of the listening user to generate the data. The voice data for each listening user is different data.

[0054] In step S4, the communication management server 1 transmits voice data to each listening user. The above processing is performed every time voice data is transmitted from the client terminal 2 used by the speaking user.

[0055] <Operation of Client Terminal 2> Referring to the flowchart of FIG. 9, the basic processing of client terminal 2 will be described.

[0056] In step S11, client terminal 2 determines whether microphone audio has been input. The microphone audio is the audio collected by the microphone provided in client terminal 2.

[0057] If it is determined in step S11 that microphone audio has been input, then in step S12, client terminal 2 transmits the audio data to communication management server 1. If it is determined in step S11 that no microphone audio has been input, the process of step S12 is skipped.

[0058] In step S13, client terminal 2 determines whether audio data has been transmitted from communication management server 1.

[0059] If it is determined in step S13 that audio data has been transmitted, then in step S14, communication management server 1 receives the audio data and outputs the voice of the speaking user.

[0060] After the voice of the speaking user is output, or if it is determined in step S13 that no audio data has been transmitted, the process returns to step S11 and the above-described processing is repeated.

[0061] <<Configuration of Each Device>> <Configuration of Communication Management Server 1> FIG. 10 is a block diagram showing an example of the hardware configuration of communication management server 1.

[0062] The communication management server 1 is composed of a computer. The communication management server 1 may be composed of one computer having the configuration shown in FIG. 10, or may be composed of a plurality of computers.

[0063] The CPU 101, ROM 102, and RAM 103 are interconnected by a bus 104. The CPU 101 executes the server program 101A and controls the overall operation of the communication management server 1. The server program 101A is a program for realizing a Tele-communication system.

[0064] An input / output interface 105 is further connected to the bus 104. An input unit 106 composed of a keyboard, a mouse, etc. and an output unit 107 composed of a display, a speaker, etc. are connected to the input / output interface 105.

[0065] Also, a storage unit 108 composed of a hard disk, a non-volatile memory, etc., a communication unit 109 composed of a network interface, etc., and a drive 110 for driving a removable medium 111 are connected to the input / output interface 105. For example, the communication unit 109 communicates with the client terminals 2 used by the respective users via the network 11.

[0066] FIG. 11 is a block diagram showing a functional configuration example of the communication management server 1. At least a part of the functional units shown in FIG. 11 is realized by the server program 101A being executed by the CPU 101 in FIG. 10.

[0067] In the communication management server 1, an information processing unit 121 is realized. The information processing unit 121 is composed of an audio reception unit 131, a signal processing unit 132, a participant information management unit 133, an audio-visual localization processing unit 134, an HRTF data storage unit 135, a system audio management unit 136, a 2ch mix processing unit 137, and an audio transmission unit 138.

[0068] The voice receiving unit 131 controls the communication unit 109 and receives voice data transmitted from the client terminal 2 used by the speaking user. The voice data received by the voice receiving unit 131 is output to the signal processing unit 132.

[0069] The signal processing unit 132 appropriately performs predetermined signal processing on the voice data supplied from the voice receiving unit 131, and outputs the voice data obtained by performing the signal processing to the sound image localization processing unit 134. For example, the signal processing unit 132 performs processing to separate the voice of the speaking user from the ambient sound. The microphone voice includes, in addition to the voice of the speaking user, ambient sounds such as noise and noise in the space where the speaking user is located.

[0070] The participant information management unit 133 controls the communication unit 109 and manages participant information, which is information about the participants in the meeting, by communicating with the client terminal 2.

[0071] FIG. 12 is a diagram showing an example of participant information.

[0072] As shown in FIG. 12, the participant information includes user information, position information, setting information, and volume information.

[0073] The user information is information about the users who participate in the meeting set by a certain user. For example, the user ID and the like are included in the user information. Other information included in the participant information is managed in association with the user information, for example.

[0074] The position information is information representing the positions of the respective users in the virtual space.

[0075] The setting information is information representing the content of settings related to the meeting, such as the setting of background sound used in the meeting.

[0076] The volume information is information representing the volume when the voices of the respective users are output.

[0077] The participant information managed by the Participant Information Management Unit 133 is supplied to the Audio-Visual Localization Processing Unit 134. The participant information managed by the Participant Information Management Unit 133 is also supplied to the System Voice Management Unit 136, the 2ch Mixing Processing Unit 137, the Voice Transmission Unit 138, etc. as appropriate. In this way, the Participant Information Management Unit 133 functions as a position management unit that manages the positions of each user in the virtual space and also functions as a background sound management unit that manages the setting of background sounds.

[0078] Based on the position information supplied from the Participant Information Management Unit 133, the Audio-Visual Localization Processing Unit 134 reads and acquires HRTF data corresponding to the positional relationship of each user from the HRTF Data Storage Unit 135. The Audio-Visual Localization Processing Unit 134 performs audio-visual localization processing on the audio data supplied from the Signal Processing Unit 132 using the HRTF data read from the HRTF Data Storage Unit 135 to generate audio data for each listening user.

[0079] In addition, the Audio-Visual Localization Processing Unit 134 performs audio-visual localization processing on the data of the system voice supplied from the System Voice Management Unit 136 using predetermined HRTF data. The system voice is a voice that is generated on the Communication Management Server 1 side and is listened to by the listening user together with the voice of the speaking user. The system voice includes, for example, background sounds such as BGM and sound effects. The system voice is a voice different from the voice of the user.

[0080] That is, in the Communication Management Server 1, voices other than the voice of the speaking user, such as background sounds and sound effects, are also processed as object audio. Audio-visual localization processing for localizing the audio at a predetermined position in the virtual space is also performed on the audio data of the system voice. For example, audio-visual localization processing for localizing the audio at a position farther away from the position of the participant is performed on the audio data of the background sound.

[0081] The audio-visual localization processing unit 134 outputs the audio data obtained by performing audio-visual localization processing to the 2ch mixing processing unit 137. To the 2ch mixing processing unit 137, the audio data of the speaking user and, as appropriate, the audio data of the system voice are output.

[0082] The HRTF data storage unit 135 stores HRTF data corresponding to a plurality of positions based on each listening position in the virtual space.

[0083] The system voice management unit 136 manages the system voice. The system voice management unit 136 outputs the audio data of the system voice to the audio-visual localization processing unit 134.

[0084] The 2ch mixing processing unit 137 performs 2ch mixing processing on the audio data supplied from the audio-visual localization processing unit 134. By performing the 2ch mixing processing, channel-based audio data including the components of the audio signals L and R of the voice of the speaking user and the system voice respectively is generated. The audio data obtained by performing the 2ch mixing processing is output to the audio transmission unit 138.

[0085] The audio transmission unit 138 controls the communication unit 109 and transmits the audio data supplied from the 2ch mixing processing unit 137 to the client terminal 2 used by each listening user.

[0086] <Configuration of Client Terminal 2> FIG. 13 is a block diagram showing a hardware configuration example of the client terminal 2.

[0087] The client terminal 2 is configured by connecting a memory 202, an audio input device 203, an audio output device 204, an operation unit 205, a communication unit 206, a display 207, and a sensor unit 208 to a control unit 201.

[0088] The control unit 201 is composed of a CPU, ROM, RAM, etc. By executing the client program 201A, the control unit 201 controls the overall operation of the client terminal 2. The client program 201A is a program for using the Tele-communication system managed by the communication management server 1. The client program 201A includes a transmission-side module 201A-1 that executes transmission-side processing and a reception-side module 201A-2 that executes reception-side processing.

[0089] The memory 202 is composed of a flash memory or the like. The memory 202 stores various information such as the client program 201A executed by the control unit 201.

[0090] The voice input device 203 is composed of a microphone. The voice collected by the voice input device 203 is output to the control unit 201 as microphone voice.

[0091] The voice output device 204 is composed of devices such as headphones and speakers. The voice output device 204 outputs the voice of the conference participants or the like based on the audio signal supplied from the control unit 201.

[0092] Hereinafter, the voice input device 203 will be described as a microphone as appropriate. Also, the voice output device 204 will be described as headphones.

[0093] The operation unit 205 is composed of various buttons and a touch panel provided over the display 207. The operation unit 205 outputs information representing the content of the user's operation to the control unit 201.

[0094] The communication unit 206 is a communication module corresponding to wireless communication of a mobile communication system such as 5G communication and a communication module corresponding to wireless LAN or the like. The communication unit 206 receives radio waves output by a base station and communicates with various devices such as the communication management server 1 via the network 11. The communication unit 206 receives information transmitted from the communication management server 1 and outputs it to the control unit 201. Further, the communication unit 206 transmits the information supplied from the control unit 201 to the communication management server 1.

[0095] The display 207 is composed of an organic EL display, an LCD, or the like. Various screens such as a remote conference screen are displayed on the display 207.

[0096] The sensor unit 208 is composed of various sensors such as an RGB camera, a depth camera, a gyro sensor, and an acceleration sensor. The sensor unit 208 outputs sensor data obtained by performing measurement to the control unit 201. Based on the sensor data measured by the sensor unit 208, the situation of the user is appropriately recognized.

[0097] FIG. 14 is a block diagram showing a functional configuration example of the client terminal 2. At least a part of the functional units shown in FIG. 14 is realized by the client program 201A being executed by the control unit 201 of FIG. 13.

[0098] In the client terminal 2, an information processing unit 211 is realized. The information processing unit 211 is composed of an audio processing unit 221, a setting information transmission unit 222, a user situation recognition unit 223, and a display control unit 224.

[0099] The information processing unit 211 is composed of an audio reception unit 231, an output control unit 232, a microphone audio acquisition unit 233, and an audio transmission unit 234.

[0100] The voice receiving unit 231 controls the communication unit 206 and receives voice data transmitted from the communication management server 1. The voice data received by the voice receiving unit 231 is supplied to the output control unit 232.

[0101] The output control unit 232 causes the voice output device 204 to output voice corresponding to the voice data transmitted from the communication management server 1.

[0102] The microphone voice acquisition unit 233 acquires voice data of the microphone voice collected by the microphone constituting the voice input device 203. The voice data of the microphone voice acquired by the microphone voice acquisition unit 233 is supplied to the voice transmission unit 234.

[0103] The voice transmission unit 234 controls the communication unit 206 and transmits the voice data of the microphone voice supplied from the microphone voice acquisition unit 233 to the communication management server 1.

[0104] The setting information transmission unit 222 generates setting information representing the contents of various settings according to the user's operation. The setting information transmission unit 222 controls the communication unit 206 and transmits the setting information to the communication management server 1.

[0105] The user status recognition unit 223 recognizes the user's status based on the sensor data measured by the sensor unit 208. The user status recognition unit 223 controls the communication unit 206 and transmits information representing the user's status to the communication management server 1.

[0106] The display control unit 224 communicates with the communication management server 1 by controlling the communication unit 206, and causes the display 207 to display a remote conference screen based on the information transmitted from the communication management server 1.

[0107] <<Use case of audio-visual localization>> The use case of audio-visual localization of various voices including the speech voices of the conference participants will be described.

[0108] <Virtual Reaction Function> The virtual reaction function is a function used when communicating one's reaction to other users. In a remote conference realized by the communication management server 1, for example, a clapping function, which is a virtual reaction function, is provided. Outputting the sound effect of clapping using the clapping function is instructed from the screen displayed as a GUI on the display 207 of the client terminal 2.

[0109] FIG. 15 is a diagram showing an example of a remote conference screen.

[0110] On the remote conference screen shown in FIG. 15, participant icons I31 to I33 representing the users participating in the conference are displayed. Assuming that the remote conference screen shown in FIG. 15 is the screen displayed on the client terminal 2A used by user A, the participant icons I31 to I33 represent users B to D, respectively. The participant icons I31 to I33 are displayed at positions corresponding to the positions of users B to D in the virtual space.

[0111] Below the participant icons I31 to I33, a virtual reaction button 301 is displayed. The virtual reaction button 301 is a button that is pressed when instructing the output of the sound effect of clapping. Similar screens are also displayed on the client terminals 2 used by users B to D.

[0112] For example, when users B and C press the virtual reaction button 301, as shown in FIG. 16, icons indicating that users B and C are using the clapping function are displayed next to the participant icons I31 and I32.

[0113] Also, the sound effect of clapping is reproduced on the communication management server 1 side as system audio and is delivered to each listening user together with the voice of the speaking user. Audio localization processing for localizing the audio image at a predetermined position is also performed on the audio data of the sound effect of clapping.

[0114] FIG. 17 is a diagram showing a process flow of outputting sound effects using a virtual reaction function.

[0115] When the virtual reaction button 301 is pressed, operation information indicating that the output of the applause sound effect has been instructed is transmitted from the client terminal 2 to the communication management server 1 as shown by the arrows A11 and A12.

[0116] When the microphone voice is transmitted from the client terminal 2 as shown by the arrows A13 and A14, in the communication management server 1, the applause sound effect is added to the microphone voice, and sound image localization processing using HRTF data according to the positional relationship is performed on each of the voice data of the speaking user and the voice data of the sound effect.

[0117] For example, sound image localization processing for localizing the sound image at the same position as the position of the user who instructed the output of the applause sound effect is performed on the voice data of the sound effect. In this case, the sound image of the applause sound effect will be felt to be localized at the same position as the position of the user who instructed the output of the applause sound effect.

[0118] When there are a plurality of users who instructed the output of the applause sound effect, sound image localization processing for localizing the sound image at the center of gravity position of the positions of the plurality of users who instructed the output of the applause sound effect is performed on the voice data of the sound effect. In this case, the sound image of the applause sound effect will be felt to be localized at the position where the users who instructed the output of the applause sound effect are concentrated. It is possible to localize the sound image of the sound effect at various positions selected based on the positions of the users who instructed the output of the applause sound effect instead of the center of gravity position.

[0119] The voice data generated by the sound image localization processing is transmitted to the client terminal 2 used by each listening user as shown by the arrow A15 and output.

[0120] In this example, when the output of the applause sound effect is instructed by a specific user, in response to an action such as the execution of the applause function, HRTF data for localizing the sound image of the applause sound effect to a predetermined position is selected. Also, based on the audio data obtained by the sound image localization process using the selected HRTF data, the applause sound effect is provided to each listening user as audio content.

[0121] Note that in FIG. 17, the microphone voices #1 to #N shown at the uppermost stage using a plurality of blocks are the voices of the speaking users detected in different client terminals 2, respectively. Also, the audio output shown at the lowermost stage using one block represents the output at the client terminal 2 used by one listening user.

[0122] As shown on the left side of FIG. 17, for example, the functions indicated by the arrows A11 and A12 regarding the instruction to send a virtual reaction are realized by the transmission-side module 201A-1. Also, the sound image localization process using HRTF data is realized by the server program 101A.

[0123] Referring to the flowchart of FIG. 18, the control process of the communication management server 1 regarding the output of the sound effect using the virtual reaction function will be described.

[0124] Regarding the content that overlaps with the content described with reference to FIG. 8 in the control process of the communication management server 1, the description will be omitted as appropriate. The same applies to FIGS. 21 and the like described later.

[0125] In step S101, the system audio management unit 136 (FIG. 11) receives operation information indicating that the output of the applause sound effect has been instructed. When the user presses the virtual reaction button 301, operation information indicating that the output of the applause sound effect has been instructed is transmitted from the client terminal 2 used by that user. The transmission of the operation information is performed, for example, by the user status recognition unit 223 (FIG. 14) of the client terminal 2.

[0126] In step S102, the voice receiving unit 131 receives voice data transmitted from the client terminal 2 used by the speaking user. The voice data received by the voice receiving unit 131 is supplied to the sound image localization processing unit 134 via the signal processing unit 132.

[0127] In step S103, the system voice management unit 136 outputs voice data of the applause sound effect to the sound image localization processing unit 134 and adds it as voice data to be the target of sound image localization processing.

[0128] In step S104, the sound image localization processing unit 134 reads and obtains HRTF data corresponding to the positional relationship between the position of the listening user and the position of the speaking user, and HRTF data corresponding to the positional relationship between the position of the listening user and the position of the applause sound effect from the HRTF data storage unit 135. The position of the applause sound effect is a predetermined position as described above selected as the position for localizing the sound image of the applause sound effect.

[0129] The sound image localization processing unit 134 performs sound image localization processing using HRTF data for the speaking voice on the voice data of the speaking user, and performs sound image localization processing using HRTF data for the sound effect on the voice data of the applause sound effect.

[0130] In step S105, the voice transmission unit 138 transmits the voice data obtained by the sound image localization processing to the client terminal 2 used by the listening user.

[0131] Through the above processing, on the client terminal 2 used by the listening user, the sound image of the voice of the speaking user and the sound image of the applause sound effect can be felt to be localized at predetermined positions respectively.

[0132] Note that instead of performing sound image localization processing on the voice data of the speaking user and the voice data of the applause sound effect respectively, sound image localization processing may be performed on the synthesized voice data obtained by synthesizing the voice data of the applause sound effect with the voice data of the speaking user. Also by this, the sound image of the applause sound effect is localized at the same position as the position of the user who instructed the output of the applause sound effect.

[0133] Through the above processing, it becomes possible to share the applause sound effect that expresses the empathy and surprise of individual users as a common voice among all users.

[0134] Also, since the sound image of the applause sound effect is localized and felt at the same position as the position of the user who instructed its output, each listening user can intuitively recognize who the user is who is showing reactions such as empathy and surprise.

[0135] The output of the voice including the microphone voice of the speaking user and the applause sound effect may be performed as follows.

[0136] (A) As shown at the tip of arrow A16 in FIG. 17, the microphone voice whose voice quality has been changed by the filter processing on the client terminal 2 side (transmission-side module 201A-1) is transmitted to the communication management server 1. For example, filter processing for changing the voice quality to that of an elderly person or a child is performed on the microphone voice of the speaking user.

[0137] (B) Depending on the number of users who simultaneously instructed the output of the sound effect, the type of sound effect reproduced as the system voice is changed. For example, when the number of users who instructed the output of the applause sound effect is equal to or more than the threshold number of people, instead of the applause sound effect, a sound effect representing the cheers of a large number of people is reproduced and delivered to the listening users. The selection of the type of sound effect is performed by the system voice management unit 136.

[0138] For sound effects representing cheers, HRTF data for localization at a predetermined position, such as a position near the listening user, an upper position, or a lower position, is selected, and sound image localization processing is performed.

[0139] The position where the sound image of the sound effect is localized may be changed according to the number of users who simultaneously instructed the output of the sound effect, or the volume may be changed.

[0140] Functions for conveying other reactions different from applause, such as a function for expressing joy or a function for expressing anger, may be prepared as virtual reaction functions. Different voice data is reproduced and output as a sound effect for each type of reaction. The position where the sound image is localized may be changed for each type of reaction.

[0141] <Whisper function> The whisper function is a function that designates one user as the listening user and allows speech. The voice of the speaking user is delivered only to the designated user and not to other users. Delivering voice to one user using the whisper function is specified from the screen displayed as a GUI on the display 207 of the client terminal 2.

[0142] FIG. 19 is a diagram showing an example of a remote conference screen.

[0143] Similar to the screen described with reference to FIG. 15, participant icons I31 to I33 representing the users participating in the conference are displayed on the remote conference screen. Assuming that the remote conference screen shown in FIG. 19 is the screen displayed on the client terminal 2A used by user A, the participant icons I31 to I33 represent users B to D, respectively.

[0144] For example, when the participant icon I31 is selected by user A using the cursor, user B is designated as the whisper target user who is the listening destination of the voice. The participant icon I31 representing user B is highlighted as shown in FIG. 19.

[0145] When user A speaks in this state, in the communication management server 1, audio-visual localization processing is performed on the voice data of user A to localize the audio-visual at the ear of user B designated as the user to be whispered to.

[0146] Note that the default state is a state where the user to be whispered to is not designated. The voice of the speaking user is delivered to all other users so that the audio-visual is localized at a position according to the positional relationship between the listening user and the speaking user.

[0147] Figure 20 is a diagram showing the flow of processing related to the output of voice using the whispering function.

[0148] When the user to be whispered to is designated by selecting the participant icon, operation information indicating that the user to be whispered to has been designated is transmitted from the client terminal 2 to the communication management server 1 as shown by arrow A21.

[0149] When an image captured by a camera is analyzed and the posture for whispering is estimated, operation information indicating that the user to be whispered to has been designated may be transmitted as shown by arrow A22.

[0150] When microphone voice is transmitted from the client terminal 2 used by the user who performed the whispering as shown by arrow A23, in the communication management server 1, audio-visual localization processing is performed on the voice data of microphone voice #1 to localize the audio-visual at the position of the ear of the user designated as the target of whispering. That is, HRTF data corresponding to the position of the ear of the user designated as the target of whispering is selected and used for the audio-visual localization processing.

[0151] In Figure 20, the microphone voice #1 indicated by arrow A23 is the voice of the user who performed the whispering, that is, the speaking user who designated one user as the user to be whispered to using the whispering function.

[0152] The voice data generated by the sound image localization process is transmitted to and output from the client terminal 2 used by the user targeted for whispering, as indicated by arrow A24.

[0153] On the other hand, as indicated by arrow A25, when microphone voice is transmitted from the client terminal 2 used by a user who is not using the whispering function, in the communication management server 1, sound image localization processing is performed using HRTF data according to the positional relationship between the listening user and the speaking user.

[0154] The voice data generated by the sound image localization process is transmitted to and output from the client terminal 2 used by the listening user, as indicated by arrow A26.

[0155] In this example, when the user targeted for whispering is instructed by a specific user, in response to an action such as the execution of the whispering function, HRTF data for localizing the sound image of the voice of the user using the whispering function to the ear of the user targeted for whispering is selected. Also, based on the voice data obtained by the sound image localization processing using the selected HRTF data, the voice of the user using the whispering function is provided as voice content to the user targeted for whispering.

[0156] With reference to the flowchart of FIG. 21, the control process of the communication management server 1 regarding the output of voice using the whispering function will be described.

[0157] In step S111, the system voice management unit 136 receives operation information indicating that the user targeted for whispering has been selected. When a user selects the user targeted for whispering, operation information indicating that the user targeted for whispering has been selected is transmitted from the client terminal 2 used by that user. The transmission of the operation information is performed, for example, by the user status recognition unit 223 of the client terminal 2.

[0158] In step S112, the voice receiving unit 131 receives voice data transmitted from the client terminal 2 used by the user who has made a whisper. The voice data received by the voice receiving unit 131 is supplied to the audio-visual localization processing unit 134.

[0159] In step S113, the audio-visual localization processing unit 134 reads and acquires HRTF data corresponding to the position near the ear of the user targeted for whispering from the HRTF data storage unit 135. Also, the audio-visual localization processing unit 134 performs audio-visual localization processing using the HRTF data on the voice data of the speaking user (the user who has made a whisper) so as to localize the sound image near the ear of the user targeted for whispering.

[0160] In step S114, the voice transmission unit 138 transmits the voice data obtained by the audio-visual localization processing to the client terminal 2 used by the user targeted for whispering.

[0161] In the client terminal 2 used by the user targeted for whispering, based on the voice data transmitted from the communication management server 1, the voice of the user who has made a whisper is output. The user selected as the target for whispering will listen to the voice of the user who has made a whisper while feeling the sound image near the ear.

[0162] Through the above processing, even when there are multiple participants in the meeting, the speaking user can specify one user and talk only to that user.

[0163] It is also possible to enable the specification of multiple users as the users targeted for whispering.

[0164] Also, for the user (listener) selected as the whispering target, the voices of other users who are speaking simultaneously may be delivered together with the voice of the user who whispered. In this case, for the voice data of the user who whispered, sound image localization processing is performed so that the sound image is localized near the listener's ear. Also, for the voice data of other users who are not whispering, sound image localization processing using HRTF data according to the positional relationship between the position of the listener and the position of the speaking user is performed.

[0165] It is possible to localize the sound image of the voice of the user who whispered at an arbitrary position near the user who is the whispering target, rather than near the ear of the user who is the whispering target. The position where the sound image is to be localized may be specified by the user who whispered.

[0166] <Focus function> The focus function is a function that designates one user as the focus target and makes it easier to hear the voice of that user. While the above-described whispering function is a function used by the speaking user, the focus function is a function used by the listening user. The user who is the focus target is specified from the screen displayed as a GUI on the display 207 of the client terminal 2.

[0167] FIG. 22 is a diagram showing an example of a remote conference screen.

[0168] Similar to the screen described with reference to FIG. 15, on the remote conference screen, participant icons I31 to I33 representing the users participating in the conference are displayed. Assuming that the remote conference screen shown in FIG. 22 is the screen displayed on the client terminal 2A used by user A, the participant icons I31 to I33 represent users B to D, respectively.

[0169] For example, when the participant icon I31 is selected by user A using the cursor, user B is designated as the focus target user. The participant icon I31 representing user B is highlighted as shown in FIG. 22.

[0170] When User B speaks in this state, in the communication management server 1, for the voice data of User B, audio-visual localization processing is performed to localize the audio-visual image near User A, who has been designated as the user to be focused on. When Users C and D, who are not designated as the focus target, speak, for the voice data of User C and User D respectively, audio-visual localization processing using HRTF data corresponding to the positional relationship with User A is performed.

[0171] Note that the default state is the state where no user to be focused on is designated. The voice of the speaking user is delivered to all other users so that the audio-visual image is localized at a position corresponding to the positional relationship between the listening user and the speaking user.

[0172] FIG. 23 is a diagram showing the flow of processing related to the output of voice using the focus function.

[0173] When the user to be focused on is designated by selecting the participant icon, operation information indicating that the user to be focused on has been designated is transmitted from the client terminal 2 to the communication management server 1 as shown by arrow A31.

[0174] Alternatively, when an image captured by a camera is analyzed and the focus target is estimated based on gaze detection or the like, operation information indicating that the user to be focused on has been designated may be transmitted as shown by arrow A32.

[0175] When microphone voice is transmitted from the client terminal 2 as shown by arrows A33 and A34, in the communication management server 1, for the voice data of the microphone voice of the user to be focused on, audio-visual localization processing is performed to localize the audio-visual image near the user. That is, HRTF data corresponding to a position near the position of the user who designated the focus target is selected and used for the audio-visual localization processing.

[0176] In addition, audio localization processing is performed on the audio data of the microphone voice of users other than the user targeted for focus to localize the sound image at a position away from the user. That is, HRTF data corresponding to a position away from the position of the user who specified the focus target is selected and used for the audio localization processing.

[0177] In the example of FIG. 23, for example, the microphone voice #1 indicated by the arrow A33 is the microphone voice of the user targeted for focus. The audio data of the microphone voice #1 is transmitted from the client terminal 2 used by the user targeted for focus to the communication management server 1.

[0178] Also, the microphone voice #N indicated by the arrow A34 is the microphone voice of a user other than the user targeted for focus. The audio data of the microphone voice #N is transmitted from the client terminal 2 used by a user other than the user targeted for focus to the communication management server 1.

[0179] The audio data generated by the audio localization processing is transmitted to and output from the client terminal 2 used by the user who specified the focus target, as shown by the arrow A35.

[0180] In this example, when the user targeted for focus is selected by a specific user, HRTF data for localizing the sound image of the voice of the user targeted for focus near the user who selected the focus target is selected according to an action such as the execution of the focus function. Also, based on the audio data obtained by the audio localization processing using the selected HRTF data, the voice of the user targeted for focus is provided as audio content to the user who selected the focus target.

[0181] With reference to the flowchart of FIG. 24, the control processing of the communication management server 1 regarding the output of voice using the focus function will be described.

[0182] In step S121, the participant information management unit 133 receives operation information indicating that the user targeted for focus has been selected. When a certain user selects the user targeted for focus, operation information indicating that the user targeted for focus has been selected is transmitted from the client terminal 2 used by that user. The transmission of the operation information is performed, for example, by the user status recognition unit 223 of the client terminal 2.

[0183] In step S122, the voice reception unit 131 receives voice data transmitted from the client terminal 2. For example, together with the voice data of the user targeted for focus, voice data of users other than the user targeted for focus (users not selected as the focus target) is received. The voice data received by the voice reception unit 131 is supplied to the audio-visual localization processing unit 134.

[0184] In step S123, the audio-visual localization processing unit 134 reads and acquires HRTF data corresponding to a position near the user who selected the focus target from the HRTF data storage unit 135. Also, the audio-visual localization processing unit 134 performs audio-visual localization processing using the acquired HRTF data on the voice data of the user targeted for focus so as to localize the audio-visual at a position near the user who selected the focus target.

[0185] In step S124, the audio-visual localization processing unit 134 reads and acquires HRTF data corresponding to a position away from the user who selected the focus target from the HRTF data storage unit 135. Also, the audio-visual localization processing unit 134 performs audio-visual localization processing using the acquired HRTF data on the voice data of users other than the user targeted for focus so as to localize the audio-visual at a position away from the user who selected the focus target.

[0186] In step S125, the voice transmission unit 138 transmits the voice data obtained by the audio-visual localization processing to the client terminal 2 used by the user who selected the focus target.

[0187] In the client terminal 2 used by the user who has selected the focus target, the voice of the speaking user is output based on the voice data transmitted from the communication management server 1. The user who has selected the focus target will listen to the voice of the focus target user while feeling that the sound image is close. Also, the user who has selected the focus target will listen to the voice of a user other than the focus target user while feeling that the sound image is at a distant position.

[0188] Through the above processing, even when there are multiple participants in the meeting, the user can specify one user and concentrate on listening to the speech of that user.

[0189] It may be possible to select multiple users as the focus target users.

[0190] Instead of selecting the focus target user, it may be possible to select the user to be avoided. In this case, for the voice data of the voice of the user selected as the user to be avoided, sound image localization processing is performed so that the sound image is localized at a position away from the listening user.

[0191] <Dynamic switching of sound image localization processing> It is dynamically switched whether the sound image localization processing, which is the processing of object audio including rendering, etc., is performed on the communication management server 1 side or on the client terminal 2 side.

[0192] In this case, at least the same configurations as the sound image localization processing unit 134, the HRTF data storage unit 135, and the 2ch mix processing unit 137 in the configuration shown in FIG. 11 of the communication management server 1 are also provided in the client terminal 2. The same configurations as the sound image localization processing unit 134, the HRTF data storage unit 135, and the 2ch mix processing unit 137 are realized, for example, by the receiving side module 201A-2.

[0193] When parameters used for audio-visual localization processing, such as the listening user's location information, are changed during a meeting and the change is reflected in the audio-visual localization processing in real time, the audio-visual localization processing is performed on the client terminal 2 side. By performing the audio-visual localization processing locally, it becomes possible to speed up the response to parameter changes.

[0194] On the other hand, when there is no parameter setting change for a certain period of time or more, the audio-visual localization processing is performed on the communication management server 1 side. By performing the audio-visual localization processing on the server, it becomes possible to suppress the data communication volume between the communication management server 1 and the client terminal 2.

[0195] Figure 25 is a diagram showing the processing flow related to the dynamic switching of the audio-visual localization processing.

[0196] When the audio-visual localization processing is performed on the client terminal 2 side, the microphone voice transmitted from the client terminal 2 as shown by the arrows A101 and A102 is transmitted directly to the client terminal 2 as shown by the arrow A103. The client terminal 2 that is the transmission source of the microphone voice is the client terminal 2 used by the speaking user, and the client terminal 2 that is the transmission destination of the microphone voice is the client terminal 2 used by the listening user.

[0197] When parameters related to the localization of audio-visual, such as the location of the listening user, are changed by the listening user as shown by the arrow A104, the audio-visual localization processing is performed on the microphone voice transmitted from the communication management server 1 in real time reflecting the change in the setting.

[0198] Voice corresponding to the audio data generated by the audio-visual localization processing on the client terminal 2 side is output as shown by the arrow A105.

[0199] In the client terminal 2, the change content of the parameter setting is saved, and information representing the change content is transmitted to the communication management server 1 as shown by the arrow A106.

[0200] When the audio-visual localization process is performed on the communication management server 1 side, for the microphone audio transmitted from the client terminal 2 as shown by the arrows A107 and A108, the audio-visual localization process is performed reflecting the changed parameters.

[0201] The audio data generated by the audio-visual localization process is transmitted to the client terminal 2 used by the listening user as shown by the arrow A109 and output.

[0202] With reference to the flowchart of FIG. 26, the control process of the communication management server 1 regarding the dynamic switching of the audio-visual localization process will be described.

[0203] In step S201, it is determined whether or not the parameter setting has not been changed for a certain period of time or more. This determination is made by the participant information management unit 133 based on information transmitted from the client terminal 2 used by the listening user, for example.

[0204] If it is determined in step S201 that there is a parameter setting change, in step S202, the voice transmission unit 138 transmits the voice data of the speaking user received by the participant information management unit 133 as it is to the client terminal 2 used by the listening user. The voice data to be transmitted becomes object audio data.

[0205] In the client terminal 2, the audio-visual localization process is performed using the changed settings and the voice is output. Also, information representing the content of the changed settings is transmitted to the communication management server 1.

[0206] In step S203, the participant information management unit 133 receives information representing the content of the setting change transmitted from the client terminal 2. Based on the information transmitted from the client terminal 2, after updating the position information of the listening user, etc., the process returns to step S201, and subsequent processes are performed. The audio-visual localization process performed on the communication management server 1 side is performed based on the updated position information.

[0207] On the other hand, if it is determined in step S201 that there is no change in the parameter settings, in step S204, the audio-visual localization process on the communication management server 1 side is performed. The process performed in step S204 is basically the same process as the process described with reference to FIG. 8.

[0208] The above processes are also performed when other parameters such as not only the change in position but also the change in the background sound setting are changed.

[0209] <Management of Acoustic Settings> Acoustic settings suitable for background sound may be stored in a database and managed by the communication management server 1. For example, for each type of background sound, a position suitable as the position for localizing the audio-visual is set, and the HRTF data corresponding to the set position is stored. Parameters related to other acoustic settings such as reverb may also be stored.

[0210] FIG. 27 is a diagram showing the flow of the process related to the management of acoustic settings.

[0211] When synthesizing background sound with the voice of the speaking user, in the communication management server 1, the background sound is played, and as shown by the arrow A121, the audio-visual localization process is performed using acoustic settings such as HRTF data suitable for the background sound.

[0212] The voice data generated by the audio-visual localization process is transmitted to the client terminal 2 used by the listening user as shown by the arrow A122 and output.

[0213] <<Modification Example>> Although the conversation conducted by multiple users is assumed to be a conversation in a remote meeting, the above-described technology is applicable to various types of conversations as long as they are conversations in which multiple people participate via the Internet, such as conversations at a meal or conversations at a lecture.

[0214] ·Regarding the program The above-described series of processes can be executed by hardware or by software. When the series of processes is executed by software, the program constituting the software is installed in a computer incorporated in dedicated hardware or in a general-purpose personal computer or the like.

[0215] The program to be installed is recorded and provided on a removable medium 111 shown in FIG. 10, which includes an optical disk (such as a CD-ROM (Compact Disc-Read Only Memory) or a DVD (Digital Versatile Disc)) or a semiconductor memory. Further, it may be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting. The program can be installed in advance in the ROM 102 or the storage unit 108.

[0216] Note that the program executed by the computer may be a program in which processing is performed in time series in accordance with the order described in this specification, or a program in which processing is performed in parallel or at a necessary timing such as when a call is made.

[0217] In the specification, a system means a collection of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and a single device in which a plurality of modules are housed in one housing are both systems.

[0218] The effects described in this specification are merely illustrative and not limiting, and there may be other effects.

[0219] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the gist of the present technology. Although headphones or speakers are used as the audio output device, other devices may be used. For example, ordinary earphones (inner ear headphones) or open-type earphones capable of capturing environmental sound can be used as the audio output device.

[0220] For example, the present technology can adopt a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.

[0221] In addition, each step described in the above flowchart can be executed by one device or can be shared and executed by a plurality of devices.

[0222] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or can be shared and executed by a plurality of devices.

[0223] · Example of configuration combination The present technology can also adopt the following configuration.

[0224] (1) A storage unit that stores HRTF data corresponding to a plurality of positions based on the listening position, An audio localization processing unit that provides audio content selected according to the action by performing audio localization processing using the HRTF data selected according to the action by a specific participant among the participants of the conversation participating via the network so that the sound image is localized at a predetermined position An information processing apparatus comprising (2) In response to the action of instructing the output of the sound effect being performed by the specific participant, the audio-visual localization processing unit provides the audio content for outputting the sound effect. The information processing apparatus according to (1) above. (3) The audio-visual localization processing unit performs the audio-visual localization processing on the voice data of the sound effecter using the HRTF data according to the relationship between the position of the participant who is the listener in the virtual space and the position of the specific participant who performed the action. The information processing apparatus according to (2) above. (4) In response to the action of selecting the participant who is the listening destination being performed by the specific participant, the audio-visual localization processing unit provides the audio content for outputting the voice of the specific participant. The information processing apparatus according to (1) above. (5) The selection of the participant who is the listening destination is performed using the visual information that visually represents the participant and is displayed on the screen. The information processing apparatus according to (4) above. (6) The audio-visual localization processing unit performs the audio-visual localization processing on the voice data of the specific participant using the HRTF data according to the position near the ear of the participant who is the listening destination in the virtual space. The information processing apparatus according to (4) or (5) above. (7) In response to the action of selecting the speaker to be focused on being performed by the specific participant, the audio-visual localization processing unit provides the audio content for outputting the voice of the speaker. The information processing apparatus according to (1) above. (8) The selection of the speaker to be focused on is performed using the visual information that visually represents the participant and is displayed on the screen. The information processing apparatus according to (7) above. (9) The audio-visual localization processing unit performs the audio-visual localization processing on the voice data of the speaker to be focused on, using the HRTF data corresponding to a position near the position of the specific participant in the virtual space. The information processing apparatus according to (7) or (8) above. (10) The information processing apparatus stores HRTF data corresponding to a plurality of positions based on the listening position, and provides the voice content selected according to the action by performing audio-visual localization processing using the HRTF data selected according to the action of a specific participant among the participants in the conversation participating via the network, so that the audio-visual is localized at a predetermined position. Information processing method. (11) Causes a computer to store HRTF data corresponding to a plurality of positions based on the listening position, and provides the voice content selected according to the action by performing audio-visual localization processing using the HRTF data selected according to the action of a specific participant among the participants in the conversation participating via the network, so that the audio-visual is localized at a predetermined position. A program for executing the processing. (12) stores HRTF data corresponding to a plurality of positions based on the listening position, performs audio-visual localization processing using the HRTF data selected according to the action of a specific participant among the participants in the conversation participating via the network, and receives the voice content obtained by performing the audio-visual localization processing, which is transmitted from an information processing apparatus that provides the voice content selected according to the action so that the audio-visual is localized at a predetermined position, and includes a voice receiving unit that outputs voice. Information processing terminal. (13) The voice receiving unit receives the voice data of the sound effect, which is transmitted in response to the action of instructing the output of the sound effect being performed by the specific participant. The information processing terminal described in (12). (14) The voice receiving unit receives the voice data of the effector obtained by performing the sound image localization process using the HRTF data according to the relationship between the position of the user of the information processing terminal and the position of the specific participant who performed the action in the virtual space. The information processing terminal described in (13). (15) The voice receiving unit receives the voice data of the specific participant that has been transmitted in response to the action of selecting the user of the information processing terminal as the participant for listening to the voice being performed by the specific participant. The information processing terminal described in (12). (16) The voice receiving unit receives the voice data of the specific participant obtained by performing the sound image localization process using the HRTF data according to the position near the ear of the user of the information processing terminal in the virtual space. The information processing terminal described in (15). (17) The voice receiving unit receives the voice data of the speaker to be focused on that has been transmitted in response to the action of selecting the speaker to be focused on being performed by the user of the information processing terminal as the specific participant. The information processing terminal described in (12). (18) The voice receiving unit receives the voice data of the speaker to be focused on obtained by performing the sound image localization process using the HRTF data according to the position near the position of the user of the information processing terminal in the virtual space. The information processing terminal described in (17). (19) The information processing terminal is Store HRTF data corresponding to a plurality of positions based on the listening position, and perform audio-visual localization processing using the HRTF data selected according to an action by a specific participant among the participants in a conversation participating via a network, so that the audio-visual is localized at a predetermined position, receive the audio content obtained by performing the audio-visual localization processing, which has been transmitted from an information processing device that provides the audio content selected according to the action, and output the audio Information processing method. (20) Cause a computer to Store HRTF data corresponding to a plurality of positions based on the listening position, and perform audio-visual localization processing using the HRTF data selected according to an action by a specific participant among the participants in a conversation participating via a network, so that the audio-visual is localized at a predetermined position, receive the audio content obtained by performing the audio-visual localization processing, which has been transmitted from an information processing device that provides the audio content selected according to the action, and output the audio Program for executing the processing.

Explanation of symbols

[0225] 1 Communication management server, 2A to 2D Client terminals, 121 Information processing unit, 131 Audio reception unit, 132 Signal processing unit, 133 Participant information management unit, 134 Audio-visual localization processing unit, 135 HRTF data storage unit, 136 System audio management unit, 137 2ch mixing processing unit, 138 Audio transmission unit, 201 Control unit, 211 Information processing unit, 221 Audio processing unit, 222 Setting information transmission unit, 223 User status recognition unit, 231 Audio reception unit, 233 Microphone audio acquisition unit

Claims

1. A storage unit that stores HRTF data corresponding to a plurality of positions based on a listening position; An audio-visual localization processing unit that provides audio content corresponding to the type of action selected by a specific participant among the participants participating in a conversation via a network on a screen during the conversation, and performs audio-visual localization processing using the HRTF data, so that the audio is localized at a predetermined position An information processing apparatus comprising:

2. The information processing apparatus according to claim 1, wherein the audio-visual localization processing unit provides the audio content for outputting the sound effect in response to the action of instructing the output of the sound effect being performed by the specific participant.

3. The information processing apparatus according to claim 2, wherein the audio-visual localization processing unit performs the audio-visual localization processing on the audio data of the sound effect using the HRTF data according to the relationship between the position of the participant who is the listener and the position of the specific participant who performed the action in the virtual space. The information processing apparatus according to claim 2.

4. The information processing apparatus according to claim 1, wherein the audio-visual localization processing unit provides the audio content for outputting the voice of the specific participant in response to the action of selecting the participant as the audio listening destination being performed by the specific participant. The information processing apparatus according to claim 1.

5. The information processing apparatus according to claim 4, wherein the selection of the participant as the listening destination is performed using visual information that visually represents the participant and is displayed on the screen. The information processing apparatus according to claim 4.

6. The information processing apparatus according to claim 4, wherein the audio-visual localization processing unit performs the audio-visual localization processing on the audio data of the specific participant using the HRTF data according to the position near the ear of the participant as the listening destination in the virtual space. The information processing apparatus according to claim 4.

7. The information processing apparatus according to claim 1, wherein the audio-visual localization processing unit provides the audio content for outputting the voice of the speaker in response to the action of selecting the speaker as the focus target being performed by the specific participant. The information processing apparatus according to claim 1.

8. The information processing apparatus according to claim 7, wherein the selection of the speaker as the focus target is performed using visual information that visually represents the participant and is displayed on the screen. The information processing apparatus according to claim 7.

9. The information processing apparatus according to claim 7, wherein the audio-visual localization processing unit performs the audio-visual localization processing on the audio data of the speaker as the focus target using the HRTF data according to the position near the position of the specific participant in the virtual space. The information processing apparatus according to claim 7.

10. The information processing apparatus stores HRTF data corresponding to a plurality of positions based on the listening position, and by performing audio-visual localization processing using the HRTF data corresponding to the type of action selected by a specific participant among the participants of the conversation participating via the network on the screen during the conversation, provides audio content corresponding to the type of action so that the audio-visual is localized at a predetermined position. An information processing method.

11. Causes a computer to store HRTF data corresponding to a plurality of positions based on the listening position, and by performing audio-visual localization processing using the HRTF data corresponding to the type of action selected by a specific participant among the participants of the conversation participating via the network on the screen during the conversation, provides audio content corresponding to the type of action so that the audio-visual is localized at a predetermined position. A program for executing the processing.

12. A voice receiving unit that receives the audio content obtained by performing the audio-visual localization processing, which is transmitted from an information processing apparatus that stores HRTF data corresponding to a plurality of positions based on the listening position, and performs audio-visual localization processing using the HRTF data corresponding to the type of action selected by a specific participant among the participants of the conversation participating via the network on the screen during the conversation, so that the audio-visual is localized at a predetermined position, and outputs the voice. An information processing terminal.

13. The voice receiving unit receives the voice data of the sound effect transmitted in response to the action of instructing the output of the sound effect being performed by the specific participant. The information processing terminal according to claim 12.

14. The voice receiving unit receives the voice data of the sound effect obtained by performing the audio-visual localization processing using the HRTF data according to the relationship between the position of the user of the information processing terminal and the position of the specific participant who performed the action in the virtual space. The information processing terminal according to claim 13.

15. The voice receiving unit receives the voice data of the specific participant transmitted in response to the action of selecting the user of the information processing terminal as the participant for listening to the voice being performed by the specific participant. The information processing terminal according to claim 12.

16. The voice receiving unit receives the voice data of the specific participant obtained by performing the sound image localization process using the HRTF data corresponding to the position near the ear of the user of the information processing terminal in the virtual space. The information processing terminal according to claim 15.

17. The voice receiving unit receives the voice data of the speaker to be focused on, which has been transmitted in response to the action of selecting the speaker to be focused on being performed by the user of the information processing terminal as the specific participant. The information processing terminal according to claim 12.

18. The voice receiving unit receives the voice data of the speaker to be focused on obtained by performing the sound image localization process using the HRTF data corresponding to the position near the position of the user of the information processing terminal in the virtual space. The information processing terminal according to claim 17.

19. An information processing terminal stores HRTF data corresponding to a plurality of positions based on a listening position, and performs a sound image localization process using the HRTF data corresponding to the type of action selected by a specific participant among the participants of a conversation participating via a network on a screen during the conversation, and receives the voice content obtained by performing the sound image localization process, which has been transmitted from an information processing device that provides voice content corresponding to the type of action so that the sound image is localized at a predetermined position, and outputs the voice. An information processing method.

20. A program causing a computer to store HRTF data corresponding to a plurality of positions based on a listening position, and perform a sound image localization process using the HRTF data corresponding to the type of action selected by a specific participant among the participants of a conversation participating via a network on a screen during the conversation, and receive the voice content obtained by performing the sound image localization process, which has been transmitted from an information processing device that provides voice content corresponding to the type of action so that the sound image is localized at a predetermined position, and output the voice. to execute the process.

Citation Information

Patent Citations

  • Digital processing circuit, headphone device and speaker using it

    JP1999331992A

  • Portable telephone terminal

    JP2006287878A

  • Voice output control device, voice output control method, program, and recording medium

    JP2014011509A

  • Spatial Audio Teleconferencing

    US20080144794A1

  • Sound Localization for an Electronic Call

    US20150373477A1