Sound control device, sound control system, and sound control method
The sound control device and method improve intuitive understanding of presenter focus in shared virtual spaces by localizing sound images based on gaze and audio information, addressing the challenge of material identification in multi-user virtual presentations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2026-03-13
AI Technical Summary
Conventional information processing devices struggle to intuitively convey which material a presenter is explaining in a virtual space when multiple users share content, especially in presentations using materials arranged over the entire sphere.
A sound control device and method that acquires gaze and audio information from a first user, identifies their gaze position, and localizes the sound image position based on this information within a shared display space, using wearable devices and a server to enhance understanding of the presenter's focus.
Facilitates easier comprehension of which material the presenter is explaining, enhancing viewer understanding in shared virtual spaces.
Smart Images

Figure 2026046247000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an audio control device, an audio control system, and an audio control method.
Background Art
[0002] When a plurality of users wearing XR (Extended Reality) glasses share content in a virtual space, there is a known technique for displaying the range displayed on their own XR glasses on the XR glasses of other users. For example, Patent Document 1 discloses an information processing device in which, when a first user and a second user view the same content, a viewing direction indication mark indicating the direction in which the viewing field of the first user exists is superimposed and displayed on the content being viewed by the second user.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, when the content in the virtual space is, for example, a presentation using a plurality of materials arranged over the entire sphere, the conventional information processing device has a problem that it is difficult for viewers to intuitively understand which material the presenter of the presentation is explaining.
[0005] The present invention has been made to solve the above problems, and an object thereof is to make it easier for viewers to understand which material in the virtual space the presenter is explaining.
Means for Solving the Problems
[0006] A preferred embodiment of the present invention provides a sound control device comprising: an acquisition unit that acquires gaze information relating to the gaze of a first user and audio information relating to the speech of the first user; an identification unit that identifies the gaze position of the first user in a display space shared by the first user and the second user based on the gaze information; and a sound control unit that localizes the sound image position of the first user's speech based on the audio information in the space to which the identified gaze position belongs among the first space and a second space separate from the first space, which are included in the display space.
[0007] A sound control system according to a preferred embodiment of the present invention includes a first wearable device worn on the head of a first user, a second wearable device worn on the head of a second user, an acquisition process for acquiring gaze information relating to the gaze of the first user and audio information relating to the speech of the first user, a identification process for identifying the gaze position of the first user in a display space shared by the first user and the second user based on the gaze information, and a sound control process for localizing the sound image position of the speech of the first user based on the audio information in the space to which the identified gaze position belongs among the first space and a second space separate from the first space, which are included in the display space.
[0008] A preferred embodiment of the present invention provides a sound control method which acquires gaze information relating to the gaze of a first user and audio information relating to the speech of the first user, identifies the gaze position of the first user in a display space shared by the first user and the second user based on the gaze information, and positions the sound image of the first user's speech based on the audio information in the space to which the identified gaze position belongs among the first space and the second space separate from the first space, which are included in the display space. [Effects of the Invention]
[0009] The sound control device and sound control method according to the present invention make it easier for viewers to understand which material in the virtual space the presenter is explaining. [Brief explanation of the drawing]
[0010] [Figure 1] This figure shows the overall configuration of a sound control system including a sound control device according to the first embodiment. [Figure 2] Figure 1 is a schematic diagram showing an example of a virtual space displayed by MR glasses. [Figure 3] Figure 1 is a perspective view showing the appearance of the MR glasses. [Figure 4] This block diagram shows an example of the configuration of MR glasses in Figure 1. [Figure 5] This is a block diagram showing an example of the configuration of a terminal device in Figure 1. [Figure 6] This block diagram shows an example of the server configuration in Figure 1. [Figure 7] This is a schematic diagram illustrating the relationship between the virtual space and the display area of the MR glasses. [Figure 8] This is a schematic diagram illustrating the relationship between the first user's line of sight and the first object. [Figure 9] This is a schematic diagram illustrating an example of how a first user places presentation materials in a virtual space. [Figure 10] This is a schematic diagram illustrating an example of a virtual space (VS) shared by a second user. [Figure 11] This schematic diagram illustrates an example of a case where the sound image is localized near the first user's gaze position in a virtual space (VS) shared by the second user. [Figure 12] Figure 6 is a flowchart showing an example of the sound control operation of the processing unit. [Modes for carrying out the invention]
[0011] 1. First Embodiment The configuration of the sound control device according to the first embodiment of the present invention will be described below with reference to Figures 1 to 12.
[0012] 1.1. Configuration of the First Embodiment 1.1.1. Overall Structure FIG. 1 is a diagram showing the overall configuration of a sound control system 1 including a sound control device according to the first embodiment. The sound control system 1 includes an MR (Mixed Reality) glass 10, a terminal device 20, a server 30, and a communication network NET. In the sound control system 1, the terminal device 20 and the server 30 are communicably connected to each other via the communication network NET. In the configuration shown in FIG. 1, the MR glass 10 is not directly connected to the communication network NET, but it may be directly connected to the communication network NET.
[0013] The MR glass 10 includes n MR glasses 10[1], 10[2],..., 10[k],..., 10[n]. Here, n is an arbitrary natural number, and k is an arbitrary natural number smaller than n. In the present embodiment, the configurations of the MR glasses 10[1] to 10[n] are the same as each other. Note that the MR glass 10 may include MR glasses with different configurations.
[0014] The user who uses the MR glass 10[1] is the user U[1], the user who uses the MR glass 10[2] is the user U[2], the user who uses the MR glass 10[k] is the user U[k], and the user who uses the MR glass 10[n] is the user U[n]. When referring to an unspecified number of users or all users, the user is also denoted as the user U.
[0015] The MR glasses 10 are wearable display devices worn on the head of the user U. The MR glasses 10 display virtual objects on display panels provided on each lens corresponding to both eyes of the user U. Each lens and display panel of the MR glasses 10 are see-through. Therefore, the user U wearing the MR glasses 10 can visually recognize the real space as an external image through each lens and display panel of the MR glasses 10. The external image may be a virtual image obtained by imaging the surroundings of the user U. The method using the real external image is called the optical see-through method. The method using the virtual external image is called the video see-through method. In the following description, it is assumed that the MR glasses 10 employ the optical see-through method. Note that the sound control system 1 can also be realized by using a head-mounted display that cannot visually recognize the external image instead of the MR glasses 10.
[0016] The terminal device 20 includes n terminal devices 20[1], 20[2], …, 20[k], …, 20[n]. Here, n is an arbitrary natural number, and k is an arbitrary natural number smaller than n. In the present embodiment, the configurations of the terminal devices 20[1] to 20[n] are the same as each other. Note that the terminal device 20 may include terminal devices with different configurations.
[0017] The user who uses the terminal device 20[1] is the user U[1], the user who uses the terminal device 20[2] is the user U[2], the user who uses the terminal device 20[k] is the user U[k], and the user who uses the terminal device 20[n] is the user U[n].
[0018] In the present embodiment, the terminal device 20 includes devices such as a PC, a tablet terminal, a smartphone, and a smartwatch. In the present embodiment, it is assumed that the terminal device 20 is a smartphone for explanation. The terminal device 20[k] displays various digital contents distributed from the server 30 as virtual objects on the corresponding MR glasses 10[k].
[0019] Server 30 is a device that provides digital content. Server 30 receives requests from terminal devices 20 via the communication network NET and delivers various digital content to terminal devices 20 in response to requests from terminal devices 20. In this embodiment, the number of MR glasses 10, terminal devices 20, and users U are each set to n, but these numbers only need to be at least 2 each.
[0020] Figure 2 is a schematic diagram showing an example of the virtual space VS displayed by the MR glasses 10[k] in Figure 1. The virtual space VS is defined as the space inside a virtual celestial sphere CS centered on the position of the user U[k]'s head. If the radius of the celestial sphere CS is R, then the virtual space VS can be said to be the space inside a sphere of radius R. In other words, the celestial sphere CS represents the entire range of the virtual space VS. The inner surface of the celestial sphere CS is hereafter referred to as the celestial sphere SS. The zenith ZE of the celestial sphere CS is the point above the user U[k]'s head on the celestial sphere SS. The naval base NA of the celestial sphere CS is the point below the user U[k]'s feet on the celestial sphere SS. The equatorial EQ is the 0-degree latitude line of the celestial sphere CS. The virtual space VS is an example of a display space.
[0021] As shown in Figure 2, one or more virtual objects VO are placed in the virtual space VS. In this example, the virtual object VO includes five virtual objects VOa to VOe. In this embodiment, each of the virtual objects VOa to VOe is placed on the celestial sphere SS. Note that although user U[k] is shown in Figure 2 for convenience, user U[k] is not actually displayed on the MR glasses 10[k].
[0022] One or more virtual objects VO are, for example, virtual objects representing data such as still images, videos, 3D CG models, HTML files, and text files, and virtual objects representing applications. Examples of text files include memos and source code. Examples of applications include browsers, applications for using social networking services, and applications for generating document files. The number of one or more virtual objects VO in Figure 2 is merely illustrative, and the number of one or more virtual objects VO is not limited to this disclosure.
[0023] 1.1.2. Configuration of MR Glasses Figure 3 is a perspective view showing the appearance of the MR glasses 10[k] in Figure 1. The MR glasses 10[k] comprises a temple 94L, a temple 94R, a bridge 96, a projection optical system 98L, a projection optical system 98R, a sound output device 17, and a sound pickup device 18. In the following description, when distinguishing between similar elements, suffixes such as "L" in temple 94L and "R" in temple 94R are used. When similar elements are not distinguished, only a common number without a suffix is used, such as temple 94.
[0024] The temple 94 is a rod-shaped component supported by the auricle. The bridge 96 is positioned between the projection optical system 98L and the projection optical system 98R. The projection optical system 98 includes a display device 16, a light guide path 981, and a half mirror 982.
[0025] The display device 16 is located within the temple 94. The display device 16 displays an image. The display device 16 has various display panels, such as a liquid crystal panel and an organic EL (Electro-Luminescence) panel. When the display device 16 displays an image, light representing the image is emitted from the display device 16. The light emitted from the display device 16 is guided to the half mirror 982 by the light guide path 981. The half mirror 982 reflects the light guided by the light guide path 981. The light reflected by the half mirror 982 is projected onto the retina of the user U[k]. Through this light, the user U[k] perceives the image. The half mirror 982 has a surface facing the user U[k]. Hereinafter, this surface will be referred to as the "display surface SC". When the display device 16 is not displaying an image, the user U[k] can see the outside world through the half mirror 982.
[0026] The sound output device 17 is positioned on the side of the temple 94. The sound output device 17 outputs sound. User U[k] can hear, for example, the voice spoken by user U[1] through the sound output device 17.
[0027] The sound pickup device 18 is positioned on the side of the temple 94. The sound pickup device 18 primarily picks up the utterances of the user U[k]. Alternatively, the sound pickup device 18 may be positioned on a microphone boom arm (not shown) instead of the side of the temple 94.
[0028] Figure 4 is a block diagram showing an example configuration of the MR glasses 10[k] shown in Figure 1. The MR glasses 10[k] comprises a processing unit 11, a storage device 12, a gaze acquisition device 13, a motion detection device 14, a communication device 15, a display device 16, a sound output device 17, and a sound pickup device 18. Each element of the MR glasses 10[k] is interconnected by one or more buses for communicating information. In this specification, the term "device" may be replaced with other terms such as circuit, device, unit, etc.
[0029] The processing unit 11 is a processor that controls the entire MR glasses 10[k] and is configured, for example, using one or more chips. The processing unit 11 is configured, for example, using a central processing unit (CPU) that includes interfaces with peripheral devices, an arithmetic unit, and registers. Some or all of the functions of the processing unit 11 may be implemented by hardware such as a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), or FPGA (Field Programmable Gate Array). The processing unit 11 executes various processes in parallel or sequentially.
[0030] The storage device 12 is a recording medium that can be read from and written to by the processing device 21. The storage device 12 includes, for example, non-volatile memory and volatile memory. Non-volatile memory is, for example, ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), and EEPROM (Electrically Erasable Programmable Read Only Memory). Volatile memory is, for example, RAM (Random Access Memory).
[0031] The storage device 12 stores multiple programs, including the control program PR1, which is executed by the processing unit 11. The storage device 12 also functions as a work area for the processing unit 11.
[0032] The gaze acquisition device 13 acquires the direction in which user U[k] is looking, i.e., the direction of user U[k]'s gaze, by tracking the movement of user U[k]'s left and right eyeballs, and outputs gaze direction information indicating the direction of user U[k]'s gaze to the processing device 11 based on the acquisition results.
[0033] More specifically, the gaze acquisition device 13 includes a pair of light sources and a pair of cameras. Each of the light sources corresponds to the left and right eyes of user U[k], and each of the cameras corresponds to the left and right eyes of user U[k]. The light source corresponding to user U's left eye illuminates user U[k]'s left eye with infrared light, and the camera corresponding to user U[k]'s left eye captures a corneal reflection image of user U[k]'s left eye as an image formed by the reflection of the illuminated infrared light. Similarly, the light source corresponding to user U[k]'s right eye illuminates user U[k]'s right eye with infrared light, and the camera corresponding to user U[k]'s right eye captures a corneal reflection image of user U's right eye.
[0034] The gaze acquisition device 13 detects the position of the inner corner of the user U[k]'s eye and the position of the iris from the acquired corneal reflection image of the left eye and the corneal reflection image of the right eye, and acquires the gaze direction information of user U[k] based on the detected position of the inner corner of the eye and the position of the iris. Note that the method of acquiring the gaze by the gaze acquisition device 13 is not limited to the method described above, and any method may be used.
[0035] The motion detection device 14 detects the movement of the MR glasses 10[k] and outputs motion information to the processing device 11. The motion information includes acceleration information indicating the acceleration in the X, Y, and Z axes, and angular acceleration information indicating the angular acceleration with the X, Y, and Z axes as the centers of rotation. The motion detection device 14 includes an acceleration sensor for detecting acceleration, an inertial sensor such as a gyro sensor for detecting angular acceleration, and a geomagnetic sensor for detecting the direction in which the MR glasses 10[k] are facing.
[0036] The accelerometer detects acceleration in the mutually orthogonal X, Y, and Z axes. The gyroscope detects angular acceleration with the X, Y, and Z axes as the center of rotation. The geomagnetic sensor detects the Earth's magnetic field in the X, Y, and Z axes to determine the direction the MR glasses 10 are facing.
[0037] The communication device 15 is hardware that acts as a transmitting and receiving device for communicating with other devices. The communication device 15 is also called, for example, a network device, network controller, network card, or communication module. The communication device 15 may be equipped with a connector for wired connection and an interface circuit corresponding to the connector. The communication device 15 may also be equipped with a wireless communication interface. Examples of connectors and interface circuits for wired connection include products compliant with wired LAN, IEEE1394, USB, etc. Examples of wireless communication interfaces include products compliant with wireless LAN, Bluetooth®, etc.
[0038] The display device 16 is a device that displays images. The display device 16 displays various images based on control by the processing device 11. The display device 16 includes a display panel for the left eye and a display panel for the right eye.
[0039] The sound output device 17 is a device that outputs sound. The sound output device 17 outputs various sounds based on control by the processing device 11. The sound output device 17 includes a speaker for the left ear and at least two speakers for the right ear.
[0040] The sound acquisition device 18 is a device that converts the voice of user U[k] and external sounds into electrical signals and outputs them as voice information to the processing device 11. The sound acquisition device 18 includes at least one microphone.
[0041] The processing unit 11 functions as an acquisition unit 111, a transmission unit 112, and a display control unit 113, for example, by reading and executing the control program PR1 from the storage device 12.
[0042] The acquisition unit 111 acquires user U[k]'s gaze direction information input from the gaze acquisition device 13, user U[k]'s movement information input from the motion detection device 14, and audio information input from the sound acquisition device 18.
[0043] The acquisition unit 111 acquires display information from the terminal device 20[k] via the communication device 15. The display information includes image information rendered by the terminal device 20[k] for the virtual object VO placed in the virtual space VS.
[0044] The transmitting unit 112 transmits the acquired user U[k]'s gaze direction information, user U[k]'s movement information, and voice information to the terminal device 20[k] via the communication device 15.
[0045] The display control unit 113 controls the display on the display device 16. The display control unit 113 displays the display image rendered on the terminal device 20[k] on the display device 16. The range of the image displayed on the display device 16 is a portion of the virtual space VS.
[0046] 1.1.3. Configuration of the terminal device Figure 5 is a block diagram showing an example configuration of the terminal device 20[k] in Figure 1. As shown in Figure 5, the terminal device 20[k] comprises a processing unit 21, a storage device 22, a communication device 23, a display device 24, and an input device 25. Each element of the terminal device 20[k] is interconnected by one or more buses for communicating information.
[0047] The processing unit 21 is a processor that controls the entire terminal device 20[k], and is configured, for example, using one or more chips. The processing unit 21 is configured, for example, using a central processing unit (CPU) that includes interfaces with peripheral devices, arithmetic units, registers, etc. Some or all of the functions of the processing unit 21 may be implemented by hardware such as a DSP, ASIC, PLD, FPGA, etc. The processing unit 21 executes various processes in parallel or sequentially.
[0048] The storage device 22 is a recording medium that can be read from and written to by the processing device 21. The storage device 22 includes, for example, non-volatile memory and volatile memory. Non-volatile memory is, for example, ROM, EPROM, and EEPROM. Volatile memory is, for example, RAM.
[0049] The storage device 22 stores multiple programs, including the control program PR2, which is executed by the processing unit 21. The storage device 22 also functions as the work area for the processing unit 21. The control program PR2 is a program that controls the entire processing unit 21.
[0050] The communication device 23 is hardware that acts as a transmitting and receiving device for communicating with other devices. The communication device 23 is also called, for example, a network device, network controller, network card, or communication module. The communication device 23 may be equipped with a connector for wired connection and an interface circuit corresponding to the connector. The communication device 23 may also be equipped with a wireless communication interface. Examples of connectors and interface circuits for wired connection include products compliant with wired LAN, IEEE1394, and USB. Examples of wireless communication interfaces include products compliant with wireless LAN and Bluetooth®.
[0051] The display device 24 is a device that displays images and text information. The display device 24 displays various images based on control by the processing device 21. For example, various display panels such as liquid crystal panels and organic EL panels are preferably used as the display device 24.
[0052] The input device 25 accepts input from user U. For example, the input device 25 includes a pointing device such as a keyboard, touchpad, touch panel, or mouse. If the input device 25 includes a touch panel, it may also function as the display device 24.
[0053] The processing unit 21 functions as an acquisition unit 211, an image processing unit 212, and a transmission unit 213 by, for example, reading and executing the control program PR2 from the storage device 22.
[0054] The acquisition unit 211 acquires gaze direction information, motion information, and audio information from the MR glasses 10[k]. The acquisition unit 211 also acquires image information of the entire virtual space VS from the server 30.
[0055] The image processing unit 212 renders the image included in the image information from the server 30 based on the motion information from the MR glasses 10[k] and generates a display image to be shown on the display device 16 of the MR glasses 10[k]. The rendering process includes cutting out a portion of the image displayed in the virtual space VS in accordance with the movement of the MR glasses 10[k].
[0056] If information regarding the presentation materials is stored in the storage device 22, the transmitting unit 213 transmits the information regarding the presentation materials to the server 30 via the communication device 23. The transmitting unit 213 also transmits gaze direction information, motion information, and audio information acquired from the MR glasses 10[k] to the server 30 via the communication device 23. The information regarding the presentation materials may be stored in an external server or the like (not shown).
[0057] Furthermore, the transmitting unit 213 transmits display image information related to the display image generated by the image processing unit 212 to the MR glasses 10[k] via the communication device 23. The transmitting unit 213 also transmits processed audio information, in which the sound image position of the sound included in the audio information is controlled by the server 30, to the MR glasses 10[k] via the communication device 23. However, if user U[1] is the presenter, the transmitting unit 213 does not need to transmit the processed audio information to the MR glasses 10[1] used by user U[1].
[0058] 1.1.4. Server Configuration Figure 6 is a block diagram showing an example configuration of the server 30 in Figure 1. The server 30 comprises a processing unit 31, a storage device 32, and a communication device 33. Each element of the server 30 is interconnected by one or more buses for communicating information. The server 30 is an example of a sound control device.
[0059] The processing unit 31 is a processor that controls the entire server 30 and is configured, for example, using one or more chips. The processing unit 31 is configured using a central processing unit (CPU) that includes interfaces with peripheral devices, arithmetic units, registers, etc. Some or all of the functions of the processing unit 31 may be implemented by hardware such as a DSP, ASIC, PLD, FPGA, etc. The processing unit 31 executes various processes in parallel or sequentially.
[0060] The storage device 32 is a recording medium that can be read from and written to by the processing device 31. The storage device 32 includes, for example, non-volatile memory and volatile memory. Non-volatile memory is, for example, ROM, EPROM, and EEPROM. Volatile memory is, for example, RAM.
[0061] The storage device 32 stores multiple programs, including the control program PR3, which is executed by the processing unit 31. The storage device 32 also functions as a work area for the processing unit 31. The control program PR3 is a program that controls the entire processing unit 31.
[0062] The communication device 33 is hardware that acts as a transmitting and receiving device for communicating with other devices. The communication device 33 is also called, for example, a network device, network controller, network card, or communication module. The communication device 33 may be equipped with a connector for wired connection and an interface circuit corresponding to the connector. The communication device 33 may also be equipped with a wireless communication interface. Examples of connectors and interface circuits for wired connection include products compliant with wired LAN, IEEE1394, and USB. Examples of wireless communication interfaces include products compliant with wireless LAN and Bluetooth®.
[0063] The processing unit 31 functions as an acquisition unit 311, a specification unit 312, a sound control unit 313, and a display control unit 314, for example, by reading and executing the control program PR3 from the storage device 32.
[0064] In the following explanation, user U[1] is the presenter giving a presentation in the virtual space VS, and the other users U[2]~U[n] are viewers watching user U[1]'s presentation. In the following explanation, user U[2] will represent the viewers, and user U[1] will be referred to as the first user U[1], and user U[2] as the second user U[2]. In addition, the MR glasses 10[1] worn on the head of user U[1] will be referred to as the first glasses 10[1], and the MR glasses 10[2] worn on the head of user U[2] will be referred to as the second glasses 10[2].
[0065] In this case, the first user U[1] is granted the authority to place their presentation materials in the virtual space VS. The granting of the authority to place presentation materials in the virtual space VS may be decided before the start of the meeting or during the meeting. The first glasses 10[1] are an example of a first wearable display device, and the second glasses 10[2] are an example of a second wearable device. However, designating user U[1] as the presenter and the other users U[2]~U[n] as viewers is for the sake of simplicity and does not limit the present invention. Any of the users U may be the presenter.
[0066] The acquisition unit 311 acquires gaze information and audio information from the terminal device 20[1] used by the presenter, the first user U[1]. The gaze information is information relating to the gaze of the first user U[1]. The gaze information includes gaze direction information acquired by the gaze acquisition device 13 of the first glasses 10[1] and motion information detected by the motion detection device 14 of the MR glasses 10[1]. The audio information is information relating to the speech of the first user U[1]. The audio information is information output from the sound pickup device 18 of the MR glasses 10[1].
[0067] The method for identifying the line of sight of the first user U[1] by the identification unit 312 will be explained below with reference to Figure 7. Figure 7 is a schematic diagram showing the relationship between the virtual space VS and the display area of the MR glasses 10[1]. As shown in Figure 7, the first object VO1 is placed in the virtual space VS inside the celestial sphere CS, and the first user U[1] is looking at the first object VO1. The first object VO1 is located to the right of the front of the first user U[1]. When the first user U[1] rotates their head to the right to look at the first object VO1, the display area DA1 of the display device 16 of the MR glasses 10[1] becomes to the right of the front of the first user U[1]. The display area DA1 can also be said to be the field of view of the first user U[1]. The line of sight of the first user U[1] is indicated by the first line of sight LV1.
[0068] The identification unit 312 identifies the display area DA1 based on motion information acquired from the MR glasses 10[1]. Furthermore, the identification unit 312 identifies the first gaze direction LV1 of the first user U[1] based on gaze direction information acquired from the MR glasses 10[1]. That is, the first gaze direction LV1 is determined by the amount of rotation of the first user U[1]'s head and the amount of rotation of the first user U[1]'s eyeballs. Since gaze direction identification can be achieved using well-known head tracking and eye tracking technologies, a detailed explanation is omitted.
[0069] This identifies the gaze of the first user U[1] in the virtual space VS. Whether or not the first user U[1] is fixating on a certain virtual object can be determined, for example, by understanding the eye movements of the first user U[1] from the gaze direction information. Human eye movements are broadly classified into retention and saccades. Retention is an eye movement in which a person keeps an object fixed in the fovea in order to obtain detailed information of interest. Saccades are eye movements in which a person rapidly moves the fovea from one place of interest to the next place of interest.
[0070] Therefore, if the eye movement is in a stationary state, the first user U[1] may be fixating on a virtual object VO, and if the eye movement is in a saccadic state, the first user U[1] is not seeing anything in particular. For example, whether the eye movement is in a stationary or saccadic state is determined based on the angular velocity of the eye.
[0071] Next, with reference to Figure 8, the method by which the identification unit 312 identifies the gaze position of the first user U[1] will be described. Figure 8 is a schematic diagram showing the relationship between the gaze of the first user U[1] and the first object VO1. As shown in Figure 8, the virtual space VS contains the first object VO1 and one or more second objects VO2. The first object VO1 is one of the virtual objects VO, and is the object that the presenter, the first user U[1], is describing. The one or more second objects VO2 are virtual objects VO separate from the first object VO1.
[0072] Figure 8 shows the first region AR1 along with the first object VO1 and one or more second objects VO2. The first region AR1 is a region that includes the first object VO1 but does not include the one or more second objects VO2. In this embodiment, the first region AR1 is a circular region, but the shape of the first region AR1 is not particularly limited to a circle. The shape of the first region AR1 may be an ellipse or a polygon.
[0073] The center of the first region AR1 is denoted by the point Pa. The coordinates of point Pa are (Xa, Ya, Za). The point O where the first user U[1] is located is the origin of the coordinate system (0, 0, 0).
[0074] The first line of sight LV1 of the first user U[1] intersects with the first object VO1 at intersection P1. Intersection P1 is a point on the first object VO1. The coordinates of intersection P1 are (X1, Y1, Z1).
[0075] The second line of sight LV2 of the first user U[1] intersects with the first region AR1 on the celestial sphere SS at intersection P2. The first region AR1 is the region containing the first object VO1. Intersection P2 is a point on the first region AR1 excluding the display range of the first object VO1. The coordinates of intersection P2 are (X2, Y2, Z2).
[0076] The third line of sight LV3 of the first user U[1] intersects the celestial sphere SS at intersection P3. Intersection P3 is a point outside the first region AR1. The coordinates of intersection P3 are (X3, Y3, Z3).
[0077] If the line of sight of the first user U[1] is the third line of sight LV3, then it cannot be said that the first user U[1] is looking at the first object VO1, and therefore the identification unit 312 does not identify point P3 as the gaze position.
[0078] Normally, the gaze of the first user U[1] is constantly changing. Therefore, even if the gaze of the first user U[1] intersects with the first object VO1, as with the first gaze LV1, if the time of intersection is temporary, it cannot be said that the first user U[1] is fixating on the first object VO1.
[0079] When the line of sight of the first user U[1] transitions from a state where it does not intersect with the first object VO1 to a state where it intersects with the first object VO1, if the line of sight of the first user U[1] remains on the first object VO1 for at least one hour after it begins to intersect with the first object VO1, the identification unit 312 identifies the position of the first object VO1 as the gaze position. In other words, the identification unit 312 identifies the position of the first object VO1 as the gaze position if the line of sight of the first user U[1] intersects with the first object VO1 for at least one hour. The first hour is not particularly limited, but for example, it is preferably about 1 to 3 seconds.
[0080] According to this embodiment, even if the line of sight of the first user U[1], who is looking at a virtual object other than the first object VO1, changes rapidly and intersects with the first object VO1, if the line of sight of the first user U[1] deviates from the first object VO1 before the first time elapses, the identification unit 312 is not likely to mistakenly identify that the first user U[1] is looking at the first object VO1.
[0081] Note that "position of the first object VO1" refers to "any point on the first object VO1". The position of the first object VO1 may be the intersection point P1 with the line of sight of the first user U[1], the center of the first object VO1, or the centroid of the first object VO1. Alternatively, the position of the first object VO1 may be the position closest to the coordinate origin on the first object VO1, or the position furthest from the coordinate origin on the first object VO1.
[0082] Furthermore, if the gaze of the first user U[1] transitions from a state where it intersects with the first object VO1 to a state where it does not intersect with the first object VO1, the identification unit 312 stops identifying the position of the first object VO1 as the gaze position if two hours or more have elapsed since the gaze of the first user U[1] ceased to intersect with the first object VO1. In other words, the identification unit 312 continues to identify the position of the first object VO1 as the gaze position until two hours have elapsed, even after the gaze of the first user U[1] ceases to intersect with the first object VO1. The second hours are not particularly limited, but are preferably, for example, 1 to 3 seconds.
[0083] According to this embodiment, even if the gaze of the first user U[1] changes rapidly and deviates from the first object VO1, if the gaze of the first user U[1] intersects with the first object VO1 again before the second time has elapsed, the identification unit 312 can determine that the first user U[1] is continuing to gaze at the first object VO1.
[0084] Referring again to Figure 6, the sound control unit 313 localizes the sound image position of the first user U[1]'s speech based on the speech information near the gaze position of the first user U[1] identified by the identification unit 312.
[0085] Controlling the sound image position can be described as intentionally changing the sound image localization. Sound image localization is generally obtained by reproducing the head-related transfer function (HRTF), which takes into account the time difference and intensity difference of signals reaching the left and right ears from the sound source, the changes in the frequency characteristics of sound waves caused by diffraction in the head and auricle, and reflections from surrounding walls. In very simple terms, by giving the left and right sound output devices 17L and 18R a time difference and intensity difference, the sound image localization is perceived as changing by the second user U[2]. When the sound image localization is accurately reproduced, the second user U[2] can experience three-dimensional sound.
[0086] Furthermore, AudioSion EP, a technology that achieves three-dimensional sound through proprietary digital signal processing without relying on HRTF reproduction, is known (https: / / av.watch.impress.co.jp / docs / series / dal / 1309076.html). By employing technologies such as AudioSion EP, which reproduces HRTF, it is possible to output three-dimensional sound using the sound output device 17.
[0087] Ideally, the gaze position of the first user U[1] and the sound image position of the first user U[1]'s speech should coincide, but it is sufficient if the sound image position of the first user U[1]'s speech is localized in the vicinity of the gaze position of the first user U[1]. Furthermore, if the virtual space VS is divided into a first space and a second space separate from the first space, it is sufficient if the sound image positions are localized so that the gaze position of the first user U[1] and the sound image position of the first user U[1]'s speech belong to at least the same space, either the first space or the second space. In other words, the sound image position may be controlled by a simple method, without relying on HRTF-based stereophonic sound, AudioSion EP-based stereophonic sound, etc.
[0088] For example, the first space is defined as the space to the left of the virtual space VS with respect to the second user U[2], and the second space is defined as the space to the right of the virtual space VS with respect to the second user U[2]. If the first object VO1 is located to the right of the second user U[2], the sound image position should be localized to the right of the second user U[2], and if the first object VO1 is located to the left of the second user U[2], the sound image position should be localized to the left of the second user U[2]. This makes it possible to at least indicate to the second user U[2] whether the first object VO1 is to the left or right from their perspective. According to this embodiment, the second user U[2] can determine at least whether to direct their gaze to the left or right.
[0089] If the gaze of the first user U[1] does not intersect with any virtual object VO in the virtual space VS, the identification unit 312 cannot determine the gaze position of the first user U[1]. In this case, the sound control unit 313 may, for example, make it so that the speech of the first user U[1] can be heard from in front of the second user U[2].
[0090] The display control unit 314 displays multiple virtual objects VO in the virtual space VS. The multiple virtual objects VO include presentation materials prepared by the first user U[1].
[0091] The following describes an example of how the first user U[1] places presentation materials in the virtual space VS, referring to Figures 9-11. Figure 9 is a schematic diagram showing an example of how the first user U[1] places presentation materials in the virtual space VS. Figure 9 is a view of the virtual space VS from the zenith ZE, and the outer perimeter of the virtual space VS corresponds to the equatorial EQ of the celestial sphere CS. The plane containing the equatorial EQ is parallel to the XY plane and is the plane with a Z coordinate of 0. The head of the first user U[1] is located at the origin O of the virtual space VS.
[0092] As shown in Figure 9, multiple virtual object VOs, including a first object VO1 and one or more second objects VO2, are arranged along the equator EQ. Each of the multiple virtual object VOs is a slide that makes up a presentation document. The first user U[1] is explaining while looking at the first object VO1.
[0093] Figure 10 is a schematic diagram showing an example of a virtual space VS shared by the second user U[2]. Similar to Figure 9, Figure 10 is a view of the virtual space VS from the zenith ZE, and the outer perimeter of the virtual space VS corresponds to the equator EQ of the celestial sphere CS. The head of the second user U[2] is located at the origin O of the virtual space VS.
[0094] As shown in Figure 10, multiple virtual objects VO, including the first object VO1, are arranged along the equator EQ. That is, the virtual space VS experienced by the first user U[1] and the virtual space VS experienced by the second user U[2] are the same, and therefore, the first user U[1] and the second user U[2] can share the virtual space VS with each other. In the state shown in Figure 10, the second user U[2] has not yet grasped that the first user U[1] is describing the first object VO1. In Figure 10, for convenience, the virtual space VS is divided into the first space SP1 on the left and the second space SP2 on the right, with the second user U[2] as the reference point.
[0095] Figure 11 is a schematic diagram showing an example of a case where the sound image position is localized near the gaze position of the first user U[1] in the virtual space VS shared by the second user U[2]. Since the first user U[1] is gazing at the first object VO1, the identification unit 312 identifies the intersection point P1, which is the position of the first object VO1, as the gaze position of the first user U[1]. The space to which the identified gaze position belongs is the second space SP2. The sound control unit 313 localizes the sound image position SI1 near the identified gaze position, i.e., the intersection point P1. In other words, the sound control unit 313 localizes the sound image position SI1 in the second space SP2. As a result, it is expected that the second user U[2] will direct their gaze in the direction of the sound image position SI1. When the second user U[2] directs their gaze in the direction of the sound image position SI1, the gaze of the second user U[2] intersects with the first object VO1.
[0096] 1.2. Operation of the control device according to the first embodiment 1.2.1. Operation of the Processing Unit Figure 12 is a flowchart showing an example of the sound control operation of the processing unit 31 in Figure 6. The sound control operation of the processing unit 31 will be explained below with reference to Figure 12.
[0097] In step S11, the processing unit 31 functions as an acquisition unit 311 to acquire gaze information related to the gaze of the first user U[1] and audio information related to the speech of the first user U[1] from the MR glasses 10[1].
[0098] In step S12, the processing unit 31, by functioning as a specific unit 312, identifies the direction of the gaze of the first user U[1] in the virtual space VS shared by the first user U[1] and the second user U[2], based on the gaze information acquired by the acquisition unit 311.
[0099] In step S13, the processing unit 31, by functioning as a specific unit 312, determines whether the line of sight of the first user U[1] has intersected with the first object VO1 for a period of time or longer.
[0100] If it is determined that the gaze of the first user U[1] has intersected with the first object VO1 for more than one hour, that is, if the determination result in step S13 is positive, the processing unit 31, in step S14, functions as a identifying unit 312 to identify the position of the first object VO1 as the gaze position of the first user U[1].
[0101] In step S15, the processing unit 31 functions as a sound control unit 313 to localize the sound image position of the first user U[1]'s speech near the gaze position of the identified first user U[1]. Specific methods for localizing the sound image position are well known.
[0102] In step S16, the processing unit 31 functions as a sound control unit 313 and outputs a sound to the terminal device 20[2] of the second user U[2], with the sound image position localized near the gaze position of the first user U[1], and then terminates this routine.
[0103] On the other hand, if it is determined that the gaze of the first user U[1] has not intersected with the first object VO1 for more than one hour, that is, if the determination result in step S13 is negative, the processing unit 31 functions as a sound control unit 313 in step S17 to output sound at the previous sound image position and terminate this routine.
[0104] 1.3. Effects of the First Embodiment As described above, the sound control device according to the first embodiment comprises an acquisition unit 311, a identification unit 312, and a sound control unit 313. The acquisition unit 311 acquires gaze information relating to the gaze of the first user U[1] and speech information relating to the speech of the first user U[1]. Based on the gaze information, the identification unit 312 identifies the gaze position of the first user U[1] in the virtual space VS shared by the first user U[1] and the second user U[2]. The sound control unit 313 localizes the sound image position of the speech of the first user U[1] based on the speech information in the space to which the identified gaze position belongs among the first space SP1 and the second space SP2, which is separate from the first space SP1, both included in the virtual space VS.
[0105] In this embodiment, the gaze position of the presenter, first user U[1], and the sound image position of the first user U[1]'s utterance belong to the same space among the first space SP1 and the second space SP2 contained in the virtual space VS. Since it is assumed that the first user U[1] will speak while gazing at the first object VO1 that is the subject of explanation, the second user U[2] can easily find the first object VO1 by looking in the direction from which the first user U[1]'s utterance is coming. Therefore, it becomes easier for the audience to understand which material in the virtual space VS the presenter is explaining.
[0106] Furthermore, the first space SP1 is the space to the left of the virtual space VS with respect to the second user U[2], and the second space SP2 is the space to the right of the virtual space VS with respect to the second user U[2].
[0107] In this embodiment, the gaze position of the presenter, first user U[1], and the sound image position of the first user U[1]'s speech belong to the same space, which is divided between the first space SP1 and the second space SP2, with respect to the viewer, second user U[2]. Therefore, if the first user U[2] hears the first user U[1]'s speech from the left, second user U[2] can easily find the first object VO1 by looking to the left. Similarly, if the first user U[2] hears the first user U[1]'s speech from the right, second user U[2] can easily find the first object VO1 by looking to the right.
[0108] Furthermore, the identification unit 312 identifies the position of the first object VO1 as the gaze position when the line of sight LV of the first user U[1] intersects with the first object VO1 in the virtual space VS for one hour or longer.
[0109] According to this embodiment, if the time during which the line of sight LV of the first user U[1] intersects with the first object VO1 is less than 1 time, the position of the first object VO1 is not identified as the position of gaze of the first user U[1]. In other words, if the line of sight of the first user U[1] temporarily moves and intersects with the first object VO1, it is prevented from being mistakenly identified as gaze at the first object VO1.
[0110] Furthermore, the sound control unit 313 localizes the sound image position near the gaze position of the identified first user U[1].
[0111] According to this embodiment, the sound image is localized near the gaze position of the identified first user U[1], so the second user U[2] can more easily find the first object VO1 by looking in the direction from which the speech of the first user U[1] is coming.
[0112] Furthermore, the sound control system according to the first embodiment includes MR glasses 10[1], MR glasses 10[2], and a processing unit 31. MR glasses 10[1] are worn on the head of the first user U[1]. MR glasses 10[2] are worn on the head of the second user U[2]. The server 30 performs acquisition processing, identification processing, and sound control processing. Acquisition processing is the process of acquiring gaze information related to the gaze of the first user U[1] and audio information related to the speech of the first user U[1]. Identification processing is the process of identifying the gaze position of the first user U[1] in the virtual space VS shared by the first user U[1] and the second user U[2], based on the gaze information. The sound control process is the process of localizing the sound image position of the first user U[1]'s speech, based on speech information, into the space to which the identified gaze position belongs, which is included in the first space SP1 and the second space SP2, which is separate from the first space SP1, both contained in the virtual space VS.
[0113] In this embodiment, the gaze position of the presenter, first user U[1], and the sound image position of the first user U[1]'s utterance belong to the same space among the first space SP1 and the second space SP2 contained in the virtual space VS. Since it is assumed that the first user U[1] will speak while gazing at the first object VO1 that is the subject of explanation, the second user U[2] can easily find the first object VO1 by looking in the direction from which the first user U[1]'s utterance is coming. Therefore, it becomes easier for the audience to understand which material in the virtual space VS the presenter is explaining.
[0114] Furthermore, the sound control method according to the first embodiment acquires gaze information relating to the gaze of the first user U[1] and speech information relating to the speech of the first user U[1], identifies the gaze position of the first user U[1] in the virtual space VS shared by the first user U[1] and the second user U[2] based on the gaze information, and localizes the sound image position of the speech of the first user U[1] based on the speech information in the space to which the identified gaze position belongs among the first space SP1 and the second space SP2, which is separate from the first space SP1, both included in the virtual space VS.
[0115] In this embodiment, the gaze position of the presenter, first user U[1], and the sound image position of the first user U[1]'s utterance belong to the same space, either the first space SP1 or the second space SP2, which are included in the virtual space VS. Since it is assumed that the first user U[1] will speak while gazing at the first object VO1, which is the subject of explanation, the second user U[2] can easily find the first object VO1 by looking in the direction from which the first user U[1]'s utterance is coming. Therefore, it becomes easier for the audience to understand which material in the virtual space VS the presenter is explaining.
[0116] 2. Variations This disclosure is not limited to the embodiments illustrated above. Specific variations are illustrated below. Two or more embodiments may be arbitrarily selected from the following examples and combined. Furthermore, the embodiments described above and the variations described below can be combined arbitrarily as long as they do not contradict each other.
[0117] 2.1. Variation 1 In the first embodiment, it is determined that the first user U[1] is fixating on the first object VO1 if the line of sight of the first user U[1] intersects with the first object VO1 for one hour or longer. However, it is also possible to determine that the first user U[1] is fixating on the first object VO1 if the line of sight of the first user U[1] intersects with the first region AR1 for one hour or longer. The first region AR1, as shown in Figure 8, is the region surrounding the first object VO1, including the first object VO1, and does not include one or more second objects VO2.
[0118] More specifically, if the gaze of the first user U[1] transitions from a state where it does not intersect with the first region AR1 to a state where it intersects with the first region AR1, the identification unit 312 may identify the position of the first region AR1 as the gaze position if the gaze of the first user U[1] remains on the first region AR1 for one hour or more after it begins to intersect with the first region AR1. In other words, the identification unit 312 identifies the position of the first object VO1 as the gaze position if the gaze of the first user U[1] intersects with the first region AR1 containing the first object VO1 in the virtual space VS for one hour or more.
[0119] The smaller the display range of the first object VO1, the more difficult it is for the first user U[1] to continue to gaze at the first object VO1. According to this embodiment, by setting a first region AR1 that includes the first object VO1, it becomes easier for the first user U[1] to continue to gaze at the first object VO1 regardless of the display range of the first object VO1. Furthermore, by making the shape of the first region AR1 circular, the first user U[1] can continue to gaze at the first object VO1 more stably.
[0120] Furthermore, if the gaze of the first user U[1] transitions from a state where it intersects with the first region AR1 to a state where it does not intersect with the first region AR1, the identification unit 312 will cease identifying the position of the first region AR1 as the identified position if two hours or more have elapsed since the gaze of the first user U[1] ceased to intersect with the first region AR1. In other words, the identification unit 312 will continue to identify the position of the first region AR1 as the gaze position until two hours have elapsed, even after the gaze of the first user U[1] ceases to intersect with the first region AR1.
[0121] According to this embodiment, even if the gaze of the first user U[1] changes rapidly and deviates from the first region AR1, if the gaze of the first user U[1] intersects with the first region AR1 again before the second time has elapsed, the identification unit 312 can determine that the first user U[1] is continuing to gaze at the first region AR1.
[0122] 2.2. Variation Example 2 In the first embodiment or the above modification 1, it is determined that the first user U[1] is gazing at the first object VO1 or the first region AR1 if the gaze of the first user U[1] remains within the range of the first object VO1 or the first region AR1 for one hour or more. However, the method for determining whether or not the first user U[1] is gazing is not limited to this.
[0123] For example, if the line of sight of the first user U[1] intersects with the first object VO1 or the first region AR1, and it is determined that the line of sight of the first user U[1] is stationary, the identification unit 312 may consider that the first user U[1] is fixating on the first object VO1, and identify the position of the first object VO1 as the fixation position. Whether or not the line of sight is stationary can be determined in the same way as whether or not eye movement is stationary. That is, whether or not the line of sight is stationary can be determined by whether or not the angular velocity of the change in the line of sight is below a threshold.
[0124] 2.3. Variation 3 In the example above, the gaze position is determined in a situation where it is assumed that the gaze of the first user U[1] is fixed on the first object VO1 or the first region AR1, but the method of determining the gaze position is not limited to this. For example, even if the gaze of the first user U[1] is a third gaze LV3, and the intersection point P3 does not intersect with either the first object VO1 or any other virtual object, the gaze position of the first user U[1] may still be determined.
[0125] In this case, the identification unit 312 may, for example, determine that the gaze of the first user U[1] is stationary, and identify the intersection point of the gaze of the first user U[1] and the virtual space VS as the gaze position. That is, in this embodiment, the identification unit 312 may identify the stationary position of the gaze of the first user U[1] as the gaze position, regardless of whether or not there is a virtual object VO including the first object VO1, if the gaze of the first user U[1] is stationary. The stationary position of the gaze may be the position of the intersection point P3 at the time it is determined that the gaze is stationary.
[0126] Furthermore, if the duration of the gaze retention exceeds the third hour, the specific unit 312 may identify the gaze retention position as the fixation position.
[0127] 2.4. Variation 4 In the first embodiment and each of its variations described above, the presenter was described as a single first user U[1], but there may be two or more presenters.
[0128] 2.5. Variation 5 In the first embodiment described above, the virtual space VS is shown to be divided into a first space SP1 on the left and a second space SP2 on the right, with respect to the second user U[2], but the method of division is not limited to this example. For example, the virtual space VS may be divided into two spaces, an upper space and a lower space, or it may be divided into two spaces, one in front of the second user U[2] and one other than the front.
[0129] 2.6. Variation 6 In the first embodiment, the second user U[2] can hear the speech of the first user U[1] using the sound output device 17 provided on the MR glasses 10[2]. However, the second user U[2] may also use earphones worn separately on their ears or speakers placed in the real world to hear the speech of the first user U[1].
[0130] 2.7. Example 7 In the first embodiment, the MR glasses 10[k] are connected to the communication network NET via the terminal device 20[k], but the MR glasses 10[k] may be connected directly to the communication network NET without going through the terminal device 20[k]. Also, in the first embodiment, the display image shown on the display device 16 of the MR glasses 10[k] is rendered by the terminal device 20[k], but the display image may be rendered by the MR glasses 10[k]. In this case, the terminal device 20[k] can be omitted. The display image may also be rendered by the server 30. In this case as well, the terminal device 20[k] can be omitted.
[0131] 3. Others (1) In the embodiments described above, the storage devices 12, 22, and 32 are exemplified by ROM and RAM, but can also be flexible disks, magneto-optical disks (e.g., compact disks, digital multipurpose disks, Blu-ray® disks), smart cards, flash memory devices (e.g., cards, sticks, key drives), CD-ROMs (Compact Disc-ROMs), registers, removable disks, hard disks, floppy® disks, magnetic strips, databases, servers, and other suitable storage media. The program may also be transmitted from a network via a telecommunications line. The program may also be transmitted from a communication network NET via a telecommunications line.
[0132] (2) In the embodiments described above, the information, signals, etc. may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be mentioned throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0133] (3) In the embodiments described above, the input and output information may be stored in a specific location (e.g., memory) or managed using a management table. The input and output information may be overwritten, updated, or appended to. The output information may be deleted. The input information may be transmitted to other devices.
[0134] (4) In the embodiments described above, the determination may be made by a value represented using 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).
[0135] (5) The processing procedures, sequences, flowcharts, etc., exemplified in the embodiments described above may be rearranged in order, as long as they do not contradict each other. For example, the methods described in this disclosure present various step elements using an exemplary order and are not limited to the specific order presented.
[0136] (6) Each function illustrated in Figures 1 to 12 is implemented by any combination of at least one of hardware and software. Furthermore, the method of implementing each function block is not particularly limited. That is, each function block may be implemented using one device that is physically or logically coupled, or it may be implemented using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A function block may also be implemented by combining the above one device or the above multiple devices with software.
[0137] (7) The programs illustrated in the embodiments described above should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether they are called software, firmware, middleware, microcode, hardware description languages or by other names.
[0138] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0139] (8) In each of the above-mentioned forms, the terms “system” and “network” shall be used interchangeably.
[0140] (9) The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values from a given value, or other corresponding information.
[0141] (10) In the embodiments described above, the MR glasses 10 and the terminal equipment 20 may be a mobile station (MS). A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or several other appropriate terms. In this disclosure, terms such as “mobile station,” “user terminal,” “user equipment (UE),” and “terminal” may be used interchangeably.
[0142] (11) In the embodiments described above, the terms “connected,” “coupled,” or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are “connected” or “coupled” with each other. The coupling or connection between elements may be a physical coupling or connection, a logical coupling or connection, or a combination thereof. For example, “connection” may be reinterpreted as “access.” As used in this disclosure, two elements may be considered to be “connected” or “coupled” with each other using at least one of one or more wires, cables and printed electrical connections, and, in some non-limiting and non-exclusive examples, electromagnetic energy having wavelengths in the radio frequency domain, microwave domain and optical (both visible and invisible) domain.
[0143] (12) In the embodiments described above, the phrase “based on” does not mean “based solely on” unless otherwise specified. In other words, the phrase “based on” means both “based solely on” and “based at least on.”
[0144] (13) The terms “determining” and “determining” as used in this disclosure may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiry (e.g., searching in a table, database or other data structure), and ascertaining. “Determining” may also include, for example, receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."
[0145] (14) Where the terms “include,” “including,” and variations thereof are used in the embodiments described above, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to be exclusive OR.
[0146] (15) In the present disclosure, if articles are added by translation, such as a, an, and the in English, the present disclosure may include the fact that the noun following these articles is plural.
[0147] (16) In this disclosure, the term “A and B are different” may mean “A and B are different from each other.” The term may also mean “A and B are each different from C.” Terms such as “separate” and “combine” may be interpreted in the same way as “different.”
[0148] (17) Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed in practice. Furthermore, notification of certain information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).
[0149] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Accordingly, the descriptions in the present disclosure are illustrative and not restrictive in any way. [Explanation of symbols]
[0150] 1...Sound control system, 10...MR glasses, 10[1]...First glasses, 10[2]...Second glasses, 20...Terminal device, 30...Server, 311...Acquisition unit, 312...Identification unit, 313...Sound control unit, LV,LV1,LV2,LV3...Line of sight, SI1...Sound image position, SP1...First space, SP2...Second space, U...User, U[1]...First user, U[2]...Second user, VO1...First object, VS...Virtual space.
Claims
1. An acquisition unit that acquires gaze information relating to the gaze of the first user and audio information relating to the speech of the first user, Based on the aforementioned gaze information, a specification unit identifies the gaze position of the first user in the display space shared by the first user and the second user, A sound control unit that localizes the sound image position of the first user's speech based on the audio information in the space to which the identified gaze position belongs among the first space and a second space separate from the first space included in the display space, A sound control device equipped with the following features.
2. The first space is the space to the left of the display space with respect to the second user, and the second space is the space to the right of the display space with respect to the second user. The sound control device according to claim 1.
3. The identifying unit identifies the position of the first object as the gaze position when the gaze of the first user intersects with the first object in the display space for one hour or more. The sound control device according to claim 1.
4. The display space contains a first object and one or more second objects. The identifying unit identifies the position of the first object as the gaze position when the gaze of the first user intersects for one hour or more with the first region surrounding the first object, which includes the first object in the display space, and which does not include the one or more second objects. The sound control device according to claim 1.
5. The sound control unit localizes the sound image position near the identified gaze position. The sound control device according to claim 1.
6. A first wearable device to be attached to the head of the first user, A second wearable device to be worn on the head of a second user, An acquisition process for acquiring gaze information relating to the gaze of the first user and audio information relating to the speech of the first user, Based on the aforementioned gaze information, a determination process is performed to determine the gaze position of the first user in the display space shared by the first user and the second user. Sound control processing to localize the sound image position of the first user's speech based on the audio information in the space to which the identified gaze position belongs, among the first space and a second space separate from the first space, which are included in the display space; A sound control device that performs the following: Sound control system including
7. The system acquires gaze information relating to the gaze of the first user and audio information relating to the speech of the first user. Based on the aforementioned gaze information, the gaze position of the first user in the display space shared by the first user and the second user is identified. The sound image position of the first user's speech, based on the audio information, is localized in the space to which the identified gaze position belongs, among the first space and a second space separate from the first space, which are included in the display space. Sound control method.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP7127645B2