Audio augmented reality object playback device, information terminal system
The audio augmented reality device facilitates easy mapping of objects in virtual spaces by using audio input and output to position information terminals, enhancing user convenience.
Patent Information
- Application Number
- JP2024513578
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-04
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-04-04
AI Technical Summary
Users face difficulties in mapping objects in audio augmented reality due to limited vision, leading to reduced convenience.
An audio augmented reality object reproduction device that maps information terminals or applications based on audio output and input, using processors to position them in a virtual space.
Enhances user convenience by allowing easy mapping of objects in virtual spaces, improving the usability of audio augmented reality systems.
Smart Images

Figure 0007781260000001 
Figure 0007781260000002 
Figure 0007781260000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio augmented reality object reproduction device and an information terminal system using the audio augmented reality object reproduction device. [Background technology]
[0002] Conventionally, an example of an audio augmented reality object playback device has been known, which is worn on the user's head, outputs audio based on stereophonic technology from an audio output device such as a speaker, and displays various information on a display screen in front of the user's eyes.
[0003] Patent Document 1 discloses a technology related to stereophonic technology. Specifically, Patent Document 1 discloses a stereophonic signal reproduction device that generates and reproduces stereophonic signals, comprising: a first processing unit that performs a Fourier transform along an azimuth angle on a head-related transfer function measured at a first distance, then performs a conversion process from the first distance to a second distance using a Hankel function, and further performs an inverse Fourier transform with the order of the Hankel function as a variable to generate a head-related transfer function at the second distance; and a second processing unit that applies the head-related transfer function at the second distance as a filter to an input acoustic signal to generate the stereophonic signal.
[0004] The technology of Patent Document 1 is said to have the effect of suppressing quality degradation caused by discontinuities and enabling high-quality stereophonic sound reproduction, even when using a method for synthesizing HRTFs at any distance on a horizontal plane. Patent Document 1 also discloses that a stereophonic signal reproduction device with a high sense of presence can be realized on the horizontal plane, where human perception is highly accurate.
[0005] On the other hand, Patent Document 2 discloses a voice processing device comprising: a microphone array having microphone elements for at least two channels; a band division unit that divides signals from the microphone array into a plurality of frequency bands for each channel; a sound source localization unit that estimates a sound source direction from the band-divided band division signals; a sound source separation unit that enhances the band division signals for each of the estimated sound source directions; a sound source overlap determination unit that determines whether the band division signals are signals from multiple sound sources or a single sound source using information on the enhanced band division signals and the estimated sound source directions; and a sound source search unit that performs a sound source search using the signal determined to be a band division signal from a single sound source.
[0006] The technology in Patent Document 2 determines whether multiple sound sources overlap and uses only the band division signals from a single sound source for sound source localization, thereby not using band components where multiple sound sources overlap and direction information of the sound source is lost. As a result, the technology in Patent Document 2 is said to be able to determine the direction from which voice or music is coming with high accuracy. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Japanese Patent Application Publication No. 2018-64227 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-227328 Summary of the Invention [Problem to be solved by the invention]
[0008] The following method is one example of how a user may use an audio augmented reality object playback device. That is, the user maps an object (such as an information terminal or an app) as a virtual object into the virtual space of the audio augmented reality object playback device, and then operates the mapped object using the audio augmented reality object playback device. However, for example, in cases where the user's vision of the outside world is limited, the user may find it difficult to perform mapping, resulting in a lack of user convenience. Note that Patent Documents 1 and 2 described above are considered not to disclose such mapping technology.
[0009] Therefore, an object of the present invention is to provide an audio augmented reality object playback device that is designed to improve user convenience and allows easy mapping, and an information terminal system that uses the audio augmented reality object playback device. [Means for solving the problem]
[0010] According to a first aspect of the present invention, there is provided an audio augmented reality object reproduction device as described below. The audio augmented reality object reproduction device is capable of mapping an object in a virtual space. The audio augmented reality object reproduction device includes a processor. The processor maps the information terminal or an application of the information terminal as an object to a position in the virtual space corresponding to the position of the information terminal, based on audio output from and input to the information terminal.
[0011] According to a second aspect of the present invention, there is provided an information terminal system as follows. The information terminal system includes one or more information terminals and an audio augmented reality object reproduction device capable of mapping an object in a virtual space. The audio augmented reality object reproduction device includes a processor. The processor maps the information terminal or an application of the information terminal as an object to a position in the virtual space corresponding to the position of the information terminal, based on audio output from and input to the information terminal. [Effects of the Invention]
[0012] According to the present invention, an audio augmented reality object reproduction device that improves user convenience and allows easy mapping, and an information terminal system that uses the audio augmented reality object reproduction device are provided. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a block diagram used to explain an example of the configuration of a head-mounted display according to a first embodiment. FIG. [Figure 2] FIG. 2 is a diagram used to explain an example of a connection for communication with an information terminal. [Figure 3] FIG. 1 is a diagram used to explain an example of the structure of a head-mounted display. [Figure 4] FIG. 1 is a diagram used to explain an example of the structure of a head-mounted display. [Figure 5] FIG. 10 is a diagram used to explain an example of a method for a user to map an object. [Figure 6] FIG. 10 is a diagram used to explain an example of a method for a user to map an object. [Figure 7] FIG. 10 is a diagram used to explain an example of a method for a user to map an object. [Figure 8] FIG. 1 is a diagram used to explain the sound source of the sound heard during mapping. [Figure 9] FIG. 1 is a diagram used to explain the sound source of the sound heard after mapping. [Figure 10] FIG. 1 is a diagram used to explain the relationship between a virtual sound source and stereophonic sound in a virtual space. [Figure 11] FIG. 1 is a diagram used to explain the position of a virtual sound source in a local coordinate system. [Figure 12] FIG. 2 is a diagram used to explain the position of a virtual sound source in a world coordinate system. [Figure 13] 10 is a flowchart used to explain an example of a mapping process. [Figure 14] 10 is a flowchart used to explain an example of a mapping process. [Figure 15]10 is a flowchart used to explain an example of a mapping process. [Figure 16] 10 is a flowchart used to explain an example of a voice operation process. [Figure 17] 10 is a flowchart used to explain an example of a voice operation process. [Figure 18] 10 is a flowchart used to explain an example of a voice operation process. [Figure 19] FIG. 10 is a diagram used to explain an example of data input / output between a head-mounted display and an information terminal in voice operation. [Figure 20] FIG. 10 is a block diagram used to explain an example of the configuration of an audio augmented reality object playback device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] An embodiment of the present invention will be described below with reference to the drawings. Note that the content described below is one embodiment of the present invention, and does not limit application to other configurations or embodiments that can perform similar processing. The mapping technology according to the present invention can contribute to "9. Build resilient infrastructure, promote inclusive and sustainable industrialization, promote innovation and build resilient infrastructure" of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0015] First, referring to FIG. 1, an example of the configuration of a head-mounted display (sometimes referred to as HMD) will be described as an example of an audio augmented reality object playback device. The audio augmented reality object playback device is a device that can perform mapping using the audio of a target and play back the audio of the mapped target. FIG. 1 is a block diagram used to explain an example of the configuration of an HMD. According to the first embodiment, the HMD 101 can map a target in a virtual space and generate an icon of the mapped target. Then, a user can select the generated icon and operate the mapped target.
[0016] As shown in FIG. 1, the HMD 101 includes a control unit 10, a ROM 11, a RAM 12, a storage unit 13, a camera 14, a display 15, a microphone 16, a speaker 17, a button 18, and a touch sensor 19.
[0017] The control unit 10 (processor) controls the entire HMD 101 in accordance with a predetermined operation program. The control unit 10 transmits and receives various commands and data to and from each component block within the HMD 101 via a system bus, which is a data communication path. The control unit 10 may be any entity that executes predetermined processing, and may be configured, for example, by a CPU (Central Processing Unit), but may also be configured using a semiconductor device such as a GPU (Graphics Processing Unit).
[0018] The ROM 11 is configured from an appropriate storage device such as a flash ROM, and stores data such as programs related to the operation and processing of the HMD 101. The RAM 12 is a memory used when the control unit 10 executes predetermined processing. The storage unit 13 can be configured from an appropriate storage device such as a hard disk drive (HDD), and can store data.
[0019] The camera 14 is provided at an appropriate position so as to be able to acquire an image of the outside. For example, the camera 14 may be provided so as to be able to acquire information outside the range of the user's field of view.
[0020] The display 15 (display unit) is provided on the front side and displays an image. For example, an image acquired by the camera 14 may be displayed on the display 15, and a user wearing the HMD 101 may visually obtain information by viewing the image acquired by the camera 14 displayed on the display 15. As will be described in detail later, the display 15 can display icons generated by performing a mapping process, but other information (for example, information related to the output volume from the HMD 101, information acquired from the outside via wireless communication, etc.) may also be displayed on the display 15 as appropriate.
[0021] The display 15 may have any suitable structure. For example, the display 15 may be a non-transmissive type or a transmissive type. The HMD 101 may have a structure in which one display 15 is placed in front of each of the user's eyes, or a structure in which one display 15 is placed to cover both of the user's eyes.
[0022] The microphone 16 is a voice input device, and in this embodiment, is provided at an appropriate position so that the voice of the user wearing the HMD 101 can be input. The microphone 16 may be provided, for example, via a member that extends to the mouth.
[0023] The speaker 17 is an audio output device that outputs information by audio. The speaker 17 is provided at an appropriate position so that the user can hear the output audio. Note that an audio output device other than the speaker 17 may be used, and for example, headphones may be provided as the audio output device.
[0024] The HMD 101 may be configured so that the user can perform various operations such as adjusting the volume and image quality, and setting communication settings using the buttons 18 and touch sensor 19. The desired operation can be realized by pressing the button 18 corresponding to the user's desired operation, and the positions and number of the buttons 18 can be set as appropriate. The touch sensor 19 is provided as appropriate so as to be able to detect the user's operation of pressing an icon or the like displayed on the display 15.
[0025] The HMD 101 includes a voice recognition unit 20. The voice recognition unit 20 is configured to include circuits used for voice recognition processing. Programs and data used for voice recognition are stored in an appropriate storage device such as the ROM 11 or the storage unit 13. Note that the processing of the voice recognition unit 20 may use a known method, and for example, may involve analyzing and recognizing input voice using an acoustic model or a language model.
[0026] The HMD 101 includes an audio input unit 21. The audio input unit 21 is configured as an audio input device to which audio output from the information terminal 102 is input, for example, in a mapping process described later. The audio input unit 21 is, for example, an audio input device that can acquire information on the direction to the source of audio, and can be configured with, for example, an array microphone 22 or a directional microphone 23, as will be described in detail later.
[0027] The HMD 101 includes a distance measurement unit 24. The distance measurement unit 24 can be configured, for example, by a sensor that measures the distance to the information terminal 102 in a mapping process described later. The distance measurement unit 24 can be configured, for example, by a distance measurement camera 25 (a stereo camera, for example), a LiDAR 26, or a distance sensor 27 that is a different sensor and can appropriately measure the distance to the information terminal 102. The distance measurement unit 24 may be configured by one or more sensors. The distance measurement unit 24 may also be configured by one or more types of sensors.
[0028] The HMD 101 includes a head tracking unit 28. The head tracking unit 28 is used to detect the tilt of the user's head when the HMD 101 is worn. The head tracking unit 28 can be configured with sensors such as an acceleration sensor 29 and a gyro sensor 30. The head tracking unit 28 may be configured with one or more sensors. The head tracking unit 28 may also be configured with one or more types of sensors.
[0029] The HMD 101 includes an eye tracking unit 31. The eye tracking unit 31 is used to detect the direction of the user's gaze when the HMD 101 is worn. The eye tracking unit 31 can be configured with a sensor such as a gaze detection sensor 32. The eye tracking unit 31 may be configured with one or more sensors. The eye tracking unit 31 may also be configured with one or more types of sensors.
[0030] The HMD 101 includes a communication processing unit 33. The communication processing unit 33 includes a circuit that performs communication processing (e.g., signal processing) in wireless communication, and in this embodiment, the HMD 101 includes a wireless LAN communication unit 34 that performs communication processing when performing communication via wireless LAN, and a close-proximity wireless communication unit 35 that performs communication processing when performing close-proximity wireless communication.
[0031] The HMD 101 also includes an interface 36 used for communication. The HMD 101 can transmit and receive data to and from the outside by performing wireless communication with the outside via the interface 36. Here, the HMD 101 may also include an antenna 37 used for wireless communication. Also, a device used for wireless communication, such as a wireless adapter, may be provided.
[0032] Next, an example of a mode of wireless communication will be described with reference to Fig. 2. As shown in Fig. 2, an HMD 101 can communicate with an information terminal 102 via a network 202, for example. In this embodiment, the information terminal 102 is a device that can output audio, and examples of the information terminal 102 include a wearable device 200 and a smartphone 201.
[0033] Next, an example of the structure of an HMD 101 in which the voice input unit 21 is configured with an array microphone 22 will be described with reference to Fig. 3. Note that in the example of Fig. 3, the HMD 101 has a structure shaped like glasses, but the structure of the HMD 101 is not limited to this example and can be modified as appropriate. Here, the description will be made based on the front-back, left-right, and up-down directions shown in Fig. 3.
[0034] 3, the HMD 101 includes a front frame unit 51 on the front side (front side), a left frame unit 52, and a right frame unit 53. Two displays 15 are attached to the front frame unit 51 so as to be positioned in front of the left and right eyes of the user when worn.
[0035] The left frame portion 52 extends rearward from the left end portion 51a of the front frame portion 51 and is located on the left side of the user's head when worn. A speaker 17 (not shown in FIG. 3) is attached to the left frame portion 52 so as to output sound toward the user's left ear. Similarly, the right frame portion 53 extends rearward from the right end portion 51b of the front frame portion 51 and is located on the right side of the user's head when worn. A speaker 17 (not shown in FIG. 3) is attached to the right frame portion 53 so as to output sound toward the user's right ear.
[0036] The HMD 101 is also provided with a first microphone 22a, a second microphone 22b, and a third microphone 22c, which constitute the array microphone 22. In the example of FIG. 3, the first microphone 22a and the second microphone 22b are arranged at the left end 51a and the right end 51b of the front frame 51. That is, the first microphone 22a is arranged at the lower right end of the front frame 51, and the second microphone 22b is arranged at the upper left end of the front frame 51. The third microphone 22c is arranged on the outside (right side) of the right frame 52. Note that, contrary to the arrangement shown in FIG. 3, the first microphone 22a may be arranged at the lower left end of the front frame 51, the second microphone 22b may be arranged at the upper right end of the front frame 51, and the third microphone 22c may be arranged on the outside (left side) of the left frame 52. Furthermore, the first microphone 22a and the second microphone 22b may be located at the end of the front frame portion 51 on the front side of the HMD 101, or on the left or right side.
[0037] When sound is input by the first microphone 22a and the second microphone 22b arranged in this manner, the direction of the sound source (directions related to the left-right and up-down directions) is identified based on the difference in timing of input to the first microphone 22a and the second microphone 22b. Also, when sound is input by the first microphone 22a and the third microphone 22c, the direction of the sound source (direction related to the front-back direction) is identified based on the difference in timing of input to the first microphone 22a and the third microphone 22c. Therefore, with the array microphone 22 arranged in this manner, the HMD 101 can easily identify the direction of the sound source.
[0038] Regarding the above-described arrangement, it is preferable that the microphones (22a, 22b, 22c) of the array microphone 22 are arranged so that the distance between the first microphone 22a and the second microphone 22b and the distance between the first microphone 22a and the third microphone 22c are approximately the same. By adopting such a positional relationship, it is possible to improve the accuracy of identifying the direction of a sound source.
[0039] Next, an example of the structure of the HMD 101 in which the voice input unit 21 is configured with a directional microphone 23 will be described with reference to Fig. 4. In the example of Fig. 4, as in the case of Fig. 3, the HMD 101 has a structure shaped like glasses, but is not limited to this structure. Here, the description will be made based on the front-back, left-right, and up-down directions shown in Fig. 4.
[0040] Similar to the configuration of the array microphone 22 described above, the HMD 101 comprises a front frame portion 51 on the front side (front side), a left frame portion 52, and a right frame portion 53, and the front frame portion 51 has a display 15 attached thereto, and the left frame portion 52 and the right frame portion 53 have speakers 17 attached thereto (not shown in FIG. 4).
[0041] In the example of FIG. 4, the directional microphone 23 is disposed on the upper end side of the central portion 51c of the front frame portion 51. The direction of a sound source is identified by using the directional microphone 23. Note that it is only necessary to be able to identify the direction of the sound source, and the directional pattern of the microphone may be set appropriately. Also, in this example, the directional microphone 23 is disposed on the upper end side of the central portion 51c of the front frame portion 51, but the directional microphone 23 may be disposed in another position. Also, a plurality of directional microphones 23 may be provided instead of a single one, and the number of microphones can be reduced, for example, by appropriately switching the directional pattern of the microphone.
[0042] The above describes the HMD 101 including the array microphone 22 and the HMD 101 including the directional microphone 23, but the HMD 101 may have the following structure. For example, the HMD 101 may be provided with both the array microphone 22 and the directional microphone 23, and the HMD 101 may identify the direction of a sound source based on audio data input to both the array microphone 22 and the directional microphone 23. The HMD 101 may also be provided with a position adjustment mechanism that adjusts the position of the microphone. As an example, the position adjustment mechanism may be a mechanism that can adjust the position of the microphone by sliding the microphone along a frame. The HMD 101 may also have a structure that allows it to be folded or unfolded between frames.
[0043] Next, an example of a method in which a user maps an object will be described with reference to Figures 5 to 7. In the example of Figures 5 to 7, the object of mapping is the information terminal 102 (more specifically, the wearable device 200, which is an example of the information terminal 102). In this example, the information terminal 102 is capable of voice input and voice output, and transitions to a mode in which mapping is performed (mapping mode) by recognizing input voice.
[0044] 5, a user wearing the HMD 101 (the operator 100 in FIG. 5) commands the HMD 101 and the wearable device 200 to start mapping by inputting a voice commanding the start of mapping into the microphone 16 of the HMD 101 and the wearable device 200. When the user inputs a voice commanding the start of mapping, for example, "start mapping," the HMD 101 and the wearable device 200 transition to mapping mode based on appropriate voice recognition.
[0045] Although an example has been described in which the HMD 101 and the information terminal 102 are simultaneously transitioned to the mapping mode, each information device (101, 102) may be transitioned to the mapping mode at different times. For example, the user may transition the HMD 101 to the mapping mode and then transition the information terminal 102 to the mapping mode.
[0046] 6, the user moves the wearable device 200 to a position where the user wants to register the wearable device 200, and causes the wearable device 200 to output sound. Here, the user causes the wearable device 200 to output sound by an appropriate method (for example, by operating a key on the wearable device 200, touching the screen, or inputting sound).
[0047] Then, as shown in FIG. 7 , the sound from the information terminal 102 is input to the HMD 101 (more specifically, the sound input unit 21 of the HMD 101), and the HMD 101 performs a process of mapping the information terminal 102 in a virtual space based on the input sound. Here, the HMD 101 identifies the direction of the sound source (i.e., the information terminal 102) based on the sound input to the sound input unit 21, and calculates the distance to the sound source. Note that the distance to the sound source may be calculated appropriately using data of the sound input to the sound input unit 21 (for example, data correlating the volume of the input sound with the distance to the sound source). Furthermore, if the HMD 101 includes a distance measurement unit 24, the measurement result of the distance to the information terminal 102 by the distance measurement unit 24 may be used. Using the measurement result of the distance measurement unit 24 improves the accuracy of mapping (particularly, the accuracy in the depth direction to the information terminal 102). Furthermore, the HMD 101 may detect the position of the information terminal 102 by wireless communication with the information terminal 102, and perform mapping using the result of the detection.
[0048] Then, based on the direction of the sound source and the distance to the sound source, the HMD 101 maps the target information terminal 102 (in this example, the wearable device 200) to a corresponding position in the virtual space, and places the mapped target virtual sound source 103. Note that in the description here, the information terminal 102 is the target of mapping, but an app held by the information terminal 102 may also be the target of mapping. In this case, the app mapping process is performed by causing the information terminal 102 holding the target app to output sound when the target app is launched or used.
[0049] Then, the HMD 101 generates an icon indicating the mapped target and displays the generated icon on the display 15. Here, the HMD 101 may display the icon at an appropriate position on the display 15, and as an example, may display the icon of the target at a position corresponding to the position of the target mapped in the virtual space. Note that the HMD 101 may display information regarding the name indicating the target (for example, text information such as "wearable device" when the target is the wearable device 200) attached to the icon.
[0050] Here, an example of the audio output of the HMD 101 during and after mapping will be described with reference to FIGS.
[0051] During mapping of the target, as shown in FIG. 8, the user can hear sound from the information terminal 102 (in this example, the wearable device 200) and sound from the speaker 17 of the HMD 101. Here, the speakers 17 of the HMD 101 (left and right speakers 17a and 17b in FIG. 8) output sound from a virtual sound source 103 located at a position considered to be the same as the information terminal 102 (i.e., a position determined based on the orientation of the information terminal 102 and the distance to the information terminal 102). Therefore, sound similar to the sound heard from the information terminal 102 (i.e., sound that sounds as if it is coming from the position of the virtual sound source 103) is output from the speaker 17 of the HMD 101. Therefore, the user can easily check whether mapping is being performed appropriately by comparing the sound actually heard from the information terminal 102 with the sound output from the speaker 17.
[0052] After the mapping, as shown in FIG. 9, even if the position of the information terminal 102 (in this example, the wearable device 200) is changed, the HMD 101 outputs sound that sounds as if it is coming from the position of the virtual sound source 103.
[0053] Here, the relationship between the virtual sound source 103 and stereophonic sound in the virtual space 300, which is the space where an object is mapped, will be described with reference to Fig. 10. Stereophonic sound is reproduced so that the direction and distance of the sound can be sensed, and in this embodiment, the HMD 101 places the virtual sound source 103 in the virtual space 300 and expresses stereophonic sound by calculating whether the sound emitted from it reaches the ears.
[0054] That is, by the user's operation as described above, the HMD 101 maps an object into a virtual space 300, which is a coordinate space centered on the position of the user (in the figure, the operator 100 wearing the HMD 101), and places virtual sound sources (103a, 103b) at the mapped positions in the virtual space. The HMD 101 then expresses stereophonic sound by outputting appropriate sound based on the direction and distance of the virtual sound sources (103a, 103b). Here, the HMD 101 can adjust the sound to suit the sound output device and output the adjusted sound. For example, if the sound output device is a speaker 17, the HMD 101 can output sound adjusted to suit the speaker 17. For example, if the sound output device is a headphone, the HMD 101 can output sound adjusted to suit the headphone.
[0055] In this embodiment, the HMD 101 can map an object in the virtual space 300 of a coordinate system (local coordinate system or world coordinate system) selected by the user. Here, the position of the virtual sound source in each coordinate system when the user moves will be described with reference to Figs. 11 and 12.
[0056] First, a case where the virtual space 300 is a local coordinate system will be described with reference to Fig. 11. The local coordinate system is a coordinate system in which the positions of the virtual sound sources (103a, 103b) move together with the user (in the figure, the operator 100), and in the case of the local coordinate system, the virtual sound sources (103a, 103b) move in accordance with the movement of the user.
[0057] As shown in FIG. 11, for example, when a user wearing the HMD 101 changes their orientation, the positions of the mapped virtual sound sources (103a, 103b) change to follow the changed orientation of the user. In the example of FIG. 11, changing the position of the virtual sound source 103a places a virtual sound source 103c in the virtual space 300, and changing the position of the virtual sound source 103b places a virtual sound source 103d in the virtual space 300. In this way, in the local coordinate system, the positions of the virtual sound sources (103a, 103b) move in the virtual space 300 so as to maintain a constant orientation and a constant distance relationship based on the position and orientation of the user (in other words, the position and orientation of the HMD 101). Note that, for example, head tracking may be used in this processing. Also, for example, a GPS receiving sensor may be provided in the HMD 101, and data based on the GPS may be used.
[0058] Therefore, in the local coordinate system, even if the user wearing HMD101 changes direction or moves, the direction of the virtual sound source relative to the user and the distance to the virtual sound source do not change, and HMD101 outputs sound from a virtual sound source that is in a fixed direction and at a fixed distance relative to the user.
[0059] In contrast, the world coordinate system is a coordinate system in which the positions of the virtual sound sources (103a, 103b) are fixed, and in the world coordinate system, the positions of the virtual sound sources (103a, 103b) do not change even if the user moves, etc. Therefore, as shown in Fig. 12, for example, when the user (operator 100 in the figure) changes his or her orientation, the orientation of the virtual sound sources (103a, 103b) relative to the user changes accordingly, and the HMD 101 outputs sound from the virtual sound sources (103a, 103b) in different directions before and after the user changes his or her orientation. Therefore, unlike the local coordinate system, in the world coordinate system, the direction of a sound heard and the sense of distance of the sound change when the user changes his or her orientation or moves.
[0060] Next, the mapping process will be described in detail with reference to the flowcharts shown in Figures 13 to 15. Figures 13 to 15 are flowcharts used to explain an example of the mapping process.
[0061] As shown in FIG. 13 , the HMD 101 waits until the user signals the start of mapping (S101). When the user utters a voice signaling the start of mapping (for example, the user utters "start mapping") (S102), the control unit 10 performs voice recognition to recognize a keyword signaling the start of mapping (S103). The HMD 101 (more specifically, the control unit 10) recognizes the keyword through voice recognition and activates a mapping mode in which the target is mapped (S104). The HMD 101 then outputs a voice notification prompting the user to select whether to perform mapping in a local coordinate system or a world coordinate system (S105). When the user utters a keyword selected to indicate which coordinate system to use for mapping (for example, the user utters "local coordinate system") (S106), the control unit 10 performs voice recognition to recognize a keyword indicating which coordinate system to use (S107). Then, the HMD 101 outputs a sound notifying the user that the mapping mode in the selected coordinate system has been activated (S108). Here, the HMD 101 outputs a sound such as "The mapping mode in the local coordinate system is starting."
[0062] Note that data such as keywords used by the HMD 101 for voice recognition in the above-described steps S101 to S108 may be stored in advance in an appropriate storage device such as the storage unit 13.
[0063] Next, the user speaks a voice to signal the information terminal 102 (in this example, the wearable device 200) to start mapping (S109). The user speaks, for example, "Start registration." At this point, the wearable device 200 recognizes a keyword by voice recognition (S110), as in the case of the HMD 101 described above, and activates a device registration mode, which is a mapping mode (S111). At this point, the wearable device 200 may output a voice notifying that the device registration mode has been activated (S112). For example, the wearable device 200 may output a voice saying, "Starting device registration mode."
[0064] As in the above case, data such as keywords used by the information terminal 102 for voice recognition in steps S109 to S112 may be stored in advance in an appropriate storage device of the information terminal 102. Also, in this example, an example has been described in which the HMD 101 and the information terminal 102 are individually set to the mapping mode, but the user may simultaneously transition the HMD 101 and the information terminal 102 to the mapping mode by inputting voice into the HMD 101 and the information terminal 102 at the same time.
[0065] In this way, preparations for the mapping process are made from S101 to S112. Then, mapping is performed by the process described below.
[0066] 14, first, the user moves the wearable device 200 to a position where mapping is desired (S201), and then the user presses a button on the wearable device 200 to output a sound (position detection sound) of the target to be mapped (S202).
[0067] Here, when the target to be mapped is the information terminal 102 (in this example, the wearable device 200), the user, for example, causes the information terminal 102 to output a sound related to the mapping mode. On the other hand, when the target to be mapped is an application held by the information terminal 102, the user operates the information terminal 102 to execute the target application and causes the information terminal 102 to output the sound of the application.
[0068] The method for causing the information terminal 102 (in this example, the wearable device 200) to output sound may be any method that can output sound appropriately, and is not limited to pressing a button, but may also be key operation, touching a screen, voice input, or other methods.
[0069] Then, when sound is output from the information terminal 102 in S202, the HMD 101 captures the sound (position detection sound) via the sound input unit 21 (S203). Here, in this example, the sound input unit 21 is the array microphone 22, but may be replaced with, for example, a directional microphone 23.
[0070] Then, the control unit 10 calculates the position (distance and direction) of the wearable device 200 from the captured sound (position detection sound) (S204). The control unit 10 then stores the calculated position information in memory (in this example, the storage unit 13) (S205). The control unit 10 then maps the object (in this example, the wearable device 200) to the calculated position in the stereophonic space (in the virtual space 300) (S206). The control unit 10 then maps the object onto the virtual space 300 based on the coordinate system recognized by the sound recognition in S107 described above. As a result, the virtual sound source 103 is set in the virtual space 300.
[0071] After mapping onto the virtual space 300, the control unit 10 outputs sound from the speaker 17 so that the sound appears to be coming from the mapped position (i.e., the virtual sound source 103) (S207). Therefore, the user can check whether the target has been properly mapped by comparing the sound output from the wearable device 200 with the sound output from the speaker 17.
[0072] The control unit 10 may determine whether the mapping is appropriate based on whether the position of the virtual sound source 103 placed in the virtual space 300 by mapping matches the position of the information terminal 102. The control unit 10 may then automatically adjust the mapped position based on the result of the mapping. That is, the control unit 10 may determine whether the direction of the information terminal 102 and the direction of the virtual sound source 103 match, and adjust the position of the virtual sound source 103 based on the result (S208). Specifically, the control unit 10 determines whether the sound direction matches based on whether the deviation in the sound direction is within a predetermined threshold. If the control unit 10 determines that the sound direction does not match, it adjusts the position information of the wearable device 200. The control unit 10 saves the adjusted position information in memory (S205), and adjusts the position of the virtual sound source 103 by performing mapping again based on this position information (S206).
[0073] Then, the user checks whether the sound direction of the information terminal 102 and the virtual sound source 103 match, and presses a button on the wearable device 200 to stop the sound output (S209). Note that, as in the case of S202 described above, the user may stop the sound output of the wearable device 200 by any appropriate method other than pressing a button.
[0074] In this way, in the processes of S201 to S206, the object is mapped onto the virtual space 300, and in the processes of S207 to S209, it is confirmed whether the mapping is appropriate. Then, through the processes described below, the mapping process ends.
[0075] As shown in FIG. 15, the user checks whether there are any other objects to be mapped, and if there are other objects to be mapped, maps those objects using the method described above (S301). Then, if the user confirms that there are no more objects to be mapped, he or she utters a voice signaling the end of mapping (S302). Here, the user utters, for example, "mapping end." Then, the control unit 10 performs voice recognition to recognize a keyword signaling the end of mapping (S303), and ends the mapping mode (S304). Then, the HMD 101 outputs a voice notifying the user that the mapping mode has ended (S305). Here, the HMD 101 outputs a voice signaling, for example, "mapping mode is ending."
[0076] In this way, the mapping process is completed through steps S301 to S305 (S306). Note that data such as keywords used by the HMD 101 for voice recognition in steps S301 to S305 may be stored in advance in an appropriate storage device such as the storage unit 13.
[0077] The HMD 101 may output a voice warning when attempting to map to a position that has already been mapped in the virtual space 300. In this case, the HMD 101 may output a voice suggesting in which direction to shift the position of the object to be mapped. The HMD 101 may then use voice recognition to recognize a keyword from the voice input by the user and shift the position of the object to be mapped in a predetermined direction. Here, the keyword (e.g., "left," "right," etc.) is stored in an appropriate storage device. The amount of shift can be set as appropriate, but, as an example, can be the minimum amount that avoids overlap. The control unit 10 may then determine the consistency of the voice direction regarding S208 described above after adding this amount of shift.
[0078] Furthermore, the HMD 101 can generate an icon indicating the mapped target. Next, an example of a method for the HMD 101 (specifically, the control unit 10) to generate an icon will be described.
[0079] The HMD 101 can use the voice output from the information terminal 102 when generating an icon for a target. That is, data such as a keyword indicating the target and the voice output when the target is activated is stored in advance in a storage device as data for voice recognition. Then, the HMD 101 performs voice recognition based on the voice input from the information terminal 102 in the above-mentioned S202 or the like, and determines the target for which an icon is to be generated.
[0080] Here, for example, if the target is the wearable device 200, which is the information terminal 102, the sound output when the wearable device 200 is started in mapping mode may be used as a keyword, and the HMD 101 may determine that the target for which an icon is to be generated is the wearable device 200 by recognizing this sound.
[0081] The HMD 101 then generates an icon for the identified target. Data such as the icon design and icon name may be stored in a storage device, and the control unit 10 can generate an icon corresponding to the identified target based on this data. As will be described in detail later, the control unit 10 can also display the generated icon on the display 15. At this time, the icon may be displayed with a name indicating the target.
[0082] An example of generating an icon when the target is an application will also be described. When the target is an application, data such as keywords indicating the application is stored in a storage device, similar to the example of information terminal 102.
[0083] For example, if the target is an app related to a weather forecast, voice that is a keyword related to the weather forecast (for example, "weather, sunny, cloudy, rainy") and voice output when the app is started may be stored in the storage device. Then, the HMD 101 performs voice recognition based on the voice of the app input from the information terminal 102 in S202 or the like, and determines the target for which an icon is to be generated.
[0084] Although an example has been described here in which a target for generating an icon is determined based on voice recognition, the HMD 101 may acquire information for determining the target through communication. For example, the HMD 101 may acquire data for determining the target (e.g., information related to the name of the target) through communication with the information terminal 102 and determine the target using the acquired information. Here, information associated with the information acquired through communication (e.g., a table containing records of information obtainable through communication and the name of the target) may be stored in a storage device, and the HMD 101 may determine the target from the information acquired through communication by referring to the stored information.
[0085] The HMD 101 can then display the generated target icon on the display 15. Here, as an example, the control unit 10 may display the icon at a position corresponding to the mapped position in the virtual space 300, with the user wearing the HMD as a reference. The display position of the icon can be moved appropriately by a user operation or the like. The HMD 101 may be configured to move the icon, for example, by a user operation to select and move a displayed icon (by drag and drop). On the other hand, as will be described in detail later, the icon may also be moved by voice input.
[0086] In the HMD 101, the target icon displayed on the display 15 can be selected by the user. The user can then appropriately select the target icon and operate the mapped target. Next, a voice operation process using icons will be described with reference to the flowcharts shown in Figs. 16 to 18. Figs. 16 to 18 are flowcharts used to explain an example of the voice operation process.
[0087] As shown in FIG. 16, the HMD 101 waits until the user signals the start of the voice operation mode (a mode in which voice operation is possible) (S401). Then, when the user utters a voice signaling the start of the voice operation mode (for example, the user utters "start operation") (S402), the control unit 10 performs voice recognition to recognize a keyword signaling the start of the voice operation mode (S403). Then, the HMD 101 (more specifically, the control unit 10) recognizes the keyword through voice recognition and activates the voice operation mode (S404). Here, the HMD 101 outputs a voice notifying that the voice operation mode has been activated (S405). The HMD 101 issues a notification such as "operation will start," for example.
[0088] In this way, the voice operation mode is activated from S401 to S405, and preparations for voice operation are made. Then, as in the example described below, the user can execute the target voice operation. In the following description, the icon generated by mapping may be referred to as a mapping icon.
[0089] First, the user vocalizes the mapping icon that the user wants to operate (S406). As an example, if the user wants to select the smartphone 201, which is the mapped information terminal 102, the user vocalizes "smartphone." Then, the control unit 10 recognizes the mapping icon vocalized by the user through voice recognition (S407). That is, the control unit 10 selects the mapping icon that corresponds to the voice input by the user. Note that "smartphone" is an abbreviation for the smartphone 201.
[0090] The HMD 101 then notifies the user of the selected mapping icon by voice (S408). Here, the HMD 101 notifies, for example, "Smartphone has been selected." The user checks whether the selected mapping icon is correct based on the content of the notification, and if it is correct, vocalizes that it is correct (for example, vocalizes "OK") (S409). This allows the HMD 101 to recognize the keyword by voice recognition, and becomes able to execute the processing of S501 described below. On the other hand, if the mapping icon selected is incorrect, the user vocalizes that it is incorrect (for example, vocalizes "NO"). The user then vocalizes the mapping icon they want to operate again, causing the HMD 101 to execute processing to recognize the mapping icon.
[0091] In this way, in steps S406 to S409, the mapping icon that the user wants to operate by voice is selected. When a mapping icon is selected, a sound may be output indicating that the mapping icon has been selected. This sound may be a simple sound such as "pop," or the name of the object that the mapping icon represents. This allows the user to understand that the mapping icon has been selected.
[0092] Furthermore, sound indicating that a mapping icon has been selected may be output from the speaker 17 so as to sound as if it is coming from the direction in which the selected mapping icon is displayed. As an example, when the selected mapping icon is displayed directly in front of the right eye of the user wearing the HMD 101, based on the center of the front side of the HMD 101, sound may be output so as to sound as if it is coming from the right side. Furthermore, when the mapping icon is displayed toward the center of the HMD 101, sound may be output so as to sound as if it is coming from the front.
[0093] The HMD 101 may also use an appropriate tracking technology when selecting a mapping icon. For example, the HMD 101 may detect the direction of the user's head using the head tracking unit 28 in addition to the user's voice input into the microphone 16, and select a mapping icon that corresponds to the voice input into the microphone 16 and is displayed in that direction. In this case, the user's desired mapping icon is selected by turning their head in the direction of the mapping icon they want to select and uttering a voice.
[0094] Furthermore, the HMD 101 may detect the user's line of sight using the eye tracking unit 31 in addition to the user's voice being input to the microphone 16, and select a mapping icon displayed in that direction that corresponds to the voice input to the microphone 16. In this case, the user's desired mapping icon is selected by directing their line of sight toward the mapping icon they wish to select and uttering a voice.
[0095] In this way, by using tracking technology, it is possible to select a mapping icon based not only on voice but also on the user's movements and gaze. Note that in S401 to S409, data such as keywords used for voice recognition may be stored in advance in an appropriate storage device such as the storage unit 13. Next, the voice operation process will be described. This voice operation is performed via wireless communication with the information terminal 102, which processes voice input from the HMD 101.
[0096] As shown in FIG. 17, the user vocalizes the operation content of the selected mapping icon (S501).
[0097] Here, various operations are conceivable as the operation contents. Examples of the operation contents include operations related to display (such as displaying a menu or selecting a menu item), displaying and moving a cursor, adjusting the volume, operations related to making and receiving calls when the target has a call function (a function for processing audio during a call) such as the smartphone 201, moving the position of a displayed icon (remapping), operating the target information terminal 102, and executing the target app (launching the app). Note that the HMD 101 can output audio from the target via the speaker 17 based on a virtual sound source in the virtual space 300. Furthermore, if the information terminal 102 has a call function, the information terminal 102 may process audio related to a call, and the microphone 16 and speaker 17 of the HMD 101 may input and output audio during a call.
[0098] Then, the control unit 10 recognizes the operation content by voice recognition (S502), and the HMD 101 notifies the recognized operation content by voice (S503).
[0099] For example, when the user wants to move the mapping icon of the selected smartphone 201 to the left, the user utters "move left." The HMD 101 then recognizes through voice recognition that the mapping icon is to be moved to the left, and notifies the user by voice, for example, "Move the smartphone to the left." In this way, in steps S501 to S503, the operation content is input to the HMD 101, and the HMD 101 recognizes the operation content.
[0100] The control unit 10 then executes an operation according to the input operation content (S504) and notifies the user of the executed operation content by voice (S505). When the control unit 10 executes an operation to move the mapping icon of the smartphone 201 to the left, the control unit 10 notifies the user by voice, for example, "Your smartphone has been moved to the left." Note that the operation of the control unit 10 here is a process before confirmation, and the user determines whether the operation content is correct (S506). If the user determines that the operation content is correct, the process described below is executed, and the operation content is confirmed. On the other hand, if the user determines that the operation content is incorrect, the user inputs the operation content again. Note that in this case, the operation content that the user determined to be incorrect is reset. In this way, the operation content input by the control unit 10 is executed from S504 to S506. Next, the process of confirming the operation content will be described.
[0101] If the user determines that the operation content is correct, the user inputs a keyword indicating this by voice (S507). For example, the user utters "OK." Then, the control unit 10 recognizes the keyword by voice recognition (S508) and confirms the operation content (S509). Then, the control unit 10 notifies the user by voice that the operation content has been confirmed (S510). As described above, when it is confirmed that the mapping icon of the smartphone 201 has been moved to the left, the control unit 10 may notify the user by voice, for example, saying, "The move to the left is confirmed."
[0102] In this way, the voice operation is confirmed from S507 to S510. Here, in the voice processing operation from S501 to S510, data such as keywords used for voice recognition may be stored appropriately in a storage device such as the storage unit 13, and the control unit 10 can use this data in the voice recognition.
[0103] In addition, when a mapping icon is moved by voice operation and overlaps with another mapping icon, the HMD 101 may output a voice warning. The HMD 101 may also output a voice suggesting in which direction to move the mapping icon to prevent the mapping icon from overlapping. The HMD 101 can recognize a keyword from the voice input by the user using voice recognition and shift the position of the mapping target in a predetermined direction. Here, the keyword (e.g., "left," "right," etc.) is stored in an appropriate storage device. The amount of shift can be set appropriately, but, as an example, can be the minimum amount that avoids overlapping.
[0104] Next, an example of a process for ending a voice processing operation (i.e., a process for ending the voice operation mode) will be described. As shown in FIG. 18, the user checks whether there is a mapping icon that the user wants to operate by voice (S601). If there is no corresponding mapping icon, the user speaks a keyword indicating that the voice operation is to be ended (S602). The user speaks, for example, "end operation." Then, the control unit 10 recognizes the keyword by voice recognition (S603), and the HMD 101 ends the voice operation mode (S604). Then, the HMD 101 notifies the user by voice that the voice mode has ended (S605). Here, the HMD 101 outputs a voice saying, for example, "end operation."
[0105] In this way, the voice operation mode ends after steps S601 to S605 (S606). Note that data such as keywords used by the HMD 101 for voice recognition in steps S601 to S605 may be stored in advance in an appropriate storage device such as the storage unit 13.
[0106] As described above, the user can perform voice operations on the information terminal 102 from the HMD 101. Here, input and output of data between the HMD 101 and the information terminal 102 during voice operations will be described with reference to FIG.
[0107] First, the HMD 101 waits for a voice input from the user regarding the operation content, and when the voice input regarding the operation content is received, starts an operation mode for the information terminal 102 (wearable device operation mode in FIG. 19) (S701). Then, when the information terminal 102 is to be operated by voice (i.e., when the operation content for the information terminal 102 is recognized in the processing of S502 described above), the control unit 10 starts up the communication unit (communication processing unit 33 and interface 36) and starts communication with the information terminal 102 (wearable device 200 in this example) (S702).
[0108] Then, the control unit 10 transmits the operation content to the wearable device 200 via the network 202 (S703), and receives the operation result from the wearable device 200 (S704). The user then checks the received operation result to see if the operation was performed correctly (S705). That is, in S705, the confirmation of S506 described above is performed. Then, if the user confirms that the operation was performed correctly, the user vocally inputs a keyword indicating that the operation was performed. Then, the control unit 10 confirms the operation content, and the operation mode for the information terminal 102 ends (S706).
[0109] According to this embodiment, a user can easily perform target mapping processing, target icon generation, and target operation based on the simple technique of inputting voice. Therefore, for example, even if a user's view of the outside world is limited, the system can be used conveniently. Furthermore, according to this embodiment, an information terminal system is realized that includes an HMD 101, which is an example of an audio augmented reality object playback device, and one or more information terminals 102. Note that, in the above description, examples have been described in which a wearable device 200 or a smartphone 201 is used as an example of the information terminal 102, but the information terminal 102 may be a different type of terminal. Furthermore, the information terminal 102 may be a terminal that can be normally operated using methods other than voice. In this case, an input to the information terminal 102 to signal the start of mapping may be made using methods other than voice.
[0110] Next, a second embodiment will be described with reference to FIG. 20. Functions similar to those in other embodiments are denoted by the same reference numerals, and descriptions thereof may be omitted. In the second embodiment, an example of an audio augmented reality object reproduction device 1001 will be described, in which the display 15 is omitted from the HMD 101 described in the first embodiment. In this audio augmented reality object reproduction device 1001, display-related processing is omitted.
[0111] As an example, the audio augmented reality object reproduction device 1001 can be a device worn on the head like headphones. The audio augmented reality object reproduction device 1001 is connected to the information terminal 102 and, similar to the above description, performs mapping on the virtual space 300 in response to audio input from the target. When the user inputs a desired operation, the audio augmented reality object reproduction device 1001 performs processing corresponding to the user's operation. Here, the user can perform various operations, such as an operation to reproduce the mapped target, similar to the above description. When reproducing the target, the audio augmented reality object reproduction device 1001 can output sound that sounds as if it is coming from the position of the virtual sound source 103 in the virtual space 300.
[0112] The first and second embodiments have been described above. Here, the HMD 101 and the audio augmented reality object reproduction device 1001 described in the embodiments may be used standalone without being connected to the information terminal 102. In this case, the HMD 101 and the audio augmented reality object reproduction device 1001 perform mapping using audio from the information terminal 102 and perform processing in response to user operations, as in the above description, but the processing using communication with the information terminal 102 is omitted.
[0113] When playing back a mapped object, the HMD 101 and the audio augmented reality object playback device 1001 store data to be played back from the mapped object in advance, and the HMD 101 and the audio augmented reality object playback device 1001 output the sound as if it were coming from a corresponding position in the virtual space 300 based on the pre-stored data.
[0114] The audio augmented reality object reproduction device (101, 1001) may be configured to be used only as a standalone device, in which case the configuration for communicating with the information terminal 102 may be omitted. The information terminal 102 may also be a terminal from which the configuration for communication is omitted.
[0115] Although the embodiments of the present invention have been described above, it goes without saying that the configurations for realizing the technology of the present invention are not limited to the above-described embodiments, and various modifications are possible. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. All of these fall within the scope of the present invention. Furthermore, numerical values, messages, etc. appearing in the text and figures are merely examples, and the effects of the present invention will not be impaired even if different ones are used.
[0116] The programs described in each processing example may be independent programs, or multiple programs may constitute a single application program. The order in which each process is performed may also be changed.
[0117] Some or all of the functions of the present invention described above may be implemented in hardware, for example, by designing them as integrated circuits. They may also be implemented in software by a microprocessor unit, CPU, or the like interpreting and executing an operating program that implements each function. Furthermore, the scope of software implementation is not limited, and hardware and software may be used together. Some or all of the functions may also be implemented by a server. The server may be, for example, a local server, a cloud server, an edge server, or an online service, as long as it can cooperate with other components via communications to execute the functions. Information such as programs, tables, and files that implement each function may be stored in a memory, a recording device such as a hard disk or solid-state drive (SSD), or a recording medium such as an IC card, SD card, or DVD, or may be stored in a device on a communications network.
[0118] Furthermore, the control lines and information lines shown in the diagram are those considered necessary for explanation, and do not necessarily represent all the control lines and information lines on the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0119] 10 Control unit (processor) 11 ROM 12 RAM 13 Storage Section 14 Camera 15 Display (display unit) 16. Mike 17 Speaker 18 buttons 19 Touch Sensor 20 Voice Recognition Unit 21 Audio input section 22 Array microphone 23 Directional microphone 24 Distance measurement unit 25 Distance Measuring Camera 26 LiDAR 27 Distance Sensor 28 Head Tracking Unit 29 Acceleration Sensor 30 Gyro sensor 31 Eye Tracking Section 32 Eye gaze detection sensor 33 Communication processing unit 34 Wireless LAN communication section 35 Near Field Communication Unit 36 Interface 37 Radio Antenna 100 Operators (Users) 101 HMD (Head Mounted Display) 102 Information terminal 103 Virtual Sound Source 200 Wearable Devices 201 Smartphone 202 Network 300 Virtual Space 1001 Audio Augmented Reality Object Playback Device
Claims
1. An audio augmented reality object playback device capable of mapping an object in a virtual space, comprising: a processor; The processor: Based on the voice output from and input to the information terminal, the information terminal or the application of the information terminal is mapped to a position in a virtual space corresponding to the position of the information terminal; determining whether the mapping is appropriate based on whether the position of the virtual sound source arranged by the mapping matches the position of the information terminal that outputs the sound; An audio augmented reality object playback device.
2. 2. The audio augmented reality object playback device according to claim 1, an array microphone for inputting voice from the information terminal; The array microphone includes: (1) Consisting of microphones arranged at the upper left and lower right ends of the front side of the audio augmented reality object playback device and on the right side of the audio augmented reality object playback device, or (2) Consisting of microphones arranged at the upper right and lower left ends of the front side of the audio augmented reality object playback device and on the left side of the audio augmented reality object playback device. An audio augmented reality object playback device.
3. 3. The audio augmented reality object playback device according to claim 2, In the array microphone, In the case of the configuration (1), when the audio augmented reality object reproduction device is worn, the microphones are arranged so that the distance between each microphone on the front side is approximately the same as the distance between the microphone at the bottom right end of the front side and the microphone on the right side; In the case of the configuration (2) above, when the audio augmented reality object playback device is worn, the microphones are arranged so that the distance between each microphone on the front side is approximately the same as the distance between the microphone at the bottom left end of the front side and the microphone on the left side. An audio augmented reality object playback device.
4. 2. The audio augmented reality object playback device according to claim 1, one or more directional microphones for inputting voice from the information terminal; An audio augmented reality object playback device.
5. An audio augmented reality object playback device according to any one of claims 1 to 4, The processor: If it is determined that the mapping is not appropriate, the position of the virtual sound source is adjusted to match the position of the information terminal. An audio augmented reality object playback device.
6. An audio augmented reality object playback device according to any one of claims 1 to 4, The processor: If the position to be mapped in the virtual space overlaps with the position of another object that has already been mapped, an audio warning is output. An audio augmented reality object playback device.
7. The audio augmented reality object playback device according to claim 6, The processor: Output a voice message suggesting the direction in which to shift the position of the target to be mapped. An audio augmented reality object playback device.
8. An audio augmented reality object playback device according to any one of claims 1 to 4, The processor: outputting a voice prompting the user to select whether the mapping will use a local coordinate system or a world coordinate system; Mapping in a coordinate system corresponding to the user's voice input, An audio augmented reality object playback device.
9. An audio augmented reality object playback device according to any one of claims 1 to 4, A display unit is provided, The processor: an icon used to operate the mapped object is displayed on the display unit; An audio augmented reality object playback device.
10. The audio augmented reality object playback device according to claim 9, The processor: Selecting a target icon corresponding to the user's voice input; An audio augmented reality object playback device.
11. The audio augmented reality object playback device according to claim 10, A head tracking unit is provided to detect the movement of the user's head. The processor: Select an icon of the direction detected by the head tracking unit. An audio augmented reality object playback device.
12. The audio augmented reality object playback device according to claim 10, An eye tracking unit is provided to detect the direction of the user's gaze, The processor: Selecting an icon of the direction detected by the eye tracking unit; An audio augmented reality object playback device.
13. The audio augmented reality object playback device according to claim 9, The processor: a name of the object acquired based on the voice from the information terminal, displayed on the display unit together with an icon of the object; An audio augmented reality object playback device.
14. The audio augmented reality object playback device according to claim 9, It has an interface for communication, The processor: a name of the object acquired based on communication with the information terminal, displayed on the display unit together with an icon of the object; An audio augmented reality object playback device.
15. one or more information terminals; an audio augmented reality object playback device capable of mapping an object in a virtual space; Equipped with The audio augmented reality object playback device includes: a processor; The processor: Based on the voice output from and input to the information terminal, the information terminal or the application of the information terminal is mapped to a position in a virtual space corresponding to the position of the information terminal; determining whether the mapping is appropriate based on whether the position of the virtual sound source arranged by the mapping matches the position of the information terminal that outputs the sound; An information terminal system comprising:
16. An information terminal system according to claim 15, The audio augmented reality object playback device includes: A display unit is provided, The processor: an icon used to operate the mapped object is displayed on the display unit; An information terminal system comprising:
17. An information terminal system according to claim 16, The processor: Selecting a target icon corresponding to the user's voice input; An information terminal system comprising:
18. An information terminal system according to claim 16, The processor: a name of the object acquired based on the voice from the information terminal, displayed on the display unit together with an icon of the object; An information terminal system comprising:
19. The information terminal system according to claim 16, The audio augmented reality object playback device includes: It has an interface for communication, The processor: a name of the object acquired based on communication with the information terminal, displayed on the display unit together with an icon of the object; An information terminal system comprising:
Citation Information
Patent Citations
Sound processor
JP2006227328A
Voice control device, voice control method, and program
JP2013101248A
Stereo acoustic signal reproduction device, stereo acoustic signal reproduction method and stereo acoustic signal reproduction program
JP2018064227A
Sound image localization device and sound image localization method
JP2018148323A
Sound data processing device and sound data processing method
JP2021150835A