Information processing apparatus and information processing method
The information processing apparatus corrects virtual speaker positions based on user interaction to address inaccuracies in HRTF, improving sound quality and usability in stereoscopic sound reproduction.
Patent Information
- Application Number
- JP2022536225
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-15
- Filing Date
- 2021-06-28
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2041-06-28
AI Technical Summary
Conventional technologies for stereoscopic sound reproduction using head-related transfer functions (HRTF) suffer from inaccuracies that impair sound quality and user usability due to individual differences in HRTF, leading to mislocalization of sound images and potential deterioration in sound quality.
An information processing apparatus that corrects the position information of virtual speakers based on user-perceived positions, utilizing a correction unit to align first position information with second position information acquired through user interaction, ensuring accurate localization of sound objects.
Improves user usability by accurately localizing sound images at intended positions, enhancing sound quality and reducing the likelihood of sound quality impairment.
Smart Images

Figure 0007711708000001 
Figure 0007711708000002 
Figure 0007711708000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a terminal device.
Background Art
[0002] A technique for stereoscopically reproducing sound images in headphones or the like is known by using a head-related transfer function (hereinafter, appropriately referred to as "Head Related Transfer Function: HRTF") that mathematically represents the way sound reaches the ears from a sound source.
[0003] Since HRTF has large individual differences, it is desirable to use the HRTF for each individual when using it. For this purpose, for example, a technique for estimating HRTF based on an image of a user's auricle is known.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the conventional technology, there is room for further promoting the improvement of user-friendliness. For example, in the conventional technology, since it is an estimation of HRTF, an error from the actual HRTF may occur, and the sound quality may be impaired when reproducing the sound image.
[0006] Therefore, the present disclosure proposes a novel and improved information processing apparatus, information processing method, and terminal device capable of further promoting the improvement of user-friendliness.
Means for Solving the Problems
[0007] According to the present disclosure, there is provided an information processing apparatus including: a correction unit that renders audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space; and an acquisition unit that acquires first position information regarding virtual positions of the virtual speakers in the space and second position information regarding positions of the virtual speakers in the space perceived by a user, wherein the correction unit corrects the first position information of at least one of the plurality of virtual speakers based on the second position information.
Brief Description of Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Best Mode for Carrying Out the Invention
[0009] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions are omitted.
[0010] The description will be made in the following order. 1. One Embodiment of the Present Disclosure 1.1. First 1.2. Configuration of the Information Processing System 2. Functions of the Information Processing System 2.1. Overview 2.2. Example of Functional Configuration 2.3. Processing of the Information Processing System 2.4. Variations of Processing 3. Example of Hardware Configuration 4. Summary
[0011] <<1. One Embodiment of the Present Disclosure>> <1.1. First> HRTF represents, as a transfer function, the change in sound caused by surrounding objects including the shape of a human pinna and head. In general, measurement data for obtaining HRTF is acquired by measuring a measurement acoustic signal (audio signal) using a microphone or dummy head microphone worn inside the human pinna.
[0012] For example, HRTF used in technologies such as 3D audio is often calculated using measurement data obtained with a dummy head microphone or the like, or the average value of measurement data obtained from a large number of humans. However, since there are large individual differences in HRTF, it is desirable to use the user's own HRTF in order to achieve a more effective acoustic performance effect.
[0013] In relation to the above technology, for example, a technology for estimating HRTF based on an image of a user's auricle is known (Patent Document 1). However, in the conventional technology, there is a possibility that the sound quality may be impaired when reproducing sound images, so there is room for further improvement in user usability.
[0014] In recent years, the development of multi-channel audio, which expands the playback capabilities of two-channel stereo in three-dimensional directions, has become widespread. In 3D-Audio in the MPEG-H 3D-Audio standard, it is possible to reproduce three-dimensional sound direction, distance, spread, etc., so more immersive playback is possible compared to conventional stereo playback.
[0015] In relation to the above technology, for example, a technology for obtaining a speaker signal (virtual speaker signal) by rendering 3D-Audio object data (for example, acoustic signals of sound objects and metadata of position information) to a plurality of virtual speakers with predetermined positions by VBAP (Vector Based Amplitude Panning), which is an example of a third-order audio panning method, is known. In VBAP, the playback space is divided into triangular regions composed of three speakers, and amplitude panning is performed by distributing the sound source signal to each speaker according to the weight coefficients. Also, in relation to the above technology, for example, a technology for obtaining a headphone signal (headphone playback signal) for each virtual speaker composed of L (Left: L) and R (Right: R) signals by applying a pre-held HRTF to the speaker signal for each virtual speaker is known. And, in relation to the above technology, for example, a technology for obtaining a headphone signal by adding (summing) the headphone signals for each virtual speaker for all virtual speakers for each L and R signal is known. Thus, in the conventional technology, 3D-Audio could be reproduced with headphones, for example, by obtaining a signal reproduced from headphones using the above technology. However, in the conventional technology, there is a case where the sound image may not be localized at a predetermined position, and there is a possibility that the sound quality may be impaired, so there is room for further improvement in user usability.
[0016] Therefore, the present disclosure proposes a novel and improved information processing apparatus, information processing method, and terminal device that can promote further improvement in user usability.
[0017] <1.2. Configuration of Information Processing System> The configuration of the information processing system 1 according to the embodiment will be described. FIG. 1 is a diagram showing a configuration example of the information processing system 1. As shown in FIG. 1, the information processing system 1 includes an information processing apparatus 10, headphones 20, and a terminal device 30. A variety of devices can be connected to the information processing apparatus 10. For example, the headphones 20 and the terminal device 30 are connected to the information processing apparatus 10, and information cooperation is performed between the devices. The information processing apparatus 10, the headphones 20, and the terminal device 30 are connected to an information communication network by wireless or wired communication so that they can perform mutual information and data communication and operate in cooperation. The information communication network can be configured by the Internet, a home network, an IoT (Internet of Things) network, a P2P (Peer-to-Peer) network, a proximity communication mesh network, or the like. For wireless, for example, technologies based on mobile communication standards such as Wi-Fi, Bluetooth (registered trademark), or 4G and 5G can be used. For wired, power line communication technologies such as Ethernet (registered trademark) or PLC (Power Line Communications) can be used.
[0018] The information processing device 10, the headphones 20, and the terminal device 30 may each be separately provided as a plurality of computer hardware devices on so-called on-premises, an edge server, or in the cloud. Alternatively, the functions of any plurality of the devices among the information processing device 10, the headphones 20, and the terminal device 30 may be provided as the same device. For example, the information processing device 10 and the headphones 20 may be provided as a device that functions integrally and communicates with the terminal device 30. Also, for example, the information processing device 10 and the terminal device 30 may be realized such that the information processing device 10 and the terminal device 30 function integrally as the same terminal such as a smartphone. Further, the user can perform information and data communication with the information processing device 10, the headphones 20, and the terminal device 30 via a user interface (including a graphical user interface: GUI) and software (configured by a computer program (hereinafter also referred to as a program)) that operate on a terminal device (not shown) (a personal device such as a display as an information display device, a PC (personal computer) including voice and keyboard input, or a smartphone).
[0019] (1) Information processing device 10 The information processing device 10 is an information processing device that performs a process of rendering audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space. Also, the information processing device 10 corrects position information regarding a virtual position of the virtual speaker in space. Thereby, since the information processing device 10 can localize the sound image of the sound object at the intended position, the possibility of deterioration of sound quality can be reduced. As a result, the information processing device 10 can promote further improvement of user usability.
[0020] In addition, the information processing apparatus 10 also has a function of controlling the overall operation of the information processing system 1. For example, the information processing apparatus 10 controls the overall operation of the information processing system 1 based on the information coordinated among the devices. Specifically, the information processing apparatus 10 corrects the position information of the virtual speaker based on the information transmitted from the terminal device 30.
[0021] The information processing apparatus 10 is realized by a PC (Personal Computer), a server, or the like. Note that the information processing apparatus 10 is not limited to a PC, a server, or the like. For example, the information processing apparatus 10 may be a computer hardware device such as a PC or a server in which the function as the information processing apparatus 10 is implemented as an application.
[0022] The information processing apparatus 10 may be any device as long as it can realize the processing in the embodiment. Also, the information processing apparatus 10 may be a device such as a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, or a PDA. Note that hereinafter, in the embodiment, it is assumed that the information processing apparatus 10 and the terminal device 30 may be realized by the same terminal such as a smartphone.
[0023] (2) Headphones 20 The headphones 20 are headphones used by the user to listen to sound. For example, the headphones 20 are headphones having a member in the configuration that can provide sound in contact with the user's ears. For example, the headphones 20 are headphones having a member in the configuration that can separate the space including the user's eardrum from the outside world. When the user plays back, the headphones 20 output, for example, a two-channel headphone signal for L and R.
[0024] The headphones 20 are not limited to headphones and may be any device that can provide sound. For example, the headphones 20 may be earphones or the like.
[0025] (3) Terminal device 30 The terminal device 30 is an information processing device used by a user. The terminal device 30 may be any device as long as it can realize the processing in the embodiment. Further, the terminal device 30 may be a device such as a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, a PDA, or the like.
[0026] <<2. Functions of the Information Processing System>> As described above, the configuration of the information processing system 1 has been explained. Next, the functions of the information processing system 1 will be explained.
[0027] In the following embodiments, a virtual speaker will be used for the explanation. However, it is not limited to a virtual speaker, and any device that provides virtual sound may be used.
[0028] In the following embodiments, the position information regarding the virtual position of the virtual speaker in space will be appropriately referred to as "first position information". Further, in the following embodiments, the position information regarding the position of the virtual speaker in space perceived by the user will be appropriately referred to as "second position information".
[0029] The HRTF according to the embodiment is not limited to the HRTF based on the measurement data actually measured as the user's HRTF. For example, the HRTF according to the embodiment may be an average HRTF based on the HRTFs of a plurality of users, which is used as the HRTF of the target user (the user to be targeted). As another example, the HRTF according to the embodiment may be an HRTF estimated from imaging information such as ear images. In the following embodiments, the HRTF will be used for the explanation. However, it is not limited to the HRTF, and it may be a BRIR (Binaural Room Impulse Response: BRIR). Further, the HRTF according to the embodiment may be any one as long as it measures the transfer characteristics of the sound reaching the user's ear from a predetermined position in space as an impulse response.
[0030] <2.1. Overview> FIG. 2 is a diagram showing an overview of an acoustic space according to an embodiment. In FIG. 2, an acoustic space is provided to user U11 using three virtual speakers (speakers SP11 to SP13). Note that user U11 is assumed to reproduce an acoustic signal from speakers SP11 to SP13 with headphones HP11. Here, speakers SP11 to SP13 are respectively assumed to be located at positions A to C. These positions A to C are the first position information of each virtual speaker. Also, data TF11 to data TF13 respectively indicate HRTFs from positions A to C. Specifically, data TF11 to data TF13 respectively indicate characteristics simulating the transfer characteristics from predetermined positions A to C to the eardrum of user U11.
[0031] In the prior art, HRTFs may be held for each of positions A to C. Note that this HRTF is, for example, the impulse response for L and R of headphones or the like. In the prior art, for example, when applying the HRTF of position A to a certain one-channel acoustic signal, a two-channel acoustic signal may be obtained. Among the two-channel acoustic signals, the signal for L is the result of performing a convolution process on the one-channel acoustic signal of the input with the impulse response for L of the HRTF. Similarly, the signal for R among the two-channel acoustic signals is the result of performing a convolution process with the impulse response for R. Here, since the HRTF is a characteristic simulating the transfer characteristics from a predetermined position to the eardrum of a human, when reproducing an acoustic signal with headphones HP11, user U11 perceives that the sound is localized at position A, for example.
[0032] FIG. 3 is a diagram showing an overview of an acoustic space according to an embodiment. Here, in FIG. 2, the case where sound is localized at a predetermined position is shown, while in FIG. 3, the case where sound is not localized at a predetermined position is shown. Note that the same explanations as those in FIG. 2 are omitted as appropriate. In FIG. 3, the speaker SP11 will be described as a virtual speaker to be rendered (hereinafter, appropriately referred to as a "virtual speaker to be reproduced"). Since the HRTF depends on the shape of a human head, the shape of an auricle, the shape of an ear canal, etc., the pre-held HRTF may not match the HRTF of a user. In FIG. 3, since the pre-held HRTF does not match the HRTF of the user U11, for example, the sound image is localized at a position A prime different from the position A. This position A prime is the second position information of the speaker SP1 perceived by the user. In this case, the user U11 perceives the speaker SP11 at the position A prime instead of the original position A. Also, when using the prior art to obtain a headphone signal by rendering, for example, a sound object TB11 having position information of the centroid position ☆ in the triangular region from position A to position C, since the perceived position of the position A of the user U11 becomes the position A prime, the perceived position of the sound object TB11 may also become the position ☆ prime instead of the position ☆. Here, the reason why the perceived position becomes the position ☆ prime is that after the gain of each virtual speaker becomes the same by VBAP, the user perceives the position A as the position A prime. Therefore, since the sound object TB11 may not be perceived at the originally intended position, there is a possibility that the sound quality may be impaired.
[0033] <2.2. Functional Configuration Example> FIG. 4 is a block diagram showing a functional configuration example of the information processing system 1 according to the embodiment.
[0034] (1) Information Processing Device 10 As shown in FIG. 4, the information processing device 10 includes a communication unit 100 and a control unit 110. Note that the information processing device 10 has at least the control unit 110.
[0035] (1-1) Communication Unit 100 The communication unit 100 has a function of communicating with an external device. For example, in communication with an external device, the communication unit 100 outputs the information received from the external device to the control unit 110. Specifically, the communication unit 100 outputs the information received from the terminal device 30 to the control unit 110. For example, the communication unit 100 outputs the second position information of the virtual speaker to the control unit 110.
[0036] In communication with an external device, the communication unit 100 transmits the information input from the control unit 110 to the external device. Specifically, the communication unit 100 transmits the information regarding the acquisition of the information about the perceived position of the virtual speaker input from the control unit 110 to the terminal device 30. The communication unit 100 is composed of a hardware circuit (such as a communication processor), and can be configured to perform processing by a computer program operating on the hardware circuit or on another processing device (such as a CPU) that controls the hardware circuit.
[0037] (1-2) Control unit 110 The control unit 110 has a function of controlling the operation of the information processing device 10. For example, based on the second position information, the control unit 110 performs processing for correcting the first position information.
[0038] In order to realize the above functions, as shown in FIG. 4, the control unit 110 includes an acquisition unit 111, a processing unit 112, and an output unit 113. The control unit 110 is composed of a processor such as a CPU, and may be configured to read software (computer program) for realizing the functions of the acquisition unit 111, the processing unit 112, and the output unit 113 from the storage unit 120 and perform processing. Also, one or more of the acquisition unit 111, the processing unit 112, and the output unit 113 can be composed of a hardware circuit (such as a processor) different from the control unit 110, and can be configured to be controlled by a computer program operating on the different hardware circuit or on the control unit 110.
[0039] · Acquisition unit 111 The acquisition unit 111 has a function of acquiring the first position information of the virtual speaker. For example, the acquisition unit 111 acquires the first position information of a plurality of virtual speakers. Further, the acquisition unit 111 acquires the second position information of the virtual speaker perceived by the user. For example, the acquisition unit 111 acquires the second position information of the virtual speaker to be played. Further, for example, the acquisition unit 111 acquires the second position information of the virtual speaker based on the input information input by the user during the reproduction of the output signal (for example, the headphone signal) from the audio output unit such as headphones.
[0040] The acquisition unit 111 acquires the user's HRTF data held at the position of the virtual speaker. For example, the acquisition unit 111 acquires the HRTF data obtained by measuring the transfer characteristics of the sound reaching the user's ear from each virtual speaker as an impulse response.
[0041] The acquisition unit 111 acquires the position information of one or more sound objects. The sound object is assumed to be located within a predetermined range configured based on a plurality of first position information. Further, the acquisition unit 111 acquires information regarding the perceived position of the sound object.
[0042] ·Processing unit 112 The processing unit 112 has a function for controlling the processing of the information processing apparatus 10. As shown in FIG. 4, the processing unit 112 includes a determination unit 1121, a correction unit 1122, and a generation unit 1123. The determination unit 1121, the correction unit 1122, and the generation unit 1123 included in the processing unit 112 may each be configured as modules of independent computer programs, or may be configured as modules of a single integrated computer program.
[0043] ·Determination unit 1121 The determination unit 1121 has a function of determining the second position information. Here, the determination of the second position information will be described by taking the following two methods as examples.
[0044] (1) The user designates the perceived position The determination unit 1121 may determine the second position information based on the line-of-sight information based on the imaging information captured while the terminal device 30 is directed toward the direction of the sound object perceived by the user. Specifically, the determination unit 1121 may determine the second position information by holding the terminal device 30 in the direction where the sound reproduced by the headphones 20 is localized while facing the imaging member such as a camera in the terminal device 30 having an imaging function toward the user's face. In this case, the determination unit 1121 may determine the second position information by calculating from the angle of the user's face in which direction the user is holding the terminal device 30.
[0045] The determination unit 1121 may determine the second position information based on the geomagnetic information detected by the terminal device 30 while the terminal device 30 having a rod shape is directed toward the direction of the sound object perceived by the user. Specifically, the determination unit 1121 may determine the second position information by holding the rod-shaped terminal device 30 equipped with a geomagnetic sensor in the direction where the sound reproduced by the headphones 20 is localized. In this case, the determination unit 1121 may determine the second position information by calculating the sensor value of the geomagnetic sensor. In this way, the determination unit 1121 may determine the second position information based on the sensor information of the terminal device 30.
[0046] The determination unit 1121 may determine the second position information based on a method capable of specifying the position intended by the user, such as GUI (Graphical User Interface) software.
[0047] FIG. 5 shows an example of a process for determining second position information using GUI software. FIG. 5(A) shows a display screen of the terminal device 30 at the time of starting up the GUI software or the like. In FIG. 5(A), the position information of the user U11 and the first position information of virtual speakers (speakers SP11 to SP13) for which HRTF has been determined in advance are three-dimensionally drawn and displayed. Thereby, the user U11 can appropriately grasp the positions of the virtual speakers by changing the angle in various ways. In FIG. 5, similar to FIG. 3, the speaker SP11 is set as the virtual speaker to be reproduced. Also, in FIG. 5, the virtual speaker to be reproduced is represented by a thick-lined circle "○" that can be moved on the screen. Here, when the user U11 operates (for example, clicks or taps) the command BB11, the terminal device 30 transmits the operation information to the information processing device 10. The information processing device 10 transmits a signal obtained by convolving an acoustic signal such as white noise with the HRTF of the position of the virtual speaker to be reproduced to the headphones 20. Then, the headphones 20 reproduce based on the signal received from the information processing device 10.
[0048] FIG. 5(B) shows a display screen of the terminal device 30 when the user U11 operates (for example, moves by dragging or tapping) the position perceived by the reproduced sound from position A to position A prime. Here, the dotted-line circle "○" shown at position A indicates the position of the speaker SP11 before the operation. Also, the solid-line circle "○" shown at position A prime indicates the position of the speaker SP11 after the operation.
[0049] Figure 5(C) shows the display screen of the terminal device 30 when the user U11 operates the command BB12. In Figure 5(C), when the user U11 operates the command BB12, the virtual speaker to be played is switched to a different speaker. Specifically, the virtual speaker to be played is switched from the speaker SP11 to the speaker SP12. In this case, the thick-lined circle "○" that can be moved on the screen becomes a circle indicating the position from the position of the speaker SP11 to the position of the speaker SP12. Then, the same processing as that for the speaker SP11 is performed. Although not shown in Figure 5, when the user U11 operates the position perceived by the user with the signals convolved with the HRTF at each position for all the virtual speakers, the determination unit 1121 determines the second position information.
[0050] (2) The user adjusts the perceived position In FIG. 3, the case where the perceived position of the sound object TB11 becomes the position ☆' when rendering is performed with the virtual speaker located at the position A as the virtual speaker to be played has been described. Also, when rendering is performed with the virtual speaker located at the position A' as the virtual speaker to be played, the perceived position of the sound object TB11 may become the position ☆. Here, the reason why the perceived position becomes the position ☆ is that the gain of the virtual speaker located at the position A' is larger than the gains of the virtual speakers at the positions B and C. Thus, when the virtual speaker located at the position A is moved to the position A', the sound object TB11 located at the position ☆' moves to the position ☆. Note that when the direction from the position A to the position A' is the downward direction, the sound object TB11 moves in the upward direction from the position ☆' to the position ☆. From this, the determination unit 1121 may determine the second position information by causing the user to adjust the position such that the sound object TB11 moves to the position ☆ using GUI software or the like. The determination unit 1121 may determine the second position information by causing the user to move the virtual speaker manually (for example, by manual input). Hereinafter, this will be described with reference to FIGS. 6 to 8.
[0051] FIG. 6 is a diagram showing an overview of the functions of the information processing apparatus 10 according to the embodiment. Note that the same description as in FIGS. 2 and 3 will be omitted as appropriate. In FIG. 6, user U11 inputs downward operation information using member GU11 (for example, a screen or a device) (S11). Then, speaker SP11 moves from position A to position A prime so as to match the input of user U11 (S12). Then, the perceived position of sound object TB11 moves from position ☆ prime to position ☆ in accordance with the movement of speaker SP11 (S13). Note that member GU11 may be, for example, a perceived position adjustment button for adjusting the perceived position of sound object TB11. In this way, determination unit 1121 may determine the second position information based on an operation of the user's GUI that moves the first position information to the second position information. In this way, determination unit 1121 may determine the second position information based on the input information input by the user during the reproduction of the output signal. Here, since the perceived position of sound object TB11 moves in the direction opposite to that of speaker SP11, it may be difficult for user U11 to adjust.
[0052] FIG. 7 is a diagram showing an overview of the functions of the information processing apparatus 10 according to the embodiment. FIG. 7 is a modification of FIG. 6. Note that the same description as in FIG. 6 will be omitted as appropriate. In FIG. 7, user U11 inputs upward operation information using member GU11 (S21). Then, speaker SP11 moves from position A to position A prime in the direction opposite to the input of user U11 (S22). Note that step S23 is the same as step S13. In this case, the perceived position of sound object TB11 moves from position ☆ prime to position ☆ so as to match the input of user U11. As a result, in FIG. 7, since the perceived position moves so as to match the input of user U11, user U11 can adjust it in a natural sense. As a result, improvement of usability can be promoted. Hereinafter, the overview of the functions of FIGS. 6 and 7 will be described with reference to FIG. 8.
[0053] FIG. 8 shows the display screen of the terminal device 30 at the time of startup of the GUI software or the like. Note that the same description as in FIG. 5 will be omitted as appropriate. In FIG. 8, the position information of the user U11 is displayed. Also, in FIG. 8, a round "○" is displayed at a position inside a triangle formed by the first position information (positions A to C) of virtual speakers (speakers SP11 to SP13) for which HRTF has been determined in advance. Note that in FIG. 8, for convenience of explanation, cases where the signs of positions A to C are displayed are shown, but they may not actually be displayed. Here, when the user U11 operates the command BB11, a signal generated based on an acoustic signal such as white noise and the position information indicated by the round "○" is reproduced by the headphones 20. The user U11 adjusts using the member GU11 so that the position perceived by the sound reproduced by the headphones 20 becomes the position indicated by the round "○". The determination unit 1121 determines the second position information based on such adjustment by the user U11. Then, when the user U11 operates the command BB13, a round "○" is displayed at a position inside a triangle formed based on the first position information of different virtual speakers that do not use the vertices of the triangle formed by positions A to C as vertices. Then, the same processing as when a round "○" is displayed at a position inside the triangle formed by positions A to C is performed.
[0054] As described above, two methods for determining the second position information of the virtual speaker according to the embodiment have been described as examples, but the present invention is not limited to these examples. For example, the determination unit 1121 may perform processing using a method in which conventional techniques are appropriately combined.
[0055] · Correction unit 1122 The correction unit 1122 has a function of rendering audio data including the position information of the sound object to a plurality of virtual speakers virtually arranged in space. Further, the correction unit 1122 corrects the first position information of at least one of the plurality of virtual speakers based on the second position information. Alternatively, the correction unit 1122 corrects the first position information of at least one of the plurality of virtual speakers based on the difference between the first position information and the second position information. For example, the correction unit 1122 corrects the first position information based on the second position information determined by the determination unit 1121. Also, for example, the correction unit 1122 corrects the first position information so that the perceived position of the sound object perceived by the user becomes a predetermined position based on the position information of the sound object.
[0056] Note that the calculation of the difference between the first position information and the second position information is, for example, performed by the correction unit 1122. For example, the correction unit 1122 calculates the difference based on the comparison of the coordinate information indicating the position information. Also, for example, the correction unit 1122 corrects the first position information based on the distance information indicating the difference.
[0057] The correction unit 1122 may correct the first position information so that the larger the difference between the first position information and the second position information, the larger the correction amount of the perceived position of the sound object. For example, the correction unit 1122 may correct the first position information based on the correction amount of the perceived position of the sound object predetermined according to the difference between the first position information and the second position information.
[0058] The correction unit 1122 may correct the first position information of the virtual speaker to be reproduced based on the perceived position of the sound object included in a predetermined range configured based on the first position information of the plurality of virtual speakers. For example, the correction unit 1122 may correct the first position information of the virtual speaker to be reproduced based on the perceived position of the sound object included in the range of a triangle configured based on the first position information of three virtual speakers.
[0059] · Generation unit 1123 The generation unit 1123 has a function of generating sound for playback. For example, the generation unit 1123 generates sound for playback by adding the sounds of all virtual speakers.
[0060] The generation unit 1123 generates an output signal for each audio output unit based on the HRTF of the user from the speaker signals for each virtual speaker generated by the correction unit 1122. For example, the generation unit 1123 may generate an output signal for each audio output unit based on the HRTF estimated from imaging information such as the user's ear image. Also, for example, the generation unit 1123 may generate an output signal for each audio output unit based on the average HRTF calculated from the HRTFs of a plurality of users.
[0061] For each virtual speaker, the generation unit 1123 generates a speaker signal by rendering with VBAP using the second position information as the first position information. Also, for each virtual speaker, the generation unit 1123 applies the pre-held HRTF to the speaker signal to generate an output signal for each virtual speaker. Then, for each virtual speaker, the generation unit 1123 adds the output signals for each virtual speaker for each L and R signal to generate an output signal.
[0062] ·Output unit 113 The output unit 113 has a function of outputting the correction result by the correction unit 1122. The output unit 113 provides information regarding the correction result to, for example, the terminal device 30 via the communication unit 100. When the terminal device 30 receives the output information provided from the output unit 113, it displays the output information via the output unit 320. The output unit 113 may provide control information for displaying the output information. Also, the output unit 113 may generate output information for the terminal device 30 to display information regarding the correction result.
[0063] The output unit 113 has a function of outputting the generation result by the generation unit 1123. The output unit 113 provides information regarding the generation result to, for example, the headphones 20 via the communication unit 100. For example, the output unit 113 provides an output signal for each audio output unit. Specifically, the output unit 113 provides an output signal obtained by adding the speaker signals for each virtual speaker for each of the L and R signals. When the headphones 20 receive the output information provided from the output unit 113, the headphones 20 output the output information via the output unit 220. The output unit 113 may provide control information for outputting the output information. Further, the output unit 113 may generate output information for outputting information regarding the generation result to the headphones 20.
[0064] (1-3) Memory unit 120 The memory unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The memory unit 120 has a function of storing a computer program and data (including one format of the program) related to the processing in the information processing apparatus 10.
[0065] FIG. 9 shows an example of the memory unit 120. The memory unit 120 shown in FIG. 9 stores the first position information of the virtual speaker. As shown in FIG. 9, the memory unit 120 may have items such as "virtual speaker ID", "user ID", "virtual speaker position", and "HRTF".
[0066] "Virtual Speaker ID" indicates identification information for identifying a virtual speaker. "User ID" indicates identification information for identifying a user. "Virtual Speaker Position" indicates the first position information of the virtual speaker. In the example shown in FIG. 9, it shows a case where conceptual information such as "Virtual Speaker Position #11" or "Virtual Speaker Position #12" is stored in "Virtual Speaker Position", but actually, coordinate information or information indicating the relative position to other virtual speakers may be stored. "HRTF" indicates a predetermined HRTF based on the first position information of the virtual speaker. In the example shown in FIG. 9, it shows a case where conceptual information such as "HRTF #11" or "HRTF #12" is stored in "HRTF", but actually, HRTF data measured by a microphone near the user's ear or the like is stored.
[0067] (2) Headphone 20 As shown in FIG. 4, the headphone 20 includes a communication unit 200, a control unit 210, and an output unit 220.
[0068] (2-1) Communication Unit 200 The communication unit 200 has a function of communicating with an external device. For example, in communication with an external device, the communication unit 200 outputs information received from the external device to the control unit 210. Specifically, the communication unit 200 outputs information received from the information processing device 10 to the control unit 210. For example, the communication unit 200 outputs information regarding the acquisition of information related to the sound for reproduction to the control unit 210. For example, the communication unit 200 outputs information regarding the acquisition of the output signal for each voice output unit to the control unit 210.
[0069] (2-2) Control Unit 210 The control unit 210 has a function of controlling the operation of the headphone 20. For example, based on the information transmitted from the information processing device 10 via the communication unit 200, the control unit 210 performs processing for reproducing sound. For example, the control unit 210 performs processing for outputting an output signal.
[0070] (2-3) Output Unit 220 The output unit 220 is realized by a member capable of outputting sound such as a speaker. The output unit 220 outputs sound. For example, the output unit 220 outputs an output signal.
[0071] (3) Terminal device 30 As shown in FIG. 4, the terminal device 30 includes a communication unit 300, a control unit 310, and an output unit 320.
[0072] (3-1) Communication unit 300 The communication unit 300 has a function of communicating with an external device. For example, in communication with an external device, the communication unit 300 outputs information received from the external device to the control unit 310. Specifically, the communication unit 300 outputs information regarding the correction result received from the information processing device 10 to the control unit 310.
[0073] (3-2) Control unit 310 The control unit 310 has a function of controlling the overall operation of the terminal device 30. For example, the control unit 310 performs a process of controlling the output of information regarding the correction result. Also, for example, the control unit 310 performs a process for moving the virtual speaker to be played back according to an operation by the user. Also, for example, the control unit 310 performs a process for moving the perceived position of the sound object perceived by the user according to the movement of the virtual speaker to be played back.
[0074] (3-3) Output unit 320 The output unit 320 has a function of outputting information regarding the correction result. The output unit 320 outputs the output information provided from the output unit 113 via the communication unit 300. For example, the output unit 320 displays the output information on the display screen of the terminal device 30. Also, the output unit 320 may output the output information based on the control information provided from the output unit 113.
[0075] The output unit 320 displays output information according to an operation by the user. For example, the output unit 320 displays information regarding the position information of the virtual speaker to be played back and the sound object.
[0076] <2.3. Processing of the Information Processing System> The functions of the information processing system 1 according to the embodiment have been described above. Next, the processing of the information processing system 1 will be described.
[0077] FIG. 10 is a flowchart showing the processing flow in the information processing apparatus 10 according to the embodiment. The information processing apparatus 10 acquires the first position information of the virtual speaker (S101). Further, the information processing apparatus 10 acquires the second position information of the virtual speaker (S102). Next, the information processing apparatus 10 calculates the difference between the first position information and the second position information (S103). For example, the information processing apparatus 10 calculates the difference based on the comparison of the coordinate information. Then, the information processing apparatus 10 corrects the first position information based on the calculated difference (S104). For example, the information processing apparatus 10 corrects the first position information so that the perceived position of the sound object perceived by the user becomes a predetermined position based on the position information of the sound object based on the calculated difference.
[0078] FIG. 11 is a flowchart showing the processing flow in the information processing apparatus 10 according to the embodiment. The information processing apparatus 10 determines whether it has received a designation from the user for all target virtual speakers (S201). When the information processing apparatus 10 determines that it has received a designation from the user for all virtual speakers (S201; YES), it ends the information processing. Also, when the information processing apparatus 10 determines that it has not received a designation from the user for all virtual speakers (S201; NO), it determines one of the un-designated virtual speakers as the reproduction target virtual speaker (S202). Next, the information processing apparatus 10 generates an output signal by convolving the HRTF of the reproduction target virtual speaker with white noise or the like (S203). Also, the information processing apparatus 10 performs processing for reproducing the output signal with headphones or the like (S204). Next, the information processing apparatus 10 designates the perceived position that the user perceives with the output signal reproduced with headphones or the like, and performs processing for shifting to another virtual speaker (S205). Specifically, when the user designates the perceived position and accepts an operation such as a "next" button, the information processing apparatus 10 performs processing for shifting to another virtual speaker. Then, it returns to the processing of step S201.
[0079] <2.4. Variations of Processing> The embodiments of the present disclosure have been described above. Subsequently, variations of the processing of the embodiments of the present disclosure will be described. Note that the variations of the processing described below may be applied to the embodiments of the present disclosure alone or in combination. Also, the variations of the processing may be applied in place of the configurations described in the embodiments of the present disclosure, or may be additionally applied to the configurations described in the embodiments of the present disclosure.
[0080] In the above-described embodiment, the general functions of the information processing apparatus 10 will be described when the number of input sound objects is N and the number of virtual speakers is M. Note that N, where there are N sound objects, may be any integer equal to or greater than 1, and M, where there are M virtual speakers, may be any integer equal to or greater than 2. FIG. 12 is a diagram showing the general functions of the information processing apparatus 10 according to a modification of the embodiment. In the above-described embodiment, as shown in FIG. 4, the processing unit 112 has been shown as having a determination unit 1121, a correction unit 1122, and a generation unit 1123. Here, as shown in FIG. 12, in addition to the configuration shown in FIG. 4, the processing unit 112 may have a user perception acquisition unit 1124, a virtual speaker rendering unit 1125, an HRTF processing unit 1126, and an addition unit 1127. The determination unit 1121, correction unit 1122, generation unit 1123, user perception acquisition unit 1124, virtual speaker rendering unit 1125, HRTF processing unit 1126, and addition unit 1127 included in the processing unit 112 may each be configured as an independent computer program module, or a plurality of functions may be configured as one integrated computer program module.
[0081] The user perception acquisition unit 1124 acquires information (second position information) regarding the perception position at which a signal to which the held HRTF is applied is perceived by the user for each of the M virtual speakers. Then, the user perception acquisition unit 1124 provides the acquired second position information to the virtual speaker rendering unit 1125 (S31).
[0082] For each of the N sound objects, the virtual speaker rendering unit 1125 performs rendering processing using VBAP with the second position information acquired by the user perception acquisition unit 1124 as the first position information, and generates N×M signals (hereinafter, appropriately referred to as "virtual speaker rendering signals"). Also, for each of the virtual speakers, the virtual speaker rendering unit 1125 adds the N virtual speaker rendering signals for each sound object. Then, the virtual speaker rendering unit 1125 provides the resulting M speaker signals to the HRTF processing unit 1126 (S32).
[0083] For each virtual speaker, the HRTF processing unit 1126 applies the HRTF stored in advance to the speaker signal provided by the virtual speaker rendering unit 1125. Then, the HRTF processing unit 1126 provides the output signal for each of the M virtual speakers (for example, a headphone signal) as a result to the addition unit 1127 (S33).
[0084] For each virtual speaker, the addition unit 1127 adds the output signal for each virtual speaker provided by the HRTF processing unit 1126 for each of the L and R signals. Then, the addition unit 1127 performs processing for outputting the output signal (S34).
[0085] <<3. Hardware Configuration Example>> Finally, with reference to FIG. 13, a hardware configuration example of the information processing apparatus according to the embodiment will be described. FIG. 13 is a block diagram showing a hardware configuration example of the information processing apparatus according to the embodiment. Note that the information processing apparatus 900 shown in FIG. 13 can realize, for example, the information processing apparatus 10, the headphones 20, and the terminal device 30 shown in FIG. 4. Information processing by the information processing apparatus 10, the headphones 20, and the terminal device 30 according to the embodiment is realized by cooperation between software (configured by a computer program) and the hardware described below.
[0086] As shown in FIG. 13, the information processing apparatus 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903. Further, the information processing apparatus 900 includes a host bus 904a, a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 910, and a communication device 911. Note that the hardware configuration shown here is an example, and some of the components may be omitted. Also, the hardware configuration may further include components other than the components shown here.
[0087] The CPU 901 functions as, for example, an arithmetic processing unit or a control unit, and controls the overall operation or a part of the operation of each component based on various computer programs recorded in the ROM 902, the RAM 903, or the storage device 908. The ROM 902 is a means for storing programs read by the CPU 901, data used for arithmetic operations, and the like. In the RAM 903, for example, programs read by the CPU 901 and data such as various parameters that change as appropriate when executing the programs (a part of the programs) are stored temporarily or permanently. These are interconnected by a host bus 904a composed of a CPU bus or the like. The CPU 901, the ROM 902, and the RAM 903 can realize the functions of the control unit 110, the control unit 210, and the control unit 310 described with reference to FIG. 4, for example, in cooperation with software.
[0088] The CPU 901, the ROM 902, and the RAM 903 are interconnected via, for example, a host bus 904a capable of high-speed data transmission. On the other hand, the host bus 904a is connected to an external bus 904b with a relatively low data transmission speed via, for example, a bridge 904. Further, the external bus 904b is connected to various components via an interface 905.
[0089] The input device 906 is realized by, for example, a device through which information is input by a listener, such as a mouse, a keyboard, a touch panel, a button, a microphone, a switch, and a lever. Further, the input device 906 may be, for example, a remote control device using infrared rays or other radio waves, or an external connection device such as a mobile phone or a PDA corresponding to the operation of the information processing device 900. Furthermore, the input device 906 may include, for example, an input control circuit that generates an input signal based on the information input using the above input means and outputs it to the CPU 901. The administrator of the information processing device 900 can input various data to the information processing device 900 or instruct a processing operation by operating the input device 906.
[0090] In addition, the input device 906 can be formed by a device that detects the user's position. For example, the input device 906 can include various sensors such as an image sensor (e.g., a camera), a depth sensor (e.g., a stereo camera), an acceleration sensor, a gyro sensor, a geomagnetic sensor, a light sensor, a sound sensor, a distance measurement sensor (e.g., a ToF (Time of Flight) sensor), a force sensor, etc. Further, the input device 906 may acquire information regarding the state of the information processing device 900 itself, such as the posture and moving speed of the information processing device 900, and information regarding the surrounding space of the information processing device 900, such as the brightness and noise around the information processing device 900. Also, the input device 906 may include a GNSS module that receives GNSS signals from GNSS (Global Navigation Satellite System) satellites (e.g., GPS (Global Positioning System) signals from GPS satellites) and measures position information including the latitude, longitude, and altitude of the device. Regarding the position information, the input device 906 may detect the position by transmitting and receiving with Wi-Fi (registered trademark), mobile phones, PHS, smartphones, etc., or by short-range communication, etc. The input device 906 can, for example, realize the functions of the acquisition unit 111 described with reference to FIG. 4.
[0091] The output device 907 is formed by a device capable of notifying the user of the acquired information visually or auditorily. Such devices include display devices such as CRT display devices, liquid crystal display devices, plasma display devices, EL display devices, laser projectors, LED projectors, and lamps, acoustic output devices such as speakers and headphones, and printer devices. The output device 907 outputs, for example, the results obtained by various processes performed by the information processing device 900. Specifically, the display device visually displays the results obtained by various processes performed by the information processing device 900 in various forms such as text, images, tables, graphs, etc. On the other hand, the audio output device converts an audio signal composed of reproduced audio data, acoustic data, etc. into an analog signal and outputs it auditorily. The output device 907 can, for example, realize the functions of the output unit 113, the output unit 220, and the output unit 320 described with reference to FIG. 4.
[0092] The storage device 908 is a data storage device formed as an example of the storage unit of the information processing device 900. The storage device 908 is realized by, for example, a magnetic storage unit device such as an HDD, a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 may include a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, and a deleting device for deleting data recorded on the storage medium. This storage device 908 stores a computer program executed by the CPU 901, various data, and various data acquired from the outside. The storage device 908 can realize the functions of the storage unit 120 described with reference to FIG. 4, for example.
[0093] The drive 909 is a reader / writer for a storage medium and is built in or externally attached to the information processing device 900. The drive 909 reads information recorded on a removable storage medium such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory that is mounted and outputs it to the RAM 903. The drive 909 can also write information to the removable storage medium.
[0094] The connection port 910 is a port for connecting an external connection device such as, for example, a USB (Universal Serial Bus) port, an IEEE 1394 port, a SCSI (Small Computer System Interface), an RS-232C port, or an optical audio terminal.
[0095] The communication device 911 is, for example, a communication interface formed by a communication device or the like for connecting to the network 920. The communication device 911 is, for example, a communication card for wired or wireless LAN (Local Area Network), LTE (Long Term Evolution), Bluetooth (registered trademark), or WUSB (Wireless USB). Further, the communication device 911 may be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various communications. This communication device 911 can transmit and receive signals, etc. in accordance with a predetermined protocol such as TCP / IP, for example, between the Internet and other communication devices. The communication device 911 can realize the functions of the communication unit 100, the communication unit 200, and the communication unit 300 described with reference to FIG. 4, for example.
[0096] Note that the network 920 is a wired or wireless transmission path for information transmitted from a device connected to the network 920. For example, the network 920 may include a public line network such as the Internet, a telephone line network, a satellite communication network, or various LANs (Local Area Networks) including Ethernet (registered trademark), WANs (Wide Area Networks), etc. Further, the network 920 may include a dedicated line network such as an IP-VPN (Internet Protocol-Virtual Private Network).
[0097] As described above, an example of the hardware configuration capable of realizing the functions of the information processing apparatus 900 according to the embodiment has been shown. Each of the above components may be realized using general-purpose members, or may be realized by hardware specialized for the functions of each component. Therefore, it is possible to appropriately change the hardware configuration to be used according to the technical level at the time of implementing the embodiment.
[0098] <<4. Summary>> As described above, the information processing apparatus 10 according to the embodiment performs a process for correcting the first position information based on the second position information. Further, the information processing apparatus 10 corrects the first position information so that the perceived position of the sound object perceived by the user becomes a predetermined position based on the position information of the sound object. Thereby, since the information processing apparatus 10 can localize the sound image of the sound object at the intended position, it is possible to promote the improvement of the sound quality when reproducing the sound image.
[0099] Therefore, it is possible to provide a novel and improved information processing apparatus, information processing method, and terminal device that can promote further improvement in user usability.
[0100] As described above, the preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings, but the technical scope of the present disclosure is not limited to such examples. It is obvious that those having ordinary knowledge in the technical field of the present disclosure can conceive of various modification examples or correction examples within the scope of the technical idea described in the claims, and these are also naturally understood to belong to the technical scope of the present disclosure.
[0101] For example, each device described in this specification may be realized as a single device, or part or all of them may be realized as separate devices. For example, the information processing apparatus 10, the headphones 20, and the terminal device 30 shown in FIG. 4 may each be realized as a single device. Further, for example, the information processing apparatus 10, the headphones 20, and the terminal device 30 may be realized as a server device connected to a network or the like. Further, the function of the control unit 110 included in the information processing apparatus 10 may be a configuration included in a server device connected to a network or the like.
[0102] In addition, a series of processes performed by each apparatus described in this specification may be implemented using any of software, hardware, and a combination of software and hardware. A computer program constituting the software is pre-stored, for example, in a recording medium (non-transitory media) provided inside or outside each apparatus. Then, each program is read into the RAM during execution by a computer, for example, and executed by a processor such as a CPU.
[0103] In addition, the processes described using flowcharts in this specification do not necessarily have to be executed in the illustrated order. Some process steps may be executed in parallel. Also, additional process steps may be adopted, and some process steps may be omitted.
[0104] Also, the effects described in this specification are merely illustrative or exemplary and not limiting. That is, the technology according to the present disclosure may exhibit other effects apparent to those skilled in the art from the description in this specification, together with or instead of the above effects.
[0105] Note that the following configurations also belong to the technical scope of the present disclosure. (1) A correction unit that renders audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space, An acquisition unit that acquires first position information regarding the virtual positions of the virtual speakers in the space and second position information regarding the positions of the virtual speakers in the space perceived by a user, Comprising The correction unit Corrects the first position information of at least one of the plurality of virtual speakers based on the second position information An information processing apparatus. (2) The apparatus further includes a generation unit that generates an output signal for each voice output unit based on the head transfer function of the user from the speaker signals for each virtual speaker generated by the correction unit. The acquisition unit acquires the second position information of the virtual speaker based on input information input by the user during reproduction of the output signal from the voice output unit. The information processing apparatus according to (1) above. (3) The generation unit generates the output signal for each voice output unit based on the head transfer function estimated from the ear image of the user. The information processing apparatus according to (2) above. (4) The generation unit generates the output signal for each voice output unit based on an average head transfer function calculated from the head transfer functions of a plurality of users. The information processing apparatus according to (2) above. (5) The correction unit corrects the first position information so that the perceived position of the sound object perceived by the user becomes a predetermined position based on the position information of the sound object. The information processing apparatus according to any one of (1) to (4) above. (6) The apparatus further includes a determination unit that determines the second position information, and the correction unit corrects the first position information based on the second position information determined by the determination unit. The information processing apparatus according to any one of (1) to (5) above. (7) The determination unit determines the second position information based on the gaze information based on the imaging information obtained by imaging the user while the user faces the terminal device in the direction of the sound object perceived by the user. The information processing apparatus according to (6) above. (8) The determination unit While facing the terminal device with a rod-shaped shape in the direction of the sound object perceived by the user, the second position information is determined based on the geomagnetic information detected by the terminal device. The information processing apparatus according to (6) above. (9) The determination unit Based on an operation of the user's GUI (Graphical User Interface) that moves the first position information to the second position information, the second position information is determined. The information processing apparatus according to (6) above. (10) The determination unit Based on the movement of the virtual speaker in the opposite direction of the operation, the second position information is determined. The information processing apparatus according to (9) above. (11) The sound object is included in a predetermined range configured based on a plurality of the first position information. The information processing apparatus according to any one of (1) to (10) above. (12) An information processing method executed by a computer, A correction step of rendering audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space, An acquisition step of acquiring first position information regarding a virtual position of the virtual speaker in the space and second position information regarding a position of the virtual speaker in the space perceived by the user, Comprising The correction step Based on the second position information, correcting the first position information of at least one of the plurality of virtual speakers. An information processing method including. (13) An output unit that outputs output information corresponding to an operation of moving first position information regarding a virtual position in space of a virtual speaker provided from an information processing apparatus to second position information regarding the position of the virtual speaker in the space that a user perceives. A terminal device comprising: The information processing apparatus corrects at least one of the first position information of a plurality of virtual speakers that have rendered audio data including position information of sound objects based on the second position information. A terminal device characterized by that.
Explanation of symbols
[0106] 1 Information processing system 10 Information processing apparatus 20 Headphones 30 Terminal device 100 Communication unit 110 Control unit 111 Acquisition unit 112 Processing unit 1121 Decision unit 1122 Correction unit 1123 Generation unit 1124 User perception acquisition unit 1125 Virtual speaker rendering unit 1126 HRTF processing unit 1127 Addition unit 113 Output unit 200 Communication unit 210 Control unit 220 Output unit 300 Communication unit 310 Control unit 320 Output unit
Claims
1. A correction unit that renders audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space; An acquisition unit that acquires first position information regarding a virtual position of the virtual speaker in the space and second position information regarding a position of the virtual speaker in the space perceived by a user; A determination unit that determines the second position information based on line-of-sight information based on imaging information obtained by imaging the user while the user faces the terminal device in the direction of the sound object perceived by the user, or based on geomagnetic information detected by the terminal device while the user faces a rod-shaped terminal device in the direction of the sound object perceived by the user; Comprising: The correction unit: Based on the second position information determined by the determination unit, corrects the first position information of at least one of the plurality of virtual speakers An information processing apparatus.
2. A correction unit that renders audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space; An acquisition unit that acquires first position information regarding a virtual position of the virtual speaker in the space and second position information regarding a position of the virtual speaker in the space perceived by a user; A determination unit that determines the second position information based on movement of the virtual speaker in a direction opposite to an operation of moving the first position information to the second position information, which is an operation of the user's GUI (Graphical User Interface); Comprising: The correction unit: Based on the second position information determined by the determination unit, corrects the first position information of at least one of the plurality of virtual speakers An information processing apparatus.
3. Further comprising a generation unit that generates an output signal for each audio output unit based on a head transfer function of the user from the speaker signal for each virtual speaker generated by the correction unit; The acquisition unit: Acquires the second position information of the virtual speaker based on input information input by the user during reproduction of the output signal from the audio output unit The information processing apparatus according to claim 1 or 2.
4. The generation unit: Generates the output signal for each audio output unit based on a head transfer function estimated from an ear image of the user The information processing apparatus according to claim 3.
5. The generation unit: Generate the output signal for each of the voice output units based on the average head transfer function calculated from the head transfer functions of a plurality of users The information processing apparatus according to claim 3
6. The correction unit Correct the first position information so that the perceived position of the sound object perceived by the user is a predetermined position based on the position information of the sound object The information processing apparatus according to claim 1 or 2
7. The determination unit Determine the second position information based on the line-of-sight information based on the imaging information obtained by imaging the user while facing the terminal device in the direction of the sound object perceived by the user The information processing apparatus according to claim 2
8. The determination unit Determine the second position information based on the geomagnetic information detected by the terminal device while facing the terminal device having a rod shape in the direction of the sound object perceived by the user The information processing apparatus according to claim 2
9. The determination unit Determine the second position information based on an operation of the user's GUI (Graphical User Interface) that moves the first position information to the second position information The information processing apparatus according to claim 1
10. The determination unit Determine the second position information based on the movement of the virtual speaker in the opposite direction of the operation The information processing apparatus according to claim 9
11. The sound object is included in a predetermined range configured based on the plurality of first position information The information processing apparatus according to claim 1 or 2
12. An information processing method executed by a computer, comprising A correction step of rendering audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space An acquisition step of acquiring first position information regarding a virtual position of the virtual speaker in the space and second position information regarding a position of the virtual speaker in the space perceived by the user A determination step of determining the second position information based on line-of-sight information based on imaging information obtained by imaging the user while facing the terminal device in the direction of the sound object perceived by the user, or based on geomagnetic information detected by the terminal device while facing the terminal device having a rod shape in the direction of the sound object perceived by the user Comprising The correction step Based on the second position information determined by the determination process, correcting the first position information of at least one of the plurality of virtual speakers An information processing method including the above.
13. An information processing method executed by a computer, comprising: A correction process of rendering audio data including position information of a sound object to a plurality of virtual speakers virtually arranged in space; An acquisition process of acquiring first position information regarding a virtual position of the virtual speaker in the space and second position information regarding a position of the virtual speaker in the space perceived by a user; A determination process of determining the second position information based on movement of the virtual speaker in a direction opposite to an operation of moving the first position information to the second position information, which is an operation of the user's GUI (Graphical User Interface); The information processing method comprising: The correction process includes: Based on the second position information determined by the determination process, correcting the first position information of at least one of the plurality of virtual speakers An information processing method including the above.
Citation Information
Patent Citations
Voice control device, voice control method, and program
JP2013101248A
Method for generating customized spatial audio with head tracking
JP2019146160A
Signal processor, acoustic processing system, and program
JP2020088632A
Sound processing device and method, and program
WO2015107926A1
Information processing device, information processing method, and information processing program
WO2020075622A1