A program and method for generating virtual acoustic signals while considering the listener's orientation, and a virtual acoustic manifestation device.

The virtual sound generation program and device address audiovisual inconsistencies by adjusting sound output based on viewer orientation, ensuring synchronized audio-visual experience through multiple speakers and sensors.

JP7839762B2Active Publication Date: 2026-04-02KDDI CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing technologies fail to accurately localize sound output based on changes in the viewer's head orientation, leading to audiovisual inconsistencies when viewing multimedia content on devices with two loudspeakers, as they do not account for the viewer's head direction.

Method used

A virtual sound generation program and device that utilizes multiple speakers, a camera, and an attitude sensor to generate acoustic signals from virtual sound sources, adjusting sound output based on the viewer's head orientation and screen orientation, ensuring synchronized audio-visual experience.

Benefits of technology

The solution effectively localizes sound output from virtual sound sources, aligning it with the viewer's head orientation, thereby eliminating audiovisual inconsistencies and providing a synchronized audio-visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839762000008
    Figure 0007839762000008
  • Figure 0007839762000009
    Figure 0007839762000009
  • Figure 0007839762000010
    Figure 0007839762000010
Patent Text Reader

Abstract

To provide a program capable of generating an acoustic signal of a virtual sound source while taking into account a direction of a listener.SOLUTION: A virtual sound generation program generates virtual sound signals for realizing sounds output from a plurality of virtual sound sources set in accordance with a plurality of input sound signals. Specifically, the program causes a computer to function as: reproduction function generating means for generating a reproduction transfer function for reproduction by a plurality of installed speakers such that a listener with the head oriented in a certain direction perceives the sound from the speakers as sound from a virtual sound source, using a head transfer function of the listener on the basis of obtained direction information relating to a direction of the listener's head or face relative to the speakers or a device including the speakers; and virtual signal generating means for generating a virtual sound signal to be input to the speakers from the sound signals using the generated reproduction transfer function.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a technique for generating acoustic signals from a virtual sound source. [Background technology]

[0002] In recent years, information devices such as smartphones have seen significant increases in storage capacity, allowing for the storage and playback of a large amount of multimedia content, including audiovisual information such as movies and promotional videos.

[0003] As a typical example, when playing multimedia content on a smartphone, the display, which is typically oriented vertically when the device is held vertically (the default orientation), and the loudspeaker, which is positioned horizontally when the device is held vertically, both output video and audio / sound in sync with each other.

[0004] Therefore, for example, when playing multimedia content containing audiovisual information adapted to a vertical display, it is preferable to view it with the device held vertically, which is the default orientation. On the other hand, when playing multimedia content containing audiovisual information adapted to a horizontal display, it is usually necessary to hold the smartphone horizontally to make the display horizontal (and to rotate the image 90 degrees within the display to match this orientation) for viewing.

[0005] However, in this case, the loudspeakers are positioned vertically, resulting in a mismatch between the position of subjects and objects perceived in the video and the localization of the sound image perceived in the audio / sound system. This creates a sense of unease (a feeling of audiovisual inconsistency) for the viewer. This problem also occurs in multimedia-capable mobile devices, which are typically held horizontally by default.

[0006] To address the problem of discrepancies between the placement direction of sound sources set in the video and the actual placement direction of loudspeakers, i.e., the localization of the sound image in relation to the subject or representation in the video and audio, Patent Document 1 proposes a technique in which three loudspeakers are placed in the upper center, lower left, and lower right of the display, and the orientation of the information device is detected to select two loudspeakers suitable for stereo playback. Patent Document 2 discloses a technique in which three loudspeakers are placed in the three corners of the display, and the orientation of the information device is detected to select two loudspeakers suitable for stereo playback.

[0007] Furthermore, the technology disclosed in Patent Document 3 attempts to solve the above problem by installing two loudspeakers at opposing corners of the display, detecting the orientation of the information device, and appropriately switching between the left and right channels. Patent Document 4 also discloses a multimedia device that filters the input signal so that the sound output from each installed loudspeaker matches the image displayed on the display at the viewing position. [Prior art documents] [Patent Documents]

[0008] [Patent Document 1] Japanese Patent Publication No. 2006-174277 [Patent Document 2] Japanese Patent Publication No. 2009-232424 [Patent Document 3] Japanese Patent Publication No. 2003-078601 [Patent Document 4] Japanese Patent Publication No. 2015-070324 [Patent Document 5] U.S. Patent Application Publication No. 2020 / 0104033 [Overview of the project] [Problems that the invention aims to solve]

[0009] However, with the conventional technologies described above, it remains difficult to properly localize the output sound and audio in accordance with changes in the viewer's head orientation, and therefore changes in the position of their ears. Consequently, for example, if a viewer changes the direction of their face while watching, a problem arises where the sound image localization between the video and the audio / audio no longer matches.

[0010] For example, the technologies disclosed in Patent Documents 1 and 2 do indeed take into account the orientation of the detected information device and select two out of three loudspeakers that are closest to the direction matching the image, thereby bringing the sound image localization of the image and audio closer together. However, no information regarding the orientation of the viewer's head is acquired, and therefore, it is not considered at all in the localization process. Consequently, if the orientation of the viewer's head changes and the position of their ears moves outside the so-called sweet spot, localization of the audio becomes impossible. Furthermore, these technologies require three loudspeakers and cannot be applied to systems based on two loudspeakers, which are the case for many mobile devices.

[0011] Furthermore, the technology disclosed in Patent Document 3 also does not take into account information related to the direction of the viewer's head, resulting in the same problems as described above. Moreover, in this technology, two loudspeakers are placed on the diagonal of the screen, and localization is performed along this diagonal, so if the information equipment is large, integrating audiovisual information becomes extremely difficult. Furthermore, the technology disclosed in Patent Document 4 also does not take into account information related to the direction of the viewer's head, resulting in the problem that an accurate filter cannot be generated if the direction of the viewer's head changes. As a result, even if the input signal is corrected using the generated filter, it becomes difficult to realize a virtual loudspeaker set to the desired position.

[0012] Furthermore, the technology disclosed in Patent Document 5 determines the orientation of the user's face from their facial image and then determines the display direction of the smartphone screen according to the determined facial orientation. However, this technology has not yet solved the problem described above, where the intended placement direction of the sound source in the video does not match the actual placement direction of the loudspeaker.

[0013] Therefore, the present invention aims to provide a program and method capable of generating acoustic signals of a virtual sound source while taking into account the direction of the listener, as well as a virtual sound realization device. [Means for solving the problem]

[0014] According to the present invention, a virtual sound generation program generates virtual sound signals to realize sounds output from multiple virtual sound sources set according to multiple input sound signals, A playback transfer function for playback using multiple installed speakers, wherein the sound produced by the speakers is perceived as sound from a virtual sound source by a listener whose head is positioned in a certain direction. to produce A means for generating a reproduction function, A virtual signal generation means that generates the virtual acoustic signal to be input to the speaker from the acoustic signal using the said reproduction transfer function. to make the computer work 、 Multiple speakers are installed in a portable device for the listener, which includes a screen capable of displaying video or images related to video or image signals input in accordance with the audio signal, a camera capable of generating image information including the listener and distance information related to the distance to the listener, and an attitude sensor capable of generating screen orientation information related to the orientation of the screen. Multiple virtual sound sources are set up within the device in accordance with the orientation of the multiple sound sources assumed in the multiple sound signals, which is the orientation within the video or image displayed on the screen. The means for generating the regeneration function is, From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. Characterized by A virtual sound generation program is provided.

[0017] Ma Ta, Multiple speakers are positioned laterally within the device in its default orientation. The video or image in question shows the device in its default position. Ra pictureIf the device is tilted 90 degrees within a plane that includes the surface, it will be displayed on the screen rotated 90 degrees in the opposite direction to the tilt of the device. It is also preferable that the above-mentioned multiple virtual sound sources be set within the device in a direction within the video or image that has been rotated 90 degrees, and in line with the arrangement direction of the assumed multiple sound sources.

[0020] Sara again The raw function generation means may also generate the regenerated transfer function only when the screen orientation information satisfies predetermined conditions.

[0022] According to the present invention, sound is also realized from multiple virtual sound sources set in accordance with multiple input sound signals. , portable for the listener A virtual sound manifestation device, Multiple speakers, A screen capable of displaying video or images related to video or image signals input in accordance with the said audio signal, A camera capable of generating image information including the listener and distance information relating to the distance to the listener, A posture sensor capable of generating screen orientation information related to the screen orientation, Multiple A reproduction transfer function for playback using a number of speakers, wherein the sound produced by the speakers is produced when the head is facing a certain direction. The A playback transfer function that allows the listener to perceive the sound as being produced by the virtual sound source. to produce A means for generating a reproduction function, Using the said reproduction transfer function, from the said acoustic signal, To realize sound output from multiple virtual sound sources A virtual signal generation means for generating the said virtual acoustic signal and It has, Multiple speakers use the virtual acoustic signal as input. , multiple Outputs sound related to the sound produced from a number of virtual sound sources. death, Multiple virtual sound sources are set up within the virtual sound embodiment device in accordance with the orientation of the multiple sound sources assumed in the multiple sound signals, which is the orientation within the video or image displayed on the screen. The means for generating the regeneration function is, From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or virtual sound embodiment device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. A virtual sound manifestation device is provided, characterized by the following features.

[0025] According to the present invention, the sound output from multiple virtual sound sources set in accordance with multiple input acoustic signals is further realized. Equipped with a virtual sound manifestation device that can be carried by the listener. It is a virtual sound manifestation system, Installed in the virtual sound manifestation device Multiple speakers, A screen provided in a virtual sound manifestation device that is capable of displaying video or images related to video or image signals input in accordance with the said audio signal, A camera installed in a virtual sound manifestation device that is capable of generating image information including the listener and distance information relating to the distance to the listener, A posture sensor provided in a virtual sound manifestation device capable of generating screen orientation information related to the orientation of the screen, Multiple A reproduction transfer function for playback using a number of speakers, wherein the sound produced by the speakers is produced when the head is facing a certain direction. The A playback transfer function that allows the listener to perceive the sound as being produced by the virtual sound source. to produce A means for generating a reproduction function, Using the said reproduction transfer function, from the said acoustic signal, To realize sound output from multiple virtual sound sources A virtual signal generation means for generating a virtual acoustic signal and It has, The above-mentioned multiple speakers use the virtual acoustic signal as input. , multiple Outputs sound related to the sound produced from a number of virtual sound sources. death, Multiple virtual sound sources are set up within the virtual sound embodiment device in accordance with the orientation of the multiple sound sources assumed in the multiple sound signals, which is the orientation within the video or image displayed on the screen. The means for generating the regeneration function is, From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or virtual sound embodiment device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. A virtual sound manifestation system characterized by the above is provided.

[0026] The present invention further provides a virtual sound generation method for generating virtual sound signals to embody sounds output from a plurality of virtual sound sources set in accordance with a plurality of input sound signals, A playback transfer function for playback using multiple installed speakers, wherein the sound produced by the speakers is perceived as sound from a virtual sound source by a listener whose head is positioned in a certain direction. to produce accomplish First Steps and Using the said reproduction transfer function, the virtual acoustic signal that is input to the speaker is generated from the acoustic signal. Second Steps and to have death, Multiple speakers are installed in a portable device for the listener, which includes a screen capable of displaying video or images related to video or image signals input in accordance with the audio signal, a camera capable of generating image information including the listener and distance information related to the distance to the listener, and an attitude sensor capable of generating screen orientation information related to the orientation of the screen. Multiple virtual sound sources are set up within the device in accordance with the orientation of the multiple sound sources assumed in the multiple sound signals, which is the orientation within the video or image displayed on the screen. In the first step, From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. This is a characteristic of A computer-based method for generating virtual sound is provided. [Effects of the Invention]

[0027] According to the virtual sound generation program and method, and virtual sound realization device of the present invention, it is possible to generate an acoustic signal of a virtual sound source while taking into account the direction of the listener. [Brief explanation of the drawing]

[0028] [Figure 1] This is a functional block diagram showing the functional configuration of one embodiment of the virtual sound realization device according to the present invention. [Figure 2] This is a schematic diagram illustrating one embodiment of the regeneration function generation process according to the present invention. [Figure 3] This is a schematic diagram illustrating another embodiment of the virtual sound generation process according to the present invention. [Figure 4] This is a flowchart illustrating one embodiment of the virtual sound generation method according to the present invention. [Modes for carrying out the invention]

[0029] Embodiments of the present invention will be described in detail below with reference to the drawings.

[0030] [Virtual Acoustic Manifestation Device] Figure 1 is a functional block diagram showing the functional configuration of one embodiment of the virtual sound realization device according to the present invention.

[0031] The smartphone 1 shown in Figure 1, as one embodiment of the virtual sound embodiment device according to the present invention, is an audio output device that, for example, when playing multimedia content received from a multimedia (MM) content distribution server 2, allows the sound (in a broad sense, including voice) output from the left loudspeaker 106-1 and the right loudspeaker 106-2 to be perceived as sound from virtual sound sources V1 and V2 by a listener (viewer in Figure 1) whose head is positioned in a certain direction.

[0032] In this embodiment, the smartphone 1 is equipped with a touch panel display 105 that is vertically oriented when the device is held vertically, which is the default orientation, and also has a left loudspeaker 106-1 and a right loudspeaker 106-2 that are arranged horizontally in the same orientation. Furthermore, in the example shown in Figure 1, the multimedia content includes audiovisual (image and sound) information adapted to the horizontally oriented display.

[0033] Therefore, the listener (viewer) holds the smartphone 1 horizontally (tilting it 90° within the screen area) to make the display landscape, and in addition, the image (in a broad sense, including video) is displayed on the screen rotated 90° in the opposite direction to the tilt of the smartphone 1 for viewing.

[0034] However, in this case, the orientation of the left and right loudspeakers 106-1 and 106-2, which were originally positioned horizontally, becomes vertical, which is completely (90 degrees) different from the "presumed orientation of multiple sound sources (within the displayed image)" in the audiovisual (image and sound) information contained in the multimedia content. Moreover, the listener's head is not necessarily facing forward but can be at any angle, and the relative positional relationship between the listener's left and right ears and the left and right loudspeakers 106-1 and 106-2 is likely to deviate more significantly from the originally assumed symmetrical position. As a result, the listener (viewer) will experience a sense of incongruity (a feeling of audiovisual incompatibility) because the position of the subject or expression perceived from the displayed image does not match the localization of the sound image perceived from the simultaneously output sound.

[0035] Therefore, the smartphone 1 realizes virtual sound sources V1 and V2 set horizontally (within the smartphone 1) in accordance with the audio signal (and image signal) of the multimedia content, and also taking into account the orientation of the listener's head (virtually displacing the left and right loudspeakers 106-1 and 106-2 in the example of Figure 1), thereby matching the position of the subject or representation with the localization of the sound image between the image and sound, and eliminating this sense of incongruity.

[0036] Specifically, smartphone 1 is designed to perform the following functions: (A) A playback function generation unit 113 generates a "playback transfer function" for playback by multiple installed loudspeakers (two left and right loudspeakers 106-1 and 106-2 in Figure 1) based on the obtained "orientation information" relating to the direction of the listener's (viewer's) head or face relative to these loudspeakers or a device including them (smartphone 1 in Figure 1), using the listener's (viewer's) "head transfer function". (B) A virtual signal generation unit 114 uses the generated "reproduction transfer function" to generate "virtual acoustic signals" that are input to these loudspeakers from multiple input acoustic signals (in the example of Figure 1, stereo signals of multimedia content to be reproduced). It is equipped with.

[0037] Here, the "reproduction transfer function" in (A) above is precisely a transfer function for virtual sound source reproduction, which ensures that the sound from multiple loudspeakers (left and right loudspeakers 106-1 and 106-2) is perceived by the listener (viewer) as sound from multiple virtual sound sources (in Figure 1, virtual sound sources V1 and V2 arranged horizontally) that are set up to match the multiple input sound signals (stereo signals of multimedia content).

[0038] Furthermore, this "reproduction transfer function" is generated based on the "orientation information" obtained in (A) above, so that the sound from the set virtual sound sources (V1, V2) is realized even for a listener (viewer) whose head is in a certain orientation. In this way, the smartphone 1 of this embodiment can generate sound signals from multiple virtual sound sources (V1 and V2) set according to the multiple sound signals (stereo signals) that are input, taking into account the orientation of the listener (viewer), and the sound output from these virtual sound sources (V1, V2) can be realized by the provided loudspeakers (106-1, 106-2).

[0039] Incidentally, the "orientation information" mentioned above, which will be explained in detail later, is generated in this embodiment by the listener information generation unit 113a of the playback function generation unit 113 based on the image information output from the RGB-D camera 104.

[0040] In another embodiment, the playback function generation unit 113 described in (A) and the virtual signal generation unit 114 described in (B) may be included in different devices. For example, the playback function generation unit 113 can be taken out of the smartphone 1 shown in Figure 1, and its function can be provided to, for example, an external server. In this case, information output from the acceleration / gyro sensor unit 103 and the RGB-D camera 104 is transmitted from the smartphone 1 to the server, and the generated "playback transfer function" is transmitted from the server to the smartphone 1. Incidentally, such a smartphone 1 and server constitute a virtual sound realization system according to the present invention.

[0041] [Device configuration, virtual sound generation program / method] The following describes in more detail the functional configuration of a smartphone 1 as one embodiment of the virtual sound realization device according to the present invention. In the functional block diagram of Figure 1, the smartphone 1 includes a communication interface unit 101, a multimedia (MM) content storage unit 102, an acceleration / gyro sensor unit 103, an RGB-D camera 104, a touch panel display 105, a left loudspeaker 106-1, a right loudspeaker 106-2, and a processor memory (a processing system with memory functionality). Here, the processor memory stores the virtual sound generation program according to the present invention and also has computer functionality, and performs virtual sound generation processing by executing this virtual sound generation program.

[0042] Herein, the virtual sound realization device according to the present invention is not limited to smartphones, but may also be a tablet computer, a wearable device, or another communication terminal capable of receiving sound signals (including audio signals of various content) equipped with the virtual sound generation program according to the present invention, or even a device dedicated to virtual sound generation processing, a personal computer (PC), a cloud server, or a non-cloud server. Incidentally, if the virtual sound realization device is a PC or a server, a separate device equipped with a loudspeaker (or display) that receives the generated virtual sound signal and outputs virtual sound may be provided.

[0043] Furthermore, the processor memory includes, as functional components, an acoustic signal adjustment unit 111 including a determination unit 111a, a video signal adjustment unit 112, a playback function generation unit 113 including a listener information generation unit 113a, a virtual signal generation unit 114 including an acoustic output adjustment unit 114a, a communication control unit 121, and a video output adjustment unit 122. These functional components can be considered as functions realized by the execution of the virtual sound generation program according to the present invention, which is stored in the processor memory. The processing flow shown by connecting the functional components of the smartphone 1 with arrows in the functional block diagram of Figure 1 can also be understood as one embodiment of the virtual sound generation method according to the present invention.

[0044] In the embodiments described below, the smartphone 1 performs virtual sound generation processing related to multimedia content. However, the smartphone 1 can also perform virtual sound generation processing related to sound signals (audio signals) other than multimedia content, that is, processing to realize sound output from multiple virtual sound sources set according to multiple input sound signals, in the same manner as described below.

[0045] <Acoustic signal adjustment means> Similarly, in the functional block diagram of Figure 1, the MM content storage unit 102 stores and manages multimedia content distributed, for example, from the MM content distribution server 2 via the communication interface unit 101 and the communication control unit 121. The acoustic signal adjustment unit 111 retrieves the multimedia content to be played from the MM content storage unit 102, performs preprocessing such as decoding on the acoustic signal, then performs a Fast Fourier Transform (FFT) to generate an FFT-processed acoustic signal capable of performing transfer function convolution, and outputs it to the virtual signal generation unit 114.

[0046] In this embodiment, the acceleration / gyro sensor unit 103 is an attitude sensor capable of generating "screen orientation information" related to the orientation of the touch panel display 105, such as information relating to the angle between the long-side direction vector and the short-side direction vector within the screen and the vertical axis. The determination unit 111a of the acoustic signal adjustment unit 111 acquires the "screen orientation information" from the acceleration / gyro sensor unit 103 during playback of the multimedia content and determines the attitude of the smartphone 1 during playback (from holding the device vertically to holding it horizontally).

[0047] Incidentally, the operating system (OS) installed in smartphones usually includes an on-device API (Application Programming Interface) for developers to provide screen orientation information and terminal orientation information. The determination unit 111a may, for example, obtain screen orientation information from this on-device API to determine the orientation of the smartphone 1, or it may obtain the orientation information of the smartphone 1 as a determination result from this on-device API.

[0048] In this embodiment, the acoustic signal adjustment unit 111 generates an FFT-processed acoustic signal only when it is determined that the terminal is being held horizontally. It is also preferable that the unit be configured to generate an FFT-processed acoustic signal when it is determined that the terminal is tilted by a predetermined angle (e.g., 45°) or more from the default vertical holding state to the horizontal holding state. Furthermore, in this embodiment, when the acoustic signal adjustment unit 111 makes such a determination, it outputs a "reproduction transfer function generation instruction" to the reproduction function generation unit 113, which will be described later, and the reproduction function generation unit 113 performs the reproduction transfer function generation process upon receiving this instruction. Incidentally, the reproduction function generation unit 113 may also be configured to perform the reproduction transfer function generation process when it receives the same instruction from the user via the touch panel display 105.

[0049] <Video signal adjustment means> Similarly, in the functional block diagram of Figure 1, the video signal adjustment unit 112 performs pre-processing such as decoding on the image signal of the multimedia content to be played back, and outputs the pre-processed image signal to the video output adjustment unit 122 to display the image to be played back on the touch panel display 105. Incidentally, in this embodiment, if the smartphone 1 is tilted by a predetermined angle (for example, 45°) or more, the video output adjustment unit 122 displays the image to be played back on the touch panel display 105 rotated 90° in the opposite direction to the tilt of the smartphone 1.

[0050] <Reproduction function generation means: Listener information generation process> Similarly, in the functional block diagram of Figure 1, the listener information generation unit 113a of the playback function generation unit 113 receives a "playback transfer function generation instruction" from the acoustic signal adjustment unit 111 (or from the user via the touch panel display 105) in this embodiment. (a) "Orientation information" relating to the direction of the viewer's (listener's) head or face relative to the smartphone 1 (including left and right loudspeakers 106-1 and 106-2), (b) "Distance information" relating to the distance from smartphone 1 (including left and right loudspeakers 106-1 and 106-2) to the viewer (listener), and (c) "Screen orientation information" (In this embodiment, this is information relating to the angle that the long-side direction vector and the short-side direction vector in the screen make with the vertical axis) Based on this, "speaker orientation information" relating to the orientation of the left and right loudspeakers 106-1 and 106-2 as seen from the left and right ears of the viewer (listener) is determined.

[0051] Specifically, the "internal speaker position information" relating to the respective positions of the left and right loudspeakers 106-1 and 106-2 within the smartphone 1 is pre-set and saved information that can be obtained. First, the position coordinates of the viewer's head within the coordinate space fixed to the smartphone 1 can be determined from this "internal speaker position information," "screen orientation information," and "distance information." Next, by adopting predetermined values ​​(for example, the average value of measurement results for many people) as the positions of the left and right ears within the viewer's head, the position coordinates of the viewer's left and right ears within this coordinate space can be determined from the "orientation information." Finally, this makes it possible to determine the "speaker orientation information" for each of the left and right loudspeakers 106-1 and 106-2 as seen from each ear.

[0052] Here, the "orientation information" in (a) above can be generated in this embodiment by applying known face recognition technology to image information including the viewer generated by the RGB-D camera 104. In addition, the OS installed in a smartphone usually has a function to provide this "orientation information" as an on-device API for developers, and the "orientation information" may be obtained from this on-device API. Furthermore, the "distance information" in (b) above can be the distance value to the viewer generated by the RGB-D camera 104.

[0053] Furthermore, if the distance from smartphone 1 to the viewer is set to a predetermined value (for example, empirically determined), and the screen orientation is set to a typical orientation for when the device is held horizontally, it is possible to determine the "speaker orientation information" based solely on the "orientation information" in (a) above. However, in this embodiment, by using the acquired "distance information" and furthermore the "screen orientation information," it is possible to generate a more accurate playback transfer function afterward, and to perform a better virtual sound generation process. <Regeneration function generation means: Regeneration transfer function generation process>

[0054] Similarly, in the functional block diagram of Figure 1, the playback function generation unit 113 uses the "speaker orientation information" determined as described above to generate a "playback transfer function" for playback by the left and right loudspeakers 106-1 and 106-2, which ensures that the sound from these speakers is perceived by a viewer with their head facing a certain direction as sound from the set virtual sound sources V1 and V2.

[0055] Specifically, in this embodiment, the regeneration function generation unit 113 is, (a) A "first head-related transfer function" set for each installed speaker (left and right loudspeakers 106-1 and 106-2 in Figure 1), which is a function of the orientation of the speaker as seen from the viewer's head or ears (ears in this embodiment), and (b) A "second head-related transfer function" set for each configured virtual sound source (virtual sound sources V1 and V2 in Figure 1), which is a function of the orientation of the virtual sound source as seen from the viewer's head or ears (ears in this embodiment). This is determined based on the "speaker orientation information" determined above and the "virtual sound source orientation information" which will be explained later. Furthermore, the "reproduction transfer function" is generated using the determined "first head-direction transfer function" and "second head-direction transfer function".

[0056] Here, "virtual sound source orientation information" is orientation information relating to the orientation of virtual sound sources V1 and V2 as seen from the left and right ears of the viewer (listener). In this embodiment, the listener information generation unit 113a generates the virtual sound source information in the same manner as the "speaker orientation information" described above. (a) "Orientation information" relating to the direction of the viewer's (listener's) head or face relative to the smartphone 1 (including virtual sound sources V1 and V2 in this embodiment), (b) "Distance information" relating to the distance from the smartphone 1 to the viewer (listener), (including virtual sound sources V1 and V2 in this embodiment), and (c) "Screen orientation information" (In this embodiment, this is information relating to the angle that the long-side direction vector and the short-side direction vector in the screen make with the vertical axis) Based on this, the "virtual sound source orientation information" is determined. Incidentally, in this embodiment, the in-terminal location information relating to the respective positions of the virtual sound sources V1 and V2 set in the smartphone 1 is precisely the setting item and is naturally acquired, so the "virtual sound source orientation information" can be determined in the same way as the "speaker orientation information" described above.

[0057] In this embodiment, the playback function generation unit 113 is configured to generate a playback transfer function only when the "screen orientation information" obtained from the acceleration / gyro sensor unit 103 (or on-device API) satisfies predetermined conditions, for example, when it is determined from the "screen orientation information" that "the smartphone 1 is tilted by a predetermined angle (e.g., 45°) or more from the default vertical orientation to the horizontal orientation." This allows for the generation of a playback transfer function to resolve discrepancies between the position of subjects or objects captured in the image and the localization of sound images captured by the sound, only when necessary, thus avoiding the creation of unnecessary virtual sound sources.

[0058] Figure 2 is a schematic diagram illustrating one embodiment of the playback function generation process according to the present invention. Incidentally, in the example shown in Figure 2, the smartphone 1 is held horizontally, and the virtual sound sources V1 and V2 are set to the positions within the smartphone 1 as shown in Figure 2(B) (positions arranged horizontally in the horizontal orientation when the device is held horizontally).

[0059] According to Figure 2(A), in this embodiment, the listener information generation unit 113a is, (a) Orientation information (θ) derived from the image information generated by the RGB-D camera 104 (or obtained from the on-device API) HEAD , φ HEAD )and, (b) Distance information d generated by the RGB-D camera 104 HEAD and, (c) Screen orientation information generated by the acceleration / gyro sensor unit 103 (or obtained from the on-device API) and Perform the calculation as described above using (R1L) Speaker orientation information (θ R1L , φ R1L ) of the left loudspeaker 106-1 as seen from the left ear of the viewer, (R1R) Speaker orientation information (θ R1R , φ R1R ) of the left loudspeaker 106-1 as seen from the right ear of the viewer, (R2L) Speaker orientation information (θ R2L , φ R2L ) of the right loudspeaker 106-2 as seen from the left ear of the viewer, (R2R) Speaker orientation information (θ R2R , φ R2R ) of the right loudspeaker 106-2 as seen from the right ear of the viewer, (V1L) Virtual sound source orientation information (θ V1L , φ V1L ) of the virtual sound source V1 as seen from the left ear of the viewer, (V1R) Virtual sound source orientation information (θ V1R , φ V1R ) of the virtual sound source V1 as seen from the right ear of the viewer, (V2L) Virtual sound source orientation information (θ V2L , φ V2L ) of the virtual sound source V2 as seen from the left ear of the viewer, and (V2R) Virtual sound source orientation information (θ V2R , φ V2R ) of the virtual sound source V2 as seen from the right ear of the viewer are determined.

[0060] Here, θ HEAD and φ HEAD of the orientation information are the θ coordinate value and the φ coordinate value in the terminal spherical coordinate system (with the position of the RGB-D camera 104 fixed to the smartphone 1 as the origin). Also, θ R** and φ R** of the speaker orientation information, as well as θ V** and φ V** of the virtual sound source orientation information are the θ coordinate value and the φ coordinate value in the listener spherical coordinate system (Figure 2) (with the position of the corresponding ear of the viewer as the origin).

[0061] Next, as shown in Figure 2(B), in this embodiment, the playback function generation unit 113 uses the head-related transfer function H(f, θ, φ), which has been acquired and stored in advance (in a form determined by specifying the coordinate values ​​(θ, φ) in the listener spherical coordinate system), to determine the above four speaker orientation information and four virtual sound source orientation information, (R1L) The first head-related transfer function HR is a function of the orientation of the left loudspeaker 106-1 as seen from the viewer's left ear. 1L (f, θ R1L , φ R1L ), (R1R) The first head-related transfer function HR is a function of the orientation of the left loudspeaker 106-1 as seen from the viewer's right ear. 1R (f, θ R1R , φ R1R ), (R2L) The first head-related transfer function HR is a function of the orientation of the right loudspeaker 106-2 as seen from the viewer's left ear. 2L (f, θ R2L , φ R2L ), (R2R) The first head-related transfer function HR is a function of the orientation of the right loudspeaker 106-2 as seen from the viewer's right ear. 2R (f, θ R2R , φ R2R ), (V1L) The second head-related transfer function HV is a function of the orientation of the virtual sound source V1 as seen from the viewer's left ear. 1L (f, θ V1L , φ V1L ), (V1R) The second head-related transfer function HV is a function of the orientation of the virtual sound source V1 as seen from the viewer's right ear. 1R (f, θ V1R , φ V1R ), (V2L) The second head-related transfer function (HV) is a function of the orientation of the virtual sound source V2 as seen from the viewer's left ear. 2L (f, θ V2L , φ V2L ), and (V2R) The second head-related transfer function (HV) is a function of the orientation of the virtual sound source V2 as seen from the viewer's right ear. 2R (f, θ V2R , φ V2R ) Determine the frequency (index). Here, f is the frequency.

[0062] Thus, the playback function generation unit 113 of this embodiment takes into account the viewer's orientation and the viewer's orientation information (θ HEAD , φ HEAD It is possible to determine the optimal first head-level transfer function HR and second head-level transfer function HV, which also incorporate the ) data.

[0063] Returning to the functional block diagram in Figure 1, the playback function generation unit 113 of this embodiment uses the four first head-related transfer functions HR (R1L) to (R2R) and the four second head-related transfer functions HV (V1L) to (V2R) to generate a playback transfer function (W) for playback by the left and right loudspeakers 106-1 and 106-2. 1L , W 1R , W 2L , W 2R ) is generated. Below, the FFT-processed left channel acoustic signal and right channel acoustic signal output from the acoustic signal adjustment unit 111 are respectively S L (f, t) and S R Let (f, t) be the time (index).

[0064] In this embodiment, the "regenerative transfer function (W 1L , W 1R , W 2L , W 2R ) These convoluted left and right channel acoustic signals (S L , S R The sound output from the left and right loudspeakers 106-1 and 106-2, which have received the sound signal (S) from these left and right channels, is the sound signal (S) from these left and right channels. L , S R Using the "condition" that "the sound output from virtual sound sources V1 and V2 that received ) matches", the playback transfer function (W 1L , W 1R , W 2L , W 2R ) generates.

[0065] Specifically, this "condition" is expressed as follows:

number

[0066] Here the regeneration transfer function (W 1L , W 1R , W 2L , W 2R To find ), from equation (1) above, the following equation

number

number

[0067] Thus, the playback function generation unit 113 of this embodiment uses the above equation (3) which expresses the above "conditions" to determine the first head-related transfer function HR and the second head-related transfer function HV, and determines that "the sound from the left and right loudspeakers 106-1 and 106-2 is directed in the direction of the head (θ HEAD , φ HEAD A playback transfer function (W) that makes it possible to "make the sound appear to the viewer as sound produced by virtual sound sources V1 and V2". 1L , W 1R , W 2L , W 2R It is possible to generate ).

[0068] <Virtual signal generation means> Also in the functional block diagram of FIG. 1, the virtual signal generation unit 114 of the present embodiment receives the reproduction transfer function (W 1L , W 1R , W 2L , W 2R ) from the reproduction function generation unit 113, and from the acoustic signals (S L , S R ) of the left and right channels after FFT processing received from the audio signal adjustment unit 111, generates virtual acoustic signals (S' L , S' R ) to be input to the left and right loudspeakers 106-1 and 106-2.

[0069] Specifically, the virtual acoustic signals (S' L , S' R ) use a circuit (FIG. 1) that implements the "product" of the "matrix" related to the second W and the "matrix" related to the third S on the left side of the above equation (1), and the following equations (4) S' L = FFT -1 [W 1L * S L + W 1R * S R (5) S' R = FFT -1 [W 2L * S L + W 2R * S R to generate. Here, FFT -1 [] is a function representing the inverse FFT (inverse fast Fourier transform) process executed by the acoustic output adjustment unit 114a.

[0070] The virtual acoustic signals (S' L , S' R ​​By inputting these signals to the left and right loudspeakers 106-1 and 106-2 respectively and outputting sound, it is possible to realize the sound output from virtual sound sources V1 and V2 (which are set to be aligned horizontally within the smartphone 1 when the device is held horizontally in Figure 1). In other words, in this embodiment, the left and right loudspeakers 106-1 and 106-2 can be virtually displaced to desired positions (positions aligned horizontally).

[0071] Incidentally, in this embodiment, the left and right channels of the FFT processed acoustic signals (S L , S R The virtual signal generation unit 114 of this embodiment generates the resulting signal by sequentially shifting each of these sequentially shifted acoustic signals (S) by a predetermined amount chg_len, taking an acoustic signal portion of a predetermined width window from the original acoustic signal, and then performing an FFT process on each of the signals obtained in this way. L , S R The above equations (4) and (5) are used to perform the processing (including inverse FFT processing) on ​​each virtual acoustic signal (S') obtained as a result. L , S' R The signals are sequentially shifted by the amount of chg_len and added together (overlap addition), and this added virtual acoustic signal is output to the left and right loudspeakers 106-1 and 106-2.

[0072] <Other embodiments of virtual sound generation processing> Figure 3 is a schematic diagram showing another embodiment of the virtual sound generation process according to the present invention.

[0073] In the smartphone 1 of the embodiment shown in Figure 3(A), two virtual sound sources V1 and V2 are set, just like in the embodiment shown in Figure 1. However, unlike in Figure 1, in addition to the two left and right loudspeakers 106-1 and 106-2, an upper loudspeaker 106-3 is provided. Even when three or more N (≧3) loudspeakers are provided in this way, the playback function generation unit 113 (Figure 1) generates a playback transfer function, and the virtual signal generation unit 114 (Figure 1) can generate a virtual acoustic signal that embodies the virtual sound sources V1 and V2 using the N loudspeakers.

[0074] Specifically, the first head transfer function (HR) of N left ears for N loudspeakers. 1L , ··, HR NL ) and the first head transfer function (HR) of the N right ears 1R , ··, HR NR After determining the following equation,

number

number

[0075] Next, in the smartphone 1 of the embodiment shown in Figure 3(B), similar to Figure 3(A), left and right loudspeakers 106-1 and 106-2, and an upper loudspeaker 106-3 are provided. Furthermore, in addition to virtual sound sources V1 and V2, virtual sound sources V1S and V2S are also set. In this case, the sound signals handled by the sound signal adjustment unit 111 (Figure 1) are four sound signals for realizing surround sound, and the four virtual sound sources (V1, V2, V1S, V2S) are set to positions that are around the viewer.

[0076] Incidentally, all of these virtual sound sources (V1, V2, V1S, V2S) can be set outside of the smartphone 1. It is also possible to set three or more virtual sound sources. In any of these cases, the playback function generation unit 113 (Figure 1) generates a playback transfer function appropriate to the case, and the virtual signal generation unit 114 (Figure 1) can generate a virtual acoustic signal that embodies multiple virtual sound sources using N loudspeakers.

[0077] Specifically, for example, four virtual sound sources (V1, V2, V1S, V2S) are set up, and the first head transfer function (HR) of the N left ears for the N loudspeakers is calculated. 1L , ··, HR NL ) and the first head transfer function (HR) of the N right ears 1R , ··, HR NR After determining the following equation,

number

number

[0078] As explained above, by increasing the number of loudspeakers that serve as input destinations for the virtual acoustic signal, or by increasing the number of virtual sound sources that are set, it becomes possible to virtually realize a wide variety of sounds and to achieve more accurate (closer to the original design) sounds.

[0079] <Virtual sound generation method> Figure 4 is a flowchart schematically showing one embodiment of the virtual sound generation method according to the present invention. The outline of each step in the flowchart is described below. Note that the virtual sound generation method in this embodiment uses the smartphone 1 shown in Figure 1.

[0080] (S1) Read the installation location information of the display (105) and loudspeakers (106-1, 106-2). (S2) Starts acquiring and playing back multimedia data. Specifically in this embodiment, a predetermined window of audio signal portion is sequentially extracted from the audio signal of the multimedia data, and each extracted audio signal portion is acquired sequentially to perform the processing in steps S3 to S8.

[0081] (S3) The screen orientation information of the display (105) is retrieved from the on-device API. (S4) Using the captured screen orientation information, it is determined whether or not to set a virtual sound source, that is, whether the smartphone 1 is held horizontally or vertically. If it is determined that a virtual sound source should not be set (the device is held vertically), the process proceeds to step S8, which will be described later. On the other hand, if it is determined that a virtual sound source should be set (the device is held horizontally), the virtual sound sources (V1, V2) are set, and the process proceeds to the next step, S5.

[0082] (S5) Capture the viewer's head orientation and distance information. (S6) The first head transfer function HR and the second head transfer function HV are determined from the acquired orientation information, distance information, and screen orientation information. (S7) Using the determined first head transfer function HR and second head transfer function HV, a regenerated transfer function W is generated. (S8) A virtual acoustic signal is generated using the generated playback transfer function W, and the captured acoustic signal is played back (sound is output from the loudspeakers (106-1, 106-2)). If it is determined in step S4 that a virtual sound source is not set (the terminal is held vertically), the acoustic signal portion extracted from the multimedia data is used for playback. Incidentally, the corresponding video signal portion in the multimedia data is also played back at this point along with the playback of the acoustic signal portion.

[0083] (S9) The processes in steps S3 to S8 are repeated until the acquisition of multimedia data is complete. When the acquisition and playback of multimedia data are complete, this virtual sound generation method is terminated.

[0084] As described in detail above, according to the present invention, it is possible to generate audio signals from multiple virtual sound sources that are set to match multiple input audio signals, taking into account the orientation of the listener, and to realize the sound output from these virtual sound sources by a provided loudspeaker. Furthermore, this makes it possible to avoid situations such as when playing multimedia content, where a listener with their head facing a certain direction experiences a sense of incongruity (a feeling of audiovisual inconsistency) because the position of the subject or expression captured from the image (video) does not match the localization of the sound image captured from the simultaneously output sound.

[0085] Furthermore, for example, children and students not only in urban areas but also in rural areas can enjoy the playback of multimedia teaching materials with less discomfort using the virtual sound embodiment device according to the present invention, such as a tablet device or smartphone. This also makes it possible to realize higher quality online classes. In other words, according to the present invention, it is possible to contribute to Goal 4 of the United Nations-led Sustainable Development Goals (SDGs), "Ensure inclusive and equitable quality education and promote lifelong learning opportunities for all."

[0086] Furthermore, adults not only in urban areas but also in rural areas can enjoy the playback of multimedia digital materials with less discomfort using the virtual sound embodiment device according to the present invention, such as a tablet or smartphone. This also makes it possible to make online work meetings more efficient and of higher quality. In other words, according to the present invention, it is possible to contribute to Goal 8 of the United Nations-led SDGs, "Promote inclusive and sustainable economic growth, employment and decent work for all."

[0087] Various changes, modifications, and omissions to the scope of the technical concept and viewpoint of the present invention can be easily made by those skilled in the art with respect to the various embodiments of the present invention described above. The foregoing description is merely illustrative and is not intended to limit the present invention in any way. The present invention is limited only to what is limited by the claims and their equivalents. [Explanation of Symbols]

[0088] 1. Smartphone (Virtual Sound Manifestation Device) 101 Communication Interface Section 102 Multimedia (MM) Content Storage Section 103 Acceleration / Gyroscope Sensor Unit 104 RGB-D Camera 105 Touch Panel Display 106-1 Left loudspeaker 106-2 Right loudspeaker 111 Sound signal adjustment section 111a Judgment part 112 Video signal adjustment unit 113 Regeneration Function Generation Unit 113a Listener information generation section 114 Virtual signal generation unit 114a Sound output adjustment section 121 Communication Control Unit 122 Video Output Adjustment Section

Claims

1. A virtual sound generation program that generates virtual sound signals to realize sounds output from multiple virtual sound sources configured to match multiple input sound signals, A playback function generation means for generating a playback transfer function for playback by multiple installed speakers, wherein the sound from the speakers is perceived as sound from a virtual sound source by a listener whose head is facing a certain direction. A virtual signal generation means that generates the virtual acoustic signal to be input to the speaker from the acoustic signal using the said reproduction transfer function. to make the computer work, The aforementioned plurality of speakers are installed in a portable device for the listener, which includes a screen capable of displaying video or images related to video or image signals input in accordance with the acoustic signal, a camera capable of generating image information including the listener and distance information related to the distance to the listener, and an attitude sensor capable of generating screen orientation information related to the orientation of the screen. The plurality of virtual sound sources are set up in the device in a direction within the video or image displayed on the screen, along with the arrangement direction of the plurality of sound sources assumed in the plurality of sound signals. The aforementioned reproduction function generation means is From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. A virtual sound generation program characterized by the following features.

2. The aforementioned multiple speakers are installed along the lateral direction within the device in its default orientation. If the device is tilted 90 degrees from its default position within the plane including the screen, the video or image will be displayed within the screen rotated 90 degrees in the opposite direction to the tilt of the device. The aforementioned plurality of virtual sound sources are set within the device in a direction within the video or image that has been rotated 90 degrees, and along the arrangement direction of the assumed plurality of sound sources. The virtual sound generation program according to feature 1.

3. The virtual sound generation program according to Claim 1, characterized in that the playback function generation means generates the playback transfer function only when the screen orientation information satisfies predetermined conditions.

4. A portable virtual sound embodiment device for the listener that embodies sound output from multiple virtual sound sources configured to match multiple input acoustic signals, Multiple speakers, A screen capable of displaying video or images related to video or image signals input in accordance with the said audio signal, A camera capable of generating image information including the listener and distance information relating to the distance to the listener, A posture sensor capable of generating screen orientation information relating to the orientation of the aforementioned screen, A playback function generation means for generating a playback transfer function for playback by the plurality of speakers such that the sound from the speakers is perceived as sound from the virtual sound source by the listener whose head is facing a certain direction, A virtual signal generation means that uses the reproduction transfer function to generate a virtual acoustic signal from the acoustic signal to embody the sound output from the plurality of virtual sound sources. It has, The plurality of speakers take the virtual sound signal as input and output sound related to the sound output from the plurality of virtual sound sources. The plurality of virtual sound sources are set within the virtual sound realization device in a direction within the video or image displayed on the screen, along with the arrangement direction of the plurality of sound sources assumed in the plurality of sound signals. The aforementioned reproduction function generation means is From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or virtual sound embodiment device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. A virtual sound manifestation device characterized by the following features.

5. A virtual sound embodiment system comprising a virtual sound embodiment device portable to the listener, which realizes sound output from multiple virtual sound sources configured to match multiple input acoustic signals, Multiple speakers provided in the virtual sound realization device, A screen provided in the virtual sound manifestation device that is capable of displaying video or images related to video or image signals input in accordance with the sound signal, A camera provided in the virtual sound manifestation device is capable of generating image information including the listener and distance information relating to the distance to the listener, A posture sensor provided in the virtual sound embodiment device, which is capable of generating screen orientation information relating to the orientation of the screen, A playback function generation means for generating a playback transfer function for playback by the plurality of speakers such that the sound from the speakers is perceived as sound from the virtual sound source by the listener whose head is facing a certain direction, A virtual signal generation means that uses the reproduction transfer function to generate a virtual acoustic signal from the acoustic signal to embody the sound output from the plurality of virtual sound sources. It has, The plurality of speakers take the virtual sound signal as input and output sound related to the sound output from the plurality of virtual sound sources. The plurality of virtual sound sources are set within the virtual sound realization device in a direction within the video or image displayed on the screen, along with the arrangement direction of the plurality of sound sources assumed in the plurality of sound signals. The aforementioned reproduction function generation means is From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or virtual sound embodiment device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. A virtual sound manifestation system characterized by the following features.

6. A virtual sound generation method that generates virtual sound signals to realize sounds output from multiple virtual sound sources configured to match multiple input sound signals, A first step of generating a playback transfer function for playback by multiple installed speakers, wherein the sound from the speakers is perceived as sound from a virtual sound source by a listener whose head is facing a certain direction. A second step involves using the reproduction transfer function to generate a virtual acoustic signal that is input to the speaker from the acoustic signal. It has, The aforementioned plurality of speakers are installed in a portable device for the listener, which includes a screen capable of displaying video or images related to video or image signals input in accordance with the acoustic signal, a camera capable of generating image information including the listener and distance information related to the distance to the listener, and an attitude sensor capable of generating screen orientation information related to the orientation of the screen. The plurality of virtual sound sources are set up in the device in a direction within the video or image displayed on the screen, along with the arrangement direction of the plurality of sound sources assumed in the plurality of sound signals. In the first step, From the generated distance information, the orientation information relating to the direction of the listener's head or face relative to the speaker or device determined from the image information, and the screen orientation information, speaker orientation information relating to the coordinate values ​​of the direction of the speaker as seen from the listener's head or ears in a spherical coordinate system fixed to the listener's head, and virtual sound source orientation information relating to the coordinate values ​​of the direction of the virtual sound source as seen from the listener's head or ears in the spherical coordinate system are determined. Using the first head-direction transfer function of the listener, which is a function set for each speaker and includes the speaker orientation information as a variable, and the second head-direction transfer function of the listener, which is a function set for each virtual sound source and includes the virtual sound source orientation information as a variable, the playback transfer function is generated from the determined speaker orientation information and virtual sound source orientation information. A computer-based method for generating virtual sound, characterized by the following features.

Citation Information

Patent Citations

  • Portable terminal

    JP2003078601A

  • Mobile terminal, stereo reproducing method, and stereo reproducing program

    JP2006174277A

  • Media reproducing apparatus

    JP2009232424A

  • Multimedia device and program

    JP2015070324A

  • Using face detection to update user interface orientation

    US20200104033A1