Signal transmission method, signal generation method, signal reproduction method, audio signal processing program, audio transmission device, audio reproduction device, and audio transmission and reproduction system
By separating and processing binaural and non-binaural audio signals in the audio transmission and reproduction device, the problem of unnatural sound and image caused by transmission delay in head tracking devices is solved, and more natural audio reproduction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2024-09-20
- Publication Date
- 2026-04-24
AI Technical Summary
In related technologies, when using head tracking devices to update HRTF, there is a problem of transmission delay leading to unnatural audio-visual presentation.
A signal transmission method is used to separate and transmit a first type of audio signal that includes binauralization and a second type of audio signal that does not include binauralization. The binauralization process is then performed by an audio transmission device and a reproduction device to generate a binaural signal.
By sharing the binaural processing, transmission delay is reduced, unnatural sound images are prevented, and the naturalness of audio signal reproduction is improved.
Smart Images

Figure CN121925869A_ABST
Abstract
Description
Technical Field
[0001] This invention relates in particular to a signal transmission method, a signal generation method, a signal reproduction method, an audio signal processing program, an audio transmission device, an audio reproduction device, and an audio transmission and reproduction system, for generating a three-dimensional audio signal to be reproduced via headphones or the like. Background Technology
[0002] Among related technologies, there are VR headsets, HMDs (head-mounted displays), etc. (hereinafter referred to as "HMDs, etc.") that are capable of reproducing content such as movies, VR (virtual reality), and AR (augmented reality).
[0003] In HMDs and similar devices, binaural audio has been reproduced to provide a wider sound field. This is a three-dimensional audio format used to provide external head sound localization by employing a head-related transfer function (HRTF) that considers the direction from the listener to the sound source. Furthermore, by measuring the listener's head orientation relative to the sound source using accelerometers, position sensors, etc. (hereinafter referred to as "head tracking"), a more accurate two-channel (2ch) three-dimensional audio format, i.e., binaural audio (hereinafter referred to as "binauralization"), is generated, resulting in a more realistic sound image.
[0004] In addition, conventionally, in the use of HRTF in binauralization, the head-related impulse response (hereinafter referred to as "HRIR"), which is the time-domain representation of the head-related transfer function, is usually used to perform arithmetic operations on the actual audio signal (hereinafter, HRTF or HRIR will be referred to as "HRTF etc").
[0005] Patent Document 1 describes a technique for directional audio capture, acoustic preprocessing, encoding, decoding, and binaural (two-ear) rendering in an audio scene. The technique of Patent Document 1 includes an apparatus adapted to modify the directional characteristics of captured directional audio in response to spatial data from a microphone system used for capturing directional audio. The apparatus is configured to modify the directional characteristics of received directional audio based on head tracking in response to received spatial data.
[0006] Existing technical documents Patent documents Patent Document 1: Japanese Patent Application Publication No. 2021-5822 Summary of the Invention
[0007] Technical issues In head-tracking devices in related technologies (such as the device described in Patent Document 1), the orientation information of the listener's head is transmitted, and the device receiving this information updates the HRTF, etc. Therefore, there is a problem that a delay equivalent to the round-trip communication time occurs during the update, resulting in an unnatural sound image.
[0008] The present invention was made in view of such circumstances, and its purpose is to solve the above-mentioned problems.
[0009] Solution to the problem According to the present invention, a signal transmission method is provided for transmitting an audio signal comprising multiple channels, the signal transmission method comprising: real-time transmission of a signal comprising a first type of audio signal that is binauralized and a second type of audio signal that is not binauralized.
[0010] In the signal transmission method according to the present invention, the first type of binauralized audio signal is the audio signal of some of these channels, and the second type of non-binauralized audio signal is the audio signal of other channels.
[0011] In the signal transmission method according to the present invention, the second type of audio signal is a signal in which auditory unnaturalness caused by positional displacement or movement is more easily perceived compared to the first type of audio signal.
[0012] In the signal transmission method according to the present invention, whether auditory unnaturalness in the audio signals of multiple channels is easily perceptible is determined based on azimuth and elevation angles.
[0013] In the signal transmission method according to the present invention, the first type of audio signal is a low-frequency component, and the second type of audio signal is a high-frequency component.
[0014] In the signal transmission method according to the present invention, the first type of audio signal is a representative point signal, which includes multiple channels of audio signal and is grouped into directions with a number less than the number of channels of the signal.
[0015] According to the present invention, a signal generation method is provided for generating an audio signal containing multiple channels, the signal generation method comprising: binauralizing a first type of audio signal only based on directional information received by a listener; and generating a signal containing the binauralized first type of audio signal and a non-binauralized second type of audio signal.
[0016] According to the present invention, a signal reproduction method for reproducing an audio signal containing multiple channels is provided. The signal reproduction method includes: acquiring a signal transmitted by a signal transmission method or a signal generated by a signal generation method; binauralizing the received second type of audio signal; and combining the binauralized second type of audio signal with a first type of audio signal that was binauralized at the time of reception to generate a binaural signal.
[0017] According to the present invention, a signal reproduction method performed by an audio generation and reproduction system is provided, the audio generation and reproduction system comprising: an audio transmission device that generates and transmits an audio signal containing multiple channels; and an audio reproduction device that reproduces the signal transmitted by the audio transmission device, the signal reproduction method comprising: grouping the audio signal containing multiple channels generated by the audio transmission device into representative point signals in a direction fewer in number than the number of channels of the signal, and transmitting the representative point signals in real time; and binauralizing the representative point signals by the audio reproduction device to generate a binaural signal.
[0018] According to the present invention, an audio signal processing program executed by an audio transmission device is provided, the audio signal processing program causing the audio transmission device to perform the following: binauralizing only a first type of audio signal based on directional information received by a listener; and generating a signal comprising the binauralized first type of audio signal and a non-binauralized second type of audio signal.
[0019] According to the present invention, an audio signal processing program executed by an audio reproduction device is provided, the audio signal processing program causing the audio reproduction device to perform the following: acquiring a signal transmitted by a signal transmission method or a signal generated by a signal generation method; binauralizing the received second type of audio signal; and combining the binauralized second type of audio signal with a first type of audio signal that was binauralized when received to generate a binaural signal.
[0020] According to the present invention, an audio transmission apparatus is provided for generating an audio signal comprising multiple channels, the audio transmission apparatus comprising: a first binauralization unit that binauralizes only a first type of audio signal based on directional information of a listener received from an audio reproduction device; and a signal generation unit that generates a signal comprising the first type of audio signal binauralized by the first binauralization unit and a second type of audio signal not binauralized.
[0021] According to the present invention, an audio reproduction apparatus is provided, the audio reproduction apparatus reproducing a signal containing multiple channels of audio signals, the audio reproduction apparatus comprising: a signal acquisition unit that acquires a signal containing a first type of audio signal that has been binauralized and a second type of audio signal that has not been binauralized; a second binauralization unit that binauralizes the second type of audio signal acquired by the signal acquisition unit; a binaural signal generation unit that combines the second type of audio signal binauralized by the second binauralization unit and the first type of audio signal to generate a binaural signal; and an audio output unit that outputs the binaural signal generated by the binaural signal generation unit.
[0022] According to the present invention, an audio transmission and reproduction system is provided as an audio generation and reproduction system, the audio transmission and reproduction system comprising: an audio transmission device for generating a signal comprising an audio signal including multiple channels; and an audio reproduction device for reproducing the signal generated and transmitted by the audio transmission device, wherein the audio transmission device comprises: a first binauralization unit for binauralizing only a first type of audio signal based on directional information of a listener received from the audio reproduction device; a signal generation unit for generating a signal comprising the first type of audio signal binauralized by the first binauralization unit and a second type of audio signal not binauralized; and a transmission unit for transmitting the signal generated by the signal generation unit in real time, and the audio reproduction device comprises: a signal acquisition unit for acquiring the signal generated by the audio transmission device; a second binauralization unit for binauralizing the second type of audio signal acquired by the signal acquisition unit; a binaural signal generation unit for combining the second type of audio signal binauralized by the second binauralization unit and the first type of audio signal to generate a binaural signal; an audio output unit for outputting the binaural signal generated by the binaural signal generation unit; and a directional transmission unit for transmitting the directional information of the listener to the audio transmission device.
[0023] Invention Effects According to the present invention, a signal transmission method can be provided in which, by transmitting in real time a transmission signal comprising a first type of audio signal that is binauralized and a second type of audio signal that is not binauralized, binauralization can be shared between the transmission side and the receiving side, and unnatural sound images can be prevented. Attached Figure Description
[0024] Figure 1 A block diagram illustrating the configuration of an audio transmission and reproduction system according to an embodiment of the present invention.
[0025] Figure 2 A conceptual diagram illustrating audio transmission and reproduction processing according to an embodiment of the present invention.
[0026] Figure 3 A flowchart illustrating audio transmission and reproduction processing according to an embodiment of the present invention is provided. Detailed Implementation
[0027] <Implementation Plan> [Control Configuration of Audio Transmission and Reproduction System X] First, refer to Figure 1 The control configuration of an audio transmission and reproduction system X according to an embodiment of the present invention is described.
[0028] The audio transmission and reproduction system X includes an audio transmission device 1 and an audio reproduction device 2.
[0029] The audio transmission device 1 is an apparatus for encoding and transmitting audio signals (hereinafter referred to as "source audio signals") from multiple channels. Furthermore, the audio transmission device 1 may be able to generate the source audio signals itself, for example, by using or synthesizing a basic audio signal.
[0030] In this embodiment, the audio transmission device 1 is, for example: a host terminal, such as a smartphone, mobile phone, or HMD; a portable gaming PC (personal computer); a consumer or commercial video game console; a server; a television; a television (video) conferencing system; a teleconferencing equipment; a hearing aid; a hearing device; other household appliances; equipment for cinemas or public viewing venues; a dedicated audio encoder; a switcher; a PA; or other broadcasting or audio reproduction equipment.
[0031] The audio reproduction device 2 is a device for reproducing the transmission signal T generated and transmitted by the audio transmission device.
[0032] In this embodiment, the audio playback device 2 is, for example: an HMD (head-mounted display) for XR (extended reality) such as VR (virtual reality), AR (augmented reality), MR (mixed reality), or SR (alternate reality) that is capable of head tracking; an ear-mounted speaker, ear-mounted headphones, or headphones equipped with a head tracking sensor (hereinafter referred to as "headphones, etc."); an ear-hook smart device; a three-dimensional audio playback device connected to headphones, etc.; a dedicated game console; a content playback device; or other audio playback devices.
[0033] In this embodiment, the audio transmission device 1 and the audio playback device 2 are connected to each other wirelessly or via a wire. For wireless connectivity, the following can be used: short-range wireless communication, including Bluetooth (registered trademark) and wireless local area networks (LAN, Wi-Fi); Li-Fi and infrared communication; cellular networks, such as 4G and 5G; mid-range wireless communication; satellite radio communication networks; and other wireless communication technologies. For wired connectivity, the connection can be established via various USB (Universal Serial Bus), wired LAN, wired WAN (Wide Area Network), HDMI (registered trademark), DP (Display Port), or other dedicated lines. Furthermore, communication between the audio transmission device 1 and the audio playback device 2 can be established via Bluetooth (registered trademark), TCP / IP, UDP, other IP networks, or other communication protocols.
[0034] Furthermore, whether to use a wireless or wired connection to connect the audio transmission device 1 and the audio playback device 2 can be selected based on factors such as the distance between the devices and convenience. Additionally, the audio transmission device 1 and the audio playback device 2 can be connected via various methods.
[0035] In this embodiment, binaural audio is generated from audio signals including channels C-1 to Cn (serving as source audio signals). Any one of the multiple channels C-1 to Cn may also be referred to as "channel C" below.
[0036] Here, in the audio transmission and reproduction system X, the audio transmission device 1 according to this embodiment classifies the source audio signal into a first category and a second category. Control is performed to binauralize only the signals of the first category among the classified signals to generate a binauralized signal (first category audio signal F). Then, the transmission signal T, which includes the binauralized first category audio signal F and the non-binauralized second category audio signal S, is transmitted to the audio reproduction device 2.
[0037] Simultaneously, the audio reproduction device 2 according to this embodiment acquires the transmission signal T transmitted by the audio transmission device 1 and controls it to binauralize the received second type audio signal S. The audio reproduction device 2 combines the binauralized second type audio signal S with the first type audio signal F that was binauralized when received to generate a binaural signal. Subsequently, the audio reproduction device 2 outputs the binaural signal.
[0038] In other words, in this embodiment, the audio transmission device 1 and the audio reproduction device are controlled to transmit and receive (transmit) data, such as transmission signal T.
[0039] In the transmission between the audio transmission device 1 and the audio reproduction device 2, there may be a delay of tens to hundreds of milliseconds or even longer. Furthermore, the length of the delay may vary depending on whether the connection is wireless or wired.
[0040] The functional configuration of each device in the audio transmission and reproduction system X will be described next.
[0041] The audio transmission device 1 according to this embodiment includes a binaural unit 11 (first binaural unit), a signal generation unit 12, and a transmission unit 13 as functional configurations.
[0042] The audio reproduction device 2 includes a signal acquisition unit 20, a binauralization unit 21 (second binauralization unit), a binaural signal generation unit 22, an audio output unit 23, and a directional transmission unit 24.
[0043] The binauralization unit 11 and binauralization unit 21 are arithmetic operation units, which perform binauralization by convolving HRTF and other signals with the audio signal.
[0044] In this embodiment, based on the listener's direction information D received from the audio reproduction device 2, the binauralization unit 11 binauralizes only the first type of signal in the source audio signal to generate the first type of audio signal F.
[0045] The binauralization unit 21 binauralizes the second type of audio signal S, which is the second type of signal acquired by the signal acquisition unit 20.
[0046] In this embodiment, the binauralization unit 11 can determine whether the auditory unnaturalness in the source audio signal is easily perceptible based on the directional information D received by the listener.
[0047] Here, the binauralization unit 11 can determine whether auditory unnaturalness is easily perceptible based on the azimuth and elevation angles for each channel C. In this case, for each channel C, the binauralization unit 11 can determine the azimuth and elevation angles based on the listener's direction information D and the direction information set for channel C.
[0048] Alternatively, for channel C of the source audio signal, the binauralization unit 11 can determine the low-frequency component as a first-type audio signal F. The low-frequency component can be an audio component divided into frequency bands, for example, 1000 to 500 Hz or lower.
[0049] The signal generation unit 12 generates a transmission signal T, which includes a first type of audio signal F that is binauralized by the binauralization unit 11 and a second type of audio signal S that is not binauralized.
[0050] The transmission unit 13 transmits the transmission signal T generated by the signal generation unit 12 in real time. Specifically, the encoded transmission signal T is transmitted by the audio transmission device 1 to the audio reproduction device 2 using the various wired or wireless methods described above.
[0051] In addition, the transmission unit 13 can also receive the listener's direction information D obtained by the audio reproduction device 2.
[0052] The signal acquisition unit 20 acquires the transmission signal T generated by the audio transmission device 1. The signal acquisition unit 20 also acquires the transmission signal T through any of the aforementioned wired or wireless methods.
[0053] The binaural signal generation unit 22 combines the audio signals from other channels binauralized by the binauralization unit 21, as well as the audio signals from some of these channels, to generate a binaural signal.
[0054] The audio output unit 23 outputs the binaural signal generated by the binaural signal generation unit 22 as binaural audio. In this embodiment, the audio output unit 23 includes, for example, a D / A converter, an amplifier for headphones, etc., and enables headphones such as HMDs to reproduce audio. Furthermore, the audio output unit 23 can drive a vibrating element for a bass speaker or a subwoofer.
[0055] Alternatively, the audio output unit 23 can encode the audio signal and output the encoded audio signal as an audio file or streaming audio for reproduction.
[0056] The directional transmission unit 24 acquires directional information D related to the listener's audio listening and transmits the directional information D to the audio transmission device 1. Specifically, the directional transmission unit 24 can acquire the listener's directional information D from a head tracking device (such as an HMD) and transmit the directional information D to the audio transmission device 1.
[0057] In this embodiment, audio signals from each direction of the content, audio signals from each content unit, audio signals from the audio source object, audio signals divided by frequency band, and multiple audio signals from remote call participants can be used as source audio signals.
[0058] This content can be, for example, a variety of things, such as movie or music data, games, audiobooks, e-book data capable of speech synthesis, television or radio broadcast data, various audio data related to operating instructions for car navigation systems and various home appliances, XR content, and other data capable of audio output. Movies also include musical performances, speeches, etc. Alternatively, content can include background music, sound effects, or MIDI files from games, voice call data from mobile phones or transceivers, or synthesized audio data from text in messaging programs. In this case, channel C can include the following as audio source objects: audio signals from objects (such as human voices, musical instruments, vehicles, and game characters); and audio signals from people (such as actors, narrators, rakugo storytellers, Kodan storytellers, and other speakers as audio sources). Furthermore, channel C can be audio signals from remote participants in one-to-one, one-to-many, or many-to-many live events or video conferencing systems between locations. For audio signals, spatial arrangement relationships are set as directional information within the content. Alternatively, the content may include: music files, such as background music, sound effects, or MIDI files from a game; voice call data from a mobile phone or transceiver; synthesized audio data from text in a messaging program; or other data that is not directly reproducible.
[0059] Alternatively, it can be used when the source audio signal is an audio signal from a remote participant, an audio signal from a microphone at the venue, or an audio signal spoken by a user (participant) on a PC (personal computer) or smartphone using various messaging programs or video conferencing application software (hereinafter referred to as "applications"). Directional information can be added by the mixing manager, such as the head orientation of participants captured by a camera, or the orientation of virtual avatars arranged in the virtual space.
[0060] Furthermore, in both cases, audio signals recorded by microphones or other devices connected via a network or directly can be used as the source audio signal. In this case, directional information can also be added to the audio signal. Alternatively, any combination of the above and the audio signal of a remote participant can be used. Moreover, even if channel C is not an audio signal, directional information can be set for channel C.
[0061] Furthermore, in this embodiment, the source audio signal can also serve as a “target signal” to reproduce the direction used for binauralization.
[0062] In this embodiment, directional information can be set for each channel C of the source audio signal. This directional information can be represented in an absolute or relative coordinate system relative to the listener. Furthermore, the details and types of the above-mentioned content can be set for each channel C.
[0063] In this embodiment, for example, in terms of wireless connectivity, the transmission signal T can be a signal of various codecs capable of transmitting and receiving multiple channels of audio signals via Bluetooth (registered trademark).
[0064] In this embodiment, the first type of audio signal F is an audio signal obtained by binauralizing a portion of the source audio signal by the audio transmission device 1.
[0065] Specifically, the first type of audio signal F can be an audio signal of some of the channels C-1 to Cn (hereinafter referred to as "some of these channels").
[0066] More specifically, the first type of audio signal F can be a signal in which the auditory unnaturalness caused by positional displacement or movement is less perceptible (almost imperceptible) than the auditory unnaturalness of the second type of audio signal S. This signal can be a part of the source audio signal in which no problem arises even if the HRTF update rate is very low.
[0067] On the other hand, the second type of audio signal S is the audio signal in the source audio signal that is binauralized by the audio reproduction device 2 but not by the audio transmission device 1.
[0068] Specifically, the second type of audio signal S can be an audio signal from other channels that are different from some of these channels.
[0069] More specifically, the second type of audio signal S can be a signal in which auditory unnaturalness caused by positional displacement or movement is more readily perceived than auditory unnaturalness in the first type of audio signal F. This signal can be a portion of the source audio signal in which the HRTF update rate needs to be high.
[0070] [Hardware configuration of audio transmission device 1 and audio playback device 2] The audio transmission device 1 and the audio reproduction device 2 include, for example, control tools (control units), such as various circuits including ASIC (Application-Specific Processor), DSP (Digital Signal Processor), CPU (Central Processing Unit), MPU (Microprocessor Unit), and GPU (Graphics Processing Unit).
[0071] Furthermore, the audio transmission device 1 and the audio reproduction device 2 may include storage units as storage tools (storage units), such as semiconductor memories (such as ROM (Read-Only Memory) or RAM (Random Access Memory)), magnetic recording media (such as HDD (Hard Disk Drive)), or optical recording media. ROM may include flash memory or other writable and appendable recording media. Alternatively, instead of HDD, SSD (Solid State Drive) may be provided. The storage unit may store control programs according to this embodiment, as well as various types of content. The control programs are programs for implementing each functional configuration and each method, including the audio signal processing program according to this embodiment. This control program includes embedded programs, such as firmware, OS (Operating System), and applications.
[0072] The contents of the storage unit according to this embodiment can be obtained by downloading files or data blocks transmitted via wire or wireless means, or by obtaining them in stages through streaming or other means.
[0073] In addition, applications such as media players for reproducing content, messaging applications, or video conferencing applications can be installed as applications according to this implementation scheme.
[0074] In addition, the audio transmission device 1 and the audio reproduction device 2 may also include direction calculation tools, including: a GNSS (Global Navigation Satellite System) receiver that calculates the direction the listener is facing; an indoor position and direction detector, an accelerometer, a gyroscope sensor, a geomagnetic sensor, etc., which are capable of head tracking; and circuitry that converts the output of the above units into direction information.
[0075] In addition, the audio transmission device 1 and the audio playback device 2 may include: a display unit, such as a liquid crystal display, an organic EL display, or an LED (light-emitting diode); an input unit, such as buttons, a keyboard, a pointing device (such as a mouse or a touch panel); and an interface unit for wirelessly or wiredly connecting to various devices. The interface unit may include: an interface for flash memory media, such as a microSD (registered trademark) card or a USB (Universal Serial Bus) memory; a LAN board; a wireless LAN board; and interfaces such as serial or parallel interfaces.
[0076] Furthermore, the audio transmission device 1 and the audio reproduction device 2 can utilize hardware resources to execute various programs primarily stored in storage via control tools to implement each method according to this embodiment. These programs include an audio signal processing program according to this embodiment.
[0077] It should be noted that some of the configurations or any combination thereof described above can be configured in hardware or circuitry using ICs, programmable logic, FPGAs (Field Programmable Gate Arrays), etc.
[0078] [Audio transmission and reproduction processing performed by audio reproduction device 2] Next, we will refer to Figure 2 and Figure 3 This describes the audio transmission and reproduction processing performed by the audio reproduction device 2 according to an embodiment of the present invention.
[0079] First, refer to Figure 2 This provides an overview of the audio transmission and reproduction processing according to this embodiment.
[0080] In the audio transmission and reproduction processing according to this embodiment, the audio signals (source audio signals) of multiple channels are classified into: a first type of audio signal F, in which a low HRTF update rate will not cause problems; and a second type of audio signal S, in which a higher HRTF update rate is required. The first type of audio signal F is then binauralized by the audio transmission device 1. On the other hand, the second type of audio signal S is binauralized by the audio reproduction device 2. Therefore, upon receiving communication, the second type of signal S, which requires a higher update rate, is updated. Thus, even if the communication time in the second type of signal S is delayed, the update of the head-related transfer function corresponding to head movement will not be delayed, thereby preventing unnatural sound imaging.
[0081] In each of the audio transmission device 1 and the audio reproduction device 2, the audio transmission and reproduction processing according to this embodiment is performed by controlling and executing control programs stored in the storage device using hardware resources in cooperation with each unit, or by each circuit directly.
[0082] In the following text, reference will be made to Figure 3 The flowchart describes the details of audio transmission and reproduction processing for each step.
[0083] (Step S101) First, the binaural unit 11 of the audio transmission device 1 performs audio separation processing.
[0084] For each channel C of the source audio signal, the binauralization unit 11 determines whether auditory unnaturalness is readily perceptible when the update of the head-related transfer function corresponding to head movement is delayed. Then, the channels C in which auditory unnaturalness is not readily perceptible (almost imperceptible) are separated as the first type of audio signal F.
[0085] In this embodiment, the binauralization unit 11 separates the source audio signal into a portion where the HRTF update rate needs to be high and a portion where no problem will occur even if the update rate is low.
[0086] Specifically, the binauralization unit 11 receives the listener's direction information D from the audio reproduction device 2 via the transmission unit 13. Furthermore, the binauralization unit 11 can acquire the direction of the audio for each channel C in the virtual space.
[0087] Then, based on the listener's orientation information D, the binaural unit 11 calculates the orientation (azimuth and elevation) of the channel C and the listener in the spatial arrangement including the virtual space.
[0088] Here, the binaural unit 11 determines whether auditory unnaturalness is easily perceptible based on azimuth and elevation angles.
[0089] As a specific example, the binauralization unit 11 can separate the signal from channel C as a first type of audio signal F, where the absolute values of the azimuth and elevation angles of the signal from channel C relative to the listener are equal to or greater than a specific angle threshold. The angle thresholds can be set separately for each of the azimuth and elevation angles based on experimental values, etc. Alternatively, the binauralization unit 11 can separate the signal from channel C as a second type of audio signal S, where the absolute values of the azimuth and elevation angles of the signal from channel C relative to the listener are less than a specific angle threshold.
[0090] In this example, the binauralization unit 11 determines that the auditory unnaturalness caused by positional displacement or movement of the signal of channel C, which is oriented upward, downward, or laterally relative to the listener, is not easily perceptible (almost imperceptible). Therefore, the binauralization unit 11 separates the signal of channel C as a first-type audio signal F and binauralizes the separated signal before it is transmitted by the audio transmission device 1.
[0091] Conversely, when the listener is positioned closer to the front, the binauralization unit 11 determines that auditory unnaturalness caused by positional displacement or movement is easily perceptible. Therefore, the binauralization unit 11 separates the signal from channel C as a second-type audio signal S.
[0092] For example, the binauralization unit 11 can separate the music performance channel C as a first-type audio signal F. This is because, since the instrument sound is diffuse in the component coming from the sides, head tracking is expected to be relatively slow.
[0093] On the other hand, the binauralization unit 11 can separate the human voice channel C as a second type of audio signal S. This is because, since the human voice is usually located in a forward horizontal direction, the auditory sensitivity caused by positional displacement or movement is relatively high. In other words, binauralization performed by the audio reproduction device 2 can achieve high-speed tracking of the listener's head movement.
[0094] Here, the binaural unit 11 can also determine the orientation of channel C in virtual space based on the details and type of the content of channel C or based on the setting information of channel C.
[0095] Alternatively, for each channel C of the source audio signal, the binauralization unit 11 can determine the low-frequency component and separate it as a first-type audio signal F. In this case, the binauralization unit 11 can determine the high-frequency component and separate it as a second-type audio signal S. That is, the binauralization unit 11 can separate the low-frequency component as the first-type audio signal F and the high-frequency component as the second-type audio signal S.
[0096] In this configuration, the binauralization unit 11 can use a bandpass filter (such as a low-pass or high-pass filter) to divide each channel C into a channel for low-frequency components and a channel for high-frequency components. The binauralization unit 11 can then identify the channel for low-frequency components as a first-type audio signal F. Furthermore, the binauralization unit 11 can identify the channel for high-frequency components as a second-type audio signal S.
[0097] Alternatively, the binaural unit 11 can calculate the ratio of low-frequency components to high-frequency components for each channel C to determine whether a channel is a channel of a first type of audio signal F or a channel of a second type of audio signal S based on the ratio.
[0098] In this example, since the auditory unnaturalness caused by positional displacement or movement is almost imperceptible, the binauralization unit 11 can separate the low-frequency component as a first-class audio signal F for binauralization by the audio transmission device 1.
[0099] On the other hand, since auditory unnaturalness caused by positional displacement or movement is easily perceived in the high-frequency components, the binauralization unit 11 can separate the high-frequency components as a second type of audio signal S for binauralization by the audio reproduction device 2. In this case, the high-frequency components have lower energy and therefore can be kept at a lower bit rate even when encoded in multiple channels. Furthermore, the high-frequency components play an important role in direction perception and can therefore be binauralized by the audio reproduction device 2 to track the listener's head movement at high speed.
[0100] (Step S102) Next, the binauralization unit 11 undergoes the first binauralization process.
[0101] Here, the binauralization unit 11 binauralizes each channel of the first type of audio signal F (a 2-channel conversion for the left and right) by convolving HRTF and the like based on the direction information of the sound source and the direction information D of the listener. That is, the binauralization unit 11 binauralizes the first type of audio signal F, which is determined to be a signal in which auditory unnaturalness is not easily perceived (almost imperceptible) even with transmission delays, etc. In the example above, many instrument sounds are separated as the first type of audio signal F and binauralized (converted to 2ch), thereby saving transmission bandwidth during encoding.
[0102] On the other hand, the binauralization unit 11 does not binauralize the second type of audio signal S, which is determined to be a signal in which auditory unnaturalness is easily perceived. That is, the second type of audio signal S is an audio signal in which auditory unnaturalness caused by delay, etc., is easily perceived, and is therefore binauralized by the binauralization unit 21 in the audio reproduction device 2, which is close to the listener. In other words, as described below, the second type of audio signal S is binauralized by the audio reproduction device 2 after the following transmission.
[0103] (Step S103) Next, the signal generation unit 12 performs transmission signal generation processing.
[0104] The signal generation unit 12 generates a transmission signal T, which includes a first type of audio signal F that has been binauralized and a second type of audio signal S that has not been binauralized.
[0105] For example, when transmitting signals via Bluetooth (registered trademark), the signal generation unit 12 can encode the first type audio signal F and the second type audio signal S using various codecs (such as AAC, AptX LL, HD, LDAC, and LHDC) to include them in the transmission signal T. Furthermore, the signal generation unit 12 can encode multiple channels (such as 2 channels (2ch), 5.1 channels (5.1ch), 7.1 channels (7.1ch), and 22.2 channels (22.2ch)) of audio signals into the transmission signal T according to audio encoding schemes (such as MPEG-2 AAC, MPEG-4 AAC, MP3, Dolby (registered trademark), and DTS (registered trademark)). Additionally, the signal generation unit 12 can bundle multiple channels of audio signals and control signals (control signals) into groups to generate the transmission signal T.
[0106] (Step S104) Next, the transmission unit 13 performs transmission signal transmission processing.
[0107] The transmission unit 13 transmits the transmission signal T in real time. At this time, the transmission unit 13 uses various wired or wireless methods to transmit the transmission signal T. Therefore, the binauralized first-type audio signal F and the non-binauralized second-type audio signal S can be transmitted in a "combined" manner.
[0108] (Step S201) The processing of the audio reproduction device 2 will be described here.
[0109] The signal acquisition unit 20 of the audio reproduction device 2 performs signal acquisition processing.
[0110] The signal acquisition unit 20 receives (acquires) the transmission signal T generated by the audio transmission device 1 and temporarily stores the transmission signal T in the storage device.
[0111] (Step S202) Next, the binauralization unit 21 undergoes a second binauralization process.
[0112] The binauralization unit 21 binauralizes the second type of audio signal S. The second type of audio signal S is an audio signal in which, as described above, auditory unnaturalness is easily perceived.
[0113] (Step S203) Next, the binaural signal generation unit 22 performs binaural signal generation processing.
[0114] The binaural signal generation unit 22 combines the audio signals from other channels binauralized by the binauralization unit 21, as well as the audio signals from some of these channels, to generate a binaural signal.
[0115] In this embodiment, the binaural signal may include an L signal for the listener's left ear and an R signal for the listener's right ear. Furthermore, the binaural signal may be a digital signal, which can be heard by the listener by decoding the digital data and reproducing the decoded data using the audio output unit 23.
[0116] (Step S204) Next, the audio output unit 23 performs audio output processing.
[0117] The audio output unit 23 outputs the audio signal generated by the binaural signal generation unit 22. This output can be, for example, a 2ch analog audio signal corresponding to the listener's left and right ears. Therefore, the audio output unit 23 can reproduce the audio signal corresponding to the virtual sound field as a 2ch audio signal via headphones. Thus, the listener can hear left and right three-dimensional audio.
[0118] (Step S205) Next, the direction transmission unit 24 performs direction transmission processing.
[0119] The orientation transmission unit 24 can acquire the orientation of the listener's head and orientation information (such as the orientation of a virtual avatar in virtual space) from a head-tracking device using a gyroscope sensor on an HMD or smartphone.
[0120] The directional transmission unit 24 transmits the directional information D acquired by the listener to the audio transmission device 1. In this case, the directional transmission unit 24 can transmit the listener's directional information D directly and separately from the transmission signal T via a wired or wireless connection. Alternatively, the directional transmission unit 24 can be configured to transmit the listener's directional information D via the signal acquisition unit 20 through the same wireless or wired connection (such as Bluetooth) as the connection to the transmission signal T.
[0121] As can be seen from the above, the audio transmission and reproduction processing according to this implementation plan is now complete.
[0122] The above configuration can achieve the following effects.
[0123] In recent years, it has become common to connect smartphones to HMDs or headphones to enjoy virtual sound fields (3D audio) for VR and AR.
[0124] However, in related technologies, in the reproduction of 3D audio headsets, in order to generate binaural signals, multiple audio source signals are convolved with the HRIR of their corresponding audio source directions on the transmission side (such as on a smartphone). That is, head-out-of-head localization is achieved, and 3D audio is generated by convolving the head-related transfer function of the audio source direction with the audio signal of each channel of each audio source signal to achieve binaural localization.
[0125] In a head-tracking system, the relative azimuth between the sound source and the listener changes when the listener moves their head. Therefore, the listener's head direction is detected, this direction information is transmitted to a smartphone or similar device, and the HRIR (Head-Responsive Audio Context) is updated accordingly. In other words, the azimuth of the HRTF (Head-Responsive Audio Context) to be used is calculated from the relationship between the listener's head direction and the sound source position, and a stationary sound source is generated by updating the azimuth in real time. The same process is applied to moving sound sources, generating moving sound sources based on the listener's head movement.
[0126] However, latency in communication between smartphones and headphones has always been a problem. Specifically, for wireless headphones using Bluetooth (registered trademark), which has a significant latency, the total time required to transmit orientation information detected by head tracking to a smartphone and then to transmit information related to HRTF updates based on the detected orientation information back to the headphones is approximately 200 milliseconds. In other words, HRTF updates are delayed by 200 milliseconds or more relative to head movement.
[0127] Therefore, headphones using head tracking in related technologies have poor performance in tracking the listener's head movement and suffer from the following problems: the sound image does not stabilize relative to the listener's movement. That is, due to round-trip communication time, there is a delay in updating HRTF, etc., and the sound image becomes unnatural due to the listener's head movement. Specifically, when the listener rotates their head, the sound image of the stationary sound source temporarily follows the head rotation.
[0128] In fact, this problem has been observed in consumer-grade headphones and other devices capable of head tracking. Unless the latency is around 60 milliseconds or less, the unnaturalness of auditory localization is noticeable.
[0129] On the other hand, (A) the signal transmission method according to an embodiment of the present disclosure is a signal transmission method for transmitting a transmission signal T containing multiple channels of audio signals, the signal transmission method comprising: real-time transmission of a transmission signal T containing a binauralized first type of audio signal F and a non-binauralized second type of audio signal S.
[0130] In signal transmission associated with an audio signal comprising three or more channels C, this configuration provides a signal transmission method for transmitting a group of signals in a state where some signals in the group are binauralized as a first type, and others are not binauralized as a second type. Therefore, by sharing the binauralization between the transmission and reception sides, the second type of audio signal can be binauralized at the listener's side, thereby suppressing unnaturalness in the sound image. That is, the effect of transmission delay can be minimized, making the unnaturalness virtually imperceptible to the listener.
[0131] Therefore, this invention is applicable to various XR application software, such as games, movies, and home TVs.
[0132] Furthermore, in related technologies, the binauralization of multiple channels of audio signals for headphone reproduction has been performed by transmission-side devices such as smartphones. In this case, the audio signal to be transmitted is a 2-channel binaural signal. Therefore, the sound quality degradation caused by encoding for transmission is minimized. However, as mentioned above, in this case, the following problem arises: the sound image becomes unnatural due to the image tracking delay caused by transmission latency.
[0133] On the other hand, to eliminate the unnaturalness of the sound image caused by such delays, when binaural processing is performed at the receiving end (such as with headphones) rather than at the transmitting end, there is no communication-induced delay in updating the HRTF, and the HRTF can be updated instantly according to the listener's head orientation. However, in such cases, the audio signals from multiple channels need to be transmitted intact to the receiving end before being converted into binaural signals. Therefore, with the increase in transmission rate (transmission capacity), efficient coding with high compression ratios is required. Consequently, sound quality degradation may occur due to coding distortion.
[0134] In contrast, the signal transmission method according to an embodiment of the present invention can transmit signals while simultaneously tracking head movement, without the degradation caused by encoding. That is, the signal-to-noise ratio (SNR) of the audio can be improved.
[0135] Therefore, by applying the signal transmission method according to this embodiment to a 3D audio reproduction system (such as a wireless headphone reproduction system connected to a smartphone), higher sound quality can be achieved than with related technologies. Furthermore, by applying the signal transmission method to smartphones, home appliances, etc., three-dimensional audio with multiple channels of audio signals can be generated with high quality, and the signal transmission method is compatible with international standards, etc.
[0136] (B) The signal transmission method according to an embodiment of the present invention is the signal transmission method according to (A), wherein the binauralized first type of audio signal F is the audio signal of some of these channels, and the non-binauralized second type of audio signal is the audio signal of other channels.
[0137] In this configuration, in signal transmission associated with an audio signal comprising three or more channels C, the first type of channel C is binauralized before transmission, the second type of channel C is binauralized after reception, and the first and second types of channel C are combined at the receiving end to generate a binaural signal. That is, the source audio signal can be classified into two categories based on channel C, including a first type of audio signal F and a second type of audio signal S, with dynamic bit allocation between these two categories. In other words, the allocation of transmission bandwidth can be dynamically changed. Thus, by performing combined transmission (where binauralization is shared between the transmitting and receiving ends), high-quality audio signal transmission can be achieved while suppressing the unnaturalness of the sound field caused by transmission delay.
[0138] In this case, for example, a configuration can be implemented in which channel C belonging to the first type of audio signal F is binauralized by the audio transmission device 1 before being encoded, and channel C belonging to the second type of audio signal S is encoded and binauralized on a channel-by-channel basis after being decoded by the audio reproduction device 2.
[0139] (C) The signal transmission method according to an embodiment of the present invention is the signal transmission method according to (A) or (B), wherein the second type of audio signal S is a signal in which auditory unnaturalness caused by position displacement or movement is more easily perceived compared to the first type of audio signal F.
[0140] In this configuration, the audio signals from multiple channels are classified into a first type of audio signal F and a second type of audio signal S. In the first type of audio signal F, since the auditory unnaturalness caused by the positional displacement of the sound source direction is not easily perceived (almost imperceptible), a low HRTF update rate is not a problem. In the second type of audio signal S, since the auditory unnaturalness caused by the positional displacement of the sound source direction is easily perceived, a higher HRTF update rate is required. Then, only the first type of audio signal F, where the auditory unnaturalness is not easily perceived (almost imperceptible), is binauralized by the audio transmission device 1, and the second type of audio signal S, where the auditory unnaturalness is easily perceived, is binauralized by the audio reproduction device 2.
[0141] Therefore, while minimizing the number of channels of the signal to be transmitted, high-speed updates of HRTF can be achieved for the necessary audio sources, thus enabling high-quality sound field reproduction with stable positioning even in the presence of head movement.
[0142] (D) The signal transmission method according to an embodiment of the present invention is the signal transmission method according to (C), wherein whether auditory unnaturalness can be easily perceived in the source audio signal that serves as an audio signal of multiple channels is determined based on azimuth and elevation angles.
[0143] In this configuration, each channel C can be grouped into two categories based on azimuth and elevation angles.
[0144] Therefore, for example, for instrument sounds with diffuse characteristics in the lateral components, tracking can be performed relatively slowly. Thus, instrument sounds can be identified as a first-class audio signal F, where auditory unnaturalness is almost imperceptible. Therefore, by using the audio transmission device 1 to binauralize the instrument sounds, signal transmission capacity can be suppressed. That is, by pre-binauralizing many instrument sounds into two channels, namely the left channel and the right channel (2ch), the capacity during encoding can be significantly reduced.
[0145] On the other hand, for example, the unnaturalness of hearing is easily perceived in human voices, which are typically positioned horizontally in front. Therefore, human voices can be identified as a second type of audio signal S that is sensitive to positional displacement. Thus, only the second type of audio signal S of human voices is transmitted as a separate signal, while the first type of audio signal F of instrumental sounds is pre-dichotomized and transmitted. Then, by using the audio reproduction device 2 to binauralize the second type of audio signal S, high-speed tracking of head movements can be achieved.
[0146] (E) The signal transmission method according to an embodiment of the present invention is the signal transmission method according to (C), wherein the first type of audio signal F is a low-frequency component and the second type of audio signal is a high-frequency component.
[0147] In this configuration, each channel C can be grouped into two classes based on the transmission bandwidth. Thus, for example, low-frequency components can be binauralized before transmission, while high-frequency components can be binauralized after reception, and these components can be combined at the receiving end to generate a binaural signal. Specifically, for example, the low-frequency components of a musical instrument sound can be binauralized before transmission, and only the high-frequency components can be encoded in a separate channel and binauralized at the receiving end. In this case, the high-frequency components of the audio signal have lower energy and therefore can be kept at a lower bit rate (transmission bandwidth) even when encoded in multiple channels. On the other hand, since high-frequency components are important for the listener's sense of direction, high-speed tracking of head movements is required. Therefore, this configuration is suitable.
[0148] (F) According to an embodiment of the present invention, the signal generation method is a signal generation method for generating a transmission signal T containing multiple channels of audio signals. The signal generation method includes: binauralizing only a first type of audio signal F based on the direction information D received by the listener; and generating a transmission signal T containing the first type of audio signal F binauralized by the binauralization unit 11 and a second type of audio signal S that is not binauralized.
[0149] In this configuration, in signal transmission associated with an audio signal containing three or more channels C, only the first type of audio signal F can be binauralized by the audio transmission device 1 (such as a smartphone) based on the listener's directional information D, where delay is tolerable.
[0150] (G) The signal reproduction method according to an embodiment of the present invention is a signal reproduction method for reproducing a transmission signal T containing an audio signal with multiple channels, the signal reproduction method comprising: acquiring a transmission signal T transmitted by a signal transmission method according to any one of (A) to (E) or a transmission signal T generated by a signal generation method according to (F); binauralizing a received second type audio signal S; and combining the binauralized second type audio signal S with a first type audio signal F that was binauralized when received to generate a binaural signal.
[0151] In this configuration, in signal reproduction associated with an audio signal containing three or more channels, a binaural signal can be generated by combining some channels C in the signal group that were binauralized before transmission with other channels C that were binauralized after reception. That is, in signal reproduction associated with the source audio signal, a binaural signal can be generated by combining a portion of the received transmitted signal T unchanged with another portion of the transmitted signal T that was binauralized after reception. Therefore, for source audio signals in applications such as content creation, point-to-point connections, point-to-multipoint connections, multipoint-to-multipoint connections, and teleconferencing, a sound field reproduction in which unnaturalness is virtually imperceptible to the listener can be achieved while minimizing the encoding distortion of the transmitted signal T. Furthermore, realistic audio can be experienced by outputting binaural audio signals via headphones or HMDs.
[0152] (H) According to an embodiment of the present invention, the audio transmission and reproduction system X is a reproduction system comprising: an audio transmission device 1 for generating a transmission signal T containing multiple channels of audio signals; and an audio reproduction device 2 for reproducing the transmission signal T generated by the audio transmission device, wherein the audio transmission device comprises: a binauralization unit 11 for binauralizing only a first type of audio signal F based on the listener's orientation information D received from the audio reproduction device 2; a signal generation unit 12 for generating a transmission signal T comprising the first type of audio signal F binauralized by the binauralization unit 11 and a non-binauralized second type of audio signal S; and a transmission unit... The audio reproduction device 2 includes: a signal acquisition unit 20, which acquires the transmission signal T generated by the audio transmission device 1; a binauralization unit 21, which binauralizes the second type of audio signal S acquired by the signal acquisition unit 20; a binaural signal generation unit 22, which combines the second type of audio signal S binauralized by the binauralization unit 21 with the second type of audio signal F to generate a binaural signal; an audio output unit 23, which outputs the audio signal generated by the binaural signal generation unit 22; and a direction transmission unit 24, which transmits the listener's direction information D to the audio transmission device 1.
[0153] This configuration minimizes both signal encoding distortion during transmission and the impact of transmission delay, making the unnaturalness virtually imperceptible to the listener.
[0154] [Other Implementation Plans] In the above implementation, examples have been described in which each channel of the source audio signal from channel C-1 to Cn is classified unchanged into a first type of audio signal F and a second type of audio signal S according to channel C, and examples have been described in which each channel from channel C-1 to Cn is separated into a low-frequency component or a high-frequency component.
[0155] However, in the first type of audio signal, the channels C-1 to Cn of the source audio signal (containing several to hundreds (n) signals existing in 3D (three-dimensional) space) can be grouped into m specific directions, fewer than the number of channels n, through clustering or translation. In this case, signals for the m directions can be generated within the audio transmission device 1 by applying an appropriate gain to each signal of channels C-1 to Cn and calculating the sum of the time-shifted signals, and then the generated signals can be grouped into representative point signals located at representative points in the m directions. The grouped representative point signals can be transmitted to the audio reproduction device 2. That is, in this example, the signals for the m directions (representative point signals) can be regarded as the first type of audio signal. Then, the signals for the m directions (representative point signals) can be binauralized by convolving HRIR, etc., using the audio reproduction device 2. Furthermore, in this case, the non-binauralized second type of audio signal can be combined with the first type of audio signal as a transmission signal for transmission. Alternatively, the second type of audio signal at this time can be a control signal, etc., for encoding the audio signal.
[0156] In this case, as in the above-described implementation, a delay may occur in the audio transmission between the audio transmission device 1 and the audio reproduction device 2. The audio reproduction device 2, which receives the transmitted signal, can group the m signals (binaurally) and simultaneously update the head orientation information and the HRIR corresponding to the azimuth angles of the m signals in real time based on the listener's head movement.
[0157] At this time, in the audio transmission device 1, the sum of the signals of channel C grouped at representative points can be obtained to generate a sum signal, and this sum signal can be transmitted. The binauralization unit 21 of the audio reproduction device 2 can binauralize the sum signal by convolving the HRIR (HRIR in the direction of the representative point) of the position of the representative point to generate a binaural signal.
[0158] That is, in a signal transmission method according to another embodiment of the present invention, the first type of audio signal can be a representative point signal, wherein the signal (source audio signal) containing multiple channels of audio signal is grouped into directions with a number less than the number of channels of the source audio signal.
[0159] Alternatively, a signal reproduction method according to another embodiment of the present invention may include: generating a representative point signal by an audio transmission device 1 (wherein the source audio signal is grouped into directions fewer than the number of channels of the source audio signal) and transmitting the representative point signal in real time; and binauralizing the representative point signal by an audio reproduction device 2 to generate a binaural signal.
[0160] With this configuration, the audio transmission device 1 can group several to several hundred (n) signals that need to be convolved with HRIR in XR or the like into signals for m directions, and can transmit the grouped signals to the audio reproduction device 2 for binaural processing. By grouping the signals into m signals, the transmission bandwidth can be suppressed, and wireless transmission can be more easily achieved. Furthermore, compared with the method in related technologies that convolves the audio source signal with HRIR one by one, the amount of computation can be reduced.
[0161] Furthermore, the binauralization achieved by the audio reproduction device 2 allows control over the azimuth angle of the virtual speaker, which is a representative direction of translation combined with the listener's head movement. Therefore, the sound source position can be updated instantly in response to head movement without being affected by transmission delay. Consequently, the HRIR can be updated instantly relative to head movement, and delay does not cause problems. In other words, by reducing the number of sound source objects to be reproduced via translation, wireless transmission from smartphones to headphones can be achieved. Therefore, highly realistic HMDs, wireless headphones, and the like can be realized.
[0162] In the above implementation scheme, an example has been described in which the listener's angular orientation in the up, down, left, and right directions is taken into account as directional information.
[0163] However, configurations that only consider the left and right directions or configurations that consider 6DOF (six degrees of freedom) can be used.
[0164] Furthermore, in the above implementation, examples have been described in which the determination based on azimuth and elevation angles or based on low-frequency and high-frequency components is used to determine whether auditory unnaturalness is easily perceptible.
[0165] However, in the binaural unit 11, the determination based on azimuth and elevation angles and the determination based on low-frequency and high-frequency components can be combined.
[0166] For example, when the human voice channel C is separated as a second type of audio signal S, the low-frequency component of the instrument sound channel C can be separated as a first type of audio signal F, and the high-frequency component of the instrument sound channel C can be separated as a second type of audio signal S in a channel different from the human voice channel C.
[0167] Furthermore, as described in the above implementation, when separating the first type of audio signal and the second type of audio signal, bit allocation is performed, that is, bandwidth allocation for transmission corresponding to the energy of each component.
[0168] However, the bandwidth of the transmitted signal can vary between the first type of audio signal and the second type of audio signal. For example, the first type of audio signal can have a narrower bandwidth compared to the second type of audio signal. In this case, only the bandwidth of the low-frequency components can be reduced, while the bandwidth of the high-frequency components is set to correspond to the same energy level as in the above embodiment.
[0169] Alternatively, the bandwidth ratio between the first type of audio signal and the second type of audio signal can be configured to vary according to the listener's directional information D, such as movement speed and acceleration.
[0170] Alternatively, the bandwidth ratio between the first and second type audio signals can be configured to vary depending on whether the connection is wireless or wired, or simply based on the magnitude of the transmission delay. In this case, when the delay is initially less than tens of milliseconds due to a high-speed wired connection, the first and second type audio signals can be configured not to be separated from each other.
[0171] This configuration saves transmission bandwidth while making the positional displacement of the sound source almost imperceptible to the listener, and enables high-quality audio transmission.
[0172] In the above implementation, it has been described that, during binauralization, HRIR is convolved with the audio signal of channel C.
[0173] However, the same process can be performed by converting the audio signal of channel C to the frequency domain and applying HRTF.
[0174] Furthermore, HRTF can also be configured to be applied only when the low-frequency and high-frequency components are separated to achieve combined transmission. Specifically, as in the embodiments described above, by using HRTFs in both the low-frequency and high-frequency ranges, referencing frequencies near or above the frequency band where human hearing sensitivity is high, more accurate synthesis can be achieved.
[0175] Furthermore, although the above implementation does not describe the differences between HRTFs, etc., in practice, the binaural unit 21 may be able to select the user's personal HRIR from the HRIR table, the HRIR generated by the HRIR database, etc.
[0176] Furthermore, when the speaker and listener are transformed into virtual avatars in a virtual space, the binauralization unit 21 can also select an HRIR accordingly. For example, when the virtual avatar has an image such as a cat or rabbit with its ears pointing upwards, an HRIR that provides a listening experience corresponding to the shape can be selected.
[0177] Furthermore, the binaural unit 21 can further enhance the sense of realism by using convolution and other methods to separately superimpose the direct sound of channel C and the reflected sound from the environment.
[0178] This configuration allows for the reproduction of more realistic and clearer sound.
[0179] Furthermore, in the above embodiments, an example has been described in which the audio output unit 23 reproduces audio in two left and right channels.
[0180] In this regard, multiple channels can also be configured to be reproduced via headphones or the like that play these channels separately. Furthermore, the binaural signal generation unit 22 can be configured to generate signals for the subwoofer and vibration respectively and reproduce the generated signals.
[0181] As described in the above implementation, the audio reproduction device 2 is configured as an integral unit.
[0182] However, the audio reproduction device 2 can be configured such that a device receiving audio is connected to a device that actually reproduces the audio, such as a headset kit, headphones, or separate left and right earphones. That is, in addition to the audio transmission device 1, the audio reproduction device 2 itself can be configured with multiple devices. Furthermore, the audio transmission device 1 itself can be configured with multiple devices.
[0183] Furthermore, in the above embodiments, an example has been described in which the distance between the audio transmission device 1 and the audio reproduction device 2 is relatively short, but this distance can also be quite long. In this case, the function of the audio transmission device 1 can also be performed via an intranet or a server on the Internet.
[0184] In addition, a configuration can be adopted in which the audio reproduction device 2 reproduces the contents stored in the optical medium or flash memory card, the generated sound source object, etc. as the source audio signal, and the first type of audio signal F in the source audio signal is transmitted to the audio transmission device 1 for binaural conversion.
[0185] Furthermore, in the above embodiments, the audio reproduction device 2 has been described as including a binaural signal generation unit 22 and an audio output unit 23. However, a configuration excluding the binaural signal generation unit 22 and the audio output unit 23 may also be used. In this case, for example, the generated audio signal (such as binaural signal) data can be stored on a recording medium. In this case, other audio reproduction devices can be any device capable of acquiring the direction of channel C in virtual space, such as devices including televisions or monitors, video telephony calls via monitors, video conferencing, or telepresence.
[0186] In the above implementation, an example in which direction information is added to the audio signal of channel C has been described.
[0187] In this regard, when the roles of speaker and listener are constantly changing (such as in the aforementioned remote conference), it is not necessary to add directional information to the audio signal of channel C. In other words, when the current listener is the speaker, the speaker's (current listener's) direction can be estimated using the spoken audio signal, and this direction can be taken as the listener's direction from the current speaker's perspective.
[0188] In this scenario, the directional transmission unit 24 calculates, for example, the directions of arrival of the audio signals as seen by the listener, the L (left) channel signal (hereinafter referred to as the "L signal") and the R (right) channel signal (hereinafter referred to as the "R signal"). At this point, the directional transmission unit 24 can obtain the intensity ratio between the L and R channels. From the intensity ratio, the direction of arrival of the signal for each frequency component can also be estimated.
[0189] Alternatively, the directional transmission unit 24 can estimate the direction of arrival of the audio signal from the relationship between the ITD (interaural time difference) of the signal at each frequency in the HRTF and the direction of arrival. As a relationship between ITD and direction of arrival, the directional transmission unit 24 can refer to the relationship between ITD and direction of arrival stored in the storage unit as a database.
[0190] Alternatively, the speaker's or listener's orientation can be estimated through facial recognition from facial image data of people such as speakers or listeners in content or video conferences. That is, orientation can be estimated even in configurations where head tracking is not performed. Similarly, the speaker's or listener's position in space can also be detected.
[0191] This configuration can respond to a variety of flexible setups. Furthermore, in an XR where the sound source position is pre-set, the direction of channel C can be obtained from the positional relationship between channel C and the listener L, without needing to estimate the sound source direction.
[0192] It goes without saying that the configuration and operation of the above implementation scheme are merely examples and can be appropriately modified within the scope of this implementation scheme.
[0193] Industrial applicability The audio transmission device according to this embodiment can provide an audio transmission and reproduction system that prevents the sound image from becoming unnatural when generating three-dimensional audio, and is industrially applicable.
[0194] List of reference numerals 1. Audio transmission device 2 Audio reproduction device 11, 21 Binaural Units 12 Signal Generation Units 13 Transmission Units 20 Signal Acquisition Units 22 Binaural Signal Generation Units 23 Audio Output Unit 24-directional transmission unit D. Directional information from the listener F. Type I audio signals S Type II audio signal T transmit signal Channels C-1 to Cn X Audio Transmission and Reproduction System
Claims
1. A signal transmission method for transmitting an audio signal comprising multiple channels, the signal transmission method comprising: The signal is transmitted in real time, comprising a first type of audio signal that is binauralized and a second type of audio signal that is not binauralized.
2. The signal transmission method according to claim 1, wherein the binauralized first type of audio signal is an audio signal of some of the channels, and the non-binauralized second type of audio signal is an audio signal of other channels.
3. The signal transmission method according to claim 1, wherein the second type of audio signal is a signal whose auditory unnaturalness caused by positional displacement or movement is more easily perceived compared to the first type of audio signal.
4. The signal transmission method according to claim 3, wherein whether the auditory unnaturalness in the audio signals of the plurality of channels is easily perceptible is determined based on the azimuth and elevation angles.
5. The signal transmission method according to claim 3, wherein the first type of audio signal is a low-frequency component, and the second type of audio signal is a high-frequency component.
6. The signal transmission method according to claim 1, wherein the first type of audio signal is a representative point signal, and the audio signal containing the plurality of channels is grouped into directions in a number less than the number of channels of the signal.
7. A signal generation method for generating an audio signal containing multiple channels, the signal generation method comprising: Based on the directional information received by the listener, only the first type of audio signal is binauralized; as well as The signal is generated, which includes the first type of audio signal that has been binauralized and the second type of audio signal that has not been binauralized.
8. A signal reproduction method for reproducing an audio signal containing multiple channels, the signal reproduction method comprising: Acquire the signal transmitted by the signal transmission method according to any one of claims 1 to 6 or the signal generated by the signal generation method according to claim 7; The received Type II audio signal is binauralized; and The binauralized second type of audio signal is combined with the binauralized first type of audio signal received to generate a binaural signal.
9. A signal reproduction method performed by an audio generation and reproduction system, the audio generation and reproduction system comprising: An audio transmission device that generates and transmits an audio signal containing multiple channels; And an audio reproduction device that reproduces the signal transmitted by the audio transmission device, the signal reproduction method comprising: The audio signal, which includes the plurality of channels, generated by the audio transmission device, is grouped into representative point signals in a direction fewer in number than the number of channels of the signal, and the representative point signals are transmitted in real time; and The audio reproduction device binauralizes the representative point signal to generate a binaural signal.
10. An audio signal processing program executed by an audio transmission device, The audio signal processing program causes the audio transmission device to perform the following: Based on the direction information received by the listener, only the first type of audio signal is binauralized; and Generate a signal that includes the binauralized first-class audio signal and the non-binauralized second-class audio signal.
11. An audio signal processing program executed by an audio playback device, The audio signal processing program causes the audio reproduction device to perform the following: Acquire the signal transmitted by the signal transmission method according to any one of claims 1 to 6 or the signal generated by the signal generation method according to claim 7; The received Type II audio signal is binauralized; and The binauralized second type of audio signal is combined with the binauralized first type of audio signal received to generate a binaural signal.
12. An audio transmission apparatus for generating an audio signal comprising multiple channels, the audio transmission apparatus comprising: The first binauralization unit binauralizes only the first type of audio signal based on the listener's orientation information received from the audio reproduction device; as well as A signal generation unit generates the signal comprising a first type of audio signal binauralized by the first binauralization unit and a second type of audio signal not binauralized.
13. An audio reproduction apparatus for reproducing an audio signal comprising multiple channels, the audio reproduction apparatus comprising: A signal acquisition unit acquires the signal comprising a binauralized first type of audio signal and a non-binauralized second type of audio signal; The second binauralization unit binauralizes the second type of audio signal acquired by the signal acquisition unit; A binaural signal generation unit combines the second type of audio signal binauralized by the second binauralization unit and the first type of audio signal to generate a binaural signal; as well as An audio output unit that outputs the binaural signal generated by the binaural signal generation unit.
14. An audio generation and reproduction system, the audio generation and reproduction system comprising: An audio transmission device that generates and transmits an audio signal containing multiple channels; as well as An audio reproduction device that reproduces the signal transmitted by the audio transmission device, wherein... The audio transmission device includes: The first binauralization unit binauralizes only the first type of audio signal based on the listener's orientation information received from the audio reproduction device; A signal generation unit that generates the signal comprising a first type of audio signal binarized by the first binauralization unit and a second type of audio signal not binauralized; and The transmission unit transmits the signal generated by the signal generation unit in real time, and The audio reproduction device includes: A signal acquisition unit acquires the signal generated by the audio transmission device; The second binauralization unit binauralizes the second type of audio signal acquired by the signal acquisition unit; A binaural signal generation unit combines the second type of audio signal binauralized by the second binauralization unit and the first type of audio signal to generate a binaural signal; An audio output unit that outputs the binaural signal generated by the binaural signal generation unit; and A directional transmission unit that transmits the listener's directional information to the audio transmission device.
Citation Information
Patent Citations
Sound processing device and sound processing method
JP2021005822A