Acoustic signal control method, acoustic device, and program
The sound enhancement system using a speaker and open-air earphones addresses the challenge of achieving immersive audio experiences by enhancing spatial perception and realism without costly equipment, providing a personalized 3D sound effect.
Patent Information
- Application Number
- PCT/JP2024/028648
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional audio technologies face challenges in providing a highly immersive audio experience without the need for expensive multiple speaker equipment and achieving personalized 3D sound effects using headphones, often resulting in localized sound images within the user's head.
A sound enhancement system using a speaker and open-air earphones that emit an auxiliary signal, processed by a sound enhancement device to create a spatially enhanced audio experience, enhancing the sense of spaciousness and envelopment.
Enhances the immersive value of audio experiences by adding a stereophonic effect to realistic speaker sound without the need for expensive equipment, improving spatial perception and realism.
Smart Images

Figure JP2024028648_12022026_PF_FP_ABST
Abstract
Description
Acoustic signal control method, audio device, and program
[0001] The disclosed technology relates to a technology that enriches the sound heard by a listener (giving the listener a sense of spaciousness in the sound or making the listener feel as if they are surrounded by the sound) by adjusting the gain and delay and superimposing sound from the same sound source from open-ear earphones on simple sound from speakers for monaural and stereo playback.
[0002] Based on the "Law of the First Wave Front," we will explain the "sense of spaciousness" and "sense of envelopment" of sound. [Law of the First Wave Front (Precedence Effect)] When perceiving the direction of a sound source (or sound image) in a space with reflected sound, the directional information of the sound that reaches the listener first is dominant. It is known that reflected sound that arrives between approximately 1 ms and 30 ms compared to direct sound does not significantly interfere with the directional perception of direct sound. This phenomenon in which a sound image is localized in the direction of the first arriving sound, even in the presence of delayed arriving sounds, is called the Law of the First Wave Front, or the precedence effect.
[0003] The following explanation will be given with reference to Figure 1. [Spaciousness] The size of the sound image perceived in the direction from which the preceding sound comes, blending with the preceding sound in both time and space due to the precedence effect, is called the "apparent sound source width" (Reference 1). The "apparent sound source width" contributes to the "spaciousness" of the sound. [Envelopment] The feeling that the listener is surrounded by sound images other than the apparent sound source is called the "envelopment" (Reference 1).
[0004] Reference 1: Masayuki Morimoto et al., "Differences between apparent sound source width and the feeling of being surrounded by sound," Journal of the Acoustical Society of Japan, Vol. 46, No. 6, pp. 449-457, 1990
[0005] Surround playback or stereophonic playback is a method for improving the sense of spaciousness and envelopment of sound to achieve a highly immersive audio experience. This requires multiple speaker equipment, such as 5.1ch surround playback (Non-Patent Document 1). On the other hand, as a stereophonic playback technology using only headphones, Apple offers spatial audio that utilizes Dolby Atmos (Non-Patent Document 2), and Sony offers 360 Reality Audio (Non-Patent Document 3).
[0006] ITU, RECOMMENDATION ITU-R BS.775-1, "Multichannel stereophonic sound system with and without accompanying picture", 1992-1994. Apple, "Dolby Atmos spatial audio", [Retrieved July 29, 2024]<https: / / www.apple.com / jp / newsroom / 2021 / 05 / apple-music-announces-spatial-audio-and-lossless-audio / > SONY, "360 Reality Audio", [Retrieved July 29, 2024], Internet<https: / / www.sony.jp / headphone / special / 360_Reality_Audio / > ,
[0007] Despite the desire for a highly immersive audio experience, conventional technologies face the following challenges: Adopting a surround playback method using only speakers requires the preparation of multiple speaker equipment, resulting in high costs for procuring and installing the equipment; Adopting a 3D sound playback method using only headphones is not easy to tune to obtain the optimal 3D sound effect for each individual user, and failure to achieve personal optimization results in, for example, sound images being localized within the user's head, resulting in a decrease in the value of the experience.
[0008] To solve the above problem, the disclosed technology provides a sound enhancement system that includes a speaker, open-air earphones, and a sound enhancement device. The speaker emits a sound signal, and the open-air earphones emit an auxiliary signal. The sound enhancement device acquires the sound signal and generates the auxiliary signal from the sound signal. The auxiliary signal is generated from the sound signal so that a user perceives the sound as being spatially enhanced when listening to both the sound signal and the auxiliary signal, compared to when the user listens to only the sound signal.
[0009] With the disclosed technology, simply by putting on earphones, a stereophonic effect is added to the realistic speaker sound. This increases the sense of immersion and realism, improving the experiential value of the content. There is also a cost advantage in that there is no need to prepare expensive equipment.
[0010] 1 is a diagram illustrating the "sense of spaciousness" and "sense of envelopment" of sound. FIG. 1 is a functional block diagram of a sound enhancement system according to a first embodiment. FIG. 2 is a functional block diagram of a sound enhancement device according to a first embodiment. FIG. 3 is a flowchart illustrating an example of the operation of a sound enhancement device according to a first embodiment. FIG. 4 is a diagram illustrating spatial transfer characteristics and presentation sound transfer characteristics. FIG. 5 is a flowchart illustrating an example of the operation of a sound enhancement device according to a first modified embodiment. FIG. 6 is a functional block diagram of a sound enhancement system according to a second embodiment. FIG. 7 is a functional block diagram of a sound enhancement device according to a second embodiment. FIG. 8 is a flowchart illustrating an example of the operation of a sound enhancement device according to a second embodiment. FIG. 9 is a diagram illustrating the positional relationship between stereo speakers and open-air earphones. FIG. 10 is a diagram illustrating an example of the functional configuration of a computer.
[0011] The following describes in detail embodiments of the disclosed technology. Components having the same functions are assigned the same numbers, and redundant explanations will be omitted. Superimposing the sound of a specific earphone on the sound of a speaker to create a "sense of spaciousness" or "envelopment" will be referred to as "sound enhancement."
[0012] [First Embodiment] FIG. 2 is a functional block diagram of a sound reinforcement system according to a first embodiment. The sound reinforcement system 2 includes a speaker 201, open-air earphones 202L and 202R, a dry source 204, an amplifier 205, and a sound reinforcement device 206. The open-air earphones allow a listener 203 to hear both the sound emitted by the earphones and the surrounding sounds. Therefore, the listener 203 can simultaneously hear the sound from the speaker and the earphones. The open-air earphones 202L and 202R are worn on the left and right ears, respectively. The term "open-air earphones" also includes wearable audio devices, such as earphones, headphones, and wearable speakers, that allow a user to hear both the sound emitted by the audio device and the surrounding sounds when wearing the audio device. The dry source is an audio signal recorded in a space without reverberation (an anechoic chamber). The amplifier 205 amplifies the dry source and outputs the sound from the speaker 201. The sound reinforcement device 206 processes and amplifies the dry source and emits the sound from the open-air earphones 202.
[0013] Fig. 3 is a functional block diagram showing a detailed configuration example of the sound reinforcement device 206. The sound reinforcement device 206 includes an actual sound propagation estimation unit 207, a sound reinforcement unit 208, and a storage unit 209. Fig. 4 is a flowchart explaining an example of the operation of the sound reinforcement device 206. The following description will be given with reference to Figs. 2, 3, and 4.
[0014] The real sound propagation estimation unit 207 acquires sound source position information and presentation point position information (step S402). As shown in FIG. 2, the position of the speaker 201 is the sound source position, and the position of the open-air earphones 202R / L is the presentation point position. In the first embodiment, the sound source position and the presentation point position are assumed to be fixed. The real sound propagation estimation unit 207 estimates the state of real space propagation of the sound source signal from the sound source position to the presentation point position (step S403). The estimated information is, for example, spatial transfer characteristics, and includes information on spatiotemporal transfer characteristics such as distance propagation delay, distance attenuation, and reverberation characteristics from the sound source position to the presentation point position. The spatial transfer characteristics are estimated for each of the distance between the presentation point 202R and the sound source position and the distance between the presentation point 202L and the sound source position. FIG. 5 (top) shows an example of spatial transfer characteristics in the form of an impulse response. The horizontal axis represents time, and the vertical axis represents the amplitude of the response. The leftmost axis represents direct sound, and the second and subsequent axes represent reflected sound and reverberation sound.
[0015] The sound enhancement unit 208 estimates the "Upper-limit of law of the first wavefront" (ULFW) from the spatial transfer characteristics (step S404). Figure 5 shows the estimated ULFW. The "upper limit" is the relative sound pressure level of the reflected and reverberant sound at which the direct sound from the sound image and the reflected and reverberant sound are perceived as blended together. The relative sound pressure level remains constant up to 20 ms away from the direct sound, but after 20 ms, the relative sound pressure level to the direct sound decreases by approximately 4 dB for every 10 ms of delay time. For details on ULFW estimation, see Reference 2.
[0016] Reference 2: M. Morimoto et al., "A chart of %-split of sound image", Journal of the Acoustical Society of Japan (E), Vol.11, 3, pp.157-160, 1990.
[0017] The audio enrichment unit 208 retrieves audio enrichment parameters from the storage unit (step S401). The audio enrichment parameters include an ASW parameter that controls the apparent audio source width (ASW), which corresponds to the sense of spaciousness, and an LEV parameter that controls the listener's sense of being surrounded by sound images other than the apparent audio source (LEV). The audio enrichment unit 208 determines the presentation sound transfer characteristics based on the spatial transfer characteristics, the ULFW, and the audio enrichment parameters (step S405). Figure 5 (bottom) shows an example of the presentation sound transfer characteristics in the form of an impulse response. The ULFW is calculated from the spatial transfer characteristics. Response components (ASW components) whose amplitude is equal to or less than the ULFW contribute to the "sense of spaciousness," while response components (LEV components) whose amplitude exceeds the ULFW contribute to the "sense of envelopment" (Reference 3: Kazuhiro Iida, Masayuki Morimoto, "Spatial Acoustics," Corona Publishing, pp. 34-35, 2010).
[0018] The time position of the response in the presentation sound transfer characteristics is delayed from the position (time) of the direct sound in the spatial transfer characteristics, inducing a precedence effect. The ASW parameter and LEV parameter have pairs of delay time from the direct sound and amplitude values, and specify what kind of subsequent sound to add. <Adding ASW Components> The ASW parameter takes values from 1 to 5, for example. As the value increases, subsequent sounds to enhance the ASW are added to the presentation sound transfer characteristics. For example, in the case of ASW, the time is specified in 10 ms intervals from a delay time of 10 ms from the direct sound, and the amplitude is the ULFW value. When ASW is "1", a response with an amplitude of ULFW (in this case, a value such that the relative sound pressure level with respect to the direct sound is 0 dB) is placed at the 10 ms delay. <Adding LEV Components> The LEV parameter takes values from 1 to 5, for example. As the value increases, subsequent sounds to enhance the LEV are added to the presentation sound transfer characteristics. For example, in the case of LEV, the time is specified in 10 ms intervals from a delay time of 60 ms from the direct sound, and the amplitude is specified to be a value exceeding the ULFW. When LEV is "1", a response is placed at a delay time of 60 ms with an amplitude that exceeds the ULFW (in this case, an amplitude where the relative sound pressure level to the direct sound is -16 dB or more).
[0019] When the sound source position and the presentation point position remain unchanged, the spatial transfer characteristics and the presentation sound transfer characteristics also remain unchanged.
[0020] The sound enhancement unit 208 acquires a sound source signal from a sound source (dry source) (step S406), convolves the presentation sound transfer characteristics with the sound source signal, and generates a presentation sound (step S407). The sound enhancement device 206 imparts a predetermined delay to the output of the sound enhancement unit 208 based on the distance propagation delay calculated by the actual propagation estimation unit, and emits the presentation sound from the open-air earphones in time with the sound source signal emitted from the speaker reaching the ear (step S408). If the spatial transfer characteristics and the presentation sound transfer characteristics are unchanged, the sound enhancement device 208 repeats steps S406 to S408 to generate and emit the presentation sound.
[0021] The above is the description of the first embodiment.
[0022] [Modification 1] In the first embodiment, the sound source position and the presentation point position are described as being fixed. The sound source position and the presentation point position (the position of the listener) may be variable. FIG. 6 is a flowchart illustrating an example of the operation of the sound enhancement device 206 when the sound source position and the presentation point position are variable. The content of the processing of each step is the same as in the first embodiment, but differs from the first embodiment in that the sound enhancement device 206 repeatedly executes steps S402 to S408 to update the sound source position and the presentation sound position using a position identification means (not shown).
[0023] Second Embodiment In the first embodiment, a case where a dry source can be obtained as a sound source signal is described. In the second embodiment, a case where a dry source cannot be obtained is described. In this case, speaker sound is observed with an external microphone, and a target signal from which noise, reflection, and reverberation have been removed or reduced is extracted and used as the sound source signal.
[0024] 7 is a functional block diagram of the sound reinforcement system according to the second embodiment. This system differs from the sound reinforcement system 2 according to the first embodiment in that the sound source is a non-dry source 701 and the sound reinforcement device input is an external microphone 702.
[0025] Fig. 8 is a functional block diagram showing a detailed configuration example of a sound enhancement device 703 according to the second embodiment. It differs from the sound enhancement device 206 according to the first embodiment in that it includes a target signal extraction unit. Fig. 9 is a flowchart explaining an example of the operation of the sound enhancement device 703. The following description will be given with reference to Figs. 7, 8, and 9. It is assumed that the sound source position and the presentation point position are fixed.
[0026] [Extraction of Target Signal] The sound reinforcement device 703 acquires an observed signal observed by the external microphone 702 (step S901). The target signal extraction unit 704 extracts a target signal from the observed signal, with noise, reflection, and reverberation removed or reduced (step S902). Existing methods can be used for extraction. Methods using signal processing (e.g., beamformer) and deep learning (e.g., DNN speech enhancement) can be used, but lightweight processing is desirable.
[0027] [Determining Presentation Sound Transfer Characteristics] Steps S401 and S402 are the same as those in the first embodiment. The sound reinforcement device 703 acquires an observation signal observed by the external microphone 702 (step S901). The real sound propagation estimation unit 207 estimates the state of real-space propagation of the sound source signal from the sound source position to the presentation point position based on the sound source position information, the presentation point position information, and the observation signal (step S903). Steps S404 and S405 are the same as those in the first embodiment.
[0028] [Generation of Presentation Sound] The sound enhancement unit 208 acquires a target signal (step S904) and convolves the target signal with the presentation sound transfer characteristics to generate a presentation sound (step S407). The sound enhancement device 206 imparts a predetermined delay to the output of the sound enhancement unit 208 based on the distance propagation delay calculated by the actual propagation estimation unit, and emits the presentation sound from the open-air earphones in time with the sound (non-dry source) emitted from the speaker reaching the ear (step S408). If the spatial transfer characteristics and the presentation sound transfer characteristics remain unchanged, the sound enhancement device 208 returns to step S901 and processes the next observation signal.
[0029] The above is a description of the second embodiment. In the second embodiment, since the presentation sound is generated from the observed speaker sound, it is desirable to position the external microphone as close to the sound source as possible. Furthermore, if the sound source position and the presentation point position move (are not fixed), the sound reinforcement device 703 may update the sound source position information and the presentation point position information using a position identification means (not shown). In this case, the flow of FIG. 9 returns to step 901 and step 402 after step S408.
[0030] [Supplementary Information] <In the Case of Stereo Speakers / Dry Source> In the first and second embodiments, monaural speakers have been used for explanation. The spatial transmission characteristics and presented sound transmission characteristics when stereo speakers are used will be explained. FIG. 10 shows the arrangement of stereo speakers and open-air earphones. The speaker on the left is 201L, and the speaker on the right is 201R. The spatial transmission characteristics between 201L and 202L are C LL , the spatial transfer characteristic between 201L and 202R is C LR , the spatial transfer characteristic between 201R and 202L is C RL , the spatial transfer characteristic between 201R and 202R is C RR , and the sound reinforcement equipment is C LL and C RL are combined to form the spatial transfer characteristic L, and C RL and C RR and generate the spatial transfer characteristic R. The sound enhancement device generates a presentation sound transfer characteristic for the open-air earphone 202L based on the spatial transfer characteristic L, and generates a presentation sound transfer characteristic for the open-air earphone 202R based on the spatial transfer characteristic R.
[0031] <Stereo Speakers / Non-Dry Source> When the sound source is a non-dry source, two external microphones are used and placed near each speaker to acquire observation signals for left and right sounds.
[0032] [Program, Recording Medium] The functions realized by the components described in this specification may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (Central Processing Units), conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in a memory.
[0033] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0034] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0035] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 11 and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0036] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0037] The program may be distributed by, for example, selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to other computers via a network, thereby distributing the program.
[0038] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the program each time a program is transferred from a server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. Furthermore, the server computer may execute the process at the terminal using a so-called SaaS (Software as a Service) service, which allows users to use part of a server computer along with the program. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that dictate computer processing).
[0039] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
Claims
1. An acoustic signal control method for allowing a user to hear both an acoustic signal emitted from a speaker and an acoustic signal emitted from an earphone, wherein a speaker emits an acoustic signal and an open-air earphone emits an auxiliary signal, and an audio amplification device acquires the acoustic signal and generates an auxiliary signal from the acoustic signal, and the auxiliary signal is generated from the acoustic signal so that the user perceives the signal as spatially expanded when listening to both the acoustic signal and the auxiliary signal compared to when the user listens to only the acoustic signal.
2. An acoustic signal control method as described in claim 1, wherein the sound amplification device estimates the spatial transfer characteristics between the speaker and the open-air earphone, estimates the upper limit of the first wave front law from the spatial transfer characteristics, determines the presentation sound transfer characteristics using the upper limit of the first wave front law, and converts the acoustic signal using the presentation sound transfer characteristics to generate the auxiliary signal.
3. An acoustic signal control method according to claim 2, wherein the presentation sound transmission characteristics have a component that causes the perception of apparent sound source width (ASW) and / or a component that causes the perception of envelopment (LEV).
4. An acoustic signal control method according to claim 3, wherein the time position and amplitude of the component that causes the perception of ASW are determined based on the time position of the direct sound component of the spatial transfer characteristics and ASW control parameters.
5. An acoustic signal control method according to claim 3, wherein the time position and amplitude of the component that causes the perception of LEV are determined based on the time position of the direct sound component of the spatial transfer characteristics and an LEV control parameter.
6. An acoustic device that allows a user to listen to both an acoustic signal emitted from a speaker and an acoustic signal emitted from an earphone, the acoustic device comprising: an acoustic enhancement unit that acquires the acoustic signal and generates the auxiliary signal from the acoustic signal, where the speaker emits the acoustic signal and the open-air earphone emits an auxiliary signal, and the auxiliary signal is generated from the acoustic signal so that the user perceives the auxiliary signal as being spatially expanded when listening to the acoustic signal and the auxiliary signal, compared to when the user listens to only the acoustic signal.
7. A program for causing a computer to execute the acoustic signal control method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Sound field creating apparatus
JP2010193105A
Sound addition device and sound addition method
JP4348886B2