Acoustic signal generation device, acoustic signal generation method, and program

The acoustic signal generation device with open headphones and surrounding speakers corrects sound signals in real time to address the challenges of sound localization in VR/AR, ensuring accurate sound reproduction despite head movements.

WO2025253638A1PCT designated stage Publication Date: 2025-12-11NT T INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/020894
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing audio reproduction technologies for virtual reality and augmented reality headsets struggle to accurately localize sound images in real time due to variations in head-related transfer functions among individuals and limitations in device size and listening positions, especially when the listener's head moves.

Method used

An acoustic signal generation device using open-type headphones and speakers arranged around the listener, which includes a processing unit to correct acoustic signals based on the listener's head movements, ensuring accurate sound localization in real time.

Benefits of technology

The solution enables real-time sound reproduction that adapts to the listener's head movements, providing accurate sound localization and enhancing the immersive experience in virtual reality and augmented reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024020894_11122025_PF_FP_ABST
    Figure JP2024020894_11122025_PF_FP_ABST
Patent Text Reader

Abstract

One aspect of the present invention is an acoustic signal generation device (1) that comprises an open headphone (20) that includes a first speaker and is worn by a user, a speaker unit (30) that includes a second speaker and is positioned in the surroundings of the user, and a processing unit (10) that includes an acoustic signal acquisition unit (111) that acquires a first acoustic signal that is a first sound source as reproduced from a second position that corresponds to a first position of the first speaker and captured at a third position that corresponds to the initial position of the head of the user and a second acoustic signal that is the first sound source as reproduced from a fifth position that corresponds to a fourth position of the second speaker and captured at the third position, a first correction unit (112) that corrects the first acoustic signal on the basis of the first acoustic signal and the second acoustic signal to generate a third acoustic signal, and a second correction unit (114) that corrects the third acoustic signal and the second acoustic signal on the basis of the initial position, the current position of the head of the user, and the fourth position.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic signal generating device, acoustic signal generating method, and program

[0001] One aspect of the present invention relates to an audio signal generating device, an audio signal generating method, and a program.

[0002] In recent years, content that provides a high level of immersion using a head-mounted display (HMD), such as virtual reality (VR) or augmented reality (AR), has become widespread. To provide a content user with a more immersive experience, it is necessary to provide an experience that is comparable to an experience in real space. To achieve this, it is necessary to reproduce a certain acoustic space in a different space. Known methods for reproducing an acoustic space include, for example, a method using headphones and a method using a speaker array consisting of multiple speakers.

[0003] In the headphone-based method, for example, an HMD equipped with headphones or an HMD with built-in speakers that do not block the ears is used. In these methods using headphones or speakers placed near the head, sound images can be localized by using a head-related transfer function. A head-related transfer function is a signal containing sound propagation information, and is capable of reproducing the volume and phase difference of sound between the two ears. However, since head-related transfer functions vary from person to person, there are problems such as incorrectly determining the front or back of the sound image position or the sound image being localized inside the head.

[0004] On the other hand, speaker array techniques include surround systems using multiple speakers arranged around the listener, and wave field synthesis systems that physically reproduce a sound field. These techniques can accurately localize a sound image for the listener. However, they have known issues, such as being dependent on the size of the device and having a limited listening position (sweet spot).

[0005] Non-Patent Document 1 describes a method for correcting signals entering both ears when a speaker and a full open-air headphone are used in combination.

[0006] Ikuichiro Kinoshita et al., "Compensation for interaural differences in two-channel reproduction using both loudspeakers and full open-air headphones," Journal of the Acoustical Society of Japan, Vol. 61, No. 4, 2005, pp. 179-191

[0007] However, the method described in Non-Patent Document 1 assumes that the listener is stationary, and it has been confirmed that errors become larger when the listener's head actually moves. This makes it difficult to respond in real time. In order to combine virtual space and real space using content using an HMD, such as VR or AR, an audio reproduction technology that responds to the listener's movements in real time is required.

[0008] This invention was made with the above-mentioned circumstances in mind, and provides an acoustic signal generation technology for reproducing sound in real time in response to the listener's head movements, using open headphones worn by the listener and speakers placed around the listener.

[0009] In order to solve the above problem, an acoustic signal generating device according to one aspect of the present invention includes open-type headphones that include a first speaker and are worn by a user; a speaker unit that includes a second speaker and is arranged around the user; an acoustic signal acquisition unit that acquires a first acoustic signal in which a first sound source is reproduced from a second position corresponding to a first position of the first speaker and picked up at a third position corresponding to an initial position of the user's head, and a second acoustic signal in which the first sound source is reproduced from a fifth position corresponding to a fourth position of the second speaker and picked up at the third position; a first correction unit that generates a third acoustic signal by correcting the first acoustic signal based on the first acoustic signal and the second acoustic signal; and a second correction unit that corrects the third acoustic signal and the second acoustic signal based on the initial position, the current position of the user's head, and the fourth position.

[0010] An acoustic signal generation method according to one aspect of the present invention is an acoustic signal generation method executed by an acoustic signal generation device including open-type headphones worn by a user, the open-type headphones including a first speaker, a speaker unit including a second speaker and arranged around the user, and a processing unit, the method comprising the steps of: acquiring a first acoustic signal in which a first sound source is reproduced from a second position corresponding to a first position of the first speaker and picked up at a third position corresponding to an initial position of the user's head; and acquiring a second acoustic signal in which the first sound source is reproduced from a fifth position corresponding to a fourth position of the second speaker and picked up at the third position; generating a third acoustic signal by correcting the first acoustic signal based on the first acoustic signal and the second acoustic signal; and correcting the third acoustic signal and the second acoustic signal based on the initial position, the current position of the user's head, and the fourth position.

[0011] According to one aspect of the present invention, an acoustic signal can be generated to reproduce sound in real time in response to the listener's head movements using open headphones worn by the listener and speakers placed around the listener.

[0012] FIG. 1 is a schematic diagram showing an example of the configuration of an acoustic signal generation device according to an embodiment. FIG. 2 is a block diagram showing an example of the hardware configuration of an acoustic signal generation device according to an embodiment. FIG. 3 is a block diagram showing an example of the functional configuration of an acoustic signal generation device according to an embodiment. FIG. 4 is a diagram showing an example of a method for collecting an acoustic signal using a dummy head. FIG. 5 is a diagram showing another example of a method for collecting an acoustic signal using a microphone. FIG. 6 is a diagram showing another example of a method for collecting an acoustic signal using a dummy head. FIG. 7 is a diagram showing a situation when the head of a listener moves in the acoustic signal generation device according to an embodiment. FIG. 8 is a flowchart showing an example of the operation of the acoustic signal generation device according to an embodiment.

[0013] Hereinafter, embodiments will be described with reference to the drawings. In the following description, components having substantially the same functions and configurations will be designated by the same reference numerals. When particularly distinguishing between elements having similar configurations, different letters or numbers may be added to the end of the same reference numerals.

[0014] 1. Embodiments An acoustic signal generating device according to an embodiment will be described below. In the following, an acoustic signal generating device including open headphones worn by a listener (user) and speakers arranged around the listener will be described as an example of the acoustic signal generating device.

[0015] 1.1 Configuration 1.1.1 Configuration of the Acoustic Signal Generation Device The configuration of the acoustic signal generation device according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a schematic diagram showing an example of the configuration of the acoustic signal generation device.

[0016] The acoustic signal generating device 1 is a device that generates at least two corrected acoustic signals by correcting at least two acoustic signals acquired from the outside. Hereinafter, the at least two acoustic signals acquired from the outside by the acoustic signal generating device 1 will be referred to as "acoustic signal P HL,HR (t)" and "acoustic signal P Li (t)" (i is an integer equal to or greater than 1). t represents time. HL,HR (t) is the acoustic signal P HL (t) and the acoustic signal P HR (t). The acoustic signal P HL,HR The acoustic signal obtained by correcting (t) is called "acoustic signal P~ HL,HR (t)". ​​The acoustic signal P~ HL,HR (t) is the acoustic signal P HL (t) and the acoustic signal P HR (t). The acoustic signal P Li The acoustic signal obtained by correcting (t) is called "acoustic signal P~ Li In the example of FIG. 1, the acoustic signal P Li (t) is the acoustic signal P L1 (t), P L2 (t), P L3 (t), .... The acoustic signal P Li (t) is the acoustic signal P L1 (t), P~ L2 (t), P~ L3 (t), .... The listener receives the acoustic signal P~ HL,HR (t) and P Li (t) is inserted.

[0017] As shown in FIG. 1 , the acoustic signal generating device 1 includes a processing unit 10 , headphones 20 , and a speaker unit 30 .

[0018] The processing unit 10 is, for example, a computer. The processing unit 10 is configured to be able to communicate with the headphones 20, the speaker unit 30, and other external devices. For example, the processing unit 10 is connected to the headphones 20, the speaker unit 30, and other external devices wirelessly.

[0019] The processing unit 10 receives an audio signal P HL,HR (t) and P Li The processing unit 10 acquires the initial head position PS0 of the listener U from the headphones 20. T and current position PS C (t) = [x C , y C , z C ] T The processing unit 10 acquires the speaker SP from the speaker unit 30. Li Position PS Li The processing unit 10 acquires the acoustic signal P HL,HR (t) and P Li (t), initial position PS0 and current position PS C (t), and speaker SP Li Position PS Li Based on this, the acoustic signal P HL,HR (t) and P Li The processing unit 10 corrects the acoustic signal P HL The acoustic signal P obtained by correcting (t) HL (t) to the speaker SP of the headphone 20 HL The processing unit 10 transmits the acoustic signal P HR The acoustic signal P obtained by correcting (t) HR (t) to the speaker SP of the headphone 20 HR The processing unit 10 transmits the acoustic signal P Li The acoustic signal P obtained by correcting (t) Li (t) is the speaker SP of the speaker unit 30 Li Send to.

[0020] The headphones 20 are, for example, open type headphones. The headphones 20 include a mounting part AP, a speaker unit 24, and a sensor 25. The mounting part AP is a main body part that is mounted on the head of a listener U. The speaker unit 24 includes a speaker SP HL and SP HR Includes speaker SP HL is the speaker for the left ear. HR is the speaker for the right ear. HL and SP HR is attached to the mounting unit AP so as not to block the ears of the listener U. The sensor 25 detects the position of the head of the listener U. The sensor 25 is disposed on the mounting unit AP, for example.

[0021] The headphones 20 detect the initial position PS0 and the current position PS of the head of the listener U detected by the sensor 25. C (t) is transmitted to the processing unit 10. The speaker SP of the headphone 20 HL and SP HR Each of the signals P is output from the processing unit 10 as an audio signal P HL,HR (t) and the acoustic signal P HL,HR (t) is output (played back).

[0022] The speaker unit 30 includes one or more speakers SP Li Includes speaker SP Li is the acoustic signal P~ Li (t). That is, the speaker SP L1 is the acoustic signal P~ L1 (t) Speaker SP L2 is the acoustic signal P~ L2 (t) Speaker SP L3 is the acoustic signal P~ L3 In other words, the speaker unit 30 outputs the acoustic signal P Li (t) Li One or more speakers SP Li The speaker unit 30 is arranged, for example, around the listener U. L1 For example, the position PS L1 = [x L1 , yL1 , z L1 ] T Speaker SP L2 For example, the position PS L2 = [x L2 , y L2 , z L2 ] T Speaker SP L3 For example, the position PS L3 = [x L3 , y L3 , z L3 ] T will be placed in.

[0023] Speaker SP of speaker unit 30 Li is output from the processing unit 10 as an acoustic signal P Li (t) and the acoustic signal P Li (t) is output (played back).

[0024] The listener U receives the acoustic signal P~ output from the headphones 20. HL,HR (t), and the acoustic signal P~ output from the speaker unit 30 Li Listen to (t).

[0025] 1.1.2 Hardware Configuration of Acoustic Signal Generation Device The hardware configuration of the acoustic signal generation device 1 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the hardware configuration of the acoustic signal generation device 1.

[0026] (Processing Unit 10) The following describes the hardware configuration of the processing unit 10. As shown in FIG.

[0027] The processor 11 is, for example, a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), etc. The processor 11 executes, for example, an acoustic signal acquisition process, a first correction process, a position information acquisition process, and a second correction process. The acoustic signal acquisition process is performed by acquiring an acoustic signal P HL,HR(t) and P Li The first correction process is a process for obtaining the acoustic signal P HL,HR (t) and P Li (t) is corrected to produce an acoustic signal P^ HL,HR (t) and P^ Li (t) is generated. HL,HR (t) is the acoustic signal P^ HL (t) and the acoustic signal P^ HR (t). HL (t) is the acoustic signal P HL The acoustic signal P^(t) is obtained by correcting the acoustic signal P^(t). HR (t) is the acoustic signal P HR The position information acquisition process is carried out by correcting the initial position PS0 and the current position PS1 of the head of the listener U. C (t), and the speaker SP of the speaker unit 30 Li Position PS Li The second correction process is a process for obtaining the acoustic signal P^ HL,HR (t) and P^ Li The corrected acoustic signal P HL,HR (t) and P Li This is a process for generating the acoustic signal P~(t). HL (t) is the acoustic signal P^ HL The acoustic signal P is obtained by correcting (t). HR (t) is the acoustic signal P^ HR The acoustic signal is obtained by correcting (t). The acoustic signal acquisition process, the first correction process, the position information acquisition process, and the second correction process will be described in detail later.

[0028] The memory 12 includes, for example, a read-only memory (ROM) and a random access memory (RAM). The ROM stores, for example, programs for causing the processor 11 to execute the acoustic signal acquisition process, the first correction process, the position information acquisition process, and the second correction process. The RAM is used as a working area for the processor 11. The RAM temporarily stores, for example, the above-mentioned programs executed by the processor 11, as well as data obtained when the acoustic signal acquisition process, the first correction process, the position information acquisition process, and the second correction process are executed.

[0029] The communication interface 13 controls communication between the processing unit 10 and the headphones 20, the speaker unit 30, and other external devices. HL,HR (t) and P Li The communication interface 13 receives the acoustic signal P HL,HR (t) and P Li (t) to the processor 11. The communication interface 13 receives the initial position PS0 and the current position PS1 of the head of the listener U from the headphones 20. C The communication interface 13 receives the initial position PS0 and the current position PS C The communication interface 13 transmits the signal (t) to the processor 11. Li From speaker SP Li Position PS Li The communication interface 13 receives the position PS Li The communication interface 13 transmits the acoustic signal P~ to the processor 11. HL,HR (t) and P Li The communication interface 13 receives the acoustic signal P HL,HR (t) to the headphones 20. The communication interface 13 transmits the acoustic signal P Li (t) is transmitted to the speaker unit 30.

[0030] (Headphones 20) The following describes the hardware configuration of the headphones 20. As shown in FIG.

[0031] The processor 21 is, for example, a CPU, an MPU, a GPU, an FPGA, etc. The processor 21 executes, for example, a position information acquisition process. The position information acquisition process acquires the initial position PS0 and the current position PS1 of the head of the listener U. C The location information acquisition process will be described in detail later.

[0032] The memory 22 includes, for example, a ROM and a RAM. The ROM stores, for example, a program for causing the processor 21 to execute the location information acquisition process. The RAM is used as a work area for the processor 21. The RAM temporarily stores, for example, the above-mentioned program executed by the processor 21 and data during execution of the location information acquisition process.

[0033] The communication interface 23 controls communication between the headphones 20 and the processing unit 10. The communication interface 23 receives the acoustic signal P~ from the processing unit 10. HL,HR The communication interface 23 receives the acoustic signal P HL,HR (t) to the processor 21. The communication interface 23 receives the initial position PS0 and the current position PS of the head of the listener U from the processor 21. C The communication interface 23 receives the initial position PS0 and the current position PS C (t) is sent to the processing unit 10.

[0034] The sensor 25 is, for example, a head tracker. The sensor 25 includes, for example, an acceleration sensor and a gyro sensor.

[0035] The headphones 20 may include a display unit (for example, a display).

[0036] (Speaker unit 30) The speaker unit 30 includes one or more speakers SP. Li Includes speaker SP Liincludes a GNSS (Global Navigation Satellite System) device (not shown). Li is the speaker SP acquired by the GNSS device Li Position PS Li is transmitted to the processing unit 10.

[0037] 1.1.3 Functional Configuration of the Acoustic Signal Generating Device The functional configuration of the acoustic signal generating device 1 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the functional configuration of the acoustic signal generating device 1. Fig. 3 also shows the communication interface 13 of the processing unit 10, the communication interface 23 of the headphones 20, the speaker unit 24, the sensor 25, and the speaker unit 30.

[0038] (Processing Unit 10) The following describes the functional configuration of the processing unit 10. As shown in Fig. 3, the processing unit 10 includes, as functional blocks, an acoustic signal acquisition unit 111, a first correction unit 112, a position information acquisition unit 113, and a second correction unit 114. The processor 11 of the processing unit 10 functions as the acoustic signal acquisition unit 111, the first correction unit 112, the position information acquisition unit 113, and the second correction unit 114.

[0039] (Acoustic Signal Acquisition Unit 111) The acoustic signal acquisition unit 111 executes an acoustic signal acquisition process. Specifically, the acoustic signal acquisition unit 111 receives an acoustic signal P HL,HR (t) and P Li The acoustic signal acquisition unit 111 acquires the acoustic signal P HL,HR (t) and P Li (t) is transmitted to the first correction unit 112.

[0040] (First Correction Unit 112) The first correction unit 112 executes a first correction process. The first correction process is a process of correcting an acoustic signal at the initial position PS0 of the head of the listener U. The first correction process will be described in detail below.

[0041] The first correction unit 112 receives the acoustic signal P HL,HR (t) and P Li (t) is received.

[0042] The first correction unit 112 corrects the acoustic signal P HL,HR (t) and P Li (t), the acoustic signal P HL,HR (t) and P Li (t) are corrected to produce the acoustic signal P^ HL,HR (t) and P^ Li Specifically, the first correction unit 112 generates the acoustic signal P HL,HR (t) and P Li (t), the acoustic signal P HL,HR (t) and the acoustic signal P Li A correction factor t that corrects the time delay between (t) L,R , and the acoustic signal P HL,HR (t) and the acoustic signal P Li (t) and the correction coefficient c L,R The first correction unit 112 calculates the correction coefficient t L,R and c L,R Based on this, the acoustic signal P̂ HL,HR (t) and P^ Li (t) is generated.

[0043] First, the speaker unit 30 is a single speaker SP Li (SP L1 ) Correction coefficient t L,R and c L,R The calculation of the speaker unit 30 is explained below. Li The acoustic signal generating device 1 in the case where the acoustic signal generating device 1 includes, for example, the acoustic signal generating device 1 shown in FIG. L2 , SP L3 , ... have been eliminated.

[0044] Speaker SP of headphone 20 HL and SP HR , and the speaker SP of the speaker unit 30 Li When the same sound source is played simultaneously from the speaker SP HL and SP HR Acoustic signal P from HL,HR Since the sound signal P (t) reaches the listener U first, the sound image may be localized at the position of the headphones 20 due to the precedence effect. HL,HR (t) and the acoustic signal PLi (t) and the time delay between them. L,R As a method for determining the above, for example, the following method using a dummy head DH can be considered.

[0045] 4 is a diagram showing an example of a method for collecting an acoustic signal using a dummy head DH. As shown in FIG. 4, for example, when the initial position PS0 of the head of a listener U is [0,0,0], T Position PS0' corresponding to [x0, y0, z0] T The dummy head DH is a recording device that mimics the shape of a human head. The dummy head DH is connected to the microphone MP L and MP R Includes microphone MP L is provided at a position corresponding to the left ear of the dummy head DH. R is provided at a position corresponding to the right ear of the dummy head DH.

[0046] For example, in a space where a dummy head DH is arranged as shown in FIG. Li '=[x Li ',y Li ', z Li '] T , position PS H1 '=[x H1 ',y H1 ', z H1 '] T , and position PS H2 '=[x H2 ',y H2 ', z H2 '] T The same sound source SS is played back simultaneously from the position PS Li ' denotes the speaker SP of the speaker unit 30. Li Position PS Li = [x Li , y Li , z Li ] T The position corresponds to the position PS H1 ' is the speaker SP of the headphone 20 HL Position PS H1 = [x H1 , y H1 , zH1 ] T The position corresponds to the position PS H2 ' is the speaker SP of the headphone 20 HR Position PS H2 = [x H2 , y H2 , z H2 ] T That is, the position of the speaker SP relative to the initial position PS0. HL , SP HR , and SP Li The positional relationship of the position PS H1 ', position PS H2 ', and position PS Li '. When collecting sound with the dummy head DH and when the listener U listens, the speaker SP HL , SP HR , and SP Li Although the absolute position of the microphone MP1 does not change, in this specification, for ease of understanding, the positions where the sound is picked up by the dummy head DH or the microphone MP1 described later are indicated with "'". When the sound source SS is played back simultaneously, the dummy head DH placed at the position PS0' (more precisely, the microphone MP1 at the position PS0') L and MP R at the position of the microphone MP L and MP R The acoustic signal is picked up by the speaker SP. HL and SP HR It is played from the microphone MP L and MP R The acoustic signal collected by the HL,HR (t). Speaker SP Li It is played from the microphone MP L and MP R The acoustic signal collected by the Li Let (t) be the time.

[0047] The collected acoustic signal P HL,HR (t), and P Li(t) is acquired by the acoustic signal acquisition unit 111 via the communication interface 13 and transmitted from the acoustic signal acquisition unit 111 to the first correction unit 112 .

[0048] The first correction unit 112 corrects the acoustic signal P HL,HR (t) and P Li The correction coefficient t is used to maximize the cross-correlation of (t). L,R For example, the first correction unit 112 determines the correction coefficient t that maximizes the cross-correlation function R using the following equation (1): L,R As a result, the acoustic signal P HL,HR (t) and the acoustic signal P Li (t) can be corrected.

[0049] Also, the speaker SP of the headphone 20 HL and SP HR and the speaker SP of the speaker unit 30 Li When the sound signal P is reproduced using different devices, the volume differs for each device, and the sound image is localized to the device with the larger volume. HL,HR (t) and the acoustic signal P Li The first correction unit 112 corrects the amplitude difference between the acoustic signal P HL,HR (t) and P Li The correction coefficient c is set so that the maximum amplitude value of (t) coincides. L,R For example, the first correction unit 112 determines the acoustic signal P HL,HR (t) and P Li The correction coefficient c that makes the maximum amplitude value of (t) coincide L,R As a result, the acoustic signal P HL,HR (t) and the acoustic signal P Li (t) can be corrected. Here, max||P Li || is the acoustic signal P Li represents the maximum amplitude value of (t). max∥P HL,HR || is the acoustic signal P HL,HR represents the maximum amplitude value of (t).

[0050] The first correction unit 112 calculates the correction coefficient t L,R and c L,R Based on this, the acoustic signal P HL,HR (t) is corrected to produce an acoustic signal P^ HL,HR (t) to generate the corrected acoustic signal P̂ HL,HR (t) is expressed as the following equation (3).

[0051] The speaker unit 30 is a single speaker SP Li If the acoustic signal P Li (t) is the acoustic signal P^ Li That is, the acoustic signal P^ (t) generated (output) by the first correction unit 112 is Li (t) is the acoustic signal P Li It is the same as (t).

[0052] Next, the speaker unit 30 is composed of I speakers SP (I is an integer of 2 or more). Li Correction coefficient t when (1≦i≦I) is included L,R and c L,R The calculation of is explained below.

[0053] A plurality of speakers SP are placed around the listener U. Li When the speakers SP are arranged, the acoustic signal received by the listener U is Li The acoustic signal P Li However, the microphone MP of the dummy head DH needs to be corrected. L and MP R Therefore, the first correction unit 112 corrects the sound pressure and phase difference of the I speakers SP so that the sound pressure and phase difference are the same at the initial position PS0 of the head of the listener U. Li The acoustic signal output from the input terminal is corrected in the same manner as in equations (1) and (2).

[0054] 5 is a diagram showing another example of a method for collecting an acoustic signal using a microphone MP1. As shown in FIG. 5, for example, when the initial position PS0 of the head of a listener U is [0,0,0], T Position PS0' corresponding to [x0, y0, z0] Tis determined, and microphone MP1 is placed at position PS0'.

[0055] For example, in a space where a microphone MP1 is arranged as shown in FIG. 5, each speaker SP of the speaker unit 30 Li (SP L1 , SP L2 , SP L3 , ...) Li '(PS L1 ', P.S. L2 ', P.S. L3 The same sound source SS is simultaneously reproduced from the microphones MP1, MP2, and MP3. Si (t) (P S1 (t), P S2 (t), P S3 (t), ...). The first correction unit 112 corrects the collected sound signal P Si (t) (P S1 (t), P S2 (t), P S3 (t), ...), P Si (t) is corrected to produce an acoustic signal P^ Si (t), i.e., the acoustic signal P Li (t) is corrected to produce an acoustic signal P^ Li Specifically, the first correction unit 112 generates the acoustic signal P Si (t) to P Max (t), and for the time delay, similarly to equation (1), the correction coefficient t that maximizes the cross-correlation function R is calculated using the following equation (4): Si (t S1 , t S2 , t S3 , ...) are calculated. Si The time delay between (t) can be corrected.

[0056] Similarly to equation (2), the first correction unit 112 also calculates the amplitude difference by using the following equation (5) to obtain the acoustic signal P Si (t) and P Max The correction coefficient c that makes the maximum amplitude value of (t) coincide Si (c S1 , c S2 , cS3 , ...) are calculated. Si The amplitude difference between (t) can be corrected. Here, max||P Max || is the acoustic signal P Max represents the maximum amplitude value of (t). max∥P Si || is the acoustic signal P Si represents the maximum amplitude value of (t).

[0057] The first correction unit 112 calculates the correction coefficient t Si and c Si Based on this, the acoustic signal P Li (t) is corrected to produce an acoustic signal P^ Li (t) is generated. Max Speaker SP from which (t) is played Li The speaker SP is located closer to the initial position PS0 than the speaker SP Li The picked-up sound signal P Si (t) is corrected. The corrected acoustic signal P^ Li (t) is expressed as the following equation (6).

[0058] 6 is a diagram showing another example of a method for collecting an acoustic signal using a dummy head DH. The dummy head DH is disposed in the same manner as in FIG.

[0059] acoustic signal P^ Li After generating (t), the speaker SP of the speaker unit 30 is placed in the space where the dummy head DH is placed as shown in FIG. Li (SP L1 , SP L2 , SP L3 , ...) Li '(PS L1 ', P.S. L2 ', P.S. L3 ', ...) Li (t) are simultaneously played back. HL Position PS H1 Position PS corresponding to H1 ', and the speaker SP of the headphone 20 HRPosition PS H2 Position PS corresponding to H2 At this time, the first correction unit 112 simultaneously reproduces the same sound source SS from the microphone MP L and MP R By using equations (1) and (2) for an acoustic signal picked up by HL,HR That is, when the speaker unit 30 includes I units (I is an integer equal to or greater than 2), the first correction unit 112 generates the acoustic signal P HL,HR (t) and the acoustic signal P Li In the correction based on (t), the correction coefficient t Si and c Si The acoustic signal P̂ Li (t), and the acoustic signal P HL,HR (t), the acoustic signal P HL,HR (t) is corrected to produce an acoustic signal P^ HL,HR (t) is generated. HL,HR (t) and a plurality of acoustic signals P Si (t) (acoustic signal P Li (t)) can be corrected for time delay and amplitude differences between them.

[0060] (Position Information Acquisition Unit 113) The position information acquisition unit 113 executes a position information acquisition process. Specifically, the position information acquisition unit 113 receives the initial position PS0 and the current position PS1 of the head of the listener U from the position information acquisition unit 211 of the headphones 20 via the communication interface 13. C The position information acquisition unit 113 acquires the initial position PS0 and the current position PS C The position information acquisition unit 113 transmits the position information (t) of the speaker SP of the speaker unit 30 via the communication interface 13. Li Position PS Li The position information acquisition unit 113 acquires the position information of the speaker SP Li Position PS Li to the second correction unit 114.

[0061] (Second Correction Unit 114) The second correction unit 114 executes the second correction process. The second correction process is a process for correcting the acoustic signal during head movement of the listener U. The second correction process will be described in detail below.

[0062] FIG. 7 shows the head of the listener U moving from the initial position PS0 to the current position PS C (t) = [x C , y C , z C ] T This is a diagram showing the situation when the camera moves to the left.

[0063] The second correction unit 114 receives the acoustic signal P^ from the first correction unit 112. HL,HR (t) and P^ Li The second correction unit 114 receives the initial position PS0 and the current position PS1 of the head of the listener U from the position information acquisition unit 113. C (t), and speaker SP Li Position PS Li Receive.

[0064] The second correction unit 114 calculates the initial position PS0 and the current position PS1 of the head of the listener U. C (t), and speaker SP Li Position PS Li Based on this, the acoustic signal P̂ HL,HR (t) and P^ Li (t) are corrected to produce the acoustic signal P ~ HL,HR (t) and P ~ Li Specifically, the second correction unit 114 generates a value (t) based on the initial position PS0 and the current position PS C (t), and position PS Li Based on this, the amplitude coefficient A of the acoustic signal Li (t), time difference TD i (t), and the time difference TD H (t) is calculated. Li (t) is the acoustic signal P ~ Li (t) (the reproduced audio signal) and the audio signal P^ Li (t) is a correction coefficient for correcting the amplitude difference between the time difference TD i (t) is the acoustic signal P ~Li (t) and the acoustic signal P^ Li (t) is a correction coefficient for correcting the time difference between the H (t) is the acoustic signal P ~ HL,HR (t) (the reproduced audio signal) and the audio signal P^ HL,HR The second correction unit 114 corrects the time difference between the amplitude coefficient A Li (t) and time difference TD i (t), the acoustic signal P ~ Li The second correction unit 114 generates the time difference TD H (t), the acoustic signal P ~ HL,HR (t) is generated.

[0065] The second correction unit 114 calculates the amplitude coefficient A based on the assumption that the sound pressure is inversely proportional to the distance from the sound source. Li (t) is calculated by the following formula (7). Here, d i (t) is the speaker SP Li represents the distance from the initial position PS0 to the i (t) = ||PS Li -PS0||. d Hi (t) is the speaker SP Li From current location PS C represents the distance to (t), and d Hi (t) = ||PS Li -PS C (t)||.

[0066] The second correction unit 114 corrects the acoustic signal P ~ Li Time difference TD of (t) i (t) is calculated using the following equation (8). Here, max(d Hi (t)) is the distance d Hi (t) represents the maximum value of the frequency component, c represents the speed of sound, and fs represents the sampling frequency.

[0067] The second correction unit 114 also corrects the acoustic signal P ~ HL,HRTime difference TD of (t) H (t) is calculated by the following equation (9). Here, d imax is max(d Hi The speaker SP selected in (t) Li represents the distance from the initial position PS0 to the

[0068] The second correction unit 114 calculates the time difference TD H (t), the acoustic signal P^ HL,HR (t) is corrected to produce an acoustic signal P ~ HL,HR (t) to generate the corrected acoustic signal P ~ HL,HR (t) is expressed as the following equation (10).

[0069] The second correction unit 114 calculates the amplitude coefficient A Li (t) and time difference TD i (t), the acoustic signal P^ Li (t) is corrected to produce an acoustic signal P ~ Li (t) to generate the corrected acoustic signal P ~ Li (t) is expressed as the following equation (11).

[0070] The second correction unit 114 corrects the acoustic signal P ~ HL,HR (t) to the headphones 20 via the communication interface 13. The second correction unit 114 corrects the acoustic signal P ~ Li (t) is transmitted to the speaker unit 30 via the communication interface 13 .

[0071] (Headphones 20) The following describes the functional configuration of the headphones 20. As shown in Fig. 3, the headphones 20 include, as a functional block, a position information acquisition unit 211. The processor 21 of the headphones 20 functions as the position information acquisition unit 211.

[0072] (Position Information Acquisition Unit 211) The position information acquisition unit 211 executes a position information acquisition process. Specifically, the position information acquisition unit 211 acquires the initial position PS0 and the current position PS1 of the head of the listener U from the sensor 25. C The position information acquisition unit 211 transmits the initial position PS0 and the current position PS1 to the position information acquisition unit 113 of the processing unit 10 via the communication interface 23. C (t) is transmitted.

[0073] 1.2 Operation The following describes the operation of the acoustic signal generation device 1. Fig. 8 is a flowchart showing an example of the operation of the acoustic signal generation device 1.

[0074] First, the acoustic signal acquisition unit 111 of the processing unit 10 receives an acoustic signal P HL,HR (t) and P Li (t) is acquired (S101). HL,HR (t) and P Li (t) is an acoustic signal acquired by, for example, the sound collection method shown in FIGS.

[0075] Next, the first correction unit 112 of the processing unit 10 executes the first correction process described above (S102). HL,HR (t) is corrected to produce an acoustic signal P^ HL,HR (t), and the acoustic signal P Li (t) is corrected to produce an acoustic signal P^ Li (t) is generated.

[0076] Next, the position information acquisition unit 113 of the processing unit 10 executes the above-mentioned position information acquisition process (S103). As a result, the initial position PS0 and the current position PS1 of the head of the listener U are obtained. C (t), and the speaker SP of the speaker unit 30 Li Position PS Li is obtained.

[0077] Next, the second correction unit 114 of the processing unit 10 executes the second correction process described above (S104). HL,HR The corrected acoustic signal P HL,HR (t), and the acoustic signal P̂ Li The corrected acoustic signal P Li(t) is generated.

[0078] Next, the second correction unit 114 of the processing unit 10 corrects the acoustic signal P HL,HR (t) is transmitted to the headphones 20, and the acoustic signal P Li (t) is transmitted to the speaker unit 30 (S105).

[0079] 1.3 Effects of the Present Embodiment In the acoustic signal generating device 1 according to the present embodiment, the speaker SP HL (and SP HR ), headphones 20 (open type) worn by a listener U, and a speaker SP Li The audio system includes speaker units 30 arranged around a listener U, and a processing unit 10. The processing unit 10 includes an acoustic signal acquisition unit 111, a first correction unit 112, and a second correction unit 114.

[0080] The acoustic signal acquisition unit 111 receives the acoustic signal P HL,HR (t) and P Li (t) is acquired. HL,HR (t) is the speaker SP of the headphone 20 HL (and SP HR ) position PS H1 (and P.S. H2 ) corresponding to the position PS H1 '(and P.S. H2 The sound source SS is reproduced from the position PS0′, and the sound signal is picked up at the position PS0′ corresponding to the initial position PS0 of the head of the listener U. Li (t) is the speaker SP of the speaker unit 30 Li Position PS Li Position PS corresponding to Li The sound source SS is reproduced from the position PS0′, and the sound signal is picked up at the position PS0′ corresponding to the initial position PS0 of the listener U's head.

[0081] The first correction unit 112 corrects the acoustic signal P HL,HR (t) and P Li (t), the acoustic signal P HL,HR (t) is corrected to produce an acoustic signal P^ HL,HR (t) is generated. This generates the signal from the speaker SP of the headphone 20. HL (and SP HR), and the speaker SP of the speaker unit 30 Li The acoustic signals reproduced from each of the speakers and heard by the listener U can be corrected according to the position of the speaker and the type of device.

[0082] The second correction unit 114 calculates the initial position PS0 and the current position PS1 of the head of the listener U. C (t), and speaker SP Li Position PS Li Based on this, the acoustic signal P̂ HL,HR (t) and P Li (t) is corrected. As a result, the speaker SP of the headphone 20 HL (and SP HR ), and the speaker SP of the speaker unit 30 Li The acoustic signals reproduced from each of the speakers and heard by the listener U can be corrected in real time according to the movements of the listener U's head.

[0083] Therefore, according to this embodiment, an acoustic signal can be generated to reproduce sound in real time in response to the listener's head movements using open headphones worn by the listener and speakers placed around the listener.

[0084] 2. Modifications, etc. In the above embodiment, an example is shown in which the sensor 25 is integrated with the headphones 20, but the sensor 25 may be separate from the headphones 20 as long as it can be worn on the head of the listener U.

[0085] In the above embodiment, the speaker SP Li Position PS Li is acquired by the GNSS device, but the speaker SP Li Position PS Li is the speaker SP Li Position PS Li As long as the above information can be obtained, it may be acquired by a method other than acquisition by a GNSS device.

[0086] Furthermore, in the flowcharts described in the above embodiments, the order of the processes can be changed as much as possible.

[0087] It should be noted that the present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by selecting and combining the multiple disclosed constituent elements.

[0088] REFERENCE SIGNS LIST 1 Acoustic signal generating device 10 Processing unit 11 Processor 12 Memory 13 Communication interface 20 Headphones 21 Processor 22 Memory 23 Communication interface 24 Speaker unit 25 Sensor 30 Speaker unit 111 Acoustic signal acquiring unit 112 First correction unit 113 Position information acquiring unit 114 Second correction unit 211 Position information acquiring unit

Claims

1. An acoustic signal generating device comprising: open-type headphones including a first speaker and worn by a user; speaker units including a second speaker and arranged around the user; an acoustic signal acquisition unit that acquires a first acoustic signal in which a first sound source is reproduced from a second position corresponding to a first position of the first speaker and picked up at a third position corresponding to an initial position of the user's head, and a second acoustic signal in which the first sound source is reproduced from a fifth position corresponding to a fourth position of the second speaker and picked up at the third position; a first correction unit that generates a third acoustic signal by correcting the first acoustic signal based on the first acoustic signal and the second acoustic signal; and a second correction unit that corrects the third acoustic signal and the second acoustic signal based on the initial position, the current position of the user's head, and the fourth position.

2. The acoustic signal generating device according to claim 1, wherein the correction by the first correction unit includes: calculating, based on the first acoustic signal and the second acoustic signal, a first correction coefficient that corrects a time delay between the first acoustic signal and the second acoustic signal, and a second correction coefficient that corrects an amplitude difference between the first acoustic signal and the second acoustic signal; and generating the third acoustic signal by correcting the first acoustic signal based on the first correction coefficient and the second correction coefficient.

3. The acoustic signal generating device according to claim 1, wherein the correction by the second correction unit includes: calculating a third correction coefficient for correcting the time difference between the third acoustic signal based on the initial position, the current position, and the fourth position; and correcting the third acoustic signal based on the third correction coefficient.

4. The acoustic signal generating device of claim 1, wherein the correction by the second correction unit includes: calculating a fourth correction coefficient for correcting a time difference with the second acoustic signal and a fifth correction coefficient for correcting an amplitude difference with the second acoustic signal based on the initial position, the current position, and the fourth position; and correcting the second acoustic signal based on the fourth correction coefficient and the fifth correction coefficient.

5. The acoustic signal generating device according to claim 1, wherein the speaker unit further includes a third speaker arranged at a sixth position closer to the initial position than the second speaker, the acoustic signal acquiring unit further acquires a fourth acoustic signal collected at the third position when the first sound source is reproduced from a seventh position corresponding to the sixth position, and the first correcting unit further generates a fifth acoustic signal by correcting the fourth acoustic signal based on the second acoustic signal and the fourth acoustic signal.

6. The acoustic signal generating device according to claim 5, wherein the correction based on the first acoustic signal and the second acoustic signal by the first correction unit is performed based on the first acoustic signal and the fifth acoustic signal.

7. An acoustic signal generation method executed by an acoustic signal generation device including open-type headphones that include a first speaker and are worn by a user, speaker units that include a second speaker and are arranged around the user, and a processing unit, the acoustic signal generation method comprising: acquiring a first acoustic signal in which a first sound source is reproduced from a second position corresponding to a first position of the first speaker and picked up at a third position that corresponds to an initial position of the user's head, and a second acoustic signal in which the first sound source is reproduced from a fifth position corresponding to a fourth position of the second speaker and picked up at the third position; generating a third acoustic signal by correcting the first acoustic signal based on the first acoustic signal and the second acoustic signal; and correcting the third acoustic signal and the second acoustic signal based on the initial position, the current position of the user's head, and the fourth position.

8. A program for causing a processor included in an acoustic signal generating device to execute processing by the acoustic signal generating device according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    JP2015206989A

  • Sound reproduction device and sound reproduction method

    WO2012042905A1