Signal processing system, signal processing method, and program

The signal processing system improves binaural playback quality by calculating and convolving transfer characteristics to correct acoustic signals, addressing the limitations of existing technologies and ensuring consistent playback across different users and devices.

JP7831876B2Active Publication Date: 2026-03-17KLEPSYDRA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing binaural recording and playback technologies, such as those described in Patent Document 1, lack improvements in quality and efficiency, particularly in accurately reproducing the sense of three-dimensionality and presence during playback.

Method used

A signal processing system and method that calculates the transfer characteristic between acoustic signals using measurement earphones and microphones, convolves the inverse characteristic with acquired signals to generate corrected acoustic signals, and utilizes control units to store and reproduce these signals for improved binaural playback, even across different users and devices.

Benefits of technology

Enhances the quality of binaural playback by reducing the need for real-time corrections and metadata distribution, allowing for high-quality playback regardless of device or user differences, and enabling simple, high-quality binaural recording in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007831876000009
    Figure 0007831876000009
  • Figure 0007831876000010
    Figure 0007831876000010
  • Figure 0007831876000011
    Figure 0007831876000011
Patent Text Reader

Abstract

[Problem] To provide a mechanism that is capable of further improving the quality of binaural reproduction. [Solution] A signal processing system comprising a first control unit that: calculates a transmission characteristic corresponding to the difference between a first acoustic signal, and a second acoustic signal, corresponding to the first acoustic signal, that is acquired by an acquisition unit that acquires acoustic signals and reproduced by a first reproduction unit that reproduces the acoustic signals; and generates a fourth acoustic signal by convolving a third acoustic signal acquired by the acquisition unit with an inverse characteristic to the calculated transmission characteristic.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a signal processing system, a signal processing method, and a program. [Background technology]

[0002] In recent years, binaural recording has been attracting attention. Binaural recording is a technique that records sound as it is transmitted to the eardrums of both ears. For example, microphones placed inside the ear canals of both ears are used for binaural recording. Playing back binaurally recorded sound is also called binaural playback. By playing back binaural sound using earphones or headphones, it is possible to reproduce a sense of three-dimensionality and presence as if one were actually present at the recording location.

[0003] Various technologies related to binaural recording and binaural playback have been developed. For example, Patent Document 1 below proposes a binaural recording device that uses a noise-canceling microphone provided on the outside of an earphone that is held in the ear by inserting the earpiece into the ear canal. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2009-49947 [Overview of the project] [Problems that the invention aims to solve]

[0005] However, the technology described in Patent Document 1 above is still relatively new, and there is room for improvement in various aspects.

[0006] Therefore, this disclosure has been made in view of the above-mentioned issues, and the purpose of this disclosure is to provide a mechanism that can further improve the quality of binaural playback. [Means for solving the problem]

[0007] In order to solve the above problems, according to an aspect of the present disclosure, a transmission characteristic corresponding to a difference between a first acoustic signal and a second acoustic signal corresponding to the first acoustic signal reproduced by a first reproduction unit that acquires and reproduces the acoustic signal is calculated, and an inverse characteristic of the calculated transmission characteristic is convolved with a third acoustic signal acquired by a second acquisition unit to generate a fourth acoustic signal, and a first control unit is provided.

[0008] The first control unit may store the generated fourth acoustic signal in a storage unit.

[0009] The signal processing system may further include a second control unit that causes a second reproduction unit that reproduces an acoustic signal to reproduce the fourth acoustic signal stored in the storage unit.

[0010] The first control unit may generate a fifth acoustic signal by convolving the characteristic of the first reproduction unit and the inverse characteristic of the characteristic of a second reproduction unit that reproduces an acoustic signal with the fourth acoustic signal.

[0011] The first control unit may store the generated fifth acoustic signal in a storage unit.

[0012] The signal processing system may further include a second control unit that causes the second reproduction unit to reproduce the fifth acoustic signal stored in the storage unit.

[0013] The first control unit and the second control unit may be mounted on different devices.

[0014] The first acquisition unit may be disposed near the eardrum of a first user, the first reproduction unit may be disposed in the auricle of the first user, and the second reproduction unit may be disposed in the auricle of a second user different from the first user.

[0015] The second reproduction unit may be different from the first reproduction unit.

[0016] The second acquisition unit may be positioned near the eardrum of the first user.

[0017] Furthermore, in order to solve the above problems, according to another aspect of this disclosure, a signal processing method is provided which includes: reproducing a first acoustic signal with a first reproduction unit that reproduces an acoustic signal; acquiring a second acoustic signal corresponding to the first acoustic signal reproduced by the first reproduction unit with a first acquisition unit that acquires an acoustic signal; calculating a transfer characteristic corresponding to the difference between the first acoustic signal and the second acoustic signal; acquiring a third acoustic signal with a second acquisition unit; and generating a fourth acoustic signal by convolving the inverse characteristic of the transfer characteristic into the third acoustic signal.

[0018] Furthermore, in order to solve the above problems, according to another aspect of this disclosure, a program is provided for a computer to function as a first control unit that calculates a transfer characteristic corresponding to the difference between a first acoustic signal and a second acoustic signal corresponding to the first acoustic signal, which is acquired by a first acquisition unit that acquires an acoustic signal and reproduced by a first reproduction unit that reproduces an acoustic signal, and generates a fourth acoustic signal by convolving the inverse characteristic of the calculated transfer characteristic into a third acoustic signal acquired by a second acquisition unit. [Effects of the Invention]

[0019] As explained above, this disclosure provides a mechanism that can further improve the quality of binaural playback. [Brief explanation of the drawing]

[0020] [Figure 1] This is a block diagram showing an example of the configuration of a signal processing system according to one embodiment of the present disclosure. [Figure 2] This figure illustrates the measurement of transfer characteristics according to this embodiment. [Figure 3] This is a diagram illustrating the binaural recording according to this embodiment. [Figure 4]This is a diagram illustrating the binaural playback according to this embodiment. [Figure 5] This sequence diagram shows an example of the processing flow related to the measurement of transfer characteristics performed by the signal processing system according to this embodiment. [Figure 6] This sequence diagram shows an example of the processing flow related to binaural recording and binaural playback performed by the signal processing system according to this embodiment. [Figure 7] This sequence diagram shows an example of the processing flow related to binaural recording and binaural playback performed by the signal processing system according to the first modified example. [Figure 8] This diagram schematically shows an example of the hardware configuration of a measurement earphone and microphone. [Figure 9] This figure shows another example of the configuration of a signal processing system. [Figure 10] This figure shows another example of the configuration of a signal processing system. [Modes for carrying out the invention]

[0021] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.

[0022] <1. Example Configuration> Figure 1 is a block diagram showing an example of the configuration of a signal processing system 1 according to one embodiment of the present disclosure. As shown in Figure 1, the signal processing system 1 according to this embodiment includes measuring earphones 10 (10A and 10B), microphones 20 (20A and 20B), recording processing device 30, playback processing device 40, and playback earphones 50 (50A and 50B). The signal processing system 1 has two measuring earphones 10, two microphones 20, and two playback earphones 50 for both ears.

[0023] - Measurement earphones 10 The measurement earphone 10 is an audio output device that reproduces an acoustic signal. The measurement earphone 10 converts the input acoustic signal into sound and emits it into the surrounding space. The measurement earphone 10 can be connected to the recording processing device 30 via various devices related to the reproduction of the acoustic signal, such as a DAC (Digital Analog Converter) and an amplifier. The measurement earphone 10 is used for measuring the transmission characteristics, which will be described later. The measurement earphone 10 is an example of the first playback unit in this embodiment. The first playback unit may consist of any audio output device other than the earphone, such as a speaker.

[0024] - Mike 20 Microphone 20 is an audio input device that acquires acoustic signals. Microphone 20 converts sounds in the surrounding space into acoustic signals and outputs the converted acoustic signals. Microphone 20 can be connected to the recording processing device 30 via various devices related to the acquisition of acoustic signals, such as an ADC (Analog Digital Converter) and an amplifier. Microphone 20 is used for measuring transfer characteristics and for binaural recording. Microphone 20 may be configured as any type of audio input device, such as a dynamic microphone, a MEMS (Micro Electro Mechanical Systems) microphone, a condenser microphone, or a laser microphone. In addition to microphones that apply an external DC voltage to the diaphragm, so-called electret condenser microphones that use electret elements in the diaphragm, back pole, or back chamber may also be used as condenser microphones.

[0025] Here, microphone 20 is an example of the first and second acquisition units in this embodiment. The first acquisition unit is an audio input device used for measuring transfer characteristics. The second acquisition unit is an audio input device used for binaural recording. That is, in this embodiment, the same microphone 20 is used for both measuring transfer characteristics and binaural recording.

[0026] - Recording processing device 30 The recording processing device 30 is a signal processing device that performs various processing related to the measurement of transmission characteristics and binaural recording. The recording processing device 30 can be implemented by any device such as a PC (Personal Computer) or a smartphone. As shown in Figure 1, the recording processing device 30 includes a communication unit 31, a storage unit 32, and a control unit 33.

[0027] The communication unit 31 is a communication interface that communicates with other devices via wired or wireless means. The communication unit 31 performs communication in accordance with any communication standard. Examples of communication standards include Wi-Fi (registered trademark), Bluetooth (registered trademark), or USB (Universal Serial Bus). For example, the communication unit 31 can communicate with the playback processing device 40 via the internet or the like. The communication unit 31 is also an audio interface. The communication unit 31 transmits and receives acoustic signals to and from the measurement earphone 10 or microphone 20.

[0028] The memory unit 32 stores various types of information. The memory unit 32 stores and reads data from a predetermined storage medium. An example of a predetermined storage medium is a non-volatile storage medium such as flash memory.

[0029] The control unit 33 functions as both an arithmetic processing unit and a control device, controlling the overall operation of the recording processing unit 30 according to various programs. The control unit 33 is implemented by electronic circuits such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor). The control unit 33 may also include a ROM (Read Only Memory) for storing the programs and calculation parameters used, and a RAM (Random Access Memory) for temporarily storing parameters that change as needed.

[0030] In particular, the control unit 33 performs various signal processing related to the measurement of transfer characteristics and binaural recording. The control unit 33 is an example of the first control unit in this embodiment.

[0031] -Recycling processing device 40 The playback processing device 40 is a signal processing device that performs various processing related to binaural playback. The playback processing device 40 can be implemented by any device such as a PC (Personal Computer) or a smartphone. As shown in Figure 1, the playback processing device 40 includes a communication unit 41, a storage unit 42, and a control unit 43.

[0032] The communication unit 41 is a communication interface that communicates with other devices via wired or wireless means. The communication unit 41 performs communication in accordance with any communication standard. Examples of communication standards include Wi-Fi®, Bluetooth®, or USB (Universal Serial Bus). For example, the communication unit 41 can communicate with the recording processing device 30 via the internet or the like. The communication unit 41 is also an audio interface. The communication unit 41 transmits and receives audio signals to and from the playback earphones 50.

[0033] The memory unit 42 stores various types of information. The memory unit 42 stores and reads data from a predetermined storage medium. An example of a predetermined storage medium is a non-volatile storage medium such as flash memory.

[0034] The control unit 43 functions as an arithmetic processing unit and control unit, and controls the overall operation within the playback processing unit 40 according to various programs. The control unit 43 is implemented by electronic circuits such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor). The control unit 43 may also include a ROM (Read Only Memory) for storing the programs and calculation parameters used, and a RAM (Random Access Memory) for temporarily storing parameters that change as needed.

[0035] In particular, the control unit 43 performs various signal processing related to binaural playback. The control unit 43 is an example of a second control unit in this embodiment.

[0036] -50 earphones for playback The playback earphone 50 is an audio output device that reproduces an acoustic signal. The playback earphone 50 converts the input acoustic signal into sound and emits it into the surrounding space. The playback earphone 50 may be connected to the playback processing device 40 via various devices related to the reproduction of the acoustic signal, such as a DAC (Digital Analog Converter) and an amplifier. The playback earphone 50 is used for binaural playback. The playback earphone 50 is an example of the second playback unit in this embodiment. The second playback unit may consist of any audio output device other than the earphone, such as a speaker.

[0037] <2. Technical Features> (1) Measurement of transfer characteristics The recording processing device 30 measures the transmission characteristics. The measurement of the transmission characteristics is performed with a human user wearing the measurement earphones 10 and microphone 20. The transmission characteristics here refer to the acoustic characteristics of the transmission path from the measurement earphones 10 to the microphone 20. The acoustic characteristics may be frequency characteristics.

[0038] The microphone 20 is positioned near the user's eardrum. On the other hand, the measurement earphone 10 is positioned on the user's auricle. This configuration makes it possible to measure the acoustic characteristics of the auricle, which greatly affect how sound is transmitted to the eardrum. As an example, the microphone 20 may be positioned in the external auditory canal, and the measurement earphone 10 may be positioned in the concha. The user who wears the measurement earphone 10 and microphone 20 to measure the transmission characteristics will be referred to as User A below. User A is an example of the first user in this embodiment.

[0039] The control unit 33 calculates the transfer characteristics based on the first acoustic signal and the second acoustic signal corresponding to the first acoustic signal reproduced by the measurement earphone 10. The calculated transfer characteristics correspond to the difference between the first acoustic signal and the second acoustic signal. The first acoustic signal is an acoustic signal reproduced for the measurement of the transfer characteristics. The first acoustic signal may be, for example, a so-called sweep signal in which the frequency changes stepwise from a low frequency to a high frequency. The second acoustic signal is the first acoustic signal that has been affected by the transmission path from the measurement earphone 10 to the microphone 20.

[0040] More specifically, the control unit 33 first outputs the first acoustic signal stored in the memory unit 32 to the measurement earphone 10, causing the measurement earphone 10 to reproduce the first acoustic signal. The microphone 20 acquires a second acoustic signal, which is the first acoustic signal reproduced from the measurement earphone 10 and is an acoustic signal originating from sound that arrived via the transmission path from the measurement earphone 10 to the microphone 20. Then, the control unit 33 calculates the transfer characteristics based on the first acoustic signal and the second acoustic signal. After that, the control unit 33 stores the calculated transfer characteristics in the memory unit 32.

[0041] The measurement of transfer characteristics will be explained with reference to Figure 2.

[0042] Figure 2 is a diagram illustrating the measurement of transmission characteristics according to this embodiment. As shown in Figure 2, the transmission path from the measurement earphone 10 to the microphone 20 includes the measurement earphone 10 and the auricle 90 of user A wearing the measurement earphone 10 and microphone 20. Therefore, the transmission characteristics to be measured are expressed by the following equation.

[0043]

number

[0044] Here is the G m (ω) is the transfer characteristic. H a(ω) represents the acoustic characteristics of the measurement earphone 10. In this specification, acoustic characteristics refer to, for example, amplitude-frequency characteristics, but other characteristics such as phase-frequency characteristics, phase-delay characteristics, and group-delay characteristics may also be used. A (ω) represents the acoustic characteristics of User A's auricle (90). ω is the angular frequency.

[0045] (2) Binaural recording The recording device 30 performs binaural recording. Binaural recording is performed with the user wearing the microphone 20.

[0046] In detail, microphone 20 acquires a third acoustic signal originating from the sound source to be binaurally recorded. Then, control unit 33 generates a fourth acoustic signal by correcting the acquired acoustic signal based on previously measured transfer characteristics. Specifically, control unit 33 generates the fourth acoustic signal by convolving the inverse characteristics of the previously measured transfer characteristics into the third acoustic signal. With this configuration, as will be described later, it is possible to improve the quality of binaural playback. After that, control unit 33 stores the generated fourth acoustic signal in storage unit 32. The fourth acoustic signal is the binaurally recorded content. Thus, according to this embodiment, corrections to improve the quality of binaural playback can be performed in advance during binaural recording.

[0047] It is desirable that the user wearing microphone 20 during binaural recording and the user wearing measurement earphones 10 and microphone 20 during transmission characteristic measurement be the same. Furthermore, it is desirable that the placement of microphone 20 during binaural recording and the placement of microphone 20 during transmission characteristic measurement be the same. In this case, it is possible to maximize the effect of correction and improve the quality of binaural playback. Of course, the user wearing microphone 20 during binaural recording and the user wearing measurement earphones 10 and microphone 20 during transmission characteristic measurement may be different. In the following, it will be assumed that binaural recording is performed with user A wearing microphone 20 in the same placement as during transmission characteristic measurement.

[0048] Binaural recording will be described with reference to FIG. 3.

[0049] FIG. 3 is a diagram for explaining binaural recording according to the present embodiment. As shown in FIG. 3, in the transmission path from the sound source 80 to the microphone 20, which is the target of binaural recording, there is the auricle 90 of the user A wearing the microphone 20. Therefore, the third acoustic signal acquired by the microphone 20 is represented by the following equation.

[0050]

Equation

[0051] Here, y rec (ω) is the third acoustic signal. x(ω) is an acoustic signal (hereinafter also referred to as a sound source signal) derived from the sound generated by the sound source 80.

[0052] The control unit 33 generates a fourth acoustic signal by convolving the inverse characteristic of the previously measured transfer characteristic G m (ω) with the third acoustic signal y rec (ω). The fourth acoustic signal is represented by the following equation.

[0053]

Equation

[0054] Here, y´(ω) is the fourth acoustic signal. G m -1 (ω) is the inverse characteristic of the transfer characteristic G m (ω). H a -1 (ω) is the inverse characteristic of the acoustic characteristic H a (ω) of the measurement earphone 10.

[0055] As shown in Equation (3), the fourth acoustic signal y´(ω) cancels the acoustic characteristic G A (ω) of the auricle 90 of the user A, and the inverse characteristic H a (ω) of the acoustic characteristic Ha -1 (ω) is the pre-convolved sound source signal x(ω). Therefore, during binaural playback, the acoustic characteristics G of user A's auricle 90 A (ω) is canceled, and the acoustic characteristics H of the measurement earphone 10 a It becomes possible to improve the quality of binaural playback without performing any corrections to cancel out (ω).

[0056] Since correction is not required during binaural playback, the overall processing load of the system can be significantly reduced in a system that distributes binaurally recorded content to multiple playback processing devices 40 in real time. Furthermore, when correction is performed during binaural playback, it may be necessary to distribute metadata for correction along with the binaurally recorded content. In this respect, according to this embodiment, the distribution of metadata for correction is not required, so the communication load can also be significantly reduced. The metadata for correction includes the acoustic characteristics G of user A's auricle 90. A (ω), and the acoustic characteristics H of the measurement earphone 10 a Examples include (ω), etc.

[0057] Furthermore, according to this embodiment, binaural recording is performed with the microphone 20 attached to a human user A. Therefore, compared to performing binaural recording using a dummy head, it becomes possible to perform simple and high-quality binaural recording in a variety of use cases. For example, binaural recording can be performed by attaching the microphone 20 to a user who is shooting video while moving with a camera in their hand. In addition, the user can perform binaural recording and monitoring (i.e., checking the recorded sound) simultaneously.

[0058] (3) Binaural playback The playback device 40 performs binaural playback. Binaural playback is performed with the user wearing playback earphones 50. The playback earphones 50 are placed on the user's auricle. For example, the playback earphones 50 may be placed in the concha.

[0059] More specifically, the control unit 43 plays back the fourth acoustic signal stored in the memory unit 32 through the playback earphones 50. For example, the control unit 43 controls the communication unit 41 to receive the fourth acoustic signal stored in the memory unit 32. Next, the control unit 43 stores the fourth acoustic signal received by the communication unit 41 in the memory unit 42. After that, the control unit 43 outputs the fourth acoustic signal stored in the memory unit 42 to the playback earphones 50, and the playback earphones 50 play back the fourth acoustic signal. As a result, the user wearing the playback earphones 50 can listen to the binaurally recorded sound.

[0060] The user wearing the microphone 20 during binaural recording and the user wearing the playback earphones 50 during binaural playback may be the same person. That is, binaural playback may be performed with user A wearing the playback earphones 50. On the other hand, the user wearing the microphone 20 during binaural recording and the user wearing the playback earphones 50 during binaural playback may be different people. That is, binaural playback may be performed with user B, who is different from user A, wearing the playback earphones 50. User B is an example of a second user in this embodiment.

[0061] Furthermore, the measurement earphone 10 and the playback earphone 50 may be the same. On the other hand, the measurement earphone 10 and the playback earphone 50 may be different.

[0062] The following describes the sounds a user will hear when binaurally recorded content is played back in three different playback environments.

[0063] -1st playback environment The first playback environment is one in which the measurement earphone 10 and the playback earphone 50 are identical, and the playback earphone 50 is worn by user A. Binaural playback in the first playback environment will be explained with reference to Figure 4.

[0064] Figure 4 is a diagram illustrating binaural playback according to this embodiment. As shown in Figure 4, the transmission path from the playback earphone 50, which is identical to the measurement earphone 10, to the eardrum of user A includes the auricle 90 of user A wearing the playback earphone 50. Therefore, the acoustic signal representing the sound heard by user A is expressed by the following equation.

[0065]

number

[0066] Here, y rep (ω) is an acoustic signal representing the sound heard by the user wearing the playback earphones 50, i.e., user A. a (ω) represents the acoustic characteristics of the playback earphone 50, which is identical to the measurement earphone 10.

[0067] As shown in equation (4), user A receives the third acoustic signal y rec (ω) can be heard. In other words, user A can hear the same sound as during binaural recording. In this way, it is possible to improve the quality of binaural playback.

[0068] -Second playback environment The second playback environment is one in which the measurement earphone 10 and the playback earphone 50 are the same, and the playback earphone 50 is worn by user B, who is different from user A.

[0069] In this playback environment, the transmission path from the playback earphone 50, which is identical to the measurement earphone 10, to user B's eardrum includes user B's auricle 90, which is wearing the playback earphone 50. Therefore, the acoustic signal representing the sound heard by user B is expressed by the following equation.

[0070]

number

[0071] Here, y rep (ω) is an acoustic signal representing the sound heard by the user wearing the playback earphones 50, i.e., user B. a (ω) represents the acoustic characteristics of the playback earphone 50, which is identical to the measurement earphone 10. B (ω) represents the acoustic characteristics of User B's auricle (90).

[0072] Referring to equation (2), the acoustic signal y represents the sound that user A hears during binaural recording. rec (ω) represents the acoustic characteristics G of user A's auricle 90 in the sound source signal x(ω). A (ω) is a convolved form. In contrast, referring to equation (5), the acoustic signal y represents the sound that user B hears during binaural playback. rep (ω) represents the acoustic characteristics G of user B's auricle 90 in the sound source signal x(ω). B (ω) is a folded form. In other words, user B can hear an acoustic signal in the binaural playback environment that represents the sound that user B would have heard if binaural recording had been performed with user B wearing microphone 20 instead of user A. In this way, user B can hear sounds as if they were present at the binaural recording site, instead of user A. This makes it possible to improve the quality of binaural playback.

[0073] However, the binaurally recorded sound source signal x(ω) may include the influence of acoustic characteristics specific to user A, in addition to the acoustic characteristics of user A's auricle 90. Such acoustic characteristics include those resulting from physical characteristics other than user A's auricle 90. The acoustic signal y represents the sound heard by user B. rep (ω) will include the influence of acoustic characteristics specific to user A, who is a different person, which may impair the naturalness of the sound.

[0074] However, when binaural recording is performed with microphone 20 attached to a human ear, it is possible to improve the quality of binaural playback compared to when microphone 20 is attached to a dummy head. When binaural recording is performed with microphone 20 attached to a dummy head, the acoustic signal y representing the sound heard by user B is rep This is because (ω) would include the acoustic characteristics of the dummy head. In that case, the naturalness of the sound would be significantly impaired due to the difference in sound reflection coefficients compared to human skin and the difference in structure compared to the human body.

[0075] -Third playback environment The third playback environment is one in which the measurement earphone 10 and the playback earphone 50 are different, and the playback earphone 50 is worn by user B, who is different from user A.

[0076] In this playback environment, the transmission path from the playback earphone 50, which is different from the measurement earphone 10, to user B's eardrum includes user B's auricle 90, which is wearing the playback earphone 50. Therefore, the acoustic signal representing the sound heard by user B is expressed by the following equation.

[0077]

number

[0078] Here, y rep (ω) is an acoustic signal representing the sound heard by the user wearing the playback earphones 50, i.e., user B. n (ω) represents the acoustic characteristics of the playback earphone 50, which is different from the measurement earphone 10. B (ω) represents the acoustic characteristics of User B's auricle (90).

[0079] Referring to equation (6), User B uses the acoustic characteristics H corresponding to the difference between the measurement earphone 10 and the playback earphone 50 to represent the acoustic signal indicating the sound that User B hears in the second playback environment described above. n (ω) / H aThe listener will hear a convoluted (ω) sound. In other words, user B will be able to hear a sound in the binaural playback environment that is similar to the sound that user B would have heard if binaural recording had been performed with user B wearing microphone 20 instead of user A. Therefore, an improvement in the quality of binaural playback is expected.

[0080] (4) Processing flow - Measurement of transfer characteristics The following describes the processing flow for measuring transfer characteristics according to this embodiment, with reference to Figure 5. Figure 5 is a sequence diagram showing an example of the processing flow for measuring transfer characteristics performed by the signal processing system 1 according to this embodiment. This sequence involves a measurement earphone 10, a microphone 20, and a recording processing device 30.

[0081] As shown in Figure 5, first, the recording processing device 30 outputs the first acoustic signal to the measurement earphone 10 (step S102).

[0082] Next, the measurement earphone 10 plays the input first acoustic signal (step S104). 。

[0083] Next, the microphone 20 acquires a second acoustic signal (step S106). The second acoustic signal is the first acoustic signal reproduced from the measurement earphone 10, and is an acoustic signal originating from the sound that arrived at the microphone 20.

[0084] Next, the microphone 20 outputs the acquired second acoustic signal to the recording processing device 30 (step S108).

[0085] Next, the recording processing device 30 calculates the transfer characteristics based on the first acoustic signal and the second acoustic signal (step S110).

[0086] Then, the recording processing device 30 stores the calculated transfer characteristics (step S112).

[0087] - Binaural recording and binaural playback The following describes the processing flow for binaural recording and binaural playback according to this embodiment, with reference to Figure 6. Figure 6 is a sequence diagram showing an example of the processing flow for binaural recording and binaural playback performed by the signal processing system 1 according to this embodiment. This sequence involves a microphone 20, a recording processing device 30, a playback processing device 40, and playback earphones 50.

[0088] As shown in Figure 6, first, microphone 20 acquires a third acoustic signal arriving from the sound source to be binaurally recorded (step S202).

[0089] Next, the microphone 20 outputs the acquired third acoustic signal to the recording processing device 30 (step S204).

[0090] Next, the recording processing device 30 generates a fourth acoustic signal by convolving the inverse characteristics of the transfer characteristics into the third acoustic signal (step S206).

[0091] Next, the recording processing device 30 stores the generated fourth acoustic signal (step S208).

[0092] The processes described above relate to binaural recording. The following sections will explain the processes related to binaural playback.

[0093] The recording device 30 transmits the stored fourth acoustic signal to the playback device 40 (step S210). For example, the recording device 30 transmits the fourth acoustic signal in response to a request from the playback device 40.

[0094] Next, the playback processing unit 40 outputs the received fourth acoustic signal to the playback earphone 50 (step S212).

[0095] Then, the playback earphone 50 plays the input fourth sound signal (step S214).

[0096] <3. Supplement> While preferred embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the present disclosure is not limited to such examples. It is clear to any person with ordinary skill in the art to which the present disclosure pertains that various modifications or alterations may be conceived within the scope of the technical ideas described in the claims, and these will naturally be understood to fall within the technical scope of the present disclosure.

[0097] <3.1. First variation> This modified example demonstrates how to perform correction during binaural recording in a third playback environment where the measurement earphone 10 and the playback earphone 50 are different. The following describes points specific to this modified example, while points common to the above embodiment are omitted.

[0098] (1) Measurement of transfer characteristics The control unit 33 measures the transmission characteristics in the same manner as in the above embodiment. Furthermore, the control unit 33 measures the acoustic characteristics of the measurement earphone 10 and the playback earphone 50, respectively. The acoustic characteristics of the measurement earphone 10 can be measured in a free space such as an anechoic chamber. Similarly, the acoustic characteristics of the playback earphone 50 can be measured in a free space such as an anechoic chamber.

[0099] (2) Binaural recording The control unit 33 generates a fourth acoustic signal in the same manner as in the above embodiment. Furthermore, the control unit 33 generates a fifth acoustic signal by correcting the fourth acoustic signal based on the acoustic characteristics of the measurement earphone 10 and the playback earphone 50, which were measured in advance. Specifically, the control unit 33 generates the fifth acoustic signal by convolving the acoustic characteristics of the measurement earphone 10 and the inverse characteristics of the acoustic characteristics of the playback earphone 50 into the fourth acoustic signal. With this configuration, as will be described later, it is possible to improve the quality of binaural playback in the third playback environment. After that, the control unit 33 stores the generated fifth acoustic signal in the storage unit 32. The fifth acoustic signal is the binaurally recorded content. Thus, with this modified example, corrections to improve the quality of binaural playback in the third playback environment can be performed in advance during binaural recording.

[0100] The fifth acoustic signal is expressed by the following equation:

[0101]

number

[0102] Here, y´´(ω) is the fifth acoustic signal. y´(ω) is the fourth acoustic signal. H a (ω) represents the acoustic characteristics of the measurement earphone 10. 1 / H n (ω) represents the acoustic characteristics H of the playback earphone 50. n It is the inverse characteristic of (ω).

[0103] As shown in equation (7), the fifth acoustic signal y''(ω) is the acoustic characteristic G of user A's auricle 90. A (ω) and acoustic characteristics H of the measurement earphone 10 a (ω) is canceled, and the acoustic characteristics H of the playback earphone 50 n (ω) inverse characteristic 1 / H n (ω) is the pre-convolved sound source signal x(ω). Therefore, during binaural playback, the acoustic characteristics G of user A's auricle 90 A (ω) is canceled, and the acoustic characteristics of the playback earphone 50 H nThis makes it possible to improve the quality of binaural playback in a third playback environment without having to perform any corrections to cancel out (ω).

[0104] (3) Binaural playback The control unit 43 reproduces the fifth acoustic signal stored in the memory unit 32 through the playback earphones 50. For example, the control unit 43 controls the communication unit 41 to receive the fifth acoustic signal stored in the memory unit 32. Next, the control unit 43 stores the fifth acoustic signal received by the communication unit 41 in the memory unit 42. After that, the control unit 43 outputs the fifth acoustic signal stored in the memory unit 42 to the playback earphones 50, and the playback earphones 50 reproduce the fifth acoustic signal. As a result, the user wearing the playback earphones 50 can listen to the binaurally reproduced sound.

[0105] When the fifth acoustic signal is reproduced in the third playback environment, the acoustic signal representing the sound heard by user B is expressed by the following formula.

[0106]

number

[0107] Referring to equation (8), user B hears the same sound as the sound heard in the second playback environment. That is, in the above embodiment, the quality of binaural playback was reduced in the third playback environment by the amount of the difference between the measurement earphone 10 and the playback earphone 50, whereas in this modified example, it is possible to avoid such a reduction in the quality of binaural playback. In this way, user B can hear sounds in the third playback environment as if they were actually present at the binaural recording site. The same applies to the first and second playback environments. In this way, it is possible to improve the quality of binaural playback.

[0108] (4) Processing flow The following describes the processing flow for binaural recording and binaural playback according to this modified example, with reference to Figure 7. Figure 7 is a sequence diagram showing an example of the processing flow for binaural recording and binaural playback performed by the signal processing system 1 according to this modified example. This sequence involves a microphone 20, a recording processing device 30, a playback processing device 40, and playback earphones 50.

[0109] As shown in Figure 7, first, the microphone 20 acquires a third acoustic signal originating from the sound source to be binaurally recorded (step S302).

[0110] Next, the microphone 20 outputs the acquired third acoustic signal to the recording processing device 30 (step S304).

[0111] Next, the recording processing device 30 generates a fourth acoustic signal by convolving the inverse characteristics of the transfer characteristics into the third acoustic signal (step S306).

[0112] Next, the recording processing device 30 generates a fifth acoustic signal by convolving the acoustic characteristics of the measurement earphone 10 and the inverse acoustic characteristics of the playback earphone 50 into a fourth acoustic signal (step S308).

[0113] Next, the recording processing device 30 stores the generated fifth acoustic signal (step S310).

[0114] The processes described above relate to binaural recording. The following sections will explain the processes related to binaural playback.

[0115] The recording device 30 transmits the stored fifth acoustic signal to the playback device 40 (step S312). For example, the recording device 30 transmits the fifth acoustic signal in response to a request from the playback device 40.

[0116] Next, the playback processing device 40 outputs the received fifth acoustic signal to the playback earphone 50 (step S314).

[0117] Then, the playback earphone 50 plays the input fifth acoustic signal (step S316).

[0118] (5) Supplement The playback earphones 50 that can be used during binaural playback may be of multiple types. In that case, the recording processing device 30 may generate a fifth acoustic signal for each of the multiple types of playback earphones 50 that can be used during binaural playback. The recording processing device 30 may then store the fifth acoustic signal for each type of playback earphone 50. During binaural playback, the recording processing device 30 may transmit the fifth acoustic signal corresponding to the playback earphones 50 used for binaural playback to the playback processing device 40. With this configuration, it is possible to improve the quality of binaural playback regardless of which type of playback earphones 50 are used for binaural playback.

[0119] <3.2. Hardware Configuration Example> The measurement earphone 10 and microphone 20 can be implemented using a variety of hardware. One example will be explained with reference to Figure 8.

[0120] Figure 8 is a schematic diagram showing an example of the hardware configuration of the measurement earphone 10 and microphone 20. As shown in Figure 8, a sound collection jig 200, which includes headphones 100 as the measurement earphone 10 and a microphone 20, is attached to the user's auricle 90.

[0121] (1) Headphones 100 The headphones 100 are an audio output device that reproduces acoustic signals. The headphones 100 are an example of a measurement earphone 10. The headphones 100 are configured as a so-called ear cuff type and are worn by the user so as to cover a part of the sound collection jig 200 worn by the user. The headphones 100 include a driver unit 110 and a frame 120.

[0122] The driver unit 110 is a device that converts the input acoustic signal into sound and emits it into the surrounding space.

[0123] The frame 120 is a component that holds the driver unit 110 to the auricle 90. When the headphones 100 are worn by the user, the frame 120 is curved so as to pass outside at least one of the helix 96 or earlobe 97 from the front to the back of the auricle 90. The driver unit 110 is connected to one end of the frame 120. The frame 120 then clamps the auricle 90 from the front and back of the auricle 90 between the driver unit 110 connected to one end of the frame 120 and the other end of the frame 120.

[0124] (2) Sound collection jig 200 The sound collection jig 200 has an insertion section 210 including a microphone 20, a first frame 220, a second frame 230, and a third frame 240.

[0125] The insertion part 210 is a component that is inserted into the user's ear canal 98. The insertion part 210 is configured as a cylindrical body having a through-hole that penetrates in the insertion direction. The microphone 20 is positioned inside the through-hole of the insertion part 210, with a gap between it and the inner wall of the through-hole. Therefore, when the insertion part 210 is inserted into the user's ear canal 98, the microphone 20 is positioned near the user's eardrum. Furthermore, sounds arriving from the outside world pass through the through-hole and reach the user's eardrum. Consequently, the user can clearly hear ambient sounds while wearing the sound-collecting jig 200.

[0126] The first frame 220 is a ring-shaped component. The first frame 220 contacts the user's concha 92 when the sound collection jig 200 is attached to the user. The first frame 220 is connected to the insertion part 210.

[0127] The second frame 230 is a component configured in the shape of a shark fin with weight-reducing cutouts. The second frame 230 contacts the user's concha 91 when the sound-collecting jig 200 is attached to the user. The second frame 230 is connected to the first frame 220.

[0128] The third frame 240 is curved to pass outside the user's helix crus 93, from the front to the back of the user's auricle 90, when the sound-collecting jig 200 is attached to the user. The third frame 240 is connected to the first frame 220.

[0129] (3) Supplement The above describes an example of the hardware configuration of the measurement earphone 10 and microphone 20. According to the example described above, the microphone 20 can be inserted into the user's ear canal 98 and positioned near the eardrum, while the measurement earphone 10 can be positioned on the user's auricle 90. Furthermore, it is possible to measure the transmission characteristics and perform binaural recording while the user's ear canal 98 remains open. Moreover, since the measurement of transmission characteristics and binaural recording can be performed while the device is worn, it becomes easy to keep the position of the microphone 20 the same during the measurement of transmission characteristics and during binaural recording. As a result, it becomes easy to maximize the effect of correction and improve the quality of binaural playback.

[0130] Although the above describes an example in which the headphones 100 and the sound-collecting jig 200 are configured as separate devices, this disclosure is not limited to such an example. The headphones 100 and the sound-collecting jig 200 may be implemented as the same device. For example, a driver unit 110 may be provided on the first frame 220. In other words, the measurement earphone 10 and the microphone 20 may be mounted on the same device.

[0131] <3.3. Network Configuration Example> The above example shows a case where the recording processing device 30 and the playback processing device 40 communicate directly, but this disclosure is not limited to such an example. As will be explained below with reference to Figures 9 and 10, communication between the recording processing device 30 and the playback processing device 40 may be relayed by other devices.

[0132] Figure 9 shows another example of the configuration of the signal processing system 1. As shown in Figure 9, the signal processing system 1 may include a server 60 in addition to the devices shown in Figure 1. The server 60 is an information processing device located on the internet. The recording processing device 30 and the playback processing device 40 may be connected via the server 60. For example, the recording processing device 30 uploads binaurally recorded content to the server 60. The playback processing device 40 then downloads the binaurally recorded content from the server 60 and plays it back using playback earphones 50. Such a communication path can be used, for example, when distributing binaurally recorded content in real time over the internet.

[0133] Figure 10 shows another example of the configuration of signal processing system 1. As shown in Figure 10, signal processing system 1 may include a server 60 and a terminal device 70 in addition to the devices shown in Figure 1. Server 60 is an information processing device located on the internet. Terminal device 70 is an information processing device operated by a user. An example of terminal device 70 is a smartphone or tablet terminal. The recording processing device 30 and the playback processing device 40 may be connected via server 60 and terminal device 70. For example, recording processing device 30 transmits binaurally recorded content to terminal device 70. Terminal device 70 uploads the received binaurally recorded content to server 60. Then, playback processing device 40 downloads the binaurally recorded content from server 60 and plays it back using playback earphones 50. Such a communication path is, for example, a binaurally recorded content on the internet Tongue It can be used when broadcasting in real time.

[0134] By including the terminal device 70 in the signal processing system 1, the communication function with the server 60 can be omitted from the recording processing device 30. Furthermore, various settings related to real-time distribution can be performed via the terminal device 70. This improves user convenience regarding real-time distribution. The terminal device 70 may also have an imaging unit such as a camera. The terminal device 70 may upload the video recorded in parallel with the binaural recording to the server 60 along with the acoustic signal obtained from the binaural recording. The playback processing device 40 may then download the video recorded in parallel with the binaural recording and play it back together with the acoustic signal obtained from the binaural recording. In this case, it becomes possible to play back the video along with the realistic sound recorded using binaural recording.

[0135] <3.4. Others> Although the above describes an example in which the control unit 33 and the control unit 43 are mounted in different devices, this disclosure is not limited to such examples. The control unit 33 and the control unit 43 may be mounted in the same device. That is, the calculation of transfer characteristics, binaural recording, and binaural playback may be performed by a single information processing device having the control unit 33 and the control unit 43.

[0136] The above describes an example of correction performed during binaural recording, but this disclosure is not limited to such examples. At least some of the corrections may be performed during binaural playback rather than during binaural recording. For example, correction based on pre-measured transfer characteristics may be performed by the playback processing device 40. As another example, correction based on pre-measured acoustic characteristics of the measurement earphone 10 and the playback earphone 50 may be performed by the playback processing device 40. Furthermore, when correction is performed during binaural playback, the measurement of transfer characteristics may be performed after binaural recording. The same applies to the measurement of the acoustic characteristics of the measurement earphone 10 and the playback earphone 50.

[0137] The above describes an example in which the measurement of transfer characteristics and binaural recording are performed with a human being fitted with measurement earphones 10 and / or microphones 20, but this disclosure is not limited to such an example. The measurement of transfer characteristics and binaural recording may also be performed with a dummy head fitted with measurement earphones 10 and / or microphones 20.

[0138] Figure 1 illustrates an example in which the signal processing system 1 has two measurement earphones 10, two microphones 20, and two playback earphones 50 for both ears, but the disclosure is not limited to this example. The signal processing system 1 may have one measurement earphone 10, one microphone 20, and one playback earphone 50 for one ear. In other words, the disclosure is applicable not only to binaural recording / playback targeting both ears, but also to binaural recording / playback targeting one ear.

[0139] Each device described herein may be implemented as a standalone device, or some or all of them may be implemented as separate devices. For example, some of the functions of the recording processing device 30 shown in Figure 1 may be provided on a device such as a server connected via a network. Specifically, at least some of the information stored by the storage unit 32 or the processing performed by the control unit 33 may be stored or executed by the server. As another example, some of the functions of the playback processing device 40 shown in Figure 1 may be provided on a device such as a server connected via a network. Specifically, at least some of the information stored by the storage unit 42 or the processing performed by the control unit 43 may be stored or executed by the server. As yet another example, the server 60 shown in Figure 9 or Figure 10 may be implemented as a standalone device, or as a combination of multiple devices. Specifically, the recording processing device 30 and the playback processing device 40 may communicate via a mesh network, i.e., via multiple devices. Furthermore, the implementation destination of some of the functions of the recording processing device 30 or playback processing device 40 shown in Figure 1 is not limited to one, but may be two or more devices. For example, some of the functions of the recording processing device 30 or playback processing device 40 shown in Figure 1 may be distributed and provided across multiple devices on a mesh network.

[0140] The above describes an example in which a first acquisition unit used for measuring transfer characteristics and a second acquisition unit used for binaural recording are implemented as a single microphone 20, but this disclosure is not limited to such an example. The first acquisition unit and the second acquisition unit may be separate. That is, different audio input devices may be used for measuring transfer characteristics and for binaural recording.

[0141] The series of processes performed by each device described herein may be implemented using software, hardware, or a combination of software and hardware. The programs constituting the software are pre-stored on a recording medium (more specifically, a non-temporary storage medium readable by a computer) located inside or outside each device. Each program is then loaded into RAM when executed by a computer controlling each device described herein, and executed by a processing circuit such as a CPU. The recording medium is, for example, a magnetic disk, an optical disk, a magneto-optical disk, or flash memory. The computer program may also be distributed via a network, for example, without using a recording medium. The computer may be an application-specific integrated circuit such as an ASIC, a general-purpose processor that performs functions by loading software programs, or a computer on a server used for cloud computing. Furthermore, the series of processes performed by each device described herein may be distributed and processed by multiple computers.

[0142] Furthermore, the processes described herein using flowcharts and sequence diagrams do not necessarily have to be executed in the order shown. Some processing steps may be executed in parallel. Additional processing steps may be adopted, and some processing steps may be omitted. [Explanation of symbols]

[0143] 1. Signal Processing System 10 (10A, 10B) Measurement earphones 20 (20A, 20B) Microphone 30 Recording Processing Device 31 Communications Department 32 Storage section 33 Control Unit 40 Regeneration Processing Equipment 41 Communications Department 42 Storage section 43 Control Unit 50 (50A, 50B) Playback Earphones 60 servers 70 Terminal devices 80 sound sources 90 Auricle

Claims

1. The transfer characteristics corresponding to the difference between the first acoustic signal and the second acoustic signal, which corresponds to the first acoustic signal and is reproduced by the first reproduction unit that reproduces the acoustic signal, are calculated. A first control unit generates a fourth acoustic signal by convolving the inverse characteristics of the calculated transfer characteristics with the third acoustic signal acquired by the second acquisition unit. Equipped with, The first acquisition unit is an audio input device positioned near the eardrum of the first user, who is either a user or a dummy user. The first playback unit is an audio output device positioned close to the ear of the first user, The aforementioned transmission characteristics are acoustic characteristics that include the acoustic characteristics of the first playback unit and the acoustic characteristics of the first user's auricle, The second acquisition unit is an audio input device positioned near the eardrum of the first user, The third acoustic signal is an acoustic signal obtained by convolving the acoustic characteristics of the first user's auricle into a sound source signal, which is an acoustic signal derived from sound generated from a sound source that is the subject of binaural recording. The fourth acoustic signal is an acoustic signal in which the inverse characteristics of the acoustic characteristics of the first playback unit are convolved into the sound source signal. Signal processing system.

2. The first control unit causes the generated fourth acoustic signal to be stored in the storage unit. The signal processing system according to claim 1.

3. The signal processing system further comprises a second control unit which causes the second playback unit, which reproduces the acoustic signal, to reproduce the fourth acoustic signal stored in the storage unit. The signal processing system according to claim 2.

4. The first control unit generates a fifth acoustic signal by convolving the characteristics of the first playback unit and the inverse characteristics of the second playback unit that reproduces the acoustic signal into the fourth acoustic signal. The signal processing system according to claim 1.

5. The first control unit causes the generated fifth acoustic signal to be stored in the storage unit. The signal processing system according to claim 4.

6. The signal processing system further includes a second control unit that causes the second playback unit to reproduce the fifth acoustic signal stored in the storage unit. The signal processing system according to claim 5.

7. The first control unit and the second control unit are mounted on different devices. The signal processing system according to claim 3 or 6.

8. The first playback unit is positioned in the auricle of the first user, The second playback unit is positioned in the auricle of a second user, who is different from the first user. The signal processing system according to any one of claims 3 to 6.

9. The second playback unit is different from the first playback unit. The signal processing system according to claim 8.

10. The first playback unit reproduces the first audio signal, The first acquisition unit acquires an acoustic signal, and the second acoustic signal corresponds to the first acoustic signal reproduced by the first playback unit. To calculate the transfer characteristics corresponding to the difference between the first acoustic signal and the second acoustic signal, The second acquisition unit acquires the third acoustic signal, A fourth acoustic signal is generated by convolving the inverse characteristics of the aforementioned transfer characteristics into the third acoustic signal. Includes, The first acquisition unit is an audio input device positioned near the eardrum of the first user, who is either a user or a dummy user. The first playback unit is an audio output device positioned close to the ear of the first user, The aforementioned transmission characteristics are acoustic characteristics that include the acoustic characteristics of the first playback unit and the acoustic characteristics of the first user's auricle, The second acquisition unit is an audio input device positioned near the eardrum of the first user, The third acoustic signal is an acoustic signal obtained by convolving the acoustic characteristics of the first user's auricle into a sound source signal, which is an acoustic signal derived from sound generated from a sound source that is the subject of binaural recording. The fourth acoustic signal is an acoustic signal in which the inverse characteristics of the acoustic characteristics of the first playback unit are convolved into the sound source signal. Signal processing method.

11. Computers, The transfer characteristics corresponding to the difference between the first acoustic signal and the second acoustic signal, which corresponds to the first acoustic signal and is reproduced by the first reproduction unit that reproduces the acoustic signal, are calculated. A first control unit generates a fourth acoustic signal by convolving the inverse characteristics of the calculated transfer characteristics with the third acoustic signal acquired by the second acquisition unit. To make it function as, The first acquisition unit is an audio input device positioned near the eardrum of the first user, who is either a user or a dummy user. The first playback unit is an audio output device positioned close to the ear of the first user, The aforementioned transmission characteristics are acoustic characteristics that include the acoustic characteristics of the first playback unit and the acoustic characteristics of the first user's auricle, The second acquisition unit is an audio input device positioned near the eardrum of the first user, The third acoustic signal is an acoustic signal obtained by convolving the acoustic characteristics of the first user's auricle into a sound source signal, which is an acoustic signal derived from sound generated from a sound source that is the subject of binaural recording. The fourth acoustic signal is an acoustic signal in which the inverse characteristics of the acoustic characteristics of the first playback unit are convolved into the sound source signal. program.

Citation Information

Patent Citations

  • Binaural recording and noise canceling headphone

    JP2009049947A

  • Information-processing device and method

    WO2017195616A1