FILTER GENERATION DEVICE, FILTER GENERATION METHOD, AND PROGRAM

The filter generation device and method address the limitations of out-of-head localization by generating filters tailored to individual user equipment and environments, ensuring high-quality sound localization across different playback devices.

JP7750003B2Active Publication Date: 2025-10-07JVC KENWOOD CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021156783
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-10-07
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

Existing out-of-head localization technologies are limited by the need for specific playback devices and variations in transfer characteristics due to different headphones, speakers, and measurement environments, leading to potential clipping of processed signals.

Method used

A filter generation device and method that includes frequency characteristic acquisition, level calculation, correction, and filter generation to create filters suitable for out-of-head localization processing, using spatial and ear canal transfer characteristics measured on the user's specific equipment and environment.

Benefits of technology

Enables effective out-of-head localization processing adaptable to various playback devices and environments, ensuring high-quality sound image localization without signal clipping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007750003000002
    Figure 0007750003000002
  • Figure 0007750003000003
    Figure 0007750003000003
  • Figure 0007750003000004
    Figure 0007750003000004
Patent Text Reader

Abstract

To provide a filter generation device and filter generation method capable of generating a filter suitable for out-of-head localization processing.SOLUTION: A processing device comprises: a frequency characteristic acquisition section for acquiring frequency characteristics based on a sound collection signal; a level calculation section for calculating a reference level in the frequency characteristics; a correction section 225 for calculating correction characteristics by correcting the frequency characteristics to be within a predetermined level range including the reference level; and a filter generation section 230 for generating a correction filter based on the correction characteristics.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure provides a filter generation device, Filter generation method and program Regarding. [Background technology]

[0002] One type of sound image localization technology is out-of-head localization, which uses headphones to localize a sound image outside the listener's head. Out-of-head localization technology localizes a sound image outside the head by canceling the characteristics from the headphones to the ears (headphone characteristics) and providing the characteristics of two lines from one speaker (monaural speaker) to the ears (spatial acoustic transfer characteristics).

[0003] In out-of-head localization playback using stereo speakers, measurement signals (impulse sounds, etc.) emitted from two-channel (hereinafter referred to as ch) speakers are recorded by microphones (hereinafter referred to as mics) placed at the ears of the listener. Then, a processing device generates a filter based on the collected signal obtained by collecting the measurement signal. Out-of-head localization playback can be achieved by convolving the generated filter with the two-channel audio signal.

[0004] Furthermore, in order to generate a filter (also called an inverse filter) that cancels the characteristics from the headphones to the ear, the characteristics from the headphones to the eardrum (also called the ear canal transfer function ECTF, or ear canal transfer characteristics) are measured using a microphone placed in the listener's own ear.

[0005] Patent Document 1 discloses a device that performs out-of-head localization processing. Furthermore, in Patent Document 1, the out-of-head localization processing performs DRC (Dynamic Range Compression) processing on the playback signal. In the DRC processing, the processing device smoothes the frequency characteristics. Furthermore, the processing device performs band division based on the smoothed characteristics. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Publication No. 2019-62430 Summary of the Invention [Problem to be solved by the invention]

[0007] In such out-of-head localization listening, it is desirable to perform processing without being limited to a specific playback device. For example, it is desirable to perform appropriate out-of-head localization processing even when a user's own headphones are used as the playback device. Alternatively, it is desirable to reproduce the spatial acoustic transfer characteristics in an environment where the user's usual speakers are installed as the playback device.

[0008] Changing the playback device may change the transfer characteristics. Therefore, it is preferable to measure the user's personal characteristics (spatial acoustic transfer characteristics and ear canal transfer characteristics) using the playback device that the user is currently using. Even when measuring personal characteristics, sharp peaks and dips may occur in the frequency characteristics, which may cause clipping of the out-of-head localization processed signal.

[0009] Peaks and dips vary depending on the characteristics of playback devices such as speakers and headphones, or the acoustic characteristics of the room in which the measurement is performed. Peaks and dips also vary depending on the shape of the user's head and ears. Therefore, the levels and frequencies of peaks and dips vary due to various factors. Depending on the playback device and measurement environment, it becomes necessary to check the characteristics and make adjustments accordingly.

[0010] The present disclosure has been made in consideration of the above points, and has an object to provide a filter generation device and a filter generation method that can generate a filter suitable for out-of-head localization processing. [Means for solving the problem]

[0011] The filter generation device according to this embodiment includes a frequency characteristic acquisition unit that acquires frequency characteristics based on a picked-up signal, a level calculation unit that calculates a reference level for the frequency characteristics, a correction unit that calculates correction characteristics by correcting the frequency characteristics so that the frequency characteristics fall within a predetermined level range including the reference level, and a filter generation unit that generates a correction filter based on the correction characteristics.

[0012] The filter generation method s according to this embodiment includes the steps of acquiring frequency characteristics based on a picked-up signal, calculating a reference level for the frequency characteristics, calculating correction characteristics by correcting the frequency characteristics so that the frequency characteristics fall within a predetermined level range including the reference level, and generating a filter based on the correction characteristics. [Effects of the Invention]

[0013] According to the present disclosure, it is possible to provide a filter generation device and a filter generation method that can generate a filter suitable for out-of-head localization processing. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a block diagram showing an out-of-head localization processing device according to an embodiment of the present invention; [Figure 2] FIG. 1 is a diagram illustrating a configuration of a measurement device for measuring spatial acoustic transfer characteristics. [Figure 3] FIG. 1 is a diagram illustrating a configuration of a measurement device for measuring ear canal transfer characteristics. [Figure 4] FIG. 2 is a control block diagram showing the configuration of the processing device. [Figure 5] 1 is a flowchart illustrating a method for generating a filter in a processing device. [Figure 6] 10 is a flowchart showing a first example of correction processing. [Figure 7] 10 is a graph showing frequency amplitude characteristics before and after correction according to processing example 1. [Figure 8]10 is a flowchart showing a second example of the correction process. [Figure 9] 10 is a graph showing frequency amplitude characteristics before and after correction according to processing example 2. [Figure 10] 10 is a flowchart showing a fourth example of the correction process. [Figure 11] 10 is a graph showing frequency bands according to processing example 4. [Figure 12] FIG. 10 is a block diagram showing a configuration of a processing device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] An overview of the sound image localization processing according to this embodiment will be described. The out-of-head localization processing according to this embodiment is performed using spatial acoustic transfer characteristics and ear canal transfer characteristics. The spatial acoustic transfer characteristics are the transfer characteristics from a sound source such as a speaker to the ear canal. The ear canal transfer characteristics are the transfer characteristics from a speaker unit of headphones or earphones to the eardrum. In this embodiment, the spatial acoustic transfer characteristics are measured when headphones or earphones are not worn, and the ear canal transfer characteristics are measured when headphones or earphones are worn, and the out-of-head localization processing is realized using these measurement data. This embodiment is characterized by a microphone system for measuring the spatial acoustic transfer characteristics or the ear canal transfer characteristics.

[0016] The out-of-head localization processing according to this embodiment is executed by a user terminal such as a personal computer, a smartphone, or a tablet PC. The user terminal is an information processing device having processing means such as a processor, storage means such as a memory or a hard disk, display means such as an LCD monitor, and input means such as a touch panel, buttons, a keyboard, or a mouse. The user terminal may have a communication function for transmitting and receiving data. Furthermore, output means (output unit) having headphones or earphones is connected to the user terminal. The connection between the user terminal and the output means may be wired or wireless.

[0017] Embodiment 1 (Extracranial stereotaxic processing device) FIG. 1 shows a block diagram of an out-of-head localization processing device 100, which is an example of a sound field reproduction device according to this embodiment. The out-of-head localization processing device 100 reproduces a sound field for a user U wearing headphones 43. To this end, the out-of-head localization processing device 100 performs sound image localization processing on Lch and Rch stereo input signals XL and XR. The Lch and Rch stereo input signals XL and XR are analog audio playback signals output from a CD (Compact Disc) player or the like, or digital audio data such as mp3 (MPEG Audio Layer-3). Note that audio playback signals or digital audio data are collectively referred to as playback signals. In other words, the Lch and Rch stereo input signals XL and XR are playback signals.

[0018] The out-of-head localization processing device 100 is not limited to a single physical device, and some of the processing may be performed by different devices. For example, some of the processing may be performed by a smartphone or the like, and the remaining processing may be performed by a DSP (Digital Signal Processor) built into the headphones 43.

[0019] The out-of-head localization processing device 100 includes an out-of-head localization processing unit 10, a filter unit 41 that stores an inverse filter Linv, a filter unit 42 that stores an inverse filter Rinv, and headphones 43. The out-of-head localization processing unit 10, the filter unit 41, and the filter unit 42 can be specifically realized by a processor or the like.

[0020] The out-of-head localization processing unit 10 includes convolution calculation units 11-12, 21-22 that store spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs, and adders 24 and 25. The convolution calculation units 11-12, 21-22 perform convolution processing using the spatial acoustic transfer characteristics. Stereo input signals XL and XR from a CD player or the like are input to the out-of-head localization processing unit 10. Spatial acoustic transfer characteristics are set in the out-of-head localization processing unit 10. The out-of-head localization processing unit 10 convolves the stereo input signals XL and XR of each channel with a filter of the spatial acoustic transfer characteristics (hereinafter also referred to as a spatial acoustic filter). The spatial acoustic transfer characteristics may be head-related transfer functions (HRTFs) measured on the head or pinnae of the subject, or head-related transfer functions of a dummy head or a third party.

[0021] A set of four spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs is defined as a spatial acoustic transfer function. Data used for convolution in convolution calculation units 11, 12, 21, and 22 becomes a spatial acoustic filter. A spatial acoustic filter is generated by cutting out the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs with a predetermined filter length.

[0022] The spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are obtained in advance by impulse response measurement or the like. For example, a user U wears a microphone on each of his or her left and right ears. Left and right speakers placed in front of the user U output impulse sounds for impulse response measurement. Then, a measurement signal such as the impulse sound output from the speakers is picked up by the microphone. The spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are obtained based on the signal picked up by the microphone. The spatial acoustic transfer characteristic Hls between the left speaker and the left microphone, the spatial acoustic transfer characteristic Hlo between the left speaker and the right microphone, the spatial acoustic transfer characteristic Hro between the right speaker and the left microphone, and the spatial acoustic transfer characteristic Hrs between the right speaker and the right microphone are measured.

[0023] The convolution calculation unit 11 then convolves the Lch stereo input signal XL with a spatial acoustic filter that corresponds to the spatial acoustic transfer characteristic Hls. The convolution calculation unit 11 outputs the convolution calculation data to the adder 24. The convolution calculation unit 21 convolves the Rch stereo input signal XR with a spatial acoustic filter that corresponds to the spatial acoustic transfer characteristic Hro. The convolution calculation unit 21 outputs the convolution calculation data to the adder 24. The adder 24 adds the two convolution calculation data and outputs the result to the filter unit 41.

[0024] The convolution calculation unit 12 convolves the left channel stereo input signal XL with a spatial acoustic filter corresponding to the spatial acoustic transfer characteristic Hlo. The convolution calculation unit 12 outputs the convolution calculation data to the adder 25. The convolution calculation unit 22 convolves the right channel stereo input signal XR with a spatial acoustic filter corresponding to the spatial acoustic transfer characteristic Hrs. The convolution calculation unit 22 outputs the convolution calculation data to the adder 25. The adder 25 adds the two convolution calculation data and outputs the result to the filter unit 42.

[0025] Inverse filters Linv and Rinv that cancel headphone characteristics (characteristics between the headphone playback unit and microphone) are set in filter units 41 and 42. The inverse filters Linv and Rinv are then convolved with the playback signal (convolution signal) that has been processed by out-of-head localization processing unit 10. Filter unit 41 convolves the inverse filter Linv of the Lch headphone characteristics with the Lch signal from adder 24. Similarly, filter unit 42 convolves the Rch signal from adder 25 with the inverse filter Rinv of the Rch headphone characteristics. When headphones 43 are worn, the inverse filters Linv and Rinv cancel the characteristics from the headphone unit to the microphone. The microphone may be placed anywhere between the entrance of the ear canal and the eardrum.

[0026] The filter unit 41 outputs the processed Lch signal YL to the left unit 43L of the headphones 43. The filter unit 42 outputs the processed Rch signal YR to the right unit 43R of the headphones 43. A user U wears the headphones 43. The headphones 43 output the Lch signal YL and the Rch signal YR (hereinafter, the Lch signal YL and the Rch signal YR are also collectively referred to as stereo signals) to the user U. This makes it possible to reproduce a sound image localized outside the head of the user U.

[0027] In this way, the out-head localization processing device 100 performs out-head localization processing using spatial acoustic filters corresponding to the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs, and inverse filters Linv and Rinv of headphone characteristics. In the following description, the spatial acoustic filters corresponding to the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs, and the inverse filters Linv and Rinv of headphone characteristics are collectively referred to as out-head localization processing filters. In the case of a 2-channel stereo playback signal, the out-head localization filters are composed of four spatial acoustic filters and two inverse filters. The out-head localization processing device 100 then performs convolution operation processing on the stereo playback signal using a total of six out-head localization filters, thereby executing out-head localization processing. It is preferable that the out-head localization filters be based on personal measurements of the user U. For example, the out-head localization filters are set based on sound signals picked up by microphones attached to the user U's ears.

[0028] In this way, the spatial acoustic filter and the headphone characteristic inverse filters Linv, Rinv are filters for audio signals. These filters are convoluted with the playback signals (stereo input signals XL, XR), so that the out-of-head localization processing device 100 executes out-of-head localization processing. In this embodiment, the process of generating the spatial acoustic filter is one of the technical features. Specifically, in the process of generating the spatial acoustic filter, the level range of the frequency characteristics is compressed.

[0029] (Spatial acoustic transfer characteristic measurement device) A measurement device 200 for measuring spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs will be described with reference to Fig. 2. Fig. 2 is a diagram schematically showing a measurement configuration for measuring a subject 1. Note that, in this example, the subject 1 is the same person as the user U in Fig. 1, but may be a different person.

[0030] As shown in Fig. 2, the measurement device 200 has a stereo speaker 5 and a microphone unit 2. The stereo speaker 5 is installed in a measurement environment. The measurement environment may be a room in the user U's home, an audio system sales store or showroom, or the like. The measurement environment is preferably a listening room equipped with speakers and acoustics.

[0031] In this embodiment, the processing device 201 of the measuring device 200 performs arithmetic processing for appropriately generating a spatial acoustic filter. The processing device 201 includes, for example, a music player such as a CD player. The processing device 201 may be a personal computer (PC), a tablet terminal, a smartphone, or the like. The processing device 201 may also be a server device itself.

[0032] The stereo speakers 5 include a left speaker 5L and a right speaker 5R. For example, the left speaker 5L and the right speaker 5R are placed in front of the subject 1. The left speaker 5L and the right speaker 5R output impulse sounds and the like for measuring impulse responses. In the following description of this embodiment, the number of speakers serving as sound sources is two (stereo speakers), but the number of sound sources used for measurement is not limited to two, and may be one or more. In other words, this embodiment can also be applied to a 1-channel monaural environment or a so-called multi-channel environment such as 5.1-channel or 7.1-channel.

[0033] The microphone unit 2 is a stereo microphone having a left microphone 2L and a right microphone 2R. The left microphone 2L is placed at the left ear 9L of the subject 1, and the right microphone 2R is placed at the right ear 9R of the subject 1. Specifically, it is preferable to place the microphones 2L and 2R at a position from the entrance of the ear canal to the eardrum of the left ear 9L and the right ear 9R. The microphones 2L and 2R collect measurement signals output from the stereo speakers 5 and obtain collected signals. The microphones 2L and 2R output the collected signals to the processing device 201. The subject 1 may be a person or a dummy head. In other words, in this embodiment, the subject 1 is a concept that includes not only a person but also a dummy head.

[0034] As described above, the impulse responses are measured by measuring the impulse sounds output from the left speaker 5L and the right speaker 5R with the microphones 2L and 2R. The processing device 201 stores the picked-up signals acquired by the impulse response measurement in a memory or the like. This allows the spatial acoustic transfer characteristic Hls between the left speaker 5L and the left microphone 2L, the spatial acoustic transfer characteristic Hlo between the left speaker 5L and the right microphone 2R, the spatial acoustic transfer characteristic Hro between the right speaker 5R and the left microphone 2L, and the spatial acoustic transfer characteristic Hrs between the right speaker 5R and the right microphone 2R to be measured. That is, the spatial acoustic transfer characteristic Hls is acquired by the left microphone 2L picking up the measurement signal output from the left speaker 5L. The spatial acoustic transfer characteristic Hlo is acquired by the right microphone 2R picking up the measurement signal output from the left speaker 5L. The spatial acoustic transfer characteristic Hro is acquired by the left microphone 2L picking up the measurement signal output from the right speaker 5R. The measurement signal output from the right speaker 5R is picked up by the right microphone 2R, and the spatial acoustic transfer characteristic Hrs is acquired.

[0035] Furthermore, the measuring device 200 may generate spatial acoustic filters according to the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs from the left and right speakers 5L and 5R to the left and right microphones 2L and 2R based on the collected sound signals. For example, the processing device 201 cuts out the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs with a predetermined filter length. The processing device 201 may correct the measured spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs.

[0036] In this way, the processing device 201 generates a spatial acoustic filter used in the convolution operation of the out-of-head localization processing device 100. As shown in Fig. 1, the out-of-head localization processing device 100 performs out-of-head localization processing using a spatial acoustic filter according to the spatial acoustic transfer characteristics Hls, Hlo, Hro, Hrs between the left and right speakers 5L, 5R and the left and right microphones 2L, 2R. In other words, the out-of-head localization processing is performed by convolving the spatial acoustic filter with the audio playback signal.

[0037] The processing device 201 performs the same processing on the collected sound signals corresponding to the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs. That is, the same processing is performed on each of the four collected sound signals corresponding to the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs. This makes it possible to generate spatial acoustic filters corresponding to the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs.

[0038] (Ear canal transfer characteristic measuring device) The ear canal transfer characteristic measuring device 300 will be described with reference to Fig. 3. Fig. 3 shows a configuration for measuring the transfer characteristic of a user U. The measuring device 300 measures the ear canal transfer characteristic in order to generate an inverse filter. The measuring device 300 includes a microphone unit 2, headphones 43, and a processing device 301. Note that here, the person being measured 1 is the same person as the user U in Fig. 1, but may be a different person.

[0039] In this embodiment, processing device 301 of measurement device 300 performs arithmetic processing to appropriately generate a filter according to measurement results. Processing device 301 is a personal computer (PC), tablet terminal, smartphone, etc., and includes a memory and a processor. The memory stores processing programs, various parameters, measurement data, etc. The processor executes the processing programs stored in the memory. Each process is performed by the processor executing the processing programs. The processor may be, for example, a central processing unit (CPU), a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a graphics processing unit (GPU).

[0040] 2. The processing device 301 in FIG. 3 is the same as the processing device in FIG. 201 The processing device may be the same as or different from the processing device shown in FIG. 2. That is, the measurements shown in FIG. 2 and FIG. 3 are not limited to being performed using the same processing device. For example, the measurement shown in FIG. 2 may be performed using a dedicated processing device 201 installed in a listening room or the like, while the measurement shown in FIG. 3 may be performed using a general-purpose processing device 301 such as a smartphone.

[0041] A microphone unit 2 and headphones 43 are connected to the processing device 301. The microphone unit 2 may be built into the headphones 43. The microphone unit 2 includes a left microphone 2L and a right microphone 2R. The left microphone 2L is worn on the left ear 9L of the user U. The right microphone 2R is worn on the right ear 9R of the user U. The processing device 301 may be the same processing device as the out-of-head localization processing device 100, or may be a different processing device. Also, earphones may be used instead of the headphones 43.

[0042] The headphones 43 include a headphone band 43B, a left unit 43L, and a right unit 43R. The headphone band 43B connects the left unit 43L and the right unit 43R. The left unit 43L outputs sound toward the left ear 9L of the user U. The right unit 43R outputs sound toward the right ear 9R of the user U. The headphones 43 may be of any type, such as a closed type, an open type, a semi-open type, or a semi-closed type. The user U wears the headphones 43 with the microphone unit 2 attached to them. That is, the left unit 43L and the right unit 43R of the headphones 43 are attached to the left ear 9L and the right ear 9R, respectively, on which the left microphone 2L and the right microphone 2R are attached. The headphone band 43B generates a biasing force that presses the left unit 43L and the right unit 43R against the left ear 9L and the right ear 9R, respectively.

[0043] The left microphone 2L picks up sound output from the left unit 43L of the headphones 43. The right microphone 2R picks up sound output from the right unit 43R of the headphones 43. The microphone portions of the left microphone 2L and right microphone 2R are placed in sound collection positions near the external ear canals. The left microphone 2L and right microphone 2R are configured so as not to interfere with the headphones 43. In other words, the user U can wear the headphones 43 with the left microphone 2L and right microphone 2R placed in appropriate positions on the left ear 9L and right ear 9R.

[0044] The processing device 301 outputs a measurement signal to the headphones 43. This causes the headphones 43 to generate an impulse sound or the like. Specifically, the impulse sound output from the left unit 43L is measured by the left microphone 2L. The impulse sound output from the right unit 43R is measured by the right microphone 2R. When the measurement signal is output, the microphones 2L and 2R acquire the picked-up signal, thereby performing impulse response measurement.

[0045] The processing device 301 performs similar processing on the sound signals picked up from the microphones 2L and 2R to generate inverse filters Linv and Rinv.

[0046] (Level range compression) At least one of the measuring devices 200 and 300 performs a process of compressing the frequency characteristics of the collected sound signal so that the frequency characteristics fall within a predetermined level range. In the following, the process of compressing the level range of the frequency characteristics of the collected sound signal corresponding to the spatial acoustic transfer characteristics Hls and Hlo in the measuring device 200 will be described. That is, the process of compressing the level range of the frequency characteristics of the collected sound signal corresponding to the spatial acoustic transfer characteristics Hro and Hrs in the measuring device 200 is similar to the process described below, and therefore the description will be omitted as appropriate. Similarly, the process of compressing the level range of the frequency characteristics of the collected sound signal for the left and right ear canal transfer characteristics in the measuring device 300 is similar to the process described below, and therefore the description will be omitted as appropriate.

[0047] 4 is a block diagram showing the configuration of the processing device 201 of the measurement device 200. The processing device 201 includes a measurement signal generating unit 211, a picked-up signal acquiring unit 212, a segmental power acquiring unit 215, a frequency characteristic acquiring unit 221, a level calculating unit 223, a level range setting unit 224, a correcting unit 225, an adjusting unit 231, and an inverse transforming unit 232. The inverse transforming unit 232 and the adjusting unit 231 function as a filter generating unit 230.

[0048] The measurement signal generation unit 211 includes a D / A converter, an amplifier, etc., and generates a measurement signal for measuring the spatial acoustic transfer characteristics and the ear canal transfer characteristics. The measurement signal is, for example, an impulse signal or a TSP (Time Stretched Pulse) signal. Here, the measuring device 200 performs impulse response measurement using an impulse sound as the measurement signal. The measurement signal generation unit 211 outputs the measurement signal to each of the stereo speakers 5. Here, an example will be described in which the measurement signal is output from the left speaker 5L to obtain picked-up sound signals corresponding to the spatial acoustic transfer characteristics Hls and Hlo.

[0049] The left microphone 2L and right microphone 2R of the microphone unit 2 each collect a measurement signal and output the collected signal to the processing device 201. The collected signal acquisition unit 212 acquires the collected signals collected by the left microphone 2L and right microphone 2R. The collected signal acquisition unit 212 may also include an A / D converter that performs A / D conversion on the collected signals from the microphones 2L and 2R. The collected signal acquisition unit 212 cuts out the collected signal at a predetermined time. That is, the collected signal acquisition unit 212 extracts a collected signal of a predetermined number of data (time width). The collected signal acquisition unit 212 may synchronously add signals obtained by multiple measurements. The collected signal acquired using the left microphone 2L is referred to as hls, and the collected signal acquired using the right microphone 2R is referred to as hlo. The collected signals hls and hlo are each signals sampled at a sampling frequency of 48 kHz. Furthermore, the cut-out collected signals hls and hlo each have a filter length (number of samples) of 4096. Of course, the sampling frequency and filter length are not limited to the above values.

[0050] The segmental power acquisition unit 215 acquires the segmental power of the sound pickup signal hls and the sound pickup signal hlo. For example, the segmental power of the sound pickup signal hls and the sound pickup signal hlo is defined as hlsP and hloP. The segmental power hlsP is the sum of the squares of the amplitude values ​​included in the sound pickup signal hls. The segmental power hloP is the sum of the squares of the amplitude values ​​included in the sound pickup signal hlo. In the time domain, if the sound pickup signal hls and the sound pickup signal hlo have 4096 pieces of data, the sum of the squares of the 4096 amplitude values ​​becomes the segmental powers hlsP and hloP.

[0051] The frequency characteristic acquisition unit 221 acquires frequency characteristics based on the picked-up sound signals hls and hlo. The frequency characteristic acquisition unit 221 calculates the frequency characteristics of the picked-up sound signals hls and hlo by discrete Fourier transform or discrete cosine transform. The frequency characteristic acquisition unit 221 calculates the frequency characteristics by, for example, performing FFT (fast Fourier transform) on the picked-up sound signal in the time domain. The frequency characteristics include an amplitude spectrum and a phase spectrum. Note that the frequency characteristic acquisition unit 221 may generate a power spectrum instead of the amplitude spectrum. The frequency amplitude characteristics of the picked-up sound signals hls and hlo are denoted by Fhls and Fhlo, respectively. The frequency characteristics Fhls and Fhlo become spectrum data of the amplitude spectrum.

[0052] The level calculation unit 223 calculates the reference level for the frequency characteristics Fhls and Fhlo. For example, the level calculation unit 223 calculates the average level (mean value) of the frequency characteristics Fhls and Fhlo and sets it as the reference level. For example, if an FFT is performed with a filter length (number of samples) T, the level value (dB) of each frequency in the frequency amplitude characteristic is calculated and the average value is obtained. The real (real part) and imag (imaginary part) after the T-point FFT are defined as real[i] and imag[i], respectively. Here, i is an integer from 0 to (T-1). The sound pressure level Amp_dB[i] at each i-point is expressed by the following equation (1). Amp_dB[i]=log10(sqrt(real[i]*real[i]+imag[i]*imag[i])) ···(1) In equation (1), i = 1 to (T / 2 + 1), and sqrt is the square root.

[0053] Furthermore, if the frequency (Hz) at point i is freq[i] and the sampling frequency is fs, freq[i] can be obtained by the following equation (2). freq[i]=(T / fs)*i (2)

[0054] The reference level A for the entire frequency band is given by the following equation (3).

number

[0055] If the reference level of the frequency characteristic Fhls is Ahls and the reference level of the frequency characteristic Fhlo is Ahlo, the reference level A is (Ahls+Ahlo) / 2.

[0056] Furthermore, the level calculation unit 223 calculates the maximum level maxL and minimum level minL of the frequency amplitude characteristics. The maximum level maxL is the maximum value among the amplitude values ​​included in the two spectrum data of the frequency characteristics Fhls and Fhlo. The minimum level minL is the minimum value among the amplitude values ​​included in the two spectrum data of the frequency characteristics Fhls and Fhlo. The reference level A, maximum level maxL, and minimum level minL are common values ​​for the two frequency characteristics Fhls and Fhlo.

[0057] The level range setting unit 224 sets the level range X to be compressed. The level range setting unit 224 inputs the level range X according to, for example, the playback device. To obtain an appropriate out-of-head localization effect, it is preferable to set X to 40 dB or more. Furthermore, if the amplifier of the playback device does not have high performance in terms of audio output efficiency or quality, X can be set to 20 dB. Although it is preferable to set X to 20 dB or more and 40 dB or less, it is not particularly limited to this range.

[0058] The correction unit 225 calculates correction characteristics by correcting the frequency characteristics Fhls and Fhlo so that they fall within a predetermined level range X that includes the reference level A. In other words, the correction unit 225 compresses the amplitude levels of the frequency characteristics Fhls and Fhlo so that the amplitude values ​​of the frequency characteristics fall within the level range X. For example, when the level range X=40 dB, the correction unit 225 corrects the frequency characteristics Fhls and Fhlo by compressing the amplitude values ​​so that they fall within a range of the reference level A±20 dB. The characteristics corrected by the correction unit 225 are called correction characteristics. The correction characteristic of the frequency characteristic Fhls is called NewFhls, and the correction characteristic of the frequency characteristic Fhlo is called NewFhlo.

[0059] Here, the amplitude value before correction at an arbitrary frequency is L, and the amplitude value after correction is NewL. In other words, the frequency characteristics Fhls and Fhlo are a set of amplitude values ​​L before correction, and the correction characteristics NewFhls and NewFhlo are a set of amplitude values ​​NewL.

[0060] For example, the corrector 225 can correct the frequency characteristics using the following equations (4) and (5). If L is greater than or equal to A NewL=A+(LX)*(X / 2) / (maxL-A) ···(4) If L is less than A NewL=A+(LX)*(X / 2) / (A-minL) ···(5)

[0061] By doing this, NewL falls within the level range X centered on the reference level. In other words, NewL has an amplitude value greater than or equal to (A-(X / 2)) and less than or equal to (A+(X / 2)). Then, the correction unit 225 calculates the corrected amplitude value NewL for all data (amplitude value L) within the correction band using the above equations (1) and (2). The set of corrected amplitude values ​​NewL becomes the correction characteristic. The correction characteristic is obtained by correcting the amplitude value of the frequency characteristic Fhls. Furthermore, by performing correction using (1) and (2), it is possible to compress the range while maintaining the spectral shape of the frequency characteristics Fhls and Fhlo before correction.

[0062] The frequency band corrected by the correction unit 225 may be the entire band or a part of the band. For example, the correction band for correcting the frequency characteristics Fhls and Fhlo can be set to 10 Hz to 20 kHz. In other words, the correction unit 225 does not correct amplitude values ​​in a band equal to or higher than the lowest frequency (e.g., 1 Hz) and lower than 10 Hz, or in a band higher than 20 kHz and lower than the highest frequency. Therefore, outside the correction band, the amplitude values ​​of the frequency characteristics Fhls and Fhlo are used as is. The correction band may be changed according to the playback band of the headphones 43 that perform out-of-head localization playback, that is, the headphones 43 in FIG. 1.

[0063] The filter generation unit 230 generates a correction filter based on the correction characteristics. Specifically, the filter generation unit 230 includes an inverse transformation unit 232 and an adjustment unit 231. The inverse transformation unit 232 inversely transforms the correction characteristics to generate a time-domain correction signal. The inverse transformation unit 232 calculates the time-domain correction signal from the correction characteristics and phase characteristics by inverse discrete Fourier transform or inverse discrete cosine transform. The inverse transformation unit 232 generates the time-domain correction signal by performing IFFT (inverse fast Fourier transform) on the correction characteristics and phase characteristics. The correction signal obtained from the correction characteristics NewFhls is defined as hls2. The correction signal obtained from the correction characteristics NewFhlo is defined as hlo2. The correction signals hls2 and hlo2 have the same filter length as the extracted picked-up signal.

[0064] The phase characteristics calculated by the frequency characteristic acquisition unit 221 can be used as they are. That is, the inverse transformation unit 232 generates the correction signal hls2 by performing an inverse Fourier transform on the phase characteristics corresponding to the frequency characteristic Fhls and the correction characteristics NewFhls. The inverse transformation unit 232 generates the correction signal hlo2 by performing an inverse Fourier transform on the phase characteristics corresponding to the frequency characteristic Fhlo and the correction characteristics NewFhlo.

[0065] The segmental power acquisition unit 215 acquires the segmental power of the corrected signal hls2 and the corrected signal hlo2. As described above, the segmental power can be the sum of squares of the amplitude values ​​of the time domain signal. The segmental power of the corrected signal hls2 is hls2P, and the segmental power of the corrected signal hlo2 is hlo2P. P Let's say.

[0066] The adjustment unit 231 adjusts the power of the correction signals hls2 and hlo2 so as to maintain the power ratio (energy ratio) between the left and right. The adjustment unit 231 amplifies the correction signals so that the power ratio before and after correction matches. For example, the adjustment unit 231 multiplies the amplitude value of the correction signal by a predetermined number. The predetermined number for the correction signal hls2 is (hlsP / hlsP2), and the predetermined number for the correction signal hlo2 is (hloP / hloP2).

[0067] The correction signals hls2 and hlo2 after adjusting the power ratio are called correction filters hls3 and hlo3. The product of the amplitude value of the correction signal hls2 and a predetermined number (hlsP / hlsP2) becomes the amplitude value of the correction filter hls3. The product of the amplitude value of the correction signal hlo2 and a predetermined number (hloP / hloP2) becomes the amplitude value of the correction filter hlo3. Therefore, the segmental power of the correction filter hls3 becomes the same as the segmental power of the picked-up signal hls. The segmental power of the correction filter hlo3 becomes the same as the segmental power of the picked-up signal hlo.

[0068] In this way, an appropriate correction filter can be generated. That is, the processing device 201 can generate a correction filter according to the playback device. The correction filters hls3 and hlo3 are set as spatial acoustic filters in the convolution calculation units 11 and 12 shown in FIG. 1. This allows the out-of-head localization processing device 100 to perform playback with a high out-of-head localization effect.

[0069] Specifically, the correction characteristics are generated so that the level falls within the level range X appropriate for the playback device. This makes it possible to perform measurements and out-of-head localization processing in a state appropriate for the playback device. This makes it possible to generate a filter suitable for out-of-head localization processing.

[0070] Furthermore, in the above embodiment, the adjustment unit 231 adjusts the left-right balance. Out-of-head localization reproduction with good left-right balance can be realized. Of course, the adjustment of the power balance by the adjustment unit 231 can be omitted. For example, when the processing device 201 processes a single picked-up sound signal hls, the processing by the adjustment unit 231 is omitted. In this case, the correction signal hls2 is set as it is in the convolution calculation unit 11 as a correction filter.

[0071] The processing device 201 can also perform similar processing on the collected sound signals indicating the spatial acoustic transfer characteristics Hro and Hrs. In this case, the filter generation unit 230 adjusts the correction signal so that the segmental power ratio of the collected sound signals indicating the spatial acoustic transfer characteristics Hro and Hrs is maintained before and after the correction. Furthermore, the processing device 201 can also perform similar processing on the ear canal transfer characteristics of both ears. In the processing device 201, the filter generation unit 230 adjusts the correction signal so that the segmental power ratio between the ear canal transfer characteristic ECTFL of the left ear and the ear canal transfer characteristic ECTFR of the right ear is maintained before and after the correction.

[0072] Next, a filter generation method according to this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the filter generation method.

[0073] First, the measurement device 200 measures the transfer characteristics using an impulse sound or the like (S101). That is, the measurement signal generation unit 211 outputs a measurement signal such as an impulse sound from the left speaker 5L. The collected sound signal acquisition unit 212 acquires the collected sound signal from the microphone unit 2 (S102). The collected sound signal acquisition unit 212 cuts out the collected sound signal from the left microphone 2L and the collected sound signal from the right microphone 2R using a predetermined filter length. As a result, the collected sound signals hls and hlo are obtained.

[0074] The segmental power acquisition unit 215 calculates the segmental power of each of the picked-up sound signals hls and hlo (S103). The frequency characteristic acquisition unit 221 performs a Fourier transform on the picked-up sound signal (S104). This obtains the frequency characteristics Fhls and Fhlo. The frequency characteristics are frequency amplitude characteristics (amplitude spectrum), but may also be frequency power characteristics (power spectrum).

[0075] The level calculation unit 223 calculates the reference level (S105). As described above, the reference level is the average value of the amplitude values ​​of the two frequency characteristics Fhls and Fhlo. Furthermore, the level calculation unit 223 calculates the maximum and minimum levels of the frequency characteristics Fhls and Fhlo. The reference level, maximum level, and minimum level may be calculated from the amplitude values ​​of the entire band, or may be calculated from the amplitude values ​​of some of the bands.

[0076] Further, the level range setting unit 224 sets the level range to be compressed (S106). The level range is set depending on the model, performance, etc. of the playback device. For example, the user or a staff member responsible for filter generation may input the level range X. Then, the correction unit 225 compresses and corrects the frequency characteristics Fhls and Fhlo so that the amplitude values ​​of the frequency characteristics Fhls and Fhlo fall within the level range X that includes the reference level (S107). As a result, the correction characteristics NewFhls and NewFhlo are obtained. The amplitude values ​​of the correction characteristics NewFhls and NewFhlo are included in the level range X.

[0077] Next, the inverse transform unit 232 performs an inverse Fourier transform on the correction characteristics (S108). In the inverse Fourier transform, the frequency amplitude characteristics are the correction characteristics, and the frequency phase characteristics are the frequency phase characteristics calculated by the Fourier transform in S104. As a result, correction signals hls2 and hlo2 in the time domain are obtained.

[0078] The adjustment unit 231 adjusts the amplitude levels of the correction signals hls2 and hlo2 so as to maintain the segmental power ratio of the picked-up sound signals hls and hlo (S109). Specifically, the adjustment unit 231 multiplies the correction signals hls2 and hlo2 by a predetermined number according to the segmental power ratio, respectively. This results in correction filters hls3 and hlo3. The adjustment unit 231 adjusts the power ratio, making it possible to generate filters with good left-right balance.

[0079] (Correction processing example 1) Next, an example of the correction step of step S107 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing a processing example 1 of the correction process by the correction unit 225.

[0080] First, the correction unit 225 determines whether the level difference of the frequency amplitude characteristics is equal to or greater than the level range X (S201). The level difference is the level difference (maxL-minL) between the maximum value (maximum level maxL) and the minimum value (minimum level minL). The maximum level and minimum level may be the maximum and minimum values ​​of the frequency amplitude characteristics in the entire band, or may be the maximum and minimum values ​​of a part of the band.

[0081] If the level difference is smaller than the level range X (NO in S201), the correction unit 225 ends the process without making any correction. The above In the case of S201 YES ), the correction unit 225 compresses the level (amplitude value) of each frequency toward the reference level (S202). As a result, the frequency characteristics are corrected so that the level at each frequency falls within the level range X.

[0082] FIG. 7 is a graph showing frequency amplitude characteristics before and after correction in Processing Example 1. That is, FIG. 7 shows the amplitude spectrum of the frequency characteristics Fhls before correction and the correction characteristics NewFhls. As shown in FIG. 7, the frequency amplitude characteristics after correction fall within a level range X centered around a reference level A. In FIG. 7, the reference level A=-9.4 dB and the level range X=20 dB. Furthermore, in FIG. 7, the correction band is set to 10 Hz to 20 kHz.

[0083] (Correction processing example 2) Next, another example of the correction step of step S107 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing processing example 2 of the correction process by the correction unit 225. In processing example 2, the correction unit 225 corrects only levels (amplitude values) that are greater than the reference level.

[0084] First, the correction unit 225 determines whether the level difference of the frequency amplitude characteristics is equal to or greater than the level range X (S301). The level difference is the difference value (maxL-minL) between the maximum value (maximum level maxL) and the minimum value (minimum level minL). The maximum level and minimum level may be the maximum and minimum values ​​of the frequency amplitude characteristics in the entire band, or may be the maximum and minimum values ​​of a part of the band.

[0085] If the level difference is smaller than the level range X (NO in S301), the correction unit 225 ends the process without making correction. If the difference is larger than the level range X (YES in S301), the correction unit 225 compresses only the levels (amplitude values) of each frequency that are larger than the reference level toward the reference level (S302). The correction unit 225 lowers levels that are higher than the reference level.

[0086] In processing example 2, the corrector 225 does not perform correction for levels lower than the reference level. Therefore, for frequencies lower than the reference level, the amplitude values ​​before and after correction match.

[0087] Furthermore, in processing example 2, the correction unit 225 corrects only levels higher than the reference level, but it may also correct only levels lower than the reference level. In other words, in processing example 2, the correction unit 225 corrects only one of levels higher than the reference level and levels lower than the reference level. It is sufficient for the correction unit 225 to correct the frequency characteristics only at levels equal to or higher than the reference level or levels equal to or lower than the reference level.

[0088] FIG. 9 is a graph showing frequency amplitude characteristics before and after correction in processing example 2. In FIG. 9, the reference level A is −9.4 dB and the level range X is 20 dB. Furthermore, in FIG. 9, the correction band is set to 10 Hz to 20 kHz. As shown in FIG. 9, for amplitude values ​​higher than the reference level A, the frequency amplitude characteristics after correction fall within the level range X. In this case, for levels lower than the reference level A, it is possible that the frequency amplitude characteristics will not fall within the level range X. In other words, in processing example 2, the frequency amplitude characteristics fall within a level range equal to or higher than the min level and equal to or lower than (A+(X / 2)).

[0089] (Correction processing example 3) In processing example 3, the frequency axis of the frequency amplitude characteristic is a logarithmic scale. The reason for converting the frequency axis to a logarithmic scale will be explained. It is generally said that human sensory quantities are converted to logarithms. For this reason, it is important to consider the frequency of audible sounds on a logarithmic scale as well. By converting the scale, data is spaced equally in the above sensory quantities, so that data can be treated equally in all frequency bands. As a result, mathematical operations, frequency band division and weighting become easier, and stable results can be obtained. Note that the frequency characteristic acquisition unit 221 is not limited to the logarithmic scale, but can also convert to a scale close to human hearing (called the auditory scale). Frequency response The auditory scale may be converted into a logarithmic scale, a mel scale, a Bark scale, an ERB (Equivalent Rectangular Bandwidth) scale, or the like.

[0090] The frequency characteristic acquisition unit 221 performs scale conversion of the spectral data on the auditory scale by data interpolation. For example, the frequency characteristic acquisition unit 221 interpolates data in the low frequency band, which has coarse data intervals on the auditory scale, to make the data in the low frequency band dense. Data that is evenly spaced on the auditory scale becomes data that is dense in the low frequency band and coarse in the high frequency band on a linear scale. In this way, the frequency characteristic acquisition unit 221 can generate axis-converted data that is evenly spaced on the auditory scale. Of course, the axis-converted data does not have to be data that is completely evenly spaced on the auditory scale. In this way, the correction unit 225 and the like perform processing on the frequency amplitude characteristics on the logarithmic scale. Furthermore, in order to match the frequency phase characteristics and the number of samples, the frequency axis may be returned to a linear scale before inverse conversion.

[0091] (Correction processing example 4) In processing example 4, the correction unit 225 does not correct the entire correction band, but corrects the amplitude value only at frequencies around the peak that exceeds the upper limit of the level range X. Processing example 4 will be described with reference to Fig. 10. Fig. 10 is a flowchart showing processing example 4.

[0092] First, the correction unit 225 determines whether the level difference of the frequency amplitude characteristics is equal to or greater than the level range X (S401). The level difference is the difference value (maxL-minL) between the maximum value (maximum level maxL) and the minimum value (minimum level minL). The maximum level and minimum level may be the maximum and minimum values ​​of the frequency amplitude characteristics in the entire band, or may be the maximum and minimum values ​​of a part of the band.

[0093] If the level difference is smaller than the level range X (NO in S401), the correction unit 225 ends the process without making correction. The above Case (S 401 YES ), the correction unit 225 compresses the amplitude value toward the reference level around the peak frequency that is the peak exceeding the upper limit value (A+X / 2) of the range (S402).

[0094] For example, the correction unit 225 determines crossover frequencies that cross the upper limit before and after the peak frequency. The correction unit 225 calculates a first crossover frequency that is lower than the peak frequency and a second crossover frequency that is higher than the peak frequency. The correction unit 225 compresses the amplitude value toward a reference level in a frequency band defined by the first crossover frequency and the second crossover frequency.

[0095] Specifically, the correction unit 225 determines a first crossover frequency that intersects with the upper limit of the range on the lower frequency side of the peak frequency. The correction unit 225 determines a second crossover frequency that intersects with the upper limit of the range on the higher frequency side of the peak frequency. The correction unit 225 corrects the amplitude values ​​in the frequency band from the first crossover frequency to the second crossover frequency. In this way, it is possible to correct amplitude values ​​that exceed the upper limit of the range around the peak.

[0096] Fig. 11 is a graph showing three frequency bands (a) to (c) defined by crossover frequencies. Frequency band (a) is a frequency band that includes a first peak P1. In other words, frequency band (a) is defined by crossover frequencies before and after the first peak P1. Frequency band (b) is a frequency band that includes a second peak P2. Frequency band (c) is a frequency band that includes a third peak P3. Furthermore, as shown in Fig. 11, one frequency band may include multiple peaks that are close to each other.

[0097] In this way, in processing example 4, only amplitude values ​​that exceed only the upper limit of the range are compressed toward the reference level. Furthermore, the correction unit 225 may correct amplitude values ​​at frequencies around a dip below the lower limit (A-(X / 2)) of the level range X. In this case, too, the correction unit 225 determines crossover frequencies that intersect with the lower limit before and after a dip below the lower limit. The correction unit 225 simply compresses amplitude values ​​in a frequency band defined by the two crossover frequencies. Of course, the correction unit 225 may compress amplitude values ​​in both a frequency band including a peak and a frequency band including a dip. Alternatively, the correction unit 225 may compress amplitude values ​​only in a frequency band including a peak, or only in a frequency band including a dip.

[0098] (Correction processing example 5) In processing example 5, the correction unit 225 performs correction using a different method. Specifically, the level of the amplitude value is corrected using smoothing processing such as moving average. The frequency characteristics (spectral data) are smoothed using methods such as moving average, Savitzky-Golay filter, smoothing spline, cepstrum transform, and cepstrum envelope. The correction unit 225 performs smoothing processing on the frequency characteristics to correct them so that they fall within the level range X.

[0099] (Correction processing example 6) In processing example 6, a collected signal for ear canal transfer characteristics is processed. That is, the measurement is performed by the measuring device 300 shown in FIG. 3. Specifically, in the processing device 201 shown in FIG. 4, the measurement signal generation unit 211 outputs a measurement signal to the headphones 43 instead of the speaker 5L. In this case, the left and right microphones 2L and 2R collect collected signals indicating the ear canal transfer characteristics of the left and right ears. The frequency amplitude characteristics are acquired. The reference level, maximum level, and minimum level are acquired from the two frequency amplitude characteristics. The contents other than those described above are the same as those of the above-mentioned embodiments and processing examples, so a description thereof will be omitted.

[0100] (Correction processing example 7) In the processing example 7, multi-channel speakers such as 5.1ch and 7.1ch are used, and the adjustment unit 231 performs adjustment so that the power ratio of the picked-up signals is maintained for each channel.

[0101] In a 5.1ch multi-channel system, left and right front speakers, left and right rear speakers, a center speaker, and a subwoofer are used. In this case, the adjustment unit 231 adjusts the correction signals so that the power ratio between the front speakers and the rear speakers is maintained. Specifically, the adjustment unit 231 multiplies each correction signal by a coefficient that makes the segmental power ratio the same before and after correction.

[0102] Specifically, the measurement device 200 sequentially performs measurements using speakers of different channels. For example, the measurement signal generator 211 generates measurement signals and outputs them sequentially to the speakers of each channel. The collected signal acquirer 212 acquires acquired signals by sequentially collecting measurement signals from the speakers of each channel. The frequency characteristic acquirer 221 acquires multiple frequency characteristics based on the collected signals obtained by collecting the measurement signals output from the speakers of different channels.

[0103] The segmental power acquisition unit 215 calculates the left and right segmental powers of the picked-up signal for each channel. The adjustment unit 231 adjusts the level of the correction signal so as to maintain the power ratio. In this way, a filter with good balance between channels can be generated. Note that the level range X may be different for each channel or may be the same for all channels.

[0104] The process of maintaining the power ratio between channels is not limited to multi-channels such as 5.1ch, but can also be applied to the 2ch measurement device shown in Fig. 2. For example, measurements may be performed on the left and right speakers, and adjustments may be made to maintain the power ratio.

[0105] The above processing examples 1 to 7 can be combined as appropriate. For example, when correcting the amplitude value around the peak frequency or the dip frequency as in processing example 4, the correction unit 225 may use the axis conversion process of the frequency axis in processing example 3 or the smoothing process in processing example 5.

[0106] As described above, according to this embodiment, the frequency characteristics are corrected so that they fall within a predetermined level range X including the reference level. Therefore, it is possible to reproduce a filter that can obtain an appropriate out-of-head localization effect even with various playback devices, equipment, and measurement environments. In other words, it is possible to automatically correct a filter that does not clip signals that have been processed for out-of-head localization. It is possible to perform out-of-head localization listening in accordance with the speakers, headphones, and measurement environment that suit the user's preferences. Furthermore, automatic correction is possible in accordance with the playback device.

[0107] Embodiment 2 An apparatus and method according to the second embodiment will be described with reference to Fig. 12. Fig. 12 is a block diagram showing the configuration of a processing device 201. One of the technical features of the second embodiment is the process of setting a level range X. Therefore, in the processing device 201 shown in Fig. 12, a determination unit 242 is added to the configuration of Fig. 4. The configuration and process other than the determination unit 242 are the same as those of the first embodiment, and therefore the description will be omitted where appropriate.

[0108] The determination unit 242 determines the performance of the playback device. For example, the determination unit 242 evaluates the performance of the amplifier of the playback device. . judgment In accordance with the determination result of the setting unit 242, the level range setting unit 224 sets the level range X. The correction unit 225 calculates the correction characteristics by correcting the frequency characteristics based on the level range X. The filter generation unit 230 generates a correction filter based on the correction characteristics.

[0109] For example, the determination unit 242 can make a determination based on the frequency characteristics acquired by the frequency characteristic acquisition unit 221. The determination unit 242 detects the level difference (maxL-minL) between the maximum level (maxL) and the minimum level (minL) of the frequency amplitude characteristics. The determination unit 242 acquires the output level (output sound pressure level) and the S / N ratio of the playback device based on the level difference. Then, the determination unit 242 determines the performance based on the output level or the S / N ratio. The determination unit 242 may determine the level range X according to the level difference between the maximum level and the minimum level of the frequency amplitude characteristics.

[0110] For example, for a playback device with a large level difference, the level range X is set to about 80% of the level difference. The determination unit 242 sets the variable to 0.8. For a playback device with a small level difference, the level range X is set to about 40% of the level difference. The level range setting unit 224 sets the level range X by multiplying the level difference by a variable according to the determination result.

[0111] Furthermore, the processing device 201 can set the level range X without using a variable. For example, the determination unit 242 calculates the level difference (maxL-minL) in a portion of the determination band. The determination band can be, for example, 100 Hz to 8 kHz. That is, the determination unit 242 finds the maximum level (maxL) and minimum level (minL) in the range of 100 Hz to 8 kHz. Then, the determination unit 242 makes a determination based on the level difference (maxL-minL). Alternatively, the determination unit 242 may have a conversion formula or a conversion table for converting the level difference into the level range X.

[0112] In this way, the determination unit 242 makes a determination based on the frequency characteristics of the picked-up signal obtained by measurement using a playback device. The determination unit 242 makes a determination based on the level difference between the maximum level and the minimum level of the frequency characteristics.

[0113] Alternatively, the determination unit 242 may acquire playback device information about the playback device and determine the performance based on the playback device information. Then, the level range setting unit 224 sets the level range X according to the performance of the playback device. For example, if the amplifier of the playback device is high performance, the level range setting unit 224 sets X=40 dB. If the amplifier is low performance, the level range setting unit 224 sets X=20 dB. Of course, the determination by the determination unit 242 is not limited to two levels, high performance and low performance, and may be three or more levels.

[0114] Furthermore, the determination unit 242 may have a table indicating the performance for each model number of the playback device. The determination unit 242 acquires playback device information indicating the model number of the playback device. The determination unit 242 determines the performance according to the model number of the playback device. The playback device information regarding the playback device may be acquired automatically or may be input by the user, for example. For example, in the case of a playback device connected via Bluetooth, the determination unit 242 can automatically acquire information regarding the playback device.

[0115] For example, the measurement device 200 or the measurement device 300 performs measurements to acquire frequency characteristics for each playback device in advance. Then, as described above, the determination unit 242 determines performance according to the level difference of the frequency characteristics and stores the determination results in a table. The determination unit 242 can then make a determination by referring to the table.

[0116] The playback device may be the speakers 5L and 5R and their amplifiers shown in Fig. 2, or the headphones 43 shown in Fig. 3. In other words, the playback device may be a playback device used during measurement. Alternatively, it may be the headphones 43 in the out-of-head localization processing device shown in Fig. 1. In other words, the playback device may be the headphones 43 or earphones used during out-of-head localization listening. In the second embodiment, any one or more of the above processing examples 1 to 7 may be used.

[0117] As described above, according to this embodiment, it is possible to automatically set the level range X according to the performance of the playback device. Then, the correction unit 225 performs correction based on the level range X. Therefore, it is possible to reproduce a filter that can obtain an appropriate out-of-head localization effect even with various playback devices, equipment, and measurement environments. In other words, it is possible to automatically correct a filter that does not clip the signal that has been processed for out-of-head localization. It is possible to perform out-of-head localization listening according to the speaker, headphones, and measurement environment that suit the user's preferences. Furthermore, automatic correction according to the playback device is possible.

[0118] Some or all of the above processes may be performed by a computer program. The above-mentioned program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.

[0119] The invention made by the inventor has been specifically described above based on an embodiment, but it goes without saying that the present invention is not limited to the above embodiment and can be modified in various ways without departing from the gist of the invention. [Explanation of symbols]

[0120] U User 1 Person to be measured 2 microphone units 2L Left microphone 2R Right Microphone 5 stereo speakers 5L Left speaker 5R Right speaker 10. Out-of-head localization processing unit 11 Convolution operation unit 12 Convolution operation unit 21 Convolution operation unit 22 Convolution operation unit 24 Adder 25 Adder 41 Filter section 42 Filter section 43 Headphones 200 Measuring Equipment 201 Processing equipment 211 Measurement signal generation unit 212 Sound signal acquisition unit 215 Segmental Power Acquisition Department 221 Frequency characteristic acquisition unit 223 Level Calculation Unit 224 Level range setting section 225 Correction Unit 230 Filter Generation Unit 231 Adjustment section 232 Inverse conversion unit 242 Judgment section

Claims

1. A frequency characteristic acquisition unit that acquires frequency characteristics based on a sound signal picked up by a microphone; a level calculation unit that calculates a reference level in the frequency characteristic; a correction unit that calculates a correction characteristic by correcting the frequency characteristic so that the frequency characteristic falls within a predetermined level range including the reference level; a filter generation unit that generates a spatial acoustic filter or an inverse filter used in out-of-head localization processing based on the correction characteristics, The frequency characteristic acquisition unit acquiring a first frequency characteristic based on a first picked-up signal picked up by a left microphone attached to the left ear of the user; acquiring a second frequency characteristic based on a second picked-up signal picked up by a right microphone attached to the right ear of the user; the level calculation unit calculates a common level for the first frequency characteristic and the second frequency characteristic; the correction unit calculates a first correction characteristic obtained by correcting the first frequency characteristic and a second correction characteristic obtained by correcting the second frequency characteristic; The filter generation unit generating a first correction signal and a second correction signal in the time domain by inversely transforming the first correction characteristic and the second correction characteristic, respectively; A filter generating device that adjusts the levels of the first correction signal and the second correction signal so as to maintain the left-right power ratio before and after correction.

2. the frequency characteristic acquisition unit acquires a plurality of frequency characteristics based on collected signals obtained by sequentially collecting measurement signals output from speakers of different channels; The filter generating device according to claim 1 , wherein the level of the correction signal is adjusted so as to maintain a power ratio of the picked-up signals for each channel.

3. A frequency characteristic acquisition unit that acquires frequency characteristics based on a sound signal picked up by a microphone; a level calculation unit that calculates a reference level in the frequency characteristic; a correction unit that calculates a correction characteristic by correcting the frequency characteristic so that the frequency characteristic falls within a predetermined level range including the reference level; a filter generation unit that generates a spatial acoustic filter or an inverse filter used in out-of-head localization processing based on the correction characteristics, The correction unit corrects the frequency characteristics only at levels equal to or higher than the reference level or at levels equal to or lower than the reference level.

4. A step of acquiring frequency characteristics based on a sound signal picked up by a microphone; calculating a reference level in the frequency characteristic; calculating a correction characteristic by correcting the frequency characteristic so that the frequency characteristic falls within a predetermined level range including the reference level; generating a spatial acoustic filter or an inverse filter used in out-of-head localization processing based on the correction characteristics; In the step of acquiring the frequency characteristics, acquiring a first frequency characteristic based on a first picked-up signal picked up by a left microphone attached to the left ear of the user; acquiring a second frequency characteristic based on a second picked-up signal picked up by a right microphone attached to the right ear of the user; In the step of calculating the reference level, a common level is calculated for the first frequency characteristic and the second frequency characteristic; In the correcting step, a first correction characteristic obtained by correcting the first frequency characteristic and a second correction characteristic obtained by correcting the second frequency characteristic are calculated, In the step of generating the spatial acoustic filter or the inverse filter, generating a first correction signal and a second correction signal in the time domain by inversely transforming the first correction characteristic and the second correction characteristic, respectively; A filter generating method, comprising adjusting levels of the first correction signal and the second correction signal so as to maintain a power ratio between left and right before and after correction.

5. A program for causing a computer to execute a filter generation method, comprising: The filter generation method includes: acquiring frequency characteristics based on a sound signal picked up by a microphone; calculating a reference level in the frequency characteristic; calculating a correction characteristic by correcting the frequency characteristic so that the frequency characteristic falls within a predetermined level range including the reference level; generating a spatial acoustic filter or an inverse filter used in out-of-head localization processing based on the correction characteristics; In the step of acquiring the frequency characteristics, acquiring a first frequency characteristic based on a first picked-up signal picked up by a left microphone attached to the left ear of the user; acquiring a second frequency characteristic based on a second picked-up signal picked up by a right microphone attached to the right ear of the user; In the step of calculating the reference level, a common level is calculated for the first frequency characteristic and the second frequency characteristic; In the correcting step, a first correction characteristic obtained by correcting the first frequency characteristic and a second correction characteristic obtained by correcting the second frequency characteristic are calculated, In the step of generating the spatial acoustic filter or the inverse filter, generating a first correction signal and a second correction signal in the time domain by inversely transforming the first correction characteristic and the second correction characteristic, respectively; a program for adjusting levels of the first correction signal and the second correction signal so as to maintain a power ratio between left and right before and after correction;

Citation Information

Patent Citations

  • Sound processing method

    JP2010141396A

  • Sound reproducing apparatus

    JP2012054863A

  • Out-of-head sound localization processing device and out-of-head sound localization processing method

    JP2017060040A

  • Processing device, processing method, and program

    JP2019062430A

  • Processing device, processing method, regeneration process, and program

    JP2020136752A