Filter generation device

The filter generation device addresses the limitations of out-of-head sound localization by generating filters that adapt to playback devices and environments, ensuring effective sound localization and balanced reproduction.

JP2025186541APending Publication Date: 2025-12-23JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025166213
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing out-of-head sound localization technologies are limited by the specific characteristics of playback devices and environments, leading to issues such as steep peaks and dips in frequency characteristics that can cause clipping and require adjustments for individual user equipment and environments.

Method used

A filter generation device that includes a frequency characteristic acquisition unit, a determination unit, a level range setting unit, and a correction unit to generate filters suitable for out-of-head localization processing, adjusting frequency characteristics to fit a predetermined level range and using spatial acoustic filters and inverse filters to cancel headphone characteristics.

Benefits of technology

Enables the generation of filters that adapt to different playback devices and environments, ensuring effective out-of-head sound localization without clipping, and maintaining balanced sound reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186541000001_ABST
    Figure 2025186541000001_ABST
Patent Text Reader

Abstract

To provide a filter generation device for generating a filter suitable for extracranial localization processing.SOLUTION: A processing device 201 includes a frequency characteristic acquisition unit 221 that acquires a frequency characteristic based on a collected sound signal collected by microphones 2L and 2R and interpolates low-frequency band data using a frequency axis as an auditory measure close to human auditory sense to generate axis conversion data at regular intervals in the auditory measure, a determination unit 242 that determines a performance of a reproduction device, a level range setting unit 224 that sets a level range according to the determination result of the determination unit 242, a correction unit 225 that calculates a correction characteristic by correcting the frequency characteristic based on the level range with respect to a frequency amplitude characteristic of an auditory measure, and a filter generation unit 230 that generates a spatial acoustic filter used in an out-of-head localization process or an inverse filter that cancels the characteristic from the reproduction device to an ear of a user based on the correction characteristic.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a filter generation device. [Background technology]

[0002] As a sound image localization technology, out-of-head sound localization is used to localize a sound image outside the listener's head using headphones. Out-of-head localization technology uses the characteristics from the headphones to the ears (headphone characteristics) Cancellation and two characteristics from one speaker (monaural speaker) to the ear (spatial acoustics) By providing a certain transfer characteristic, the sound image is positioned outside the head.

[0003] In out-of-head localization playback using stereo speakers, two channels (hereafter referred to as ch) are used. The measurement signal (impulse sound, etc.) emitted from the speaker was placed at the listener's ear. The measurement signal is recorded using a microphone (hereafter referred to as the microphone). The processing unit generates a filter based on the collected sound signal. By convolving this with the audio signal, out-of-head localization playback can be achieved.

[0004] Furthermore, a filter (also called an inverse filter) is used to cancel the characteristics from the headphones to the ears. ) to generate the characteristic from the headphones to the ear and eardrum (ear canal transfer function ECTF The ear canal transfer characteristics (also called ear canal transfer characteristics) are measured using a microphone placed in the listener's ear.

[0005] Patent Document 1 discloses an apparatus for performing out-of-head localization processing. The out-of-head localization processing applies DRC (Dynamic Range Comparator) to the playback signal. In DRC processing, the processing device smooths the frequency characteristics. Furthermore, the processing unit performs band division based on the smoothed characteristics. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Publication No. 2019-62430 Summary of the Invention [Problem to be solved by the invention]

[0007] In such out-of-head localization listening, it is desirable to process without being limited to a specific playback device. For example, even if the user uses headphones as a playback device, It is desirable to perform appropriate out-of-head localization processing. It is desirable to reproduce the spatial acoustic transmission characteristics in the environment in which the camera is installed as a playback device.

[0008] If you change the playback device, the transfer characteristics may change. The individual characteristics of the user (spatial acoustic transfer characteristics and ear canal conduction characteristics) were measured using the playback equipment used by the user. Even when measuring personal characteristics, it is preferable to measure the frequency characteristics. Steep peaks and dips occur, which can cause the out-of-head localization processed signal to clip. be.

[0009] Peaks and dips are caused by the characteristics of playback devices such as speakers and headphones, or by the measurement environment. The peaks and dips vary depending on the acoustic characteristics of the room. The levels and frequencies of the peaks and dips vary depending on the shape of the ear and the area of ​​the ear. It varies depending on the cause of the problem. Check the characteristics of the playback device and measurement environment. It becomes necessary to make adjustments according to the equipment, measurement environment, etc.

[0010] The present disclosure has been made in consideration of the above points, and provides a method for generating a filter suitable for out-of-head localization processing. The object of the present invention is to provide a filter generating device that can [Means for solving the problem]

[0011] The filter generation device according to this embodiment includes a frequency characteristic acquisition unit that acquires frequency characteristics based on a signal picked up by a microphone, and generates axis-transformed data at equal intervals on an auditory scale by interpolating low-frequency band data on a frequency axis that is an auditory scale similar to human hearing; a determination unit that determines the performance of a playback device; a level range setting unit that sets a level range in accordance with the determination result of the determination unit; a correction unit that calculates correction characteristics by correcting the frequency characteristics with respect to the frequency amplitude characteristics of the auditory scale based on the level range; and a filter generation unit that generates, based on the correction characteristics, a spatial acoustic filter used in out-of-head localization processing or an inverse filter that cancels the characteristics from the playback device to the user's ears.

[0012] The filter generation device according to this embodiment includes a frequency characteristic acquisition unit that acquires frequency characteristics based on a sound signal collected by a microphone, and generates axis-transformed data at equal intervals on an auditory scale by interpolating low-frequency band data on a frequency axis that is an auditory scale similar to human hearing; a level calculation unit that calculates a reference level in the frequency characteristics; a correction unit that calculates correction characteristics by correcting the frequency characteristics so that the frequency amplitude characteristics of the auditory scale fall within a predetermined level range including the reference level; and a filter generation unit that generates a spatial acoustic filter or an inverse filter used in out-of-head localization processing based on the correction characteristics. [Effects of the Invention]

[0013] According to the present disclosure, a filter generation method capable of generating a filter suitable for out-of-head localization processing is provided. An apparatus can be provided. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a block diagram showing an out-of-head localization processing device according to an embodiment of the present invention; [Figure 2] FIG. 1 is a diagram illustrating a configuration of a measurement device for measuring spatial acoustic transfer characteristics. [Figure 3] FIG. 1 is a diagram illustrating a configuration of a measurement device for measuring ear canal transfer characteristics. [Figure 4] FIG. 2 is a control block diagram showing the configuration of the processing device. [Figure 5] 1 is a flowchart illustrating a method for generating a filter in a processing device. [Figure 6] 10 is a flowchart showing a first example of correction processing. [Figure 7] 10 is a graph showing frequency amplitude characteristics before and after correction according to processing example 1. [Figure 8] 10 is a flowchart showing a second example of the correction process. [Figure 9] 10 is a graph showing frequency amplitude characteristics before and after correction according to processing example 2. [Figure 10] 10 is a flowchart showing a fourth example of the correction process. [Figure 11] 10 is a graph showing frequency bands according to processing example 4. [Figure 12] FIG. 10 is a block diagram showing a configuration of a processing device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] An outline of the sound image localization processing according to this embodiment will be described. The out-of-head localization processing uses the spatial acoustic transfer characteristics and ear canal transfer characteristics to perform out-of-head localization processing. The spatial acoustic transfer characteristics are the transfer characteristics from a sound source such as a speaker to the ear canal. The transfer characteristic is the transfer characteristic from the speaker unit of the headphones or earphones to the eardrum. In this embodiment, the spatial acoustic transmission characteristics are measured when headphones or earphones are not worn. The ear canal transfer characteristics are measured while wearing headphones or earphones. The measurement data is used to realize out-of-head localization processing. It is characterized by a microphone system for measuring the transfer characteristics or ear canal transfer characteristics.

[0016] The out-of-head localization processing according to this embodiment can be performed on a personal computer, a smartphone, a tablet, or the like. The user terminal is executed on a user terminal such as a laptop PC. Storage means such as memory or hard disk, display means such as LCD monitor, touch panel, buttons A user terminal is an information processing device having input means such as a keyboard and a mouse. The user terminal may have a communication function for transmitting and receiving data. An output means (output unit) having earphones is connected. The connection may be a wired or wireless connection.

[0017] Embodiment 1 (Extracranial stereotaxic processing device) 1 is a block diagram of an out-of-head localization processing device 100, which is an example of a sound field reproduction device according to this embodiment. The diagram is shown in FIG. 1. The out-of-head localization processing device 100 outputs the following to a user U wearing headphones 43: Therefore, the out-of-head localization processing device 100 receives left and right stereo inputs. The left and right stereo input signals X and XR are processed for sound image localization. L and XR are analog audio signals output from a CD (Compact Disc) player, etc. It is a playback signal or digital audio data such as mp3 (MPEG Audio Layer-3). The audio playback signal and digital audio data are collectively referred to as the playback signal. In other words, the Lch and Rch stereo input signals XL and XR are the playback signals.

[0018] The head localization processing device 100 is not limited to a single physical device. For example, some of the processing may be performed by a smartphone, etc. The remaining processing is done by a DSP (Digital Signal Processor) built into the headphones 43. It may be performed by

[0019] The head-outside localization processing device 100 includes a head-outside localization processing unit 10, a filter for storing an inverse filter Linv, and a The device includes a filter unit 41, a filter unit 42 that stores an inverse filter Rinv, and a headphone 43. The out-of-head localization processing unit 10, the filter unit 41, and the filter unit 42 are specifically This can be achieved by using a processor or the like.

[0020] The out-of-head localization processing unit 10 stores spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs. The convolution calculation unit 11-12, 21-22, and adders 24 and 25 are provided. The calculation units 11 to 12 and 21 to 22 perform convolution processing using spatial acoustic transfer characteristics. The localization processing unit 10 receives stereo input signals XL and XR from a CD player or the like. The out-of-head localization processing unit 10 is set with spatial acoustic transfer characteristics. is a filter of spatial acoustic transfer characteristics (hereinafter, The spatial acoustic transfer characteristics are calculated by convolving the sound waves generated by the subject's head and pinna. It can be a measured head-related transfer function (HRTF), a dummy head, or a third-party head-related transfer function. It's okay to have it.

[0021] A set of four spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs is used for spatial The data used for convolution in the convolution calculation units 11, 12, 21, and 22 are acoustic transfer functions. The spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are given. A spatial acoustic filter is generated by cutting out the signal at a fixed filter length.

[0022] The spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are impulse response measurements. For example, if a user U wears a microphone in each ear, The left and right speakers placed in front of the user U are used to measure the impulse response. The impulse sounds output from the speakers are The measurement signal is picked up by a microphone. Based on the signal picked up by the microphone, the spatial acoustic transfer characteristic Hl s, Hlo, Hro, Hrs are obtained. Spatial acoustic transfer between the left speaker and the left microphone. The spatial acoustic transfer characteristic Hls between the left speaker and the right microphone, Hlo between the right speaker and the left microphone, Spatial acoustic transfer characteristics Hro between the left speaker and the right microphone, and spatial acoustic transfer characteristics Hro between the right speaker and the right microphone Hrs is measured.

[0023] The convolution calculation unit 11 then performs spatial acoustic transfer on the Lch stereo input signal XL. The convolution calculation unit 11 performs convolution calculation using a spatial acoustic filter according to the characteristic Hls. The convolution operation unit 21 outputs the data to the adder 24. The convolution operation unit 21 converts the Rch stereo input signal XR A spatial acoustic filter corresponding to the spatial acoustic transfer characteristic Hro is convolved with the 21 outputs the convolution operation data to the adder 24. The adder 24 performs two convolution operations. The data is added and output to the filter unit 41.

[0024] The convolution calculation unit 12 calculates the spatial acoustic transfer characteristic Hl for the Lch stereo input signal XL. The convolution calculation unit 12 performs convolution calculation on the spatial acoustic filter corresponding to o. , and outputs it to the adder 25. The convolution calculation unit 22 calculates the Rch stereo input signal XR as follows: The convolution calculation unit 22 performs convolution with a spatial acoustic filter according to the spatial acoustic transfer characteristic Hrs. The adder 25 outputs the two convolution operation data. The data is added and output to the filter unit 42.

[0025] The filter sections 41 and 42 are for detecting headphone characteristics (characteristics between the headphone playback unit and the microphone). The inverse filters Linv and Rinv are set to cancel the noise. The reproduced signal (convolution signal) processed by the phase processor 10 is filtered by an inverse filter Linv. The filter unit 41 convolves the Lch signal from the adder 24 with Rinv. Similarly, the filter unit 42 is convoluted with the inverse filter Linv of the headphone characteristic. The Rch signal from 5 is convolved with the inverse filter Rinv of the Rch headphone characteristics. The inverse filters Linv and Rinv are used to detect the The microphone cancels the characteristics from the ear canal entrance to the eardrum. If so, you can place it anywhere.

[0026] The filter unit 41 outputs the processed Lch signal YL to the left unit 43L of the headphone 43. The filter unit 42 outputs the processed Rch signal YR to the right unit of the headphone 43. The user U is wearing the headphones 43. The headphones 43 are ch signal YL and Rch signal YR (hereinafter, Lch signal YL and Rch signal YR are collectively referred to as The headphone amplifier outputs a localized signal (also called a headphone signal) to the user U. The resulting sound image can be reproduced.

[0027] In this way, the out-of-head localization processing device 100 calculates the spatial acoustic transfer characteristics Hls, Hlo, Hro, The spatial acoustic filter according to Hrs and the inverse filter Linv and Rinv of the headphone characteristics are In the following explanation, the spatial acoustic transfer characteristics Hls and H Spatial acoustic filters according to lo, Hro, and Hrs, and an inverse filter Lin for headphone characteristics v and Rinv are collectively used as the out-of-head localization processing filter. In this case, the out-of-head localization filter is composed of four spatial acoustic filters and two inverse filters. The out-of-head localization processing device 100 then outputs a total of six out-of-head signals to the stereo playback signal. By performing convolution processing using a localization filter, out-of-head localization processing is performed. The position filter is preferably based on the user U's personal measurements. For example, The out-of-head localization filter is set based on the sound signal picked up by the microphone attached to the U's ear. are.

[0028] In this way, the spatial acoustic filter and the inverse filters Linv and Rinv of the headphone characteristics are These filters are for the playback signal (stereo input signal XL , XR), the out-of-head localization processing device 100 executes out-of-head localization processing. In this embodiment, one of the technical features is the process of generating a spatial acoustic filter. Specifically, in the process of generating a spatial acoustic filter, the level range of the frequency characteristics is compressed. It is decorated with:

[0029] (Spatial acoustic transfer characteristic measurement device) Measurement equipment for measuring spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs using Figure 2 2 is a schematic diagram of a measurement configuration for measuring a subject 1. In this case, the subject 1 is the same person as the user U in FIG. However, they may be different people.

[0030] As shown in FIG. 2, the measurement device 200 includes a stereo speaker 5 and a microphone unit 2. The stereo speakers 5 are installed in the measurement environment. The measurement environment can be a room, a sales store or showroom of an audio system, etc. A listening room with good lighting and acoustics is preferable.

[0031] In this embodiment, the processing device 201 of the measurement device 200 appropriately generates a spatial acoustic filter. The processing unit 201 performs calculations to generate the sound from, for example, a CD player. The processing device 201 includes a personal computer (PC), The processing device 201 may be a tablet terminal, a smartphone, or the like. It may be itself.

[0032] The stereo speaker 5 includes a left speaker 5L and a right speaker 5R. The left speaker 5L and the right speaker 5R are installed in front of the attendant 1. The speaker 5R outputs impulse sounds for measuring impulse responses. In the embodiment, the number of speakers serving as sound sources will be explained as two (stereo speakers). The number of sound sources used for measurement is not limited to two, but may be one or more. Or, in the so-called multi-channel environment such as 5.1ch, 7.1ch, etc., this implementation The form can be applied.

[0033] The microphone unit 2 is a stereo microphone having a left microphone 2L and a right microphone 2R. The left microphone 2L is placed at the left ear 9L of the subject 1, and the right microphone 2R is placed at the left ear 9L of the subject 1. Specifically, the ear canal entrances of the left ear 9L and right ear 9R are connected to the eardrum. It is recommended to place microphones 2L and 2R at the position shown in Fig. 1. The measurement signal output from the speaker 5 is picked up and the picked-up signal is obtained. The subject 1 may be a person or a dummy head. That is, in this embodiment, the subject 1 includes not only a human but also a dummy head. It is a concept.

[0034] As shown above, the impulse sound output from the left speaker 5L and the right speaker 5R is recorded by microphone 2. The impulse response is measured by measuring L and 2R. The sound pickup signal obtained by response measurement is stored in memory, etc. This allows the left speaker 5L and the left microphone 2L, and the spatial acoustic transfer characteristic Hls between the left speaker 5L and the right microphone 2R. The spatial acoustic transfer characteristic Hlo between the right speaker 5R and the left microphone 2L is ro, and the spatial acoustic transfer characteristic Hrs between the right speaker 5R and the right microphone 2R is measured. That is, the measurement signal output from the left speaker 5L is picked up by the left microphone 2L, and the spatial The acoustic transfer characteristic Hls is acquired. The measurement signal output from the left speaker 5L is input to the right microphone 2 The spatial acoustic transfer characteristic Hlo is obtained by picking up the sound from the right speaker 5R. The left microphone 2L picks up the measurement signal thus obtained, and the spatial acoustic transfer characteristic Hro is obtained. The measurement signal output from the right speaker 5R is picked up by the right microphone 2R, and spatial acoustic transmission is performed. The characteristic Hrs is obtained.

[0035] Furthermore, the measuring device 200 outputs left and right signals from the left and right speakers 5L and 5R based on the collected sound signals. Spatial acoustic transmission characteristics Hls, Hlo, Hro, Hrs for microphones 2L and 2R For example, the processing unit 201 may generate an acoustic filter. The processor 201 extracts Hlo, Hro, and Hrs with a predetermined filter length. The inter-sound transmission characteristics Hls, Hlo, Hro, and Hrs may be corrected.

[0036] In this way, the processing device 201 performs the convolution operation of the out-of-head localization processing device 100. As shown in FIG. 1, the out-of-head localization processing device 100 generates the spatial acoustic filter to be used. The spatial acoustic transfer characteristic Hl between the left and right speakers 5L, 5R and the left and right microphones 2L, 2R is The out-of-head localization process is performed using spatial acoustic filters according to s, Hlo, Hro, and Hrs. That is, by convolving a spatial acoustic filter into the audio playback signal, out-of-head localization processing is performed. conduct.

[0037] The processing device 201 performs the processing for each of the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs. The same processing is performed on the corresponding picked-up signal. The same processing is performed on the four collected signals corresponding to Hlo, Hro, and Hrs. This allows for spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs to be Inter-acoustic filters can be generated respectively.

[0038] (Ear canal transfer characteristic measuring device) The ear canal transfer characteristic measuring device 300 will be described with reference to FIG. 3. The measurement device 300 is configured to measure the transfer characteristics of the inverse filter. The measurement device 300 measures the ear canal transfer characteristics to generate the sound. The device is equipped with headphones 43 and a processing device 301. Here, the subject 1 Although the user is the same person as user U in FIG. 1, the user may be a different person.

[0039] In this embodiment, the processing device 301 of the measurement device 300 selects a filter according to the measurement result. The processing unit 301 performs calculations to appropriately generate the These include personal computers (PCs), tablet devices, smartphones, etc., and are equipped with memory and processors. The memory stores processing programs, various parameters, measurement data, and so on. The processor executes the processing program stored in the memory. Each process is executed by executing the RAM. Processing Unit), FPGA (Field-Programmable Gate Array), DSP (Digital Si gnal Processor), ASIC (Application Specific Integrated Circuit), or GPU (Graphics Processing Unit), etc.

[0040] Also, the processing device 301 in FIG. 3 is physically the same processing device as the processing device 301 in FIG. In other words, the measurements in Figures 2 and 3 may be performed using the same processing device. For example, the measurement shown in FIG. The measurement shown in FIG. 3 is performed by a dedicated measurement processing device 201 installed in a room or the like. The processing may be performed by a general-purpose processing device 301 such as a computer.

[0041] The processing device 301 is connected to a microphone unit 2 and a headphone 43. The microphone unit 2 may be built into the headphones 43. The left microphone 2L is connected to the left ear 9L of the user U. The right microphone 2R is attached to the right ear 9R of the user U. The processing device 301 The processing device may be the same as the extracranial localization processing device 100 or may be a different processing device. It is also possible to use earphones instead of the headphones 43.

[0042] The headphones 43 are composed of a headphone band 43B, a left unit 43L, and a right unit 43 The headphone band 43B has a left unit 43L and a right unit 43R. The left unit 43L outputs sound to the left ear 9L of the user U. The headphones 43R output sound toward the right ear 9R of the user U. The headphones 43 are closed type, open type, Any type of headphones may be used, including open, semi-open, or semi-closed. With the terminal 2 attached to the user U, the user U puts on the headphones 43. The left ear 9L and right ear 9R are fitted with left microphone 2L and right microphone 2R, respectively. The headphone band 43B is attached to the left unit 43L and the right unit 43R. The biasing force that presses the knit 43L and the right unit 43R against the left ear 9L and the right ear 9R, respectively, is occurs.

[0043] The left microphone 2L picks up the sound output from the left unit 43L of the headphones 43. The microphone 2R picks up the sound output from the right unit 43R of the headphones 43. The microphone parts of the left microphone 2L and the right microphone 2R are placed at sound collection positions near the external ear canal. The left microphone 2L and the right microphone 2R are configured so as not to interfere with the headphones 43. That is, the left microphone 2L and the right microphone 2R are placed at appropriate positions on the left ear 9L and the right ear 9R. In this state, the user U can wear the headphones 43 .

[0044] The processing device 301 outputs a measurement signal to the headphones 43. The left unit 43L generates impulse sounds. Impulse sound is measured with the left microphone 2L. Impulse sound output from the right unit 43R The right microphone 2R measures the sound. When the measurement signal is output, microphones 2L and 2R pick up the sound. By doing so, an impulse response measurement is performed.

[0045] The processing device 301 performs the same processing on the collected signals from the microphones 2L and 2R. Then, the inverse filters Linv and Rinv are generated.

[0046] (Level range compression) In at least one of the measurement devices 200 and 300, the frequency characteristics of the picked-up signal are The range is compressed so that it fits within a predetermined level range. In the measurement device 200, the frequencies of the picked-up signals corresponding to the spatial acoustic transfer characteristics Hls and Hlo are The process of compressing the level range of the characteristics will be described. The level range of the frequency characteristics of the sound pickup signal corresponding to the spatial acoustic transfer characteristics Hro and Hrs is The compression process is the same as the process described below, and therefore the description will be omitted where appropriate. Similarly, in the measuring device 300, the frequency characteristics of the picked-up signal with respect to the transfer characteristics of the left and right ear canals are The process of compressing the level range is similar to the process described below, so The explanation will be omitted.

[0047] 4 is a block diagram showing the configuration of the processing device 201 of the measurement device 200. The device 201 includes a measurement signal generating unit 211, a picked-up signal acquiring unit 212, and a segmental power Acquisition unit 215, frequency characteristic acquisition unit 221, level calculation unit 223, and level range setting unit The image processing system includes a conversion unit 224, a correction unit 225, an adjustment unit 231, and an inverse conversion unit 232. The conversion unit 232 and the adjustment unit 231 function as a filter generation unit 230 .

[0048] The measurement signal generator 211 includes a D / A converter, an amplifier, and the like, and measures the spatial acoustic transfer characteristic The measurement signal is used to measure the impulse response and ear canal transfer characteristics. These include the time-stretched pulse (TSP) signal and the time-stretched pulse (TSP) signal. The measurement device 200 performs impulse response measurement using an impulse sound as a measurement signal. The measurement signal generating unit 211 outputs the measurement signal to each of the stereo speakers 5 . Here, to obtain the sound pickup signals corresponding to the spatial acoustic transfer characteristics Hls and Hlo, An example in which a measurement signal is output from the meter 5L will be described.

[0049] The left microphone 2L and the right microphone 2R of microphone unit 2 pick up the measurement signal, and The collected signal acquisition unit 212 outputs the collected signal to the processing device 201. The collected signal acquisition unit 212 acquires the collected sound signals from the microphones 2L and 2R. The collected sound signal acquisition unit 212 may include an A / D converter that converts the collected sound signal from the input signal to a digital signal. In other words, the sound signal acquisition unit 212 extracts the sound signal at a predetermined time. The collected signal acquisition unit 212 extracts the collected signal of the number of data (time width). The signals obtained by the left microphone 2L may be synchronously added. s, and the sound signal acquired using the right microphone 2R is the sound signal hlo. ls and hlo are signals sampled at a sampling frequency of 48 kHz. The cut-out audio signals hls and hlo have a filter length (number of samples) of 409 Of course, the sampling frequency and filter length are not limited to the above values. It is not something that can be done.

[0050] The segmental power acquisition unit 215 acquires the segmental power of the picked-up sound signal hls and the picked-up sound signal hlo. For example, the segmental power of the sound pickup signal hls and the sound pickup signal hlo is obtained as h The segmental power hlsP is the amplitude contained in the picked-up signal hls. The segmental power hloP is the sum of the squares of the amplitude values ​​contained in the picked-up signal hlo. In the time domain, the sound pickup signal hls and the sound pickup signal hlo have 4096 data points. Then, the sum of the squares of the 4096 amplitude values ​​becomes the segmental powers hlsP and hloP.

[0051] The frequency characteristic acquisition unit 221 acquires frequency characteristics based on the collected sound signals hls and hlo. The frequency characteristic acquisition unit 221 performs a discrete Fourier transform or a discrete cosine transform on the collected sound signal. The frequency characteristics acquisition unit 221 calculates the frequency characteristics of hls and hlo, respectively. The frequency characteristics are calculated by performing an FFT (Fast Fourier Transform) on the time domain picked-up signal. The frequency characteristics include an amplitude spectrum and a phase spectrum. The acquisition unit 221 may generate a power spectrum instead of an amplitude spectrum. The frequency amplitude characteristics of ls and hlo are Fhls and Fhlo, respectively. s and Fhlo are the spectrum data of the amplitude spectrum.

[0052] The level calculation unit 223 calculates the reference levels for the frequency characteristics Fhls and Fhlo. For example, the level calculation unit 223 calculates the average level (average value) of the frequency characteristics Fhls and Fhlo. ) is calculated and used as the reference level. For example, if FFT is performed with a filter length (number of samples) T, If so, calculate the level value (dB) of each frequency in the frequency amplitude characteristics and calculate the average value. Let real (real part) and imag (imaginary part) after T-point FFT be real[i] and imag[i], respectively. where i is an integer between 0 and (T-1). The sound pressure level Amp_dB[i] at each point i is calculated using the following formula ( 1) Amp_dB[i]=log10(sqrt(real[i]*real[i]+imag[i]*imag[i])) ···(1) In (1), i = 1 ~ (T / 2 + 1) and sqrt is the square root.

[0053] Also, let freq[i] be the frequency (Hz) at point i, and let fs be the sampling frequency. and freq[i] is obtained by the following equation (2). freq[i]=(T / fs)*i (2)

[0054] The reference level A for the entire frequency band is given by the following equation (3).

number

[0055] The reference level of the frequency characteristic Fhls is Ahls, and the reference level of the frequency characteristic Fhlo is If Ahlo, the reference level A is (Ahls+Ahlo) / 2.

[0056] Furthermore, the level calculation unit 223 calculates the maximum level maxL and minimum level mi of the frequency amplitude characteristic. The maximum level maxL is calculated based on the two spectrums of the frequency characteristics Fhls and Fhlo. The minimum level minL is the maximum amplitude value included in the signal data. It is the minimum value among the amplitude values ​​contained in the two spectrum data of ls and Fhlo. Level A, maximum level maxL, and minimum level minL are two frequency characteristics: Fhls and Fh This is a common value for lo.

[0057] The level range setting unit 224 sets the level range X to be compressed. The unit 224 inputs a level range X according to, for example, a playback device. To obtain the desired results, it is preferable to set X to 40 dB or more. If the efficiency and quality of the audio output are not high, X=20 dB can be used. It is preferable that X is between 20 dB and 40 dB, but it is not limited to this range. There is no.

[0058] The correction unit 225 adjusts the frequency characteristics so that the frequency falls within a predetermined level range X including the reference level A. The correction characteristic is calculated by correcting the frequency characteristics Fhls and Fhlo. The correction unit 225 adjusts the frequency characteristics Fhls and Fhl so that the value falls within the level range X. o Compresses the amplitude level. For example, when the level range X is 40 dB, the amplitude value is By compressing the frequency response so that it falls within the range of Fhl ±20 dB, the correction unit 225 The characteristics corrected by the correction unit 225 are used as the correction characteristics. The correction characteristic of Fhls is NewFhls, and the correction characteristic of the frequency characteristic Fhlo is NewFh Let's call it lo.

[0059] Here, the amplitude value before correction at an arbitrary frequency is L, and the amplitude value after correction is NewL. In other words, the frequency characteristics Fhls and Fhlo are a set of amplitude values ​​L before correction, and the corrected characteristics The properties NewFhls, NewFhlo are sets of amplitude values ​​NewL.

[0060] For example, the correction unit 225 corrects the frequency characteristics using the following equations (4) and (5): can be done. If L is greater than or equal to A NewL=A+(LX)*(X / 2) / (maxL-A) ···(4) If L is less than A NewL=A+(LX)*(X / 2) / (A-minL) ···(5)

[0061] By doing this, NewL will be placed within the level range X centered on the reference level. In other words, NewL has an amplitude greater than or equal to (A-(X / 2)) and less than or equal to (A+(X / 2)). Then, the correction unit 225 calculates the following for all data (amplitude value L) within the correction band: The corrected amplitude value NewL is calculated using the above equations (1) and (2). The set of NewL is the correction characteristic. By correcting the amplitude value of the frequency characteristic Fhls, the correction In addition, by correcting using (1) and (2), the frequency characteristics before correction can be obtained. The range can be compressed while maintaining the spectral shape of Fhls and Fhlo.

[0062] The frequency band corrected by the correction unit 225 may be the entire band, or a part of the band. For example, the correction band for correcting the frequency characteristics Fhls and Fhlo is 10Hz to 20k. Hz. In other words, the correction unit 225 can set the minimum frequency (for example, 1 Hz) or higher. Amplitude values ​​are corrected in the band below 10 Hz and in the band above 20 kHz and below the maximum frequency. Therefore, outside the correction band, the amplitude values ​​of the frequency characteristics Fhls and Fhlo remain unchanged. The compensation band is used for headphones 43 that reproduce sounds from outside the head, i.e., the headphones shown in FIG. 43 may be changed depending on the playback band.

[0063] The filter generation unit 230 generates a correction filter based on the correction characteristics. The filter generation unit 230 includes an inverse transformation unit 232 and an adjustment unit 231. The inverse transform unit 232 inversely transforms the correction characteristic to generate a correction signal in the time domain. By using the inverse discrete Fourier transform or inverse discrete cosine transform, the time domain is calculated from the correction characteristics and phase characteristics. The inverse transform unit 232 calculates the correction signal. The inverse transform unit 232 performs IFFT (Inverse Fast Fourier Transform) on the correction characteristic and the phase characteristic. The time domain correction signal is generated by performing a time domain correction (transformation). The correction signal hlo2 obtained from the correction characteristic NewFhlo is called hls2. The correction signals hls2 and hlo2 are filtered with the same filter length as the extracted sound signal. It is called "ta".

[0064] The phase characteristic calculated by the frequency characteristic acquisition unit 221 can be used as it is. That is, the inverse conversion unit 232 can convert the phase characteristic corresponding to the frequency characteristic Fhls and the By performing an inverse Fourier transform on the positive characteristic NewFhls, the correction signal hls2 is generated. The inverse conversion unit 232 generates a phase characteristic corresponding to the frequency characteristic Fhlo and a correction characteristic New A correction signal hlo2 is generated by performing an inverse Fourier transform on Fhlo.

[0065] The segmental power acquisition unit 215 acquires the segmental power of the corrected signal hls2 and the corrected signal hlo2. As mentioned above, the segmental power is the power of the time domain signal. The segmental power of the correction signal hls2 can be calculated as the sum of squares of the amplitude values. P, and the segmental power of the correction signal hlo2 is hlo2.

[0066] The adjustment unit 231 adjusts the correction signal hls so as to maintain the power ratio (energy ratio) between the left and right. 2, adjust the power of hlo2. The adjustment unit 231 adjusts the power ratio before and after the correction so that the power ratio is the same. For example, the adjustment unit 231 multiplies the amplitude value of the correction signal by a predetermined number. The predetermined number for the correction signal hls2 is (hlsP / hlsP2), and the correction signal hlo The predetermined number for 2 is (hloP / hloP2).

[0067] After adjusting the power ratio, the correction signals hls2 and hlo2 are passed through the correction filters hls3 and hlo The product of the amplitude value of the correction signal hls2 and a predetermined number (hlsP / hlsP2) is the correction filter. The amplitude value of the correction signal hlo2 is multiplied by a predetermined number (hloP / hloP 2) is the amplitude value of the correction filter hlo3. Therefore, the product of the correction filter hls3 The segmental power is the same as the segmental power of the picked-up signal hls. Correction filter The segmental power of hlo3 is the same as the segmental power of the picked-up signal hlo.

[0068] In this way, an appropriate correction filter can be generated. The device 201 can generate a correction filter suitable for the playback device. s3 and hlo3 are set as spatial acoustic filters in the convolution calculation units 11 and 12 shown in FIG. This allows the out-of-head localization processing device 100 to perform playback with a high out-of-head localization effect. can.

[0069] Specifically, the correction characteristics are generated so that the signal falls within the level range X that corresponds to the playback device. Therefore, measurements and out-of-head localization processing can be performed in a state suitable for playback equipment. It is possible to generate a filter suitable for out-of-head localization processing.

[0070] Furthermore, in the above embodiment, the adjustment unit 231 adjusts the left and right balance. It is possible to realize well-balanced out-of-head localization reproduction. The power balance adjustment can be omitted. For example, for a single picked-up signal hls, When the processing unit 201 performs the processing, the processing of the adjustment unit 231 is omitted. In this case, the correction signal hls2 is set as it is in the convolution operation unit 11 as a correction filter.

[0071] The processing device 201 also processes the picked-up sound signal that exhibits the spatial acoustic transfer characteristics Hro and Hrs. In this case, the spatial acoustic transfer characteristics Hro and Hrs of the picked-up sound signal are The filter generating unit 230 generates the correction signal so that the segmental power ratio is maintained before and after the correction. Furthermore, the processing device 201 also processes the ear canal transfer characteristics of both ears in the same way. The processing device 201 can calculate the ear canal transfer characteristics ECTFL of the left ear and the ear canal transfer characteristics ECTFL of the right ear. The filter generation is performed so that the segmental power ratio with the ECTFR is maintained before and after correction. The compensation unit 230 adjusts the compensation signal.

[0072] Next, a filter generation method according to this embodiment will be described with reference to FIG. 1 is a flowchart illustrating a filter generation method.

[0073] First, the measurement device 200 measures the transfer characteristics using an impulse sound or the like (S101 That is, the measurement signal generating unit 211 generates a measurement signal such as an impulse sound from the left speaker 5L. The collected sound signal acquisition unit 212 acquires the collected sound signal from the microphone unit 2 (S10 2) The collected sound signal acquisition unit 212 receives the collected sound signal from the left microphone 2L and the collected sound signal from the right microphone 2R. The signal is cut out using a predetermined filter length, resulting in the collected sound signals hls and hlo.

[0074] The segmental power acquisition unit 215 acquires the segmental power of the picked-up signals hls and hlo. The frequency characteristic acquisition unit 221 performs a Fourier transform on the collected sound signal ( S104). This gives the frequency characteristics Fhls and Fhlo. The amplitude spectrum is the frequency amplitude characteristic, but the power spectrum is the frequency power characteristic. That's fine.

[0075] The level calculation unit 223 calculates the reference level (S105). As described above, the reference level is The level calculation unit 2 calculates the average value of the amplitudes of the two frequency characteristics Fhls and Fhlo. 23 calculates the maximum and minimum levels of the frequency characteristics Fhls and Fhlo. The maximum and minimum levels may be calculated from the amplitude values ​​of the entire band, or may be calculated from the amplitude values ​​of a part of the band. It may be calculated from the amplitude value.

[0076] Furthermore, the level range setting unit 224 sets the level range to be compressed (S106). The bell range is set according to the model and performance of the playback device. The staff for generating the filter may input the level range X. Then, the correction unit 225 The amplitude values ​​of the frequency characteristics Fhls and Fhlo are set to fall within the level range X, which includes the reference level. As shown in Fig. 1, the frequency characteristics Fhls and Fhlo are compressed and corrected (S107). Corrective characteristics NewFhls and NewFhlo are obtained. The amplitude value of hlo is included in the level range X.

[0077] Next, the inverse transform unit 232 performs an inverse Fourier transform on the correction characteristics (S108). In this case, the frequency amplitude characteristic is a correction characteristic, and the frequency phase characteristic is a Fourier transform of S104. This is the frequency-phase characteristic calculated by the conversion. The signal hlo2 is obtained.

[0078] The adjustment unit 231 adjusts the segmental power ratio of the picked-up signals hls and hlo so as to maintain the ratio. The amplitude levels of the positive signal hls2 and the correction signal hlo2 are adjusted (S109). The adjustment unit 231 adjusts a predetermined number according to the segmental power ratio to the correction signal hls2 and the correction signal h As a result, the correction filter hls3 and the correction filter hlo3 are The adjustment unit 231 adjusts the power ratio to generate a filter with good left-right balance. It is possible.

[0079] (Correction processing example 1) Next, an example of the correction step in step S107 will be described with reference to FIG. 10 is a flowchart showing a first example of correction processing by the correction unit 225.

[0080] First, the correction unit 225 determines whether the level difference of the frequency amplitude characteristics is equal to or greater than the level range X. The level difference is calculated by dividing the maximum value (maximum level maxL) by the minimum value (minimum level m The maximum and minimum levels are the level difference (maxL - minL) between the entire band. The maximum and minimum values ​​of the frequency amplitude characteristic in the frequency band may be used. may be.

[0081] If the level difference is smaller than the level range X (NO in S201), the correction unit 225 If the difference is greater than the level range X (NO in S201), the correction is not performed and the process is terminated. The correcting unit 225 compresses the level (amplitude value) of each frequency toward the reference level (S202). This compensates for the frequency response so that the level at each frequency falls within the level range X. Be corrected.

[0082] FIG. 7 is a graph showing frequency amplitude characteristics before and after correction in Processing Example 1. That is, FIG. 10 shows the amplitude spectrum of the frequency characteristic before correction Fhls and the corrected characteristic NewFhls. As shown in Figure 7, the frequency amplitude characteristic after correction is within the level range X centered on the reference level A. In Figure 7, the reference level A is -9.4 dB and the level range X is 20 dB. Furthermore, in Figure 7, the correction band is set to 10 Hz to 20 kHz. do.

[0083] (Correction processing example 2) Next, another example of the correction step of step S107 will be described with reference to FIG. FIG. 8 is a flowchart showing a second example of the correction process performed by the correction unit 225. In this case, the correction unit 225 corrects only the levels (amplitude values) that are greater than the reference level.

[0084] First, the correction unit 225 determines whether the level difference of the frequency amplitude characteristics is equal to or greater than the level range X. The level difference is calculated by dividing the maximum value (maximum level maxL) and the minimum value (minimum level mi The maximum and minimum levels are the difference between the maximum and minimum levels (maxL - minL) of the entire band. The maximum and minimum values ​​of the frequency amplitude characteristic in a band may be used. It is also possible.

[0085] If the level difference is smaller than the level range X (NO in S301), the correction unit 225 If the difference is greater than the level range X (YES in S301), The correction unit 225 corrects only the level (amplitude value) of each frequency that is greater than the reference level. The correction unit 225 compresses the signal to a level higher than the reference level (S302). Lower.

[0086] In the processing example 2, the correction unit 225 corrects the level that is smaller than the reference level. Therefore, for frequencies lower than the reference level, the amplitude values ​​before and after correction are the same. .

[0087] In the processing example 2, the correction unit 225 corrects only the level higher than the reference level. In other words, in the processing example 2, only the level lower than the reference level may be corrected. The unit 225 detects only one of a level higher than the reference level and a level lower than the reference level. The correction unit 225 corrects the level equal to or higher than the reference level or the level equal to or lower than the reference level. It is sufficient to correct the frequency characteristics of only one of the channels.

[0088] FIG. 9 is a graph showing frequency amplitude characteristics before and after correction in Processing Example 2. In FIG. The level range is A = -9.4 dB and the level range is X = 20 dB. The positive band is set to 10 Hz to 20 kHz. As shown in Figure 9, the reference level A The amplitude values ​​higher than this are within the level range X after correction. In this case, levels lower than reference level A may not fall within level range X. In other words, in the processing example 2, when the frequency amplitude characteristic is above the min level, (A+(X / 2)) It falls within the following level ranges:

[0089] (Correction processing example 3) In the processing example 3, the frequency axis of the frequency amplitude characteristic is set to a logarithmic scale. It is generally said that human sensory quantities are converted into logarithms. Therefore, it is important to think of the frequency of audible sounds on a logarithmic scale. By doing so, the data will be equally spaced for the above sensory quantities, so data will be collected in all frequency bands. This allows the data to be treated equally. As a result, mathematical operations, frequency band division and weighting are This makes it easier to obtain stable results. The envelope data is not limited to the logarithmic scale, but is converted to a scale close to human hearing (called the auditory scale). The auditory scale can be a logarithmic scale, a mel scale, or a bar scale. Axis transformation is performed using Bark scale, ERB (Equivalent Rectangular Bandwidth) scale, etc. That's fine.

[0090] The frequency characteristic acquisition unit 221 scales the spectrum data on an auditory scale by data interpolation. For example, the frequency characteristic acquisition unit 221 converts low-frequency By interpolating data from several bands, the data in the low frequency band is made dense. On a linear scale, the data is dense in the low frequency range and sparse in the high frequency range. In this way, the frequency characteristic acquisition unit 221 obtains data on axes at equal intervals on the auditory scale. Of course, the axis-transformed data can be perfectly accurate on the auditory scale. The data does not have to be at completely equal intervals. By doing so, the correction unit 225 etc. can calculate the logarithmic scale. The frequency amplitude characteristic is processed. Also, the frequency phase characteristic and the number of samples are adjusted. To do this, the frequency axis may be restored to a linear scale before the inverse transform.

[0091] (Correction processing example 4) In the processing example 4, the correction unit 225 corrects the entire correction band, but corrects the level range X The amplitude value is corrected only at the frequency around the peak that exceeds the upper limit value. 10 is a flowchart showing the fourth processing example.

[0092] First, the correction unit 225 determines whether the level difference of the frequency amplitude characteristics is equal to or greater than the level range X. The level difference is calculated by dividing the maximum value (maximum level maxL) and the minimum value (minimum level mi The maximum and minimum levels are the difference between the maximum and minimum levels (maxL - minL) of the entire band. The maximum and minimum values ​​of the frequency amplitude characteristic in a band may be used. It is also possible.

[0093] If the level difference is smaller than the level range X (NO in S401), the correction unit 225 If the difference is greater than the level range X (NO in S201), the correction is not performed and the process is terminated. The positive part 225 is the frequency around the peak frequency that exceeds the upper limit of the range (A+X / 2). , the amplitude value is compressed toward a reference level (S402).

[0094] For example, the correction unit 225 determines the crossover frequencies that cross the upper limit before and after the peak frequency. The correction unit 225 calculates a first crossover frequency that is lower than the peak frequency and a second crossover frequency that is higher than the peak frequency. The correction unit 225 calculates the first and second crossover frequencies. In a specified frequency band, the amplitude value is compressed toward a reference level.

[0095] Specifically, the correction unit 225 corrects the upper limit of the range on the lower frequency side than the peak frequency. The correction unit 225 calculates a first crossover frequency at which the frequency crosses the peak frequency. The correction unit 225 calculates a second crossover frequency at which the first crossover frequency intersects with the upper limit of the range. The amplitude value is corrected in the frequency band from the wave number to the second crossover frequency. This makes it possible to correct amplitude values ​​that exceed the upper limit of the range around the peak.

[0096] FIG. 11 is a graph showing three frequency bands (a) to (c) defined by crossover frequencies. The frequency band (a) is a frequency band including the first peak P1. The frequency band (a) is defined by the crossover frequencies before and after the first peak P1. Frequency band (c) is the frequency band including the third peak P3. As shown in FIG. 11, one frequency band may have multiple peaks that are close to each other. may also include:

[0097] In this way, in processing example 4, only amplitude values ​​that exceed the upper limit of the range are set to the reference level. The correction unit 225 also compresses the signal toward the lower limit (A-(X / 2 The amplitude value may be corrected at a frequency around the dip below . 25 calculates the crossover frequency at which the frequency crosses the lower limit value before and after the dip below the lower limit value. 25 can be achieved by compressing the amplitude value in the frequency band defined by the two crossover frequencies. Of course, the correction unit 225 corrects both the frequency band including the peak and the frequency band including the dip. Alternatively, the correction unit 225 may compress the amplitude value only in the frequency band including the peak. The amplitude values ​​may be compressed, or the amplitude values ​​may be compressed only in the frequency band including the dip.

[0098] (Correction processing example 5) In the processing example 5, the correction unit 225 performs correction using a different method. The amplitude level is corrected using smoothing processes such as moving average. Using techniques such as olay filters, smoothing splines, cepstrum transforms, and cepstrum envelopes The correction unit 225 smoothes the frequency characteristics (spectrum data) by By performing smoothing processing, the frequency characteristics are corrected to fall within level range X.

[0099] (Correction processing example 6) In Processing Example 6, the collected sound signal is processed for the ear canal transfer characteristics. Specifically, the measurement is performed by the measurement device 300 shown in FIG. The measurement signal generating unit 211 outputs the measurement signal to the headphones 43 instead of the speaker 5L. In this case, the left and right microphones 2L and 2R collect the sound signals that indicate the ear canal transfer characteristics of the left and right ears. The frequency amplitude characteristics are acquired. The reference level, maximum level, and minimum level are two frequencies. The content other than the above is the same as the above embodiment and processing example. Therefore, the description will be omitted.

[0100] (Correction processing example 7) In processing example 7, multi-channel speakers such as 5.1ch and 7.1ch are used. Then, the adjustment unit 231 adjusts the power ratio of the collected signals for each channel so that the power ratio is maintained. We are carrying out the following.

[0101] In 5.1ch multi-channel, the left and right front speakers, the left and right rear speakers, In this case, the front and rear speakers are used. The adjustment unit 231 adjusts the correction signal so that the power ratio of the speakers is maintained. The adjustment unit 231 adjusts the coefficients for each correction such that the segmental power ratios before and after correction are the same. Multiply the positive signal.

[0102] Specifically, the measurement device 200 performs measurements using speakers of different channels in sequence. For example, the measurement signal generating unit 211 generates a measurement signal and transmits it to the speakers of each channel in order. The collected signal acquisition unit 212 sequentially collects measurement signals from the speakers of each channel. The frequency characteristic acquisition unit 221 acquires an acquisition signal by sounding the signals of different channels. Based on the collected signal obtained by collecting the measurement signal output from the speaker, Obtain the wavenumber characteristics.

[0103] The segmental power acquisition unit 215 acquires the left and right segmental powers of the picked-up sound signal for each channel. The adjustment unit 231 adjusts the level of the correction signal so as to maintain the power ratio. This allows you to create a filter with good balance between channels. The level range X can be different for each channel, and can be the same for all channels. may be.

[0104] Note that the process of maintaining the power ratio between channels is limited to multi-channels such as 5.1ch. It can also be applied to the 2-channel measurement device shown in Figure 2. For example, may be measured and adjusted to maintain the power ratio.

[0105] The above processing examples 1 to 7 can be combined as appropriate. For example, processing example 4 When correcting the amplitude value around the peak frequency or dip frequency, the correction unit 2 25 may use the axis conversion process of the frequency axis in the processing example 3 or the smoothing process in the processing example 5.

[0106] In this way, according to this embodiment, the level falling within a predetermined level range X including the reference level is Therefore, even with various playback devices, equipment, and measurement environments, It is possible to reproduce a filter that can obtain an appropriate out-of-head localization effect. It is possible to automatically correct the filter so that the localized signal does not clip. To provide out-of-head localization listening according to the user's preferences, speakers, headphones, and measurement environment. Furthermore, automatic correction according to the playback device is possible.

[0107] Embodiment 2 The apparatus and method according to the second embodiment will be described with reference to FIG. 12. 1 is a block diagram showing the configuration of the device 201. In the second embodiment, when a level range X is set, Therefore, the processing device 201 shown in FIG. 4, a determination unit 242 is added. The principle is the same as in the first embodiment, so the explanation will be omitted where appropriate.

[0108] The determination unit 242 determines the performance of the playback device. For example, the determination unit 242 determines the performance of the playback device. The level range setting unit 224 evaluates the performance of the amplifier according to the result of the determination by the determination unit 242. Then, the level range setting unit 224 sets the level range X. The correction unit 225 The correction characteristic is calculated by correcting the frequency characteristic based on the range X. 30 generates a correction filter based on the correction characteristics.

[0109] For example, the determining unit 242 determines, based on the frequency characteristics acquired by the frequency characteristics acquiring unit 221, The determining unit 242 determines the maximum level (maxL) of the frequency amplitude characteristic. The determination unit 242 detects the level difference (maxL-minL) of the minimum level (minL). Based on the level difference, the output level (output sound pressure level) and S / N ratio of the playback device are obtained. Then, the determining unit 242 determines the performance based on the output level or the S / N ratio. The level level control unit 242 adjusts the level level in accordance with the difference between the maximum level and the minimum level of the frequency amplitude characteristic. A range X may be determined.

[0110] For example, in the case of a playback device with a large level difference, the level range X should be set to about 80% of the level difference. The determination unit 242 sets the variable to 0.8. In the case of a playback device with a small level difference The level range X is set to about 40% of the level difference. The level range X is set by multiplying the level difference by a variable depending on the result.

[0111] Furthermore, the processing device 201 can set the level range X without using a variable. For example, the determination unit 242 calculates the level difference (maxL-minL) in a part of the determination band as The determination band can be set to, for example, 100 Hz to 8 kHz. The determination unit 242 determines the maximum level (maxL) and minimum level ( Then, the determining unit 242 determines the level difference (maxL-minL) based on the level difference (maxL-minL). Alternatively, the determining unit 242 may convert the level difference into a level range X. The device may have a conversion formula or conversion table for this purpose.

[0112] In this way, the judgment is made based on the frequency characteristics of the picked-up signal obtained by measurement using a playback device. The determination unit 242 determines the maximum and minimum levels of the frequency characteristics. The judgment is based on the level difference between

[0113] The determination unit 242 also acquires playback device information about the playback device and determines whether the playback device is a The level range setting unit 224 may determine the performance based on the performance of the playback device. For example, if the playback device has a high-performance amplifier, set the level range X accordingly. The level range setting unit 224 sets X=40 dB. The range setting unit 224 sets X=20 dB. The number of stages is not limited to two, high performance and low performance, but may be three or more stages.

[0114] The determination unit 242 may also have a table showing the performance for each model number of the playback device. The determination unit 242 acquires playback device information indicating the model number of the playback device. The performance of the playback device is determined according to the model number of the playback device. It may be acquired by the Bluetooth connection or input by the user. In the case of a connected playback device, the determination unit 242 can automatically acquire information about the playback device. can.

[0115] For example, the measurement device 200 or the measurement device 300 repeats the measurement for acquiring the frequency characteristics. Then, as described above, the determination unit 242 determines the level of the frequency characteristics. The performance is judged according to the difference in performance, and the judgment result is stored in a table. The determination can be made by referring to the table.

[0116] The playback device may be the speakers 5L and 5R and their amplifiers shown in FIG. 2, or the It may be headphones 43. In other words, the playback device is the playback device used during measurement. Alternatively, the headphones 43 in the out-of-head localization processing device shown in FIG. In other words, the playback device is a headphone43 or earphone used for out-of-head localization listening. In the second embodiment, any one or more of the above processing examples 1 to 7 may be used. This can be done.

[0117] In this way, according to this embodiment, the level range X is automatically set according to the performance of the playback device. Then, the correction unit 225 performs correction based on the level range X. Therefore, it is possible to obtain appropriate out-of-head localization effects even with various playback devices, equipment, and measurement environments. In other words, the signal processed for out-of-head localization can be reproduced. It can automatically correct the filter so that it does not slip. It is possible to perform out-of-head localization listening depending on the speaker, headphones, and measurement environment. Automatic correction according to the device becomes possible.

[0118] Some or all of the above processes may be performed by a computer program. The above-mentioned program, when loaded into a computer, performs the operations described in the embodiments. or a set of instructions (or software code) that causes a computer to perform more functions. The program is stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example, and not limitation, a computer-readable medium or tangible storage medium may , random-access memory (RAM), read-only memory (ROM), flash memory, solid- solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD) , Blu-ray (registered trademark) disc or other optical disc storage, magnetic cassette, magnetic This includes magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a transitory computer-readable medium or a communication medium. By way of example, transitory computer-readable media or communication media may include electrical, optical, acoustic, or other forms of propagated signals.

[0119] The invention made by the present inventor has been specifically described above based on the embodiments. The present invention is not limited to the above-described embodiment, and various modifications are possible without departing from the spirit of the present invention. It goes without saying that this is the case. [Explanation of symbols]

[0120] U User 1 Person to be measured 2 microphone units 2L Left microphone 2R Right Microphone 5 stereo speakers 5L Left speaker 5R Right speaker 10. Out-of-head localization processing unit 11 Convolution operation unit 12 Convolution operation unit 21 Convolution operation unit 22 Convolution operation unit 24 Adder 25 Adder 41 Filter section 42 Filter section 43 Headphones 200 Measuring Equipment 201 Processing equipment 211 Measurement signal generation unit 212 Sound signal acquisition unit 215 Segmental Power Acquisition Department 221 Frequency characteristic acquisition unit 223 Level Calculation Unit 224 Level range setting section 225 Correction Unit 230 Filter Generation Unit 231 Adjustment section 232 Inverse conversion unit 242 Judgment section

Claims

1. a frequency characteristic acquisition unit that acquires frequency characteristics based on a sound signal collected by a microphone, and interpolates low-frequency band data using a frequency axis as an auditory scale similar to human hearing, thereby generating axis-transformed data at equal intervals on the auditory scale; a determination unit that determines the performance of the playback device; a level range setting unit that sets a level range according to the determination result of the determination unit; a correction unit that calculates a correction characteristic by correcting the frequency characteristic of the frequency amplitude characteristic of the hearing scale based on the level range; and a filter generation unit that generates, based on the correction characteristics, a spatial acoustic filter used in out-of-head localization processing or an inverse filter that cancels characteristics from the playback device to the user's ears.

2. a frequency characteristic acquisition unit that acquires frequency characteristics based on a sound signal collected by a microphone, and interpolates low-frequency band data using a frequency axis as an auditory scale similar to human hearing, thereby generating axis-transformed data at equal intervals on the auditory scale; a level calculation unit that calculates a reference level in the frequency characteristic; a correction unit that calculates a correction characteristic by correcting the frequency characteristic so that the frequency amplitude characteristic of the hearing scale falls within a predetermined level range including the reference level; and a filter generation unit that generates a spatial acoustic filter or an inverse filter used in out-of-head localization processing based on the correction characteristics.

Citation Information

Patent Citations

  • Processing device, processing method, and program

    JP2019062430A