Processing device and processing method

By acquiring and compressing the peaks or troughs in the external positioning process, the problems of unstable sound image positioning and uneven sound quality in the existing technology are solved, thus improving the user experience.

CN115914978BActive Publication Date: 2025-11-21JVC KENWOOD CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210749119.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-06
Filing Date
2022-06-29
Publication Date
2025-11-21
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

In external head positioning processing, existing technologies have difficulty effectively compressing the steep peaks or troughs of the external auditory canal transmission characteristics, resulting in unstable sound image positioning and uneven sound quality. Furthermore, the characteristics need to be re-measured when frequently changing headphones, which puts a burden on users.

Method used

By acquiring the frequency characteristics of the input signal, extracting the extreme values, calculating the kurtosis of the extreme values ​​and comparing it with a threshold, determining whether to compress the peaks or troughs, and using processing devices and methods to appropriately compress the peaks or troughs.

Benefits of technology

It achieves stable compression of peaks or troughs, improves the stability of sound image localization and sound quality balance, and reduces the frequency of repeated measurements by users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115914978B_ABST
    Figure CN115914978B_ABST
Patent Text Reader

Abstract

The present application provides a processing device and a processing method capable of appropriately compressing a peak or a valley. The processing device of the present embodiment includes: a frequency characteristic acquisition section (214) that acquires a frequency characteristic of an input signal; an extreme value extraction section (216) that extracts an extreme value of frequency spectrum data; a kurtosis calculation section (217) that calculates a value for evaluating a peak or a valley corresponding to the extreme value from the frequency spectrum data, and calculates kurtosis of the peak or the valley based on a plurality of values calculated by changing a calculation width; a determination section (218) that determines whether to compress the peak or the valley based on a comparison result of the kurtosis and a threshold value; and a compression section (219) that compresses the peak or the valley of the extreme value determined to be compressed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to processing apparatus and processing methods. Background Technology

[0002] As a sound image localization technology, there is an extra-head localization technology that uses headphones to localize the sound image to the outside of the listener's head. In extra-head localization technology, the characteristics from the headphones to the ears (headphone characteristics) are eliminated, and two characteristics from a single loudspeaker (mono loudspeaker) to the ears (spatial acoustic transmission characteristics) are assigned, thereby localizing the sound image outside the head.

[0003] In stereo loudspeaker head positioning reproduction, a microphone (hereinafter referred to as a microphone) placed in the listener's (receiver's) ear records measurement signals (pulse tones, etc.) emitted from loudspeakers on two channels (hereinafter referred to as ch). Furthermore, a processing device generates a filter based on the picked-up audio signals obtained from the measurement signals. By convolving the generated filter with the audio signals of 2ch, head positioning reproduction can be achieved.

[0004] Furthermore, in order to generate a filter (also known as an inverse filter) that eliminates the characteristics from the headphones to the ear, a microphone placed in the listener's own ear is used to measure the characteristics from the headphones to the ear and then to the eardrum (also known as the external auditory canal transfer function ECTF, external auditory canal transfer characteristics).

[0005] Patent Document 1 discloses an external head positioning processing device that uses a filter for external head positioning processing. In Patent Document 1, a measuring microphone disposed in the user's external auditory canal picks up pulsed sounds. This allows for the measurement of the transmission characteristics from the speaker unit of the headphones to the microphone within the external auditory canal.

[0006] Existing technical documents

[0007] Patent documents

[0008] Patent document 1: Japanese Patent Application Publication No. 2020-136752. Summary of the Invention

[0009] The problem that the invention aims to solve

[0010] When performing head-to-head localization processing, it is preferable to use a microphone placed in the listener's own ear to measure characteristics. When measuring the transmission characteristics of the external auditory canal, impulse response measurements are performed while the listener is wearing a microphone and headphones. By utilizing the listener's own characteristics, a filter suitable for the listener can be generated. To generate such a filter, it is desirable to appropriately process the measured pickup signal.

[0011] Inverse filters can be generated using algorithms such as the least squares method, but due to the nature of these algorithms, it is difficult to create a completely inverse characteristic across all frequencies. Furthermore, when the ECTF itself has steep peaks or troughs, these peaks or troughs become steep peaks or troughs in the inverse filter to generate its inverse characteristic. Moreover, if the inverse filter generates inverse filters with control points at different locations for the head-mounted sound field reproduction, unexpected peaks may be produced.

[0012] Furthermore, when a user re-wears the headphones, the frequencies that generate peaks before and after re-wearing may shift. This can sometimes negatively impact positioning or sound quality equalization. Ideally, the ECTF would be measured each time the headphones are re-weared, but this would be burdensome for the user. Therefore, it is desirable to pre-compress the steep peaks and troughs of the frequency response obtained through user measurements.

[0013] This disclosure was made in view of the above-mentioned problems, and its purpose is to provide a processing apparatus and processing method capable of appropriately compressing wave crests or troughs.

[0014] means for solving problems

[0015] The processing apparatus in this embodiment includes: a frequency characteristic acquisition unit for acquiring the frequency characteristics of an input signal; an extreme value extraction unit for extracting extreme values ​​of spectral data based on the frequency characteristics; a kurtosis calculation unit for calculating an evaluation value for a peak or trough corresponding to the extreme value based on spectral data in a computational width including the extreme value, and calculating the kurtosis of the peak or trough based on multiple evaluation values ​​calculated by changing the computational width; a determination unit for determining whether to compress the peak or trough based on a comparison result between the kurtosis and a threshold; and a compression unit for compressing the peak or trough that is determined to be the extreme value to be compressed.

[0016] The processing method of this embodiment includes: a step of acquiring the frequency characteristics of an input signal; a step of extracting the extreme values ​​of the frequency characteristics; a step of calculating an evaluation value for the extreme value based on data in the operational width containing the extreme value, and a step of calculating the kurtosis of the extreme value based on multiple evaluation values ​​calculated by changing the operational width; a step of determining whether to compress the extreme value based on a comparison result between the kurtosis and a threshold; and a step of compressing the extreme values ​​that are determined to be compressed.

[0017] Invention Effects

[0018] According to this disclosure, a processing apparatus and processing method can be provided that can appropriately compress wave crests or troughs. Attached Figure Description

[0019] Figure 1This is a block diagram illustrating the external head positioning processing device of this embodiment.

[0020] Figure 2 It is a diagram schematically showing the structure of the measuring device.

[0021] Figure 3 It is a block diagram showing the structure of the processing device.

[0022] Figure 4 This is a graph representing an example of frequency-amplitude characteristics.

[0023] Figure 5 It is a graph representing the peaks extracted from the spectrum after axis transformation.

[0024] Figure 6 This is a diagram illustrating an example of the process for calculating kurtosis.

[0025] Figure 7 This is a diagram used to illustrate the process of merging close peaks in Implementation Method 2.

[0026] Figure 8 This is a flowchart illustrating the processing method of the implementation method. Detailed Implementation

[0027] The general outline of the acoustic localization processing in this embodiment will be explained. The head-out localization processing in this embodiment utilizes spatial acoustic transmission characteristics and external auditory canal transmission characteristics. Spatial acoustic transmission characteristics refer to the transmission characteristics from a sound source such as a loudspeaker to the external auditory canal. External auditory canal transmission characteristics refer to the transmission characteristics from the speaker unit of a headset or in-ear headphone to the eardrum. In this embodiment, the spatial acoustic transmission characteristics are measured when the headset or in-ear headphone is not worn, and the external auditory canal transmission characteristics are measured when the headset or in-ear headphone is worn. These measurement data are used to achieve head-out localization processing. This embodiment features a microphone system used to measure the spatial acoustic transmission characteristics and external auditory canal transmission characteristics.

[0028] In this embodiment, the head-mounted positioning process is performed by a user terminal such as a personal computer, smartphone, or tablet PC. The user terminal is an information processing device comprising: a processing unit such as a processor; a storage unit such as a memory or hard disk; a display unit such as an LCD monitor; and an input unit such as a touch panel, buttons, keyboard, or mouse. The user terminal may also have communication functions for sending and receiving data. Furthermore, an output unit with either a headset or an in-ear headphone is connected to the user terminal. The connection between the user terminal and the output unit can be wired or wireless.

[0029] Implementation Method 1

[0030] (External head positioning processing device)

[0031] Figure 1 This is a block diagram showing an external head positioning processing device 100, which is an example of a sound field reproduction device according to this embodiment. The external head positioning processing device 100 reproduces the sound field for a user U wearing headphones 43. Therefore, the external head positioning processing device 100 performs sound image positioning processing on the stereo input signals XL and XR of Lch and Rch. The stereo input signals XL and XR of Lch and Rch are analog audio reproduction signals output from a CD (Compact Disc) player or the like, or digital audio data such as mp3 (MPEG Audio Layer-3). Furthermore, audio reproduction signals or digital audio data are collectively referred to as reproduction signals. That is, the stereo input signals XL and XR of Lch and Rch are reproduction signals.

[0032] Furthermore, the head positioning processing device 100 is not limited to a single physical device; some processing can be performed using different devices. For example, some processing can be performed using a smartphone, while the remaining processing can be performed using a DSP (Digital Signal Processor) built into the headset 43.

[0033] The external head positioning processing device 100 includes an external head positioning processing unit 10, a filter unit 41 storing an inverse filter Linv, a filter unit 42 storing an inverse filter Rinv, and a headset 43. The external head positioning processing unit 10, the filter unit 41, and the filter unit 42 can be implemented by a processor or the like.

[0034] The external head positioning processing unit 10 includes convolution operation units 11-12, 21-22 that store spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs, as well as adders 24 and 25. The convolution operation units 11-12, 21-22 perform convolution processing using the spatial acoustic transfer characteristics. Stereo input signals XL and XR from a CD player, etc., are input to the external head positioning processing unit 10. Spatial acoustic transfer characteristics are set in the external head positioning processing unit 10. The external head positioning processing unit 10 convolves the stereo input signals XL and XR of each channel with a filter (hereinafter also referred to as a spatial acoustic filter) that utilizes the spatial acoustic transfer characteristics. The spatial acoustic transfer characteristics can be the head transfer function HRTF measured through the head or auricle of the subject, or the head transfer function of a simulated head or a third party.

[0035] The four spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are grouped together as a spatial acoustic transfer function. The data used for convolution in the convolution operation units 11, 12, 21, and 22 becomes a spatial acoustic filter. The spatial acoustic filter is generated by cutting out the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs with a predetermined filter length.

[0036] Spatial acoustic transmission characteristics Hls, Hlo, Hro, and Hrs are pre-acquired through impulse response measurements, etc. For example, user U has microphones installed in both ears. Left and right speakers positioned in front of user U output impulse tones for impulse response measurements. The microphones pick up the measurement signals, such as the impulse tones, output from the speakers. Based on the microphone signals, the spatial acoustic transmission characteristics Hls, Hlo, Hro, and Hrs are acquired. The spatial acoustic transmission characteristics Hls between the left speaker and the left microphone, Hlo between the left speaker and the right microphone, Hro between the right speaker and the left microphone, and Hrs between the right speaker and the right microphone are measured.

[0037] Then, the convolution operation unit 11 convolves the stereo input signal XL of Lch with the spatial acoustic filter corresponding to the spatial acoustic transfer characteristic Hls. The convolution operation unit 11 outputs the convolution operation data to the adder 24. The convolution operation unit 21 convolves the stereo input signal XR of Rch with the spatial acoustic transfer characteristic Hro corresponding to the spatial acoustic filter. The convolution operation unit 21 outputs the convolution operation data to the adder 24. The adder 24 adds the two convolution operation data and outputs the result to the filter unit 41.

[0038] Convolution operation unit 12 convolves the stereo input signal XL of Lch with a spatial acoustic filter corresponding to the spatial acoustic transfer characteristic Hlo. Convolution operation unit 12 outputs the convolution operation data to adder 25. Convolution operation unit 22 convolves the stereo input signal XR of Rch with a spatial acoustic filter corresponding to the spatial acoustic transfer characteristic Hrs. Convolution operation unit 22 outputs the convolution operation data to adder 25. Adder 25 adds the two convolution operation data and outputs the result to filter unit 42.

[0039] Inverse filters Linv and Rinv, which eliminate the characteristics of headphones (the characteristics between the headphone's playback unit and microphone), are provided in filter sections 41 and 42. Then, the playback signal (convolution operation signal) processed in the head positioning processing section 10 is convolved with the inverse filters Linv and Rinv. Filter section 41 convolves the Lch signal from adder 24 with the inverse filter Linv for the headphone characteristics on the Lch side. Similarly, filter section 42 convolves the Rch signal from adder 25 with the inverse filter Rinv for the headphone characteristics on the Rch side. When the headphones 43 are worn, the inverse filters Linv and Rinv eliminate the characteristics from the headphone unit to the microphone. The microphone can be positioned between the entrance to the external auditory canal and the tympanic membrane.

[0040] Filter unit 41 outputs the processed Lch signal YL to the left unit 43L of the headset 43. Filter unit 42 outputs the processed Rch signal YR to the right unit 43R of the headset 43. User U wears the headset 43. The headset 43 outputs the Lch signal YL and the Rch signal YR to user U (hereinafter, the Lch signal YL and the Rch signal YR are collectively referred to as stereo signals). Thus, the sound image positioned outside the user U's head can be reproduced.

[0041] Thus, the external head positioning processing device 100 performs external head positioning processing using spatial acoustic filters corresponding to the spatial acoustic transmission characteristics Hls, Hlo, Hro, and Hrs, and inverse filters Linv and Rinv representing the characteristics of headphones. In the following description, the spatial acoustic filters corresponding to the spatial acoustic transmission characteristics Hls, Hlo, Hro, and Hrs, and the inverse filters Linv and Rinv representing the characteristics of headphones, are collectively referred to as the external head positioning processing filters. In the case of a 2ch stereo reproduction signal, the external head positioning filters consist of four spatial acoustic filters and two inverse filters. Then, the external head positioning processing device 100 performs external head positioning processing by convolving the stereo reproduction signal using a total of six external head positioning filters. The external head positioning filters are preferably based on measurements taken by the user U. For example, the external head positioning filters are set based on the pickup signal picked up by the microphone worn in the user U's ear.

[0042] Thus, the spatial acoustic filter and the inverse filter of the headphone characteristics, Linv and Rinv, are filters used for the audio signal. By convolving these filters with the reproduced signal (stereo input signal XL, XR), the head positioning processing device 100 performs head positioning processing. In this embodiment, the processing for generating the inverse filters Linv and Rinv is one of the technical features. The processing for generating the inverse filters will be described below.

[0043] (Measuring apparatus)

[0044] use Figure 2 The measuring device 200 will be described. Figure 2 This describes the structure used to measure the transmission characteristics of user U. To generate an inverse filter, the measuring device 200 measures the transmission characteristics of the external auditory canal. The measuring device 200 includes a microphone unit 2, a headset 43, and a processing device 201. Additionally, here, the subject 1 is... Figure 1 The user U can be the same person, but can also be a different person.

[0045] In this embodiment, the processing unit 201 of the measuring device 200 performs computational processing to appropriately generate a filter based on the measurement results. The processing unit 201 is a personal computer (PC), tablet terminal, smartphone, etc., and includes a memory and a processor. The memory stores processing programs, various parameters, measurement data, etc. The processor executes the processing program stored in the memory. The processor executes the processing program, thereby performing various processes. The processor may be, for example, a CPU (Central Processing Unit), FPGA (Field-Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), or GPU (Graphics Processing Unit), etc.

[0046] A microphone unit 2 and a headset 43 are connected to the processing device 201. Alternatively, the microphone unit 2 can be built into the headset 43. The microphone unit 2 includes a left microphone 2L and a right microphone 2R. The left microphone 2L is worn in the left ear 9L of the user U. The right microphone 2R is installed in the right ear 9R of the user U. The processing device 201 can be the same as the external head positioning processing device 100, or it can be a different processing device. Alternatively, an in-ear headphone can be used instead of the headset 43.

[0047] The headphones 43 include a headphone strap 43B, a left earpiece 43L, and a right earpiece 43R. The headphone strap 43B connects the left earpiece 43L and the right earpiece 43R. The left earpiece 43L outputs sound to the left earpiece 9L of the user U. The right earpiece 43R outputs sound to the right earpiece 9R of the user U. The headphones 43 can be closed-back, open-back, semi-open, or semi-closed, regardless of the type of headphones. The user U wears the headphones 43 with the microphone 2 worn by the user U. That is, the left earpiece 43L and the right earpiece 43R of the headphones 43 are respectively installed in the left earpiece 9L and the right earpiece 9R, where the left microphone 2L and the right microphone 2R are installed. The headphone strap 43B generates a force that presses the left earpiece 43L and the right earpiece 43R against the left earpiece 9L and the right earpiece 9R, respectively.

[0048] The left microphone 2L picks up sound output from the left earpiece 43L of the headphones 43. The right microphone 2R picks up sound output from the right earpiece 43R of the headphones 43. The microphones of the left microphone 2L and the right microphone 2R are positioned near the ear canal. The left microphone 2L and the right microphone 2R are configured not to interfere with the headphones 43. That is, with the left microphone 2L and the right microphone 2R positioned appropriately in the left ear 9L and right ear 9R, the user U can wear the headphones 43.

[0049] The processing unit 201 outputs a measurement signal to the headset 43. As a result, the headset 43 generates impulse sounds, etc. Specifically, the impulse sounds output from the left unit 43L are measured by the left microphone 2L, and the impulse sounds output from the right unit 43R are measured by the right microphone 2R. When outputting the measurement signal, the pickup signals are acquired through microphones 2L and 2R, and impulse response measurement is performed.

[0050] The processing unit 201 generates inverse filters Linv and Rinv by performing the same processing on the pickup signals from microphones 2L and 2R. Specifically, it performs processing to compress the peaks of the frequency characteristics of the pickup signals.

[0051] The processing device 201 of the measuring apparatus 200 and its processing will be described in detail below. Figure 3 This is a control block diagram representing the processing device 201. The processing device 201 includes a measurement signal generation unit 211, a pickup signal acquisition unit 212, an inverse filter generation unit 213, a frequency characteristic acquisition unit 214, an axis transformation unit 215, an extremum extraction unit 216, a kurtosis calculation unit 217, and a filter generation unit 221.

[0052] The measurement signal generation unit 211 includes a D / A converter, an amplifier, etc., and generates a measurement signal for measuring the transmission characteristics of the external auditory canal. The measurement signal may be, for example, a pulse signal or a TSP (Time Stretched Pulse) signal. Here, a pulse tone is used as the measurement signal, and the measurement device 200 performs impulse response measurement.

[0053] Microphone unit 2's left microphone 2L and right microphone 2R respectively pick up the measured signal and output the picked-up signal to processing device 201. The pickup signal acquisition unit 212 acquires the pickup signals picked up by the left microphone 2L and right microphone 2R. Furthermore, the pickup signal acquisition unit 212 may also include an A / D converter for performing A / D conversion on the pickup signals from microphones 2L and 2R. The pickup signal acquisition unit 212 may also perform synchronous addition operations on signals obtained through multiple measurements. The time-domain pickup signal is referred to as ECTF.

[0054] The inverse filter generation unit 213 generates an inverse filter that eliminates the transmission characteristics of the external auditory canal as the input signal based on the pickup signal. The inverse filter generation unit 213 calculates the frequency characteristics of the pickup signal using a discrete Fourier transform or a discrete cosine transform. For example, the inverse filter generation unit 213 calculates the frequency characteristics by performing an FFT (Fast Fourier Transform) on the input signal in the time domain. The frequency characteristics include the amplitude spectrum and the phase spectrum. Alternatively, the inverse filter generation unit 213 can generate a power spectrum instead of the amplitude spectrum.

[0055] The inverse filter generation unit 213 requires the elimination of inverse characteristics such as amplitude spectrum. Inverse characteristics are amplitude spectra with filter coefficients that eliminate amplitude spectrum characteristics. The inverse filter generation unit 213 calculates the time-domain signal based on the inverse and phase characteristics using inverse discrete Fourier transform or inverse discrete cosine transform. The inverse filter generation unit 213 generates a time signal by performing IFFT (Inverse Fast Fourier Transform) on the inverse and phase characteristics. The inverse filter generation unit 213 calculates the inverse filter by cutting the generated time signal to a predetermined filter length. The inverse filter generation unit 213 can also perform windowing to generate the inverse filter. The inverse filter generation unit 213 outputs the generated inverse filter as an input signal to the frequency response acquisition unit 214.

[0056] The frequency response acquisition unit 214 acquires the frequency response of the input signal. The frequency response acquisition unit 214 calculates the frequency response of the input signal using a Discrete Fourier Transform (DFT) or a Discrete Cosine Transform (DCT). For example, the frequency response acquisition unit 214 calculates the frequency response by performing an FFT (Fast Fourier Transform) on the input signal in the time domain. The frequency response includes the amplitude spectrum and the phase spectrum. Alternatively, the frequency response acquisition unit 214 may generate a power spectrum instead of the amplitude spectrum. Thus, the frequency response acquisition unit 214 acquires the frequency response of the input signal as an inverse filter.

[0057] The axis transformation unit 215 transforms the frequency axis of the frequency characteristics acquired by the frequency characteristic acquisition unit 214 through data interpolation. The axis transformation unit 215 changes the scale of the frequency amplitude characteristic data on the logarithmic axis in such a way that the discrete spectral data becomes equally spaced. The frequency amplitude characteristic data (also called gain data) obtained by the frequency characteristic acquisition unit 214 becomes equally spaced on the frequency axis. That is, the gain data is equally spaced on the linear frequency axis, and therefore not equally spaced on the logarithmic frequency axis. Therefore, the axis transformation unit 215 performs interpolation processing on the gain data so that the gain data becomes equally spaced on the logarithmic frequency axis.

[0058] In gain data, on the logarithmic axis, the lower the frequency domain, the sparser the intervals between adjacent data points, and the higher the frequency domain, the denser the intervals between adjacent data points. Therefore, the axis transformation unit 215 interpolates the data in the sparsely spaced low-frequency band. Specifically, the axis transformation unit 215 calculates discrete gain data arranged at equal intervals on the logarithmic axis through interpolation processing such as three-dimensional spline interpolation. The gain data that has undergone axis transformation is used as axis-transformed data. The axis-transformed data is a spectrum that establishes a correspondence between frequency and amplitude values ​​(gain values).

[0059] The rationale for transforming the frequency axis to a logarithmic scale is explained. It is generally believed that human sensory quantities are transformed into logarithms. Therefore, the frequencies of heard sounds are also important on a logarithmic axis. By performing a scale transformation, the data in the aforementioned sensory quantities become equally spaced, thus enabling equivalent processing of data across the entire frequency band. As a result, mathematical operations, frequency band segmentation, and weighting become easier, leading to stable results. Furthermore, the axis transformation unit 215 is not limited to a logarithmic scale; it can transform the envelope data to a scale close to human hearing (called the auditory scale). As an auditory scale, axis transformation can also be performed using a logarithmic scale (Log scale), a mel scale, a Bark scale, an ERB (Equivalent Rectangular Bandwidth) scale, etc.

[0060] The axis transformation unit 215 performs scale transformation on the gain data at an auditory scale through data interpolation. For example, the axis transformation unit 215 interpolates the low-frequency band data with wide data intervals at the auditory scale, thereby making the low-frequency band data denser. Data that is equally spaced at the auditory scale becomes data with dense low-frequency bands and sparse high-frequency bands at the linear scale (linear scale). Thus, the axis transformation unit 215 can generate equally spaced axis-transformed data at the auditory scale. Of course, the axis-transformed data may not be data that is completely equally spaced at the auditory scale.

[0061] In this way, the axis transformation data has spectral data based on the frequency characteristics of the input signal. Figure 4This is a graph representing an example of the spectral data of axis transformation data. In Figure 4 In the graph, the horizontal axis represents frequency [Hz], and the vertical axis represents amplitude (gain) [dB]. Additionally, in... Figure 4 The image shows the waveforms of the spectrum before and after peak compression processing (before and after peak suppression processing) in this embodiment.

[0062] The extremum extraction unit 216 calculates the extrema of the spectral data based on its frequency characteristics. That is, the extremum extraction unit 216 calculates the extrema of the axis-transformed data. The extrema of the spectral data correspond to the peaks or troughs of the spectral data. Specifically, maxima correspond to peaks, and minima correspond to troughs.

[0063] use Figure 5 This illustrates an example of how the extreme value extraction unit 216 calculates maxima (peaks), but it can also extract minima (troughs). For example... Figure 5 As shown, the extremum extraction unit 216 extracts six maxima as peaks P1 to P6. Peaks P1 to P6 each have peak frequency (the center frequency of the peak, i.e., the frequency at the maximum value) and gain data at the peak frequency.

[0064] Alternatively, the extremum extraction unit 216 can extract the extrema of the entire frequency band of the spectrum, but it can also extract only the extrema of a portion of the frequency band. Here, the extremum extraction unit 216 extracts only the maxima of a portion of the frequency amplitude characteristic. For example... Figure 5 As shown, the extreme value search band, which is the target of the compression process, is preset. The extreme value extraction unit 216 only searches for peaks within the extreme value search range. That is, the extreme value extraction unit 216 does not extract extreme values ​​outside the extreme value search range. Therefore, peak compression, which will be described later, is not performed outside the extreme value search range.

[0065] Furthermore, in the above description, the extremum extraction unit 216 extracts the extrema of the amplitude spectrum of the axis-transformed data. However, the extremum extraction unit 216 can also extract the extrema of the frequency amplitude characteristics before the axis transformation is performed by the axis transformation unit 215. The spectrum data processed by the extremum extraction unit 216 is not limited to axis-transformed data, as long as it is frequency-based spectrum data. For example, the extremum extraction unit 216 can also extract the extrema of spectrum data for which the frequency characteristics of the input signal or the axis-transformed data have been smoothed.

[0066] The kurtosis calculation unit 217 calculates the kurtosis of each of the peaks P1 to P6. Kurtosis is an indicator of the steepness of a peak. For example, the higher the kurtosis, the steeper the peak; the lower the kurtosis, the wider the peak. An example of a method for calculating kurtosis will be explained below.

[0067] In this embodiment, the kurtosis calculation unit 217 uses a kurtosis function to calculate kurtosis. The kurtosis function is a function that calculates an evaluation value used to evaluate a peak. Specifically, the kurtosis function uses a frequency function and a gain function to calculate the evaluation value. The evaluation value is a value used to evaluate the peak corresponding to an extreme value. Specifically, the extreme value is a value used to evaluate the shape and kurtosis of the peak.

[0068] The evaluation value is represented by the product of the gain function and the frequency function (gain function * frequency function). Then, the kurtosis calculation unit 217 calculates the kurtosis function based on the evaluation value. Specifically, the kurtosis calculation unit 217 obtains the kurtosis function using the following equation (1).

[0069] Kurtosis function = max{gain function * frequency function}……(1)

[0070] The gain function and frequency function are calculated separately for each peak. Examples of calculating the gain function and frequency function are explained below. Furthermore, the following explanation will focus on the case where the position on the frequency axis (frequency position) is represented by the order (integers) of the data. For example, in the spectrum of discrete axis-transformed data, the order of the data counted starting from the lowest frequency represents the frequency position.

[0071] The values ​​of the gain function and frequency function of a peak are determined by the computational width W. n And it changes. The computational width W n This represents how far away from the peak frequency (extreme value). Since the frequency position is represented by integers, W... n 1 or higher and W std (W std Integers in the range below 2 (where 2 is an integer). std It represents the operation width W. n The integer representing the maximum width (maximum value). W std This can be set by the user. std This is a parameter related to the frequency width of the peak to be detected. Users simply need to set W based on the maximum amplitude of the peak they want to detect. std That's all.

[0072] Set the frequency at the maximum value as the peak frequency f. p From the peak frequency f p Leave the operation width W n The frequency is set to f n Since the positions on the frequency axis of discrete spectral data are represented in the order of the data, therefore f n =(f p +W n ) or f n =(f p -W nFor example, in discrete spectral data, if the peak frequency f p Let's say it's the 1000th data point starting from the low-frequency side, when W... n When f = 100, then f n =900 or f n =1100.

[0073] The kurtosis calculation unit 217 gradually changes the calculation width W. n The values ​​of the gain function and frequency function are calculated simultaneously. That is, the kurtosis calculation unit 217 calculates according to 1, 2, ..., W. std Change the order of operations width W n Calculate the gain function and frequency function within their respective operational widths.

[0074] An example of the gain function is represented by the following equation (2).

[0075] Gain function = {(G p -G n ) / G std}^G m ……(2)

[0076] G p It is the peak frequency f p The gain [dB] at the peak (maximum) is the gain [dB]. n It is frequency f n Gain [dB] at lower values. When the operation width W is changed. n At that time, the gain G n Changes. Generally speaking, if the computation width W n If it increases, the gain G n It gets smaller.

[0077] G std This is a parameter related to the gain intensity you want to detect (gain reference value). When G std When the value increases, low peaks cannot be detected; when the value decreases, low peaks can be detected. Among them, G... std The closer the gain function is to 0, the closer it gets to infinity. Therefore, the balance between the gain function and the frequency function is disrupted, which requires attention. m This is the parameter (gain multiplier) that determines the gain sensitivity of the peak to be detected. If G m An increase in gain will be more sensitively reflected in areas with a steeper slope of gain.

[0078] G n By changing the operation width W n And the variable that changes. G p It is a fixed value determined for each wave peak. G m It is a fixed value set by the user. For example, G stdThe user determines the maximum width W. std In this case, a fixed value is determined for each peak. Furthermore, G std It doesn't have to be a fixed value determined by each peak. For example, G std It can be a value set by the user.

[0079] use Figure 6 The operation of the gain function is illustrated with examples. Figure 6 W is shown n Example of operation when =100. Figure 6 It means Figure 5 The diagram shows the spectrum surrounding peak P2. As mentioned above, the peak frequency f p The gain at point G is p The peak frequency f will be... p Leave 100 (=W) n Let the frequency of each be f. n(n=100) The frequency f n(n=100) The gain is set to G. n(n=100) .

[0080] exist Figure 6 In the middle, f n(n=100) =(f p +W n Therefore, (f) p +W n The gain at point G is n(n=100) That is, do not use (f) p -W n The gain at f. n(n=100) Let it be (f) p +W n ), or set as (f p -W n This can be determined simply by comparing their respective gains. Specifically, (f) p +W n ) and (f p -W n The frequency of the gain component with larger gain is taken as f. n(n=100) For example, in Figure 6 In the middle, (f p +W n The gain at point (f) is greater than (f) p -W n The gain when f is reached. Therefore, let f be the gain when f is reached. n(n=100) =(f p +W n ), in (f p +W n Set the gain to gain G at point ) n .

[0081] Similarly, the maximum width W std The frequency below is set to f n(n=Wstd) The frequency f n(n=Wstd) The gain is set to G. n(n=Wstd) .exist Figure 6 In the middle, f n(n=Wstd) =(f p +W std That is, do not use (f) p -W std The gain in ). Therefore, (f p +W std The gain at point G is n(n=Wstd) .

[0082] f n(n=Wstd) Let it be (f) p +W std ), or set to (f) p -W std This can be determined simply by comparing their respective gains. Specifically, (f) p +W std ) and (f p -W std The frequency of the one with the larger gain in the equation is taken as f. n(n=Wstd) For example, in Figure 6 In the middle, (f p +W std The gain at point (f) is greater than (f) p -W std The gain under f. Therefore, let f n(n=Wstd) =(f p +W std ), with (f p +W std Set the gain to gain G. n(n=Wstd) G std =G p -G n(n=Wstd) .

[0083] An example of a frequency function is represented by the following equation (3).

[0084] Frequency function = {(W std -W n ) / W std}^W m ……(3)

[0085] W m It determines the computational width W of the peak to be detected. n The sensitivity parameter (frequency multiplier). If W m Increasing W only results in a narrow peak. m When it becomes smaller, it will also be reflected in the width of the peak. Wm It can be set to a fixed value. W std This is a parameter related to the frequency width of the peak to be detected. Users simply need to set W based on the maximum amplitude of the peak they want to detect. std That's all.

[0086] Users can preset W according to the desired peak shape for compression. m G m W std G std Parameters such as W are used. Specifically, the user adjusts W based on the steepness of the wave peak to be compressed. m G m W std G std .

[0087] Kurtosis calculation unit 217 changes calculation width W n Calculate the frequency function and gain function using G. m With W m The shape of the peak with higher kurtosis is determined by the balance.

[0088] Kurtosis calculation unit 217 changes calculation width W n The frequency function and gain function are calculated. That is, the kurtosis calculation unit 217 will be combined with the computational width W. n The corresponding frequency f n and its gain G n Substitute the value of W into equations (2) and (3). Here, W m G m W std It is a fixed value set by the user. When the user sets W... std At that time, G is determined according to each peak. std Of course, users can also set G. std The value of .

[0089] Therefore, the kurtosis calculation unit 217 calculates the values ​​of the frequency function and the gain function for a given operational width. The kurtosis calculation unit 217 adjusts the operational width W... n In W1~W std Increase each range by 1 to calculate W. std The value of the frequency function and W std The value of the gain function. Typically, when the computational width W is increased... n When the frequency function decreases, the gain function increases.

[0090] The kurtosis calculation unit 217 calculates the product of the frequency function and the gain function as the evaluation value. std Each evaluation value. The kurtosis calculation unit 217, as shown in equation (1), calculates W. stdThe maximum value among the evaluation values ​​is used as the kurtosis.

[0091] Thus, the kurtosis calculation unit 217 calculates based on the frequency f. n Gain G in n The evaluation value of the evaluation peak is calculated. Then, the kurtosis calculation unit 217 calculates the kurtosis value based on the change in the calculation width W. n The kurtosis of the peaks is calculated from the calculated evaluation values. Furthermore, the kurtosis calculation unit 217 calculates the kurtosis for each of the peaks P1 to P6. Here, the kurtosis calculation unit 217 extracts the six peaks P1 to P6, and therefore calculates six kurtosis values.

[0092] The determination unit 218 determines whether to compress the peaks on a per-peak basis based on kurtosis. The determination unit 218 compares the kurtosis of the peaks with a threshold and determines whether to compress based on the comparison result. If the kurtosis of a peak is above the threshold, the determination unit 218 determines that the peak should be compressed. If the kurtosis of a peak is below the threshold, the determination unit 218 determines that the peak should not be compressed.

[0093] The compression unit 219 compresses the peaks determined to be compressed. That is, the compression unit 219 compresses peaks with a kurtosis of 1 / 3 or higher. For example, the compression unit 219 uses a polynomial curve to replace the peak with a characteristic that attenuates the peak. As a result, steep peaks can be suppressed.

[0094] For example, compression unit 219 suppresses peaks by substituting a Bezier curve, which is calculated based on three points obtained by multiplying the two endpoints of the peak and the maximum point of the peak by a predetermined attenuation coefficient. This reduces the gain of the peak. Furthermore, this method is an example of substituting, and the characteristics of the substituting are not limited to the calculation results obtained from the Bezier curve. The two endpoints of the peak can, for example, be set by the calculation width when the evaluation value reaches its maximum.

[0095] Furthermore, while the above description describes gain compression of peaks, gain compression of troughs (minimum values) can also be performed. In this case, the extremum extraction unit 216 extracts the minimum value as the trough. The kurtosis calculation unit 217 calculates the kurtosis for each extracted trough. The kurtosis calculation unit 217 calculates the kurtosis by performing peak processing with positive and negative reversals. As a result, the gain of the troughs increases.

[0096] The axis transformation unit 220 performs axis transformation through data interpolation, etc., to transform the frequency axis of the spectral data with compressed peaks. The processing in the axis transformation unit 220 is the reverse of the processing in the axis transformation unit 215. By performing axis transformation, the axis transformation unit 220 returns the frequency axis of the spectral data to the frequency axis before the axis transformation in the axis transformation unit 215. For example, processing is performed to return the frequency axis, which was logarithmic in the axis transformation unit 215, to linear. The spectral data with compressed peaks is divided into equally spaced data along the frequency linear axis. As a result, the frequency amplitude characteristics of the frequency axis are the same as those of the frequency phase characteristics obtained by the frequency characteristic acquisition unit 214. That is, the frequency axis (data interval) of the spectral data with frequency phase characteristics and frequency amplitude characteristics are consistent.

[0097] The filter generation unit 221 generates a filter using the spectral data that has undergone axis transformation by the axis transformation unit 220. The filter generation unit 221 transforms the frequency characteristics represented by the compressed amplitude spectrum into a time-domain signal. Here, the frequency characteristics have frequency amplitude characteristics and frequency phase characteristics. The frequency amplitude characteristics can be obtained using the compressed amplitude spectrum. The frequency phase characteristics can be obtained using the frequency phase shift characteristics obtained by the frequency characteristics acquisition unit 214.

[0098] The filter generation unit 221 generates a filter for the reproduced signal based on the peaked spectral data compressed by the compression unit 219. For example, the filter generation unit 221 calculates the time-domain signal based on the frequency amplitude and phase characteristics using an inverse discrete Fourier transform or a discrete cosine transform. The filter generation unit 221 generates a time signal by performing an IFFT (Inverse Fast Fourier Transform) on the frequency amplitude and phase characteristics. The filter generation unit 221 cuts the generated time signal to a predetermined filter length, thereby calculating the filter. The filter generation unit 221 can also perform windowing to generate the filter.

[0099] The filter generated by the filter generation unit 221 is set as an inverse filter. Figure 1 The processing unit 201 generates an inverse filter Linv by performing the above-described processing on the pickup signal picked up by the left microphone 2L. The processing unit 201 also generates an inverse filter Rinv by performing the above-described processing on the pickup signal picked up by the right microphone 2R. The inverse filters Linv and Rinv are respectively set in... Figure 1 The filter sections 41 and 42.

[0100] Thus, in this embodiment, the processing device 201 changes the calculation width W through the kurtosis calculation unit 217. n Multiple evaluation values ​​are calculated for a single peak. Furthermore, the kurtosis calculation unit 217 calculates values ​​based on changing the computation width W. nThe kurtosis is calculated using multiple evaluation values. For example, the kurtosis calculation unit 217 calculates the maximum value of the multiple evaluation values ​​as the kurtosis. This allows for appropriate compression of peaks. Because peaks of various shapes can be evaluated appropriately, steep peaks can be removed. This provides stable sound quality and sound field. It also enables the generation of a robust filter that does not become unstable even when headphones are worn again.

[0101] The external auditory canal transmission characteristics of subject 1 are measured using microphone unit 2 and headphones 43. Furthermore, the processing device 201 can be a smartphone or the like. Therefore, the measurement settings may differ for each measurement. Additionally, the wearing position of headphones 43 and microphone unit 2 may also introduce deviations. For example, the wearing position of headphones 43 during measurement may sometimes differ from the wearing position of headphones 43 during external head positioning listening. The processing device 201 suppresses peaks or troughs as described above. Thus, deviations caused by measurement, etc., can be suppressed, generating an inverse filter for the external auditory canal transmission characteristics.

[0102] In the compression unit 219, the filter generation unit 221 generates a filter using the corrected spectrum for peak suppression. This effectively suppresses the peaks generated in the inverse filters Linv and Rinv. Therefore, more appropriate inverse filters Linv and Rinv can be generated.

[0103] Implementation Method 2

[0104] use Figure 7 The processing apparatus and processing method of this embodiment will be described. Figure 7 This is spectral data used to illustrate the processing in this embodiment. Specifically, Figure 7 This is a magnified graph showing the periphery of two closely spaced peaks, P4 and P5. The horizontal axis represents frequency [Hz], and the vertical axis represents amplitude (gain) [dB].

[0105] In Implementation 2, in addition to the processing in Implementation 1, a process for merging two closely spaced peaks is added. Except for the merging process, it is the same as in Implementation 1, therefore, the description is omitted.

[0106] The processing unit 201 calculates the frequency distance between the peaks of the two maxima extracted by the extremum extraction unit. When the frequency distance between the peaks is below a frequency threshold, the processing unit 201 merges the two peaks. Specifically, in Figure 7 In this context, the frequency at which peak P4 becomes a maximum is set as f4, and the frequency at which peak P5 becomes a maximum is set as f5. The frequency distance is (f5-f4). f5-f4 is represented by integers indicating the order of the data. Furthermore, the frequency distance is the distance on the frequency axis after axis transformation by the axis transformation unit 215.

[0107] The frequency distance (f5-f4) between adjacent peaks P4 and P5 is below a threshold. Therefore, the processing device 201 merges peaks P4 and P5 into one peak. The peak frequency of the merged peak can be the frequency between peak frequency f4 and peak frequency f5, or it can be the same as peak frequency f4 or peak frequency f5.

[0108] The interpolation method here can be linear interpolation or polynomial interpolation. Of course, other interpolation methods besides linear or polynomial interpolation can also be used. This allows for appropriate compression of the peaks.

[0109] Figure 8 This is a flowchart illustrating the processing method of this embodiment. First, the frequency characteristic acquisition unit 214 acquires the frequency characteristics of the input signal (S801). For example, the input signal in the time domain is transformed into the frequency domain using FFT or the like. The input signal is, for example, an inverse filter that eliminates the transmission characteristics of the external auditory canal. The axis transformation unit 215 performs an axis transformation on the frequency characteristics (S802). As a result, spectral data in which the frequency axis of the pickup signal is transformed into a logarithmic axis is obtained.

[0110] The extreme value extraction unit 216 extracts the extreme values ​​(S803). This extracts the peak corresponding to the maximum value. Next, the kurtosis calculation unit 217 calculates the kurtosis of the peak (S804). As described above, the kurtosis calculation unit 217 changes the calculation width W... n The value of the kurtosis function is calculated as an evaluation value. The kurtosis calculation unit 217 calculates the kurtosis of the peak.

[0111] The determination unit 218 determines whether the kurtosis is 0.5 or higher (S805). Here, the threshold for kurtosis is 0.5, but the threshold is not limited to 0.5. Preferably, it is a value of 0.5 or higher, and more preferably, it is a value of 0.5 to 0.8. If the kurtosis is not 0.5 or higher (S805 "No"), the processing device 201 determines whether the processing for all peaks has ended (S809). If the processing for all peaks has not ended (S809 "No"), the process returns to S804 and repeats the processing. That is, the kurtosis calculation unit 217 calculates the kurtosis of the next peak.

[0112] If the kurtosis of the wave crest is 0.5 or higher (S805 "Yes"), it is determined whether the frequency distance between wave crests is 100 or lower (S806). Here, the frequency threshold for the frequency distance between wave crests is 100, but the frequency threshold can also be a value other than 100.

[0113] When the frequency distance between wave peaks is 100 or less (S806 "Yes"), the compression unit 219 merges the wave peaks (S807). Thus, two wave peaks are merged into one. Then, the compression unit 219 compresses the merged wave peak (S808). When the frequency distance between wave peaks is not 100 or less (S806 "No"), the compression unit 219 compresses the wave peak (S808).

[0114] In S808, when the compression section 219 compresses the peak, the processing device 201 determines whether the processing for all peaks has ended (S809). If the processing for all peaks has not ended (S809 "No"), the process returns to S804 and repeats. If the processing for all peaks has ended (S809 "Yes"), the processing device 201 ends the processing.

[0115] Therefore, peaks can be appropriately compressed. Furthermore, in the above explanation, when the kurtosis is above a threshold and the frequency distance between peaks is below a frequency threshold, the compression unit 219 merges the peaks. However, peaks can also be merged when the kurtosis is less than the threshold. That is, when the frequency distance between peaks is below the frequency threshold, the compression unit 219 can merge peaks regardless of kurtosis. Additionally, as in Embodiment 1, steps S806 and S807 are omitted when merging two peaks is not performed.

[0116] Other implementation methods

[0117] In embodiments 1 and 2, the processing device 201 compresses the peaks of the spectral data based on the pickup signal, but it can also compress the troughs corresponding to the minimum values. In this case, the processing device 201 extracts the minimum values ​​and calculates the same kurtosis for the troughs corresponding to the minimum values. At this time, the processing device 201 performs positive and negative reversal processing.

[0118] Furthermore, in embodiments 1 and 2, the processing device 201 processes the spectral data of the input signal corresponding to the inverse filter of the external auditory canal transmission characteristics. However, it can also process the spectral data of the input signal based on the spatial acoustic transmission characteristics Hls, Hlo, Hro, and Hrs. In addition, although the processing device 201 generates an external head positioning processing filter, it can also generate other filters. For example, the processing device 201 can also generate a noise suppression filter that suppresses peaks or troughs. Furthermore, peak or trough compression processing can also be applied to processing other than filter generation. For example, the processing device 201 can also perform noise reduction processing without using a filter.

[0119] The external head positioning processing device 100 or processing device 201 is not limited to a single physical device, but can also be distributed among multiple devices connected via a network or the like. In other words, the external head positioning processing method or processing method of this embodiment can also be implemented in multiple distributed devices.

[0120] Some or all of the above processes can also be executed by a computer program. The program described above includes a set of commands (or software code) for causing the computer to perform one or more functions described in the embodiments, when read into a computer. The program can be stored in a non-transitory computer-readable medium or a physical storage medium. Not limited to this, examples of computer-readable media or physical storage media include random access memory (RAM), read-only memory (ROM), flash memory, solid-state drives (SSDs) or other memory technologies, CD-ROMs, digital versatile optical discs (DVDs), Bluetooth (Blu-ray, registered trademark) discs or other optical disc storage, magnetic disks, magnetic tapes, disk storage, or other magnetic storage devices. The program can be transmitted on a temporary computer-readable medium or communication medium. Not limited to this, examples of temporary computer-readable media or communication media include electrical, optical, acoustic, or other formats of propagated signals.

[0121] The invention described above is based on specific embodiments, but the invention is not limited to the above embodiments, and various modifications can be made without departing from its spirit.

[0122] Symbol Explanation

[0123] U User

[0124] 1. Test subject

[0125] 2 microphone units

[0126] 2L Left Microphone

[0127] 2R Right Microphone

[0128] 5 stereo speakers

[0129] 5L Left Speaker

[0130] 5R Right Speaker

[0131] 10-head external positioning processing unit

[0132] 11 Convolution Operation Unit

[0133] 12 Convolution Operation Unit

[0134] 21 Convolution Operation Unit

[0135] 22 Convolution Operation Unit

[0136] 24 Adders

[0137] 25 Adders

[0138] 41 Filter Section

[0139] 42 Filter Section

[0140] 43. Headphones

[0141] 200 Measuring device

[0142] 201 Processing Unit

[0143] 211 Measurement Signal Generation Unit

[0144] 212 Sound signal acquisition unit

[0145] 213 Inverse Filter Generation Unit

[0146] 214 Frequency Response Acquisition Unit

[0147] 215-axis transformation section

[0148] 216 Extreme Value Extraction Section

[0149] 217 Kurtosis Calculation Department

[0150] 218 Judgment Department

[0151] 219 Compression Section

[0152] 220 axis conversion unit

[0153] 221 Filter Generation Unit

Claims

1. A processing apparatus comprising: a frequency characteristic acquisition section that acquires a frequency characteristic of an input signal; an extreme value extraction section that extracts an extreme value of spectral data based on the frequency characteristic; a kurtosis calculation section that calculates a kurtosis of a peak or a valley corresponding to the extreme value from spectral data in a calculation width including the extreme value, and calculates the kurtosis based on a plurality of evaluation values calculated by changing the calculation width; a determination section that determines whether to compress the peak or the valley according to a comparison result of the kurtosis and a threshold value; a compression section that compresses the peak or the valley of the extreme value determined to be compressed; an inverse filter generation section that generates an inverse filter that cancels an external ear canal transfer characteristic as the input signal based on a picked-up signal picked up by a microphone worn on an ear of a subject; and a filter generation section that generates a filter applied to a reproduced signal based on spectral data having the peak or the valley compressed by the compression section.

2. The processing apparatus according to claim 1, wherein in a case where a frequency distance between two maximum values extracted by the extreme value extraction section or a frequency distance between two minimum values is equal to or less than a frequency threshold value, the two peaks or the two valleys are merged.

3. The processing apparatus according to claim 1 or 2, further comprising: a first axis transformation section that transforms a frequency axis of the frequency characteristic acquired by the frequency characteristic acquisition section by data interpolation; and a second axis transformation section that transforms a frequency axis of the spectral data having the compressed peak or the valley by data interpolation, the filter generation section generating the filter based on the spectral data subjected to the axis transformation by the second axis transformation section.

4. A processing method comprising the steps of: acquiring a frequency characteristic of an input signal; extracting an extreme value of the frequency characteristic; calculating an evaluation value that evaluates the extreme value from data in a calculation width including the extreme value, and calculating a kurtosis of the extreme value based on a plurality of evaluation values calculated by changing the calculation width; determining whether to compress the extreme value according to a comparison result of the kurtosis and a threshold value; compressing the extreme value determined to be compressed; generating an inverse filter that cancels an external ear canal transfer characteristic as the input signal based on a picked-up signal picked up by a microphone worn on an ear of a subject; and generating a filter applied to a reproduced signal based on spectral data having the compressed extreme value.

5. The processing method according to claim 4, wherein in a case where a frequency distance between two maximum values extracted by the extreme value extraction section or a frequency distance between two minimum values is equal to or less than a frequency threshold value, the two peaks or the two valleys are merged. ​

Citation Information

Patent Citations

  • Processing device, processing method, regeneration process, and program

    JP2020136752A

  • Noise suppressor for removing irregular noise

    CN101131819A

  • Filter generation device, filter generation method, and program

    CN110301142A