Processing device and processing method

The processing device addresses the challenge of generating well-balanced filters for out-of-head localization by using a combination of frequency characteristic acquisition, smoothing, compression, and filter generation units, resulting in balanced sound image localization and natural sound quality.

JP2025085766APending Publication Date: 2025-06-05JVC KENWOOD CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025044723
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing out-of-head localization processing technologies face challenges in generating well-balanced filters due to steep peaks and dips in frequency amplitude characteristics, which can result in signal clipping and loss of individual characteristics, leading to imbalanced sound image localization.

Method used

The proposed processing device includes a frequency characteristic acquisition unit, a smoothing processor, an adjustment level calculation unit, a compression unit, and a filter generating unit, which work together to acquire, smooth, and compress frequency characteristics, and generate a filter that balances the spectral data while maintaining individual characteristics.

Benefits of technology

This approach enables the generation of well-balanced filters that prevent signal clipping and maintain the balance of sound image localization, ensuring natural sound quality and accurate out-of-head localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085766000001_ABST
    Figure 2025085766000001_ABST
Patent Text Reader

Abstract

To provide a processing device and a processing method capable of generating a well-balanced filter.SOLUTION: A processing device includes a frequency characteristic acquisition unit 214 for acquiring the frequency characteristics of a sound pickup signal, a smoothing processing unit 215 for generating smoothed spectral data by smoothing spectral data based on the frequency characteristics, an adjustment level calculation unit 217 for calculating an adjustment level on the basis of the smoothed spectral data in a first band B1, a compression unit 218 for generating compressed spectral data by compressing the smoothed spectral data in a second band B2 including the lower frequency and upper frequency of the first band B1 using the adjustment level, and a filter generation unit 221 for generating a filter on the basis of the compressed spectral data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a processing device and a processing method. [Background technology]

[0002] As a sound image localization technology, out-of-head (out-of-head) sound localization is a technique that uses headphones to localize a sound image outside the listener's head. There is a technology called localization. In out-of-head localization technology, the characteristics from the headphones to the ears (headphone characteristics) are used Cancels out the noise from one speaker (monaural speaker) to the ears, and reduces the noise generated by the two speakers (spatial acoustics) By giving the head a special characteristic called a "transfer characteristic," the sound image is positioned outside the head.

[0003] In out-of-head localization playback using stereo speakers, two channels (hereafter referred to as ch) are used. The measurement signal (impulse sound, etc.) emitted from the speaker was placed at the listener's ear. The measurement signal is recorded by a microphone. The processor generates a filter based on the collected sound signal. The generated filter is applied to the 2ch audio. By convolving this with the audio signal, out-of-head localization playback can be achieved.

[0004] In addition, a filter (also called an inverse filter) is used to cancel the characteristics from the headphones to the ears. In order to generate the characteristic from the headphones to the ear and eardrum (ear canal transfer function ECTF The ear canal transfer characteristics (also called ear canal transfer characteristics) are measured using a microphone placed in the listener's own ear.

[0005] Patent Document 1 discloses an apparatus for performing extra-head localization processing. The out-of-head localization processing applies DRC (Dynamic Range Comparator) to the playback signal. However, before the DRC processing, the processing device The signal characteristics are smoothed. The processing device then performs band division based on the smoothed characteristics. It is. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] JP 2019-62430 A Summary of the Invention [Problem to be solved by the invention]

[0007] The out-of-head localization processing uses a spatial acoustic filter obtained from the spatial acoustic transfer characteristics of the number of speakers. It uses an inverse filter calculated from the headphones' ECTF. To achieve this, it is advisable to use a spatial acoustic filter that is as measured as possible and an accurate inverse filter. This is the ideal.

[0008] However, the frequency amplitude characteristics measured using a microphone showed a steep peak. Peaks (narrow bands with very high levels) and dips (narrow bands with very low levels) occur. This often results in the processed signal clipping.

[0009] The levels and frequencies of the peaks and dips vary depending on a number of factors. For example, The level may vary depending on the characteristics of the speaker at the position, the acoustic characteristics of the room, the characteristics of the headphones, etc. The level and frequency also change depending on the shape of an individual's head and ears. For this reason, we check the characteristics of the equipment used at the time of measurement and make adjustments to suit that equipment. I had to listen to it while checking it.

[0010] Therefore, if the correction amount (compression amount) in the compression process is too large, the individual characteristics of each individual may be lost. The balance is lost. Therefore, the balance of the localization is lost, and the effect of out-of-head localization is lost. There is a risk that this will be damaged.

[0011] The present disclosure has been made in consideration of the above points, and aims to generate a well-balanced filter. The object of the present invention is to provide a processing apparatus and a processing method capable of doing so. [Means for solving the problem]

[0012] The processing device according to the present embodiment includes a frequency characteristic acquisition unit for acquiring the frequency characteristic of a picked-up signal. and smoothing the spectrum data based on the frequency characteristics to obtain a smoothed spectrum data. a smoothing processor for generating a smoothed spectrum data in a first band; an adjustment level calculation unit that calculates an adjustment level for the first band using the adjustment level; The smoothed spectrum data in a second band including a lower limit frequency and an upper limit frequency of a compression unit for generating compressed spectral data by compressing the spectral data; and a filter generating unit that generates a filter based on the

[0013] The processing method according to the present embodiment includes the steps of acquiring a frequency characteristic of a picked-up signal; The spectrum data based on the frequency characteristics is smoothed to obtain smoothed spectrum data. generating a smoothed spectral data in a first band based on the smoothed spectral data in the first band; calculating an adjustment level; and adjusting a lower limit frequency and a lower limit frequency of a first band using the adjustment level. and compressing the smoothed spectral data in a second band including an upper frequency limit. generating compressed spectral data; and and generating a filter using the resulting signal. Effect of the Invention

[0014] According to the present disclosure, there is provided a processing device capable of generating a well-balanced filter, and A method can be provided. [Brief description of the drawings]

[0015] [Figure 1] 1 is a block diagram showing an extra-head localization processing device according to an embodiment of the present invention; [Diagram 2] FIG. 2 is a diagram illustrating a schematic configuration of a measuring device. [Diagram 3] FIG. 2 is a block diagram showing a configuration of a processing device. [Figure 4] 11 is a graph showing an example of spectrum data obtained from frequency amplitude characteristics. [Diagram 5] FIG. 13 is a diagram for explaining a process of compressing smoothed spectral data. [Figure 6] FIG. 11 is a diagram for explaining a process of correcting a third band and a fourth band. [Figure 7] 1 is a flowchart illustrating a processing method according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] An overview of the sound image localization process according to this embodiment will be described. The out-of-head localization process uses the spatial acoustic transfer characteristics and the ear canal transfer characteristics to perform out-of-head localization process. The spatial acoustic transfer characteristic is the transfer characteristic from a sound source such as a speaker to the ear canal. The transfer characteristic is the transfer characteristic from the speaker unit of the headphones or earphones to the eardrum. In this embodiment, the spatial acoustic transmission characteristics when headphones or earphones are not worn are measured. The transmission characteristics of the ear canal are measured while wearing headphones or earphones. The measurement data is used to realize out-of-head localization processing. It is characterized by a microphone system for measuring the transfer characteristic or ear canal transfer characteristic.

[0017] The out-of-head localization processing according to the present embodiment is carried out on a personal computer, a smartphone, a tablet, The user terminal is implemented by a processing means such as a processor, a memory, etc. Storage means such as a memory or hard disk, display means such as an LCD monitor, touch panel, buttons A user terminal is an information processing device having input means such as a keyboard and a mouse. The user terminal may further include a headphone or An output means (output unit) having an earphone is connected. The connection may be a wired or wireless connection.

[0018] Embodiment 1 (Extracranial Localization Processing Device) A block diagram of an out-of-head localization processing device 100, which is an example of a sound field reproduction device according to the present embodiment. The figure is shown in FIG. 1. The out-of-head localization processing device 100 performs the following operations on a user U wearing headphones 43: To this end, the out-of-head localization processing device 100 uses Lch and Rch stereo inputs. The Lch and Rch stereo input signals X are processed for sound image localization. L and XR are analog audio signals output from a CD (Compact Disc) player, etc. It is a playback signal or digital audio data such as mp3 (MPEG Audio Layer-3). The audio playback signal and digital audio data are collectively referred to as the playback signal. In other words, the Lch and Rch stereo input signals XL and XR are the playback signals.

[0019] The head localization processing device 100 is not limited to a single physical device. For example, some of the processing may be performed by a smartphone. The remaining processing is done by a DSP (Digital Signal Processor) built into the headphones 43. It may be performed by

[0020] The head-external localization processing device 100 includes a head-external localization processing unit 10, a filter that stores an inverse filter Linv, and The inverse filter Rinv is stored in a filter unit 42, and a headphone 43 is provided. The head outside localization processing unit 10, the filter unit 41, and the filter unit 42 are specifically This can be realized by using a processor or the like.

[0021] The head-outside localization processing unit 10 stores the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs. The convolution calculation unit 11-12, 21-22, and adders 24 and 25 are provided. The calculation units 11 to 12 and 21 to 22 perform convolution processing using the spatial acoustic transfer characteristics. The localization processing unit 10 receives stereo input signals XL and XR from a CD player or the like. The out-of-head localization processing unit 10 is set with spatial acoustic transfer characteristics. For the stereo input signals XL and XR of each channel, a filter of spatial acoustic transfer characteristics (hereinafter, The spatial acoustic transfer characteristic is determined by the head and pinna of the subject. It can be a measured head-related transfer function (HRTF), a dummy head, or a third-party head-related transfer function. It's fine.

[0022] A set of four spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs is The data used for convolution in the convolution calculation units 11, 12, 21, and 22 are the acoustic transfer functions. The spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are given. A spatial audio filter is generated by cutting out the signal at a fixed filter length.

[0023] The spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs are impulse response measurements. For example, when a user U wears a microphone on each of his / her left and right ears, The left and right speakers placed in front of the user U are used to measure the impulse response. The impulse sounds output from the speakers are The measurement signal is picked up by a microphone. Based on the signal picked up by the microphone, the spatial acoustic transfer characteristic Hl s, Hlo, Hro, Hrs are obtained. The spatial acoustic transfer between the left speaker and the left microphone The spatial acoustic transfer characteristic Hls between the left speaker and the right microphone, and the spatial acoustic transfer characteristic Hlo between the right speaker and the left microphone. Spatial acoustic transmission characteristic Hro between the left speaker and the right microphone, and spatial acoustic transmission characteristic Hro between the right speaker and the right microphone Hrs is measured.

[0024] The convolution calculation unit 11 performs spatial acoustic transfer on the Lch stereo input signal XL. The convolution calculation unit 11 performs convolution calculation using a spatial acoustic filter according to the characteristic Hls. The convolution calculation unit 21 outputs the data to the adder 24. A spatial audio filter corresponding to the spatial audio transfer characteristic Hro is convolved with the 21 outputs the convolution operation data to the adder 24. The adder 24 performs two convolution operations. The data is added and output to the filter unit 41.

[0025] The convolution calculation unit 12 calculates a spatial acoustic transfer characteristic Hl for the Lch stereo input signal XL. The convolution calculation unit 12 convolves the spatial acoustic filter according to o. , and outputs it to the adder 25. The convolution calculation unit 22 outputs the following to the Rch stereo input signal XR: The convolution calculation unit 22 convolves a spatial acoustic filter according to the spatial acoustic transfer characteristic Hrs. The adder 25 outputs the two convolution operation data. The result is then added to the signal, and output to the filter unit 42.

[0026] The filter sections 41 and 42 are for detecting headphone characteristics (characteristics between the playback unit of the headphones and the microphone). The inverse filters Linv and Rinv are set to cancel the head deviation. The playback signal (convolution signal) processed by the phase processor 10 is filtered by an inverse filter Linv. The filter unit 41 convolves the Lch signal from the adder 24 with Rinv. Similarly, the filter unit 42 convolves the inverse filter Linv of the headphone characteristic of The Rch signal from 5 is convolved with the inverse filter Rinv of the Rch headphone characteristics. The inverse filters Linv and Rinv are used to detect the The microphone cancels the characteristics from the ear canal entrance to the eardrum. If so, you can place it anywhere.

[0027] The filter unit 41 outputs the processed Lch signal YL to the left unit 43L of the headphone 43. The filter unit 42 outputs the processed Rch signal YR to the right unit of the headphone 43. The user U is wearing the headphones 43. The headphones 43 are ch signal YL and Rch signal YR (hereinafter, Lch signal YL and Rch signal YR are collectively referred to as the stage The headphone amplifier outputs a stereo signal (also called a headphone signal) to the user U. This allows the headphone amplifier to localize the position outside the user U's head. The resulting sound image can be reproduced.

[0028] In this way, the out-of-head localization processing device 100 calculates the spatial acoustic transfer characteristics Hls, Hlo, Hro, The spatial acoustic filter according to Hrs and the inverse filter Linv, Rinv of the headphone characteristics are In the following description, the spatial acoustic transfer characteristics Hls, H Spatial acoustic filters according to lo, Hro, and Hrs, and an inverse filter Lin for headphone characteristics v and Rinv are collectively used as the out-of-head localization processing filter. In the case of a 2-channel stereo playback signal, In this case, the out-of-head localization filter is composed of four spatial acoustic filters and two inverse filters. The out-of-head localization processing device 100 outputs a total of six out-of-head localization signals to the stereo playback signal. The localization filter is used to perform convolution processing to perform out-of-head localization processing. The position filter is preferably based on the personal measurements of the user U. For example, The out-of-head localization filter is set based on the sound signal picked up by the microphone attached to the U's ear. is.

[0029] In this way, the spatial acoustic filter and the inverse filter Linv and Rinv of the headphone characteristics are These filters are for the playback signal (stereo input signal XL , XR), the head-outside localization processing device 100 executes head-outside localization processing. In this embodiment, the process of generating a spatial acoustic filter is one of the technical features. Specifically, in the process of generating a spatial acoustic filter, the spectrum in the frequency characteristic Level range control processing that compresses the range of the data gain level (Level R The LRC processing is applied to the frequency response. The level range between the minimum and maximum gain levels of the spectrum data is called the level It's called Lurange.

[0030] (Spatial acoustic transmission characteristic measuring device) Measurement device for measuring spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs using Figure 2 FIG. 2 is a schematic diagram of a measurement configuration for measuring a subject 1. In this case, the subject 1 is the same person as the user U in FIG. However, they may be different people.

[0031] As shown in FIG. 2, the measurement device 200 has a stereo speaker 5 and a microphone unit 2. The stereo speakers 5 are installed in the measurement environment. The measurement environment is the home of a user U. The measurement environment may be a room, a sales store or a showroom of an audio system. It is preferable to have a listening room with good lighting and acoustics.

[0032] In this embodiment, the processing device 201 of the measurement device 200 appropriately generates a spatial acoustic filter. The processing unit 201 performs calculations to generate a sound from, for example, a CD player. The processing device 201 includes a personal computer (PC), The processing device 201 may be a tablet terminal, a smartphone, or the like. It may be itself.

[0033] The stereo speaker 5 includes a left speaker 5L and a right speaker 5R. The left speaker 5L and the right speaker 5R are installed in front of the attendant 1. The speaker 5R outputs impulse sounds for measuring the impulse response. In the embodiment, the number of speakers serving as sound sources will be described as two (stereo speakers). The number of sound sources used for measurement is not limited to two, but may be one or more. Or, in a so-called multi-channel environment such as 5.1ch or 7.1ch, this implementation The form can be applied.

[0034] The microphone unit 2 is a stereo microphone having a left microphone 2L and a right microphone 2R. The left microphone 2L is placed at the left ear 9L of the subject 1, and the right microphone 2R is placed at the Specifically, the left ear 9L and the right ear 9R are placed from the entrance of the ear canal to the eardrum. It is preferable to place microphones 2L and 2R at the position shown in Fig. 1. The measurement signal output from the speaker 5 is picked up to obtain a picked-up signal. The collected sound signal is output to the processing device 201. The subject 1 may be a human or a dummy head. That is, in this embodiment, the subject 1 includes not only a human but also a dummy head. It is a concept.

[0035] As shown above, the impulse sound output from the left speaker 5L and the right speaker 5R is The impulse response is measured by measuring L and 2R. The pickup signal obtained by the response measurement is stored in a memory etc. This allows the left speaker 5L and the spatial acoustic transfer characteristic Hls between the left speaker 5L and the right microphone 2R. The spatial acoustic transfer characteristic Hlo between the right speaker 5R and the left microphone 2L is H ro, and the spatial acoustic transfer characteristic Hrs between the right speaker 5R and the right microphone 2R are measured. That is, the measurement signal output from the left speaker 5L is picked up by the left microphone 2L, The acoustic transfer characteristic Hls is obtained. The measurement signal output from the left speaker 5L is input to the right microphone 2 The spatial acoustic transfer characteristic Hlo is obtained by picking up the sound from the right speaker 5R. The left microphone 2L picks up the measurement signal thus obtained, and the spatial acoustic transfer characteristic Hro is obtained. The measurement signal output from the right speaker 5R is picked up by the right microphone 2R, and spatial acoustic transmission is performed. The characteristic Hrs is obtained.

[0036] Based on the collected sound signals, the measuring device 200 also measures left and right signals from the left and right speakers 5L and 5R. Spatial acoustic transmission characteristics Hls, Hlo, Hro, Hrs for microphones 2L and 2R For example, the processing device 201 may generate an acoustic filter. The processor 201 extracts Hlo, Hro, and Hrs with a predetermined filter length. The inter-channel acoustic transfer characteristics Hls, Hlo, Hro, and Hrs may be corrected.

[0037] In this way, the processing device 201 performs the convolution operation of the head localization processing device 100. As shown in FIG. 1, the head-outside localization processing device 100 generates a spatial acoustic filter to be used. The spatial acoustic transfer characteristic Hl between the left and right speakers 5L, 5R and the left and right microphones 2L, 2R The head-outside localization process is performed using spatial acoustic filters according to s, Hlo, Hro, and Hrs. That is, by convolving the spatial acoustic filter with the audio playback signal, out-of-head localization processing is performed. conduct.

[0038] The processing device 201 performs the following for each of the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs: The same processing is performed on the corresponding pickup signal. The same processing is performed on the four pickup signals corresponding to Hlo, Hro, and Hrs. This allows the spatial acoustic transfer characteristics Hls, Hlo, Hro, and Hrs to be Inter-acoustic filters can be generated respectively.

[0039] The processing device 201 of the measuring device 200 and its processing will be described in detail below. 2 is a control block diagram showing a processing device 201. The processing device 201 includes a measurement signal generating unit 21. 1, a pickup signal acquisition unit 212, a frequency characteristic acquisition unit 214, a smoothing processing unit 215, and an axis A conversion unit 216, an adjustment level calculation unit 217, a compression unit 218, a correction processing unit 219, and an axis conversion unit The filter generating unit 221 includes a conversion unit 220 .

[0040] The measurement signal generating unit 211 includes a D / A converter, an amplifier, and the like, and measures the ear canal transfer characteristics. A measurement signal is generated for measurement. The measurement signal is, for example, an impulse signal or a TSP (T Here, the measurement signal is an I-mode Stretched Pulse signal. Using an impulse sound, the measurement device 200 performs an impulse response measurement.

[0041] The left microphone 2L and the right microphone 2R of the microphone unit 2 pick up the measurement signal, and the picked-up signal The pickup signal acquisition unit 212 outputs a pickup signal to the processing device 201. The collected sound signal acquisition unit 212 acquires the collected sound signals from the microphones 2L and 2R. The collected sound signal acquisition unit 212 may include an A / D converter that converts the collected sound signal from the collected sound signal into a digital signal. Alternatively, signals obtained by multiple measurements may be synchronously added.

[0042] The frequency characteristic acquisition unit 214 acquires the frequency characteristics of the picked-up signal. 4 calculates the frequency characteristics of the picked-up signal using discrete Fourier transform and discrete cosine transform. The frequency characteristic acquisition unit 214 performs, for example, FFT (Fast Fourier Transform) on the picked-up sound signal in the time domain. The frequency characteristics are calculated by the amplitude spectrum and the phase spectrum. The frequency characteristic acquisition unit 214 includes a power spectrum instead of an amplitude spectrum. A vector may be generated.

[0043] The smoothing processor 215 performs a smoothing process on the spectrum data based on the frequency characteristics. The smoothing processor 215 uses a moving average, a Savitzky-Golay filter, a smoothing spline, Smoothing of the spectral data using techniques such as the cepstrum transform and the cepstrum envelope The spectrum data smoothed by the smoothing processing unit 215 is called smoothed spectrum data. The smoothing processor 215 smoothes the spectrum data based on the frequency characteristics to obtain a smoothed spectrum. Generate smoothed spectral data.

[0044] The axis conversion unit 216 converts the frequency axis of the smoothed spectrum data by data interpolation. The axis converter 216 converts the discrete spectrum data into equal intervals on the logarithmic axis. The scale of the data of the frequency amplitude characteristic is changed. Amplitude characteristic spectrum data and smoothed spectrum data (hereinafter also referred to as gain data) ) are equally spaced in frequency. That is, the gain data is Since the frequency is evenly spaced, the logarithmic frequency axis is not evenly spaced. 216 is a frequency logarithmic axis, the gain data is spaced equally on the frequency logarithmic axis. Then, the interpolation process is performed.

[0045] In gain data, the lower the frequency range is on the logarithmic axis, the larger the gap between adjacent data points is. The intervals are coarse, and the higher the frequency, the closer the adjacent data intervals become. The axis conversion unit 216 interpolates data in the low frequency band where the data intervals are coarse. The axis conversion unit 216 performs an interpolation process such as a three-dimensional spline interpolation to convert the logarithmic axis into a uniform Obtain discrete gain data arranged in intervals. The gain data that has undergone axis transformation is The axis conversion data is a data in which frequency and amplitude value (gain value) are associated with each other. The axis-transformed data is the smoothed spectrum data that has been axis-transformed. .

[0046] The reason for converting the frequency axis to a logarithmic scale is explained below. Generally, human sensory quantities are converted to logarithmic scales. It is said that the frequency of the sound is converted. Therefore, the frequency of the sound that we hear can also be thought of on a logarithmic axis. It is important to convert the scale so that the data are equally spaced for the above sensory quantities. This allows data to be treated equally across all frequency bands. This makes it easier to divide and weight the regions, and makes it possible to obtain stable results. The unit 216 converts the envelope data into a scale close to human hearing (called a hearing scale) instead of the logarithmic scale. The data can be converted into a logarithmic scale, a mel scale, etc. ) scale, Bark scale, ERB (Equivalent Rectangular Bandwidth) scale, etc. You can also perform axis conversion using

[0047] The axis conversion unit 216 performs scale conversion of the gain data on an auditory scale by data interpolation. For example, the axis conversion unit 216 interpolates data in a low frequency band where the data intervals are sparse in the auditory scale. By doing so, the data in the low frequency band is made dense. Data that is equally spaced on the auditory scale is linearly scaled. In the linear scale, the low frequency band is dense and the high frequency band is sparse. By doing so, the axis transformation unit 216 can generate axis transformation data that is equally spaced on the auditory scale. Of course, axis-transformed data do not have to be perfectly evenly spaced on the auditory scale. good.

[0048] The adjustment level calculation unit 217 calculates the smoothed spectrum data in the first band B1 based on the smoothed spectrum data in the first band B2. The adjustment level is calculated by, for example, adjusting the smoothed spectrum in the first band B1. In other words, the adjustment level calculation unit 217 can calculate the average level of the vector data. Calculate the sum of the gains of the smoothed spectrum data included in band B1 of 1. Adjustment level The calculation unit 217 calculates the adjustment level by dividing the sum by the number of data included in the first band B1. Calculate the rule.

[0049] An example of calculation of the adjustment level is shown in FIG. 4. FIG. 4 shows the smoothed spectrum data A sm and adjustment level RuA ave 4 is a graph showing a schematic diagram of the frequency [Hz] of the sine wave. The amplitude value (gain) is [dB]. Here, the smoothed spectrum data A sm year The axis conversion data converted by the axis conversion unit 216 is used, but the axis conversion process is omitted. That is, the axis-transformed smoothed spectrum data A sm The average level of A aveFor example, adjustment level A ave = 3 dB. In the band B1 of 1, the smoothed spectrum data A sm The average gain is 3 dB. is.

[0050] The first band B1 can be, for example, 5 kHz to 10 kHz. The lower limit frequency f of band B1 1S is 5kHz, and the upper frequency f 1E is 10kHz As will be described later, the average level is based on the sound signals picked up by the left and right microphones 2L and 2R. Alternatively, the average value of the smoothed spectrum data may be used.

[0051] The compression unit 218 adjusts the audio signal at adjustment level A. ave Using the smoothed spectrum in the second band B2, The smoothed spectrum data compressed by the compression unit 218 is compressed. For example, the compressor 218 may calculate the gain of the smoothed spectrum data as Adjustment level A ave Then, a predetermined compression coefficient is applied to the difference value. The compression unit 218 multiplies the smoothed spectrum in the second band B2 to calculate a compressed value. The compression value is subtracted from the gain of the spectral data. is generated.

[0052] FIG. 5 is a graph for explaining the LRC processing in the compression unit 218. Data lrc , smoothed spectrum data A sm [dB], adjust the level to A ave [dB] and the compression coefficient is lrcRate. The LRC processing in the compression unit 218 is expressed by the following formula ( 1) is shown. A lrc =A sm -lrcRate*(A sm -A ave ) · · · (1)

[0053] The compression unit 218 calculates the gain of the smoothed spectral data included in the second band B2 using the formula (1 ) is compressed based on the smoothed spectrum data A sm Since the gain value differs for each frequency, Therefore, compressed spectrum data A lrc The gain value for each frequency is different. Rate*(A sm -A ave The compression unit 218 calculates the difference between the frequency bands. For each frequency, compressed spectrum data A lrc Calculate the gain value of The compressor 218 compresses the smoothed spectrum data with a different compression value for each frequency.

[0054] The compression factor lrcRate can be a constant value. For example, the compression factor lrcRa te can be a value greater than 0 and less than or equal to 1. Here, the compression factor lrcRate = 0.5. Adjustment level A ave = 3[dB]. A sm = 5[dB] The compression value is 0.5*(5-3)=1[dB], and the compressed spectrum data A lrc =5 -1=4[dB].

[0055] In this way, the compressor 218 adjusts the smoothed spectral data in the second band B2. Level A ave In other words, the compression unit 218 corrects the smoothed spectrum The data is compressed to approximate the adjustment level. In other words, the compressed spectrum data is smoothed. The value is between the calculated spectral data and the adjustment level.

[0056] At frequencies where the smoothed spectrum data is greater than the adjustment level, the compressed spectrum data The smoothed spectrum data is adjusted to a smaller value than the smoothed spectrum data. At frequencies smaller than the threshold, the compressed spectral data is larger than the smoothed spectral data. This allows the smoothed spectrum data to be generated while maintaining the individual characteristics. In the second band B2, the compression factor lrcRate is constant. Therefore, the greater the difference value from the adjustment level, the greater the compression.

[0057] The second band B2 can be, for example, 1 kHz to 20 kHz. The lower limit frequency f of band B2 2S is 1kHz, and the upper frequency f 2E is 20kHz It should be noted that the first band B1 and the second band B2 are not limited to the above ranges. For example, the first band B1 is a band in which gain fluctuations due to individual frequency characteristics are particularly large among the second band B2. This allows the frequency band to be adjusted so that it is possible to balance the individual characteristics of each individual. The range of spectral data can be compressed without compromising the quality of the data.

[0058] The second band B2 may be a band that coincides with the first band B1 or may be a different band. The first band B1 and the second band B2 may be bands that partially overlap each other. The first band B1 may be included in the second band B2. number f 1S is the lower limit frequency f of the second band B2 2S Above, the upper limit frequency f 2E The following shall be done: The upper limit frequency f 1Eis the lower limit frequency f of the second band B2 2S Below Upper, upper limit frequency f 2E It can be as follows:

[0059] The correction processing unit 219 adjusts the gain in the vicinity of the second band B2 compressed by the compression unit 218. The compressed spectrum data is corrected so that the difference does not change suddenly. The processing unit 219 processes the compressed spectrum data (smoothed) in the third band B3 and the fourth band B4. The gain of the spectral data is corrected.

[0060] As shown in FIG. 6, the third band B3 is an offset band on the lower frequency side of the second band B2. The third band B3 is adjacent to the second band B2. The fourth band B4 is The fourth band B4 is an offset band on the higher frequency side of the second band B2. It is the band adjacent to 2.

[0061] For example, the third band B3 is 900 Hz to 1 kHz, and the fourth band B4 is 20 k The lower limit frequency f of the third band B3 is Hz to 21 kHz. 3S is 900Hz, and the upper limit frequency f 3E = 1 kHz. The upper limit frequency f 3E is the second band B2 The lower limit frequency f 2S The lower limit frequency f 4S = 20kHz Yes, upper limit frequency f 4E = 21 kHz. The lower limit frequency f 4S The Upper frequency f of band B2 2E It matches.

[0062] The correction processor 219 corrects the smoothed spectrum data of the third band B3. Specifically, the lower limit frequency f 2S The gain does not change suddenly around As shown above, the correction processor 219 corrects the gain of the third band B3.

[0063] For example, the correction processor 219 detects a lower limit frequency f 3S and upper frequency f 3E Between The gain is corrected so that the spectral data is smoothly connected. is the lower limit frequency f 3S and upper frequency f 3E The curve between is interpolated using a sine curve or other curve. Specifically, the correction processor 219 calculates the lower limit frequency f 3S Gain and upper frequency f 3E in The gain of the third band B3 is connected by a curve such as a sine function or a polynomial curve. Alternatively, the correction processor 219 interpolates the lower limit frequency f 3S Gain and upper limit wave number f 3E Linear interpolation may be used to connect the gains at the lower limit in a straight line. frequency f 3S to upper frequency f 3E As the gain increases, Alternatively, the gain is corrected so as to gradually decrease.

[0064] Alternatively, the compression coefficient lrcRate in Eq. (1) may be changed gradually. Band B3 may be compressed. In this case, the lower limit frequency f 3S to upper frequency f 3E Towards The correction processor 219 performs smoothing by using a gradually increasing compression coefficient lrcRate. Compresses spectral data, e.g., the lower frequency limit f 3S The compression factor at is 0, and the upper limit wave number f 3EThe compression factor at is 0.5. In this case, the lower limit frequency f 3S From upper limit wave number f 3E Set the compression factor lrcRate to gradually increase from 0 to 0.5 toward the Determine.

[0065] lower limit frequency f 3S to upper frequency f 3E The correction process is performed so that the image is gradually compressed toward the In other words, the upper limit frequency f 3E to the lower limit frequency f 3S To The correction processor 219 corrects the gain so that compression is gradually reduced toward the center of the input signal.

[0066] Similarly, the correction processor 219 corrects the smoothed spectrum data of the fourth band B4. Specifically, the upper limit frequency f 2E The gain changes rapidly around To prevent this, the correction processor 219 corrects the gain of the fourth band B4.

[0067] For example, the correction processor 219 detects a lower limit frequency f 4S and upper frequency f4 E Between The gain is corrected so that the spectral data is smoothly connected. is the lower limit frequency f 4S and upper frequency f 4E The distance between the points is interpolated using straight lines and curves. The lower limit frequency f 4S to upper frequency f 4E The gain gradually increases as The gain is corrected so that it increases or decreases gradually.

[0068] Alternatively, the compression coefficient lrcRate in Eq. (1) may be changed gradually. Band B3 may be compressed. In this case, the lower limit frequency f 4S to upper frequency f 4E Towards The correction processing unit 219 smoothes the spectrum data by using a compression coefficient that gradually decreases. In this way, the upper limit frequency f 4E to the lower limit frequency f 4S Gradually compressed towards In other words, the correction processor 219 adjusts the gain so that the lower limit frequency f 4S to upper frequency f 4E The correction processing unit 219 gradually reduces compression toward the Correct the gain.

[0069] The spectrum data corrected by the correction processing unit 219 is called corrected spectrum data. The corrected spectral data of the third band B3 and the fourth band B4 are added to the smoothed spectral data. The data is corrected by the correcting processor 219. The spectrum data is the same as the compressed spectrum data. The corrected spectrum data is the gain value generated by the compression process of the compression unit 218. The lower limit frequency f of the third band B3 3S In the lower frequency band, the corrected spectrum data The data is the same as the smoothed spectrum data. f 4E In the higher frequency band, the corrected spectrum data is It is the same data.

[0070] The axis conversion unit 220 converts the frequency axis of the corrected spectrum data by data interpolation or the like. The process in the axis converter 220 is the same as that in the axis converter 216. The axis conversion unit 220 converts the axis of the corrected spectrum data. The wave number axis returns to the frequency axis before axis conversion by the axis conversion unit 216. The frequency axis is converted to a linear scale and then converted to a linear scale. The data is spaced equally on the wavenumber linear axis. It is possible to obtain the same frequency amplitude characteristic on the frequency axis as the frequency phase characteristic. The frequency axis (data interval) of the spectrum data of the phase characteristics and frequency amplitude characteristics coincides.

[0071] The filter generating unit 221 uses the corrected spectrum data that has been axis-converted by the axis conversion unit 220. The filter generating unit 221 generates a filter based on the corrected spectrum data. The filter generator 221 generates a filter to be applied to the reproduced signal. For example, the filter generator 221 generates an inverse discrete filter. The time domain signal is calculated from the amplitude and phase characteristics by the Lambertian transform or the inverse discrete cosine transform. The filter generation unit 221 performs IFFT (inverse fast Fourier transform) on the amplitude and phase characteristics. ) to generate a time signal. The filter generating unit 221 filters the generated time signal by a predetermined The filter generating unit 221 calculates a spatial acoustic filter by cutting out the signal with a filter length of Windowing may be performed to generate a spatial audio filter.

[0072] The filter generation unit 221 generates a filter by converting a measurement signal from the left speaker 5L into a signal picked up by the left microphone 2L. By performing the above processing on the sound signal, the spatial audio filter corresponding to the spatial audio transfer characteristic Hls is obtained. The filter generation unit 221 generates a filter by applying a measurement signal from the left speaker 5L to the right microphone 2. By carrying out the above processing on the sound signal picked up at L, the spatial acoustic transfer characteristic Hlo is obtained. The spatial acoustic filter is generated.

[0073] The filter generation unit 221 generates a filter by converting a measurement signal from the right speaker 5R into a signal picked up by the left microphone 2L. By performing the above processing on the sound signal, the spatial acoustic transfer characteristic Hro is obtained. The filter generation unit 221 generates a filter by filtering the measurement signal from the right speaker 5R with the right microphone 2. By carrying out the above processing on the sound signal picked up by R, the spatial acoustic transfer characteristic Hrs is obtained. The spatial acoustic filter is generated.

[0074] By doing this, the frequency characteristics can be compressed in a well-balanced manner. It is possible to generate a filter suitable for image localization. In other words, while maintaining individual characteristics, It is possible to compress the frequency characteristics of the user. It is possible to prevent the loss of balance in the sound image localization. It is possible to localize a well-balanced sound image. Well-balanced sound quality This allows the generation of filters that are adjusted to ensure a natural sound quality. Cut.

[0075] The adjustment level calculation unit 217 calculates the signal from the left microphone 2L and the right microphone 2R. The adjustment level may be calculated from the spectrum data based on the frequency characteristics of the As shown in the figure, the left microphone 2L and the right microphone 2R measure the pickup signal. 212 acquires two pickup signals in one measurement. For example, measurement from the left speaker 5L When the signal is output, the picked-up sound signal acquisition unit 212 outputs the picked-up sound signal corresponding to the spatial acoustic transfer characteristic Hls. The signal and the pickup signal corresponding to the spatial acoustic transfer characteristic Hlo are obtained. Then, the adjustment level The calculation unit 217 calculates a common adjustment level for the left and right from the smoothed spectrum data of the two collected sound signals. Calculate.

[0076] The frequency characteristic acquisition unit 214 calculates the frequency characteristics of the two picked-up signals. The smoothed spectrum data of the sound signal picked up by the 2L is smL and recorded with the right microphone 2R. The smoothed spectrum data of the recorded sound is smR Smoothed spectrum data A smL The adjustment level obtained from A aveL Then, the smoothed spectrum data A smR from The obtained adjustment level is A aveR Here, adjustment level A aveL The first band Smoothed spectrum data A in area B1 smL Adjustment level A aveR is the smoothed spectrum data A in the first band B1. smR The average value of the If it is possible to calculate a stable adjustment level that can be obtained regardless of the frequency balance of individual characteristics, For example, the adjustment level may be a representative value such as the median, or a statistical value. The adjustment level may be a combination of statistics, such as the mean plus the standard deviation.

[0077] The adjustment level calculation unit 217 calculates the left and right adjustment levels A aveL , A aveR Overall adjustment from Level A ave For example, the left and right common adjustment level A ave is expressed by the following formula (2): As shown. A ave =(A aveL +A aveR ) / twenty two)

[0078] The adjustment level for the spectrum data based on the sound signal picked up by the left speaker 5L and the right speaker The adjustment level for the spectrum data based on the sound pickup signal of the 5R will be the same. This makes it possible to more appropriately compress the second band B2.

[0079] Compressed spectrum data based on the signal picked up by the left microphone 2L is A lrcL Then, the right microphone The compressed spectrum data obtained from the sound signal of Q2R is lrcR Then, the compression unit 21 The LRC processing in 8 is shown in the following equations (3) and (4). A lrcL =A smL -lrcRate*(A smL -A ave ) · · · (3) A lrcR =A smR -lrcRate*(A smR -A ave ) · · · (4)

[0080] The filter generation unit 221 generates the compressed spectrum data A lrcL Based on spatial acoustic transmission The filter generating unit 221 generates a filter corresponding to the characteristic Hls. Data A lrcR Based on this, a filter corresponding to the spatial acoustic transfer characteristic Hlo is generated.

[0081] Similarly, for the measurement using the right speaker 5R, the adjustment level calculation unit 217 calculates a common left and right The filter generation unit 221 calculates the adjustment level of the compressed spectrum data A lrcL To The filter generating unit 22 generates a filter corresponding to the spatial acoustic transfer characteristic Hro based on the 1 is compressed spectrum data A lrcR Based on this, the spatial acoustic transfer characteristic Hrs is In this way, the adjustment level calculation unit 217 generates a filter for the left microphone 2L and the right microphone The adjustment level is calculated from the smoothed spectrum data of the sound signal picked up by 2R. , a more appropriate adjustment level can be used to compress the spectral data. A balanced filter can be generated.

[0082] FIG. 7 is a flowchart showing a processing method according to this embodiment. The acquisition unit 214 acquires the frequency characteristics of the picked-up signal acquired by the picked-up signal acquisition unit 212 (S 701). For example, the time domain signal is transformed into the frequency domain by FFT. Next, the smoothing processor 215 performs a smoothing process on the spectrum data (S702). This results in smoothed spectral data.

[0083] The axis conversion unit 216 converts the axis of the smoothed spectrum data (S703). The spectrum data is obtained by converting the frequency axis of the collected sound signal into a logarithmic axis. The axis conversion process by the axis conversion unit 220 described later can be omitted. No conversion process is required.

[0084] Next, the adjustment level calculation unit 217 calculates the first band The average level of the left and right adjustment levels A and B is calculated (S704). aveL , A aveR Next, the adjustment level calculation unit 217 calculates the average of the left and right Adjust the average level Level A ave (S705). This allows for a common adjustment for both the left and right. Adjustment level A ave If different adjustment levels are used for the left and right, step S7 Step 05 can be omitted.

[0085] Next, the compressor 218 uses the adjustment level to compress the smoothed spectral data of the second band B2. Specifically, the compression unit 218 compresses the data according to the above formulas (3) and (4) as follows: Based on this, compressed spectral data is generated.

[0086] The correction processing unit 219 corrects the offset band (S707). 19 corrects the compressed spectrum data of the third band B3 and the fourth band B4. The axis converter 220 converts the axis of the corrected spectrum data. The filter generating unit 221 performs the conversion based on the corrected spectrum data after the axis conversion (S708). A filter is generated based on the spatial acoustic transfer characteristics Hls and Hlo (S709). A spatial acoustic filter corresponding to the spatial acoustic transfer characteristics Hro and Hrs is generated. In this way, a well-balanced filter can be generated.

[0087] Furthermore, in the first and second embodiments, the processing device 201 determines the spatial acoustic transfer characteristics Hls, Hlo The spectral data of the collected sound signal showing Hro and Hrs was processed, but the ear canal transfer characteristics were not shown. The processing device 201 may process the spectrum data of the collected sound signal. However, other filters may be generated. By using the filters generated by this method, it is possible to localize a balanced sound image. do.

[0088] The head localization processing device 100 is not limited to a single physical device, but may be connected via a network or the like. In other words, the head-external determination in this embodiment may be distributed among multiple devices connected to the head. The position processing method may be carried out in a distributed manner by a plurality of devices.

[0089] Some or all of the above processes may be executed by a computer program. The above-mentioned program, when loaded into a computer, performs the steps described in the embodiment. or more functions on a computer The program is stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, a computer readable medium or tangible storage medium may be , random-access memory (RAM), read-only memory (ROM), flash memory, solid- solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD) , Blu-ray (registered trademark) disc or other optical disc storage, magnetic cassette, magnetic This includes magnetic tapes, magnetic disk storage or other magnetic storage devices. The ram may be transmitted over a transitory computer readable medium or a communication medium. By way of example only, transitory computer-readable media or communication media may be electrical, optical, acoustic, or any other form of propagated signal.

[0090] The present invention has been described above in detail based on the embodiments. The present invention is not limited to the above-described embodiment, and various modifications can be made without departing from the spirit and scope of the present invention. It goes without saying that this is the case. [Explanation of symbols]

[0091] U User 1 Person to be measured 2 Microphone Unit 2L Left microphone 2R Right Microphone 5 Stereo speakers 5L Left speaker 5R Right speaker 10. Extra-head localization processing unit 11 Convolution Calculation Unit 12 Convolution Calculation Unit 21 Convolution Calculation Unit 22 Convolution Calculation Unit 24 Adder 25 Adder 41 Filter section 42 Filter section 43 Headphones 200 Measuring Equipment 201 Processing equipment 211 Measurement signal generator 212 Sound collection signal acquisition unit 214 Frequency characteristic acquisition section 215 Smoothing processing unit 216 Axis conversion unit 217 Adjustment level calculation section 218 Compression section 219 Correction processing section 220 Axis conversion unit 221 Filter Generation Unit B1 First band B2 Second band B3 Third Band B4 Fourth band

Claims

1. A frequency characteristic acquisition unit for acquiring a frequency characteristic of a picked-up sound signal; The spectrum data based on the frequency characteristics is smoothed to obtain smoothed spectrum data. A smoothing unit that generates a smoothing signal; An adjustment level is calculated based on the smoothed spectral data in a first band. A level calculation unit; a compression unit that generates compressed spectral data by compressing the smoothed spectral data in a second band that includes a lower limit frequency and an upper limit frequency of the first band using the adjustment level; a filter generating unit that generates a filter based on the compressed spectral data.

2. The processing device according to claim 1 , wherein the first band is a band in which a frequency of gain fluctuations occurring in spectrum data based on the frequency characteristics is high among gain fluctuations in the entire second band.

3. a first axis conversion unit that converts a frequency axis of the smoothed spectrum data by data interpolation; 、 a second axis conversion unit that converts a frequency axis of the compressed spectrum data by data interpolation; Further equipped with The filter generating unit generates a filter based on the compressed spectrum data that has been axis-converted by the second axis conversion unit. The processing device according to claim 1 or 2, wherein the filter is generated based on the above.

4. In order to prevent the gain from suddenly changing on the high frequency side and the low frequency side of the second band, a correction processing unit that corrects the compressed spectral data in an offset band provided The processing apparatus according to any one of claims 1 to 3, further comprising:

5. The collected sound signals are collected by microphones attached to the left and right ears of the subject. And, The adjustment level calculation unit calculates the smoothed spectrum of the sound signals picked up by the left and right microphones.

5. The processing device according to claim 1, wherein the adjustment level is calculated from torque data. 。

6. A step of acquiring a frequency characteristic of a picked-up sound signal; The spectrum data based on the frequency characteristics is smoothed to obtain smoothed spectrum data. generating data; A step of calculating an adjustment level based on the smoothed spectral data in a first band. Tep and Using the adjustment level, a second frequency band including a lower frequency band and an upper frequency band is generated. compressing the smoothed spectral data in a band to generate compressed spectral data; generating a filter based on the compressed spectral data. Law.

Citation Information

Patent Citations

  • Sound field correction device, control method thereof, and program

    JP2015139060A

  • Processing device, processing method, and program

    JP2019062430A