Sound processing device, sound processing method and program

The sound processing device stabilizes acoustic transfer function estimation by calculating spatial spectra, removing outliers, and updating based on reliable signals, enhancing sound source localization and separation performance in dynamic environments.

JP7721089B2Active Publication Date: 2025-08-12HONDA MOTOR CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022197160
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-08-12
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing microphone array processing systems face challenges in accurately estimating acoustic transfer functions in dynamic usage environments due to changes in acoustic conditions, which affect sound source localization and separation performance, and remeasuring these functions is impractical.

Method used

A sound processing device that stabilizes the estimation of acoustic transfer functions by using a first function stored for each sound source direction, calculating spatial spectra, identifying a representative direction, removing outliers, and updating the transfer function based on reliable acoustic signals to adapt to changing environments.

Benefits of technology

The device enables stable estimation of acoustic transfer functions, improving sound source localization and separation performance by updating the transfer function based on reliable acoustic signals, reducing sudden fluctuations and maintaining system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007721089000006
    Figure 0007721089000006
  • Figure 0007721089000007
    Figure 0007721089000007
  • Figure 0007721089000008
    Figure 0007721089000008
Patent Text Reader

Abstract

To stably estimate an acoustic transfer function in an actual acoustic environment.SOLUTION: A sound source direction estimation unit is configured to calculate, for each frame, spatial spectrum in each sound source direction based on a first acoustic transfer function and a conversion coefficient in a frequency domain of acoustic signals in each channel, and estimate, as an estimated sound source direction, a sound source direction having the maximum spatial spectrum. A representative estimated sound source direction determination unit is configured to determine a representative estimated sound source direction which is a representative value of the estimated sound source directions, on the basis of a frequency distribution of the estimated sound source directions in an observation period, consisting of a plurality of frames. An outlier removal unit removes conversion coefficient of a frame in which the estimated sound source direction is out of an allowable range determined in advance from the representative estimated sound source direction. An acoustic transfer function estimation unit estimates, as a second acoustic transfer function, a representative value of acoustic transfer functions from a sound source to a sound collector of the acoustic signals in the observation period, on the basis of the conversion coefficients of the acoustic signals of the remaining frames, and updates the first acoustic transfer function for the representative estimated sound source direction, with the second acoustic transfer function.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound processing device, a sound processing method, and a program. [Background technology]

[0002] Microphone array processing is a fundamental technology in acoustic signal processing. Microphone array processing uses multi-channel acoustic signals collected using a microphone array, and examples of this include sound source localization and sound source separation. Sound source localization is a method for estimating the direction of a sound source from multi-channel acoustic signals. Sound source separation is a method for extracting components coming from individual sound sources from multi-channel acoustic signals. Sound source localization and sound source separation are useful for identifying individual sounds, such as when speech is spoken in a noisy environment or when there are multiple sound sources. Sound source localization and sound source separation are applied to a variety of applications, including robot audition, smart speakers, teleconferencing systems, and meeting minutes creation.

[0003] Microphone array processing uses an acoustic transfer function that indicates the transfer characteristics of sound from a sound source to a sound receiving point. The sound source transfer function may be geometrically calculated using a mathematical model assuming a free sound field, or may be measured in advance in a laboratory using sound sources installed in multiple directions. However, such an acoustic transfer function differs from that measured in the acoustic environment in which microphone array processing is actually used (sometimes referred to as the "usage environment" in this application). This may result in a decrease in the performance of sound source localization or sound source separation. It may be possible to measure the acoustic transfer function in advance in the usage environment to ensure the performance of sound source localization or sound source separation. The acoustic transfer function changes from the time of measurement as the acoustic environment changes, but re-measuring the acoustic transfer function requires a lot of time and effort. Therefore, re-measuring the acoustic transfer function in the usage environment is not practical. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Kazuhiro Nakadai, Masayuki Takigahira, Kumasuke Kawai, Hiroshi Nakajima, "Improvement of Sound Source Localization and Separation by Continuous Online Adaptation of Transfer Functions," Japanese Society for Artificial Intelligence, Second Type Research Meeting Materials, AI Challenge Research Group, SIG-Challenge-058-07,<https: / / doi.org / 10.11517 / jsaisigtwo.2021.Challenge-058_07> , released on December 16, 2021 Summary of the Invention [Problem to be solved by the invention]

[0005] Non-Patent Document 1 describes a method for estimating and sequentially updating an acoustic transfer function using acoustic signals acquired by a microphone array. This method allows for online acquisition of an acoustic transfer function using any sound source without the need to set up equipment for remeasurement. However, acoustic signals acquired in a usage environment are not necessarily suitable for acquiring an acoustic transfer function. For example, significant noise may be temporarily mixed into the acoustic signal. As a result, it has sometimes been difficult to stably estimate the acoustic transfer function.

[0006] The present embodiment has been made in consideration of the above points, and aims to provide a sound processing device, a sound processing method, and a program that can stably estimate an acoustic transfer function in a usage environment. [Means for solving the problem]

[0007] (1) The present application has been made to solve the above-mentioned problems, and one aspect of this embodiment is an audio processing device including: a storage unit that stores a first acoustic transfer function for each sound source direction; a sound source direction estimation unit that calculates, for each frame, a spatial spectrum for each sound source direction based on a transform coefficient in the frequency domain of an audio signal for each channel and the first acoustic transfer function, and estimates, as an estimated sound source direction, the sound source direction where the spatial spectrum is maximum; a representative estimated sound source direction determination unit that determines a representative estimated sound source direction that is a representative value of the estimated sound source directions based on a frequency distribution of the estimated sound source directions in an observation period consisting of a plurality of frames; an outlier removal unit that removes transform coefficients of frames where the estimated sound source direction falls outside a predetermined tolerance range from the representative estimated sound source direction; an acoustic transfer function estimation unit that estimates, as a second acoustic transfer function, a representative value of the acoustic transfer function from a sound source to a sound collection unit for the audio signal in the observation period based on the transform coefficients of the audio signal in the remaining frames; and an acoustic transfer function update unit that updates the first acoustic transfer function for the representative estimated sound source direction using the second acoustic transfer function.

[0008] (2) Another aspect of this embodiment is the sound processing device of (1), Representative estimated sound source direction determination unit The estimated sound source direction having the maximum frequency may be determined as the representative estimated sound source direction.

[0009] (3) Another aspect of this embodiment is the acoustic processing device of (1), wherein the acoustic transfer function update unit may update the weighted average value of the first acoustic transfer function and the second acoustic transfer function to a new first acoustic transfer function.

[0010] (4) Another aspect of this embodiment is the sound processing device of (1), wherein the acoustic transfer function update unit determines the reliability of the representative estimated sound source direction based on the frequency with which the estimated sound source direction during the observation period falls within the tolerance range, and the higher the reliability, the higher the ratio of the second acoustic transfer function to the first acoustic transfer function.

[0011] (5) Another aspect of the present embodiment is the sound processing device of (1), wherein the allowable range is equal to the representative estimated sound source direction and does not include a direction different from the representative estimated sound source direction.

[0012] (6) Another aspect of this embodiment is the sound processing device of (1), wherein the sound source direction estimation unit may calculate the spatial spectrum by multiplying an input vector including the conversion coefficient for each channel by a pseudo-inverse matrix of an acoustic transfer function vector including the first acoustic transfer function for each channel.

[0013] (7) Another aspect of this embodiment may be a program for causing a computer to function as the sound processing device of (1).

[0014] (8) Another aspect of the present embodiment is an acoustic processing method in an acoustic processing device including a storage unit that stores a first acoustic transfer function for each sound source direction, wherein the acoustic processing device executes the following acoustic processing steps: a sound source direction estimation step of calculating, for each frame, a spatial spectrum for each sound source direction based on a transform coefficient in the frequency domain of an acoustic signal for each channel and the first acoustic transfer function, and estimating the sound source direction where the spatial spectrum is maximum as an estimated sound source direction; a representative sound source direction determination step of determining a representative estimated sound source direction that is a representative value of the estimated sound source directions based on a frequency distribution of the estimated sound source directions in an observation period consisting of a plurality of frames; an outlier removal step of removing transform coefficients of frames where the estimated sound source direction falls outside a predetermined tolerance range from the representative estimated sound source direction; an acoustic transfer function estimation step of estimating, as a second acoustic transfer function, a representative value of the acoustic transfer function from a sound source to a sound collection unit for the acoustic signal in the observation period based on the transform coefficients of the acoustic signal in the remaining frames; and an acoustic transfer function updating step of updating the first acoustic transfer function for the representative estimated sound source direction using the second acoustic transfer function. [Effects of the Invention]

[0015] According to this embodiment, it is possible to stably estimate an acoustic transfer function in a real acoustic environment. According to the above-described configurations (1), (7), and (8), the second acoustic transfer function is calculated based on the conversion coefficient of the acoustic signal that gives an estimated sound source direction within a predetermined range from the representative estimated sound source direction, and the calculated second acoustic transfer function can be used to update the first acoustic transfer function in association with the representative estimated sound source direction. A representative value of the acoustic transfer function obtained based on the acoustic signal that statistically gives the representative estimated sound source direction or an estimated sound source direction approximate thereto is updated as the second acoustic transfer function in association with the representative estimated sound source direction, so that a first acoustic transfer function with a stable correspondence relationship with the sound source direction can be obtained.

[0016] According to the above-described configuration (2), the estimated sound source direction with the maximum frequency within the observation period is determined as the representative estimated sound source direction, and therefore the most likely estimated sound source direction is simply determined as the representative estimated sound source direction.

[0017] According to the above-mentioned configuration (3), when the observation period is changed, the first acoustic transfer function is not completely replaced by the second acoustic transfer function through updating, but some of its components remain. This avoids sudden fluctuations in the first acoustic transfer function, thereby ensuring system stability.

[0018] According to the above-described configuration (4), the first acoustic transfer function can be updated using the second acoustic transfer function, with emphasis being placed on acoustic signals that provide highly reliable estimated sound source directions, thereby improving the reliability of the updated first acoustic transfer function.

[0019] According to the above-mentioned configuration (5), it is possible to simply determine whether or not to exclude the transform coefficient of the acoustic signal that gives the estimated sound source direction, depending on whether or not the estimated sound source direction is equal to the representative estimated sound source direction.

[0020] According to the above-mentioned configuration (6), the sound source direction can be estimated based on the spatial spectrum calculated by a simple matrix operation. Since it does not require many calculation resources, it can be realized economically. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a schematic block diagram illustrating an example of the configuration of a sound processing system according to an embodiment of the present invention. [Figure 2] 10 is a data flowchart showing an example of acoustic transfer function adaptation processing according to the present embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of the configuration of a sound collection unit. [Figure 4] FIG. 2 is an explanatory diagram illustrating an acoustic transfer function. [Figure 5] FIG. 1 illustrates a laboratory. [Figure 6] FIG. 10 is a diagram illustrating an example of a success rate for each observation period. [Figure 7] FIG. 10 is a diagram illustrating an example of the success rate for each type of first acoustic transfer function. [Figure 8] FIG. 1 is a schematic block diagram showing an example of the configuration of a sound processing system according to a first modified example of the present embodiment. [Figure 9] FIG. 10 is a schematic block diagram showing an example of the configuration of a sound processing system according to a second modified example of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0022] (First embodiment) A first embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a schematic block diagram showing an example of the configuration of a sound processing system S1 according to this embodiment. The sound processing system S1 includes a sound processing device 10 and a sound collection unit 20.

[0023] A first acoustic transfer function indicating the transfer characteristics of sound from a sound source is stored in advance for each sound source direction in the sound processing device 10. In the present application, the acoustic transfer function stored in advance or temporarily in the sound processing device 10 is referred to as the "first acoustic transfer function," and the acoustic transfer function estimated based on the acoustic signal acquired from the sound collection unit 20 is referred to as the "second acoustic transfer function," thereby distinguishing between the two.

[0024] The sound processing device 10 acquires acoustic signals of multiple channels from the sound collection unit 20. The sound processing device 10 calculates a transform coefficient in the frequency domain of the acoustic signal for each channel for each frame, and calculates a spatial spectrum for each sound source direction based on the calculated transform coefficient and a first acoustic transfer function. The sound processing device 10 estimates the sound source direction where the spatial spectrum has a local maximum as the estimated sound source direction (sound source localization). The sound processing device 10 accumulates the estimated sound source direction for each frame and the acoustic signals used to estimate the estimated sound source direction, and generates a frequency distribution (histogram) of the estimated sound source direction over an observation period consisting of multiple frames. The sound processing device 10 determines a representative value of the estimated sound source directions as the representative estimated sound source direction based on the frequency distribution of the estimated sound source directions.

[0025] The sound processing device 10 identifies frames of the acoustic signal that provide an estimated sound source direction that is outside a predetermined tolerance range from the representative estimated sound source direction, and removes the transform coefficients of the acoustic signal of those frames as outliers. The sound processing device 10 calculates an acoustic transfer function from the sound source to the sound collection unit 20 based on the transform coefficients that remain without being removed. The sound processing device 10 estimates a representative value of the acoustic transfer function during the observation period as a second acoustic transfer function. The sound processing device 10 updates the first acoustic transfer function for the determined representative estimated sound source direction using the estimated second acoustic transfer function.

[0026] The sound processing device 10 may perform microphone array processing using the updated first acoustic transfer function. The microphone array processing may include processing such as sound source separation in addition to sound source localization. The sound source separation includes processing of extracting sound components from individual sound sources as sound source components from the multi-channel sound signals acquired from the sound collection unit 20 based on the estimated sound source directions.

[0027] The sound processing device 10 may use one or both of the sound source direction estimated by performing sound source localization and the sound source signal indicating the sound source component extracted by performing sound source separation for other processing within the device itself, or may output them to another device (not shown, sometimes referred to as an "output destination device" in the present application) that serves as an output destination. As another processing, the sound processing device 10 may, for example, estimate the presence of an object in the estimated sound source direction. The sound processing device 10 may perform speech recognition processing on the sound source component or sound source signal from a specific sound source direction (speaker), and may acquire a spoken text indicating the content of the utterance or estimate the speaker.

[0028] The sound collection unit 20 has a plurality of microphones 20-1 to 20-M and functions as a microphone array. The number of microphones M is an integer equal to or greater than 2. The number of microphones M corresponds to the number of channels. Each microphone is arranged at a different position and has an actuator that collects sound waves arriving at that microphone. The actuator converts the arriving sound waves into an acoustic signal. The converted acoustic signal is output to the sound processing device 10 wirelessly or via a wired connection. Each microphone corresponds to a channel of the acoustic signal.

[0029] The arrangement of the multiple microphones may be fixed or variable. The positions of the multiple microphones may be different from one another. The sound collection unit 20 illustrated in FIG. 3 is configured as an eight-channel circular microphone array. In FIG. 3, each microphone is indicated by a black circle. The eight microphones are arranged at equal intervals on a circumference. The eight microphones are arranged on the side of a housing that forms a body of revolution that has rotational symmetry with respect to a rotation axis parallel to the vertical direction, and their relative positions are fixed. The sound collection unit 20 includes an output interface (not shown). The output interface aggregates the eight-channel acoustic signals recorded by the individual microphones and outputs them in parallel via a wired connection to the sound processing device 10. Note that the number and arrangement of the microphones are not limited to this. The number M of microphones may be two or more and seven or less, or nine or more. The positions of the individual microphones may be arranged in a straight line, as illustrated in FIG. 4.

[0030] The sound processing device 10 may be configured as a general-purpose information and communication device such as a PC (Personal Computer) or a multi-function mobile phone, or may be configured as a dedicated device such as a measuring instrument or a monitoring device. In the following description, the sound collection unit 20 is configured separately from the sound processing device 10 as an example, but may be separate from the sound processing device 10.

[0031] Next, an example of the functional configuration of the sound processing device 10 according to this embodiment will be described. The sound processing device 10 includes an input / output unit 110, a control unit 120, and a storage unit 150. The input / output unit 110 is connected wirelessly or by wire to other devices so as to input and output various types of data. The input / output unit 110 outputs M-channel acoustic signals from the sound collection unit 20 to the control unit 120 as input data. When output data is input from the control unit 120, the input / output unit 110 can output the input output data to an output destination device (not shown). The output data may include information obtained by microphone array processing. Such information includes, for example, estimated sound source direction information indicating an estimated sound source direction, a sound source signal indicating sound source components arriving from a sound source, and the like. The input / output unit 110 may be, for example, an input / output interface, a communication interface, or a combination thereof.

[0032] The control unit 120 executes processes for realizing the functions of the sound processing device 10, processes for controlling the functions, etc. The control unit 120 may be configured using dedicated components as a whole or for each individual function, or may be configured as a computer system including a processor such as a CPU (Central Processing Unit) and various storage media. The processor reads out a predetermined program stored in advance in a storage medium and executes processes instructed by various commands written in the read program to realize the functions of the control unit 120. The functional configuration of the control unit 120 will be described later.

[0033] The storage unit 150 includes a storage medium that temporarily or permanently stores various types of data. The storage unit 150 stores various types of data (including parameters, etc.) used by the control unit 120, and various types of data acquired by the control unit 120 or other functional units (including input data input from the outside, intermediate data during processing, and generated data generated as a processing result). The storage unit 150 stores an acoustic transfer function set. The acoustic transfer function set includes a first acoustic transfer function for each microphone (channel) for each frequency for each sound source direction. In this application, the sound source direction associated with the first acoustic transfer function in the acoustic transfer function set may be referred to as a "target direction."

[0034] The initial value of the first acoustic transfer function, H T A pre-measured acoustic transfer function may be used as , or an acoustic transfer function pre-calculated using a predetermined geometric model may be used. As the geometric model, a plane wave model assuming the propagation of a plane wave in a free sound field, a spherical wave model assuming the propagation of a spherical wave from a sound source present at a predetermined distance from the sound pickup unit 20, or the like may be used. The initial acoustic transfer function set is formed by setting an initial value H of the first acoustic transfer function for each sound source direction for each channel and frequency. T (θ1)~H T (θ N ) as an element. Initial value H T (θ1), etc. can be calculated using a geometric model assuming that sound waves arrive from a sound source direction such as θ1. N indicates the number of predetermined sound source directions. The interval between adjacent sound source directions directly affects the accuracy of the sound source direction estimated by sound source localization. The greater the number of sound source directions, the more accurate the sound source direction is expected to be, but the amount of calculation required to calculate the spatial spectrum in sound source localization increases.

[0035] Next, an example of the functional configuration of the control unit 120 according to this embodiment will be described. The control unit 120 includes a frequency analysis unit 122, a sound source direction estimation unit 124, a representative estimated sound source direction determination unit 126, an outlier removal unit 128, a representative acoustic signal determination unit 130, an acoustic transfer function estimation unit 132, and an acoustic transfer function update unit 134. Unless otherwise specified, the processes of the sound source direction estimation unit 124, the representative estimated sound source direction determination unit 126, the outlier removal unit 128, the representative acoustic signal determination unit 130, the acoustic transfer function estimation unit 132, and the acoustic transfer function update unit 134 may be executed independently for each frequency.

[0036] The frequency analysis unit 122 receives M-channel acoustic signals from the sound collection unit 20 via the input / output unit 110. The acquired M-channel acoustic signals represent a time series (waveform) of amplitudes at each sample time in the time domain. The frequency analysis unit 122 performs frequency analysis on the time-domain acoustic signals for each channel for each frame of a predetermined period (e.g., 20 ms to 100 ms) and converts the signals into transform coefficients for each frequency in the frequency domain. For each channel, a set of transform coefficients between frequencies represents a frequency spectrum. For frequency analysis, the frequency analysis unit 122 can use techniques such as a short-time Fourier transform (STFT) or a discrete Fourier transform (DFT). The frequency analysis unit 122 outputs input signal information indicating the transform coefficients obtained by the transform to the sound source direction estimation unit 124 and the acoustic transfer function estimation unit 132.

[0037] The sound source direction estimation unit 124 refers to the acoustic transfer function set stored in the storage unit 150, and calculates the spatial spectrum S for each frequency using the conversion coefficient of each channel indicated in the input signal information input from the frequency analysis unit 122. sp (θ) is calculated. sp (θ) can be considered as an index indicating the degree of possibility that a sound source exists for each target direction θ based on the position of the sound collection unit 20. The sound source direction estimation unit 124 calculates the first acoustic transfer function H EThe sound source direction estimation unit 124 can estimate the direction in which the spatial spectrum is maximized as the estimated sound source direction φ, as exemplified by equation (1). The spatial spectrum S sp A specific example of a method for calculating (θ) will be described later. The sound source direction estimation unit 124 outputs estimated sound source direction information indicating the estimated sound source direction to the representative estimated sound source direction deciding unit 126 and the outlier removal unit 128. Furthermore, the sound source direction estimation unit 124 outputs input signal information to the outlier removal unit 128 in association with the estimated sound source direction information.

[0038]

number

[0039] The representative estimated sound source direction determiner 126 receives input of estimated sound source direction information from the sound source direction estimator 124. The representative estimated sound source direction determiner 126 generates a frequency distribution of the estimated sound source direction φ indicated in the estimated sound source direction information for each preset observation period. The observation period is a period including a plurality of frames for which a representative estimated sound source direction is determined at one time. More specifically, the representative estimated sound source direction determiner 126 identifies the estimated sound source direction indicated in the estimated sound source direction information for each frame included in the observation period, and counts the number of frames for each estimated sound source direction by adding 1 to the number of frames for the identified estimated sound source direction (increment). The representative estimated sound source direction determiner 126 can acquire the number of frames (frequency) for each estimated sound source direction counted in the observation period as a frequency distribution (histogram) of the estimated sound source direction.

[0040] The representative estimated sound source direction determining unit 126 determines the estimated sound source direction (mode) having the largest number of frames in the identified section as the representative estimated sound source direction θ'. The representative estimated sound source direction determining unit 126 outputs representative estimated sound source direction information indicating the determined representative estimated sound source direction θ' to the outlier removing unit 128.

[0041] The outlier removal unit 128 receives input of estimated sound source direction information and input signal information from the sound source direction estimation unit 124, and receives input of representative estimated sound source direction information from the representative estimated sound source direction determination unit 126. The outlier removal unit 128 determines, for example, whether the estimated sound source direction is equal to the representative estimated sound source direction. When the estimated sound source direction is within a predetermined range from the representative estimated sound source direction, the outlier removal unit 128 adopts input signal information corresponding to estimated sound source direction information indicating the estimated sound source direction. When the estimated sound source direction differs from the representative estimated sound source direction, the outlier removal unit 128 removes and discards the input signal information corresponding to the estimated sound source direction information indicating the estimated sound source direction as an outlier. In this way, the outlier removal unit 128 functions as a mode filter. The outlier removal unit 128 outputs the adopted input signal information to the representative sound signal determination unit 130.

[0042] The representative acoustic signal determination unit 130 determines, for each observation period, a representative value (e.g., average value) between frames of the transform coefficients for each channel indicated in the input signal information input from the outlier removal unit 128 as a representative transform coefficient for each frequency. The representative acoustic signal determination unit 130 outputs representative input signal information indicating the determined representative transform coefficient to the acoustic transfer function estimation unit 132.

[0043] The acoustic transfer function estimation unit 132 receives representative input signal information from the representative acoustic signal determination unit 130. For each frequency, the acoustic transfer function estimation unit 132 estimates the acoustic transfer function from the sound source to the microphone corresponding to that channel as a second acoustic transfer function H' based on the representative conversion coefficient for each channel indicated in the representative input signal information. The second acoustic transfer function H' corresponds to a representative value of the acoustic transfer functions estimated during the observation period. When estimating the second acoustic transfer function H', the acoustic transfer function estimation unit 132, for example, normalizes the amplitude and phase of the representative conversion coefficient for each channel between channels.

[0044] The acoustic transfer function estimation unit 132 can calculate the second acoustic transfer function H' according to, for example, equation (2). In the example of equation (2), the representative input vector X is divided by its norm |X| to normalize the amplitude of the representative transform coefficient. For example, the square root of the sum of squares can be applied as the norm. The representative input vector X is divided by the representative transform coefficient X for each channel m at a certain frequency. m The normalized amplitude is a real value between 0 and 1. The representative transform coefficient X m The sum of the channels Σ m X m its absolute value |Σ m X m The phase of the representative transform coefficient is normalized by multiplying it by the complex conjugate of the quotient obtained by dividing by |. By normalizing the phase, the average value of the phase between channels weighted by the amplitude of the representative transform coefficient of each channel becomes 0. In this embodiment, the acoustic transfer function may be a value in which the amplitude and phase between channels are relativized, and does not necessarily have to be an absolute value. The acoustic transfer function estimation unit 132 outputs second acoustic transfer function information indicating the estimated second acoustic transfer function H' to the acoustic transfer function update unit 134.

[0045]

number

[0046] The acoustic transfer function update unit 134 receives second acoustic transfer function information from the acoustic transfer function estimation unit 132 and receives representative estimated sound source direction information from the representative estimated sound source direction determination unit 126. For each frequency, the acoustic transfer function update unit 134 specifies the second acoustic transfer function H' for each channel indicated by the input second acoustic transfer function information as the second acoustic transfer function corresponding to the representative estimated sound source direction indicated in the representative estimated sound source direction information. Using the specified second acoustic transfer function H', the acoustic transfer function update unit 134 determines the first acoustic transfer function H corresponding to the representative estimated sound source direction from the acoustic transfer function set stored in the storage unit 150. E Update.

[0047] The acoustic transfer function update unit 134 uses, for example, exponential smoothing to update the second acoustic transfer function H′ at that time and the first acoustic transfer function H E (θ') is weighted averaged to obtain the newly updated first acoustic transfer function H E In the example of equation (3), the weighting coefficient β multiplied by the second acoustic transfer function H' is a predetermined positive real number whose maximum value is 1. The first acoustic transfer function H before updating is calculated as follows: E (θ') is multiplied by a weighting factor (1-β). The weighting factors β and (1-β) are the second acoustic transfer function H' and the first acoustic transfer function H E (θ'). Therefore, the larger the weighting coefficient β, the E The time average value of the acoustic transfer function is obtained so that the second acoustic transfer function H' is weighted as (θ'). When the weighting coefficient β is 1, the second acoustic transfer function H' is weighted more heavily than the first acoustic transfer function H for each frame. E That is, the larger the weighting coefficient β, the less the influence of the presence or absence of sound from the sound source included in the second acoustic transfer function H', temporary changes in the acoustic environment, erroneous estimation of the sound source direction, etc. is reflected in the first acoustic transfer function H E The smaller the weighting coefficient β, the smoother the temporal fluctuations of the second acoustic transfer function H′.

[0048] The acoustic transfer function update unit 134 updates the original unupdated first acoustic transfer function H E (θ') is replaced by a new first acoustic transfer function H E (θ′) is stored in the storage unit 150 in association with the estimated sound source direction θ′.

[0049]

number

[0050] The acoustic transfer function updating unit 134 may determine the weighting coefficient β for the second acoustic transfer function H' so that the weighting coefficient β increases as the reliability of the representative estimated sound source direction θ' increases. The acoustic transfer function updating unit 134 may use the ratio of the number of frames in which the estimated sound source direction φ is within a predetermined range from the representative estimated sound source direction θ' to the total number of frames as the reliability of the representative estimated sound source direction θ'. More specifically, the acoustic transfer function updating unit 134 may determine the weighting coefficient β as L / 2K, where L is the number of frames in the observation period in which the estimated sound source direction φ becomes the representative estimated sound source direction θ', and K is the total number of frames in the observation period. When L=K, the weighting coefficient β has a maximum value of 0.5.

[0051] The representative estimated sound source direction determination unit 126 may identify an estimated sound source direction in which the number of counted frames is greater than a preset lower limit of the number of frames as a significant estimated sound source direction, and may determine a section consisting of a plurality of significant sound source directions adjacent to each other. The representative estimated sound source direction determination unit 126 may determine a representative value of estimated sound source directions in which the number of frames is maximum among the identified sections as the representative estimated sound source direction. In this way, a specifically isolated estimated sound source direction is excluded and is not selected as the representative estimated sound source direction.

[0052] When the estimated sound source direction is within a predetermined tolerance range (for example, ±3 to 5°) from the representative estimated sound source direction, the outlier removal unit 128 adopts input signal information corresponding to estimated sound source direction information indicating the estimated sound source direction, and when the estimated sound source direction is outside the predetermined range from the representative estimated sound source direction, removes and discards the input signal information corresponding to the estimated sound source direction information indicating the estimated sound source direction as an outlier. As a result, a conversion coefficient for an acoustic signal that gives an estimated sound source direction that is close to the representative estimated sound source direction is also used to estimate the second acoustic transfer function. A case where the tolerance range is 0° corresponds to the above-mentioned mode filter.

[0053] The number of sound sources from which sound is emitted at one time is not necessarily limited to one, but may be two or more, or the sound source may not emit sound temporarily or continuously. spOnly when one direction in which (θ) is maximized and is greater than a predetermined spatial spectrum threshold is detected, estimated sound source direction information indicating the detected direction as an estimated sound source direction θ' may be output to the representative estimated sound source direction determiner 126 and the outlier remover 128. As described above, the acoustic transfer function updater 134 updates the first acoustic transfer function H E (θ') can be updated using the second acoustic transfer function H'. In this case, when determining the weighting coefficient β, the acoustic transfer function updating unit 134 may determine the weighting coefficient β by setting the number of frames in which the number of estimated sound source directions in the observation period is one as L and the number of frames in which the one estimated sound source direction φ falls within the representative estimated sound source direction θ' (or within an allowable range) as K.

[0054] In other words, the sound source direction estimation unit 124 calculates the spatial spectrum S sp When two or more directions are detected in which (θ) is maximum and is greater than the predetermined spatial spectrum threshold, or when the spatial spectrum S sp If no direction is detected in which (θ) is maximized and exceeds a predetermined spatial spectrum threshold, the estimated sound source direction information is not output to the representative estimated sound source direction determining unit 126 and the outlier removing unit 128. Therefore, when determining the second acoustic transfer function, input signal information acquired when the number of estimated sound source directions is two or more or when no estimated sound source direction is detected is not used. On the other hand, when two or more estimated sound source directions are detected, sounds arriving from multiple sound sources are superimposed at the microphone, so the ratio of the conversion coefficients between channels does not provide the ratio of the acoustic transfer functions for the sound source direction associated with a specific sound source. If no sound source direction is detected, no significant sound arrives at the microphone from the sound source in the first place. Therefore, by limiting sound source localization, acoustic transfer function estimation, and update when only one sound source is detected, deterioration in the estimation accuracy of the acoustic transfer function can be suppressed.

[0055] In a two-dimensional space, the sound source direction may be defined by an azimuth angle from a representative point (e.g., the center of gravity) of the sound collection unit 20. In this case, the arrangement of the sound source directions associated with the individual first acoustic transfer functions constituting the acoustic transfer function set may be, for example, a one-dimensional array distributed on a circumference parallel to a horizontal plane centered on the position of the sound collection unit 20. In a three-dimensional space, the sound source direction may be defined by a pair of an azimuth angle and an elevation angle. The arrangement of the sound source directions may be a two-dimensional array distributed on a spherical surface centered on the position of the sound collection unit 20. The acoustic transfer function set may be configured to include a first acoustic transfer function for each sound source position. In this case, the arrangement of the sound source positions is a three-dimensional distribution in the three-dimensional space. The sound source position is expressed in three-dimensional coordinates based on the position of the sound collection unit 20 and corresponds to a combination of the sound source direction and the distance from the reference position. However, although this embodiment will mainly describe a case where the distribution of the sound source positions is a one-dimensional array, the present invention can also be applied to a two-dimensional or three-dimensional array.

[0056] When the acoustic transfer function set includes a first acoustic transfer function for each sound source position, the sound source direction estimation unit 124 can estimate the sound source position as information to be estimated. The sound source direction estimation unit 124 may calculate a spatial spectrum for each sound source position instead of the sound source direction, and identify the sound source position where the spatial spectrum is maximized (or at its maximum). The acoustic transfer function update unit 134 may set the identified sound source position as the estimated sound source position, and update the first acoustic transfer function for the estimated sound source position using the second acoustic transfer function estimated by the sound source direction estimation unit 124 using the above-described method.

[0057] (Example of sound source localization) Next, an example of sound source localization will be described. The sound source direction estimation unit 124 can use, for example, a beamforming method in sound source localization. The sound source direction estimation unit 124 estimates a spatial spectrum S sp The sound source direction θ that gives the maximum value of (θ) can be calculated as the estimated sound source direction. sp (θ) is the pseudo-inverse matrix H(θ) of the acoustic transfer function vector H(θ) for the input vector X. +The acoustic transfer function vector H(θ) is a vector [H1(θ),H2(θ),…,H M (θ)] T is.

[0058]

number

[0059] The sound source direction estimation unit 124 may use a method other than the beamforming method for sound source localization, such as a MUSIC (Multiple Signal Classification) method or a delay and sum method.

[0060] (Acoustic Transfer Function Adaptation Processing) Next, the acoustic transfer function adaptation process according to this embodiment will be described. Fig. 2 is a data flowchart showing an example of the acoustic transfer function adaptation process according to this embodiment. In the following description, an example will be taken where the outlier remover functions as a mode filter. (Step S102) The frequency analysis unit 112 receives an M-channel acoustic signal x from the sound collection unit 20. (Step S104) The frequency analysis unit 112 performs frequency analysis on the acoustic signal for each frame for each channel and converts it into an input vector X (input signal information) that indicates a conversion coefficient in the frequency domain. K ] form the input signal set Z.

[0061] (Step S106) The sound source direction estimation unit 124 refers to the acoustic transfer function set and calculates the spatial spectrum S for each frequency using the conversion coefficients indicated in the input vector X for each frame. sp The sound source direction for which (θ) is maximized is calculated as the estimated sound source direction φ. The set of estimated sound source directions for each frame in the observation period [φ1, φ2, ..., φ K ] form a group of orientation directions Φ. (Step S108) The representative estimated sound source direction determiner 126 counts the frequency distribution indicating the frequency (number of frames) of each estimated sound source direction for each observation period, and determines the most frequently occurring estimated sound source direction as the representative estimated sound source direction θ′.

[0062] (Step S110) The outlier removal unit 128 removes, as an outlier, an input vector X″ of a frame that gives an estimated sound source direction φ different from the representative estimated sound source direction θ′ from among the input vectors X for each frame. L '] form the outlier-removed input signal set Z'.

[0063] (Step S112) The representative acoustic signal determination unit 130 determines a representative input vector X′ representing a representative value of the transform coefficients for each channel indicated in the input vector X′ of the frame in which the outliers are not removed and which remains in the observation section, as a representative transform coefficient. <x>Generate. (Step S114) The acoustic transfer function estimation unit 132 normalizes the representative conversion coefficient for each channel indicated in the representative input signal information for each frequency between channels, and estimates the acoustic transfer function from the sound source to the microphone corresponding to that channel as a second acoustic transfer function H'. (Step S116) The acoustic transfer function update unit 134 updates the first acoustic transfer function H E and the weighted average value (1-β)H of the second acoustic transfer function H' E +βH' is the new first acoustic transfer function H E Update as.

[0064] The process of FIG. 2 may be repeated for each observation period, or may be repeated at a cycle shorter than the observation period, for example, for each frame.

[0065] (Evaluation experiment) Next, an evaluation experiment performed to evaluate the effectiveness of the above-described embodiment will be described. The evaluation experiment was performed in a laboratory with an acoustic environment similar to that of a typical office. The laboratory was approximately rectangular in shape. The dimensions of the laboratory were 7.0 m in width (x direction), 4.0 m in length (y direction), and 3.0 m in height (z direction) (see FIG. 5). Tables were installed in the center and on the periphery of the laboratory, with multiple chairs placed around the tables. A circular microphone array (see FIG. 3) was installed as the sound collection unit 20 on the central table, and the height of the table from the floor was 0.9 m. A notebook personal computer and other items were placed on the peripheral tables.

[0066] Prior to the evaluation experiment, a sound source signal was acquired. A male voice selected from the Corpus of Spontaneous Japanese (CSJ) was used as the sound source. The loudspeaker was moved slowly clockwise along the circumference of a circle at a distance of 0.78 m from the center of the microphone array, and sound was emitted in accordance with the sound source signal. The height of the center of the loudspeaker from the floor was set to 1.0 m. Under these conditions, eight-channel acoustic signals representing the sound coming from the sound source and picked up by the microphone array were acquired for 20 minutes. The sampling frequency was set to 16 kHz. During this time, the loudspeaker made three revolutions around the circumference of the circle. In addition, the initial value of the first acoustic transfer function H T Assuming a free sound field model, an acoustic transfer function was calculated based on the positional relationship between the loudspeaker installed in the target direction and the microphone array and was set in advance.

[0067] The frequency analysis unit 122 performed STFT in the frequency analysis. In the STFT, the frame length and shift width were set to 512 points and 256 points, respectively. The Hann window (Hanning window) was used as the window function. Frames with an average sound pressure of -24 dB or more were considered valid frames. The acoustic signals in the valid frames were adopted, and other frames were discarded as silent intervals. Each time a time equivalent to the observation period elapsed, a new observation period was set. In other words, the shift width of the observation period was set to a period equivalent to the observation period.

[0068] The evaluation experiment consisted of two verification items. In the first verification, the relationship between the length of the mode filter (corresponding to the observation period) and sound source localization performance was investigated. In the first verification, the first acoustic transfer function was updated by performing the process illustrated in Figure 2 using acoustic signals acquired in advance for each of several observation periods. The observation period was set to 11 different intervals of 60 frames, ranging from 60 frames (0.96 seconds) to 600 frames (9.6 seconds).

[0069] The directional resolution of the sound source direction for the first acoustic transfer function was set to 5°. The directional resolution corresponds to the interval between the sound source directions associated with each first acoustic transfer function. The maximum and minimum frequencies of the frequency band to be processed were set to 300 Hz and 6000 Hz, respectively. Sound source localization was then performed using an acoustic transfer function set including the first acoustic transfer function obtained by updating, and the estimated sound source direction was compared with the target sound source direction. The target sound source direction corresponds to the known sound source direction that is the correct answer. The delay and sum method was used as the sound source localization method.

[0070] The success rate was calculated as an evaluation index for sound source localization performance. The success rate corresponds to the ratio of the number of successful frames to the number of valid frames. A successful frame corresponds to a frame in which sound source localization was successful. In this verification, frames in which the estimated sound source direction was within a predetermined range (for example, 5°) from the target sound source direction were counted as successful frames. Therefore, a higher success rate means better sound source localization performance. In this verification, the number of valid frames was 4,322.

[0071] Figure 6 illustrates the success rate for each observation period. Figure 6 shows that the highest success rate, 90.42%, was achieved when the observation period was 120 frames (equivalent to 1.92 seconds). Considering that the speakers were moved during this test, the performance of sound source localization could be maintained even if the speakers were stationary for approximately 1.92 seconds. Overall, the shorter the observation period, the higher the success rate. This is likely due to the fact that the longer the observation period, the higher the frequency of outliers due to changes in the target sound source direction caused by movement, and the lower the update frequency of the first acoustic transfer function. On the other hand, a shorter observation period reduces the statistical reliability of the distribution of estimated sound source directions used to determine the representative estimated sound source direction. This actually contributes to a lower success rate. The fact that the success rate reaches its highest when the observation period is 120 frames can be seen as an indication of the increase and decrease that occur depending on the observation period.

[0072] In the first verification, the relationship between reliability and the type of first acoustic transfer function was investigated. A success rate was determined for each of the following three types of first acoustic transfer function: (1) this embodiment: a weighted sum of the first acoustic transfer function and the second acoustic transfer function is updated to a new first acoustic transfer function; (2) existing method (method described in Non-Patent Document 1): the first acoustic transfer function is updated for each frame; and (3) no update: the first acoustic transfer function is not updated. However, the observation period was set to 120 frames.

[0073] FIG. 7 is a diagram illustrating the success rate for each type of first acoustic transfer function. FIG. 7 shows the success rate in the order of no update, existing method, and this embodiment. The success rate increases in the order of existing method, no update, and this embodiment. The success rate was 85.59% without updating, 80.63% with the existing method, and 90.42% with this embodiment. This result supports the effectiveness of this embodiment. One reason why the success rate with the existing method is lower than the success rate without updating is presumably because the first acoustic transfer function is updated uniformly for each frame, which causes outliers and results in cases where an estimated sound source direction with low reliability is not rejected.

[0074] (Variation) Next, a modification of this embodiment will be described. In the following description, differences from the above-described embodiment will be mainly described, and unless otherwise specified, the same reference numerals as those in the above-described embodiment will be used to refer to the description thereof. FIG. 8 is a schematic block diagram showing an example configuration of a sound processing system S1 according to a first modification of this embodiment. The sound processing system S1 according to this modification is applied to sound source separation. In the sound processing system S1, the control unit 120 of the sound processing device 10 includes a frequency analysis unit 122, a sound source direction estimation unit 124, a representative estimated sound source direction determination unit 126, an outlier removal unit 128, a representative sound signal determination unit 130, an acoustic transfer function estimation unit 132, and an acoustic transfer function update unit 134, as well as a sound source separation unit 136 and a sound source signal generation unit 138.

[0075] The frequency analysis unit 122 outputs the input signal information to the sound source direction estimation unit 124, the acoustic transfer function estimation unit 132, and also to the sound source separation unit 136. The sound source direction estimation unit 124 outputs the estimated sound source direction information to the representative estimated sound source direction determination unit 126 and the outlier removal unit 128 as well as to the sound source separation unit 136 . The acoustic transfer function estimation unit 132 outputs the second acoustic transfer function information to the acoustic transfer function update unit 134 as well as to the sound source separation unit 136 .

[0076] The sound source separation unit 136 receives input signal information from the frequency analysis unit 122 and inputs estimated sound source direction information from the sound source direction estimation unit 124. The sound source separation unit 136 extracts sound source components arriving from the estimated sound source direction from the conversion coefficients for each channel indicated in the input signal information. For example, the sound source separation unit 136 refers to the acoustic transfer function set stored in the storage unit 150, and extracts a first acoustic transfer function H E From the separation matrix W(H E , θ′) to the input vector X. As shown in equation (5), the sound source separation unit 136 calculates the separation matrix W(H E , θ') to calculate an output vector Y (separated sound source) indicating, for each frequency, the output value estimated as the sound source component arriving from the sound source present in the estimated sound source direction θ'. The input vector X is a vector including, as elements, the conversion coefficients for each channel indicated in the input signal information. When multiple estimated sound source directions are detected, the sound source separation unit 136 can determine an output value for each sound source (estimated sound source direction). The sound source separation unit 136 outputs output signal information indicating the output value determined for each frequency for each sound source to the sound source signal generation unit 138.

[0077]

number

[0078] The sound source signal generation unit 138 converts the output value for each frequency indicated in the output signal information input from the sound source separation unit 136 for each sound source into a time series of amplitude for each sample time in the time domain. When converting the output value for each frequency in the frequency domain into a time series of amplitude, the sound source signal generation unit 138 can use the inverse process of frequency analysis, for example, an inverse discrete Fourier transform. The sound source signal generation unit 138 can generate a sound source signal by concatenating the time series of amplitude obtained for each frame for each sound source between frames. The sound source signal generation unit 138 may output the generated sound source signal to an output destination device via the input / output unit 110, or may store it in the storage unit 150.

[0079] The sound source direction estimation unit 124 calculates the spatial spectrum S sp (θ) becomes a local maximum and may detect multiple directions in which it exceeds a predetermined spatial spectrum threshold. In such a case, the sound source direction estimation unit 124 may output estimated sound source direction information indicating each of the multiple sound source directions as an estimated sound source direction to the sound source separation unit 136. This is because in such a case, it is estimated that there are multiple significant sound sources. The sound source localization in the sound source direction estimation unit 124 and the sound source separation in the sound source separation unit 136 may be performed simultaneously with the update of the first acoustic voltage function in the acoustic transfer function update unit 134, but they do not necessarily have to be synchronized. In other words, when two or more sound sources are detected, the estimated sound source direction information is not output to the representative estimated sound source direction determiner 126 and the outlier remover 128, and even if the representative estimated sound source direction information is not input from the representative estimated sound source direction determiner 126 to the acoustic transfer function update unit 134, the sound source separation in the sound source separation unit 136 is permitted.

[0080] The sound source separation unit 136 can apply, for example, the above-mentioned beamforming method as a sound source separation method. In this case, the sound source separation unit 136 calculates the pseudo-inverse matrix H(θ′) of the acoustic transfer function vector H(θ′) relating to the estimated sound source direction θ′ estimated using the beamforming method. + (θ') can be used as the separation matrix. As another source separation method, for example, the GHDSS (Geometric-contrained High-order Decorrelation-based Source Separation) method can be used. The GHDSS method includes a process of adaptively calculating the separation matrix W so as to minimize the cost function J(W). The cost function J(W) is the separation sharpness J SS (W) and Geometric Constraint J GC (W) and the weighted sum of the separation sharpness J SS (W) is an index value that indicates the degree to which the sound source component Y of a certain sound source is mixed with the component of another sound source. GC (W) is an index value that indicates the degree of error between the output sound source signal and the original sound source signal emitted from the sound source.

[0081] Next, a second modified example of this embodiment will be described. In the following description, differences from the above-described embodiment and modified example will be mainly described, and unless otherwise specified, the same reference numerals as those in the above-described embodiment will be used to refer to the description thereof. The sound processing system S1 according to this modified example forms part of a robot system (not shown). FIG. 9 is a schematic block diagram showing an example configuration of the sound processing system S1 according to this modified example. One or both of the sound processing device 10 and the sound collection unit 20 constituting the sound processing system S1 may be built into the housing of the robot.

[0082] In the sound processing device 10, the control unit 120 includes a frequency analysis unit 122, a sound source direction estimation unit 124, a representative estimated sound source direction determination unit 126, an outlier removal unit 128, a representative sound signal determination unit 130, an acoustic transfer function estimation unit 132, an acoustic transfer function update unit 134, a sound source separation unit 136, and a sound source signal generation unit 138, as well as an operation control unit 140. That is, in the sound processing system S1, the sound source direction estimation unit 124, the sound source separation unit 136, and the sound source signal generation unit 138 may function as a robot audition functional block that realizes robot audition.

[0083] The control unit 120 may further include a speech recognition processing unit (not shown). The speech recognition processing unit may identify the type of sound source by performing known speech recognition processing on sound source components related to each sound source (sound source identification). The speech recognition processing unit may identify a speaker who is a person as the type of sound source. The sound source direction estimation unit 124 may notify another device of estimated sound source direction information indicating the estimated sound source direction for the identified type of sound source, or may output a sound source signal converted from output signal information for the identified type of sound source to another device. The sound source direction estimation unit 124 is capable of estimating the sound source position as described above, and outputs estimated sound source direction information indicating the estimated sound source position to the representative estimated sound source direction determination unit 126, the outlier removal unit 128, the sound source separation unit 136, and also to the operation control unit 140.

[0084] The movement control unit 140 receives estimated sound source direction information from the sound source direction estimation unit 124 and receives output signal information indicating sound source components from the sound source separation unit 136. The movement control unit 140 controls the movement of the movement mechanism 40 using one or both of the estimated sound source position and the sound source components. The movement control unit 140 may, for example, perform self-localization and environmental map creation based on the estimated sound source position and the sound source components (SLAM: Simultaneous Localization and Mapping). The movement control unit 140 can estimate the presence of an object (including a person) that is the sound source at the estimated sound source position by performing sound source identification. The movement control unit 140 may determine the existence probability of the object that is the sound source using a predetermined density function model so that the probability increases as the object approaches the estimated sound source position. The movement control unit 140 can, for example, create an environmental map by superimposing the spatial distribution of the existence probability of each object between objects. In path planning, the movement control unit 140 may determine a travel path so as not to pass through an area where the existence probability of an object is higher than a predetermined existence probability. The travel path is represented by a target position at each time. The movement control unit 140 may determine the estimated direction of a predetermined type of sound source as the target direction relative to the front of the robot. The movement control unit 140 outputs a control signal to the movement mechanism 40 indicating either or both of the target position and the target direction at that time.

[0085] The movement mechanism 40 is built into the robot's housing and controls the movement of the robot based on control signals input from the movement control unit 140. The movement mechanism 40 is equipped with a motor (not shown) that serves as a power source and an encoder (not shown) that detects the position and direction of the movement mechanism 40. The motor moves the robot so that it approaches a target position or direction specified by the control signal. The encoder sequentially outputs movement information indicating the position and direction detected at that time as the movement state to the movement control unit 140.

[0086] As described above, the sound processing device 10 according to this embodiment includes: a storage unit 150 that stores a first acoustic transfer function for each sound source direction; a sound source direction estimation unit 124 that calculates a spatial spectrum for each sound source direction for each frame based on the transform coefficients in the frequency domain of the acoustic signal for each channel and the first acoustic transfer function, and estimates the sound source direction for which the spatial spectrum is maximum as the estimated sound source direction; a representative estimated sound source direction determination unit 126 that determines a representative estimated sound source direction that is a representative value of the estimated sound source directions based on the frequency distribution of the estimated sound source directions in an observation period consisting of a plurality of frames; an outlier removal unit 128 that removes transformation coefficients of frames whose estimated sound source directions fall outside a predetermined tolerance range from the representative estimated sound source direction; an acoustic transfer function estimation unit 132 that estimates, as the second acoustic transfer function, a representative value of the acoustic transfer function from the sound source to the sound collection unit for the acoustic signal in the observation period based on the transformation coefficients of the acoustic signal in the remaining frames; and an acoustic transfer function update unit 134 that updates the first acoustic transfer function for the representative estimated sound source direction using the second acoustic transfer function. With this configuration, the second acoustic transfer function is calculated based on the conversion coefficient of the acoustic signal that gives an estimated sound source direction within a predetermined range from the representative estimated sound source direction, and the calculated second acoustic transfer function can be used to update the first acoustic transfer function in association with the representative estimated sound source direction. A representative value of the acoustic transfer function obtained based on the acoustic signal that statistically gives the representative estimated sound source direction or an estimated sound source direction that is close to the representative estimated sound source direction is updated as the second acoustic transfer function in association with the representative estimated sound source direction, thereby obtaining a first acoustic transfer function with a stable correspondence with the sound source direction. By using such a first acoustic transfer function, it is possible to improve the reliability of sound source localization, sound source separation, and other microphone array processing using any acoustic signal online.

[0087] Furthermore, the sound source direction estimating unit 124 may determine the estimated sound source direction φ having a local maximum (for example, the largest) frequency as the representative estimated sound source direction θ′. According to this configuration, the estimated sound source direction with the maximum frequency within the observation period is determined as the representative estimated sound source direction, so that the most likely estimated sound source direction is simply determined as the representative estimated sound source direction.

[0088] The acoustic transfer function update unit 134 also updates the first acoustic transfer function and the second acoustic transfer function by a weighted average value (for example, βH'+(1-β)H E ) into a new first acoustic transfer function H E may be updated to. With this configuration, when the observation period is changed, the first acoustic transfer function is not completely replaced by the second acoustic transfer function through updating, but some of its components remain. This avoids sudden fluctuations in the first acoustic transfer function, thereby ensuring system stability.

[0089] Furthermore, the acoustic transfer function update unit 134 may determine the reliability (e.g., L / K, where K is the number of frames in the observation period) of the representative estimated sound source direction based on the frequency (e.g., the number of frames L) that the estimated sound source direction falls within an acceptable range during the observation period, and may increase the ratio β of the second acoustic transfer function to the first acoustic transfer function as the reliability increases (e.g., L / 2K). With this configuration, the first acoustic transfer function can be updated using the second acoustic transfer function, with emphasis being placed on acoustic signals that provide more reliable estimated sound source directions, thereby improving the reliability of the updated first acoustic transfer function.

[0090] Furthermore, the allowable range of the estimated sound source direction is equal to the representative estimated sound source direction and does not need to include directions different from the representative estimated sound source direction. According to this configuration, it is possible to simply determine whether or not to exclude the transform coefficient of the acoustic signal that gives the estimated sound source direction, depending on whether or not the estimated sound source direction is equal to the representative estimated sound source direction.

[0091] The sound source direction estimation unit 124 also calculates a pseudo-inverse matrix H of an acoustic transfer function vector including a first acoustic transfer function for each channel, for the input vector X including a conversion coefficient for each channel. + The spatial spectrum may be calculated by multiplying the This configuration allows the sound source direction to be estimated based on the spatial spectrum calculated by a simple matrix operation, which does not require many computational resources and can be implemented economically.

[0092] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0093] S1...sound processing system, 10...sound processing device, 20...sound collection unit, 40...operation mechanism, 110...input / output unit, 120...control unit, 122...frequency analysis unit, 124...sound source direction estimation unit, 126...representative estimated sound source direction determination unit, 128...outlier removal unit, 130...representative sound signal determination unit, 132...acoustic transfer function estimation unit, 134...acoustic transfer function update unit, 136...sound source separation unit, 138...sound source signal generation unit, 140...operation control unit, 150...storage unit< / x>

Claims

1. a storage unit that stores the first acoustic transfer function for each sound source direction; For each frame, a spatial spectrum is calculated for each sound source direction based on a transform coefficient in the frequency domain of the acoustic signal for each channel and the first acoustic transfer function; a sound source direction estimation unit that estimates the sound source direction in which the spatial spectrum is maximized as an estimated sound source direction; a representative estimated sound source direction determining unit that determines a representative estimated sound source direction that is a representative value of the estimated sound source directions based on a frequency distribution of the estimated sound source directions in an observation period consisting of a plurality of frames; an outlier removal unit that removes transform coefficients of frames whose estimated sound source directions are outside a predetermined tolerance range from the representative estimated sound source direction; an acoustic transfer function estimation unit that estimates, as a second acoustic transfer function, a representative value of an acoustic transfer function from a sound source to a sound collection unit of the acoustic signal during the observation period based on a transformation coefficient of the acoustic signal of the remaining frames; an acoustic transfer function update unit that updates a first acoustic transfer function for the representative estimated sound source direction using the second acoustic transfer function; An acoustic processing device comprising:

2. The representative estimated sound source direction determining unit determines the estimated sound source direction with the maximum frequency as the representative estimated sound source direction. The sound processing device according to claim 1 .

3. The acoustic transfer function update unit updates the first acoustic transfer function to a new first acoustic transfer function by calculating a weighted average value of the first acoustic transfer function and the second acoustic transfer function. The sound processing device according to claim 1 .

4. the acoustic transfer function update unit determines a reliability of the representative estimated sound source direction based on a frequency with which the estimated sound source direction falls within the tolerance range during the observation period; The higher the reliability, the higher the ratio of the second acoustic transfer function to the first acoustic transfer function. The sound processing device according to claim 3 .

5. The tolerance range is equal to the representative estimated sound source direction and does not include directions different from the representative estimated sound source direction. The sound processing device according to claim 1 .

6. The sound source direction estimation unit The sound processing device according to claim 1 , wherein the spatial spectrum is calculated by multiplying an input vector including the conversion coefficients for each channel by a pseudo-inverse matrix of an acoustic transfer function vector including the first acoustic transfer function for each channel.

7. A program for causing a computer to function as the sound processing device according to claim 1.

8. A sound processing method in a sound processing device including a storage unit that stores a first acoustic transfer function for each sound source direction, The sound processing device comprises: For each frame, a spatial spectrum is calculated for each sound source direction based on a transform coefficient in the frequency domain of the acoustic signal for each channel and the first acoustic transfer function; a sound source direction estimating step of estimating the sound source direction in which the spatial spectrum is maximized as an estimated sound source direction; a representative sound source direction determining step of determining a representative estimated sound source direction that is a representative value of the estimated sound source directions based on a frequency distribution of the estimated sound source directions in an observation period consisting of a plurality of frames; an outlier removal step of removing transform coefficients of frames in which the estimated sound source direction is outside a predetermined tolerance range from the representative estimated sound source direction; an acoustic transfer function estimating step of estimating, as a second acoustic transfer function, a representative value of an acoustic transfer function from a sound source to a sound collection unit of the acoustic signal during the observation period based on a transformation coefficient of the acoustic signal of the remaining frames; an acoustic transfer function updating step of updating a first acoustic transfer function for the representative estimated sound source direction using the second acoustic transfer function; An acoustic processing method that performs the above.

Citation Information

Patent Citations

  • Electronic apparatus, sensitivity difference correction method, and program

    JP2015122591A

  • Voice processor and voice processing method

    JP2017067948A

  • Transfer function generation device, transfer function generation method, and program

    JP2020036271A

  • Electronic device, arrival angle estimation system, and signal processing method

    JP2020071123A

  • Acoustic Transfer Function Personalization Using Sound Scene Analysis and Beamforming

    JP2022521886A