Sound source localization method, sound source localization device, and program

The sound source localization method and device enhance accuracy by correcting parameters and suppressing non-target sound sources, enabling precise positioning and identification of specific sound sources in environments with multiple sound sources.

JP2026057125APending Publication Date: 2026-04-02PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing sound source localization technologies struggle to accurately locate and identify a specific sound source, such as a target truck, in environments with multiple sound sources, leading to inaccuracies in positioning and identification.

Method used

A sound source localization method and device that includes sound source localization, tracking, and identification processes, with parameter correction based on acoustic signals, using an array microphone and signal processing units to enhance accuracy.

Benefits of technology

Improves the accuracy of sound source localization and identification by correcting parameters and suppressing non-target sound sources, ensuring precise positioning and type recognition of the target sound source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026057125000001_ABST
    Figure 2026057125000001_ABST
Patent Text Reader

Abstract

This improves the accuracy of sound source localization processing, targeting the desired sound source within the sound pickup space. [Solution] The sound source localization device includes: performing a sound source localization process to estimate the position of a sound source based on the acoustic signals of the sound pickup space picked up by a sound pickup device placed in the sound pickup space where multiple sound sources may exist; performing a sound source tracking process to track the position of the sound source within the sound pickup space based on the results of the sound source localization process; performing a sound source identification process to identify the type of sound source based on the results of the sound source tracking process; and performing the sound source localization process by correcting parameters based on the acoustic signals used for the sound source localization process.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a sound source localization method, a sound source localization device, and a program. [Background technology]

[0002] Patent Document 1 discloses a sound source identification device that improves the detection accuracy of a specific sound source. This sound source identification device comprises a sound pickup unit composed of multiple microphones, a sound source localization unit that localizes the sound source based on the acoustic signals picked up by the sound pickup unit, a sound source separation unit that separates the sound source based on the localized signal, and a sound source identification unit that identifies the type of sound source based on the separated result. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2017-44916 [Overview of the project] [Problems that the invention aims to solve]

[0004] Here, we consider a scenario where, after identifying the type of sound source, the identified sound source is not the type of sound source we were looking for. For example, suppose in a yard with a large area, where multiple types of vehicles (e.g., trucks) exist, we want to track a target truck (hereinafter referred to as "specific truck"), but we are unable to locate (identify) the position of that specific truck, which is the sound source. Patent Document 1 does not disclose the technical details for cases where the identified sound source is not the type of sound source we were looking for. Therefore, under the above scenario, it is difficult to locate the position of the specific truck, which is the sound source, and there is room for improvement in terms of improving the accuracy of sound source localization processing.

[0005] This disclosure aims to provide a sound source localization method, a sound source localization device, and a program that have been devised in view of the conventional circumstances described above and that improve the processing accuracy of sound source localization targeting a target sound source present in a sound pickup space. [Means for solving the problem]

[0006] This disclosure provides a sound source localization method comprising: performing a sound source localization process to estimate the position of a sound source based on acoustic signals of a sound-collecting space picked up by a sound-collecting device placed in a sound-collecting space in which multiple sound sources may exist; performing a sound source tracking process to track the position of the sound source within the sound-collecting space based on the result of the sound source localization process; performing a sound source identification process to identify the type of sound source based on the result of the sound source tracking process; and performing the sound source localization process by correcting parameters based on the acoustic signals used in the sound source localization process.

[0007] Furthermore, this disclosure provides a sound source localization device comprising: a sound source localization processing unit that performs sound source localization processing to estimate the position of a sound source based on acoustic signals of the sound pickup space picked up by a sound pickup device placed in the sound pickup space in which multiple sound sources may exist; a sound source tracking processing unit that performs sound source tracking processing to track the position of the sound source in the sound pickup space based on the result of the sound source localization processing; and a sound source identification processing unit that performs sound source identification processing to identify the type of sound source based on the result of the sound source tracking processing, wherein the sound source localization processing unit corrects parameters based on the acoustic signals used for the sound source localization processing and performs the sound source localization processing.

[0008] Furthermore, this disclosure provides a program for a computer-based sound source localization device that performs the following: a process of estimating the position of a sound source based on acoustic signals of a sound-collecting space picked up by a sound-collecting device placed in a sound-collecting space in which multiple sound sources may exist; a process of tracking the position of the sound source within the sound-collecting space based on the result of the sound source localization process; a process of identifying the type of sound source based on the result of the sound source tracking process; and a process of correcting parameters based on the acoustic signals used in the sound source localization process and performing the sound source localization process.

[0009] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or recording media, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]

[0010] According to this disclosure, it is possible to improve the accuracy of sound source localization processing targeting a target sound source present in the sound pickup space. [Brief explanation of the drawing]

[0011] [Figure 1] Block diagram showing system configuration examples of the sound source localization system according to each embodiment. [Figure 2] Block diagram showing an example configuration of a sound source localization device according to Embodiment 1. [Figure 3] This figure shows an example of peak detection and correction processing for spatial spectra. [Figure 4] A flowchart showing an example of the operation procedure of the sound source localization device according to Embodiment 1, in chronological order. [Figure 5] Block diagram showing an example configuration of a sound source localization device according to Embodiment 2. [Figure 6] This figure shows an example of peak detection and rejection processing for spatial spectra. [Figure 7]A flowchart showing an example of the operation procedure of the sound source localization device according to Embodiment 2, in chronological order. [Figure 8] Block diagram showing an example configuration of a sound source localization device according to Embodiment 3. [Figure 9] This diagram shows various parameters related to the likelihood correction process and their explanations. [Figure 10] A flowchart showing an example of the operation procedure of the sound source localization device according to Embodiment 3, in chronological order. [Figure 11] Block diagram showing an example configuration of a sound source localization device according to Embodiment 4. [Figure 12] A flowchart showing an example of the operation procedure of the sound source localization device according to Embodiment 4, in chronological order. [Figure 13] Block diagram showing an example configuration of a sound source localization device according to Embodiment 5. [Figure 14] A flowchart showing an example of the operation procedure of the sound source localization device according to Embodiment 5, in chronological order. [Modes for carrying out the invention]

[0012] Hereinafter, with appropriate reference to the drawings, embodiments specifically disclosing the sound source localization method, sound source localization device, and program according to this disclosure will be described in detail. However, unnecessarily detailed explanations may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The accompanying drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter of the claims.

[0013] In the following embodiments, a yard with a large area and where multiple sound sources may exist within that area is used as an example of an acoustic signal collection space, and an example of a sound source localization system that estimates the location of a sound source based on the acoustic signal collected by a sound collection device placed within that space is described. It goes without saying that the sound collection space is not limited to a yard, although a yard is used as an example of a sound collection space.

[0014] (Embodiment 1) 1. Configuration of the sound source localization system First, an example of the system configuration of the sound source localization system 100 will be described with reference to Figure 1. Figure 1 is a block diagram showing an example of the system configuration of the sound source localization system 100 according to each embodiment. The sound source localization system 100 includes a sound pickup device AM0, an audio interface AI1, sound source localization devices 10, 10a, 10b, 10c, and 10d, and an output device 50. Hereinafter, when there is no need to distinguish between the sound source localization devices 10, 10a, 10b, 10c, and 10d, they will be collectively referred to as sound source localization devices 10 to 10d. The sound pickup device AM0 and the audio interface AI1, the audio interface AI1 and the sound source localization devices 10 to 10d, and the sound source localization devices 10 to 10d and the output device 50 are connected at least via data signal lines. This connection via data signals also enables wireless communication of data signals between them.

[0015] The sound pickup device AM0 is placed within the sound pickup space SPA1, such as a yard. The sound pickup device AM0 is an array microphone with multiple microphones, and is, for example, an ambisonic microphone capable of picking up sound from all directions within the sound pickup space SPA1. The sound pickup device AM0 picks up sound within the sound pickup space SPA1, generates an acoustic signal, and outputs it to the audio interface AI1. Note that the sound pickup device AM0 is not limited to an ambisonic microphone; for example, it may be an array microphone with multiple microphones arranged regularly in a circular pattern. In Figure 1, only one sound pickup device AM0 is shown, but multiple sound pickup devices AM0 may be placed within the sound pickup space SPA1, in which case the acoustic signals picked up by each sound pickup device AM0 are input to the audio interface AI1.

[0016] The audio interface AI1 includes, for example, an AD converter (ADC), which performs AD conversion on the acoustic signal from the sound pickup device AM0, amplifies it, and outputs it to the sound source localization devices 10-10d. The configuration of the audio interface AI1 may also be included in the sound pickup device AM0 or the sound source localization devices 10-10d. The sound pickup device AM0 and the audio interface AI1 constitute the sound pickup unit AM1 (Figures 2, 5, 8, 11, and 13).

[0017] The sound source localization devices 10 to 10d acquire acoustic signals from the sound collection unit AM1 and perform sound source localization processing to estimate the positions of multiple sound sources that may exist within the sound collection space SPA1 by performing various processing using these acoustic signals. Furthermore, as will be described in detail later, in addition to sound source localization processing, the sound source localization devices 10 to 10d can also perform sound source tracking processing based on the results of sound source localization processing, beamforming processing based on the results of sound source tracking processing, and sound source identification processing based on the results of beamforming processing. Note that beamforming processing is a process for extracting sound sources in the estimated direction, and instead of beamforming processing, generally known signal processing techniques such as sound source separation techniques may be used to extract desired sound sources.

[0018] The output device 50 is, for example, a speaker and is capable of outputting acoustic signals from the sound source localization devices 10 to 10d. Alternatively, the output device 50 may be a display such as a Liquid Crystal Display, in which case it can display a screen (not shown) that notifies the processing results of the sound source localization devices 10 to 10d.

[0019] 2. Configuration of the sound source localization device Next, the configuration of the sound source localization device according to Embodiment 1 will be described with reference to Figures 2 and 3. Figure 2 is a block diagram showing an example of the configuration of the sound source localization device 10 according to Embodiment 1. Figure 3 is a diagram showing an example of the peak detection processing and correction processing of the spatial spectrum. The sound source localization device 10 comprises a time-frequency conversion unit 11, a spatial frequency conversion unit 12, a sound source localization processing unit 13, a sound source tracking processing unit 14, a beamforming processing unit 15, and a sound source identification processing unit 16. In the following description, bold text in mathematical formulas indicates vector components, and in the following description, "is a vector component" may be written before the corresponding parameter to emphasize that it is a vector component.

[0020] Each of the time-frequency conversion unit 11, spatial frequency conversion unit 12, sound source localization processing unit 13, sound source tracking processing unit 14, beamforming processing unit 15, and sound source identification processing unit 16 is composed of a semiconductor chip or dedicated hardware on which at least one electronic device such as a Central Processing Unit (CPU), Digital Signal Processor (DSP), Graphical Processing Unit (GPU), or Field Programmable Gate Array (FPGA) and memory such as Random Access Memory (RAM) and Read Only Memory (ROM) are implemented.

[0021] The time-frequency conversion unit 11 receives the acoustic signal from the sound pickup unit AM1 and converts the time-series signal of this acoustic signal into a time-frequency domain signal by applying predetermined signal processing (e.g., discrete Fourier transform). For example, if the sound pickup device AM0 is a first-order ambisonic microphone equipped with four directional microphones arranged in four different directions, an acoustic signal is obtained in which sounds from the four known directions are picked up. The time-frequency conversion unit 11 converts this acoustic signal into a time-frequency domain signal. The time-frequency conversion unit 11 outputs the converted time-frequency domain signal to the spatial frequency conversion unit 12. Note that the time-frequency conversion unit 11 may be omitted when processing the input signal from the sound pickup unit AM1 without dividing it into each frequency band, that is, when performing sound source localization processing on a full-band signal including all frequency bands.

[0022] The spatial frequency conversion unit 12 receives a time-frequency domain signal from the time-frequency conversion unit 11 and converts the time-frequency domain signals from the multiple microphone elements of the sound-collecting device AM0 into spatial frequency domain signals by applying predetermined signal processing (e.g., spatial Fourier transform, circular harmonic transform, spherical harmonic function transform, etc.). This enables the sound source localization processing unit 13 to perform sound source localization processing using this spatial frequency domain signal. The spatial frequency conversion unit 12 outputs the converted spatial frequency domain signal to the sound source localization processing unit 13. Note that the order in which the time-frequency conversion unit 11 and the spatial frequency conversion unit 12 are executed may be reversed, and the time-frequency conversion may be applied to the spatial frequency-converted signal. Furthermore, if sound source localization processing is performed without applying a spatial Fourier transform to the input signal from the sound-collecting unit AM1 or the input signal from the time-frequency conversion unit 11, the spatial frequency conversion unit 12 may be omitted.

[0023] The sound source localization processing unit 13 receives a spatial frequency domain signal from the spatial frequency conversion unit 12 and performs sound source localization processing using this spatial frequency domain signal to estimate the position of each of the multiple sound sources that may exist within the sound pickup space SPA1. The sound source localization processing unit 13 outputs the direction estimate value θ (azimuth angle θ, elevation angle φ), which is a vector component obtained as a result of the sound source localization processing, to the sound source tracking processing unit 14. The sound source localization processing unit 13 may also output the corrected spatial spectrum P'(θ), which is a vector component, to the sound source tracking processing unit 14 instead of the direction estimate value θ (azimuth angle θ, elevation angle φ), which is a vector component. Depending on the microphone shape of the sound pickup unit AM1 (for example, a linear microphone array), only the azimuth angle θ component of the direction estimate value θ, which is a vector component, may be obtained, and the elevation angle φ component may not be obtained, resulting in one-dimensional data, but the following explanation is still valid in this case. When the sound source localization processing unit 13 receives the sound identification probability value p(θ^) and the sound source direction estimate value θ^, which are vector components obtained as a result of the sound source identification processing from the sound source identification processing unit 16, it corrects the parameters (e.g., spatial spectrum) used for sound source localization processing based on the results. The sound source localization processing unit 13 includes a spatial spectrum calculation unit 131, a spatial spectrum correction unit 132, a peak detection unit 133, a correction amount calculation unit 134, and a correction filter calculation unit 135.

[0024] Here, we will describe a detailed configuration example of the sound source localization processing unit 13.

[0025] The spatial spectrum calculation unit 131 takes the spatial frequency domain signal from the spatial frequency conversion unit 12 as input and calculates the search range information {θ}, which is a pre-set vector component. i The spatial spectrum P(θ), which is a vector component, corresponding to} is calculated. A generally known sound source localization algorithm can be used as the method for calculating the spatial spectrum. For example, in equation (1), the spatial spectrum calculation unit 131 calculates the spatial spectrum according to the MUSIC method. Here, in equation (1), a(θ) is a steering vector that represents the transfer characteristics of a plane wave arriving from direction θ (for example, the estimated sound source direction value θ^ which is a vector component), and the superscript H represents the Hermitian transpose. Also, the denominator em This is the m-th eigenvector obtained by performing eigenvalue decomposition on a matrix obtained by averaging and smoothing the spatial correlation matrix calculated from the input signal, which is a vector component, in the time direction, where M is the number of eigenvalues ​​and L is the last eigenvalue number for the signal subspace. The spatial spectrum calculation unit 131 generates a spatial spectrum map SPS1 (see Figure 3) with the azimuth angle θ and elevation angle φ as two-dimensional axes, or matrix data corresponding to the spatial spectrum map SPS1, as a result of calculating the spatial spectrum P(θ), which is a vector component. Hereinafter, the spatial spectrum map and the matrix data corresponding to the spatial spectrum map will be collectively referred to as the spatial spectrum map. This spatial spectrum map SPS1 is data that visually shows in which direction and what kind of spatial spectrum peak values ​​that could be sound sources appear in the sound collection space SPA1, as viewed from the placement position of the sound collection device AM0. Search range information {θ i} is information that is pre-stored in the spatial spectrum calculation unit 131, for example, and indicates which directional range of sound sources in the sound collection space SPA1 should be searched for, relative to the placement position of the sound collection device AM0. In equation (1), θ represents a sound source direction vector that includes azimuth and elevation angles as components. The spatial spectrum calculation unit 131 outputs the spatial spectrum P(θ), which is a vector component calculated according to equation (1), to the spatial spectrum correction unit 132.

[0026]

number

[0027] The spatial spectrum correction unit 132 uses the correction filter (i.e., a correction filter based on correction gain and correction width) calculated by the correction filter calculation unit 135, which is a vector component, and the estimated sound source direction value θ^, which is a vector component, to perform a correction process that suppresses the peak component corresponding to the estimated sound source direction value θ^ and the surrounding components on the spatial spectrum map (e.g., spatial spectrum map SPS2) corresponding to the current frame from the spatial spectrum calculation unit 131. Figure 3 shows the spatial spectrum map SPS3 after the spatial spectrum map SPS2 has been corrected by the spatial spectrum correction unit 132 following the sound source tracking process. In this spatial spectrum map SPS3, of the two peak components 1 and 2 that appeared in spatial spectrum map SPS2, peak component 1 and the surrounding components have been suppressed. If no new correction filter is input from the correction filter calculation unit 135, the spatial spectrum correction unit 132 either outputs the spatial spectrum from the spatial spectrum calculation unit 131 to the peak detection unit 133 without correction, or performs the correction process using the latest correction filter acquired by the spatial spectrum correction unit 132 at that time. The spatial spectrum calculation unit 131 outputs the spatial spectrum P(θ) or the corrected spatial spectrum P'(θ) after correction processing to the peak detection unit 133.

[0028] The peak detection unit 133 searches for and detects peak components appearing in the spatial spectrum map (see Figure 3) of the spatial spectrum P(θ), which is a vector component from the spatial spectrum correction unit 132, or the corrected spatial spectrum P'(θ), which is a vector component after correction processing. The peak detection unit 133 outputs the direction estimate value θ, which is a vector component including the azimuth angle θ and elevation angle φ, which are the detection results, to the sound source tracking processing unit 14. Note that if the calculation result or correction result of the spatial spectrum is input to the sound source tracking processing unit 14, the peak detection unit 133 may be omitted.

[0029] The correction amount calculation unit 134 receives the sound identification probability value p(θ^) corresponding to the frame before the sound source identification processing unit 16 and the sound source direction estimation value θ^, which is a vector component, and calculates the correction gain and correction width as correction amounts from the sound identification probability value p(θ^) to suppress the peak component of the spatial spectral map and its surroundings, and outputs them together with the sound source direction estimation value θ^, which is a vector component, to the correction filter calculation unit 135. The correction gain corresponds to the amount of suppression, for example, when suppressing peak component L1 of the peak components L1 and L2 appearing in the spatial spectral map SPS1. The correction width corresponds to the suppression width (i.e., the suppression width of azimuth angle θ and elevation angle φ), for example, when suppressing peak component L1 of the peak components L1 and L2 appearing in the spatial spectral map SPS1.

[0030] The correction filter calculation unit 135 calculates a correction filter by updating the parameters of the correction filter for suppressing the peak components and their surroundings in the spatial spectral map, using the sound source direction estimate θ^ and the correction amount (i.e., correction gain and correction width) from the correction amount calculation unit 134. The correction filter calculation unit 135 outputs the calculated correction filter along with the sound source direction estimate θ^ to the spatial spectral correction unit 132. Although the correction filter calculation unit 135 is provided as a separate component from the correction amount calculation unit 134, it may be included in the correction amount calculation unit 134.

[0031] Here, we will explain an example of how the correction filter is calculated by the correction filter calculation unit 135. Two examples of correction filter calculation are described here. The correction filter calculated by the correction filter calculation unit 135 is designed and generated using parameters obtained from the correction amount calculation unit 134.

[0032] (1) First calculation example One example of implementing a correction filter is a design based on the reciprocal of a probability distribution on a sphere, such as the Von-Mises-Fisher distribution. In this design example, the estimated sound source direction vector is μ^=[cosθ^cosφ^, sinθ^sinφ^, sinφ^], based on the vector component sound source direction estimate θ^=[θ^, φ^]. TWhen calculated as shown above, the Von-Mises-Fisher distribution in 3D space is calculated as shown in equation (2). The bolded letters in equation (2) indicate vector components.

[0033]

number

[0034] Here, x is the correction direction θ = [θ, φ] T This is converted into a three-dimensional sound source direction vector, similar to the sound source direction estimate, and C3(κ) on the right-hand side of equation (2) is a normalization coefficient, as shown in equation (3). Furthermore, κ is called the concentration parameter and is a parameter for controlling the steepness of the distribution, so it can be determined based on the correction range (Q value) obtained from the correction amount calculation unit 134. The reciprocal of this distribution is shown as in equation (4).

[0035]

number

[0036]

number

[0037] In equation (4), Θ is the angle between x and μ^, and ε is an infinitesimal value to prevent division by zero. After modifying this function to vary in the range [0, 1] by min-max normalization, the linear gain G = 10 is calculated from the correction gain A [dB] obtained from the correction amount calculation unit 134. (A / 20) By converting the range to the range [G, 1] based on this, it becomes possible to calculate a correction filter (mask) that provides an attenuation of A [dB] in the direction of the estimated sound source, and no attenuation occurs at positions away from the estimated direction. Note that a negative value of A results in attenuation, but if set to a positive value, it can be designed as a correction mask that emphasizes the direction of the estimated sound source by increasing the gain in the direction of the estimated sound source, and no attenuation occurs at positions away from the estimated direction.

[0038] (2) Second calculation example Also, as another example of realizing a correction filter, within the internal region of the angular width α centered on the estimated sound source direction vector, there is a constant attenuation gain G in = 10 (A / 20) and in the external region, there is a different gain value G out = 10 (B / 20) It may be realized as a correction filter having such characteristics. Here, the angular width α is an angle in radians and can be determined based on the correction width (Q value) obtained from the correction amount calculation unit 134. Also, B is a parameter for determining the dB value of the attenuation amount in the external region, and when it is 0 dB, the spatial spectrum in the external region does not change. Such a correction filter can be calculated as shown in Equation (5) using the step function u(x) that becomes 1 when x is 0 or more and 0 otherwise.

[0039]

Equation

[0040] In the spatial spectrum correction unit 132, the correction filter given by the correction amount calculation unit 134 is applied to obtain the corrected spatial spectrum P’(θ) as shown in Equation (6).

[0041]

Equation

[0042] The sound source tracking processing unit 14 receives the direction estimation value θ or the calculation or correction result of the spatial spectrum from the sound source localization processing unit 13, and performs sound source tracking processing to track the position of a sound source that can move within the sound pickup space SPA1 by performing known signal processing on this direction estimation value θ or the calculation or correction result of the spatial spectrum. The sound source tracking processing unit 14 outputs the direction tracking value θ^ of the sound source as a result of the sound source tracking processing to the beamforming processing unit 15, and may also output the sound source localization result to the output device 50. In addition, the sound source tracking processing unit 14 may simultaneously output multiple direction estimation values ​​θ^ to the beamforming processing unit 15.

[0043] The beamforming processing unit 15 receives the direction tracking value θ^ from the sound source tracking processing unit 14 and the acoustic signal picked up by the sound pickup device AM0, and performs beamforming processing to emphasize the acoustic signal in the direction indicated by this direction tracking value θ^. The beamforming processing unit 15 outputs the emphasized audio signal after beamforming processing to the sound source identification processing unit 16, and may further output the emphasized audio result to the output device 50. Note that the configuration of the beamforming processing unit 15 may be omitted. Also, if multiple direction estimate values ​​θ^ are input to the beamforming processing unit 15, the i-th direction estimate value θ i ^ Corresponding enhanced audio signal S i The output may be sent to the sound source identification processing unit 16.

[0044] The sound source identification processing unit 16 inputs the amplified audio signal from the beamforming processing unit 15 to a pre-trained model such as Artificial In2telligence (AI) that has already been generated by the learning process, and performs a sound source identification process to identify what type of sound source the input amplified audio signal originated from. The sound source identification processing unit 16 identifies the type of sound source as a result of the sound source identification process and determines whether the identified type of sound source matches a pre-set target type of sound source. If the sound source identification processing unit 16 determines that the identified type of sound source matches a pre-set target type of sound source, it generates a screen notifying the user of this fact and outputs it to the output device 50. Meanwhile, the sound source identification processing unit 16 outputs the sound identification probability value p obtained as a result of the sound source identification process for the j-th sound source type. j (θ i ^) and estimated sound source direction (θ i The value ^) is output to the correction amount calculation unit 134 of the sound source localization processing unit 13.

[0045] 3. Operation of the sound source localization device Next, with reference to Figure 4, an example of the operation of the sound source localization device 10 according to Embodiment 1 will be described. Figure 4 is a flowchart showing an example of the operation procedure of the sound source localization device 10 according to Embodiment 1 in chronological order.

[0046] In Figure 4, the time-frequency conversion unit 11 receives an acoustic signal from the sound pickup unit AM1 (see Figure 2) and converts this acoustic signal into a time-frequency domain signal by applying predetermined signal processing (e.g., discrete Fourier transform) (step St1). The spatial frequency conversion unit 12 receives the time-frequency domain signal from the time-frequency conversion unit 11 and converts this time-frequency domain signal into a spatial frequency domain signal by applying predetermined signal processing (e.g., spatial Fourier transform, circular harmonic transform, spherical harmonic function transform, etc.) (step St1). The sound source localization processing unit 13 takes the spatial frequency domain signal generated in step St1 as input and calculates a spatial spectrum corresponding to the search range information, which is a pre-set vector component (step St2).

[0047] If the sound source localization processing unit 13 has a correction filter generated based on the processing in step St8 as input, it uses that correction filter to perform a correction process that suppresses the peak component corresponding to the estimated sound source direction on the spatial spectral map calculated in step St2 and the area around that peak component (step St3). If the sound source localization processing unit 13 does not have a correction filter, it does not perform the correction process on the spatial spectrum calculated in step St2, and in this case, step St3 is omitted. Alternatively, if the sound source localization processing unit 13 does not have a new correction filter based on the processing in step St8, it performs the correction process using the currently held correction filter (step St3). Furthermore, the sound source localization processing unit 13 may apply a forgetting process to the correction filter generated and held based on the processing in step St8 to gradually make it closer to a filter that allows the signal to pass through.

[0048] The sound source localization processing unit 13 searches for and detects peak components appearing in the spatial spectrum map (see Figure 3) of the spatial spectrum calculated in step St2 or the corrected spatial spectrum processed in step St3 (step St4). The sound source tracking processing unit 14 performs sound source tracking processing to track the position of a sound source that can move within the sound pickup space SPA1 by performing known signal processing on the direction estimate values ​​corresponding to the peaks detected in step St4 (step St5). Furthermore, the beamforming processing unit 15 receives the direction tracking values ​​from the sound source tracking processing unit 14 and the acoustic signal picked up by the sound pickup device AM0, and performs beamforming processing to emphasize the acoustic signal in the direction indicated by these direction tracking values ​​(step St5).

[0049] The sound source identification processing unit 16 receives the enhanced audio signal from the beamforming processing unit 15 and performs sound source identification processing to identify the type of sound source from which the input enhanced audio signal originated (step St6). The sound source identification processing unit 16 outputs the sound identification probability value and sound source direction estimate obtained as a result of the sound source identification processing to the sound source localization processing unit 13. The sound source localization processing unit 13 receives the sound identification probability value and sound source direction estimate, which are the results of the sound source identification processing in step St6, and calculates a correction gain and correction width as correction amounts from the sound identification probability value to suppress the peak component and its surroundings in the spatial spectral map (step St7). The sound source localization processing unit 13 uses the correction amounts calculated in step St7 to calculate a correction filter that suppresses the peak component and its surroundings corresponding to the sound source direction estimate on the spatial spectral map (step St8). By applying the result of the correction filter obtained in step St8 to the spatial spectral correction unit 132 for the next frame processing, a spatial spectral map of the sound source that reflects the result of the sound source identification processing can be obtained.

[0050] As described above, the sound source localization device 10 according to Embodiment 1 feeds back the results of the sound source identification process to the processing of the next frame, thereby suppressing sound source estimation results other than the target type of sound source in the entire spatial spectral map, and suppressing the peak components and surrounding components of the parameters (e.g., spatial spectrum) used in the sound source localization process. As a result, the sound source localization device 10 can improve the accuracy of sound source localization processing for sound sources of the target type present in the sound pickup space SPA1.

[0051] (Embodiment 2) The configuration and operation of the sound source localization device 10a according to Embodiment 2 are the same as those of the sound source localization device 10 according to Embodiment 1, except that the content of the sound source localization processing differs. The configuration of the sound source localization system 100 according to Embodiment 2 is the same as that of the sound source localization system 100 according to Embodiment 1 (see Figure 1), so its description is omitted. In addition, in the description of the sound source localization device 10a according to Embodiment 2, the same reference numerals are used for parts that overlap with the configuration of the sound source localization device 10 according to Embodiment 1, and the explanation is simplified or omitted, while different parts are described.

[0052] 4. Configuration of the sound source localization device First, the configuration of the sound source localization device according to Embodiment 2 will be described with reference to Figures 5 and 6. Figure 5 is a block diagram showing an example of the configuration of the sound source localization device according to Embodiment 2. Figure 6 is a diagram showing an example of the peak detection and rejection processing of the spatial spectrum. The sound source localization device 10a includes a time-frequency conversion unit 11, a spatial frequency conversion unit 12, a sound source localization processing unit 13a, a sound source tracking processing unit 14, a beamforming processing unit 15, and a sound source identification processing unit 16. The sound source localization processing unit 13a is composed of a semiconductor chip or dedicated hardware on which at least one electronic device such as a CPU, DSP, GPU, FPGA, and memory such as RAM and ROM are mounted.

[0053] The sound source localization processing unit 13a, upon receiving the sound identification probability value p(θ^) and the estimated sound source direction value θ^, which is a vector component, as input from the sound source identification processing unit 16, corrects the parameters used for sound source localization processing (e.g., spatial spectrum) based on the results (e.g., rejects peak components). The sound source localization processing unit 13a includes a spatial spectrum calculation unit 131, a peak detection unit 133, and a peak rejection processing unit 136. The processing contents of the spatial spectrum calculation unit 131 and the peak detection unit 133 are the same as in Embodiment 1, so their explanation is omitted here.

[0054] The peak rejection processing unit 136 processes the peak direction θ, which is a vector component appearing in the spatial spectral map (see Figure 6) detected by the peak detection unit 133. ― The azimuth angle θ and elevation angle φ indicating the peak direction PK1 (for example, in Figure 6) are input as detection results. When the peak rejection processing unit 136 receives the sound identification probability value p(θ^) and the estimated sound source direction value θ^, which is a vector component, from the sound source identification processing unit 16, the peak rejection processing unit 136 receives the sound identification probability value p(θ^) and the peak direction θ^, which is a vector component of the spatial spectral map of the current frame, from the peak detection unit 133. ― If the peak components L1 and L2 in Figure 6 (for example) satisfy predetermined conditions (see Figure 7), the detection result of the peak direction, which is the vector component (for example, peak component L1), is rejected and not output to the sound source tracking processing unit 14.

[0055] The predetermined conditions referred to here include, for example, the peak direction θ, which is a vector component. ― This corresponds to a situation where the sound is in the vicinity of the sound recognition probability value p(θ^) and that sound recognition probability value p(θ^) is less than a predetermined threshold. Figure 6 shows the spatial spectral map SPS3 after the corresponding spatial spectral map SPS2 has been processed by the peak rejection processing unit 136 after sound source tracking processing. In this spatial spectral map SPS3, of the two peak components L1 and L2 that appeared in spatial spectral map SPS2, the peak component L1 that satisfies the predetermined conditions has been rejected.

[0056] Furthermore, if the peak rejection processing unit 136 does not receive a sound identification probability value p(θ^) corresponding to the result of the sound source identification processing one frame prior from the sound source identification processing unit 16, and a sound source direction estimate value θ^ which is a vector component, the peak direction θ^ which is a vector component appearing in the spatial spectral map (see Figure 6) detected by the peak detection unit 133 will be used. ―The azimuth angle θ and elevation angle φ indicating the peak direction PK1 (for example, in Figure 6) are output to the sound source tracking processing unit 14 as the direction estimate value θ. The peak rejection processing unit 136 outputs the direction estimate value θ, which is a vector component including the azimuth angle θ and elevation angle φ that are the detection results of the peak remaining after the rejection process, to the sound source tracking processing unit 14.

[0057] 5. Operation of the sound source localization device Next, with reference to Figure 7, an example of the operation of the sound source localization device 10a according to Embodiment 2 will be described. Figure 7 is a flowchart showing an example of the operation procedure of the sound source localization device 10a according to Embodiment 2 in chronological order. In the explanation of Figure 7, steps that overlap with the operation of the sound source localization device 10 shown in Figure 4 will be given the same step number to simplify or omit the explanation, while different content will be explained.

[0058] In Figure 7, after step St2, the sound source localization processing unit 13a searches for and detects peak components, which are vector components, appearing in the spatial spectrum calculated in step St2 (see Figure 6) (step St11). When the sound source localization processing unit 13a receives the sound identification probability value p(θ^) and the estimated sound source direction value θ^, which is a vector component, as input as a result of the processing in step St6 (see Figure 4), it identifies the peak direction θ^, which is a vector component of the spatial spectrum map of the current frame corresponding to the sound identification probability value p(θ^). ― It is determined whether the current is in the vicinity of the direction estimate θ^ of the sound source localization processing unit 13a from the previous frame (step St12). The peak direction θ is a vector component. ― If it is determined that the location is not in the vicinity of the direction estimate θ^ of the sound source localization processing unit 13a from the previous frame (step St12, NO), the processing of the sound source localization device 10a proceeds to step St5.

[0059] On the other hand, the sound source localization processing unit 13a controls the peak direction θ, which is a vector component. ―If it is determined that the sound source is in the vicinity of the direction estimate value θ^, which is a vector component of the sound source localization processing unit 13a from the previous frame (step St12, YES), it is determined whether the direction estimate value θ^, which is a vector component, is of the desired (target) type of sound source based on the sound identification probability value p(θ^) (step St13). If it is determined that the direction estimate value θ^, which is a vector component, is not of the desired (target) type of sound source (step St13, NO), the processing of the sound source localization device 10a proceeds to step St5.

[0060] On the other hand, if the sound source localization processing unit 13a determines that the direction estimate value θ^, which is a vector component, is of the desired (target) type of sound source (step St13, YES), it determines whether the sound identification probability value p(θ^) is less than a predetermined threshold as a result of the processing in step St6 (step St14). If it is determined that the sound identification probability value p(θ^) is greater than or equal to the predetermined threshold (step St14, NO), the processing of the sound source localization device 10a proceeds to step St5.

[0061] On the other hand, if the sound source localization processing unit 13a determines that the sound identification probability value p(θ^) is less than a predetermined threshold (step St14, YES), then the peak direction θ, which is the vector component of the spatial spectral map of the current frame corresponding to that sound identification probability value p(θ^), is determined by the sound source localization processing unit 13a. ― A rejection process is performed to remove the peak in the spatial spectral map of (for example, peak component L1) (step St15). After step St15, the processing of the sound source localization device 10a proceeds to step St5. Although not shown in Figure 7, similar to Embodiment 1, the sound source identification processing unit 16 outputs the sound identification probability value p(θ^) and the sound source direction estimate value θ^, which is a vector component, obtained as a result of the sound source identification processing in step St6, to the sound source localization processing unit 13a.

[0062] As described above, according to the sound source localization device 10a of Embodiment 2, by feeding back the results of the sound source identification process to the processing of the next frame, sound source estimation results other than those of the target type are suppressed in the entire spatial spectral map, and the peak components of the parameters (e.g., spatial spectrum) used in the sound source localization process are rejected. As a result, the sound source localization device 10a can improve the processing accuracy of sound source localization targeting sound sources of the target type present in the sound pickup space SPA1.

[0063] (Embodiment 3) In the sound source localization device 10b according to Embodiment 3, unlike the sound source localization device 10 according to Embodiment 1 and the sound source localization device 10a according to Embodiment 2, the sound identification probability value p(θ^) and the sound source direction estimation value θ^, which is a vector component, corresponding to the sound one frame prior from the sound source identification processing unit 16 are input to the sound source tracking processing unit 14b instead of the sound source localization processing units 13 and 13a (see Figure 8). The configuration of the sound source localization system 100 according to Embodiment 3 is the same as the sound source localization system 100 according to Embodiment 1 (see Figure 1), so the explanation is omitted. In addition, in the explanation of the sound source localization device 10b according to Embodiment 3, the same reference numerals are used for parts that overlap with the configuration of the sound source localization device 10 according to Embodiment 1, and the explanation is simplified or omitted, while different content is explained.

[0064] 6. Configuration of the sound source localization device First, the configuration of the sound source localization device according to Embodiment 3 will be described with reference to Figures 8 to 9. Figure 8 is a block diagram showing an example configuration of the sound source localization device according to Embodiment 3. Figure 9 is a diagram showing various parameters related to likelihood correction processing and their explanations. The sound source localization device 10b comprises a time-frequency conversion unit 11, a spatial frequency conversion unit 12, a sound source localization processing unit 13b, a sound source tracking processing unit 14b, a beamforming processing unit 15, and a sound source identification processing unit 16. The sound source localization processing unit 13b and the sound source tracking processing unit 14b are each composed of a semiconductor chip or dedicated hardware on which at least one electronic device such as a CPU, DSP, GPU, FPGA, and memory such as RAM and ROM are mounted.

[0065] Here, we will explain the concept of sound source tracking processing by the sound source tracking processing unit 14b according to Embodiment 3.

[0066] In sound source tracking, as with general dynamic systems, a moving sound source can be modeled using a motion model (in other words, a process equation) and an observation system model (in other words, an observation equation). The state vector at discrete time k is x k , the error in the motion model is u k-1 , the function in the process equation is f k-1 (·) The observation vector, which has the observed values ​​from the sensor as its elements, is z k , the error in the observation system is v k , the function in the observation equation is h k-1 If (·) is assumed, the motion model can be modeled as shown in equation (7), and the observation system model can be modeled as shown in equation (8).

[0067]

number

[0068]

number

[0069] Furthermore, if we express the motion model and the observation system model in the form of probability density functions, they can be modeled as shown in equations (9) and (10). The state vector x at discrete time k k It is composed of a vector that includes the position state vector of the sound source (see equation (11)) and the angular velocity state vector of the sound source (see equation (12)), and is defined as shown in equation (13). In equation (11), θ represents the azimuth angle and φ represents the elevation angle.

[0070]

number

[0071]

number

[0072]

number

[0073]

number

[0074]

number

[0075] The state vector may contain any state to be tracked, such as angular acceleration, in addition to angular velocity. Alternatively, instead of angular information, it may contain position in Cartesian coordinates (see equation (14)) along with its velocity and acceleration as states. The observed vector z at discrete time k. k This is the observed position vector θ^ of the sound source. k Consists of, z k =[θ^ k ] is defined as follows. In Embodiment 3, the observed position vector θ of the sound source is defined as follows. k This corresponds to the direction estimation value θ, which is the processing result of the sound source localization processing unit 13.

[0076]

number

[0077] The problem of sound source tracking is based on this probability density function model, using observed values ​​Z from the past to the present. 1:k =[z1, z2, ..., z k Given ], the current state vector x is recursively obtained. k The posterior probability density p(x k |Z 1:k This reduces to the problem of finding the posterior probability density. Once the posterior probability density is determined, it becomes possible to estimate the state, such as the current position of the sound source, using parameter estimation methods. The current state vector x kIt is known that the posterior probability density of can be formulated as shown in equation (15). Equation (15) gives the transition probability density p(x k |x k-1 ) and likelihood p(z k |x k If we obtain the following, then the prior probability density p(x) at discrete time k-1 can be determined. k |Z 1:k-1 ) from the posterior probability density p(x) at discrete time k k |Z 1:k ) can be calculated recursively. However, the normalization term is omitted in equation (15).

[0078]

number

[0079] Therefore, in Embodiment 3, when the sound source tracking processing unit 14b receives the result from the sound source identification processing unit 16, it uses a particle filter to calculate the posterior probability density (calculation of equation (15)) and includes the observed value z k (That is, the sound source direction estimation value θ^, which is the sound source localization direction input from the sound source identification processing unit 16) k-1 Likelihood p(z) k |x k A process is added to correct for () and estimate the position of a movable sound source. Note that the likelihood p(z k |x k ) The spatial spectrum calculated by the spatial spectrum calculation unit 131 in the sound source localization processing unit 13a or the sound source localization processing unit 13a may be used as ). Note that the vector x used in Embodiment 3 k、 z k、 θ^ k Regarding this, instead of repeatedly stating that it is a vector, simply x k、 z k、 θ^ k It should be written as follows.

[0080] The sound source localization processing unit 13b receives a spatial frequency domain signal from the spatial frequency conversion unit 12 and performs sound source localization processing to determine the direction in which a sound source that can move within the sound pickup space SPA1 exists by performing known signal processing on this spatial frequency domain signal. The sound source localization processing unit 13 outputs a sound source direction tracking value θ^ as a result of the sound source localization processing to the sound source tracking processing unit 14b, and may further output the sound source localization result to the output device 50.

[0081] The sound source tracking processing unit 14b includes a sampling processing unit 141, a likelihood calculation unit 142, a likelihood correction / weight calculation unit 143, a weight normalization unit 144, a resampling processing unit 145, a posterior probability density calculation unit 146, and a state estimation processing unit 147. The sound source tracking processing unit 14b has N p Using a particle filter having a certain number of particles and a direction estimation value θ from the sound source localization processing unit 13b, a sound source tracking process is performed to track the position of a sound source that can move within the sound pickup space SPA1.

[0082] The sampling processing unit 141 processes the N data, which is the result of processing by the resampling processing unit 145. p Input the resampled state vector of individual particles i, and calculate N at discrete time k-1. p The state vector x of individual particle i k-1 (i) The output of the posterior probability density calculation unit 146 at discrete time k-1 is the posterior probability density p(x k-1 (i) |Z 1:k-1 The prior probability density p(x) is treated as the prior probability density at discrete time k. The sampling processing unit 141 calculates the prior probability density p(x) k-1 (i) |Z 1:k-1 The calculation result of ) and the transition probability density p(x) read from the memory (not shown) of the sound source tracking processing unit 14b (see equation (9)) are used. k |x k-1 (For example, when applying a particle filter, any probability density function can be applied.) Using this, the prior probability density p(x k-1 (i) |Z 1:k-1) to generate the sample x which is a vector component k (i) (sampling). Note that this sample x which is a vector component k (i) has a distribution that approximates the predicted probability density p(x k (i) |Z 1:k-1 ). Then, the sampling processing unit 141 outputs the sample value x which is a vector component for each of the N p particles i at the discrete time k, and the weight 1 / Np to the likelihood calculation unit 142 and the resampling processing unit 145 respectively. k (i)

[0083] The likelihood calculation unit 142 uses the calculation result of the sample value x for each of the N p particles i at the discrete time k from the sampling processing unit 141 and the direction estimation value θ (= z k (i) k : = [θ k ) from the sound source localization processing unit 13b to calculate the likelihood p(z p |x k k (i) k k (i) ) for each of the Nparticles i. The likelihood calculation unit 142 outputs the calculation result of the likelihood p(z k |x k (i) ) to the likelihood correction / weight calculation unit 143. Alternatively, it is also possible to apply the spatial spectrum P(θ) or the corrected spatial spectrum P’(θ) as an input instead of the direction estimation value θ.

[0084] When the likelihood correction / weight calculation unit 143 inputs the sound discrimination probability value p(θ^ k-1 ) and the sound source direction estimation value θ^ k-1 from the sound source discrimination processing unit 16, the calculation result of the likelihood p(z k |x k (i) ) from the likelihood calculation unit 142 is combined with the sound discrimination probability value p(θ^ k-1 ) and the sound source direction estimation value θ^ k-1The correction is performed using this. As a result, the likelihood correction / weight calculation unit 143 calculates N at discrete time k according to equation (16). p Weight w for each individual particle i k (i) The likelihood correction and weight calculation unit 143 calculates N at discrete time k. p Weight w for each individual particle i k (i) The calculation result is output to the weight normalization unit 144. In equation (16), g is a processing term that corrects the likelihood p. In other words, the likelihood correction / weight calculation unit 143 calculates the estimated sound source direction θ^ k-1 And the state vector x in particle i k (i) The parameter θ^ k (i) Based on the angle difference, the sound recognition probability value p(θ^ k-1 The correction term is calculated using the parameter θ^ k-1 The sound source direction is estimated θ^ k-1 If it is in the vicinity of, the sound recognition probability value p(θ^ k-1 Based on this, the case is more likely to be rejected, while on the other hand, if the angle difference is large, a value is calculated in which the value of the correction term is close to 1.

[0085]

number

[0086] The weight normalization unit 144 calculates N at discrete time k from the likelihood correction / weight calculation unit 143. p Weight w for each individual particle i k (i) The calculation result is N p By dividing by the sum of the weights of all particles i, we obtain the weight w for each particle i. k (i) The weights are normalized. The weight normalization unit 144 calculates the weights w for each particle i. k (i) w is the normalized result ~ k (i) The output is sent to the resampling processing unit 145.

[0087] The resampling unit 145 processes N at discrete time k from the sampling unit 141. p The sample value x is the vector value for each individual particle i. k (i) The calculation results and the weights w for each particle i from the weight normalization unit 144 k (i) w is the normalized result ~ k (i) Using this, N at discrete time k p Sample x is a vector value with a uniform weight of 1 / Np for each particle i. ~ k (i) Recalculate (resampling). Note that sample x ~ k (i) This is an approximation of the posterior probability density. The resampling processing unit 145 calculates N at discrete time k after resampling. p Sample x per individual particle i ~ k (i) The system outputs the uniform weight 1 / Np after resampling to the posterior probability density calculation unit 146.

[0088] The posterior probability density calculation unit 146 calculates N at the discrete time k after resampling from the resampling processing unit 145. p Sample x is the vector component for each individual particle i. ~ k (i) And using a uniform weight 1 / Np, according to equation (17), N p Posterior probability density p(x) for each individual particle i k (i) |Z 1:k Calculate ).

[0089]

number

[0090] The posterior probability density calculation unit 146 calculates N pPosterior probability density p(x) for each individual particle i k (i) |Z 1:k The calculation result is output to the state estimation processing unit 147. It is also used as the prior probability density for the sampling processing unit 141 when processing the next frame.

[0091] The state estimation processing unit 147 uses an existing parameter estimation method to calculate the N from the posterior probability density calculation unit 146. p Posterior probability density p(x) for each individual particle i k (i) |Z 1:k From the calculation results, the estimated state vector of the sound source at discrete time k is x^ k The state estimation processing unit 147 estimates the state estimation result vector x^ of the sound source at discrete time k. k The output is sent to the beamforming processing unit 15. The beamforming processing unit forms a beam based on the sound source position estimation vector θ^, which is one element of the state estimation vector.

[0092] 7. Operation of the sound source localization device Next, with reference to Figure 10, an example of the operation of the sound source localization device 10b according to Embodiment 3 will be described. Figure 10 is a flowchart showing an example of the operation procedure of the sound source localization device 10b according to Embodiment 3 in chronological order. In the explanation of Figure 10, the same step number is assigned to parts that overlap with the operation of the sound source localization device 10 shown in Figure 4, and the explanation is simplified or omitted, while different content is explained.

[0093] In Figure 10, after step St1, the sound source localization processing unit 13b receives the spatial frequency domain signal from the spatial frequency conversion unit 12 and performs sound source localization processing to identify the direction in which a sound source that can move within the sound pickup space SPA1 exists by performing known signal processing on this spatial frequency domain signal (step St21). The sound source tracking processing unit 14b obtains the direction estimate, which is the result of the sound source localization processing in step St21 (step St22). The sound source tracking processing unit 14b obtains the N, which is the processing result of the resampling processing unit 145. pInput the resampled state vector of individual particles i, and calculate N at discrete time k-1. p The state vector x of individual particle i k-1 (i) The output of the posterior probability density calculation unit 146 at discrete time k-1 is the posterior probability density p(x k-1 (i) |Z 1:k-1 The sound source tracking unit 14b treats this prior probability density p(x) as the prior probability density at discrete time k. k-1 (i) |Z 1:k-1 The calculation result of ) and the transition probability density p(x) read from the memory (not shown) of the sound source tracking processing unit 14b (see equation (9)) are used. k |x k-1 Using ), the prior probability density p(x k-1 (i) |Z 1:k-1 ) from which the vector component is sample x k (i) This occurs (step St23).

[0094] The sound source tracking processing unit 14b determines N at discrete time k calculated in step St23. p Sample value x for each individual particle i k (i) The calculation result and the direction estimation value θ(=z) from the sound source localization processing unit 13b obtained in step St22 k :=[θ k Using ]), N p Likelihood p(z) for each individual particle i k |x k (i) The sound source tracking unit 14b calculates the sound identification probability value p(θ^) obtained from the sound source identification unit 16 in step St6 described later. k-1 ) and estimated sound source direction θ^ k-1 When this is entered, the likelihood calculation unit 142 calculates the likelihood p(z k |x k (i) The result of the calculation is the sound identification probability value p(θ^) from the sound source identification processing unit 16. k-1 ) and estimated sound source direction θ^ k-1Correction is performed using (step St25). As a result, the sound source tracking processing unit 14b determines N at discrete time k. p Weight w for each individual particle i k (i) The sound source tracking processing unit 14b calculates N at discrete time k in step St25. p Weight w for each individual particle i k (i) The calculation result is N p By dividing by the sum of the weights of all particles i, we obtain the weight w for each particle i. k (i) Normalize the result (step St26).

[0095] The sound source tracking processing unit 14b receives N at discrete time k from the sampling processing unit 141. p The calculation results of the predicted probability density for each individual particle i and the weights w for each particle i from the weight normalization unit 144. k (i) w is the normalized result ~ k (i) Using this, N at discrete time k p Predicted probability density for each particle i (i.e., p(x) k (i) |Z 1:k-1 )) is recalculated (step St27). The sound source tracking processing unit 14b calculates N at discrete time k after resampling in step St27. p Predicted probability density for each particle i (i.e., p(x) k (i) |Z 1:k-1 )) and the weight per particle i in step St26 w k (i) w is the normalized result ~ k (i) Using N p Posterior probability density p(x) for each individual particle i k (i) |Z 1:kThe sound source localization device 10b calculates the sound source (step St28). After step St28, the processing of the sound source localization device 10b proceeds to step St5. Although not shown in Figure 10, if the type of sound source identified as a result of the processing in step St6 does not match the type of target sound source set in advance, the sound source identification processing unit 16 outputs the sound identification probability value p(θ^) and the sound source direction estimate value θ^ obtained as a result of the sound source identification processing to the sound source tracking processing unit 14b.

[0096] As described above, the sound source localization device 10b according to Embodiment 3 corrects the parameters used in the sound source tracking process (for example, the likelihood for calculating the posterior probability density) for a sound source of the target type by feeding back the result of the sound source identification process to the processing of the next frame. As a result, the sound source localization device 10b can improve the accuracy of sound source tracking for a sound source of the target type that exists within the sound pickup space SPA1.

[0097] (Embodiment 4) The sound source localization device 10c according to Embodiment 4 has a configuration that combines the configurations of the sound source localization device 10 according to Embodiment 1 and the sound source localization device 10b according to Embodiment 3. The configuration of the sound source localization system 100 according to Embodiment 3 is the same as the sound source localization system 100 according to Embodiment 1 (see Figure 1), so its description is omitted. In addition, in the description of the sound source localization device 10c according to Embodiment 4, components that overlap with the configurations of the sound source localization device 10 according to Embodiment 1 and the sound source localization device 10b according to Embodiment 3 are given the same reference numerals to simplify or omit the description, and different contents are described.

[0098] 8. Configuration and Operation of the Sound Source Localization Device The configuration and operation of the sound source localization device according to Embodiment 4 will be described with reference to Figures 11 and 12. Figure 11 is a block diagram showing an example configuration of the sound source localization device 10c according to Embodiment 4. Figure 12 is a flowchart showing an example of the operation procedure of the sound source localization device 10c according to Embodiment 4 in chronological order. The sound source localization device 10c comprises a time-frequency conversion unit 11, a spatial frequency conversion unit 12, a sound source localization processing unit 13 (see Figure 2), a sound source tracking processing unit 14b (see Figure 8), a beamforming processing unit 15, and a sound source identification processing unit 16.

[0099] In other words, as shown in Figures 11 and 12, in the sound source localization device 10c according to Embodiment 4, the direction estimate θ, which is the result of the sound source localization processing by the sound source localization processing unit 13, is input to the likelihood calculation unit 142 of the sound source tracking processing unit 14b. As explained with reference to Figure 10, the direction estimate θ, which is the result of the sound source localization processing, is N in step St24. p Likelihood p(z) for each individual particle i k |x k (i) It is used to calculate ).

[0100] Furthermore, the sound identification probability value p(θ^) and the sound source direction estimate value θ^, which are the sound source identification results from the sound source identification processing unit 16, are input to the correction amount calculation unit 134 of the sound source localization processing unit 13 (see the processing in the order of steps St6, St7, St8, and St3 in Embodiment 1), and are also input to the likelihood correction / weight calculation unit 143 of the sound source tracking processing unit 14b (see the processing in the order of steps St6 and St25 in Embodiment 3).

[0101] As described above, the sound source localization device 10c according to Embodiment 4 suppresses the peak component and surrounding components of the parameters (e.g., spatial spectrum) used for sound source localization processing for the target type of sound source by feeding back the result of the sound source identification processing to the processing of the next frame. Furthermore, the sound source localization device 10c corrects the parameters (e.g., likelihood for calculating the posterior probability density) used for sound source tracking processing for the target type of sound source by feeding back the result of the sound source identification processing to the processing of the next frame. As a result, the sound source localization device 10c can not only improve the processing accuracy of sound source localization targeting the target type of sound source present in the sound pickup space SPA1, but also improve the processing accuracy of sound source tracking targeting the target type of sound source present in the sound pickup space SPA1.

[0102] (Embodiment 5) The sound source localization device 10d according to Embodiment 5 has a configuration that combines the configurations of the sound source localization device 10a according to Embodiment 2 and the sound source localization device 10b according to Embodiment 3. The configuration of the sound source localization system 100 according to Embodiment 5 is the same as the sound source localization system 100 according to Embodiment 1 (see Figure 1), so the explanation is omitted. In addition, in the description of the sound source localization device 10d according to Embodiment 5, the same reference numerals are used to indicate parts that overlap with the configurations of the sound source localization device 10a according to Embodiment 2 and the sound source localization device 10b according to Embodiment 3, simplifying or omitting the explanation, while describing different contents.

[0103] 9. Configuration and Operation of the Sound Source Localization Device The configuration and operation of the sound source localization device according to Embodiment 5 will be described with reference to Figures 13 and 14. Figure 13 is a block diagram showing an example configuration of the sound source localization device 10d according to Embodiment 5. Figure 14 is a flowchart showing an example of the operation procedure of the sound source localization device 10d according to Embodiment 5 in chronological order. The sound source localization device 10d comprises a time-frequency conversion unit 11, a spatial frequency conversion unit 12, a sound source localization processing unit 13a (see Figure 5), a sound source tracking processing unit 14b (see Figure 8), a beamforming processing unit 15, and a sound source identification processing unit 16.

[0104] In other words, as shown in Figures 13 and 14, in the sound source localization device 10d according to Embodiment 5, the direction estimate θ, which is the result of the sound source localization processing by the sound source localization processing unit 13a, is input to the likelihood calculation unit 142 of the sound source tracking processing unit 14b. As explained with reference to Figure 10, the direction estimate θ, which is the result of the sound source localization processing, is N in step St24. p Likelihood p(z) for each individual particle i k |x k (i) It is used to calculate ).

[0105] Furthermore, the sound identification probability value p(θ^) and the sound source direction estimate value θ^, which are the sound source identification results from the sound source identification processing unit 16, are input to the peak rejection processing unit 136 of the sound source localization processing unit 13a (see the processing in the order of steps St6 and St12 in Embodiment 2), and are also input to the likelihood correction / weight calculation unit 143 of the sound source tracking processing unit 14b (see the processing in the order of steps St6 and St25 in Embodiment 3).

[0106] As described above, the sound source localization device 10d according to Embodiment 5 suppresses peak components of the parameters (e.g., spatial spectrum) used for sound source localization processing for the target type of sound source by rejecting them, by feeding back the result of the sound source identification processing to the processing of the next frame. Furthermore, the sound source localization device 10c corrects the parameters (e.g., likelihood for calculating the posterior probability density) used for sound source tracking processing for the target type of sound source by feeding back the result of the sound source identification processing to the processing of the next frame. As a result, the sound source localization device 10c can not only improve the processing accuracy of sound source localization targeting the target type of sound source present in the sound pickup space SPA1, but can also improve the processing accuracy of sound source tracking targeting the target type of sound source present in the sound pickup space SPA1.

[0107] (Note) Based on the above description of the embodiments, the following technical concepts are disclosed.

[0108] (Item 1) Based on the acoustic signals of the sound-gathering space (SPA1) that are picked up by a sound-gathering device (AM0) placed within the sound-gathering space where multiple sound sources may exist, a sound source localization process is performed to estimate the position of the sound sources. Based on the results of the sound source localization process, a sound source tracking process is performed to track the position of the sound source within the sound pickup space. Based on the results of the sound source tracking process, a sound source identification process is performed to identify the type of sound source. The process involves correcting the parameters based on the acoustic signal used for the sound source localization process and then performing the sound source localization process. Sound source localization method. As a result, according to the sound source localization method, by feeding back the results of the sound source identification process to the processing of the next frame, corrections to the parameters based on the acoustic signal used for sound source localization processing are applied to the type of target sound source, thereby improving the accuracy of sound source localization processing for target sound sources present in the sound pickup space.

[0109] (Item 2) The method further includes applying beamforming processing to the acoustic signal based on the results of the sound source tracking process, The sound source identification process is executed by inputting the acoustic signal after the beamforming process. The sound source localization method described in item 1. As a result, according to the sound source localization method, the acoustic signal in the direction from which the sound source is located, as viewed from the sound acquisition device, is relatively emphasized by beamforming processing, thereby improving the accuracy of sound source localization processing for target sound sources located within the sound acquisition space.

[0110] (Item 3) The parameter is a spatial spectrum indicating the direction of arrival of the plurality of sound sources as seen from the sound-collecting device within the sound-collecting space. If the result of the sound source identification process does not match the type of the target sound source, the peak component appearing in the spatial spectrum corresponding to the direction of arrival of the sound source that does not match the type of the target sound source, and the area surrounding that peak component, are suppressed based on the result of the sound source identification process. The sound source localization method described in item 1 or item 2. As a result, according to the sound source localization method, peak components corresponding to sound sources that do not match the type of target sound source can be suppressed in the spatial spectrum generated during the sound source localization process, thereby improving the accuracy of sound source localization processing for target sound sources present in the sound pickup space.

[0111] (Item 4) Based on the results of the sound source identification process, a correction filter is generated to suppress the peak component and the area surrounding the peak component that appear in the spatial spectrum in the direction of arrival of sound sources that do not match the type of sound source of the target. The spatial spectrum is corrected using the correction filter. The sound source localization method described in item 3. As a result, according to the sound source localization method, unwanted peak components in the spatial spectrum generated during sound source localization processing can be easily suppressed, and by updating the correction filter based on the results of the sound source identification processing, the accuracy of sound source localization processing targeting the desired sound source present in the sound pickup space can be improved.

[0112] (Item 5) The parameter is a spatial spectrum indicating the direction of arrival of the plurality of sound sources as seen from the sound-collecting device within the sound-collecting space. If the result of the sound source identification process does not match the type of the target sound source, the peak component appearing in the spatial spectrum corresponding to the direction of arrival of the sound source that does not match the type of the target sound source is rejected based on the result of the sound source identification process. The sound source localization method described in item 1 or 2. As a result, according to the sound source localization method, peak components corresponding to sound sources that do not match the type of target sound source can be rejected in the spatial spectrum generated during the sound source localization process, thereby improving the accuracy of sound source localization processing for target sound sources present in the sound pickup space.

[0113] (Item 6) In the sound source tracking process, the likelihood calculation result is corrected based on observed values ​​showing the results of the sound source localization process obtained from the start of sound acquisition by the sound acquisition device to the present, and the position of the sound source is identified by calculating the posterior probability density of the current state of the sound source using a particle filter. The sound source localization method described in any one of items 3 to 5. As a result, according to the sound source localization method, the processing term that corrects the likelihood used in the sound source tracking process using a particle filter can be updated based on the results of the sound source identification process, and appropriate weights can be calculated. This makes it possible to appropriately track the direction of a moving sound source and improves the accuracy of the sound source localization process for target sound sources present in the sound pickup space.

[0114] (Item 7) Based on the results of the sound source identification process, the particle filter is corrected and the sound source tracking process is executed. The sound source localization method described in item 6. As a result, according to the sound source localization method, by feeding back the results of the sound source identification process to the processing of the next frame, it becomes possible to perform likelihood correction used in sound source tracking processing using a particle filter for the type of target sound source, thereby appropriately improving the accuracy of sound source localization processing targeting the target sound source present in the sound pickup space.

[0115] (Item 8) A sound source localization processing unit (13) performs sound source localization processing to estimate the position of a sound source based on the acoustic signals of the sound pickup space (SPA1) that are picked up by a sound pickup device (AM0) placed in the sound pickup space where multiple sound sources may exist, A sound source tracking processing unit (14) performs a sound source tracking process to track the position of the sound source within the sound pickup space based on the results of the sound source localization processing, The system includes a sound source identification processing unit (16) that performs a sound source identification process to identify the type of sound source based on the results of the sound source tracking process, The sound source localization processing unit corrects the parameters based on the acoustic signal used for the sound source localization processing and performs the sound source localization processing. Sound source localization device. As a result, the sound source localization device feeds back the results of the sound source identification process to the processing of the next frame, thereby applying parameter corrections based on the acoustic signal used for sound source localization processing to the type of target sound source. This improves the accuracy of sound source localization processing for target sound sources present in the sound pickup space.

[0116] (Item 9) The sound source localization device (10), which is a computer, A process for estimating the position of a sound source based on the acoustic signals of the sound-collecting space, which are collected by a sound-collecting device placed in the sound-collecting space where multiple sound sources may exist, and a process for performing sound source localization processing, Based on the results of the sound source localization process, a process is performed to track the position of the sound source within the sound collection space. A process to perform a sound source identification process to identify the type of sound source based on the results of the sound source tracking process, A process to implement the process of correcting the parameters based on the acoustic signal used for the sound source localization process and executing the sound source localization process, program. As a result, according to the program, the results of the sound source identification process are fed back to the computer-based sound source localization device for processing the next frame. This applies parameter corrections based on the acoustic signal used for sound source localization processing to the type of target sound source, thereby improving the accuracy of sound source localization processing for target sound sources present in the sound pickup space.

[0117] While embodiments have been described above with reference to the attached drawings, this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments described above can be combined in any way without departing from the spirit of the invention. [Industrial applicability]

[0118] This disclosure is useful as a sound source localization method, sound source localization device, and program for improving the processing accuracy of sound source localization targeting a target sound source present in a sound pickup space. [Explanation of symbols]

[0119] 10 Sound source localization device 10a Sound source localization device 10b Sound source localization device 10c Sound source localization device 10d sound source localization device 11. Time-based frequency conversion section 12 Spatial frequency conversion section 13. Sound source localization processing unit 13a Sound source localization processing unit 13b Sound source localization processing unit 14. Sound source tracking processing unit 14b Sound source tracking processing unit 15 Beamforming Processing Unit 16 Sound Source Identification Processing Unit 50 Output Devices 100 Sound Source Localization Systems 131 Spatial Spectrum Calculation Unit 132 Spatial Spectrum Correction Unit 133 Peak detection unit 134 Correction amount calculation unit 135 Correction filter calculation unit 136 Peak rejection processing unit 141 Sampling Processing Unit 142 Likelihood Calculation Unit 143 Likelihood Correction and Weight Calculation Unit 144 Weight normalization section 145 Resampling Processing Unit 146 Posterior probability density calculation unit 147 State Estimation Processing Unit AI1 Audio Interface AM0 sound pickup device AM1 sound collection section

Claims

1. Based on the acoustic signals of the sound-collecting space picked up by a sound-collecting device placed in the sound-collecting space where multiple sound sources may exist, a sound source localization process is performed to estimate the position of the sound sources. Based on the results of the sound source localization process, a sound source tracking process is performed to track the position of the sound source within the sound collection space. Based on the results of the sound source tracking process, a sound source identification process is performed to identify the type of sound source. The process involves correcting the parameters based on the acoustic signal used for the sound source localization process and then performing the sound source localization process. Sound source localization method.

2. The method further includes applying beamforming processing to the acoustic signal based on the results of the sound source tracking process, The sound source identification process is performed by inputting the acoustic signal after the beamforming process. The sound source localization method according to claim 1.

3. The parameter is a spatial spectrum indicating the direction of arrival of the plurality of sound sources as seen from the sound-collecting device within the sound-collecting space. If the result of the sound source identification process does not match the type of the target sound source, the peak component appearing in the spatial spectrum corresponding to the direction of arrival of the sound source that does not match the type of the target sound source, and the area surrounding that peak component, are suppressed based on the result of the sound source identification process. The sound source localization method according to claim 1.

4. Based on the results of the sound source identification process, a correction filter is generated to suppress the peak component and the area surrounding the peak component that appear in the spatial spectrum in the direction of arrival of sound sources that do not match the type of sound source of the target. The spatial spectrum is corrected using the correction filter. The sound source localization method according to claim 3.

5. The parameter is a spatial spectrum indicating the direction of arrival of the plurality of sound sources as seen from the sound-collecting device within the sound-collecting space. If the result of the sound source identification process does not match the type of the target sound source, the peak component appearing in the spatial spectrum corresponding to the direction of arrival of the sound source that does not match the type of the target sound source is rejected based on the result of the sound source identification process. The sound source localization method according to claim 1.

6. In the sound source tracking process, the likelihood calculation result is corrected based on observed values ​​showing the results of the sound source localization process obtained from the start of sound acquisition by the sound acquisition device to the present, and the position of the sound source is identified by calculating the posterior probability density of the current state of the sound source using a particle filter. The sound source localization method according to claim 3 or 5.

7. Based on the results of the sound source identification process, the particle filter is corrected and the sound source tracking process is executed. The sound source localization method according to claim 6.

8. A sound source localization processing unit performs sound source localization processing to estimate the position of a sound source based on the acoustic signals of the sound pickup space picked up by a sound pickup device placed in the sound pickup space where multiple sound sources may exist. A sound source tracking processing unit performs a sound source tracking process to track the position of the sound source within the sound pickup space based on the results of the sound source localization processing, The system includes a sound source identification processing unit that performs a sound source identification process to identify the type of sound source based on the results of the sound source tracking process, The sound source localization processing unit corrects the parameters based on the acoustic signal used for the sound source localization processing and performs the sound source localization processing. Sound source localization device.

9. The sound localization device is a computer. A process for estimating the position of a sound source based on the acoustic signals of the sound pickup space, which are picked up by a sound pickup device placed in the sound pickup space where multiple sound sources may exist, and a process for performing sound source localization processing. Based on the results of the sound source localization process, a process is performed to track the position of the sound source within the sound collection space. A process to perform a sound source identification process to identify the type of sound source based on the results of the sound source tracking process, A process to implement the process of correcting the parameters based on the acoustic signal used for the sound source localization process and executing the sound source localization process, program.

Citation Information

Patent Citations

  • Sound source identifying apparatus and sound source identifying method

    JP2017044916A