Sound source localization device, sound source localization method, and program

By using techniques such as sound signal vector generation, subspace recognition and candidate vector recognition in sound source positioning equipment, the problem of large calculations of traditional variational inference methods is solved, and real-time and high-accurate sound source positioning in the conference environment is achieved.

CN114616483BActive Publication Date: 2025-05-27AUDIO TECHNICA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180005551.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-05
Filing Date
2021-09-16
Publication Date
2025-05-27
Estimated Expiration
2041-09-16

AI Technical Summary

Technical Problem

The traditional variational inference method is computationally expensive in real-time positioning of sound sources in conferences, making it difficult to achieve real-time positioning.

Method used

The sound source positioning device is adopted to shorten the time for sound source positioning through steps such as sound signal vector generation, subspace recognition, candidate vector recognition and direction recognition, and techniques such as delay and array method and stochastic gradient descent method.

Benefits of technology

It realizes high accuracy positioning of sound sources in a short period of time, ensures real-time performance, and is suitable for multi-speaker environments such as meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114616483B_ABST
    Figure CN114616483B_ABST
Patent Text Reader

Abstract

A sound source localization device (2) includes: a sound signal vector generation unit (21) that generates a sound signal vector based on a plurality of electrical signals output from a plurality of microphones (11) that receive sound generated by a sound source; a subspace identification unit (22) that identifies a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; a candidate identification unit (23) that identifies one or more candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and a direction identification unit (24) that identifies, as the direction of the sound source, the direction indicated by a sound source direction vector searched for using an initial solution based on at least one of the one or more candidate vectors, based on an optimization objective function including the sum of squares of the inner products of the signal subspace and the noise subspace.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a sound source localization device, a sound source localization method, and a program for identifying the position of a sound source. Background Art

[0002] Conventionally, methods for identifying the direction of a sound source have been studied. Patent Document 1 discloses a method for estimating the position of a sound source by minimizing an objective function of a variational inference method that estimates various parameters to minimize the difference between a posterior distribution representing the direction of the sound source and a variational function.

[0003] [Prior Art Documents]

[0004] [Patent Documents]

[0005] Patent Document 1: Japanese Patent No. 6623185 Summary of the Invention

[0006] [Problems to be Solved by the Invention]

[0007] When using the variational inference method as in the conventional method, the estimated value and the variables used to obtain the estimated value are random variables, so there are multiple unknown parameters. Since a large amount of calculation is required to estimate multiple variables, the conventional method using variational inference is not applicable to real-time localization of sound sources in a meeting.

[0008] The present disclosure focuses on this point, and an object of the present disclosure is to shorten the time required to localize a sound source.

[0009] Means for Solving the Problems

[0010] One aspect of the present disclosure provides a sound source localization device, including: a sound signal vector generation unit that generates a sound signal vector based on a plurality of electrical signals output from a plurality of microphones that receive sound generated by a sound source; a subspace identification unit that identifies a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; a candidate identification unit that identifies one or more candidate vectors indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and a direction identification unit that identifies, as the direction of the sound source, the direction indicated by a sound source direction vector searched using an initial solution based on at least one of the one or more candidate vectors, based on an optimization objective function including the sum of squares of inner products of the signal subspace and the noise subspace.

[0011] The candidate identification unit may identify, from among the one or more candidate vectors identified by applying the delay and sum array method to the sound signal vector, the initial solution for which the sum of the squares of the inner products with the signal subspace vectors corresponding to the signal subspace satisfies a predetermined reliability condition.

[0012] The candidate identification unit may perform the process of identifying the one or more candidate vectors in parallel with the process by which the subspace identification unit identifies the signal subspace and the noise subspace.

[0013] The sound signal vector generation unit may generate the sound signal vector by performing a Fourier transform on the plurality of electrical signals, and the direction identification unit may identify the direction of the sound source for each frame of the Fourier transform.

[0014] The direction identification unit may identify the direction of the sound source based on an average direction vector, which is obtained by averaging a plurality of sound source direction vectors corresponding to a plurality of frequency ranges generated using the Fourier transform.

[0015] The candidate identification unit may identify the one or more candidate vectors by performing interval rejection on the frequency ranges in such a way that the calculation of the one or more candidate vectors can be completed within one frame of the Fourier transform applied to the plurality of electrical signals.

[0016] The direction identification unit identifies the sound source direction vector by using the stochastic gradient descent method of the optimization objective function represented by the following formula,

[0017]

[0018] where, is the direction, is the virtual steering vector assuming that a target sound source exists in the θ and directions, t is the frame number, k is the frequency range number, and Q N (t,k) is the noise subspace vector.

[0019] The subspace identification unit may identify the signal subspace according to an orthogonality objective function, which is based on the difference between the sound signal vector and the vector obtained by projecting the sound signal vector onto the signal subspace.

[0020] The subspace identification unit identifies the signal subspace based on the orthogonality objective function represented by the following equation,

[0021]

[0022] where β is a forgetting function, t is the frame number, k is the frequency bin number, Q S (t,k) is the signal subspace vector, Q PS H (l - 1,k) is the estimated result of the signal subspace vector in the previous frame, and X is the sound signal.

[0023] A second aspect of the present invention provides a sound source localization method, including the following steps executed by a computer: generating a sound signal vector based on a plurality of electrical signals output by a plurality of microphones, the plurality of microphones receiving the sound generated by a sound source; identifying a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; identifying a plurality of candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and identifying, based on a first objective function including the sum of squares of the inner products of the signal subspace and the noise subspace, the direction indicated by a sound source direction vector selected from the directions indicated by the plurality of candidate vectors as the direction of the sound source.

[0024] A third aspect of the present invention provides a program for causing a computer to execute the following steps: generating a sound signal vector based on a plurality of electrical signals output by a plurality of microphones, the plurality of microphones receiving the sound generated by a sound source; identifying a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; identifying a plurality of candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and identifying, based on a first objective function including the sum of squares of the inner products of the signal subspace and the noise subspace, the direction indicated by a sound source direction vector selected from the directions indicated by the plurality of candidate vectors as the direction of the sound source.

[0025] Effects of the Invention

[0026] According to the present disclosure, the time required for localizing a sound source can be shortened. Description of the Drawings

[0027] Figure 1 is a diagram for showing an overview of a microphone system.

[0028] Figure 2 shows a design model of a microphone array.

[0029] Figure 3 shows a configuration of a sound source localization device.

[0030] Figure 4 is a flowchart of a process in which a sound source localization device executes a sound source localization method.

[0031] Figure 5 It is a flowchart of the processing of the direction recognition unit that recognizes the direction of the sound source. Detailed implementation manners

[0032] [Overview of the microphone system S]

[0033] Figure 1 It is a diagram for showing an overview of the microphone system S, and the microphone system S includes a microphone array 1, a sound source localization device 2, and a beamformer 3. The microphone system S is a system for collecting voices generated by multiple speakers H ( Figure 1 the speakers H-1 to H-4 among them) in a space such as a conference room or a hall.

[0034] The microphone array 1 has a plurality of microphones 11 represented by Figure 1 the black circles in, and they are installed on the ceiling, wall surface, or floor surface of the space where the speaker H stays. The microphone array 1 inputs a plurality of sound signals (for example, electrical signals) based on the voices input to the plurality of microphones 11 to the sound source localization device 2.

[0035] The sound source localization device 2 analyzes the sound signals input from the microphone array 1 to identify the direction of the sound source (i.e., the speaker H) that generates the voice. As will be described in detail later, the direction of the sound source is represented by the direction centered on the microphone array 1. For example, the sound source localization device 2 includes a processor, and the processor executes a program to identify the direction of the sound source.

[0036] The beamformer 3 performs beamforming processing by adjusting the weight factors of the plurality of sound signals corresponding to the plurality of microphones 11 based on the sound source direction identified by the sound source localization device 2. For example, the beamformer 3 makes the sensitivity to the voice generated by the speaker H greater than the sensitivity to the sounds from directions other than the direction where the speaker H is located. The sound source localization device 2 and the beamformer 3 can be implemented by the same processor.

[0037] Figure 1 It shows a state where the speaker H-2 is generating a voice. In Figure 1 the state shown, the sound source localization device 2 identifies that the voice is generated from the direction of the speaker H-2, and the beamformer 3 performs beamforming processing so that the main lobe of the directional characteristic of the microphone array 1 faces the speaker H-2.

[0038] When the microphone system S is used to separate the voices of speakers or identify speakers during a meeting, the sound source localization device 2 needs to change to another speaker or move along with the speaker and identify the direction of the speaker in the speech within a short period of time. Therefore, the sound source localization device 2 preferably completes the sound source localization process within one frame of the Fourier transform applied to the sound signal to ensure real-time performance. In addition, in order to separate the voices of a large number of speakers without error, it is required that the sound source localization device 2 can accurately identify the direction of the sound source.

[0039] Multiple Signal Classification (MUSIC), one of the sound source localization methods, is a high-resolution localization method based on the orthogonality between the signal subspace and the noise subspace. This method requires eigenvalue decomposition, and when the number of microphones 11 is assumed to be M, the computational order of MUSIC is O(M 3 ). Therefore, it is difficult to achieve real-time high-speed processing using MUSIC. In addition, if MUSIC is used to identify the direction of the sound source, even if there is a single sound source, MUSIC may identify a direction different from the correct direction as the sound source direction due to the influence of reflection, reverberation, aliasing, etc. Therefore, MUSIC is insufficient in terms of accuracy.

[0040] To solve such problems, the sound source localization device 2 according to the present embodiment uses Projection Approximation Subspace Tracking (PAST) to calculate the signal subspace without performing eigenvalue decomposition, thereby greatly reducing the computational amount. In this method, by using the Recursive Least Squares (RLS) method, the signal subspace is sequentially updated for each frame to which the Fourier transform is applied. Therefore, even if the speaker changes to another or the speaker moves, the sound source localization device 2 can calculate the signal subspace at high speed.

[0041] Furthermore, the sound source localization device 2 solves the minimization problem with the denominator term of the MUSIC spectrum as the objective function, reducing the computational amount from O(M 3 ) to O(M). Specifically, the sound source localization device 2 uses Nesterov Accelerated Adaptive Moment Estimation (Nadam), which is one of the Stochastic Gradient Descent methods. Nadam is a method that incorporates Nesterov's accelerated gradient method into Adam and improves the convergence speed to the solution by using the gradient information after one iteration. The sound source localization system 2 uses the direction searched by the Delay-Sum Array (DSA) method as the initial solution to reduce the number of Nadam iterations.

[0042] Specifically, the sound source localization device 2 obtains multiple initial solution candidates obtained by the delay and array method, and calculates the inner product of each solution among the multiple initial solution candidates and the signal subspace obtained by PAST. The sound source localization system 2 uses the initial solution candidate with the largest inner product among the multiple initial solution candidates as the initial solution of Nadam, so that it is possible to search for solutions within a range around the true direction of the sound source. The sound source localization device 2 calculates candidates for the direction as the initial solution in this way using the delay and array method, thereby reducing the number of iterations of the processing in Nadam and converging the minimization problem in a short time to identify the direction of the sound source.

[0043] [Sound Source Localization Method]

[0044] (Design Model)

[0045] Figure 2 The design model of the microphone array 1 is shown. In Figure 2 , it is assumed that the Y-shaped microphone array receives signals from the direction of the fixed sound source s(n). As Figure 2 shown, each microphone 11 is arranged at a position with a distance of d1 or d2 from the center point. The angle between the three directions in which the microphones 11 are arranged is 120 degrees. Here, when the distance between the sound source and the microphone array 1 is large enough, the sound signal can be regarded as a plane wave near the microphone array 1. In this case, the received sound signal X(t,k) can be expressed as a sound signal vector in the frequency domain by the following equation.

[0046] [Equation 1]

[0047] X(t, k) = S(t, k)a k (θ L , φ L ) + Γ(t, k) (1)

[0048] [Equation 2]

[0049]

[0050] [Equation 3]

[0051] Γ(t, k) = [Γ 0 (t, k), Γ 1 (t, k), …, Γ M-1 (t, k)] T (3)

[0052] In the above equations, t represents the frame number in the Fourier transform, k represents the frequency bin number, τ mIndicates the time difference of arrival of microphone m relative to a reference microphone (e.g., microphone 11-0), S(t,k) represents the frequency display of the sound source signal, Γ(t,k) represents the frequency display of the observed noise, and T represents the transpose. Sound source localization is the process of obtaining an estimated value of the sound source direction vector from the received sound signal X(t,k) in a certain frame t. of .

[0053] (Calculation of the signal subspace)

[0054] The sound source localization method uses MUSIC and PAST to calculate the signal subspace. MUSIC is a method for estimating the direction from which the sound signal comes. MUSIC is based on the orthogonality between (a) the signal subspace vector and (b) the noise subspace vector Q N (t,k), where the signal subspace vector is established using the eigenvectors calculated by the eigenvalue decomposition of the correlation matrix R(t,k) = E[X(t,k)X H (t,k)]. E[·] represents the calculation of the expected value, and H represents the Hermitian transpose. MUSIC uses the MUSIC spectrum represented by Equation (4) as the objective function.

[0055] [Equation 4]

[0056]

[0057] is the virtual steering vector assuming that the target sound source is in the direction. Due to the orthogonality of the signal subspace and the noise subspace, when , the denominator of Equation (4) becomes 0, and indicates the maximum value (peak).

[0058] It is necessary to calculate the maximum value of Equation (4) for each frame. Performing eigenvalue decomposition for each frame increases the computational load. Therefore, the sound source localization device 2 uses PAST to sequentially update Q s (t,k) for each frame without performing eigenvalue decomposition. That is, the sound source localization device 2 calculates Q s (t,k) while reducing the computational load, and calculates Q N (t,k) that minimizes the denominator of Equation (4). PAST is a process of obtaining Q s (t,k) that minimizes J(Q s (t,k)) in Equation (5).

[0059] [Equation 5]

[0060]

[0061] Equation (5) is an orthogonality objective function whose value decreases when the orthogonality between the signal subspace vector and the noise subspace vector is large. In Equation (5), β is a forgetting coefficient, and Q PS H (l - 1, k) is the estimated result of the signal subspace vector in the previous frame. s X(l, k) is the sound signal vector, and S (t, k) PS H (l - 1, k)X(l, k) is the vector obtained by projecting the sound signal vector onto the signal subspace. The sound source localization device 2 uses the s (t, k) estimated based on Equation (5) to calculate N (t, k) N H (t, k) = I - Q S (t, k)Q S H (t, k). In addition, the sound source localization device 2 calculates the MUSIC spectrum by applying the calculated value to Equation (4). Here, I represents the identity matrix.

[0062] By having the sound source localization device 2 calculate N (t, k) N H (t, k) using PAST, the order of magnitude of the calculations required to calculate the MUSIC spectrum is reduced from the conventional O(M 3 ) to O(2M). Therefore, the sound source localization device can significantly shorten the processing time for identifying the signal subspace vector.

[0063] After identifying the noise subspace vector, the sound source localization device 2 uses Nadam, which is one of the stochastic gradient descent methods. The optimization objective function of Nadam uses the following equation. In the following equation, is the denominator of Equation (4), and the solution that minimizes this denominator corresponds to the direction vector z e .

[0064] [Equation 6]

[0065]

[0066] [Equation 7]

[0067]

[0068] When searching for the When obtaining the minimum solution, the sound source localization device 2 uses the delay and array method to estimate the initial solution candidates, thereby reducing the number of search iterations. The spatial spectrum obtained by the delay and array method is represented by Equation (8).

[0069] [Equation 8]

[0070]

[0071] R DS (t,k) = E[X DS (t,k)X DS H (t,k)] is the correlation matrix used in the delay and array method, and is the steering vector. The sound source localization device 2 identifies the direction for which the value obtained by integrating as shown in Equation (9) below is equal to or greater than a predetermined value as an initial solution candidate.

[0072] [Equation 9]

[0073]

[0074] The initial solution candidates do not require high accuracy. Therefore, in order to reduce the load of calculating the initial solution candidates, the sound source localization device 2 can intermittently eliminate the frequency interval k and the direction so that the roughness of the frequency interval k and the direction is set to a level that allows the calculation of the initial solution candidates to be completed within one frame.

[0075] However, according to the calculation result of Equation (9), peaks may appear at positions far from the true peak direction . This occurs when the spatial spectrum is affected by aliasing, reflection, reverberation, etc. Therefore, the sound source localization device 2 can obtain R peaks as initial solution candidates during the peak search based on Equation (9), and use Equation (10) to calculate the reliability.

[0076] [Equation 10]

[0077]

[0078] Equation (10) is the sum of the squares of the inner products of the initial solution candidates and the signal subspace vectors obtained using PAST. The result obtained from Equation (10) indicates that the direction with a larger value is closer to the signal subspace vector established by Q s (t,k). As shown in Equation (11), the sound source localization device 2 identifies the peak with the highest reliability as the initial solution z'.

[0079] [Equation 11]

[0080]

[0081] After determining the initial solution as z’, the sound source localization device 2 assumes that z e = z’, and uses the Nadam method using Equation (6) to calculate the The sound source localization device 2 estimates the value obtained by averaging the corresponding to each of these frequency ranges as the new sound source direction vector z e . By obtaining the initial solution z’ using the above method before using the Nadam search for the solution, the sound source localization device 2 can search for a solution close to the signal subspace established by the sound source signal in a short time.

[0082] [Configuration of Sound Source Localization Device 2]

[0083] Figure 3 Fig. shows the configuration of the sound source localization device 2. The operations of each unit of the sound source localization method performed by the sound source localization device 2 will be described below with reference to Figure 3 The sound source localization device 2 includes a sound signal vector generation unit 21, a subspace identification unit 22, a candidate identification unit 23, and a direction identification unit 24. The candidate identification unit 23 includes a delay and array processing unit 231, a reliability calculation unit 232, and an initial solution identification unit 233. The sound source localization device 2 functions as the sound signal vector generation unit 21, the subspace identification unit 22, the candidate identification unit 23, and the direction identification unit 24 by a processor executing a program stored in a memory.

[0084] The sound signal vector generation unit 21 generates a sound signal vector. The sound signal vector is generated based on a plurality of electrical signals output from a plurality of microphones 11 that receive speech from a sound source. Specifically, the sound signal vector generation unit 21 generates a sound signal vector in the frequency domain by performing a Fourier transform (e.g., a fast Fourier transform) on the plurality of electrical signals input from the plurality of microphones 11. The sound signal vector generation unit 21 inputs the generated sound signal vector to the subspace identification unit 22 and the candidate identification unit 23.

[0085] The subspace identification unit 22 identifies (a) a signal subspace corresponding to signal components included in the sound signal vector and (b) a noise subspace corresponding to noise components included in the sound signal vector. The subspace identification unit 22 identifies the signal subspace vector and the noise subspace vector, for example, by using PAST. The signal subspace vector and the noise subspace vector are identified based on the orthogonality objective function shown in Equation (5), which is based on the difference between the sound signal vector and the vector obtained by projecting the sound signal vector onto the signal subspace.

[0086] The candidate identification unit 23 identifies one or more candidate vectors by applying a delay and array method to the sound signal vector. The one or more candidate vectors correspond to one or more directions assumed to be the sound source direction (i.e., the direction from which the sound signal comes). Then, the candidate identification unit 23 identifies, among the one or more identified candidate vectors, the candidate vectors for which the sum of the squares of the inner products with the signal subspace vector satisfies a predetermined reliability condition. The reliability condition is, for example, that the sum of the squares of the inner products of the candidate vector and the signal subspace vector is equal to or greater than a threshold value. The reliability condition is that (a) the probability distribution of the sound signal arriving from the predicted direction and (b) the likelihood of the sum of the squares of the inner products of the direction indicated by the candidate vector are relatively large. When the direction identification unit 24 performs the process of searching for the sound source direction, the identified candidate vectors are used as the initial solution. The candidate identification unit 23 may perform the operation of identifying one or more candidate vectors in parallel with the process performed by the subspace identification unit 22, or may perform the operation of identifying one or more candidate vectors after the subspace identification unit 22 performs the processes of identifying the signal subspace vector and the noise subspace vector.

[0087] To reduce the load for calculating the initial candidate solution, the candidate identification component 23 may determine the frequency range k and the direction so as to complete the calculation of one or more candidate vectors as the initial solution candidates within one frame of the Fourier transform applied to the sound signal. The candidate identification unit 23, for example, intermittently eliminates a plurality of frequency ranges generated by the Fourier transform of the sound signal to determine the frequency range k and the direction

[0088] The delay and array processing unit 231 estimates a plurality of candidate vectors indicating a plurality of possible directions from which the sound signal comes, based on the time differences of the sound signals emitted from the sound source arriving at the respective microphones 11, using a known delay and array method. Subsequently, the reliability calculation unit 232 calculates the reliability of each direction corresponding to the plurality of candidate vectors estimated by the delay and array processing unit 231 using Equation (10). The initial solution identification unit 233 inputs the candidate vector with the highest reliability calculated by the reliability calculation unit 232 to the direction identification unit 24 as the initial solution for the search process performed by the direction identification unit 24.

[0089] The direction recognition unit 24 recognizes the direction of the sound source based on the optimized objective function represented by Equation (6), which includes the sum of the squares of the inner products of the signal subspace vector and the noise subspace vector recognized by the subspace recognition unit 22. The direction recognition unit 24 recognizes the direction indicated by the sound source direction vector searched by using an initial solution based on at least any one of one or more candidate vectors recognized by the subspace recognition unit 22 as the direction of the sound source. Specifically, the direction recognition unit 24 uses the stochastic gradient descent method of the optimized objective function represented by Equation (6) to recognize the sound source direction vector.

[0090] The direction recognition unit 24 recognizes the direction of the sound source for each frame of the Fourier transform. Then, the direction recognition unit 24 recognizes the direction of the sound source based on the average direction vector. The average direction vector is obtained by averaging a plurality of sound source direction vectors corresponding to a plurality of frequency bands generated by the Fourier transform.

[0091] [Flowchart of the processing of the sound source localization device 2]

[0092] Figure 4 This is a flowchart of the processing in which the sound source localization device 2 executes the sound source localization method. When the sound signal vector generation unit 21 obtains the electrical signal corresponding to the sound signal X(t, k) from the microphone array 1 (step S1), the sound signal vector generation unit 21 initializes each variable (step S2). The sound signal vector generation unit 21 performs a fast Fourier transform on the sound signal X(t, k) (step S3) to generate a sound signal vector in the frequency domain formed by the frequency band k (k is a natural number) (step S4).

[0093] Subsequently, the subspace recognition unit 22 projects the sound signal vector onto the signal subspace to generate a projection vector (step S5). The subspace recognition unit 22 updates the eigenvalue based on Equation (5) (step S6), and updates the signal subspace vector Q s (t, k) (step S7). The subspace recognition unit 22 determines whether the processing from steps S5 to S7 has been executed a specified number of times (step S8). When the subspace recognition unit 22 determines that the processing has been executed a specified number of times, the subspace recognition unit 22 inputs the latest signal subspace vector to the direction recognition unit 24.

[0094] In parallel with the processing of steps S5 to S8, the candidate recognition unit 23 generates a correlation matrix R(t, k) = E[X(t, k)X H(t, k)] (step S9), and uses equation (9) to calculate the sum of each frequency interval (step S10). The candidate identification unit 23 identifies the vector indicating the direction in which the value obtained by calculation satisfies the predetermined condition (for example, the direction is equal to or greater than the threshold) as a candidate for the initial solution (step S11). In addition, the candidate identification unit 23 calculates the reliability of the identified initial solution candidate using equation (10) (step S12), and determines the initial solution candidate with the highest reliability as the initial solution (step S13).

[0095] The candidate identification section 23 notifies the determined initial solution to the direction identification section 24. The direction identification section 24 identifies the direction of the sound source by using the optimization objective function shown in equation (6) based on the signal subspace vector notified from the subspace identification section 22 and the initial solution notified from the candidate identification section 23 (step S14).

[0096] Figure 5 24 is a flowchart of the process (step S14) of the direction recognition unit 24 for recognizing the direction of the sound source. First, the direction recognition unit 24 calculates Q N (t,k)Q N H (t,k)=IQ S (t,k)Q S H (t, k) (step S141), and based on the calculation result, calculate the equation (6) Then, the direction identification unit 24 calculates the primary moment m for Nadam processing. i and the second moment n i (Step S143), and the adaptive learning rate is calculated by Nesterov's accelerated gradient method (Step S144). The direction identification unit 24 updates the solution of the direction vector based on the calculated adaptive learning rate (Step S145).

[0097] The direction identification section 24 repeats the processing of steps S142 to S145 until the processing is performed a prescribed number of times, and calculates the average value of the solutions of the direction vectors obtained for all frequency bins, thereby identifying the direction of the sound source (step S147 ).

[0098] [Results of real environment experiments]

[0099] To demonstrate the effectiveness of the sound source localization method according to this embodiment, a real environment experiment was conducted. Conference Room 1 and Conference Room 2 at the headquarters of Audio-Technica Corporation were used as the sound signal recording environments. Conference Room 1 has dimensions of 5.3 [m] * 4.7 [m] * 2.6 [m] and a reverberation time of 0.17 seconds. Conference Room 2 has dimensions of 12.9 [m] * 6.0 [m] * 4.0 [m] and a reverberation time of 0.80 seconds. The exhaust sound of a personal computer and the air conditioner sound exist as environmental noise.

[0100] A male voice was played as the sound source through speakers installed in each conference room, and at the same time, the voice was recorded using microphone array 1. Table 1 shows the true value of the sound source direction and the distance S between microphone array 1 and the speaker d d .

[0101] [Table 1]

[0102]

[0103] In this experiment, the number of microphones M = 7, d 1 = 15 [mm], d 2 = 43 [mm], the sampling frequency f s = 12 [kHz], the frame length K = 128, 50% overlap, the used frequency band is 2 [kHz] to 5 [kHz], the forgetting factor β of PAST = 0.96, and R = 2. In addition, the step size of Nadam is 0.1. A computer with an Intel Core (registered trademark) i7 - 7700HQ CPU (2.80 GHz) and 16 GB of RAM was used to measure the processing time. The mean absolute error shown in Equation (12) was used to evaluate the deviation value between the estimated direction of the sound source and the true value. is the true value direction of the sound source.

[0104] [Equation 12]

[0105]

[0106] As the sound source localization method, the sound source localization method according to this embodiment (hereinafter, this method is referred to as "this method"), Comparative Method 1, and Comparative Method 2 were used. Comparative Method 1 is the same as this method except that the reliability is not checked using Equation (10). Comparative Method 2 is a method of performing peak search on the MUSIC spectrum through eigenvalue decomposition. It should be noted that when calculating the evaluation value of Equation (12), the evaluation value is calculated except for the silent sections.

[0107] Table 2 shows the mean absolute error δ of the results measured using each method. It can be confirmed from Table 2 that when using this method and Comparative Method 2, the error relative to the true value is less than 5 [°]. On the other hand, the error of Comparative Method 1 for which reliability confirmation was not performed is larger than the error of this method. Since Comparative Method 2 directly performs peak search on the MUSIC spectrum, its error tends to be slightly smaller than that of this method.

[0108] [Table 2]

[0109]

[0110] In addition, the average calculation time per second of the signal length for each method was compared, RTF = S c / S 1 . Here, S c is the calculation time (seconds), and S 1 is the signal length (seconds). If the average calculation time is less than 1 (second), then real-time sound source localization can be performed.

[0111] Table 3 shows the average calculation times of the respective methods. In this method and Comparative Method 1, the average calculation time is much less than 1 (second), indicating that real-time performance can be ensured. On the other hand, in Comparative Method 2, the average calculation time is much higher than 1 (second), indicating that real-time performance cannot be ensured.

[0112] [Table 3]

[0113] Average calculation time [seconds] This method 0.21 Comparison method 1 0.20 Comparison method 2 5.20

[0114] Based on the above experimental results, it was found that the sound source localization method according to this embodiment can localize the sound source in real time while ensuring sufficient accuracy.

[0115] [Effect of the sound source localization device 2 according to this embodiment]

[0116] As described above, the sound source localization device 2 according to this embodiment calculates the signal subspace at high speed by using PAST to calculate the eigenvector for MUSIC without performing eigenvalue decomposition. The sound source localization device 2 identifies a candidate for the initial solution by using the delay and array method before calculating the optimal solution using Nadam with the denominator of the MUSIC spectrum as the objective function. The sound source localization device 2 determines the initial solution based on the reliability of the candidate for the initial solution identified by the delay and array method, thereby shortening the search time for the optimal solution. Through real environment experiments, it was confirmed that the sound source localization method performed by the sound source localization device 2 can ensure real-time performance and suppress the localization error to less than 5°.

[0117] It should be noted that in the above description, a fixed sound source is used to confirm the operation. However, the sound source localization method according to an embodiment of the present invention can be applied even when the sound source is moving. The sound source localization method according to an embodiment of the present invention can search for an optimal solution at high speed. Therefore, the sound source localization method according to this embodiment enables high-speed and high-accuracy sound source tracking. In addition, in the above description, Nadam is exemplified as a means for searching for an optimal solution, but Nadam is not the only means for searching for an optimal solution, and other means for solving the minimization problem can also be used.

[0118] The present invention is explained based on exemplary embodiments. The technical scope of the present invention is not limited to the scope explained in the above embodiments, and various changes and modifications can be made within the scope of the present invention. For example, all or part of the device can be configured using any unit that is functionally or physically dispersed or integrated. In addition, new exemplary embodiments generated by any combination thereof are also included in the exemplary embodiments of the present invention. In addition, the effects of the new exemplary embodiments brought about by the combination also have the effects of the original exemplary embodiments.

[0119] [Description of Symbols]

[0120] 1 Microphone array

[0121] 2 Sound source localization device

[0122] 3 Beamformer

[0123] 11 Microphone

[0124] 21 Sound signal vector generation unit

[0125] 22 Subspace identification unit

[0126] 23 Candidate identification unit

[0127] 231 Delay and array processing unit

[0128] 232 Reliability calculation unit

[0129] 233 Initial solution identification unit

[0130] 24 Direction identification unit

Claims

1. A sound source localization device, comprising: a sound signal vector generation unit that generates a sound signal vector based on a plurality of electrical signals output from a plurality of microphones, the plurality of microphones receiving sound generated by a sound source; a subspace identification unit that identifies a signal subspace corresponding to signal components included in the sound signal vector and a noise subspace corresponding to noise components included in the sound signal vector; a candidate identification unit that identifies one or more candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and a direction identification unit that identifies, as the direction of the sound source, the direction indicated by a sound source direction vector searched for using an initial solution based on at least one of the one or more candidate vectors, based on an optimization objective function including the sum of squares of inner products of the signal subspace and the noise subspace, wherein the candidate identification unit identifies, among the one or more candidate vectors identified by applying the delay and array method to the sound signal vector, the initial solution in which the sum of squares of inner products of a signal subspace vector corresponding to the signal subspace satisfies a predetermined reliability condition.

2. The sound source localization device according to claim 1, wherein in parallel with the process of the subspace identification unit identifying the signal subspace and the noise subspace, the candidate identification unit performs the process of identifying the one or more candidate vectors.

3. The sound source localization device according to claim 1, wherein the sound signal vector generation unit generates the sound signal vector by performing a Fourier transform on the plurality of electrical signals, and the direction identification unit identifies the direction of the sound source for each frame of the Fourier transform.

4. The sound source localization device according to claim 3, wherein the direction identification unit identifies the sound source direction vector by using a stochastic gradient descent method of the optimization objective function represented by the following formula, [Equation 13] Among them, is the direction, is the virtual steering vector assuming that there is a target sound source in the θ and directions, t is the frame number, k is the frequency bin number, and Q N (t, k) is the noise subspace vector.

5. The sound source localization device according to claim 3, wherein the subspace identification unit identifies the signal subspace according to an orthogonality objective function based on the difference between the sound signal vector and a vector obtained by projecting the sound signal vector onto the signal subspace.

6. The sound source localization device according to claim 5, wherein the subspace identification unit identifies the signal subspace based on the orthogonality objective function represented by the following equation, [Equation 14] where β is a forgetting function, t is the frame number, k is the frequency bin number, Q S (t,k) is the signal subspace vector, Q PS H (l - 1,k) is the estimated result of the signal subspace vector in the previous frame, and X is the sound signal.

7. A sound source localization device, comprising: a sound signal vector generation unit that generates a sound signal vector based on a plurality of electrical signals output from a plurality of microphones, the plurality of microphones receiving sound generated by a sound source, wherein the sound signal vector generation unit generates the sound signal vector by performing a Fourier transform on the plurality of electrical signals; a subspace identification unit that identifies a signal subspace corresponding to signal components included in the sound signal vector and a noise subspace corresponding to noise components included in the sound signal vector; A candidate identification unit that identifies one or more candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and A direction identification unit that identifies the direction indicated by the sound source direction vector searched using an initial solution based on at least one of the one or more candidate vectors as the direction of the sound source based on an optimization objective function including the sum of squares of the inner products of the signal subspace and the noise subspace, wherein, the direction identification unit identifies the direction of the sound source for each frame of the Fourier transform, and the direction identification unit identifies the direction of the sound source based on an average direction vector, which is obtained by averaging a plurality of the sound source direction vectors corresponding to a plurality of frequency intervals generated by the Fourier transform.

8. The sound source localization device according to claim 7, wherein, the candidate identification unit identifies the one or more candidate vectors by performing interval rejection on the frequency intervals in such a way that the calculation of the one or more candidate vectors can be completed within one frame of the Fourier transform applied to the plurality of electrical signals.

9. The sound source localization device according to claim 7, wherein, the direction identification unit identifies the sound source direction vector by using the stochastic gradient descent method of the optimization objective function represented by the following formula, [Equation 13] Among them, is the direction, is the virtual steering vector assuming the presence of a target sound source in the θ and directions. t is the frame number, k is the frequency bin number, and Q N (t, k) is the noise subspace vector.

10. The sound source localization device according to claim 7, wherein, the subspace identification unit identifies the signal subspace according to an orthogonality objective function based on the difference between the sound signal vector and the vector obtained by projecting the sound signal vector onto the signal subspace.

11. The sound source localization device according to claim 10, wherein, the subspace identification unit identifies the signal subspace based on the orthogonality objective function represented by the following equation, [Equation 14] where β is the forgetting function, t is the frame number, k is the frequency bin number, Q S (t,k) is the signal subspace vector, Q PS H (l - 1,k) is the estimated result of the signal subspace vector in the previous frame, and X is the sound signal.

12. A sound source localization method, including the following steps performed by a computer: Generating a sound signal vector based on a plurality of electrical signals output by a plurality of microphones that receive sound generated by a sound source; Identifying a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; Identifying a plurality of candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector, wherein, among the one or more candidate vectors identified by applying the delay and array method to the sound signal vector, identifying an initial solution in which the sum of squares of the inner products of the signal subspace vectors corresponding to the signal subspace satisfies a predetermined reliability condition; and Based on a first objective function including the sum of squares of the inner products of the signal subspace and the noise subspace, identifying the direction indicated by the sound source direction vector selected from the directions indicated by the plurality of candidate vectors as the direction of the sound source.

13. A sound source localization method, comprising the following steps performed by a computer: Generating a sound signal vector based on a plurality of electrical signals output by a plurality of microphones, the plurality of microphones receiving sound generated by a sound source, wherein, generating the sound signal vector by performing a Fourier transform on the plurality of electrical signals; Identifying a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; Identifying a plurality of candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and Based on a first objective function including the sum of squares of the inner products of the signal subspace and the noise subspace, identifying the direction indicated by the sound source direction vector selected from the directions indicated by the plurality of candidate vectors as the direction of the sound source, wherein, identifying the direction of the sound source for each frame of the Fourier transform, and identifying the direction of the sound source based on an average direction vector, the average direction vector being obtained by averaging a plurality of the sound source direction vectors corresponding to a plurality of frequency intervals generated by the Fourier transform.

14. A recording medium storing a program for causing a computer to perform the following steps: Generating a sound signal vector based on a plurality of electrical signals output by a plurality of microphones, the plurality of microphones receiving sound generated by a sound source; Identifying a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; Identifying a plurality of candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector, wherein, among one or more candidate vectors identified by applying the delay and array method to the sound signal vector, identifying an initial solution in which the sum of squares of the inner products of the signal subspace vectors corresponding to the signal subspace satisfies a predetermined reliability condition; and Based on a first objective function including the sum of squares of the inner products of the signal subspace and the noise subspace, identifying the direction indicated by the sound source direction vector selected from the directions indicated by the plurality of candidate vectors as the direction of the sound source.

15. A recording medium storing a program for causing a computer to perform the following steps: Generating a sound signal vector based on a plurality of electrical signals output by a plurality of microphones, the plurality of microphones receiving sound generated by a sound source, wherein, generating the sound signal vector by performing a Fourier transform on the plurality of electrical signals; Identifying a signal subspace corresponding to a signal component included in the sound signal vector and a noise subspace corresponding to a noise component included in the sound signal vector; Identifying a plurality of candidate vectors for indicating the direction of the sound source by applying a delay and array method to the sound signal vector; and Based on a first objective function that is the sum of squares of inner products including the signal subspace and the noise subspace, identify the direction indicated by the sound source direction vector selected from the directions indicated by the plurality of candidate vectors as the direction of the sound source, wherein, identify the direction of the sound source for each frame of the Fourier transform, and identify the direction of the sound source based on an average direction vector, which is obtained by averaging a plurality of the sound source direction vectors corresponding to a plurality of frequency intervals generated by the Fourier transform.