Sound Source Tracking Method, Device and Storage Medium for Distributed Microphone Array

The method enhances sound source tracking accuracy and reduces network load in distributed microphone arrays by selecting nodes based on signal energy and using distributed Kalman filtering for fusion estimation.

CN116386668BActive Publication Date: 2025-07-15WENZHOU ELECTRIC POWER BUREAU
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310299620.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2025-07-15
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

The existing distributed microphone array ignores the difference in signal-to-noise ratios between each node when tracking the sound source, resulting in poor estimation performance and using all nodes to process increases the network burden.

Method used

By obtaining the sound source signal received by the microphone node, dynamically selecting the target microphone node and its neighboring nodes, constructing a local node set, establishing a sound source kinematic model, using distributed volume Kalman filtering for local tracking, and obtaining a global estimate through local estimation fusion, and adjusting the fusion weight to reduce the network burden.

Benefits of technology

Improves the accuracy of sound source tracking while reducing the burden on the microphone network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386668B_ABST
    Figure CN116386668B_ABST
Patent Text Reader

Abstract

The present invention discloses a sound source tracking method, device and storage medium for a distributed microphone array. The method obtains sound source signals received by each microphone node in the distributed microphone array, and determines a target microphone node according to the sound source signals received by each of the microphone nodes; constructs a local node set according to the target microphone node and its neighboring microphone nodes; establishes a sound source kinematic model of the sound source signal; according to the sound source kinematic model, uses local nodes in the local node set to perform sound source tracking to obtain a local estimate of the sound source position by the corresponding local nodes; and fuses the local estimates of the sound source position by each of the local nodes to obtain a global estimate of the sound source position, so that the accuracy of sound source tracking can be improved, and at the same time, the burden on the entire microphone network can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech signal processing, and in particular, to a sound source tracking method, device and storage medium for a distributed microphone array. Background Art

[0002] The problem of sound source tracking has always been one of the research hotspots in the field of speech processing. It has been widely used in many aspects, such as audio and video conferencing systems, human-computer interaction and speech enhancement.

[0003] Traditional acoustic tracking methods usually require the microphone array to have a regular geometric structure and generally adopt a centralized data processing method. This method is usually unreliable, which will lead to an excessive communication load. At the same time, any failure of the central processor will cause the entire network to be unable to perform tracking. Therefore, a distributed microphone array came into being. The distributed microphone array has no strict restrictions on the arrangement of microphones and is a network composed of multiple nodes arbitrarily distributed in space. Usually, each node contains a set of microphones. At present, many distributed methods have been developed for sound source tracking. Instead of using a central processor for centralized data processing, all nodes perform local filtering of global state estimation through local data exchange between adjacent nodes. For example, distributed extended Kalman filter (DEKF), distributed unscented Kalman filter (DUKF), distributed cubature Kalman filter (DCKF) and various particle filters (PF), etc.

[0004] Since the signal energies received by each microphone node are different, the signal-to-noise ratios (SNRs) of each node are also different. Therefore, there are differences in the estimation performances of each node. However, the existing sound source tracking methods for distributed microphone arrays often ignore this difference when tracking the sound source, still use all the nodes in the array, and assign them the same weights, which makes the effect of the distributed microphone array on sound source tracking sometimes not very ideal. At the same time, using all the nodes will also cause a large burden on the entire microphone network. Summary of the Invention

[0005] Embodiments of the present invention provide a sound source tracking method, device and storage medium for a distributed microphone array, which improve the accuracy of sound source tracking and at the same time reduce the burden on the entire microphone network.

[0006] In a first aspect, embodiments of the present invention provide a sound source tracking method for a distributed microphone array, including:

[0007] Obtain the sound source signals received by each microphone node in the distributed microphone array, and determine a target microphone node according to the sound source signals received by each of the microphone nodes;

[0008] Construct a local node set according to the target microphone node and its neighboring microphone nodes;

[0009] Establish a kinematic model of the sound source for the sound source signal;

[0010] According to the kinematic model of the sound source, use the local nodes in the local node set to track the sound source, and obtain the local estimation of the sound source position by the corresponding local nodes;

[0011] Fuse the local estimations of the sound source position by each of the local nodes to obtain a global estimation of the sound source position.

[0012] As an improvement to the above solution, the determining of the target microphone node according to the sound source signals received by each of the microphone nodes includes:

[0013] Calculate the short-time energy of the sound source signals received by each of the microphone nodes;

[0014] Sort each of the microphone nodes in descending order according to the corresponding short-time energy, and select several microphone nodes with the top short-time energy as the target microphone nodes.

[0015] As an improvement to the above solution, the using the local nodes in the local node set to track the sound source according to the kinematic model of the sound source and obtaining the local estimation of the sound source position by the corresponding local nodes includes:

[0016] For each of the local nodes, use distributed cubature Kalman filtering to track the sound source signal according to the kinematic model of the sound source, and obtain the local estimation of the sound source position by the corresponding local node.

[0017] As an improvement to the above solution, the fusing the local estimations of the sound source position by each of the local nodes to obtain a global estimation of the sound source position includes:

[0018] Calculate the global position according to the local estimations of the sound source position by all the local nodes;

[0019] Calculate the root mean square error between the local estimations of the sound source position by each of the local nodes and the global position;

[0020] Calculate the fusion weight of each corresponding local node according to the root mean square error corresponding to each local node and the short-time energy of the sound source signal received by the corresponding local node;

[0021] Calculate the global estimation of the sound source position according to the local estimations of the sound source position by each of the local nodes and their fusion weights.

[0022] As an improvement to the above solution, establishing the kinematic model of the sound source for the sound source signal includes:

[0023] Establishing the kinematic model of the sound source for the sound source signal through the Langevin model;

[0024] Setting the state of the sound source signal at time k as Then there is:

[0025]

[0026] where, [x k , y k T and respectively represent the position and moving speed of the sound source signal; a = e -βΔT , β represents a preset rate constant, represents a preset steady-state speed parameter, I s represents the s-order identity matrix, represents the Kronecker inner product, ΔT represents the sampling period of position estimation, u k-1 represents zero-mean Gaussian white noise with a unit covariance matrix.

[0027] As an improvement to the above solution, each of the local nodes is configured with two microphones;

[0028] Then, for each of the local nodes, according to the kinematic model of the sound source, using distributed cubature Kalman filtering to track the sound source signal to obtain the local estimation of the sound source position corresponding to the local node, includes:

[0029] For each of the local nodes, using the generalized correlation function to calculate the time delay difference between the sound source signals received by the two microphones in the local node as the time delay difference observation;

[0030] Taking the sound source signal received by the local node as the signal observation;

[0031] Inputting the signal observations, time delay difference observations corresponding to each of the local nodes and the kinematic model of the sound source into the distributed cubature Kalman filtering for sound source tracking to obtain the local estimation of the sound source position corresponding to the local node.

[0032] As an improvement to the above solution, the method further includes:

[0033] Simulating and generating multiple room acoustic impulse responses under different signal-to-noise ratios and different reverberation times;

[0034] Convolving the room acoustic impulse response with the speech signal to obtain a convolved speech signal;

[0035] Add Gaussian white noise to the convolutional speech signal to generate the sound source signal.

[0036] As an improvement to the above solution, calculating the fusion weight of each local node according to the minimum root mean square difference corresponding to each local node and the short-time energy of the sound source signal received by the corresponding local node includes:

[0037] Calculate the quotient of the short-time energy of the sound source signal received by each local node and the minimum root mean square difference corresponding to the corresponding local node as the fusion influence parameter of the corresponding local node;

[0038] Calculate the sum of the fusion influence parameters of all local nodes, and calculate the quotient of the fusion influence parameter of each local node and the sum of the fusion influence parameters of all local nodes respectively to obtain the fusion weight of the corresponding local node.

[0039] In a second aspect, an embodiment of the present invention provides a sound source tracking device for a distributed microphone array, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the sound source tracking method for a distributed microphone array as described in any item of the first aspect.

[0040] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the sound source tracking method for a distributed microphone array as described in any item of the first aspect.

[0041] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows: By obtaining the sound source signals received by each microphone node in the distributed microphone array, and determining the target microphone node according to the sound source signals received by each microphone node; constructing a local node set according to the target microphone node and its neighboring microphone nodes; establishing a kinematic model of the sound source of the sound source signal; according to the kinematic model of the sound source, using the local nodes in the local node set to perform sound source tracking to obtain a local estimate of the sound source position by the corresponding local node; fusing the local estimates of the sound source positions by each local node to obtain a global estimate of the sound source position; the embodiments of the present invention can improve the accuracy of sound source tracking and reduce the burden on the entire microphone network at the same time. Description of the Drawings

[0042] To more clearly illustrate the technical solution of the present invention, the accompanying drawings to be used in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a flowchart of a sound source tracking method for a distributed microphone array provided by an embodiment of the present invention;

[0044] Figure 2 It is a time-domain diagram and a spectrogram of a sound source signal obtained by one of the nodes of the distributed microphone array provided by an embodiment of the present invention;

[0045] Figure 3 It is a deployment diagram of the distributed microphone array and a sound source trajectory provided by an embodiment of the present invention;

[0046] Figure 4(a) is the error of sound source tracking under different signal-to-noise ratios provided by an embodiment of the present invention;

[0047] Figure 4(b) is the error of sound source tracking under different reverberation times provided by an embodiment of the present invention;

[0048] Figure 5 It is a schematic diagram of a sound source tracking device for a distributed microphone array provided by an embodiment of the present invention. Specific embodiments

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0050] Embodiment 1

[0051] Please refer to Figure 1 , which is a flowchart of a sound source tracking method for a distributed microphone array provided by an embodiment of the present invention. The sound source tracking method for a distributed microphone array includes:

[0052] S1: Obtain the sound source signals received by each microphone node in the distributed microphone array, and determine the target microphone node according to the sound source signals received by each of the microphone nodes;

[0053] In the implementation of the present invention, room simulated reverberation is performed by the virtual sound source method to generate sound source signals. Specifically:

[0054] Simulate and generate multiple room acoustic impulse responses (RIRs) at different signal-to-noise ratios and different reverberation times;

[0055] Convolve the room acoustic impulse response with the speech signal to obtain a convolved speech signal;

[0056] Superimpose white Gaussian noise on the convolved speech signal to generate the source signal.

[0057] By convolving the room acoustic impulse response reflecting different reverberation times T 60 with the speech signal, and then adding white Gaussian noise with a determined mean and covariance, a received microphone signal (i.e., the source signal) mixed with reverberation and noise is generated. The time-domain graph and frequency-spectrum graph of the source signal are as Figure 2 shown.

[0058] S2: Construct a local node set according to the target microphone node and its neighboring microphone nodes;

[0059] S3: Establish a source kinematic model for the source signal;

[0060] In the embodiment of the present invention, the source kinematic model of the source signal is established through the Langevin model;

[0061] Set the state of the source signal at time k as Then there is:

[0062]

[0063] where, [x k , y k T and respectively represent the position and moving speed of the source signal; a = e -βΔT , β represents a preset rate constant, represents a preset steady-state speed parameter, I s represents the s-order identity matrix, represents the Kronecker inner product, ΔT represents the sampling period of position estimation, u k-1 represents zero-mean Gaussian white noise with a unit covariance matrix. It can be understood that x k , y k represent the position variables of the source signal at time k, represents the speeds of the source signal in the x and y directions at time k; [] T represents the transpose. By transposing the position and moving speed of the source signal, the state unity with the source can be achieved. ​

[0064] S4: According to the sound source kinematic model, use the local nodes in the local node set to perform sound source tracking, and obtain the local estimation of the sound source position by the corresponding local nodes;

[0065] Furthermore, for each of the local nodes, according to the sound source kinematic model, use distributed volume Kalman filtering to track the sound source signal, and obtain the local estimation of the sound source position by the corresponding local node.

[0066] S5: Fuse the local estimations of the sound source position by each of the local nodes to obtain the global estimation of the sound source position.

[0067] In an embodiment of the invention, based on the sound source signals received by each microphone node in a distributed microphone array, target microphone nodes are dynamically selected, and a local node set is formed based on the target microphone nodes and their neighboring microphone nodes for sound source tracking. Then, by fusing the local estimations of the sound source position by the local nodes, the global estimation of the sound source position is obtained as the final sound source position estimation, so as to accurately track the sound source and reduce the burden on the entire microphone network at the same time.

[0068] In an alternative embodiment, the determining the target microphone nodes according to the sound source signals received by each of the microphone nodes includes:

[0069] Calculate the short-time energy of the sound source signals received by each of the microphone nodes;

[0070] Exemplarily, for the sound source signal received by each of the microphone nodes, convert the sound source signal into a digital signal and perform frame-by-frame windowing processing; for example, in an embodiment of the present invention, select the Hamming window as the window function, and the short-time energy of the sound source signal received by the microphone node is:

[0071] E p,n = x p 2 (n)*h(n) (2);

[0072] where h(n) = ω 2 (n), w(n) represents the formula of the hamming window function, and x p (n) represents the nth frame after the sound source signal is intercepted by the window function; E p,n represents the short-time energy of the nth frame of the sound source signal received by microphone node p. The short-time energy can be regarded as the output of the square of the speech signal passing through a linear filter. At this time, h(n) represents the unit impulse response of the linear filter.

[0073] Sort each of the microphone nodes in descending order according to their respective short-time energies, and select several microphone nodes with the top short-time energies as the target microphone nodes.

[0074] In an alternative embodiment, each of the local nodes is configured with two microphones;

[0075] Then, for each of the local nodes, according to the sound source kinematic model, distributed cubature Kalman filtering is used to track the sound source signal, and a local estimate of the sound source position corresponding to the local node is obtained, including:

[0076] For each of the local nodes, the generalized correlation function is used to calculate the time delay difference between the sound source signals received by the two microphones within the local node as the time delay difference observation;

[0077] In an embodiment of the present invention, each microphone node of the distributed microphone array contains two microphones, and the time delay difference can be calculated through the generalized cross-correlation function. For example, assuming that S1(k) and S2(k) are the sound source signals received by the two microphones within a microphone node at time k, the Fourier transform of the sound source signal is performed to obtain S l (f) = FFT{S l (k)}, l = 1, 2; then the generalized cross-correlation function of the sound signals received by the microphones is:

[0078]

[0079] where S1(f) and S2(f) respectively represent the frequency-domain information of the sound signals received by the two microphones, and * represents the complex conjugate operation. Therefore, the time delay difference estimate is:

[0080]

[0081] τmax is the maximum time delay estimate.

[0082] Use the sound source signal received by the local node as the signal observation;

[0083] Input the signal observations, time delay difference observations, and the sound source kinematic model corresponding to each of the local nodes into the distributed cubature Kalman filter for sound source tracking, and obtain a local estimate of the sound source position corresponding to the local node.

[0084] In an embodiment of the present invention, the signal observations, time delay difference observations, and the sound source kinematic model corresponding to the selected local nodes are input into the distributed cubature Kalman filter for sound source tracking, Figure 3 shows the deployment of the distributed microphone array nodes and the set sound source trajectory.

[0085] In an alternative embodiment, the fusion of the local estimates of the sound source position by each of the local nodes to obtain a global estimate of the sound source position includes:

[0086] Calculate the global position based on the local estimates of the sound source position by all the local nodes;

[0087] In the embodiment of the present invention, by calculating the mean value of the local estimates of the sound source position by all the local nodes as the global position.

[0088] Exemplarily, define the local estimate of the sound source position by local node p at time k as x p,k (p = 1, 2, …, N), where N represents the number of local nodes in the local node set; [x k , y k T represents the sound source position estimated by local node p at time k, defined as r p,k , and define r N,k as the global position calculated by the distributed microphone array, then there is:

[0089]

[0090] Calculate the root mean square error between the local estimates of the sound source position by each of the local nodes and the global position;

[0091] Calculate the root mean square error between the local estimate r p, k of the sound source position by each local node and the above global position r N,k , and the specific calculation formula is:

[0092]

[0093] Calculate the fusion weight of each local node according to the root mean square error corresponding to each local node and the short-time energy of the sound source signal received by the corresponding local node;

[0094] Furthermore, by calculating the quotient of the short-time energy of the sound source signal received by each local node and the root mean square error corresponding to the corresponding local node as the fusion influence parameter of the corresponding local node;

[0095] Exemplarily, combined with the energy E p of local node p at time k, calculate the fusion influence parameter of local node p, and the specific calculation formula is:

[0096]

[0097] ​Calculate the sum of the fusion influence parameters of all the local nodes, and calculate the quotient of the fusion influence parameter of each local node and the sum of the fusion influence parameters of all the local nodes respectively to obtain the fusion weight of the corresponding local node. The specific calculation formula is as follows.

[0098]

[0099] η p represents the fusion weight of local node p during global fusion.

[0100] Calculate the global estimate of the sound source position based on the local estimates of the sound source position of each local node and its fusion weight.

[0101] Finally, according to η of each local node p Perform a globally consistent estimate on the local estimate obtained by it. The specific calculation formula is:

[0102]

[0103] In the embodiment of the present invention, dynamic selection of microphone nodes and adjustment of the weighting coefficient of distributed microphone array data fusion are adopted to achieve accurate tracking of the sound source, and at the same time, the burden on the entire microphone network can be reduced. As shown in FIGS. 4(a) and 4(b), simulation experiments are respectively carried out under different signal-to-noise ratios (10 dB - 30 dB) and different reverberation times (100 ms - 500 ms). The sound source tracking method of the embodiment of the present invention is used for sound source tracking, and the sound source tracking accuracy is evaluated according to RMSE (root mean square error), and the result is the average value of 100 Monte Carlo runs.

[0104] In the embodiment of the present invention, for a distributed microphone array, target microphone nodes and their neighboring microphone nodes are selected according to the signal energy magnitude of the received sound source signal, and distributed cubature Kalman filtering is used to track the sound source; for the local estimate of the sound source position calculated by the selected microphone nodes, it is proposed to adjust the weighting coefficient of distributed microphone array data fusion by combining the root mean square error (RMSE) of the local estimates of the sound source position of each microphone node and the received sound source signal energy, which is applicable to the sound source tracking neighborhood in the distributed microphone array, can achieve accurate tracking of the sound source, and reduces the burden on the entire microphone network.

[0105] Embodiment 2

[0106] See Figure 5, which is a schematic diagram of a sound source tracking device for a distributed microphone array provided by an embodiment of the present invention. The sound source tracking device for a distributed microphone array in this embodiment includes: a processor 100, a memory 200, and a computer program stored in the memory 200 and executable on the processor 100, such as a sound source tracking program for a distributed microphone array. When the processor 100 executes the computer program, it implements the steps in each of the above-mentioned embodiments of the sound source tracking method for a distributed microphone array, such as Figure 1 The steps S1 - S5 shown.

[0107] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the sound source tracking device for a distributed microphone array.

[0108] The sound source tracking device for a distributed microphone array can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The sound source tracking device for a distributed microphone array can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the sound source tracking device for a distributed microphone array, and does not constitute a limitation on the sound source tracking device for a distributed microphone array. It can include more or fewer components than those shown, or combine certain components, or different components. For example, the sound source tracking device for a distributed microphone array can also include input / output devices, network access devices, a bus, etc.

[0109] The processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor is the control center of the sound source tracking device for a distributed microphone array, and connects various parts of the entire sound source tracking device for a distributed microphone array through various interfaces and lines.

[0110] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by invoking the data stored in the memory, the processor implements various functions of the sound source tracking device for the distributed microphone array. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0111] Among them, if the modules / units integrated in the sound source tracking device for the distributed microphone array are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0112] Embodiment III

[0113] The embodiment of the present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the sound source tracking method for the distributed microphone array as described in any one of Embodiment I.

[0114] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.

[0115] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, many improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A sound source tracking method for a distributed microphone array, characterized in that, Including: Obtain the sound source signals received by each microphone node in the distributed microphone array, and determine a target microphone node according to the sound source signals received by each of the microphone nodes; Construct a local node set according to the target microphone node and its neighboring microphone nodes; Establish a sound source kinematic model of the sound source signal; According to the sound source kinematic model, use the local nodes in the local node set to perform sound source tracking to obtain a local estimate of the sound source position by the corresponding local nodes; Fuse the local estimates of the sound source position by each of the local nodes to obtain a global estimate of the sound source position; The fusing the local estimates of the sound source position by each of the local nodes to obtain a global estimate of the sound source position includes: Calculate a global position according to the local estimates of the sound source position by all the local nodes; Calculate the root mean square error of the local estimates of the sound source position by each of the local nodes and the global position; Calculate the fusion weight of each corresponding local node according to the root mean square error corresponding to each local node and the short-time energy of the sound source signal received by the corresponding local node; Calculate the global estimate of the sound source position according to the local estimates of the sound source position by each of the local nodes and their fusion weights.

2. The sound source tracking method for a distributed microphone array according to claim 1, characterized in that, The determining a target microphone node according to the sound source signals received by each of the microphone nodes includes: Calculate the short-time energy of the sound source signals received by each of the microphone nodes; Sort each of the microphone nodes in descending order according to the corresponding short-time energy, and select several microphone nodes with the short-time energy in the front row as the target microphone nodes.

3. The sound source tracking method for a distributed microphone array according to claim 1, characterized in that, The using the local nodes in the local node set to perform sound source tracking according to the sound source kinematic model to obtain a local estimate of the sound source position by the corresponding local nodes includes: For each of the local nodes, use distributed cubature Kalman filtering to track the sound source signal according to the sound source kinematic model to obtain a local estimate of the sound source position by the corresponding local node.

4. The sound source tracking method for a distributed microphone array according to claim 1, characterized in that, The establishing a sound source kinematic model of the sound source signal includes: Establish the sound source kinematic model of the sound source signal through the Langevin model; Set the state of the sound source signal at time k as , then we have: (1); wherein, and respectively represent the position and moving speed of the sound source signal; , , represent preset rate constants, represents a preset steady-state speed parameter, represents an s-order identity matrix, represents the Kronecker inner product, represents the sampling period of the position estimation, represents zero-mean Gaussian white noise with a unit covariance matrix.

5. The sound source tracking method for a distributed microphone array according to claim 3, wherein, Each of the local nodes is configured with two microphones; Then, the using the local nodes in the local node set to perform sound source tracking according to the sound source kinematic model to obtain a local estimate of the sound source position by the corresponding local nodes includes: For each of the local nodes, use the generalized correlation function to calculate the time delay difference between the sound source signals received by the two microphones in the local node as the time delay difference observation; Use the sound source signal received by the local node as the signal observation; Input the signal observations, time delay difference observations corresponding to each of the local nodes, and the sound source kinematic model into the distributed cubature Kalman filtering for sound source tracking to obtain a local estimate of the sound source position by the corresponding local nodes.

6. The sound source tracking method for a distributed microphone array according to claim 1, wherein, The method further includes: Simulate and generate multiple room acoustic impulse responses under different signal-to-noise ratios and different reverberation times; Convolve the room acoustic impulse response with the speech signal to obtain a convolved speech signal; A Gaussian white noise is superimposed on the convolutional speech signal to generate the sound source signal.

7. The sound source tracking method for a distributed microphone array according to claim 1, characterized in that, Calculating the fusion weight of each corresponding local node according to the minimum root mean square difference corresponding to each local node and the short-time energy of the sound source signal received by the corresponding local node includes: Calculating the quotient of the short-time energy of the sound source signal received by each local node and the minimum root mean square difference corresponding to the corresponding local node as the fusion influence parameter of the corresponding local node; Calculating the sum of the fusion influence parameters of all the local nodes, and respectively calculating the quotient of the fusion influence parameter of each local node and the sum of the fusion influence parameters of all the local nodes to obtain the fusion weight of the corresponding local node.

8. A sound source tracking device for a distributed microphone array, characterized in that, It includes: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the sound source tracking method for a distributed microphone array according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the sound source tracking method for a distributed microphone array according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fully distributed wireless sensor network robustness multi-sound-source positioning method

    CN104977562A