Method and system for separating and positioning rotating sound source and static sound source

By dividing the sound source plane into multiple scanning grids and using multiple beamforming algorithms, a joint propagation model is established, and the alternating direction multiplier method is used to solve the separation and positioning of the rotating sound source and the stationary sound source, solving the problem of underdetermined use restrictions and blind source separation in the prior art.

CN120065128APending Publication Date: 2025-05-30ZHEJIANG SHANGFENG SPECIAL BLOWER IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411935027.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, when separating and positioning stationary sound sources and rotating sound sources, the microphone array cannot use a ring array, making it difficult to collect sound source signals. The separation of blind sources is usually an underdetermined problem and requires the resolution of the statistical characteristics or prior knowledge of the signal.

Method used

By dividing the sound source plane into multiple scanning grids, the acoustic signals of each scanning grid are collected, and the stationary beamforming algorithm and the modal composition beamforming algorithm are used to calculate the stationary sound source output and the modal sound source output, establish a stationary sound source model and a modal sound source model, and obtain a joint propagation model, and use the alternating direction multiplier method to solve the minimum absolute shrinkage and selection operator model to obtain the energy estimate of each scanning grid.

Benefits of technology

Effective separation and positioning of rotating sound sources and stationary sound sources are realized, and the separation and positioning results of the sound sources are clearly displayed through visual energy estimation, which solves the problem of underdetermined microphone array usage restrictions and blind source separation in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065128A_ABST
    Figure CN120065128A_ABST
Patent Text Reader

Abstract

The invention relates to a rotating sound source and static sound source separating and positioning method and system.The rotating sound source and static sound source separating and positioning method comprises the steps that firstly, a sound source plane is divided into a plurality of scanning grids, and signals of each scanning grid are collected; then calculating static sound source output of the signal of each scanning grid and modal sound source output of the signal of each scanning grid through a static beam forming algorithm and a modal composition beam forming algorithm respectively, and then modeling the static sound source output and the modal sound source output respectively to obtain a static sound source model and a modal sound source model; a static sound source model and a modal sound source model are combined to obtain a joint propagation model, finally, an alternating direction multiplier method is adopted to solve a minimum absolute contraction and selection operator model, and energy estimation of each scanning grid is obtained. The sound source separation and positioning result can be clearly displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of signal processing, and particularly to a method and system for separating and positioning a rotating sound source and a stationary sound source. Background Art

[0002] The beamforming algorithm can accurately identify a sound source with a single motion pattern. For a stationary sound source, the stationary beamforming algorithm can achieve accurate positioning in a short time. For a rotating sound source, the algorithms used include the rotating source identifier algorithm in the time domain, the virtual rotating array algorithm in the frequency domain, and the modal component beamforming algorithm. The modal component beamforming algorithm in the frequency domain not only has a fast calculation speed but also has good resolution. In addition, there is a high-resolution positioning method based on modal component beamforming.

[0003] However, these methods are only applicable to identifying a single rotating sound source and are no longer applicable when facing a sound source with multiple motion patterns. Therefore, it is necessary to separate multiple sound sources.

[0004] The main sound source separation methods include: principal component analysis, independent component analysis, and sparse component analysis. These methods all rely on the statistical characteristic transformation of signals, perform linear transformation on signals, and extract key features in order to represent and analyze signals in a new space.

[0005] Blind source separation relies on a certain statistical independence between signal sources and is described using a linear mixing model, where the observed mixed signal is a linear combination of source signals. Since it is usually an underdetermined problem, it needs to rely on the statistical characteristics or prior knowledge of signals to solve. Moreover, when separating and positioning stationary and rotating sound sources in the prior art, a circular array cannot be used for the microphone array, resulting in difficulty in collecting sound source signals. Summary of the Invention

[0006] Based on this, in view of the fact that blind source separation is usually an underdetermined problem that needs to rely on the statistical characteristics or prior knowledge of signals to solve, and in the prior art, when separating and positioning stationary and rotating sound sources, a circular array cannot be used for the microphone array, resulting in difficulty in collecting sound source signals, it is necessary to provide a method and system for separating and positioning a rotating sound source and a stationary sound source.

[0007] On the one hand, the present application provides a method for separating and positioning a rotating sound source and a stationary sound source, including:

[0008] Dividing the sound source plane into multiple scanning grids, and collecting the sound signals of each scanning grid; the sound signal of each scanning grid is the sound source beam corresponding to that scanning grid;

[0009] Calculating the output of the stationary sound source of the sound signal of each scanning grid through a stationary beamforming algorithm;

[0010] The modal sound source output of the acoustic signal of each scanning grid is calculated by a modal component beamforming algorithm;

[0011] Modeling the stationary sound source output and the modal sound source output respectively to obtain a stationary sound source model and a modal sound source model, and combining the stationary sound source model and the modal sound source model to obtain a joint propagation model;

[0012] The joint propagation model is constrained by L1 norm, and the minimum absolute shrinkage and selection operator model is constructed;

[0013] The least absolute shrinkage and selection operator models are solved using the alternating direction multiplier method to obtain energy estimates for each scanned grid.

[0014] On the other hand, the present application also provides a system for separating and locating a rotating sound source and a stationary sound source, the system for separating and locating a rotating sound source and a stationary sound source comprising:

[0015] A collection component, the collection component includes a plurality of microphones, and the sound source signal of the sound source plane is collected by the collection component;

[0016] The processing component is communicatively connected with the collection component, the sound source signal of the sound source plane collected by the collection component is transmitted to the processing component, and the processing component performs the method for separating and locating the rotating sound source and the stationary sound source as described in any one of claims 1 to 9 on the sound source signal.

[0017] The present application relates to a method and system for separating and locating a rotating sound source and a stationary sound source. The method for separating and locating a rotating sound source and a stationary sound source first divides a sound source plane into a plurality of scanning grids, collects the signal of each scanning grid, and then respectively calculates the stationary sound source output of the signal of each scanning grid and the modal sound source output of the signal of each scanning grid by a stationary beamforming algorithm and a modal composition beamforming algorithm, and then respectively models the stationary sound source output and the modal sound source output to obtain a stationary sound source model and a modal sound source model, and the stationary sound source model and the modal sound source model are combined to obtain a joint propagation model, and finally the alternating direction multiplier method is used to solve the minimum absolute shrinkage and selection operator model to obtain the energy estimation of each scanning grid. The present application can clearly display the separation and positioning results of the sound source through the visualized energy estimation on each scanning grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic flow chart of a method for separating and locating a rotating sound source and a stationary sound source provided in one embodiment of the present application.

[0019] Figure 2A schematic diagram of the structure of a system for separating and locating a rotating sound source and a stationary sound source provided in one embodiment of the present application.

[0020] Reference numerals:

[0021] 100, acquisition component; 101, microphone; 200, processing component. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0023] The present application provides a method and system for separating and locating a rotating sound source and a stationary sound source.

[0024] like Figure 1 As shown, in one embodiment of the present application, a method for separating and locating a rotating sound source and a stationary sound source is provided, and the method for separating and locating a rotating sound source and a stationary sound source includes:

[0025] S100, dividing the sound source plane into a plurality of scanning grids, and collecting the sound signal of each scanning grid; the sound signal of each scanning grid is the sound source beam corresponding to the scanning grid;

[0026] S200, calculating the stationary sound source output of the sound signal of each scanning grid by a stationary beamforming algorithm;

[0027] S300, calculating the modal sound source output of the sound signal of each scanning grid by a modal composition beamforming algorithm;

[0028] S400, respectively modeling the stationary sound source output and the modal sound source output to obtain a stationary sound source model and a modal sound source model, and combining the stationary sound source model and the modal sound source model to obtain a joint propagation model;

[0029] S500, L1 norm constraints are applied to the joint propagation model to construct the minimum absolute shrinkage and selection operator model;

[0030] S600, the least absolute shrinkage and selection operator models are solved using an alternating direction multiplier method to obtain an energy estimate for each scanned grid.

[0031] In this embodiment, the sound source plane is first divided into multiple scanning grids, and the signal of each scanning grid is collected. Then, the stationary sound source output of the signal of each scanning grid and the modal sound source output of the signal of each scanning grid are calculated by the stationary beamforming algorithm and the modal composition beamforming algorithm respectively. Then, the stationary sound source output and the modal sound source output are modeled respectively to obtain the stationary sound source model and the modal sound source model, and the stationary sound source model and the modal sound source model are combined to obtain the joint propagation model. Finally, the alternating direction multiplier method is used to solve the minimum absolute shrinkage and selection operator model to obtain the energy estimation of each scanning grid. The present application can clearly display the separation and positioning results of the sound source through the visualized energy estimation on each scanning grid.

[0032] In one embodiment of the present application, the S100 includes:

[0033] S110, creating a signal database;

[0034] S120, setting the side length of a single scanning grid; the scanning grid is a square area;

[0035] Specifically, when collecting sound source signals, the sound source plane can be regarded as a two-dimensional plane. The sound source plane includes stationary sound sources and rotating sound sources. The signal of the stationary sound source collected by the collection component is the stationary sound source signal, and the signal of the rotating sound source collected by the collection component is the rotating sound source signal.

[0036] The reason why the sound source plane is divided into multiple scanning grids is that the beamforming output result on each scanning grid represents the calculated energy estimate at the grid point. The scanning grid often overlaps with the equivalent source grid point to facilitate calculation and analysis.

[0037] S130, dividing the sound source plane into a plurality of scanning grids with equal side lengths according to the side length of a single scanning grid, and treating each scanning grid as an equivalent sound source;

[0038] Specifically, the collection distance is the shortest distance from the microphone array to the sound source plane.

[0039] The acquisition angle in this application is the polar angle. In the polar coordinate system, the angle between the line connecting any point on the plane to the pole and the polar axis is called the polar angle.

[0040] In the actual process of collecting sound source signals, it is also necessary to set the collection frequency and collection time, so as to collect multiple sound source signals within the collection time.

[0041] S140, setting a stationary sound source collection distance, a stationary sound source collection angle, a rotating sound source collection distance, and a rotating sound source collection angle;

[0042] S150. At a stationary sound source acquisition distance from the sound source plane, collect multiple stationary sound source signals according to multiple stationary sound source acquisition angles.

[0043] S160. At a rotating sound source acquisition distance from the sound source plane, collect multiple rotating sound source signals according to multiple rotating sound source acquisition angles.

[0044] S170. Put all the collected stationary sound source signals and rotating sound source signals into the signal database.

[0045] In this embodiment, the process of dividing the sound source plane into equivalent sources first involves defining the two-dimensional area where the sound source is located, and then creating uniform discrete grid points in this area, which is the scanning grid. Each grid point represents a potential equivalent sound source.

[0046] Collecting parameters are collectively referred to as the stationary sound source acquisition distance, stationary sound source acquisition angle, rotating sound source acquisition distance, rotating sound source acquisition angle, acquisition duration, and acquisition frequency.

[0047] Since multiple stationary sound sources and rotating sound sources are set when collecting sound source signals, it is necessary to set the collection parameters for each stationary sound source and rotating sound source respectively to ensure that the collected sound source signals are clear and complete enough.

[0048] For example, there are 3 stationary sound sources and 3 rotating sound sources, and the set acquisition duration is 10 seconds for all, and the acquisition frequency is 50000 for all.

[0049] Before collecting the sound source signals through the microphone array, it is first necessary to divide each sound source plane into multiple uniform scanning grids. If the sound source plane is a square plane of 0.9 meters * 0.9 meters, then the side length of the scanning grid can be set to 0.05 meters * 0.05 meters, so as to divide one sound source plane into 1296 scanning grids.

[0050] After that, set the collection parameters for each sound source plane respectively. For the 3 stationary sound sources, set the acquisition distance of the microphone array to 0.8 meters, and the polar angles of the acquisition are 0 * π, (2 / 3) * π, (3 / 4) * π respectively. For the 3 rotating sound sources, set the acquisition distance of the microphone array to 1.0 meters, and the polar angles of the acquisition are (1 / 2) * π, (7 / 6) * π, (11 / 6) * π respectively. In this way, multiple stationary sound source signals and multiple rotating sound source signals can be obtained respectively. Finally, all the collected sound source signals are put into the signal database for subsequent processing.

[0051] In an embodiment of the present application, the S200 includes:

[0052] S210. Calculate the sound pressure distribution of the sound field on the sound source plane based on each stationary sound source signal through Equation 1;

[0053]

[0054] where, is the sound pressure distribution of the sound field on the sound source plane, is the sound intensity of the sound source at a certain scanning grid, is the position of the scanning grid, t' represents the moment when the sound source occurs, V is all the equivalent sound sources on the sound source plane, is the Green's function in the free sound field;

[0055] Specifically, since each scanning grid represents an equivalent sound source, the sound pressure distribution of the sound field of the stationary sound source signal can be expressed as the superposition of the sound pressure contributions from all the equivalent sources to the field point.

[0056] Since there is a certain distance for the sound signal to be transmitted from the sound source plane to the acquisition component, that is, the acquisition distance set in the acquisition parameters, there is a time delay in the process of acquiring the sound source signal, and the time delay = acquisition distance / sound speed.

[0057] S220. Calculate the sound pressure distribution of each stationary sound source signal in the signal database through Equation 2;

[0058]

[0059] where, is the sound pressure distribution of the sound field on the sound source plane, is the sound source distribution function, is the sound field space coordinate vector, t is the time variable, c is the sound speed constant, is the divergence operator;

[0060] Specifically, in the stationary free sound field, the sound pressure distribution generated by the sound source satisfies the inhomogeneous acoustic wave equation, that is, Equation 2.

[0061] S230. Introduce the medium motion to optimize Equation 2 into Equation 3;

[0062]

[0063] where, is the sound pressure distribution of the sound field on the sound source plane, is the sound source distribution function, is the sound field space coordinate vector, t is the time variable, c is the sound speed constant, is the divergence operator;

[0064] Specifically, due to the existence of the flowing medium, the partial derivative with respect to time is replaced by the convective derivative

[0065] Among them, is a uniform and smooth flow velocity vector, is the divergence operator.

[0066] S240, based on Equation 3, optimize Equation 1 to obtain Equation 4;

[0067]

[0068] Among them, is the sound pressure distribution of the sound field on the sound source plane, is the sound intensity of the sound source at a certain scanning grid, is the position of the scanning grid, t' represents the moment when the sound source occurs, V is all equivalent sound sources on the sound source plane, is the Green's function in the flowing sound field of the medium motion.

[0069] Specifically, after dividing the sound source plane into a series of scanning grids, the sound pressure of the sound field can be expressed as the superposition of the sound pressure contributions from all equivalent sources to the field point, and its superposition form is Equation 4.

[0070] In this embodiment, analyze the characteristics of the actual sound source and select a suitable sound source model, such as a stationary sound source or a rotating sound source, and calculate the contribution of each scanning grid to the overall sound field. Through sound field reconstruction, optimize the parameters of the equivalent source to match the actual sound source, and then use experimental or simulation data to verify the accuracy of the model. Finally, apply the adjusted equivalent source model for sound source localization and acoustic analysis, and display the analysis results through visualization means for further application and decision-making.

[0071] In an embodiment of the present application, the S200 further includes:

[0072] S250, randomly select a stationary sound source signal from the signal database;

[0073] S260, calculate the stationary beam corresponding to the stationary sound source signal through Equation 5 for this signal;

[0074] P1(f) = G1(f) * S(f) + E(f) Equation 5;

[0075] Among them, P1(f) is the representation of the stationary sound source signal in the frequency domain, G1(f) is the Green's function in the free sound field, S(f) is the frequency representation of the equivalent sound source corresponding to the scanning grid, f is the frequency of interest, and E(f) is the noise;

[0076] Specifically, the Green's function in the free field can be expressed as an equation;

[0077]

[0078] where x m is the position coordinate of the m-th microphone, x n is the position coordinate of the n-th scanning grid, j is the imaginary unit, and c is the sound source;

[0079] S270. Add a weight vector to the stationary sound source signal through Formula 6 to obtain the beamforming output of the stationary sound source signal;

[0080]

[0081] where Y n (f) is the beamforming output after summing the stationary sound source signal with different weights, W m,n is the weight from the m-th microphone to the n-th scanning grid, P m (f) is the frequency-domain representation of the signal collected by the m-th microphone, is the vector representation of W m,n , is the vector representation of P m (f), and (·) H represents conjugate transpose;

[0082] Specifically, adding the weight vector is to minimize the error between the sound source signal after weighting and the original signal.

[0083] S280. Square the beamforming output through Formula 7 to obtain the stationary sound source output of the signal;

[0084]

[0085] where Y c is the stationary sound source output of the stationary sound source signal, Y n (f) is the beamforming output after summing the stationary sound source signal with different weights, R(f) is the cross-spectral matrix, is the weight vector, and H is the symbol of conjugate transpose;

[0086] Specifically, represents the weight for the cross-spectral matrix and is a parameter vector to be solved.

[0087] S290. Return to S250 to obtain the stationary sound source output corresponding to each stationary sound source signal.

[0088] In this embodiment, the stationary beamforming algorithm is used to calculate each stationary sound source signal collected by the microphone respectively to obtain the beamforming output on each scanning grid, which is used to estimate the energy output of the stationary sound source at this grid point.

[0089] In one embodiment of the present application, the S300 includes:

[0090] S310, randomly select a rotating sound source signal from the signal database;

[0091] S320, calculate the modal composition beam corresponding to the rotating sound source signal through Formula 8 for the rotating sound source signal;

[0092] P2(f) = G2(f)*S(f) + E(f) Formula 8;

[0093] where P2(f) is the representation of the rotating sound source signal in the frequency domain, G2(f) is the Green's function in the flowing sound field of the medium motion, S(f) is the frequency representation of the equivalent sound source corresponding to the scanning grid, f is the frequency of interest, and E(f) is the noise;

[0094] S330, calculate the Green's function from the m-th microphone to the n-th grid point through Formula 9;

[0095]

[0096] where G2 m,n (f,x m ,x n ) is the Green's function from the m-th microphone to the n-th grid point, x m is the position coordinate of the m-th microphone, x n is the position coordinate of the n-th grid, m 0 is different modes, and is the transfer function in a certain mode;

[0097] Specifically, the mode represents different vibration modes or propagation modes that may exist in the sound field or vibration system. S340, calculate the output of the modal composition beam through Formula 10;

[0098]

[0099] where b n (f) is the output of the modal composition beam corresponding to the rotating sound source signal, M is the total number of microphones, P m (f) is the frequency domain representation of the signal collected by the m-th microphone, G2 m,n (f,x m ,x n ) is the Green's function from the m-th microphone to the n-th grid point;

[0100] S350, square the output of the modal composition beam corresponding to the rotating sound source signal through the formula to obtain the modal sound source output of the modal composition beam corresponding to the rotating sound source signal, and the modal sound source output is as shown in Formula 11;

[0101]

[0102] Among them, Y r is the modal sound source output of the beam formed by the mode corresponding to the rotating sound source signal, and b n (f) is the output of the beam formed by the mode corresponding to the rotating sound source signal, M is the total number of microphones, and P m (f) is the frequency-domain representation of the signal collected by the m-th microphone, and G2 m,n (f, x m , x n ) is the Green's function from the m-th microphone to the n-th grid point;

[0103] S360, randomly select a rotating sound source signal from the signal database until all the rotating sound source signals in the signal database are selected, and obtain the modal sound source output of each rotating sound source signal.

[0104] In this embodiment, for a rotating sound source, due to the Doppler effect generated during its movement, the frequency shifts within a certain range. Therefore, the Green's function of the rotating sound source is different from that of the stationary sound source.

[0105] Since the sound source plane includes a stationary sound source and a rotating sound source, the collected sound source signal is actually a superposition of the signals propagated by two different motion-mode sound sources.

[0106] Due to the signal superposition, the energy output on each scanning grid is actually the beamforming output obtained by using the stationary beamforming algorithm for the mixed signal from the signal emitted by the stationary sound source and the signal emitted by the rotating sound source.

[0107] Using the modal component beamforming algorithm, the beamforming output on each scanning grid is calculated respectively, which is used to estimate the energy output of the stationary sound source at this grid point. In fact, it is the beamforming output obtained by using the stationary beamforming algorithm for the mixed signal from the signal emitted by the stationary sound source and the signal emitted by the rotating sound source.

[0108] In an embodiment of the present application, the S400 includes:

[0109] S410, define the stationary sound source output of the sound signal of each scanning grid through Formula 12;

[0110] Y c = A sr * X r + A ss * X s Formula 12;

[0111] Among them, Y c is the stationary sound source output of the acoustic signal of the scanning grid, X s is the stationary sound source signal of the acoustic signal of the scanning grid, X r is the rotating sound source signal of the acoustic signal of the scanning grid, A ss is the point spread function of the stationary sound source signal of the scanning grid using stationary beamforming, A sr is the point spread function of the rotating sound source signal of the scanning grid using stationary beamforming;

[0112] S420, define the modal sound source output of the acoustic signal of all scanning grids through Formula 13;

[0113] Y r = A rr * X r + A rs * X s Formula 13;

[0114] Among them, Y r is the modal sound source output of all scanning grids, X s is the stationary sound source signal of all scanning grids, X r is the rotating sound source signal of all scanning grids, A rs is the point spread function of the stationary sound source signal of all scanning grids using modal component beamforming, A rr is the point spread function of the rotating sound source signal of all scanning grids using modal component beamforming.

[0115] In this embodiment, the beamforming output of the beam microphone signal obtained by using the traditional beamforming algorithm can be expressed as Formula 12, and the beamforming output of the beam microphone signal obtained by using the modal component beamforming algorithm can be expressed as Formula 13.

[0116] In an embodiment of the present application, the S400 further includes:

[0117] S430, construct a complete cross-point spread function through Formula 14;

[0118]

[0119] Among them, A is the complete cross-point spread function, A rs is the point spread function of the stationary sound source signal of all scanning grids using modal component beamforming, A rr is the point spread function of the rotating sound source signal of all scanning grids using modal component beamforming, A ss is the point spread function of the stationary sound source signal of all scanning grids using stationary beamforming, A sris the point spread function of the rotating sound source signal of all scanned grids using stationary beamforming;

[0120] Specifically, for the stationary sound source signal, the point spread function A of stationary beamforming ss The element a in the i-th row and j-th column of ij , can be solved by Equation 21;

[0121]

[0122] where a ij is the element in the i-th row and j-th column of the point spread function A of the stationary sound source signal using stationary beamforming ss , that is, the point spread function from the i-th stationary equivalent source to the j-th scanned grid, is the number of microphones, G1 is the Green's function, is the cross-spectral matrix of the signals from the i-th stationary equivalent source to the j-th scanned grid;

[0123] Similarly, the other point spread functions in the complete cross-point spread function can also be solved by equations.

[0124] S440, based on the complete cross-point spread function, construct a joint propagation model through Equation 15;

[0125]

[0126] where Y r is the modal sound source output of all scanned grids, Y c is the stationary sound source output of all scanned grids, X s is the stationary sound source signal of all scanned grids, X r is the rotating sound source signal of all scanned grids, A rs is the point spread function of the stationary sound source signal of all scanned grids using modal component beamforming, A rr is the point spread function of the rotating sound source signal of all scanned grids using modal component beamforming, A ss is the point spread function of the stationary sound source signal of all scanned grids using stationary beamforming, A sr is the point spread function of the rotating sound source signal of all scanned grids using stationary beamforming;

[0127] S450, simplify Equation 15 to Equation 16;

[0128] Y = A * X equation;

[0129] where Y is the joint matrix, Y=(Y r ,Y c ) T, X is the sound source signal of all scanned grids, X is the combined sound source and X = (X r , X S ) T , (·) T represents matrix transpose.

[0130] In this embodiment, by calculating the stationary beamforming output and the modal component beamforming output from the signals collected by the microphone, Y r and Y c can be obtained. According to the complete cross-point diffusion function and the obtained beamforming, the proposed rotating stationary sound source energy propagation model, i.e., the combined propagation model, can be obtained.

[0131] In an embodiment of the present application, the S500 includes:

[0132] S510, respectively reconstruct the rotating sound source signal and the stationary sound source signal to obtain a rotating sound source column vector and a stationary sound source column vector;

[0133] S520, combine the rotating sound source column vector and the stationary sound source column vector to obtain a combined sound source;

[0134] S530, apply the L1 norm constraint to the combined sound source through Formula 17 and assign a suitable regularization parameter to it to form a regularization term, so that the energy on the scanned grid is distributed in a coefficient manner, thereby constructing a least absolute shrinkage and selection operator model;

[0135]

[0136] where A is the complete cross-point diffusion function, X is the combined sound source, and Y is the combined matrix;

[0137] Specifically, the regularization parameter is a strategy for preventing model overfitting. By adding a regularization term to the loss function, the complexity of the model is reduced. The role of the regularization parameter is to control the model complexity and prevent overfitting. By adjusting the size of the regularization parameter, the fitting ability and generalization ability of the model can be balanced.

[0138] The L1 norm (L1 norm) refers to the sum of the absolute values of the elements in the vector. It is also called the "sparse rule operator" (Lasso regularization). The L1 norm is a representation method for measuring the sparsity of a vector and is commonly used in machine learning. By adding the L1 norm to the cost function, the obtained result of learning satisfies sparsity, which is convenient for people to extract features.

[0139] S540, standardize Formula 17 through Formula 18 to obtain a standardized least absolute shrinkage and selection operator model;

[0140]

[0141] Among them, is the standardized least absolute shrinkage and selection operator model, is the standardized complete cross-point diffusion function, Y is the joint matrix, is the standardized joint matrix.

[0142] In this embodiment, after standardizing the joint matrix Y and the complete cross-point diffusion function A, the standardized joint matrix and the standardized cross-point diffusion function are obtained, and then the model is solved to improve the stability of the algorithm.

[0143] By imposing an L1 norm constraint on the energy values on each scanning grid, a least absolute shrinkage and selection operator model for solving the system of equations is constructed.

[0144] In an embodiment of the present application, the S600 includes:

[0145] S610, introducing an auxiliary variable into the least absolute shrinkage and selection operator model, and the formula for introducing the auxiliary variable is as shown in Formula 19;

[0146]

[0147] Among them, z is the auxiliary variable, A is the complete cross-point diffusion function, X is the joint sound source, and Y is the joint matrix;

[0148] S620, constructing a Lagrangian function by Formula 20 and adding an augmented Lagrangian term;

[0149]

[0150] Among them, u is the Lagrange multiplier, ρ is the penalty parameter and ρ > 0, z is the auxiliary variable, A is the complete cross-point diffusion function, X is the joint sound source, and Y is the joint matrix;

[0151] S630, fixing the Lagrange multiplier, the auxiliary variable and the joint sound source respectively to solve the other two parameters in Formula 20;

[0152] Specifically, when solving the Lagrange multiplier, the auxiliary variable and the joint sound source, soft threshold operation can be used for solving.

[0153] The soft threshold operation is such that for each element in the input signal, if its amplitude value exceeds a preset threshold, the element remains unchanged or is slightly adjusted according to certain rules; if its amplitude value is lower than the threshold, it is adjusted to zero or reduced to a value close to zero in a certain way. This operation helps to remove the small-amplitude components in the signal while retaining the large-amplitude components, thus achieving the purpose of signal processing.

[0154] S640, set the error threshold;

[0155] S650, respectively determine whether both the auxiliary variable and the joint sound source X are less than or equal to the error threshold;

[0156] S660, if both the auxiliary variable and the joint sound source X are less than or equal to the error threshold, then exit the update;

[0157] S670, if at least one of the auxiliary variable and the joint sound source X is greater than the error threshold, then return to separately fix the Lagrange multiplier, the auxiliary variable, and the joint sound source X until both the auxiliary variable and the joint sound source X are less than or equal to the error threshold.

[0158] In this embodiment, the alternating direction multiplier method is used to solve the constructed least absolute shrinkage and selection operator model to obtain the estimated energy on each scanning grid, and the estimated energy on each scanning grid is visualized to obtain the result of sound source separation and localization.

[0159] As Figure 2 shown, in an embodiment of the present application, a system for separating and localizing a rotating sound source and a stationary sound source is provided. The system for separating and localizing a rotating sound source and a stationary sound source includes an acquisition component 100 and a processing component 200.

[0160] The acquisition component 100 includes a plurality of microphones 101, and the sound source signal on the sound source plane is acquired through the acquisition component 100. The acquisition component 100 is communicatively connected to the processing component 200, and the sound source signal on the sound source plane acquired by the acquisition component 100 is transmitted into the processing component 200, and the processing component 200 performs the method for separating and localizing a rotating sound source and a stationary sound source as described in the foregoing embodiment on the sound source signal.

[0161] In this embodiment, a microphone array is formed by a plurality of microphones 101, and acquisition parameters are set for the microphone array, so as to acquire the sound source signal on the sound source plane through the microphone array. The sound source signal acquired by the microphone array is transmitted into the processing component 200, and the processing component 200 performs the method for separating and localizing a rotating sound source and a stationary sound source as described in the foregoing embodiment on the sound source signal.

[0162] The technical features of the above-described embodiments can be combined arbitrarily, and there is no limitation on the execution order of the method steps. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0163] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for separating and locating a rotating sound source and a stationary sound source, characterized in that: The method for separating and locating the rotating sound source and the stationary sound source comprises: The sound source plane is divided into a plurality of scanning grids, and the sound signal of each scanning grid is collected; the sound signal of each scanning grid is the sound source beam corresponding to the scanning grid; The stationary sound source output of the acoustic signal of each scanning grid is calculated by a stationary beamforming algorithm; The modal sound source output of the acoustic signal of each scanning grid is calculated by a modal component beamforming algorithm; Modeling the stationary sound source output and the modal sound source output respectively to obtain a stationary sound source model and a modal sound source model, and combining the stationary sound source model and the modal sound source model to obtain a joint propagation model; The joint propagation model is constrained by L1 norm, and the minimum absolute shrinkage and selection operator model is constructed; The least absolute shrinkage and selection operator models are solved using the alternating direction multiplier method to obtain energy estimates for each scanned grid.

2. The method for separating and locating a rotating sound source and a stationary sound source according to claim 1, characterized in that: The step of dividing the sound source plane into a plurality of scanning grids and collecting the sound signal of each scanning grid comprises: Create a signal database; Set the side length of a single scanning grid; the scanning grid is a square area; The sound source plane is divided into a plurality of scanning grids with equal side lengths according to the side length of a single scanning grid, and each scanning grid is regarded as an equivalent sound source; Set the static sound source collection distance, static sound source collection angle, rotating sound source collection distance and rotating sound source collection angle; At a stationary sound source collection distance from the sound source plane, multiple stationary sound source signals are collected according to multiple stationary sound source collection angles; At a rotating sound source collection distance from the sound source plane, a plurality of rotating sound source signals are collected according to a plurality of rotating sound source collection angles; All collected stationary sound source signals and rotating sound source signals are placed in the signal database.

3. The method for separating and locating a rotating sound source and a stationary sound source according to claim 2, characterized in that: The step of calculating the stationary sound source output of the sound signal of each scanning grid by using a stationary beamforming algorithm comprises: Based on each stationary sound source signal, the sound pressure distribution of the sound field on the sound source plane is calculated by formula 1; in, is the sound pressure distribution of the sound field on the sound source plane, is the sound intensity of the sound source at a certain scanning grid, is the position of the scanning grid, t' represents the time when the sound source occurs, V represents all equivalent sound sources in the sound source plane, is the Green's function in the free acoustic field; The sound pressure distribution of the sound field of each stationary sound source signal in the signal database is calculated by formula 2; in, is the sound pressure distribution of the sound field on the sound source plane, is the sound source distribution function, is the spatial coordinate vector of the sound field, t is the time variable, c is the sound speed constant, is the divergence operator; Introducing medium movement, thereby optimizing Formula 2 to Formula 3; in, is the sound pressure distribution of the sound field on the sound source plane, is the sound source distribution function, is the spatial coordinate vector of the sound field, t is the time variable, c is the sound speed constant, is the divergence operator; Based on formula 3, formula 2 is optimized to obtain formula 4; in, is the sound pressure distribution of the sound field on the sound source plane, is the sound intensity of the sound source at a certain scanning grid, is the position of the scanning grid, t' represents the time when the sound source occurs, V represents all equivalent sound sources in the sound source plane, It is the Green's function in the sound field with fluid motion.

4. The method for separating and locating a rotating sound source and a stationary sound source according to claim 3, characterized in that: The stationary sound source output of the sound signal of each scanning grid is calculated by the stationary beamforming algorithm, and further includes: Randomly select a stationary sound source signal from the signal database; Calculate the stationary beam corresponding to the stationary sound source signal using Formula 5 for the signal; P1(f)=G1(f)*S(f)+E(f) Formula 5; Where P1(f) is the representation of the stationary sound source signal in the frequency domain, G1(f) is the Green's function in the free sound field, S(f) is the frequency representation of the equivalent sound source corresponding to the scanning grid, f is the frequency of interest, and E(f) is the noise; A weight vector is added to the stationary sound source signal through Formula 6 to obtain the beamforming output of the stationary sound source signal; Among them, Y n (f) is the beamforming output after the stationary sound source signal is summed with different weights, Wm,n is the weight from the mth microphone to the nth scanning grid, P m (f) is the frequency domain representation of the signal collected by the mth microphone, It is W m,n The vector representation of YesP m The vector representation of (f), (·) H represents conjugate transpose; The stationary source output of the signal is obtained by squaring the beamforming output using Formula 7; Y c =(Y n (f) 2 = W n H R(f)W n Formula 7: Among them, Yc is the stationary sound source output of the stationary sound source signal, Yn(f ) is the beamforming output after the stationary sound source signal is summed with different weights, R ( f ) is the cross-spectral matrix, WnH is the weight vector, and H is the sign of the conjugate transpose; Return to randomly select a stationary sound source signal from the signal database until all stationary sound source signals in the signal database are selected, and obtain the stationary sound source output corresponding to each stationary sound source signal.

5. The method for separating and locating a rotating sound source and a stationary sound source according to claim 4, characterized in that: The modal sound source output of the acoustic signal of each scanning grid is calculated by the modal component beamforming algorithm, including: Randomly select a rotating sound source signal from the signal database; The modal component beam corresponding to the rotating sound source signal is calculated by formula 8; P2(f)=G2(f)*S(f)+E(f) Formula 8; Among them, P2(f) is the representation of the rotating sound source signal in the frequency domain, G2(f) is the Green's function in the flow sound field of the moving medium, S(f) is the frequency representation of the equivalent sound source corresponding to the scanning grid, f is the frequency of interest, and E(f) is the noise; The Green's function from the mth microphone to the nth grid point is calculated by formula 9; Formula 9; Among them, G2 m,n (f,x m ,x n ) is the Green's function from the mth microphone to the nth grid point, xm is the position coordinate of the mth microphone, xn is the position coordinate of the nth grid, m0 is different modes (different vibration modes or propagation modes that may exist in the sound field or vibration system), and M0 is, is the transfer function under a certain mode, f is the frequency of interest; The output of the modal component beam corresponding to the rotating sound source signal is calculated by formula 10; Among them, b n (f) is the output of the beam formed with the rotating sound source signal mode, M is the total number of microphones, P m (f) is the frequency domain representation of the signal collected by the mth microphone, G2 m,n (f,x m ,x n ) is the Green's function from the mth microphone to the nth grid point; The output of the modal component beam corresponding to the rotating sound source signal is squared by the formula to obtain the modal sound source output of the modal component beam corresponding to the rotating sound source signal. The modal sound source output is shown in Formula 11. Where Yr is the modal sound source output of the modal component beam corresponding to the rotating sound source signal, b n (f) is the output of the modal beamforming wave corresponding to the rotating sound source signal, M is the total number of microphones, and P m (f) is the frequency domain representation of the signal collected by the mth microphone, G2 m,n (f,x m ,x n ) is the Green's function from the mth microphone to the nth grid point; Return to randomly select a rotating sound source signal from the signal database until all rotating sound source signals in the signal database are selected, and obtain the modal sound source output of each rotating sound source signal.

6. The method for separating and locating a rotating sound source and a stationary sound source according to claim 5, characterized in that: The stationary sound source output and the modal sound source output are modeled respectively to obtain a stationary sound source model and a modal sound source model, and the stationary sound source model and the modal sound source model are combined to obtain a joint propagation model, including: The stationary sound source output of the acoustic signal of each scanning grid is defined by formula 12; Yc=Asr*Xr+Ass*Xs Formula 12; Where Yc is the stationary sound source output of the acoustic signal scanning the grid, X s is the stationary sound source signal of the acoustic signal scanning the grid, X r is the rotating sound source signal of the acoustic signal of the scanning grid, A ss is the point spread function of the stationary source signal of the scanning grid using stationary beamforming, A sr is the point spread function of the rotating sound source signal of the scanning grid using a stationary beamformed; The modal sound source output of the acoustic signal of all scanning grids is defined by formula 13; Yr=Arr*Xr+Ars*Xs Formula 13; Where Yr is the modal sound source output of all scanned grids, X s is the stationary sound source signal of all scanning grids, X r is the rotational sound source signal of all scanning grids, A rs is the point spread function of the stationary source signal of all scanning grids using modal composition beamforming, A rr is the point spread function of the rotational source signal of all scanned grids using modal component beamforming.

7. The method for separating and locating a rotating sound source and a stationary sound source according to claim 6, characterized in that: The method of modeling the stationary sound source output and the modal sound source output respectively to obtain a stationary sound source model and a modal sound source model, and combining the stationary sound source model and the modal sound source model to obtain a joint propagation model, further includes: The complete cross-point spread function is constructed by formula 14; Where A is the complete cross-point spread function, A rs is the point spread function of the stationary source signal of all scanning grids using modal composition beamforming, A rr is the point spread function of the rotating sound source signal of all scanning grids using the modal component beamforming, A ss is the point spread function of the stationary source signal of all scanning grids using stationary beamforming, A sr is the point spread function of the rotating sound source signal of all scanning grids using stationary beamforming; Based on the complete cross-point diffusion function, the joint propagation model is constructed through formula 15; Where Yr is the modal sound source output of all scanned grids, Yc is the stationary sound source output of all scanned grids, and X s is the stationary sound source signal of all scanning grids, X r is the rotational sound source signal of all scanning grids, A rs is the point spread function of the stationary source signal of all scanning grids using modal composition beamforming, A rr is the point spread function of the rotating sound source signal of all scanning grids using the modal component beamforming, A ss is the point spread function of the stationary source signal of all scanning grids using stationary beamforming, A sr is the point spread function of the rotating sound source signal of all scanning grids using stationary beamforming; Simplify formula 15 to formula 16; Y=A*X Formula 16; Where Y is the sound source output of all scanning grids, Y = (Y r ,Y c ) T , X is the sound source signal of all scanning grids, X=(X r ,X S )T , ( ·) T Represents matrix transpose.

8. The method for separating and locating a rotating sound source and a stationary sound source according to claim 7, characterized in that: The L1 norm constraint is applied to the joint propagation model to construct the minimum absolute contraction and selection operator model, including: Reconstructing the rotating sound source signal and the stationary sound source signal respectively to obtain a rotating sound source column vector and a stationary sound source column vector; Substitute the rotating sound source column vector and the stationary sound source column vector to obtain X; The L1 norm constraint is imposed on X by formula 17, and a suitable regularization parameter is assigned to it to form a regularization term, so that the energy on the scanning grid is distributed in coefficients, thereby constructing the minimum absolute shrinkage and selection operator model; Where A is the complete cross-point spread function, X is the joint sound source, and Y is the joint matrix; Formula 17 is normalized by formula 18 to obtain a standardized minimum absolute shrinkage and selection operator model; in, is the standardized least absolute shrinkage and selection operator model, is the normalized complete cross-point spread function, Y is the joint matrix, is the normalized joint matrix.

9. The method for separating and locating a rotating sound source and a stationary sound source according to claim 8, characterized in that: The alternating direction multiplier method is used to solve the minimum absolute shrinkage and selection operator model to obtain the energy estimate of each scan grid, including: An auxiliary variable is introduced into the minimum absolute shrinkage and selection operator model, and the formula for introducing the auxiliary variable is shown in Formula 19; Where z is the auxiliary variable, A is the complete cross-point spread function, Y is the joint matrix, and X is the joint sound source; Construct the Lagrangian function through formula 20 and add the augmented Lagrangian term; in, u is the Lagrange multiplier, ρ is the penalty parameter and ρ>0, z is the auxiliary variable, A is the complete cross-point spread function, Y is the joint matrix, and X is the joint sound source; Fix the Lagrange multiplier, auxiliary variable, and X respectively to solve the other two parameters in Equation 20; Set the error threshold; Determine whether the auxiliary variable and the joint sound source X are both less than or equal to the error threshold; If the auxiliary variable and the joint sound source X are both less than or equal to the error threshold, then exit the update; If at least one of the auxiliary variable and the joint sound source X is greater than the error threshold, the Lagrange multiplier, the auxiliary variable, and the joint sound source X are fixed respectively until the auxiliary variable and the joint sound source X are both less than or equal to the error threshold.

10. A system for separating and locating rotating sound sources and stationary sound sources, characterized in that: The separation and positioning system of the rotating sound source and the stationary sound source comprises: A collection component, the collection component includes a plurality of microphones, and the sound source signal of the sound source plane is collected by the collection component; The processing component is communicatively connected with the collection component, the sound source signal of the sound source plane collected by the collection component is transmitted to the processing component, and the processing component performs the method for separating and locating the rotating sound source and the stationary sound source as described in any one of claims 1 to 9 on the sound source signal.