Multivariate microphone array sound source positioning method based on GS-RBF and time delay estimation

By combining a two-dimensional multi-element circular microphone array with a GCC_PHAT and GS-optimized RBF neural network model, the problems of large computational load and insufficient accuracy of traditional methods in complex environments are solved, and efficient and accurate detection of latent faults in transmission lines is achieved.

CN119861333BActive Publication Date: 2026-05-01CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2024-12-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional microphone array sound source localization methods are computationally intensive and lack accuracy in complex environments, making it difficult to meet the high-efficiency detection requirements for hidden faults in power transmission lines.

Method used

A two-dimensional multi-element circular microphone array is used in conjunction with the generalized cross-correlation phase transform algorithm GCC_PHAT and the grid search algorithm GS to optimize the RBF neural network model. By calculating the time difference of arrival (TDOA) matrix and optimizing the grid search, the sound source localization is improved, thereby enhancing noise resistance and spatial resolution.

Benefits of technology

It enables rapid and accurate fault sound source localization in complex environments, improving the efficiency and reliability of latent fault detection in transmission lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119861333B_ABST
    Figure CN119861333B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on GS-RBF and time delay estimation multi-element microphone array sound source positioning method, belong to sound source positioning technical field, comprising the following steps: S1: the near-field model of fault sound source is established, and microphone array design is carried out;S2: based on generalized cross-correlation phase transformation algorithm GCC_PHAT calculation time difference of arrival TDOA matrix;S3: construct the sound source positioning model based on RBF neural network;S4: based on grid search algorithm GS, the sound source positioning model is optimized, and the model after optimization is used to locate sound source.The present application guarantees the rapidity and accuracy of fault sound source positioning when unmanned aerial vehicle inspection, and the anti-noise performance and spatial resolution capability in complex environment are considered, and an efficient, reliable technical solution is provided for the detection of transmission line hidden fault.
Need to check novelty before this filing date? Find Prior Art

Description

A Multi-Microphone Array Sound Source Localization Method Based on GS-RBF and Time Delay Estimation Technical Field

[0001] This invention belongs to the field of sound source localization technology, and relates to a multi-microphone array sound source localization method based on GS-RBF and time delay estimation. Background Technology

[0002] The safe and stable operation of transmission lines is the fundamental guarantee for the reliability of the power system. Traditional transmission line inspections rely heavily on manual checks, which are inefficient and limited by the experience of the inspectors and the complexity of the on-site environment, posing certain safety hazards and risks of missed inspections.

[0003] To address the aforementioned issues, acoustic source localization technology has gradually become one of the key methods for detecting latent faults in power transmission lines. The acoustic signals generated by partial discharge exhibit strong specificity. Acoustic source localization technology, through spatial localization and time delay analysis of these signals, can achieve precise location of partial discharge faults, significantly improving the detection capability of latent defects. The design of the acoustic field sensor array is crucial to the performance of the acoustic source localization system, with the array's geometry directly affecting the accuracy and applicability of the localization. Existing research indicates that a reasonable array geometry design can improve the robustness and accuracy of localization in environments with complex background noise and multipath interference.

[0004] There are many traditional microphone array-based sound source localization methods, such as Time Delay Difference (TDOA) sound source localization algorithm, Gaussian Convolutional Beamforming (GCC) sound source localization method, and High-Resolution Spectral Estimation (MUSIC) sound source localization method. Among them, the TDOA sound source localization algorithm is widely used in many fields due to its simple principle and real-time localization capability. However, the TDOA localization algorithm usually requires solving nonlinear equations, especially in the case of multiple sensors, where the computational load can be very large. With the development of machine learning, the emergence of neural networks has provided a good solution for complex nonlinear problems. In 2006, Harbin Institute of Technology proposed a neural network-based sound source localization method. By simulating the sound source positions of multiple microphone arrays, the neural network is used to estimate the sound source position, achieving high-precision sound source localization. In 2014, Shanghai Jiao Tong University proposed using convolutional neural networks (CNNs) for sound source localization. Research shows that this method can significantly improve the sound source localization accuracy in complex environments. In 2021, the Korea Advanced Institute of Science and Technology (KAIST) proposed a sound source localization method based on radial basis function (RBF) neural networks. By training the RBF network, the estimation of sound source location was optimized, especially performing well in environments with high noise levels. RBF neural networks are characterized by short training time and high nonlinear mapping characteristics, making them a highly efficient solution. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a multi-microphone array sound source localization method based on GS-RBF and time delay estimation. A two-dimensional multi-microphone array is employed to adapt to the characteristics of high intermediate frequency bandwidth and signal-to-noise ratio in power transmission line detection scenarios. For near-field spherical sound source models, a sound source localization method combining the Generalized Cross-Correlation Phase Transform (GCC_PHAT) algorithm is proposed. Accurate localization is achieved by calculating the Time Difference of Arrival (TDOA) matrix and combining it with a grid search algorithm (GS) optimized RBF neural network model. This method balances noise resistance and spatial resolution in complex environments, providing an efficient and reliable technical solution for the detection of latent faults in power transmission lines.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for localizing sound sources using a multi-microphone array based on GS-RBF (Grid Search-Radial basis function) and time delay estimation includes the following steps:

[0008] S1: Establish a near-field model of the fault sound source and design a microphone array;

[0009] S2: Calculate the Time Difference of Arrival (TDOA) matrix based on the Generalized Cross Correlation Phase Transform (GCC_PHAT) algorithm;

[0010] S3: Construct a sound source localization model based on RBF neural network;

[0011] S4: The sound source localization model is optimized based on the grid search algorithm GS, and the optimized model is used to locate the sound source.

[0012] Furthermore, the establishment of a near-field model of the fault sound source and the design of the microphone array in step S1 specifically includes the following steps:

[0013] S11: The Fresnel zone formula is used to divide the application range of near-field and far-field sources, and a near-field model of the fault sound source is constructed:

[0014]

[0015] In the formula: r is the distance from the sound source to the sensor; D is the array aperture; λ is the wavelength of the sound source;

[0016] S12: Design of a circular microphone array configuration based on a fault sound source model:

[0017]

[0018] Where: N is the number of microphones; R is the array radius; (x n ,y n Let be the Cartesian coordinates of the nth microphone.

[0019] Furthermore, the Time Difference of Arrival (TDOA) matrix is ​​calculated based on the Generalized Cross-Correlation Phase Transform (GCC_PHAT) algorithm, including the following steps:

[0020] S21: Utilizing an N-element microphone array (M1, M2...M... N Simultaneously acquire sound source signals in space to obtain the corresponding time-domain discrete signals x1(n), x2(n)……x N (n);

[0021] S22: Time-domain signal x for each pair of microphones i (n) and x j Perform a Fast Fourier Transform (FFT) on (n)(i,j[1,8],i≠j) to convert it to the frequency domain, and obtain the frequency domain signal representation:

[0022]

[0023] In the formula Fourier transform;

[0024] S23: Then calculate the cross-power spectral function:

[0025]

[0026] In the formula For X j The conjugate of (ω);

[0027] S24: Smooth coherent weighting of the cross-power spectrum in the frequency domain to suppress noise and reverberation interference, and finally perform inverse Fourier transform to obtain the cross-correlation function:

[0028]

[0029] In the formula It is a weighting function in the frequency domain;

[0030] S25: Introduce a suitable weighting function based on the characteristics of the real environment and the received signal;

[0031] S26: In the cross-correlation function R ij The delay time τ corresponding to the maximum peak value is searched in (τ). max TDOA, representing the time delay difference between the two signals:

[0032]

[0033] In the formula, τ ij The time delay outputs are for paths i and j. The operation to find the maximum value of the objective function corresponding to τ;

[0034] S27: Based on the number of microphone array elements, estimate the time delay τ for all microphone pairs. ij The organization is a symmetric TDOA matrix T. N×N ,in:

[0035]

[0036] Furthermore, step S25 introduces the PHAT weighting function, the specific expression of which is as follows:

[0037]

[0038] Furthermore, step S3, which involves constructing a sound source localization model based on an RBF neural network, specifically includes the following steps:

[0039] S31: Extract the upper triangular part of the TDOA matrix to obtain the TDOA features between microphone pairs, and use the TDOA features as network input samples; the expression for the network input samples is:

[0040]

[0041] In the formula, For the k-th input sample, x kn This refers to the nth element extracted from the upper triangular part of the TDOA matrix in the kth input sample.

[0042] S32: After determining the input samples, select the Gaussian function as the hidden layer function:

[0043]

[0044] In the formula, Euclidean norm, where σ is the variance of the Gaussian function. Let be the center vector of the radial basis function of the i-th hidden node;

[0045] S33: The network output can be obtained by weighted summation of the activation values ​​in the hidden layer.

[0046]

[0047] In the formula, y is the actual output value of the neural network, and w i Let be the weight from the i-th node in the hidden layer to the output layer, and m be the number of hidden nodes;

[0048] S34: Set the width spread hyperparameter of the radial basis function, input the input samples into the RBF neural network for network training to establish a sound source localization model.

[0049] Furthermore, step S4, which involves optimizing the sound source localization model based on the grid search algorithm GS, specifically includes the following steps:

[0050] S41: List the possible values ​​of the hyperparameter spread to be tuned in the RBF neural network;

[0051] S42: Evaluate model performance using the mean squared error of model predictions (MSE), expressed as follows:

[0052]

[0053] In the formula yi Let i be the true value of the i-th sample. Let be the predicted value of the i-th sample, and n be the number of samples.

[0054] S43: Iterate through all spread values ​​and perform cross-validation for each spread value;

[0055] S44: Based on the cross-validation results, select the spread value with the lowest mean square error (MSE) of the model prediction as the optimal spread value, and use it as the final hyperparameter of the sound source localization model based on the RBF neural network.

[0056] The beneficial effects of this invention are as follows: This invention ensures the speed and accuracy of locating fault sound sources during UAV inspections, while taking into account noise resistance and spatial resolution in complex environments, providing an efficient and reliable technical solution for detecting latent faults in power transmission lines.

[0057] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0059] Figure 1 is a flowchart of the sound source localization method of the present invention;

[0060] Figure 2 shows the microphone array configuration;

[0061] Figure 3 is a block diagram of the GCC_PHAT algorithm;

[0062] Figure 4 shows the structure of the RBF neural network;

[0063] Figure 5 illustrates the grid search (GS) algorithm process for 3000 simulated sound sources;

[0064] Figure 6 shows the prediction results of the BP neural network for 3000 simulated sound sources;

[0065] Figure 7 shows the prediction results of the RBF neural network for 3000 simulated sound sources. Detailed Implementation

[0066] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0067] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0068] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0069] Please refer to Figures 1 to 7. Figure 1 is a flowchart of the method of the present invention. As shown in Figure 1, the present invention provides a sound source localization method for a multi-microphone array based on a GS-RBF neural network and a time delay estimation algorithm, including the following steps:

[0070] Step S1: Establish a two-dimensional octet circular microphone array configuration and a near-field fault sound source model:

[0071]

[0072] In the formula: R is the array radius; (x n ,y n Let be the Cartesian coordinates of the nth microphone.

[0073] Step S2: Using an 8-element microphone array (M1, M2...M8), obtain the corresponding discrete-time signals x1(n), x2(n)...x8(n), and obtain the TDOA matrix based on the GCC_PHAT algorithm. This specifically includes the following steps:

[0074] S21: Use an N-element microphone array (M1, M2……M8) to synchronously acquire sound source signals in the space to obtain the corresponding time-domain discrete signals x1(n), x2(n)……x8(n).

[0075] S22: Time-domain signal x for each pair of microphones i (n) and x j Perform a Fast Fourier Transform (FFT) on (n)(i,j∈[1,8],i≠j) to transform it to the frequency domain, and obtain the frequency domain signal representation:

[0076]

[0077] In the formula This is a Fourier transform.

[0078] S23: Then the cross-power spectrum function can be obtained:

[0079]

[0080] In the formula For X j (ω) conjugate.

[0081] S24: To sharpen the peak value of the cross-correlation function, a smooth coherent weighting of the cross-power spectrum is performed in the frequency domain to suppress noise and reverberation interference. Finally, an inverse Fourier transform is performed to obtain the cross-correlation function:

[0082]

[0083] In the formula It is a weighting function in the frequency domain.

[0084] S25: In practical applications, a suitable weighting function is introduced based on the characteristics of the real-world environment and the received signal. Here, the PHAT weighting function is introduced, which has strong anti-interference and anti-noise capabilities and can improve the accuracy of time delay estimation when the signal-to-noise ratio is high. The specific expression is as follows:

[0085]

[0086] S26: In the cross-correlation function R ij The delay time τ corresponding to the maximum peak value is searched in (τ). max The time delay difference (TDOA) between the two signals:

[0087]

[0088] In the formula, τ ij The time delay outputs are for paths i and j. The operation is to find the value of τ corresponding to the maximum value of the objective function.

[0089] S27: Finally, based on the number of microphone array elements, estimate the time delay τ for all microphone pairs. ij The organization is a symmetric TDOA matrix T. 8×8 ,in:

[0090]

[0091] Step S3: Establish a sound source localization model based on the RBF neural network, which specifically includes the following steps:

[0092] S31: Extract the upper triangular portion of the TDOA matrix (i.e., remove the main diagonal and elements below it) to obtain the TDOA features between microphone pairs. Use these TDOA features as network input samples. The network input sample expression is:

[0093]

[0094] In the formula, For the k-th input sample, x kn It is the nth element extracted from the upper triangular part of the TDOA matrix in the kth input sample.

[0095] S32: After determining the input samples, select the Gaussian function as the hidden layer function:

[0096]

[0097] In the formula, Euclidean norm, where σ is the variance of the Gaussian function. Let be the center vector of the radial basis function (Gaussian function) of the i-th hidden node.

[0098] S33: The network output can be obtained by weighted summation of the activation values ​​in the hidden layer.

[0099]

[0100] In the formula, y is the actual output value of the neural network, and w i Let be the weight from the i-th node in the hidden layer to the output layer, and m be the number of hidden nodes.

[0101] S34: Set the "spread" hyperparameter of the radial basis function, input the input samples into the RBF neural network for network training to establish a sound source localization model.

[0102] Step S4: Optimize the sound source localization model based on the grid search algorithm (GS), specifically including the following steps:

[0103] S41: List the possible values ​​of the hyperparameter spread to be adjusted in the RBF neural network, and set spread to take 30 possible values ​​in the range of 0.1 to 2.

[0104] S42: Select to use the mean squared error of model prediction (MSE) to evaluate model performance. The expression for mean squared error (MSE) is:

[0105]

[0106] In the formula y i Let i be the true value of the i-th sample. Let be the predicted value of the i-th sample, and n be the number of samples.

[0107] S43: Iterate through all spread values ​​and perform cross-validation for each spread value. The cross-validation fold number is set to 5, i.e., perform 5-fold cross-validation.

[0108] S44: Based on the results of 5-fold cross-validation, the spread value with the lowest mean square error (MSE) of the model prediction is selected as the optimal spread value, which is used as the final hyperparameter of the sound source localization model based on the RBF neural network.

[0109] Experimental verification:

[0110] 500, 1000, 1500, 2000, and 3000 simulated sound source samples were randomly generated respectively to establish a sound source localization model based on an RBF neural network. The simulated sound source samples were randomly generated in the Oxyz three-dimensional space: x∈[-2,2], y∈[-2,2], z∈[0,2]. For all simulated sound source samples, the upper triangular part of the TDOA matrix was extracted as the TDOA feature after calculating the TDOA matrix based on GCC_PHAT. The data was divided into a 70% training set, 15% validation set, and 15% test set, and the input samples were normalized. After obtaining the optimal spread value using a grid search optimization algorithm, the input samples were fed into the RBF neural network for training. Simultaneously, sound source localization models based on a BP neural network were established for the above five sample sizes. The root mean square error of the sound source localization model based on the GS-RBF neural network was compared with that of the sound source localization model based on the BP neural network to further demonstrate the accuracy of the sound source localization of this invention.

[0111] Table 1 lists the optimal spread values ​​obtained by the Grid Search (GS) algorithm under five simulated sound source numbers. It can be seen that the optimal spread value decreases with a larger sample size. Figure 5, using 3000 simulated sound sources as an example, illustrates the Grid Search (GS) algorithm process, showing how the model prediction error changes with the spread value. Table 2 lists the root mean square error (RMSE) values ​​of the sound source localization models based on two neural networks under five simulated sound source numbers. It can be seen that the RMSE of the sound source localization model based on the GS-RBF neural network is lower than that of the sound source localization model based on the BP neural network under all five simulated sound source numbers. Figures 6 and 7, using 3000 simulated sound sources as an example, show the training results of the sound source localization models based on the two neural networks.

[0112] Table 1

[0113] Sample size: 500 1000 1500 2000 3000 Optimal spread value: 0.23 1 0.36 2 1 0.1 0.1 0.1655 surface

[0114] Table 2

[0115]

[0116] In the above embodiments, the reference to "this embodiment" in the specification indicates that a specific feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple appearances of "this embodiment" do not necessarily refer to the same embodiment.

[0117] In the above embodiments, although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. The embodiments of the invention are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims.

[0118] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods in this embodiment.

[0119] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0120] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to cause the terminal to perform any of the methods in this embodiment.

[0121] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0122] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal performs the steps of the above method.

[0123] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0124] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0125] This invention can be used in a wide range of general-purpose or special-purpose computing system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0126] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for localizing sound sources using a multi-microphone array based on GS-RBF and time delay estimation, characterized in that: Includes the following steps: S1: Establish a near-field model of the fault sound source and design a microphone array; S2: Calculate the Time Difference of Arrival (TDOA) matrix based on the Generalized Cross-Correlation Phase Transform (GCC_PHAT) algorithm; S3: Construct a sound source localization model based on an RBF neural network; S4: Optimize the sound source localization model based on the Grid Search (GS) algorithm, and use the optimized model to locate the sound source; Calculating the Time Difference of Arrival (TDOA) matrix based on the Generalized Cross-Correlation Phase Transform (GCC_PHAT) algorithm includes the following steps: S21: Using Meta microphone array Synchronously acquire sound source signals in space to obtain corresponding discrete-time signals. S22: Time-domain signal for each pair of microphones and Perform Fast Fourier Transform (FFT). Transform it to the frequency domain to obtain the frequency domain signal representation: In the formula For Fourier transform; S23: Then calculate the cross-power spectral density function: In the formula for S24: Smooth coherent weighting of the cross-power spectrum in the frequency domain to suppress noise and reverberation interference, and finally perform inverse Fourier transform to obtain the cross-correlation function: In the formula S25: Introduce a suitable weighting function based on the real-world environment and the characteristics of the received signal; S26: In the cross-correlation function The latency corresponding to the maximum peak value in the search TDOA, representing the time delay difference between the two signals: In the formula, for The two corresponding delay outputs, To find the maximum value of the objective function S27: Based on the number of microphone array elements, estimate the time delay of all microphone pairs. The organization is a symmetric TDOA matrix. ,in: Step S25 introduces the PHAT weighting function, the specific expression of which is as follows: Step S3, which involves constructing a sound source localization model based on an RBF neural network, specifically includes the following steps: S31: Extract the upper triangular portion of the TDOA matrix to obtain the TDOA features between microphone pairs, and use these TDOA features as network input samples; the network input sample expression is: In the formula, For the first One input sample, For the first The nth sample extracted from the upper triangular part of the TDOA matrix from the nth input sample One element; S32: After determining the input sample, select the Gaussian function as the hidden layer function: In the formula, Euclidean norm The variance of the Gaussian function. For the first S33: The center vector of the radial basis function of each hidden node; S33: The network output can be obtained by weighted summation of the activation values ​​in the hidden layer. In the formula The actual output value of the neural network. For the hidden layer The weights from each node to the output layer. S34: Set the width spread hyperparameter of the radial basis function, and input the input samples into the RBF neural network for network training to establish a sound source localization model; Step S4 optimizes the sound source localization model based on the grid search algorithm GS, specifically including the following steps: S41: List the possible values ​​of the hyperparameter spread to be adjusted in the RBF neural network; S42: Use the model prediction mean square error (MSE) to evaluate the model performance, the expression is: In the formula For the first The true value of each sample For the first The predicted value for each sample, S43: Iterate through all spread values ​​and perform cross-validation for each spread value; S44: Based on the cross-validation results, select the spread value with the lowest mean square error (MSE) of the model prediction as the optimal spread value, which will be used as the final hyperparameter of the sound source localization model based on the RBF neural network.

2. The method for localizing sound sources using a multi-microphone array based on GS-RBF and time delay estimation according to claim 1, characterized in that: Step S1, which involves establishing a near-field model of the fault sound source and designing a microphone array, specifically includes the following steps: S11: Using the Fresnel zone formula to divide the usage range of near-field and far-field sources, and constructing a near-field model of the fault sound source: In the formula: The distance from the sound source to the microphone array; For array aperture; S12: Design a circular microphone array configuration based on a fault sound source model. In the formula: Number of microphones; The array radius; For the first Cartesian coordinates of a microphone.