Sound source positioning method based on off-line array element coordinate library building

By building an offline library of array element coordinates, a library of fractional and integer time delay coefficients is pre-established, which solves the problem of high computational complexity in existing sound source localization methods in broadband 3D or low-to-mid-frequency sound source localization scenarios, and achieves more efficient sound source localization and more accurate signal alignment.

CN122085217BActive Publication Date: 2026-07-21TIANJIN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-04-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing sound source localization methods have high computational complexity and resource consumption in broadband 3D sound source or low-mid frequency sound source localization scenarios, and insufficient time delay compensation accuracy leads to a decrease in spatial spectrum quality, making it difficult to achieve accurate coherent alignment of multi-element channel signals.

Method used

The sound source localization method based on offline library construction of array element coordinates pre-establishes fractional delay coefficient library and integer delay library, obtains fractional delay coefficient and integer sampling delay through offline calculation, directly performs delay compensation on scanning signal, reduces online calculation burden, and achieves coarse compensation through integer delay and fine compensation through fractional delay coefficient, thus finely aligning multi-element channel signals.

Benefits of technology

It reduces the consumption of online computing resources, improves the accuracy and efficiency of sound source localization, achieves accurate coherent alignment of multi-element channel signals in the target direction, and improves the quality of the spatial spectrum.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122085217B_ABST
    Figure CN122085217B_ABST
Patent Text Reader

Abstract

The application provides a sound source positioning method based on offline database construction of array element coordinates, which can be applied to the technical field of sound source positioning. The method is applied to a positioning array, the positioning array comprises a plurality of array elements, the plurality of array elements comprise a reference array element, and the method comprises the following steps: for any scanning direction of each of the plurality of array elements in a scanning range: acquiring a scanning signal corresponding to each of the plurality of array elements in the scanning direction; calling a fractional time delay coefficient library and an integer time delay library corresponding to each of the plurality of array elements; determining a fractional time delay coefficient and an integer sampling delay corresponding to each of the plurality of array elements in the scanning direction respectively; performing time delay compensation on the scanning signal to obtain a scanning compensation signal corresponding to each of the plurality of array elements in the scanning direction; and obtaining a signal space spectrum corresponding to any scanning direction according to the scanning compensation signal corresponding to each of the plurality of array elements in any scanning direction; and determining a target direction of a sound source to be positioned in the scanning range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source localization technology, and more specifically, to a sound source localization method based on offline database construction of array element coordinates. Background Technology

[0002] Sound source localization, which involves analyzing the physical characteristics of sound reaching an observation point to determine the spatial location of a sound source, is a fundamental problem in array signal processing and intelligent acoustic sensing. It is widely used in scenarios such as voice interaction, conferencing systems, robot hearing, smart terminals, security monitoring, industrial diagnostics, and acoustic imaging. However, sound source localization methods in these technologies require significant resources. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a sound source localization method based on offline library construction of array element coordinates.

[0004] One aspect of this invention provides a sound source localization method based on offline library construction of array element coordinates, applied to a localization array. The localization array includes multiple array elements, among which a reference array element is included. The sound source localization method includes: for any scanning direction within the scanning range of each of the multiple array elements: acquiring the scanning signal corresponding to each of the multiple array elements in the scanning direction; calling the fractional delay coefficient library and the integer delay library corresponding to each of the multiple array elements, wherein the fractional delay coefficient library and the integer delay library are established based on the relative positions between the array elements and the reference array element; determining the fractional delay coefficient and the integer sampling delay corresponding to each of the multiple array elements in the scanning direction from the fractional delay coefficient library and the integer delay library respectively; performing time delay compensation on the scanning signal according to the fractional delay coefficient and the integer sampling delay corresponding to each of the multiple array elements in the scanning direction to obtain the scanning compensation signal corresponding to each of the multiple array elements in the scanning direction; obtaining the signal spatial spectrum corresponding to any scanning direction according to the scanning compensation signal corresponding to each of the multiple array elements in any scanning direction; and determining the target direction of the sound source to be located within the scanning range according to the signal spatial spectrum corresponding to any scanning direction.

[0005] According to embodiments of the present invention, a fractional delay coefficient library and an integer delay library are pre-established based on the relative positions of multiple array elements in the positioning array with a reference array element. After acquiring the scanning signal, the corresponding fractional delay coefficient and integer sampling delay can be directly obtained by calling the fractional delay coefficient library and the integer delay library, thereby directly performing time delay compensation on the scanning signal to obtain a scan-compensated signal. Furthermore, a signal spatial spectrum is established, and the target direction of the sound source to be located is determined based on the signal spatial spectrum. Since the relatively complex fractional delay coefficient and integer sampling delay process is completed offline, there is no need to recalculate the fractional delay coefficient and integer sampling delay in the online stage; only the fractional delay coefficient library and the integer delay library need to be called, eliminating the need for online generation and reducing the online occupation of computing resources. Furthermore, using integer delay to achieve coarse compensation of the scanning signal and using fractional delay coefficients to achieve fine compensation of the scanning signal allows for more precise coherent alignment of signals from multiple array element channels in the target direction. Attached Figure Description

[0006] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, which will be explained in conjunction with the drawings.

[0007] Figure 1 An exemplary system architecture for applying a sound source localization method based on offline library construction using array element coordinates, according to an embodiment of the present invention, is shown.

[0008] Figure 2 A flowchart of a sound source localization method based on offline library construction using array element coordinates according to an embodiment of the present invention is shown.

[0009] Figure 3 A schematic diagram of a positioning array according to an embodiment of the present invention is shown.

[0010] Figure 4 An embodiment of the present invention is shown. Figure 3 The diagram shows a schematic of the spatial spectrum of the signal obtained by the positioning array.

[0011] Figure 5 An embodiment of the present invention is shown. Figure 4 An azimuth slice of the signal spatial spectrum.

[0012] Figure 6 An embodiment of the present invention is shown. Figure 4 Elevation slice of the signal spatial spectrum.

[0013] Figure 7 A block diagram of a sound source localization device based on offline database construction of array element coordinates according to an embodiment of the present invention is shown.

[0014] Figure 8A block diagram of an electronic device suitable for implementing the methods described above, according to an embodiment of the present invention, is shown. Detailed Implementation

[0015] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0016] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0017] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0018] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0019] In the embodiments of this invention, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to maintain the security of user personal information and network security.

[0020] In the embodiments of the present invention, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0021] Sound source localization methods in related technologies include those based on time delay estimation, such as using generalized cross-correlation and its variants to locate sound sources by estimating the signal arrival time difference between array elements and combining it with the localization array to solve for the sound source direction. There are also methods based on beamforming, which perform delay compensation and energy focusing on multi-channel signals in candidate directions, and then determine the direction of arrival from the spatial spectrum peak value. Finally, there are methods based on subspace decomposition, such as using multiple signal classification (MUSIC) algorithms and their extended forms to locate sound sources, estimating the sound source direction through the orthogonality of the signal subspace and the noise subspace.

[0022] However, for broadband 3D sound sources, the actual propagation delay is usually not an integer multiple of the sampling period. If only integer sampling point shifting is used for completion delay compensation, non-integer delay approximation errors will be introduced, making it difficult for signals from different array elements to achieve strict coherent superposition in the target direction. This leads to beam output energy loss, main lobe broadening, peak shift, and angle estimation errors. Although beamforming methods have advantages such as intuitiveness, robustness, and suitability for engineering deployment, they typically require scanning the spatial spectrum point-by-point along dense candidate directions. Furthermore, the performance of beamforming methods is affected in low signal-to-noise ratio environments, and they have higher computational complexity compared to subspace decomposition-based methods. When the scanning direction dimension is expanded from two dimensions to a joint azimuth-elevation search, the online computational burden of beamforming methods will further increase. The relationship between the positioning array aperture and the acoustic wavelength determines the spatial resolution. For size-constrained compact positioning arrays, especially in the low-frequency band, the effective aperture of the positioning array is relatively small compared to the wavelength, naturally limiting the spatial resolution and easily leading to problems such as a broad main lobe, difficulty in distinguishing adjacent directions, and enhanced side lobes and spurious peak interference. In beamforming methods, the angular resolution of the positioning array is closely related to the array's aperture and operating frequency; resolution tends to decrease under low-frequency conditions. Therefore, in broadband or low-to-mid-frequency sound source positioning scenarios, insufficient time delay compensation accuracy or improper scanning implementation will further degrade the spatial spectrum quality.

[0023] In view of this, embodiments of the present invention provide a sound source localization method based on offline library construction of array element coordinates. A fractional delay coefficient library and an integer delay library are pre-established based on the relative positions of multiple array elements in the localization array with a reference array element. After acquiring the scanning signal, the corresponding fractional delay coefficients and integer sampling delays can be directly obtained from the fractional delay coefficient libraries and integer delay libraries, thereby directly compensating for the time delay of the scanning signal to obtain a scan-compensated signal. Furthermore, a signal spatial spectrum is established, and the target direction of the sound source to be located is determined based on the signal spatial spectrum. Since the relatively complex fractional delay coefficient and integer sampling delay process is completed offline, there is no need to recalculate the fractional delay coefficients and integer sampling delays in the online stage; only the fractional delay coefficient library and integer delay library need to be called, eliminating the need for online generation and reducing online computational resource consumption. Furthermore, using integer delays for coarse compensation of the scanning signal and fractional delay coefficients for fine compensation allows for more precise coherent alignment of signals from multiple array element channels in the target direction.

[0024] Figure 1 An exemplary system architecture for applying a sound source localization method based on offline library construction using array element coordinates, according to an embodiment of the present invention, is shown. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of the present invention, in order to help those skilled in the art understand the technical content of the present invention, but do not mean that embodiments of the present invention cannot be used in other devices, systems, environments or scenarios.

[0025] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a positioning array 101 and an array processor 102. The array processor 102 can execute a sound source localization method based on offline library construction of array element coordinates based on multiple array elements in the positioning array 101.

[0026] The positioning array 101 can be a two-dimensional planar array or a three-dimensional array. The positioning array can be a regular array or an irregular array. The topology of the multiple array elements of the positioning array can be configured according to the actual application scenario, hardware deployment method and spatial constraints, including but not limited to uniform circular arrays, non-uniform circular arrays, double-ring arrays, multi-ring arrays, spiral arrays, sparse arrays, spherical arrays, rectangular arrays and combinations of arrays with different topologies.

[0027] Fractional delay coefficient libraries and integer delay libraries can be established based on the relative position differences between each element in the positioning array and the reference element. In the far-field propagation model, the signal received by the positioning array can be considered a plane wave; therefore, the fractional delay coefficient libraries and integer delay libraries can be constructed based on the far-field propagation model.

[0028] The integer delay library can be obtained through the following operations: Based on the individual array element coordinates of multiple array elements and the reference array element's coordinates, determine the positional difference between each array element and the reference array element; based on the positional difference between each array element and the reference array element, determine the signal arrival time difference between each array element and the reference array element in any scanning direction; based on the signal arrival time difference between each array element and the reference array element in any scanning direction and the sampling frequency of the positioning array, determine the sampling point delay between each array element and the reference array element in any scanning direction; round the sampling point delay between each array element and the reference array element in any scanning direction to obtain the candidate integer sampling delay corresponding to each array element in any scanning direction; based on the correspondence between the candidate integer sampling delay of each array element in any scanning direction and the multiple array elements and scanning directions, obtain the integer delay library.

[0029] In an embodiment of the present invention, it can be assumed that the number of multiple array elements in the positioning array is M, where M is an integer greater than 1. The three-dimensional coordinates of the multiple array elements in the positioning array spatial coordinate system are as follows (1).

[0030] (1);

[0031] in, This represents the coordinates of the m-th element out of M array elements. , and Let m and m represent the coordinates of the m-th array element on the x-axis, y-axis, and z-axis, respectively, where 1 < m ≤ M and m is an integer.

[0032] The coordinates of the M array elements together constitute the coordinate set of the positioning array, as shown in equation (2).

[0033] (2);

[0034] in, This represents the set of coordinates for the positioning array. These represent the coordinates of the 1st array element, the 2nd array element, ..., the Mth array element, respectively.

[0035] Select one of the array elements of the positioning array as the reference array element. The coordinates of the reference array element are as follows (3).

[0036] (3);

[0037] in, Indicates the coordinates of the reference array element. , and These represent the coordinates of the reference array element on the x-axis, y-axis, and z-axis, respectively.

[0038] The positional difference between the m-th array element and the reference array element is as follows (4).

[0039] (4);

[0040] in, This represents the positional difference between the m-th array element and the reference array element.

[0041] exist hour, .

[0042] Taking the positioning array as a uniform circular array without a center as an example, when the reference array element is located at a predetermined position on the circumference, the center of the positioning array is as follows (5).

[0043] (5);

[0044] in, denoted by r, which represents the center of the positioning array.

[0045] The coordinates of the m-th array element are as follows (6).

[0046] (6);

[0047] in, Let be the circular angle parameter of the m-th array element.

[0048] Multiple array elements can be numbered and arranged in a clockwise or counterclockwise direction. This embodiment is only used to illustrate how to construct the three-dimensional coordinates of array elements according to the topology of the actual positioning array, and does not constitute a limitation on the structure of the positioning array in this invention.

[0049] The scanning range of the positioning array can be in three-dimensional space, so the set of scanning azimuth angles and the set of scanning elevation angles can be defined as follows (7).

[0050] (7);

[0051] in, Represents the set of scan azimuth angles. They represent the 1st, 2nd, ..., 1st, 2nd, ..., 3rd. One scanning azimuth angle Represents the set of scanning pitch angles. They represent the 1st, 2nd, ..., 1st, 2nd, ..., 3rd. Each scan pitch angle.

[0052] The scanning range is then defined in equation (8).

[0053] (8);

[0054] in, Indicates the scan range. This represents the azimuth angle of the i-th scan. This represents the pitch angle of the j-th scan. Indicates the scanning direction within the scanning range.

[0055] As an example, the positioning array can be a centerless uniform circular array of 16 elements, where r = 0.008m. The elements can be distributed sequentially along the circumference. The first element in the positioning array can be selected as the reference element, with coordinates as follows: The unit is meters. Then the set of three-dimensional coordinates of all array elements is... The step size for the scanning azimuth and scanning elevation angles can be 0.5°, and the set of scanning azimuth angles is as follows: The scanning elevation angle scanning set is set as follows: .

[0056] For any scanning direction Angles can be converted into radians, as shown in equation (9).

[0057] (9);

[0058] in, Let represent the azimuth angle of the i-th scan in radians. The pitch angle of the j-th scan is represented in radians.

[0059] The unit vector pointing from the positioning array to the scanning direction is as follows (10).

[0060] (10);

[0061] in, This represents the unit vector pointing from the positioning array in the scanning direction.

[0062] Based on the position difference between the m-th array element and the reference array element The m-th array element is obtained in the scanning direction. The signal arrival time difference between the upper and reference array elements is given by the following equation (11).

[0063] (11);

[0064] in, This indicates that the m-th array element is in the scanning direction. The time difference of signal arrival between the upper and reference array elements, where c represents the speed of sound. Indicates the transposed form .

[0065] The negative sign in equation (11) is used to maintain consistency with the time delay compensation symbol convention of the present invention, that is: when the array element receives the sound wave from the scanning direction earlier than the reference array element, the signal arrival time difference relative to the reference array element is defined as a positive value, so that a corresponding positive delay needs to be applied to the array element during online compensation so that the signals of each channel can achieve coherent alignment in the target direction.

[0066] Based on the sampling frequency of the positioning array, the signal arrival time difference is converted into sampling point delay, and the sampling point delay of the m-th array element relative to the reference array element is obtained as shown in the following equation (12).

[0067] (12);

[0068] in, This indicates that the m-th array element is in the scanning direction. Sampling point delay between the upper and reference array elements, The sampling frequency.

[0069] It is possible to scan the m-th array element in the scanning direction. The sampling point between the upper and reference array elements is delayed and rounded to obtain the result with respect to the m-th array element in the scanning direction. The sampling delay of the corresponding candidate integer is as follows (13).

[0070] (13);

[0071] in, This indicates that the m-th array element is in the scanning direction. The corresponding candidate integer sampling delay.

[0072] The candidate integer sampling delays of multiple array elements in any scanning direction, as well as the correspondence between multiple array elements and scanning directions, can be stored in a certain arrangement to obtain an integer delay library.

[0073] The fractional delay coefficient library can be obtained through the following operations: Based on the sampling point delay between multiple array elements and the reference array element in any scanning direction, and the candidate integer sampling delay of each array element in any scanning direction, determine the candidate fractional sampling delay of each array element in any scanning direction; based on the candidate fractional sampling delay of each array element in any scanning direction, determine the candidate fractional delay coefficients in the fractional delay filter corresponding to each array element in any scanning direction; and based on the correspondence between the candidate fractional delay coefficients of the fractional delay filter corresponding to each array element in any scanning direction and the multiple array elements and scanning directions, obtain the fractional delay coefficient library.

[0074] Based on the m-th element in the scanning direction The sampling point delay between the upper and reference array elements and the m-th array element in the scanning direction The candidate integer sampling delay is used to obtain the m-th array element in the scanning direction. The corresponding candidate score sampling delay is expressed as shown in equation (14).

[0075] (14);

[0076] in, This indicates that the m-th array element is in the scanning direction. The corresponding candidate score sampling delay.

[0077] It can be determined that the m-th array element is in the scanning direction The corresponding fractional delay filter is used to determine the position of the m-th array element in the scanning direction based on this fractional delay filter. The corresponding candidate fractional delay coefficients are obtained. This leads to the correspondence between the candidate fractional delay coefficients of the fractional delay filters corresponding to each of the multiple array elements in any scanning direction and the multiple array elements and scanning directions. This correspondence is then stored in a specific format to obtain a fractional delay coefficient library.

[0078] According to an embodiment of the present invention, determining the candidate fractional time delay coefficients in a fractional time delay filter corresponding to each of the array elements in any of the scanning directions based on the candidate fractional sampling delays of each of the array elements in any of the scanning directions may include: obtaining a frequency domain prototype sequence of the fractional time delay filter based on the sampling frequency of the positioning array within the target frequency band; converting the frequency domain prototype sequence of the fractional time delay filter into a time domain prototype sequence of the fractional time delay filter; performing interpolation processing on each tap index in the tap index set of the time domain prototype sequence of the fractional time delay filter to obtain an interpolation polynomial corresponding to each tap index of the fractional time delay filter; and determining the candidate fractional time delay coefficients in the fractional time delay filter corresponding to each of the array elements in any of the scanning directions based on the candidate fractional sampling delays of each of the array elements in any of the scanning directions and the interpolation polynomials corresponding to each tap index of the fractional time delay filter.

[0079] A frequency domain prototype sequence of a fractional time delay filter can be established in the frequency domain. The frequency domain prototype sequence can represent the ideal passband distribution of the fractional time delay filter on the discrete frequency axis. The frequency domain prototype sequence can serve as the basis for generating a subsequent time domain prototype sequence.

[0080] The frequency domain of a fractional time delay filter is discrete. Let N be the number of points in the discrete frequency domain; then the frequency resolution of the fractional time delay filter is... .

[0081] The target frequency band can be the frequency band of the sound source to be located, and can be determined according to the type of sound source. For example, if the type of sound source to be located is human voice, then the target frequency band can be 300Hz~3400Hz.

[0082] The low cutoff frequency within the target frequency band can be expressed as The high cutoff frequency can be expressed as ,and, By mapping the low and high cutoff frequencies to the frequency domain indices of the discrete frequency domain of the fractional time delay filter, we can obtain... , Then the number of frequency points corresponding to the effective passband width is M2=M H -M1, where M1 represents the frequency domain index of the discrete frequency point corresponding to the low cutoff frequency, M H M1 represents the frequency domain index of the discrete frequency point corresponding to the high cutoff frequency, and M2 represents the passband frequency width.

[0083] Based on this, a frequency domain prototype sequence of a fractional time delay filter with N frequency domain points is constructed. The frequency domain prototype sequence takes values ​​of 1 within the target passband and 0 outside the passband. In accordance with the requirement of conjugate symmetry of the spectrum of the real sequence of discrete Fourier transform, passband frequency points are symmetrically arranged at corresponding positions of positive and negative frequencies, thereby ensuring that the time domain prototype sequence obtained by the subsequent inverse transform is a real number sequence.

[0084] When the frequency domain index M1=0 corresponding to the low cutoff frequency, the frequency domain prototype sequence is a special low-pass symmetrical structure containing a DC component, which can be expressed as follows: M2 frequency points with an amplitude of 1 are continuously set at the low frequency end, zero frequency points are set in the middle frequency band, and M2-1 frequency points with an amplitude of 1 are set at the high frequency mirror position. At this time, the frequency domain prototype sequence can be expressed as follows (15).

[0085] (15);

[0086] in, This represents the frequency domain prototype sequence.

[0087] When the frequency domain index M1 > 0 corresponding to the low cutoff frequency, the frequency domain prototype sequence is a general bandpass structure. That is, first, M1 zero-value frequency points are set in the low-frequency starting segment, then M2 frequency points with an amplitude of 1 are set in the positive frequency passband region, zero-value frequency points are set in the middle, and M2 frequency points with an amplitude of 1 are set in the mirror region corresponding to the negative frequency. Finally, M1-1 zero-value frequency points are added to form a bandpass frequency domain template with a symmetrical distribution about the positive and negative frequencies. At this time, the frequency domain prototype sequence can be expressed as follows (16).

[0088] (16).

[0089] The frequency domain prototype sequence obtained by equation (16) can be used to construct the time domain prototype sequence of the fractional time delay filter. The frequency domain prototype sequence is determined based on the sampling frequency of the array element, the target frequency band, and the number of frequency points. It can constrain the main response of the subsequent fractional time delay filter within the target frequency band, providing a frequency band selection basis for time delay compensation in broadband sound source localization.

[0090] The time-domain prototype sequence of the fractional time delay filter can be obtained by performing an inverse discrete Fourier transform on the time-domain prototype sequence of the fractional time delay filter, as shown in equation (17).

[0091] (17);

[0092] in, This represents the time-domain prototype sequence of a fractional delay filter. This represents the inverse discrete Fourier transform.

[0093] To reduce the ringing effect and sidelobe rise caused by frequency domain truncation, a window function can be used to smooth the time-domain prototype sequence. A Hanning window of length N is chosen. And define a sequence of all 1s of length N. The constructed smooth window function is as follows (18).

[0094] (18);

[0095] in, This represents the smoothing window function. Hanning window With all-1 sequences Linear convolution.

[0096] Smooth window function Amplitude normalization is performed, and the time-domain prototype sequence is rearranged into an expansion sequence of length 2N-1 that is symmetric about the vicinity of zero. The sequence is then multiplied point by point with the smoothing window function to obtain the processed time-domain prototype sequence, as shown in equation (19).

[0097] (19);

[0098] in, This represents the processed time-domain prototype sequence.

[0099] The processed time-domain prototype sequence can be used to continuously approximate any fractional time delay. Since each sampling node is an integer, the processed time-domain prototype sequence can be interpolated based on the integer sampling nodes. For example, cubic spline interpolation can be performed on each integer sampling node to obtain a piecewise continuous function. The sampling nodes for integers are as follows (20).

[0100] (20);

[0101] in, This represents the sampling node for the nth integer.

[0102] Over the interval of adjacent integer sampling nodes, the result of cubic spline interpolation can be expressed as a cubic polynomial. Let the cubic polynomial corresponding to the q-th interval be as follows (21).

[0103] (twenty one);

[0104] in, This represents the cubic polynomial corresponding to the q-th interval. ...

[0105] The tap indices in the tap index set correspond to the sampling nodes of integers. Cubic spline interpolation is performed on the intervals of all integer sampling nodes to obtain the coefficient matrix of the cubic polynomial corresponding to each interval. As shown in equation (22).

[0106] (twenty two);

[0107] in, These are the coefficients of the cubic polynomial corresponding to the first interval. These are the coefficients of the cubic polynomial corresponding to the second interval. These are the coefficients of the cubic polynomial corresponding to the 2N-2th interval.

[0108] The set of tap indices of the time-domain prototype sequence of the fractional delay filter can be the set of integer sampling nodes after removing the two boundaries, and the set of tap indices k is as follows (23).

[0109] (twenty three).

[0110] The length L of the fractional delay filter is then 2N-3. As an example, the sampling frequency of the positioning array... The frequency can be set to 16000Hz, the speed of sound c can be set to 343m / s, and the low cutoff frequency. It can be 300Hz, with a high cutoff frequency. The frequency can be 3400Hz, and the number of points N in the frequency domain can be 65. Therefore, the length L of the fractional delay filter is 127.

[0111] Cubic polynomials can be extended to higher orders to reduce passband errors. Continuous fitting can be replaced by other piecewise polynomial fitting methods besides cubic spline interpolation, but the consistency of group time delay within the frequency band of the sound source to be located should be maintained.

[0112] The interpolation polynomial corresponding to the tap index in the tap index set can be a cubic polynomial. Then the interpolation polynomial at the k-th tap index position is as follows (24).

[0113] (twenty four);

[0114] in, This represents the interpolation polynomial at the k-th tap index position. The coefficients of the interpolation polynomial at the k-th tap index position are given.

[0115] It can make ,in, The fractional sampling delay can be expressed as follows (25).

[0116] (25);

[0117] in, This indicates that the k-th tap index has a fractional sampling delay of . The interpolation polynomial of the fractional delay filter under the given condition, These represent the k-th tap index with a fractional sampling delay of 1. The coefficients of the corresponding interpolation polynomials under the given conditions.

[0118] Arrange the coefficients of the interpolation polynomials corresponding to each tap index in the tap index set in order of the tap indexes to form a unified coefficient matrix. As shown in equation (26).

[0119] (26);

[0120] in, This represents the coefficients of the interpolation polynomial corresponding to the (N+2)th tap index. This represents the coefficients of the interpolation polynomial corresponding to the (-N+3)th tap index. This represents the coefficients of the interpolation polynomial corresponding to the (N-2)th tap index.

[0121] coefficient matrix It is determined solely by the results of the frequency domain prototype sequence and the interpolation polynomial, so it only needs to be calculated once in the offline stage.

[0122] Delay for sampling any fraction We can construct a fractional time delay power vector as shown in equation (27).

[0123] (27);

[0124] in, Indicates fractional sampling delay The corresponding fractional delay power vector.

[0125] The tap vector of the fractional delay filter corresponding to the fractional sampling delay is as follows (28).

[0126] (28);

[0127] in, This represents the tap vector of the fractional delay filter corresponding to the fractional sampling delay.

[0128] That is, the interpolation polynomial corresponding to the k-th tap index is .

[0129] The m-th element is in the scanning direction Corresponding candidate score sampling delay Substituting the coefficients into the interpolation polynomial expression, we can obtain the corresponding candidate fractional time delay coefficients, as shown in equation (29).

[0130] (29);

[0131] in, This indicates that the m-th array element is in the scanning direction. The corresponding candidate score delay coefficient, .

[0132] The candidate fractional delay coefficients for each array element in each scanning direction can be stored according to "element index – azimuth index – elevation index" to obtain a fractional delay coefficient library. Simultaneously, the candidate integer sampling delays for each array element in all scanning directions are stored according to "element index – azimuth index – elevation index", resulting in an integer delay library. .

[0133] Therefore, this embodiment of the invention can generate taps for a fractional time delay filter by extracting the coefficient matrix of the fractional time delay filter and then substituting it with the fractional sampling delay. This preserves the accuracy of fractional time delay compensation and facilitates the subsequent batch generation of offline filter coefficient libraries for different scanning directions and array elements. Both the fractional time delay coefficient library and the integer time delay library are processed offline, need to be generated only once, and can be directly called upon when performing online sound source localization.

[0134] Figure 2A flowchart of a sound source localization method based on offline library construction using array element coordinates according to an embodiment of the present invention is shown.

[0135] like Figure 2 As shown, the method of this embodiment is applied to a positioning array, which includes multiple array elements, including a reference array element, and includes operations S210 to S230.

[0136] In operation S210, for any scanning direction within the scanning range of multiple array elements: acquire the scanning signal corresponding to each of the multiple array elements in the scanning direction; call the fractional delay coefficient library and integer delay library corresponding to each of the multiple array elements; determine the fractional delay coefficient and integer sampling delay corresponding to each of the multiple array elements in the scanning direction from the fractional delay coefficient library and integer delay library respectively; perform time delay compensation on the scanning signal according to the fractional delay coefficient and integer sampling delay corresponding to each of the multiple array elements in the scanning direction to obtain the scanning compensation signal corresponding to each of the multiple array elements in the scanning direction.

[0137] The fractional delay coefficient library and the integer delay library are established based on the relative positions between the array elements and the reference array elements.

[0138] In operation S220, the signal spatial spectrum corresponding to any scanning direction is obtained based on the scanning compensation signal corresponding to each of the multiple array elements in any scanning direction.

[0139] In operation S230, the target direction of the sound source to be located within the scanning range is determined based on the signal spatial spectrum corresponding to any scanning direction.

[0140] The multiple elements in the positioning array can be sensors capable of converting sound signals into electrical signals, such as microphones. Because the multiple elements in the positioning array are located at different positions, the reception time of signals at the same moment will differ during signal scanning; therefore, time delay compensation is required for the scanning signal.

[0141] Multiple array elements can each receive signals in any scanning direction within the scanning range to obtain scanning signals.

[0142] Based on the array element and scanning direction corresponding to the scanning signal, the corresponding candidate fractional delay coefficient can be determined from the fractional delay coefficient library as the fractional delay coefficient for each of the multiple array elements in the scanning direction, and the candidate integer sampling delay can be determined from the integer delay library as the integer sampling delay.

[0143] The scanning signal can be compensated for the integer part of the time delay based on the integer sampling delay, and the scanning signal can be compensated for the fractional part of the time delay based on the fractional time delay coefficient. The scanning compensation signal corresponding to each of the multiple array elements in any scanning direction.

[0144] A signal spatial spectrum corresponding to a scanning direction can be constructed based on the scanning compensation signals corresponding to each of the multiple array elements in any scanning direction. In the signal spatial spectrum, the magnitude of energy can characterize the intensity of sound. The scanning direction corresponding to the position with the maximum energy value in the spatial spectrum can be determined as the target direction of the sound source to be located within the scanning range.

[0145] According to embodiments of the present invention, a fractional delay coefficient library and an integer delay library are pre-established based on the relative positions of multiple array elements in the positioning array with a reference array element. After acquiring the scanning signal, the corresponding fractional delay coefficient and integer sampling delay can be directly obtained by calling the fractional delay coefficient library and the integer delay library, thereby directly performing time delay compensation on the scanning signal to obtain a scan-compensated signal. Furthermore, a signal spatial spectrum is established, and the target direction of the sound source to be located is determined based on the signal spatial spectrum. Since the relatively complex fractional delay coefficient and integer sampling delay process is completed offline, there is no need to recalculate the fractional delay coefficient and integer sampling delay in the online stage; only the fractional delay coefficient library and the integer delay library need to be called, eliminating the need for online generation and reducing the online occupation of computing resources. Furthermore, using integer delay to achieve coarse compensation of the scanning signal and using fractional delay coefficients to achieve fine compensation of the scanning signal allows for more precise coherent alignment of signals from multiple array element channels in the target direction.

[0146] According to an embodiment of the present invention, performing time delay compensation on the scanning signal based on the fractional time delay coefficients and integer sampling delays corresponding to each of the multiple array elements in the scanning direction to obtain a scanning compensation signal corresponding to each of the multiple array elements in the scanning direction includes: performing fractional time delay compensation on the scanning signal based on the fractional time delay coefficients corresponding to each of the multiple array elements in the scanning direction to obtain a fractional time delay compensation signal corresponding to each of the multiple array elements in the scanning direction; and performing integer time delay compensation on the fractional time delay compensation signal based on the integer sampling delays corresponding to each of the multiple array elements in the scanning direction to obtain a scanning compensation signal corresponding to each of the multiple array elements in the scanning direction.

[0147] The scanning signal of the m-th array element can be expressed as follows (30).

[0148] (30);

[0149] in, Let v represent the scan signal of the m-th array element, and v represent the discrete sampling point index of the scan signal. The frame length of the scan signal.

[0150] For any scanning direction within the scanning range Fractional time delay coefficients can be used First, fractional time delay compensation is performed on the scanning signal. Fractional time delay compensation can be obtained by linearly convolving the fractional time delay coefficient with the scanning signal, as shown in the following formula (31).

[0151] (31);

[0152] in, This represents the fractional delay compensation signal. This represents the result of the linear convolution of the fractional time delay coefficient and the scanning signal.

[0153] The integer sampling delay is then used to perform integer delay compensation on the fractional delay compensation signal to obtain the scan compensation signal.

[0154] According to an embodiment of the present invention, fractional delay compensation is performed on the scanning signal based on the fractional delay coefficients corresponding to each of the multiple array elements in the scanning direction to obtain fractional delay compensation signals corresponding to each of the multiple array elements in the scanning direction. This includes: performing zero-boundary expansion processing on the scanning signal based on the fractional delay coefficients corresponding to each of the multiple array elements in the scanning direction to obtain an expanded scanning signal; and performing linear convolution between the fractional delay coefficients corresponding to each of the multiple array elements in the scanning direction and the expanded scanning signal to obtain fractional delay compensation signals corresponding to each of the multiple array elements in the scanning direction.

[0155] When the length of the tap index corresponding to the fractional delay coefficient is greater than the length of the scan signal, the scan signal can be zero-boundary extended. That is, the value at the corresponding position where the tap index exceeds the range of the scan signal can be 0, thus obtaining an extended scan signal. Based on this extended scan signal, a linear convolution is performed with the fractional delay coefficient to obtain the fractional delay compensation signal.

[0156] According to an embodiment of the present invention, integer delay compensation is performed on the fractional delay compensation signal based on the integer sampling delay corresponding to each of the multiple array elements in the scanning direction to obtain the scan compensation signal corresponding to each of the multiple array elements in the scanning direction. This includes: performing integer delay compensation on the fractional delay compensation signal based on the frame length of the scan signal according to the integer sampling delay corresponding to each of the multiple array elements in the scanning direction to obtain the scan compensation signal corresponding to each of the multiple array elements in the scanning direction, wherein the frame length of the scan compensation signal is the same as the frame length of the scan signal.

[0157] Based on the frame length of the scan signal, the fractional delay compensation signal can be segmented to compensate, so that the frame length of the output scan compensation signal is the same as that of the scan signal.

[0158] The number of tap indices of the fractional delay filter can be L=2N-3, and its corresponding group delay is as follows (32).

[0159] (32);

[0160] in, This represents the group delay of the fractional delay filter.

[0161] The fractional delay compensation signal can be segmented and compensated based on group delay and integer sampling delay, so that the frame length of the output scan compensation signal is the same as that of the scan signal. The m-th array element is in the scanning direction... The scanning compensation signal is as follows (33).

[0162] (33);

[0163] in, This indicates that the m-th array element is in the scanning direction. The scan compensation signal.

[0164] In the process of integer time delay compensation, when the starting or ending position of the segment exceeds the effective range of the convolution output, zeros can be pre-padded at the input or at the end after convolution to ensure that the frame length of the final scan compensation signal is equal to the frame length of the scan signal.

[0165] According to an embodiment of the present invention, by performing fractional delay compensation followed by integer delay compensation on the scanning signal, it is more convenient to extend the scanning signal in fractional delay compensation. In integer delay compensation, the signal is truncated based on the frame length of the scanning signal, so that the frame length of the scanning compensation signal is the same as the frame length of the scanning signal, thereby ensuring that the length of the output signal is consistent with the length of the input scanning signal.

[0166] According to an embodiment of the present invention, obtaining a signal spatial spectrum corresponding to any scanning direction based on the scanning compensation signals corresponding to each of the plurality of array elements in any scanning direction includes: averaging the scanning compensation signals corresponding to each of the plurality of array elements in any scanning direction to obtain an average scanning signal corresponding to any scanning direction; and obtaining a signal spatial spectrum corresponding to any scanning direction based on the average scanning signal corresponding to any scanning direction.

[0167] The scanning signals of all array elements in any scanning direction can be averaged based on the number of array elements to obtain the average scanning signal corresponding to any scanning direction, as shown in equation (34).

[0168] (34);

[0169] in, Indicates the scanning direction The corresponding average scan signal.

[0170] The signal spatial spectrum is constructed using mean square energy, as shown in equation (35).

[0171] (35);

[0172] in, Indicates the scanning direction The corresponding signal spatial spectrum.

[0173] The signal spatial spectrum corresponding to all scanning directions within the scanning range is traversed to obtain a two-dimensional spatial spectrum matrix. The scanning direction where the maximum value is located is determined within the scanning range, thus determining the target direction of the sound source to be located within the scanning range. As shown in equation (36).

[0174] (36);

[0175] in, Let be the azimuth angle of the sound source to be located within the target direction of the scanning range. The pitch angle of the sound source to be located in the target direction within the scanning range.

[0176] According to an embodiment of the present invention, the output signal-to-noise ratio (SNR) can be improved by averaging the scan compensation signals corresponding to each of the multiple array elements in any scanning direction. This is because if the scan compensation signals are not averaged, the signal amplitude will vary with the number of array elements, which is not conducive to energy comparison under different positioning array sizes and different scanning directions. However, if the scan compensation signals are not averaged and only the signal corresponding to a single array element is taken, the spatial gain and directivity of the positioning array will be lost. Since the signals corresponding to the target direction of the sound source to be located will achieve coherent superposition after compensation, while noise is incoherently superimposed, the output SNR can be improved by averaging the scan compensation signals.

[0177] The sound source localization method based on offline library construction of array element coordinates proposed in this invention aims to locate the azimuth and elevation angles of the sound source to be located relative to a reference array element. It does not limit the specific topology between multiple array elements in the localization array; as long as the three-dimensional spatial coordinates of each array element in the localization array can be obtained, a fractional delay coefficient library and an integer delay library corresponding to the actual localization array can be established offline based on the actual localization array parameters. The localization of the sound source to be located can then be achieved online by calling these libraries.

[0178] In this embodiment of the invention, a reference array element is used as a unified time delay reference for the positioning array. The sampling point delay is determined based on the position difference of each array element relative to the reference array element, without having to consider the positional relationship between the array elements.

[0179] For broadband acoustic source signals, the propagation delay between different array elements is usually not an integer multiple of the sampling period. If only integer sampling point shifting is used for compensation, residual delay errors are easily introduced, leading to a decrease in beam output focusing capability. In this embodiment of the invention, the sampling point delay is decomposed into two parts: integer delay compensation and fractional delay compensation, which improves broadband alignment accuracy. Through integer delay compensation and fractional delay compensation, coherent alignment of multi-channel broadband signals in the target direction can be achieved more precisely, thereby improving the peak focusing effect of the spatial spectrum and enhancing the stability of direction estimation.

[0180] The two-dimensional scanning grid constructed using azimuth and elevation angles in this invention embodiment enables three-dimensional directional search, making it suitable for sound source localization. The sound source localization method of this invention embodiment can also be extended to near-field localization, different frequency band constraint designs, and different scanning strategies. In near-field sound source localization scenarios, the scanning direction can be replaced by spatial coordinates instead of azimuth and elevation angles, and a spherical wave propagation model can be used to calculate the sampling time difference between multiple array elements and a reference array element.

[0181] This invention generates a fractional delay coefficient library and an integer delay library during the offline phase, providing a clear, fixed, and traceable calling basis for the online localization process. Compared to methods that dynamically generate parameters online, the sound source localization method of this invention allows the localization array to maintain consistency in the localization process across different scenarios.

[0182] Figure 3 A schematic diagram of a positioning array according to an embodiment of the present invention is shown.

[0183] like Figure 3 As shown, the positioning array comprises 64 elements arranged in a spiral pattern. The black pentagram shape represents the reference element, and the black dots represent the other elements besides the reference element. Based on this positioning array, fractional delay coefficient libraries and integer delay libraries are established with a step angle of 1° for both azimuth and elevation angles. The azimuth angle ranges from -180° to less than 180°, and the elevation angle ranges from greater than 0° to less than 90°. Positioning tests are performed on sound source signals ranging from 1500 to 2000 Hz, positioned at an azimuth angle of 90° and an elevation angle of 45° within the positioning array.

[0184] Figure 4 An embodiment of the present invention is shown. Figure 3 The diagram shows a schematic of the spatial spectrum of the signal obtained by the positioning array. Figure 5 An embodiment of the present invention is shown. Figure 4 An azimuth slice of the signal spatial spectrum. Figure 6 An embodiment of the present invention is shown. Figure 4 Elevation slice of the signal spatial spectrum.

[0185] like Figure 4 As shown, based on Figure 3 The positioning array shown performs sound source localization, obtaining a signal spatial spectrum. Different colors in the signal spatial spectrum represent the relative power of the positioning array's output, measured in dB. This relative power is obtained by normalizing the average power of the positioning array's output and then performing a logarithmic transformation. The positions in the signal spatial spectrum corresponding to azimuth angles of 90° and elevation angles of 45° exhibit higher focusing ability. For example... Figure 5 and Figure 6 As shown, in the azimuth and elevation slices of the signal spatial spectrum, the main lobe is narrow and the side lobe level is low, which can accurately determine the target direction of the sound source.

[0186] Figure 7 A block diagram of a sound source localization device based on offline database construction of array element coordinates according to an embodiment of the present invention is shown.

[0187] like Figure 7 As shown, the sound source localization device 700 based on offline database construction of array element coordinates is applied to a localization array. The localization array includes multiple array elements, among which a reference array element is included. The sound source localization device 700 includes a first obtaining module 710, a second obtaining module 720, and a determining module 730.

[0188] The first obtaining module 710 is used for any scanning direction within the scanning range of multiple array elements: acquiring the scanning signal corresponding to each of the multiple array elements in the scanning direction; calling the fractional delay coefficient library and integer delay library corresponding to each of the multiple array elements, wherein the fractional delay coefficient library and integer delay library are established based on the relative position between the array element and the reference array element; determining the fractional delay coefficient and integer sampling delay corresponding to each of the multiple array elements in the scanning direction from the fractional delay coefficient library and integer delay library respectively; and performing time delay compensation on the scanning signal according to the fractional delay coefficient and integer sampling delay corresponding to each of the multiple array elements in the scanning direction to obtain the scanning compensation signal corresponding to each of the multiple array elements in the scanning direction.

[0189] The second obtaining module 720 is used to obtain the signal spatial spectrum corresponding to any scanning direction based on the scanning compensation signal corresponding to each of the multiple array elements in any scanning direction.

[0190] The determination module 730 is used to determine the target direction of the sound source to be located within the scanning range based on the signal spatial spectrum corresponding to any scanning direction.

[0191] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present invention, or at least part of the functions of any one or more of them, can be implemented in a single module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present invention can be implemented by being divided into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present invention can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, and firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present invention can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0192] For example, any plurality of the first obtaining module 710, the second obtaining module 720, and the determining module 730 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of the present invention, at least one of the first obtaining module 710, the second obtaining module 720, and the determining module 730 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first obtaining module 710, the second obtaining module 720, and the determining module 730 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0193] Figure 8 A block diagram of an electronic device suitable for implementing the methods described above, according to an embodiment of the present invention, is shown. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0194] like Figure 8As shown, an electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0195] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.

[0196] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.

[0197] According to embodiments of the present invention, the method flow according to embodiments of the present invention can be implemented as a computer software program. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by processor 801, it performs the functions defined in the system of the embodiments of the present invention. According to embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0198] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0199] According to embodiments of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0200] For example, according to embodiments of the present invention, a computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 as described above.

[0201] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of the present invention. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of the present invention.

[0202] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this embodiment of the invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0203] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0204] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or pairings fall within the scope of this invention.

[0206] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

Claims

1. A sound source localization method based on offline library construction of array element coordinates, characterized in that, The sound source localization method is applied to a positioning array, wherein the positioning array includes multiple array elements, and the multiple array elements include a reference array element. For each of the multiple array elements in any scanning direction within the scanning range: Acquire the scanning signal corresponding to each of the plurality of array elements in the scanning direction; The fractional delay coefficient library and the integer delay library corresponding to each of the multiple array elements are invoked, wherein the fractional delay coefficient library and the integer delay library are established based on the relative position between the array element and the reference array element; Determine the fractional delay coefficient and integer sampling delay corresponding to each of the array elements in the scanning direction from the fractional delay coefficient library and the integer delay library corresponding to each of the array elements, respectively. Based on the fractional delay coefficient and the integer sampling delay corresponding to each of the array elements in the scanning direction, the scanning signal is time-delayed to obtain the scanning compensation signal corresponding to each of the array elements in the scanning direction. Based on the scan compensation signal corresponding to each of the plurality of array elements in any of the scan directions, the signal spatial spectrum corresponding to any of the scan directions is obtained; Based on the signal spatial spectrum corresponding to any of the scanning directions, determine the target direction of the sound source to be located within the scanning range; The fractional delay coefficient library is obtained through the following operations: Based on the sampling point delay between each of the array elements and the reference array element in any of the scanning directions and the candidate integer sampling delay of each of the array elements in any of the scanning directions, the candidate fractional sampling delay of each of the array elements in any of the scanning directions is determined; Based on the candidate fractional sampling delay of each of the multiple array elements in any of the scanning directions, the candidate fractional delay coefficients in the fractional delay filter corresponding to each of the multiple array elements in any of the scanning directions are determined; The fractional delay coefficient library is obtained based on the correspondence between the candidate fractional delay coefficients of the fractional delay filters corresponding to each of the array elements in any of the scanning directions and the array elements and the scanning directions. The step of determining the candidate fractional delay coefficients in the fractional delay filter corresponding to each of the multiple array elements in any of the scanning directions based on the candidate fractional sampling delays of each of the multiple array elements includes: Based on the sampling frequency of the positioning array within the target frequency band, the frequency domain prototype sequence of the fractional time delay filter is obtained; The frequency domain prototype sequence of the fractional delay filter is converted into the time domain prototype sequence of the fractional delay filter. Interpolation is performed on each tap index in the tap index set of the time-domain prototype sequence of the fractional delay filter to obtain the interpolation polynomial corresponding to each tap index of the fractional delay filter. Based on the interpolation polynomial corresponding to the candidate fractional sampling delay of each of the array elements in any of the scanning directions and the tap index of the fractional delay filter, the candidate fractional delay coefficients in the fractional delay filter corresponding to each of the array elements in any of the scanning directions are determined. Specifically, the coefficients of the interpolation polynomials corresponding to each tap index in the tap index set are arranged in the order of the tap indices to form a unified coefficient matrix. ,in, ; in, This represents the coefficients of the interpolation polynomial corresponding to the (N+2)th tap index. This represents the coefficients of the interpolation polynomial corresponding to the (-N+3)th tap index. This represents the coefficients of the interpolation polynomial corresponding to the (N-2)th tap index; The coefficient matrix is ​​determined solely by the results of the frequency domain prototype sequence and the interpolation polynomial, and only needs to be calculated once in the offline stage.

2. The sound source localization method according to claim 1, characterized in that, The step of performing time delay compensation on the scanning signal based on the fractional time delay coefficient and the integer sampling delay corresponding to each of the plurality of array elements in the scanning direction to obtain a scanning compensation signal corresponding to each of the plurality of array elements in the scanning direction includes: The scanning signal is subjected to fractional delay compensation based on the fractional delay coefficient corresponding to each of the array elements in the scanning direction, thereby obtaining a fractional delay compensation signal corresponding to each of the array elements in the scanning direction. The fractional delay compensation signal is compensated for integer delay based on the integer sampling delay corresponding to each of the array elements in the scanning direction, thereby obtaining the scanning compensation signal corresponding to each of the array elements in the scanning direction.

3. The sound source localization method according to claim 2, characterized in that, The step of obtaining the signal spatial spectrum corresponding to any one of the scanning directions based on the scanning compensation signals corresponding to each of the plurality of array elements in any one of the scanning directions includes: The average scan signal corresponding to each of the array elements in any of the scanning directions is obtained by averaging the scan compensation signals of each of the array elements in any of the scanning directions. Based on the average scan signal corresponding to any of the scan directions, the signal spatial spectrum corresponding to any of the scan directions is obtained.

4. The sound source localization method according to claim 2 or 3, characterized in that, The step of performing fractional time delay compensation on the scanning signal according to the fractional time delay coefficient corresponding to each of the plurality of array elements in the scanning direction to obtain a fractional time delay compensation signal corresponding to each of the plurality of array elements in the scanning direction includes: The scan signal is subjected to zero-boundary expansion processing based on the fractional time delay coefficients corresponding to each of the array elements in the scan direction to obtain an expanded scan signal; The fractional delay coefficients corresponding to each of the array elements in the scanning direction are linearly convolved with the extended scanning signal to obtain the fractional delay compensation signals corresponding to each of the array elements in the scanning direction.

5. The sound source localization method according to claim 2 or 3, characterized in that, The step of performing integer time delay compensation on the fractional time delay compensation signal based on the integer sampling delay corresponding to each of the plurality of array elements in the scanning direction to obtain the scanning compensation signal corresponding to each of the plurality of array elements in the scanning direction includes: Based on the frame length of the scanning signal, the fractional delay compensation signal is subjected to integer delay compensation according to the integer sampling delay corresponding to each of the plurality of array elements in the scanning direction, to obtain the scanning compensation signal corresponding to each of the plurality of array elements in the scanning direction, wherein the frame length of the scanning compensation signal is the same as the frame length of the scanning signal.

6. The sound source localization method according to any one of claims 1 to 3, characterized in that, The integer delay library is obtained through the following operations: Based on the individual array element coordinates of the plurality of array elements and the reference array element coordinates of the reference array element, determine the position difference between each of the plurality of array elements and the reference array element; Based on the positional differences between each of the array elements and the reference array element, the signal arrival time difference between each of the array elements and the reference array element in any of the scanning directions is determined; Based on the signal arrival time difference between each of the array elements and the reference array element in any of the scanning directions and the sampling frequency of the positioning array, the sampling point delay between each of the array elements and the reference array element in any of the scanning directions is determined; The sampling point delay between each of the plurality of array elements and the reference array element in any of the scanning directions is rounded to obtain the candidate integer sampling delay corresponding to each of the plurality of array elements in any of the scanning directions; The integer delay library is obtained based on the correspondence between the candidate integer sampling delay of each of the array elements in any of the scanning directions and the array elements and the scanning directions.

7. The sound source localization method according to any one of claims 1 to 3, characterized in that, The positioning array includes a two-dimensional planar array or a three-dimensional array.

8. The sound source localization method according to claim 1, characterized in that, The topology of the multiple array elements in the positioning array includes one of the following: double-ring array, multi-ring array, spiral array, sparse array, spherical array, or rectangular array.