A sound source orientation method and device based on a quadratic cross-correlation function, equipment and storage medium
By employing a sound source localization method based on a quadratic cross-correlation function, microphone array recording is used to perform generalized cross-correlation phase transformation, a quadratic cross-correlation function is constructed, time delay estimates are fitted, and outliers are eliminated. This solves the problem of unstable sound source localization in existing technologies and achieves more efficient and reliable sound source localization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MALANSHAN AUDIO & VIDEO LABORATORY
- Filing Date
- 2026-03-02
- Publication Date
- 2026-05-29
AI Technical Summary
Existing sound source localization techniques are unstable due to factors such as reverberation, multipath reflection, and environmental noise. Furthermore, single cross-correlation is easily affected by out-of-band noise, which affects the reliability of azimuth calculation.
A sound source localization method based on the quadratic cross-correlation function is adopted. By driving the microphone array to record and performing generalized cross-correlation phase transformation, a quadratic cross-correlation function is constructed, the time delay estimate is fitted, and the inverse cosine function is used to determine the sound source incident angle. Outliers are eliminated to improve the localization accuracy.
It improves the efficiency and stability of sound source localization, thus enhancing the user experience.
Smart Images

Figure CN122109981A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of acoustics, and in particular to a method, apparatus, device, and storage medium for sound source localization based on a quadratic cross-correlation function. Background Technology
[0002] Currently, sound source localization, also known as sound source orientation estimation or angle of arrival (DOA) estimation, is a crucial foundational capability in acoustic measurement and speech interaction, conference audio pickup, and machine hearing. Existing technologies typically employ microphone arrays to acquire multi-channel signals and then calculate the sound source's incident direction based on the time difference of arrival (TDOA) or spatial spectrum estimation methods (such as phase difference-based estimation, beamforming, and subspace-based methods).
[0003] In practical engineering, due to factors such as reverberation, multipath reflection, environmental noise, uneven frequency response of loudspeakers and microphones, and limited array size, the time delay estimation between channels is prone to peak ambiguity or peak drift, resulting in unstable orientation results. When using broadband excitations such as frequency sweep, if only the peak position of a single cross-correlation is relied upon as the time delay estimate, it may still be affected by out-of-band noise and reverberation, thus affecting the reliability of azimuth angle calculation.
[0004] As can be seen from the above, how to improve the efficiency of sound source localization in the process of sound source localization based on the quadratic cross-correlation function is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a sound source localization method, apparatus, device, and storage medium based on a quadratic cross-correlation function, which can improve the efficiency of sound source localization in the process of sound source localization based on a quadratic cross-correlation function. The specific solution is as follows: Firstly, this application provides a sound source localization method based on a quadratic cross-correlation function, comprising: A linear array including several microphones is driven to synchronously record and digitize a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency, thereby obtaining a corresponding digital recording signal. Then, a generalized cross-correlation phase transformation is performed on each of the digital recording signals based on the broadband linear sweep frequency signal to obtain a first cross-correlation sequence including time delay information. Generalized cross-correlation is performed on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation functions. The peak position is located in each of the second cross-correlation functions to determine the target sampling point curve based on the peak position. Then, the target sampling point curve is fitted to obtain the time delay estimate. The target sampling point curve is a sampling point curve constructed based on a preset number of sampling points within a preset range near the peak position. Data points are constructed by using the index difference between every two channels as the x-axis and the estimated time delay as the y-axis. A current time delay point set is constructed based on each data point and the origin. Then, the current weight corresponding to each data point in the current time delay point set is determined, and each data point is fitted based on the current weight to obtain the current fitted line and the current line slope. Determine the residual of each data point relative to the current fitted line, and determine whether the residual with the largest value among the residuals is greater than a preset threshold. If it is greater, set the corresponding data point as an outlier and remove the outlier to obtain a new current time delay point set. Then, jump back to the step of determining the current weight of each data point in the current time delay point set until the current time delay point set meets the preset condition, and set the slope of the current line as the target line slope. The incident angle of the sound source loudspeaker relative to the normal direction of the linear array is determined by using the inverse cosine function and based on the slope of the target straight line, the preset sampling frequency, the microphone spacing and the preset sound speed, so as to determine the sound source orientation result based on the incident angle.
[0006] Optionally, the drive includes a linear array of several microphones and synchronously records and digitizes a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a corresponding digitized recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transform is performed on each of the digitized recording signals to obtain a first-order cross-correlation sequence including time delay information, including: A linear array comprising several microphones is constructed to drive the linear array and synchronously record a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a recording signal to be processed. Then, the recording signal to be processed is digitized to obtain a digital recording signal corresponding to each microphone channel. The broadband linear sweep frequency signal is set as the reference reference signal, and the reference reference signal and each of the digital recording signals are transformed to the frequency domain to obtain the corresponding frequency domain representations of the reference signal and the frequency domain representations of the recording signal. The frequency domain representations of each of the recorded signals are subjected to conjugate processing to obtain a conjugate processing result. The conjugate processing result is then multiplied with the frequency domain representation of the reference signal to obtain a frequency domain cross-correlation function. The amplitude corresponding to the frequency domain cross-correlation function is normalized to obtain a normalized result. Then, the normalized result is transformed to the time domain by inverse Fourier transform to obtain a first cross-correlation sequence including time delay information. The time delay information is used to characterize the time delay information of each microphone channel relative to the broadband linear sweep frequency signal. The peak position of the first cross-correlation sequence is used to characterize the relative time delay between the sound wave emitted by the sound source loudspeaker and the reference signal when the sound wave reaches the corresponding microphone.
[0007] Optionally, the step of performing generalized cross-correlation on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation functions, locating the peak position in each of the second cross-correlation functions, determining the target sampling point curve based on the peak position, and then fitting the target sampling point curve to obtain the time delay estimate includes: Perform generalized cross-correlation phase transformation on the first cross-correlation sequences corresponding to any two different channels to obtain the second cross-correlation function corresponding to each channel, and determine the corresponding peak position in each second cross-correlation function; The peak position is set as a reference, and a preset number of sampling points near the peak position are extracted based on the reference. Then, the target sampling point curve is determined based on each of the sampling points. A quadratic function fitting process is performed on each sampling point in the target sampling point curve to obtain the estimated time delay between the two channels.
[0008] Optionally, the step of constructing data points with the index difference between every two channels as the x-axis and the estimated delay value as the y-axis, and constructing a current delay point set based on each data point and the origin, includes: Determine the peak position corresponding to each of the channels, and determine the corresponding index estimate based on the peak position. Then determine the position index difference between each index estimate. The location index difference is set as the horizontal axis, and the corresponding time delay estimate is set as the vertical axis. Two-dimensional data points are constructed based on the horizontal and vertical axes, and then the two-dimensional data points are aggregated to obtain the aggregated result. The origin of the coordinate system is set as the reference point, and the current time delay point set is constructed based on the reference point and the aggregation result.
[0009] Optionally, determining the current weight corresponding to each data point in the current delay point set, and fitting each data point based on the current weight to obtain the current fitted line and the slope of the current line, includes: Determine the x-coordinate value corresponding to each data point in the current delay point set; If the horizontal coordinate value is zero, the weight of the corresponding data point is set to 1. If the horizontal coordinate value is not zero, the difference between the total number of microphones corresponding to the linear array and the horizontal coordinate value is determined, and the reciprocal of the difference is set as the weight corresponding to the corresponding data point. Based on each data point in the current time delay point set and its corresponding weight value, a weighted least squares line fitting is performed to obtain the current fitted line, and the slope of the current line corresponding to the current fitted line is determined.
[0010] Optionally, the step of determining the residual of each data point relative to the current fitted line, and judging whether the residual with the largest value among the residuals is greater than a preset threshold, if it is greater, then setting the corresponding data point as an outlier and removing the outlier to obtain a new current time delay point set, and then jumping back to the step of determining the current weight corresponding to each data point in the current time delay point set, until the current time delay point set meets the preset condition, and setting the current line slope as the target line slope, includes: Determine the residuals of each data point in the current time delay point set relative to the current fitted line, and determine the target residual with the largest value among all residuals. Then determine whether the target residual is greater than a preset residual threshold. If the target residual is greater than the preset residual threshold, the corresponding data point of the target residual is marked as an outlier, and the outlier is removed from the current delay point set. Then the current delay point set is updated to obtain a new current delay point set, so as to jump back to the step of determining the current weight of each data point in the current delay point set, until the target residual corresponding to the current delay point set is not greater than the residual threshold, or the number of outliers removed reaches the preset upper limit value, then the current straight line slope is set as the target straight line slope.
[0011] Optionally, the step of determining the incident angle of the sound source loudspeaker relative to the normal direction of the linear array using the inverse cosine function and based on the slope of the target straight line, the preset sampling frequency, the microphone spacing, and the preset sound velocity, so as to determine the sound source orientation result based on the incident angle, includes: A preset spacing between each adjacent microphone in the linear array is determined, and a first ratio is determined based on the slope of the target straight line and the preset sampling frequency. Then, the product result is determined based on the first ratio and the preset sound speed. A second ratio is determined based on the product result and the preset spacing. The incident angle is determined using the inverse cosine function and based on the negative of the second ratio. The directionality of the sound source loudspeaker relative to the linear array is determined based on the incident angle.
[0012] Secondly, this application provides a sound source direction finding device based on a quadratic cross-correlation function, comprising: A first-order cross-correlation sequence generation module is used to drive a linear array including several microphones and synchronously record and digitize a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a corresponding digital recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transformation is performed on each of the digital recording signals to obtain a first-order cross-correlation sequence including time delay information. The delay estimation module is used to perform generalized cross-correlation on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation functions, locate the peak position in each second cross-correlation function, determine the target sampling point curve based on the peak position, and then fit the target sampling point curve to obtain the delay estimation value. The current delay point set construction module is used to construct data points with the index difference between every two channels as the horizontal axis and the estimated delay value as the vertical axis, and construct the current delay point set based on each data point and the origin of the coordinate system. Then, it determines the current weight corresponding to each data point in the current delay point set, and fits each data point based on the current weight to obtain the current fitted line and the current line slope. The residual determination module is used to determine the residual of each data point relative to the current fitted line, and to determine whether the residual with the largest value among the residuals is greater than a preset threshold. If it is greater, the corresponding data point is set as an outlier and the outlier is removed to obtain a new current time delay point set. Then, the module jumps back to the step of determining the current weight of each data point in the current time delay point set until the current time delay point set meets the preset condition, and the slope of the current line is set as the slope of the target line. The sound source orientation result determination module is used to determine the incident angle of the sound source loudspeaker relative to the normal direction of the linear array using the inverse cosine function and based on the slope of the target straight line, the preset sampling frequency, the microphone spacing and the preset sound speed, so as to determine the sound source orientation result based on the incident angle.
[0013] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned sound source localization method based on a quadratic cross-correlation function.
[0014] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned sound source localization method based on a quadratic cross-correlation function.
[0015] As can be seen from the above, before performing sound source localization based on the quadratic cross-correlation function, this application needs to drive a linear array including several microphones and synchronously record and digitize the broadband linear sweep frequency signal played by the sound source speaker based on a preset sampling frequency to obtain the corresponding digitized recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transformation is performed on each digitized recording signal to obtain a first-order cross-correlation sequence including time delay information. A generalized cross-correlation is then performed on the first-order cross-correlation sequences corresponding to any two different channels to obtain the corresponding quadratic cross-correlation function. The peak position is located in each quadratic cross-correlation function to determine the target sampling point curve based on the peak position. Then, the target sampling point curve is fitted to obtain the time delay estimate. The index difference between the two channels is used as the x-axis, and the estimated time delay is used as the y-axis to form data points. A current time delay point set is constructed based on each data point and the origin. Then, the current weight corresponding to each data point in the current time delay point set is determined, and each data point is fitted based on the current weight to obtain the current fitted line and the current line slope. The residual of each data point relative to the current fitted line is determined, and it is judged whether the residual with the largest value is greater than a preset threshold. If it is greater, the corresponding data point is set as an outlier and removed, resulting in a new current time delay point set. The process then jumps back to the step of determining the weight corresponding to each data point in the current time delay point set until the current time delay point set meets the preset conditions, and the current line slope is set as the target line slope. The incident angle of the sound source speaker relative to the normal direction of the linear array is determined using the inverse cosine function and based on the target line slope, preset sampling frequency, microphone spacing, and preset sound velocity, so as to determine the sound source orientation result based on the incident angle.
[0016] Therefore, this application first needs to drive a linear array including several microphones to synchronously record and digitize a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency, obtaining the corresponding digitized recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transformation is performed on each digitized recording signal to obtain a first-order cross-correlation sequence including time delay information. Second, a generalized cross-correlation is performed on the first-order cross-correlation sequences corresponding to any two different channels to obtain the corresponding second-order cross-correlation functions. The peak position is located in each second-order cross-correlation function to determine the target sampling point curve based on the peak position. Then, the target sampling point curve is fitted to obtain the time delay estimate. Finally, data points are constructed with the index difference between each two channels as the abscissa and the time delay estimate as the ordinate, and the current time delay is constructed based on each data point and the origin. The system first sets a set of data points, then determines the current weights for each data point in the current delay point set, and fits the data points based on these weights to obtain the current fitted line and its slope. Next, it determines the residuals of each data point relative to the current fitted line and checks if the largest residual exceeds a preset threshold. If it does, the corresponding data point is designated as an outlier and removed, resulting in a new current delay point set. The process then returns to determining the weights for each data point in the current delay point set until the current delay point set meets the preset conditions, and the current line slope is set as the target line slope. Finally, using the inverse cosine function and based on the target line slope, preset sampling frequency, microphone spacing, and preset sound velocity, it determines the incident angle of the speaker relative to the normal direction of the linear array, and uses this incident angle to determine the sound source localization result. This improves the efficiency of sound source localization based on the quadratic cross-correlation function, thereby enhancing the user experience. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This is a flowchart of a sound source localization method based on a quadratic cross-correlation function disclosed in this application; Figure 2 This is a flowchart of a specific sound source localization method based on a quadratic cross-correlation function disclosed in this application; Figure 3 This is a schematic diagram illustrating a specific point set construction and line fitting process disclosed in this application; Figure 4This is a schematic diagram of a sound source directional device based on a quadratic cross-correlation function disclosed in this application; Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Currently, sound source localization is a crucial foundational capability in acoustic measurement and speech interaction, conference audio pickup, and machine hearing. In practical engineering, factors such as reverberation, multipath reflections, environmental noise, uneven frequency response of speakers and microphones, and limited array size can easily lead to peak ambiguity or peak drift in inter-channel time delay estimation, resulting in unstable localization results. When using broadband excitations such as frequency sweeping, relying solely on the peak position of a single cross-correlation as the time delay estimate can still be affected by out-of-band noise and reverberation, thus impacting the reliability of azimuth calculation. Therefore, this application provides a sound source localization method based on a quadratic cross-correlation function, which can improve the efficiency of sound source localization in quadratic cross-correlation function-based localization processes.
[0021] See Figure 1 As shown, this embodiment of the invention discloses a sound source localization method based on a quadratic cross-correlation function, comprising: Step S11: Drive a linear array including several microphones and synchronously record and digitize the broadband linear sweep frequency signal played by the sound source speaker based on a preset sampling frequency to obtain the corresponding digital recording signal. Then, perform a generalized cross-correlation phase transformation on each of the digital recording signals based on the broadband linear sweep frequency signal to obtain a first cross-correlation sequence including time delay information.
[0022] In this embodiment, the flowchart for sound source localization based on the quadratic cross-correlation function is as follows: Figure 2 As shown, firstly, sound source excitation and acquisition are required, that is, a linear sweep frequency signal is played using a sound source loudspeaker. This linear sweep frequency signal includes a start frequency, an end frequency, and a duration. Subsequently, multi-channel synchronous recording is performed simultaneously using a linear array consisting of N microphones, resulting in N channels of recorded signals. The spacing between adjacent microphones in the array is [missing information]. The sampling frequency is .
[0023] Furthermore, in this embodiment of the application, a wideband reference signal is used as a reference. The embodiment calculates a generalized cross-correlation (GCC-PHAT) once for each channel of the recording signal, resulting in N single-order cross-correlation sequences. .
[0024] One typical GCC-PHAT calculation method can be expressed as: ; in, For the frequency domain representation of the reference signal, The frequency domain representation of a recording for a specific channel, indicated by the superscript. Indicates conjugate. Indicates the inverse Fourier transform. This is used to prevent extremely small positive numbers with a denominator of zero.
[0025] It is worth mentioning that the goal of a single cross-correlation is to suppress environmental noise, while suppressing amplitude effects under broadband conditions and highlighting phase consistency corresponding to time delay.
[0026] Specifically, the process involves driving a linear array of microphones to synchronously record and digitize a broadband linear sweep signal played from a sound source speaker based on a preset sampling frequency, obtaining corresponding digitized recording signals. Then, a generalized cross-correlation phase transform is performed on each digitized recording signal based on the broadband linear sweep signal to obtain a first-order cross-correlation sequence including time delay information. This process may include: constructing a linear array of microphones to drive the linear array and synchronously record a broadband linear sweep signal played from a sound source speaker based on a preset sampling frequency, obtaining a recording signal to be processed; then digitizing the recording signal to be processed to obtain digitized recording signals corresponding to each microphone channel; setting the broadband linear sweep signal as a reference signal, and using the reference signal... The reference signal and each digitized recording signal are transformed to the frequency domain to obtain the corresponding frequency domain representations of the reference signal and the recording signal, respectively. The frequency domain representations of each recording signal are then conjugated to obtain the conjugation result, which is multiplied by the frequency domain representation of the reference signal to obtain the frequency domain cross-correlation function. The amplitude of the frequency domain cross-correlation function is normalized to obtain the normalized result, which is then subjected to an inverse Fourier transform to the time domain, resulting in a first-order cross-correlation sequence including time delay information. The time delay information is used to characterize the time delay of each microphone channel relative to the broadband linear sweep signal. The peak position of the first-order cross-correlation sequence is used to characterize the relative time delay between the sound wave emitted by the speaker and the reference signal when the sound wave reaches the corresponding microphone.
[0027] Step S12: Perform generalized cross-correlation on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation functions, locate the peak position in each second cross-correlation function, determine the target sampling point curve based on the peak position, and then fit the target sampling point curve to obtain the time delay estimate; the target sampling point curve is a sampling point curve constructed based on a preset number of sampling points within a preset range near the peak position.
[0028] In this embodiment, the present application requires the construction of a time delay point set based on quadratic cross-correlation. That is, the aim is to extract features that more robustly reflect the relative time delay relationship between channels from N primary cross-correlation sequences: any two channels of the primary cross-correlation results of the N channels... , ( A generalized cross-correlation function, known as quadratic cross-correlation, is performed, and then the peak index of the quadratic cross-correlation function is located. Furthermore, to improve accuracy, embodiments of this application obtain a sub-sampling precision fractional index estimate by performing curve fitting (such as quadratic function fitting) on sampling points near the peak. .
[0029] Specifically, generalized cross-correlation is performed on the primary cross-correlation sequences corresponding to any two different channels to obtain the corresponding secondary cross-correlation functions. The peak positions are located in each secondary cross-correlation function, and the target sampling point curve is determined based on the peak positions. Then, the target sampling point curve is fitted to obtain the time delay estimate. This can include: performing generalized cross-correlation phase transformation on the primary cross-correlation sequences corresponding to any two different channels to obtain the secondary cross-correlation functions corresponding to each channel, and determining the corresponding peak positions in each secondary cross-correlation function; setting the peak positions as the benchmark, and extracting a preset number of sampling points near the peak positions based on the benchmark, and then determining the target sampling point curve based on each sampling point; performing quadratic function fitting on each sampling point in the target sampling point curve to obtain the time delay estimate between the two channels.
[0030] Step S13: Construct data points with the index difference between every two channels as the horizontal axis and the estimated time delay as the vertical axis, and construct the current time delay point set based on each data point and the origin of the coordinate system. Then, determine the current weight corresponding to each data point in the current time delay point set, and fit each data point based on the current weight to obtain the current fitted line and the current line slope.
[0031] In this embodiment, data points are constructed using the index difference between every two channels as the horizontal axis and the delay estimate as the vertical axis. A current delay point set is then built based on each data point and the origin. This process may include: determining the peak position corresponding to each channel, determining the corresponding index estimate based on the peak position, and then determining the position index difference between each index estimate; setting the position index difference as the horizontal axis and the corresponding delay estimate as the vertical axis, and constructing two-dimensional data points based on the horizontal and vertical axes, and then aggregating the two-dimensional data points to obtain the aggregation result; setting the origin as the reference point, and constructing the current delay point set based on the reference point and the aggregation result.
[0032] In this embodiment, the point set {(x,y)} obtained in this application embodiment should be distributed on a straight line, and its slope is related to the direction angle of arrival of the sound source. To resist interference in the actual environment (such as erroneous peaks caused by noise and reverberation), this application embodiment needs to use an iterative weighted robust fitting strategy for processing: for the current point set Perform weighted least squares linear fitting The weighting function is set as follows: ; It is worth mentioning that the above weight design ensures that all channels with the same channel index difference... The points contribute equally to the total fit, thus offsetting the fitting bias that might result from a larger number of short-interval point pairs.
[0033] Specifically, determining the current weights corresponding to each data point in the current delay point set, and fitting each data point based on the current weights to obtain the current fitted line and the current line slope can include: determining the abscissa value corresponding to each data point in the current delay point set; if the abscissa value is zero, setting the weight of the corresponding data point to 1; if the abscissa value is not zero, determining the difference between the total number of microphones corresponding to the linear array and the abscissa value, and setting the reciprocal of the difference as the weight corresponding to the corresponding data point; performing weighted least squares line fitting based on each data point in the current delay point set and the corresponding weight value to obtain the current fitted line, and determining the current line slope corresponding to the current fitted line.
[0034] Step S14: Determine the residual of each data point relative to the current fitted line, and determine whether the residual with the largest value among the residuals is greater than a preset threshold. If it is greater, set the corresponding data point as an outlier and remove the outlier to obtain a new current time delay point set. Then, jump back to the step of determining the current weight of each data point in the current time delay point set until the current time delay point set meets the preset condition, and set the current line slope as the target line slope.
[0035] In this embodiment, the residual of each data point relative to the fitted straight line needs to be calculated: ; Then, record the maximum error. .
[0036] Among them, if If a point is found to be outlier, it is removed from the point set. The above steps are then repeated until a termination condition is met (e.g., the maximum residual is below a threshold). (or the number of points removed reaches a preset upper limit Z), and finally a robust estimate of the straight slope k is obtained.
[0037] It is worth mentioning that this step, through a weighted model and iterative elimination mechanism, significantly reduces the impact of random errors on the final direction estimation, thereby improving the system's stability. The corresponding point set construction and line fitting process flowcharts are shown below. Figure 3 As shown.
[0038] Specifically, the residuals of each data point relative to the current fitted line are determined, and it is determined whether the residual with the largest value among all residuals is greater than a preset threshold. If it is greater, the corresponding data point is set as an outlier and removed, resulting in a new current time delay point set. The process then jumps back to the step of determining the current weights corresponding to each data point in the current time delay point set, until the current time delay point set meets the preset conditions, and the slope of the current line is set as the slope of the target line. This may include: determining the residuals of each data point in the current time delay point set relative to the current fitted line, and determining the values of all residuals. The maximum target residual is determined, and then it is determined whether the target residual is greater than the preset residual threshold. If the target residual is greater than the preset residual threshold, the corresponding data point of the target residual is marked as an outlier, and the outlier is removed from the current delay point set. Then the current delay point set is updated to obtain a new current delay point set, so as to jump back to the step of determining the current weight of each data point in the current delay point set, until the target residual corresponding to the current delay point set is not greater than the residual threshold, or the number of outliers removed reaches the preset upper limit value, then the current line slope is set as the target line slope.
[0039] Step S15: Using the inverse cosine function and based on the slope of the target line, the preset sampling frequency, the microphone spacing, and the preset sound velocity, determine the incident angle of the sound source loudspeaker relative to the normal direction of the linear array, so as to determine the sound source orientation result based on the incident angle.
[0040] In this embodiment, the application requires fitting the slope and calculating the angle of incidence: that is, based on the slope of the fitted line... Calculate the angle of incidence And the corresponding calculation formula is: ; in, The speed of sound (e.g., 340 m / s), d is the sampling frequency of the signal, and d is the spacing between adjacent microphones in the microphone array.
[0041] Specifically, the incident angle of the sound source loudspeaker relative to the normal direction of the linear array is determined using the inverse cosine function and based on the slope of the target line, the preset sampling frequency, the microphone spacing, and the preset sound velocity. The sound source orientation result is then determined based on the incident angle. This can include: determining the preset spacing between each adjacent microphone in the linear array, and determining a first ratio based on the slope of the target line and the preset sampling frequency; then determining the product result based on the first ratio and the preset sound velocity; determining a second ratio based on the product result and the preset spacing; and using the inverse cosine function and based on the negative of the second ratio to determine the incident angle. The sound source loudspeaker orientation result relative to the linear array is then determined based on the incident angle.
[0042] As can be seen from the above, the embodiments of this application first need to drive a linear array including several microphones and synchronously record and digitize the broadband linear sweep frequency signal played by the sound source speaker based on a preset sampling frequency to obtain the corresponding digitized recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transformation is performed on each digitized recording signal to obtain a first-order cross-correlation sequence including time delay information. Next, a generalized cross-correlation is performed on the first-order cross-correlation sequences corresponding to any two different channels to obtain the corresponding second-order cross-correlation function. The peak position is located in each second-order cross-correlation function to determine the target sampling point curve based on the peak position. Then, the target sampling point curve is fitted to obtain the time delay estimate. Then, data points are constructed with the index difference between each two channels as the abscissa and the time delay estimate as the ordinate, and the current coordinate system is constructed based on each data point and the origin. The system first sets a time delay point set, then determines the current weight of each data point in the current time delay point set, and fits each data point based on the current weight to obtain the current fitted line and the current line slope. Next, it determines the residual of each data point relative to the current fitted line, and checks whether the largest residual is greater than a preset threshold. If it is, the corresponding data point is marked as an outlier and removed, resulting in a new current time delay point set. The system then jumps back to the step of determining the weights of each data point in the current time delay point set until the current time delay point set meets the preset conditions, and the current line slope is set as the target line slope. Finally, using the inverse cosine function and based on the target line slope, preset sampling frequency, microphone spacing, and preset sound velocity, it determines the incident angle of the sound source speaker relative to the normal direction of the linear array, and uses the incident angle to determine the sound source localization result. This improves the efficiency of sound source localization in the process based on the quadratic cross-correlation function, thereby enhancing the user experience.
[0043] Accordingly, see Figure 4As shown, this application also provides a sound source direction finding device based on a quadratic cross-correlation function, comprising: A first cross-correlation sequence generation module 11 is used to drive a linear array including several microphones and synchronously record and digitize a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a corresponding digital recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transformation is performed on each of the digital recording signals to obtain a first cross-correlation sequence including time delay information. The delay estimation module 12 is used to perform generalized cross-correlation on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation function, locate the peak position in each second cross-correlation function, determine the target sampling point curve based on the peak position, and then fit the target sampling point curve to obtain the delay estimation value. The current delay point set construction module 13 is used to construct data points with the index difference between every two channels as the horizontal axis and the estimated delay value as the vertical axis, and construct the current delay point set based on each data point and the origin of the coordinate system. Then, it determines the current weight corresponding to each data point in the current delay point set, and fits each data point based on the current weight to obtain the current fitted line and the current line slope. The residual determination module 14 is used to determine the residual of each data point relative to the current fitted line, and to determine whether the residual with the largest value among the residuals is greater than a preset threshold. If it is greater, the corresponding data point is set as an outlier and the outlier is removed to obtain a new current time delay point set. Then, the module jumps back to the step of determining the current weight of each data point in the current time delay point set until the current time delay point set meets the preset condition, and the slope of the current line is set as the slope of the target line. The sound source orientation result determination module 15 is used to determine the incident angle of the sound source loudspeaker relative to the normal direction of the linear array using the inverse cosine function and based on the slope of the target straight line, the preset sampling frequency, the microphone spacing and the preset sound speed, so as to determine the sound source orientation result based on the incident angle.
[0044] In some specific embodiments, the first cross-correlation sequence generation module 11 may specifically include: A linear array construction unit is used to construct a linear array including several microphones, drive the linear array and synchronously record a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a recording signal to be processed, and then digitize the recording signal to be processed to obtain a digital recording signal corresponding to each microphone channel. The recording signal conversion unit is used to set the broadband linear sweep frequency signal as a reference signal, and to convert the reference signal and each of the digital recording signals to the frequency domain to obtain the corresponding reference signal frequency domain representation and the recording signal frequency domain representation. The conjugate processing result determination unit is used to perform conjugate processing on the frequency domain representations of each of the recorded signals to obtain a conjugate processing result, and to perform a product operation between the conjugate processing result and the frequency domain representation of the reference signal to obtain a frequency domain cross-correlation function. A first cross-correlation sequence generation subunit is used to normalize the amplitude corresponding to the frequency domain cross-correlation function to obtain a normalized result. Then, the normalized result is transformed to the time domain by inverse Fourier transform to obtain a first cross-correlation sequence including time delay information. The time delay information is used to characterize the time delay information of each microphone channel relative to the broadband linear sweep signal. The peak position of the first cross-correlation sequence is used to characterize the relative time delay between the sound wave emitted by the sound source loudspeaker and the reference signal when it reaches the corresponding microphone.
[0045] In some specific embodiments, the delay estimation value determination module 12 may specifically include: The cross-correlation phase transformation unit is used to perform generalized cross-correlation phase transformation on the first cross-correlation sequence corresponding to any two different channels to obtain the second cross-correlation function corresponding to each channel, and to determine the corresponding peak position in each second cross-correlation function. The sampling point extraction unit is used to set the peak position as a reference, extract a preset number of sampling points near the peak position based on the reference, and then determine the target sampling point curve based on each of the sampling points; The time delay estimation value determination subunit is used to perform quadratic function fitting on each of the sampling points in the target sampling point curve to obtain the time delay estimate between the two channels.
[0046] In some specific embodiments, the current delay point set construction module 13 may specifically include: An index estimation value determination unit is used to determine the peak position corresponding to each of the channels, determine the corresponding index estimation value based on the peak position, and then determine the position index difference between each index estimation value; The aggregation result generation unit is used to set the location index difference as the horizontal axis and the corresponding time delay estimate as the vertical axis, and construct two-dimensional data points based on the horizontal axis and the vertical axis, and then aggregate the two-dimensional data points to obtain the aggregation result; The current delay point set construction sub-unit is used to set the origin of the coordinate system as the reference point and construct the current delay point set based on the reference point and the aggregation result.
[0047] In some specific embodiments, the current delay point set construction module 13 may specifically include: The horizontal coordinate value determination unit is used to determine the horizontal coordinate value corresponding to each data point in the current delay point set; The weight determination unit is used to set the weight of the corresponding data point to 1 if the horizontal coordinate value is zero, and to determine the difference between the total number of microphones corresponding to the linear array and the horizontal coordinate value if the horizontal coordinate value is not zero, and to set the reciprocal of the difference as the weight corresponding to the corresponding data point. The current line slope determination unit is used to perform weighted least squares line fitting based on each data point in the current time delay point set and its corresponding weight value to obtain the current fitted line and determine the current line slope corresponding to the current fitted line.
[0048] In some specific embodiments, the residual determination module 14 may specifically include: The target residual determination unit is used to determine the residual of each data point in the current time delay point set relative to the current fitted line, and to determine the target residual with the largest value among all residuals, and then to determine whether the target residual is greater than a preset residual threshold. The straight line slope determination unit is used to mark the corresponding data points of the target residual as outliers if the target residual is greater than the preset residual threshold, remove the outliers from the current delay point set, update the current delay point set to obtain a new current delay point set, and then jump back to the step of determining the current weights corresponding to each data point in the current delay point set, until the target residual corresponding to the current delay point set is not greater than the residual threshold, or the number of outliers removed reaches a preset upper limit value, and then set the current straight line slope as the target straight line slope.
[0049] In some specific embodiments, the sound source localization result determination module 15 may specifically include: The product result generation unit is used to determine the preset spacing between each adjacent microphone in the linear array, and to determine a first ratio based on the slope of the target straight line and the preset sampling frequency, and then to determine the product result based on the first ratio and the preset sound speed; The sound source orientation result determination subunit is used to determine a second ratio based on the product result and the preset distance, and to determine the incident angle using the inverse cosine function and based on the opposite of the second ratio, so as to determine the orientation result of the sound source loudspeaker relative to the linear array based on the incident angle.
[0050] Furthermore, embodiments of this application also disclose an electronic device, Figure 5This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the sound source localization method based on the quadratic cross-correlation function disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be a computer.
[0051] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0052] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0053] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the sound source localization method based on a quadratic cross-correlation function as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0054] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned sound source localization method based on a quadratic cross-correlation function. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0055] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0056] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0057] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0058] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0059] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A sound source localization method based on a quadratic cross-correlation function, characterized in that, include: A linear array including several microphones is driven to synchronously record and digitize a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency, thereby obtaining a corresponding digital recording signal. Then, a generalized cross-correlation phase transformation is performed on each of the digital recording signals based on the broadband linear sweep frequency signal to obtain a first cross-correlation sequence including time delay information. Generalized cross-correlation is performed on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation functions. The peak position is located in each of the second cross-correlation functions to determine the target sampling point curve based on the peak position. Then, the target sampling point curve is fitted to obtain the time delay estimate. The target sampling point curve is a sampling point curve constructed based on a preset number of sampling points within a preset range near the peak position. Data points are constructed by using the index difference between every two channels as the x-axis and the estimated time delay as the y-axis. A current time delay point set is constructed based on each data point and the origin. Then, the current weight corresponding to each data point in the current time delay point set is determined, and each data point is fitted based on the current weight to obtain the current fitted line and the current line slope. Determine the residual of each data point relative to the current fitted line, and determine whether the residual with the largest value among the residuals is greater than a preset threshold. If it is greater, set the corresponding data point as an outlier and remove the outlier to obtain a new current time delay point set. Then, jump back to the step of determining the current weight of each data point in the current time delay point set until the current time delay point set meets the preset condition, and set the slope of the current line as the target line slope. The incident angle of the sound source loudspeaker relative to the normal direction of the linear array is determined by using the inverse cosine function and based on the slope of the target straight line, the preset sampling frequency, the microphone spacing and the preset sound speed, so as to determine the sound source orientation result based on the incident angle.
2. The sound source localization method based on a quadratic cross-correlation function according to claim 1, characterized in that, The drive includes a linear array of several microphones and synchronously records and digitizes a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a corresponding digitized recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transform is performed on each of the digitized recording signals to obtain a first-order cross-correlation sequence including time delay information, including: A linear array comprising several microphones is constructed to drive the linear array and synchronously record a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a recording signal to be processed. Then, the recording signal to be processed is digitally processed to obtain a digital recording signal corresponding to each microphone channel. The broadband linear sweep frequency signal is set as the reference reference signal, and the reference reference signal and each of the digital recording signals are transformed to the frequency domain to obtain the corresponding frequency domain representations of the reference signal and the frequency domain representations of the recording signal. The frequency domain representations of each of the recorded signals are subjected to conjugate processing to obtain a conjugate processing result. The conjugate processing result is then multiplied with the frequency domain representation of the reference signal to obtain a frequency domain cross-correlation function. The amplitude corresponding to the frequency domain cross-correlation function is normalized to obtain a normalized result. Then, the normalized result is transformed to the time domain by inverse Fourier transform to obtain a first cross-correlation sequence including time delay information. The time delay information is used to characterize the time delay information of each microphone channel relative to the broadband linear sweep frequency signal. The peak position of the first cross-correlation sequence is used to characterize the relative time delay between the sound wave emitted by the sound source loudspeaker and the reference signal when the sound wave reaches the corresponding microphone.
3. The sound source localization method based on the quadratic cross-correlation function according to claim 1, characterized in that, The step of performing generalized cross-correlation on the primary cross-correlation sequences corresponding to any two different channels to obtain the corresponding secondary cross-correlation functions, locating the peak position in each of the secondary cross-correlation functions, determining the target sampling point curve based on the peak position, and then fitting the target sampling point curve to obtain the time delay estimate includes: Perform generalized cross-correlation phase transformation on the first cross-correlation sequences corresponding to any two different channels to obtain the second cross-correlation function corresponding to each channel, and determine the corresponding peak position in each second cross-correlation function; The peak position is set as a reference, and a preset number of sampling points near the peak position are extracted based on the reference. Then, the target sampling point curve is determined based on each of the sampling points. A quadratic function fitting process is performed on each sampling point in the target sampling point curve to obtain the estimated time delay between the two channels.
4. The sound source localization method based on the quadratic cross-correlation function according to claim 3, characterized in that, The process of constructing data points using the index difference between every two channels as the x-axis and the estimated delay value as the y-axis, and building a current delay point set based on each data point and the origin, includes: Determine the peak position corresponding to each of the channels, and determine the corresponding index estimate based on the peak position. Then determine the position index difference between each index estimate. The location index difference is set as the horizontal axis, and the corresponding time delay estimate is set as the vertical axis. Two-dimensional data points are constructed based on the horizontal and vertical axes, and then the two-dimensional data points are aggregated to obtain the aggregated result. The origin of the coordinate system is set as the reference point, and the current time delay point set is constructed based on the reference point and the aggregation result.
5. The sound source localization method based on a quadratic cross-correlation function according to claim 1, characterized in that, The step of determining the current weight corresponding to each data point in the current delay point set, and fitting each data point based on the current weight to obtain the current fitted line and the slope of the current line, includes: Determine the x-coordinate value corresponding to each data point in the current delay point set; If the horizontal coordinate value is zero, the weight of the corresponding data point is set to 1. If the horizontal coordinate value is not zero, the difference between the total number of microphones corresponding to the linear array and the horizontal coordinate value is determined, and the reciprocal of the difference is set as the weight corresponding to the corresponding data point. Based on each data point in the current time delay point set and its corresponding weight value, a weighted least squares line fitting is performed to obtain the current fitted line, and the slope of the current line corresponding to the current fitted line is determined.
6. The sound source localization method based on the quadratic cross-correlation function according to claim 1, characterized in that, The step of determining the residual of each data point relative to the current fitted line, and judging whether the residual with the largest value among the residuals is greater than a preset threshold, if it is greater, then the corresponding data point is set as an outlier and the outlier is removed to obtain a new current time delay point set, and then jumping back to the step of determining the current weight corresponding to each data point in the current time delay point set, until the current time delay point set meets the preset condition, and the slope of the current line is set as the target line slope, includes: Determine the residuals of each data point in the current time delay point set relative to the current fitted line, and determine the target residual with the largest value among all residuals. Then determine whether the target residual is greater than a preset residual threshold. If the target residual is greater than the preset residual threshold, the corresponding data point of the target residual is marked as an outlier, and the outlier is removed from the current delay point set. Then the current delay point set is updated to obtain a new current delay point set, so as to jump back to the step of determining the current weight of each data point in the current delay point set, until the target residual corresponding to the current delay point set is not greater than the residual threshold, or the number of outliers removed reaches the preset upper limit value, then the current straight line slope is set as the target straight line slope.
7. The sound source localization method based on a quadratic cross-correlation function according to any one of claims 1 to 6, characterized in that, The step of determining the incident angle of the sound source loudspeaker relative to the normal direction of the linear array using the inverse cosine function and based on the slope of the target line, the preset sampling frequency, the microphone spacing, and the preset sound velocity, and determining the sound source orientation result based on the incident angle, includes: A preset spacing between each adjacent microphone in the linear array is determined, and a first ratio is determined based on the slope of the target straight line and the preset sampling frequency. Then, the product result is determined based on the first ratio and the preset sound speed. A second ratio is determined based on the product result and the preset spacing. The incident angle is determined using the inverse cosine function and based on the negative of the second ratio. The directionality of the sound source loudspeaker relative to the linear array is determined based on the incident angle.
8. A sound source directional device based on a quadratic cross-correlation function, characterized in that, include: A first-order cross-correlation sequence generation module is used to drive a linear array including several microphones and synchronously record and digitize a broadband linear sweep frequency signal played by a sound source speaker based on a preset sampling frequency to obtain a corresponding digital recording signal. Then, based on the broadband linear sweep frequency signal, a generalized cross-correlation phase transformation is performed on each of the digital recording signals to obtain a first-order cross-correlation sequence including time delay information. The delay estimation module is used to perform generalized cross-correlation on the first cross-correlation sequences corresponding to any two different channels to obtain the corresponding second cross-correlation functions, locate the peak position in each second cross-correlation function, determine the target sampling point curve based on the peak position, and then fit the target sampling point curve to obtain the delay estimation value. The current delay point set construction module is used to construct data points with the index difference between every two channels as the horizontal axis and the estimated delay value as the vertical axis, and construct the current delay point set based on each data point and the origin of the coordinate system. Then, it determines the current weight corresponding to each data point in the current delay point set, and fits each data point based on the current weight to obtain the current fitted line and the current line slope. The residual determination module is used to determine the residual of each data point relative to the current fitted line, and to determine whether the residual with the largest value among the residuals is greater than a preset threshold. If it is greater, the corresponding data point is set as an outlier and the outlier is removed to obtain a new current time delay point set. Then, the module jumps back to the step of determining the current weight of each data point in the current time delay point set until the current time delay point set meets the preset condition, and the slope of the current line is set as the slope of the target line. The sound source orientation result determination module is used to determine the incident angle of the sound source loudspeaker relative to the normal direction of the linear array using the inverse cosine function and based on the slope of the target straight line, the preset sampling frequency, the microphone spacing and the preset sound speed, so as to determine the sound source orientation result based on the incident angle.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the sound source localization method based on a quadratic cross-correlation function as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the sound source localization method based on a quadratic cross-correlation function as described in any one of claims 1 to 7.