An anti-reverberation sound source positioning method, device, equipment and medium

By employing cross-correlation calculations and iterative filtering strategies using multi-channel microphone arrays, combined with linear fitting methods, the problem of spurious peak interference in high reverberation environments was solved, thereby improving the accuracy and robustness of sound source localization.

CN121522573BActive Publication Date: 2026-03-27MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In high-reverberation indoor environments, existing sound source localization methods are susceptible to spurious peak interference, resulting in low localization accuracy and robustness.

Method used

Synchronous recording is performed using a multi-channel microphone array. Cross-correlation calculations and peak detection are then conducted. An iterative screening strategy is used to determine the stable peak position. Linear fitting is performed and the fitting error is calculated to determine whether the sound source localization is successful. Finally, the target angle and distance are determined based on the fitting results and the delay position.

Benefits of technology

It effectively suppresses spurious peak interference, improves the accuracy and robustness of sound source localization, enables precise calculation of sound source angle and distance, and enhances the stability and accuracy of sound source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121522573B_ABST
    Figure CN121522573B_ABST
Patent Text Reader

Abstract

The application discloses an anti-reverberation sound source positioning method and device, equipment and medium, and relates to the technical field of sound source positioning. The method comprises the following steps: synchronously recording original sweep frequency signals by a preset multi-channel microphone array to obtain a plurality of recording signals corresponding to the number of channels, and performing cross-correlation operation on the recording signals and the original sweep frequency signals respectively to obtain cross-correlation sequences; performing peak value detection on the cross-correlation sequences to obtain peak value detection results, and determining target stable peak value positions based on the peak value detection results by using a preset iteration screening strategy to obtain each target delay position corresponding to the number of channels; performing straight line fitting based on the target delay positions to calculate fitting errors according to the fitting results, and judging whether the current sound source positioning is successful by the fitting errors; if successful, determining a target angle and a target distance based on the fitting results and the target delay positions to position the target sound source.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound source positioning, in particular to an anti-reverberation sound source positioning method, device, equipment and medium. BACKGROUND

[0002] The sound source positioning algorithm based on cross-correlation method (including distance and direction) is widely used in the field of Time Difference of Arrival (TDOA) and Direction of Arrival (DOA) estimation. The traditional method often directly selects the maximum peak of the cross-correlation curve as the sound source arrival delay. In a real indoor environment, reverberation can easily cause this strategy to fail: the sound is reflected by walls, ceilings, furniture and other objects multiple times, forming multiple paths of propagation, and the signal received by the microphone is superimposed by the direct component and multiple reflection components with different delays and amplitudes, resulting in multiple local peaks in the cross-correlation curve. Strong reflections can mask the direct peak, or multiple secondary peaks and noise can jointly lift the background, causing the main peak and the pseudo-peak to have similar amplitudes. Existing improved techniques that enhance robustness through spectral preprocessing may still have multiple similar peaks under high reverberation, with low anti-reverberation performance.

[0003] In summary, how to suppress pseudo-peak interference in a high-reverberation indoor environment to improve the accuracy and robustness of sound source positioning is a technical problem that needs to be solved at present. SUMMARY

[0004] Therefore, the purpose of the present application is to provide an anti-reverberation sound source positioning method, device, equipment and medium, which can suppress pseudo-peak interference in a high-reverberation indoor environment to improve the accuracy and robustness of sound source positioning. The specific scheme is as follows:

[0005] In a first aspect, the present application provides an anti-reverberation sound source positioning method, comprising:

[0006] synchronously recording the original sweep signal through a pre-set multi-channel microphone array to obtain a plurality of recording signals corresponding to the number of channels, and performing cross-correlation operation on the recording signals and the original sweep signal respectively to obtain corresponding cross-correlation sequences; the original sweep signal is a sweep signal emitted by a target sound source;

[0007] performing peak value detection on the cross-correlation sequences to obtain corresponding peak value detection results, and determining a target stable peak value position based on the peak value detection results using a pre-set iteration screening strategy to obtain each target delay position corresponding to the number of channels;

[0008] performing linear fitting based on the target delay positions to calculate a fitting error according to the fitting result, and determining whether the current sound source positioning is successful by the fitting error to obtain a corresponding determination result;

[0009] If the judgment result represents that the current sound source positioning is successful, a target angle and a target distance are determined based on the fitting result and the target delay positions, so as to position the target sound source according to the target angle and the target distance; the target angle and the target distance are an angle and a distance of the target sound source relative to a center point of the multi-channel microphone array.

[0010] Optionally, the cross-correlation operation of the recording signal and the original sweep signal respectively to obtain a corresponding cross-correlation sequence comprises:

[0011] The signal length of the original sweep signal is determined.

[0012] Based on the signal length and a preset discrete cross-correlation function, the discrete cross-correlation operation is performed on the original sweep signal and the recording signal under different time offsets respectively, so as to generate a cross-correlation sequence based on the obtained operation result.

[0013] Optionally, the peak value detection of the cross-correlation sequence to obtain a corresponding peak value detection result comprises:

[0014] A preset peak value detection method is used to identify a local peak in the cross-correlation sequence to determine a corresponding set of local peak value positions; the local peak is a sequence index with the maximum amplitude in a preset neighborhood;

[0015] For each target peak value in the set of local peak value positions, an interpolation operation is performed in the preset neighborhood to obtain a reconstructed local continuous curve;

[0016] The maximum value in the local continuous curve is identified, and the maximum value is used to replace the target peak value to obtain an optimized set of local peak value positions;

[0017] The global maximum peak value position with the maximum amplitude is determined from the optimized set of local peak value positions.

[0018] Optionally, the target stable peak value position is determined based on the peak value detection result by using a preset iterative screening strategy, comprising:

[0019] The global maximum peak value position is determined as a current stable peak value position;

[0020] From the optimized set of local peak value positions, a left adjacent candidate peak position of the stable peak value position is determined.

[0021] The amplitude ratio between the stable peak value position and the left adjacent candidate peak position is calculated, and the amplitude ratio is compared with a preset amplitude ratio threshold to obtain a corresponding comparison result.

[0022] If the amplitude ratio is not greater than the preset amplitude ratio threshold, the left neighboring candidate peak position of the stable peak position is determined as a new stable peak position, and the step of determining the left neighboring candidate peak position of the stable peak position from the optimized local peak position set is jumped to until the amplitude ratio is greater than the preset amplitude ratio threshold, and the current stable peak position is determined as a target stable peak position.

[0023] Optionally, the linear fitting based on the target delay positions is performed to calculate a fitting error according to a fitting result, and the fitting error includes:

[0024] The target delay positions are integrated to obtain a corresponding delay position point set;

[0025] Linear fitting is performed based on the delay position point set to obtain a corresponding fitting straight line, and a slope and an intercept of the fitting straight line are calculated by a least square method; the fitting result includes the slope and the intercept;

[0026] A preset fitting error formula is calculated for each point in the delay position point set by using the fitting result to obtain each fitting error corresponding to the delay position point set.

[0027] Optionally, the fitting error is used to determine whether the current sound source positioning is successful to obtain a corresponding determination result, and the determination result includes:

[0028] If there is a point in the delay position point set with a fitting error greater than a preset fitting error threshold, a confidence degree of the current sound source positioning is set to zero; if there is no point in the delay position point set with a fitting error greater than the preset fitting error threshold, a preset confidence degree formula is calculated by using an iteration number to obtain a calculation result, and the calculation result is determined as the confidence degree of the current sound source positioning; the iteration number is a number of iterations for determining the target stable peak position;

[0029] If the confidence degree is not less than a preset confidence degree threshold, a determination result that the current sound source positioning is successful is obtained;

[0030] If the confidence degree is less than the preset confidence degree threshold, a determination result that the current sound source positioning fails is obtained.

[0031] Optionally, the target angle and the target distance are determined based on the fitting result and the target delay positions, and the target angle and the target distance include:

[0032] A microphone interval of adjacent microphones in the multi-channel microphone array is determined;

[0033] A preset angle formula is calculated by using the microphone interval and the fitting result to obtain a target angle;

[0034] The target delay positions are averaged to obtain corresponding position averages, and a preset distance formula is calculated using the position averages to obtain a target distance.

[0035] In a second aspect, the present application provides an anti-reverberation sound source positioning device, comprising:

[0036] The cross-correlation operation module is configured to perform synchronous recording on the original sweep frequency signal through a preset multi-channel microphone array to obtain a plurality of recording signals corresponding to the number of channels, and perform cross-correlation operation on the recording signals and the original sweep frequency signal respectively to obtain corresponding cross-correlation sequences; the original sweep frequency signal is a sweep frequency signal emitted by a target sound source;

[0037] The position determination module is configured to perform peak value detection on the cross-correlation sequences to obtain corresponding peak value detection results, and determine a target stable peak value position based on the peak value detection results using a preset iteration screening strategy to obtain each target delay position corresponding to the number of channels;

[0038] The straight line fitting module is configured to perform straight line fitting based on the target delay positions to calculate a fitting error according to a fitting result, and determine whether the current sound source positioning is successful by the fitting error to obtain a corresponding determination result;

[0039] The distance determination module is configured to determine a target angle and a target distance based on the fitting result and the target delay positions if the determination result indicates that the current sound source positioning is successful, and to position the target sound source according to the target angle and the target distance; the target angle and the target distance are an angle and a distance of the target sound source relative to a center point of the multi-channel microphone array, respectively.

[0040] In a third aspect, the present application provides an electronic device, comprising:

[0041] The memory is configured to save a computer program;

[0042] The processor is configured to execute the computer program to implement the anti-reverberation sound source positioning method described above.

[0043] In a fourth aspect, the present application provides a computer readable storage medium configured to save a computer program; when the computer program is executed by a processor, the anti-reverberation sound source positioning method described above is implemented.

[0044] In the present application, the original sweep signal emitted by the target sound source is synchronously recorded by a preset multi-channel microphone array to obtain a plurality of recording signals corresponding to the number of channels, and the recording signals are respectively correlated with the original sweep signal to obtain corresponding cross-correlation sequences. The original sweep signal is the sweep signal emitted by the target sound source. Peak detection is performed on the cross-correlation sequences to obtain corresponding peak detection results, and a target stable peak position is determined based on the peak detection results using a preset iterative screening strategy to obtain each target delay position corresponding to the number of channels. Linear fitting is performed based on the target delay positions to calculate a fitting error according to the fitting result, and it is judged whether the current sound source positioning is successful by the fitting error to obtain a corresponding judgment result. If the judgment result indicates that the current sound source positioning is successful, the target angle and the target distance are determined based on the fitting result and the target delay positions to locate the target sound source according to the target angle and the target distance. The target angle and the target distance are the angle and the distance of the target sound source relative to the center point of the multi-channel microphone array. As can be seen from the above, the original sweep signal emitted by the target sound source is synchronously recorded by a preset multi-channel microphone array to obtain a plurality of recording signals matching the number of channels, and each recording signal is correlated with the original sweep signal to generate a corresponding cross-correlation sequence. Peak detection is performed on the cross-correlation sequences, and a target stable peak position is determined based on a preset iterative screening strategy to obtain each target delay position corresponding to the number of channels. Linear fitting is performed based on the target delay positions, and a fitting error is calculated according to the fitting result. It is judged whether the sound source positioning is successful by the fitting error. If it is determined that the positioning is successful, the target angle and the target distance of the target sound source relative to the center point of the microphone array are calculated based on the fitting result and the target delay positions, thereby completing the positioning of the target sound source. In this way, through the above process of the present application, the sound source positioning scheme combining cross-correlation operation, iterative peak screening and linear fitting method can effectively reduce the influence of environmental noise and signal interference on delay position detection, verify the reliability of the positioning result through the fitting error, and accurately measure the angle and distance of the sound source, thereby improving the accuracy and stability of the sound source positioning, and suppressing the pseudo-peak interference in the high reverberation indoor environment to improve the accuracy and robustness of the sound source positioning. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating any inventive labor.

[0046] Figure 1 A flow chart of an anti-reverberation sound source positioning method disclosed by the present application;

[0047] Figure 2 A flow chart of an anti-reverberation sound source positioning method disclosed by the present application;

[0048] Figure 3 A flow chart of determining a target stable peak position disclosed by the present application;

[0049] Figure 4 A structural schematic diagram of an anti-reverberation sound source positioning device disclosed by the present application;

[0050] Figure 5 A structural diagram of an electronic device disclosed by the present application. DETAILED DESCRIPTION

[0051] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0052] In the traditional method, the maximum peak of the cross-correlation curve is directly selected as the sound source arrival delay. In a real indoor environment, reverberation easily causes the failure of this strategy: the sound forms multi-path propagation through multiple reflections on walls, ceilings, furniture, etc., and the signal received by the microphone is superimposed by the direct component and multiple reflection components with different delays and amplitudes, so that the cross-correlation curve has multiple local peaks. Strong reflections can mask the direct peak, or multiple secondary peaks and noise can jointly lift the background, causing the amplitude of the main peak to be close to that of the false peak. Existing improved techniques for improving robustness through spectral preprocessing may still have multiple similar peaks under high reverberation, and the anti-reverberation performance is low.

[0053] In order to overcome the above technical problems, the present application provides an anti-reverberation sound source positioning method, which can suppress false peak interference in a high-reverberation indoor environment to improve the accuracy and robustness of sound source positioning.

[0054] Referring to Figure 1 The embodiments of the present application disclose an anti-reverberation sound source positioning method, which comprises:

[0055] In step S11, a preset multi-channel microphone array is used to synchronously record an original frequency sweeping signal to obtain a plurality of recording signals corresponding to the number of channels, and the recording signals and the original frequency sweeping signal are respectively subjected to cross-correlation operation to obtain corresponding cross-correlation sequences; the original frequency sweeping signal is a frequency sweeping signal emitted by a target sound source.

[0056] In this embodiment, the original sweep signal emitted by the target sound source is synchronously recorded by a preset multi-channel microphone array to obtain a plurality of recording signals corresponding to the number N of microphone channels, and then the recording signal of each channel is respectively cross-correlated with the original sweep signal to generate a corresponding N cross-correlation sequence. Figure 2 Fig. 1 shows a flowchart of an anti-reverberation sound source positioning method provided by the present application.

[0057] It should be noted that the processing procedure of respectively performing cross-correlation operation on the recording signal and the original sweep signal is as follows: determining the signal length of the original sweep signal; based on the signal length and a preset discrete cross-correlation function, performing discrete cross-correlation operation on the original sweep signal and the recording signal under different time offsets respectively to generate a cross-correlation sequence based on the obtained operation result. That is, first, the signal length of the original sweep signal is determined, and then based on the signal length and a preset discrete cross-correlation function, discrete cross-correlation operation is performed on the original sweep signal and the recording signal under different time offsets respectively, and a cross-correlation sequence is generated according to the operation result. The expression formula of the preset discrete cross-correlation function is as follows:

[0058]

[0059] wherein, represents the preset discrete cross-correlation function; x represents the original sweep signal played by the target sound source; y represents the recording signal; k represents the index of the cross-correlation sequence; and L is the signal length of the original sweep signal. It should be further pointed out that, in order to meet the real-time or embedded platform constraints, the anti-reverberation sound source positioning method of the present application can be implemented in a lightweight manner, specifically, FFT (Fast Fourier Transform) acceleration is used when performing cross-correlation operation, a runtime parameter adjustment interface (adjusting threshold, neighborhood size) is provided, and key statistics are recorded for online self-adaptation and offline analysis. In this way, the signal processing method based on multi-channel synchronous recording and cross-correlation operation in this embodiment can accurately capture the propagation delay characteristics of the original sweep signal in different microphone channels, effectively suppress the interference of environmental noise on the signal, provide a signal analysis basis for subsequent sound source positioning, and improve the early signal processing accuracy and reliability of sound source positioning; cross-correlation operation combined with signal length and discrete cross-correlation function can accurately match the timing characteristics of the original sweep signal and the recording signal, effectively quantify the correlation of the two under different time offsets, and improve the calculation accuracy and pertinence of the cross-correlation sequence.

[0060] ​Step S12, peak detection is performed on the cross-correlation sequence to obtain a corresponding peak detection result, and a preset iteration screening strategy is used to determine a target stable peak value position based on the peak detection result, to obtain each target delay position corresponding to the number of channels.

[0061] In this embodiment, peak detection is performed on each cross-correlation sequence to obtain a corresponding peak detection result, and then a preset iteration screening strategy is used to screen a target stable peak value position based on the peak detection result, and then each target delay position corresponding to the number of channels of the microphone is obtained.

[0062] It should be noted that the process of performing peak detection on the cross-correlation sequence to obtain a corresponding peak detection result is as follows: a preset peak detection method is used to identify a local peak in the cross-correlation sequence to determine a corresponding local peak position set; the local peak is a sequence index with the maximum amplitude in a preset neighborhood; for each target peak value in the local peak position set, interpolation operation is performed in the preset neighborhood to obtain a reconstructed local continuous curve; the maximum value in the local continuous curve is identified, and the maximum value is used to replace the target peak value to obtain an optimized local peak position set; and a global maximum peak position with the maximum amplitude is determined from the optimized local peak position set. That is, a preset peak detection method, such as a threshold judgment method based on amplitude and slope or a discrimination algorithm based on local extreme value and threshold, is used to identify a characteristic peak in a signal to identify a local peak in the cross-correlation sequence to determine a corresponding local peak position set , wherein the local peak is a sequence index with the maximum amplitude in a preset neighborhood, interpolation operation (such as parabolic fitting) is performed in the preset neighborhood for each target peak value in the local peak position set to reconstruct a local continuous curve, the maximum value in the local continuous curve is identified and used to replace the target peak value to obtain an optimized local peak position set, and finally a global maximum peak position with the maximum amplitude is determined from the optimized set to identify the global maximum peak.

[0063] It should be further noted that, as Figure 3The diagram illustrates a process for determining a target stable peak position according to this application. The process of determining the target stable peak position based on the peak detection results using a preset iterative screening strategy is as follows: The global maximum peak position is determined as the current stable peak position; from the optimized set of local peak positions, the left adjacent candidate peak position of the stable peak position is determined; the amplitude ratio between the stable peak position and the left adjacent candidate peak position is calculated, and the amplitude ratio is compared with a preset amplitude ratio threshold to obtain a corresponding comparison result; if the comparison result indicates that the amplitude ratio is not greater than the preset amplitude ratio threshold, then the left adjacent candidate peak position is determined as the new stable peak position, and the process jumps to the step of determining the left adjacent candidate peak position of the stable peak position from the optimized set of local peak positions, until the amplitude ratio is greater than the preset amplitude ratio threshold, and the current stable peak position is determined as the target stable peak position. That is, determining the global maximum peak position as the current stable peak position means letting... Let the stable peak be equal to the global maximum peak, and find the stable peak position from the optimized set of local peak positions. The position of the left adjacent candidate peak That is, let the candidate peak be equal to the left adjacent peak of the stable peak, and calculate the amplitude ratio between the two. The specific formula for calculating the amplitude ratio is as follows:

[0064] ;

[0065] in, Characterizes the amplitude ratio. The amplitude ratio is compared with a preset amplitude ratio threshold (preset amplitude ratio threshold). The value can be set according to the specific environment (1.1~1.5 is recommended). If the amplitude ratio is not greater than the threshold, the position of the left adjacent candidate peak is updated to the new stable peak position, i.e., let... , the stable peak = the candidate peak, and the above-mentioned left adjacent candidate peak position determination and amplitude ratio comparison step is repeated until the amplitude ratio is greater than the threshold value, the current stable peak position is determined as the target stable peak position, and the stable peak is output. In addition, the embodiment can also stop the cyclic repetition of the step when there is no more candidate local peak. It should be noted that the maximum iteration screening number Q is specified by the user, and when the iteration screening number reaches Q, it is determined that the iteration of the channel fails, and the channel can be marked so as not to use the channel information in subsequent processing. In this way, the embodiment detects the peak value first, and then locks the stable peak value matched with the true signal delay characteristic through the iteration screening based on the adjacent peak value amplitude ratio, which can effectively suppress the false peak caused by reflection (reverb) and guarantee the accuracy and reliability of the target delay position. The local peak identification is performed first, then the curve is reconstructed by interpolation, then the peak value is optimized, and finally the global peak value is extracted to obtain the peak value detection result, which can make up for the precision limitation of the discrete cross-correlation sequence, improve the sub-sampling positioning precision of the peak value position through interpolation fitting, filter out the local interference small peak value, and accurately lock the global maximum peak value matched with the signal delay characteristic.

[0066] In step S13, a straight line fitting is performed based on the target delay positions, a fitting error is calculated according to the fitting result, and it is determined whether the current sound source positioning is successful by using the fitting error, so as to obtain a corresponding determination result.

[0067] In the embodiment, based on the target delay positions corresponding to each channel A straight line fitting is performed based on the target delay positions, a fitting error is calculated according to the fitting result, and the fitting error is used as a determination basis to determine whether the current sound source positioning operation is successful, so as to obtain a corresponding determination result.

[0068] It should be noted that the processing procedure of performing a straight line fitting based on the target delay positions to calculate a fitting error according to the fitting result is as follows: the target delay positions are integrated to obtain a corresponding delay position point set; a straight line fitting is performed based on the delay position point set to obtain a corresponding fitting straight line, and the slope and intercept of the fitting straight line are calculated by using the least square method; the fitting result includes the slope and the intercept; and a preset fitting error formula is used to calculate a fitting error corresponding to each point in the delay position point set by using the fitting result. That is, the target delay positions corresponding to each channel are integrated to obtain a delay position point set A straight line fitting is performed based on the delay position point set to generate a fitting straight line, and the slope and intercept of the fitting straight line are calculated by using the least square method. The calculation formula of the slope is as follows:

[0069] ;

[0070] wherein, characterizes the slope; m characterizes the channel point of each channel. The calculation formula of the intercept is specifically as follows:

[0071] ;

[0072] wherein, characterizes the intercept. After obtaining the slope and the intercept, a fitting error corresponding to each point in the set of delay position points is calculated respectively by using a preset fitting error formula containing the slope and the intercept, by using a fitting result of the slope and the intercept. The fitting error is a point-by-point absolute fitting error, which is used to quantify the deviation of each channel point from the fitting straight line. The expression formula of the preset fitting error formula is specifically as follows:

[0073] ;

[0074] wherein, characterizes the fitting error.

[0075] It needs to be further pointed out that the processing procedure for judging whether the current sound source positioning is successful by using the fitting error is as follows: if there is a point in the set of delay position points whose fitting error is greater than a preset fitting error threshold, the confidence of the current sound source positioning is set to zero; if there is no point in the set of delay position points whose fitting error is greater than the preset fitting error threshold, the confidence of the current sound source positioning is determined by using an iteration number calculation preset confidence formula; the iteration number is the number of iterations for determining the target stable peak position; if the confidence is not less than a preset confidence threshold, a judgment result that the current sound source positioning is successful is obtained; if the confidence is less than the preset confidence threshold, a judgment result that the current sound source positioning fails is obtained. The confidence is used to evaluate the reliability of the current positioning result. That is, a preset fitting error threshold can be set, if there is a point in the set of delay position points whose fitting error is greater than the preset fitting error threshold, that is, there is a certain m that makes , the confidence of the current sound source positioning is set to zero, that is, ; if there is no such point, that is, all , the confidence of the current sound source positioning is calculated by using a preset confidence formula combined with the iteration number for determining the target stable peak position, and the expression formula of the preset confidence formula is specifically as follows:

[0076] ;

[0077] wherein, characterizes the iteration number for determining the target stable peak position of the mth channel. A preset confidence threshold (referencing a value of 0.5 to 0.9), the confidence is compared with the preset confidence threshold, if the confidence is not less than the threshold, i.e. , it is determined that the current sound source positioning is successful, otherwise it is determined that the positioning fails. In this way, the embodiment takes the straight line fitting error of the target delay position as the verification standard for whether the positioning is successful or not, which can effectively verify the consistency and reliability of the delay position data, exclude abnormal delay data caused by signal interference or detection deviation, and improve the accuracy and reliability of the sound source positioning result; the straight line fitting based on the least square method and the point-by-point error calculation method can accurately quantify the fitting degree of the delay position point set and the fitted straight line, reflect the consistency and reliability of the target delay position, and provide data basis for judging whether the sound source positioning is successful or not through the fitting error; the confidence determination method combining the fitting error verification and the iteration number weighting can comprehensively evaluate the effectiveness of the positioning result from two dimensions of data consistency and peak value screening reliability, and effectively exclude the positioning deviation caused by abnormal delay data and unstable peak value screening.

[0078] Step S14, if the judgment result represents that the current sound source positioning is successful, the target angle and the target distance are determined based on the fitting result and the target delay positions of each channel, so as to position the target sound source according to the target angle and the target distance; the target angle and the target distance are the angle and the distance of the target sound source relative to the center point of the multi-channel microphone array.

[0079] In the embodiment, if the judgment result of the current sound source positioning is successful, the target angle and the target distance of the target sound source relative to the center point of the multi-channel microphone array are calculated based on the fitting result and the target delay positions corresponding to each channel, so as to complete the accurate positioning of the target sound source.

[0080] It should be noted that the processing procedure for determining the target angle and the target distance based on the fitting result and the target delay positions is as follows: the microphone spacing of adjacent microphones in the multi-channel microphone array is determined; the preset angle formula is calculated by using the microphone spacing and the fitting result, so as to obtain the target angle; the target delay positions are averaged to obtain the corresponding position average value, and the preset distance formula is calculated by using the position average value, so as to obtain the target distance. That is, the microphone spacing of adjacent microphones in the multi-channel microphone array is determined, the target angle is calculated by using the microphone spacing and the fitting result through the preset angle formula, and the calculation formula of the target angle is as follows:

[0081] ;

[0082] wherein, is the target angle; c is the speed of sound; is a sampling frequency of a signal; d is the microphone spacing; arccos is an inverse cosine function. A position average value is obtained by averaging the target delay positions corresponding to each channel, and a calculation formula of the position average value is specifically as follows:

[0083] ;

[0084] wherein, is the position average value; M represents a position quantity of the target delay position of a current effective channel. A target distance is calculated based on the position average value and by using a preset distance formula. A calculation formula of the preset distance formula is specifically as follows:

[0085] ;

[0086] wherein, D is the target distance. It can be understood that, considering temperature interference, the embodiment can add temperature compensation to the sound velocity to correct the sound velocity, and at this time, an expression formula of the sound velocity is specifically as follows:

[0087] ;

[0088] wherein, T is an air temperature, and a unit is degree Celsius (°C). In this way, the embodiment combines the fitting parameters and the actual delay position to calculate the sound source positioning parameters, which can accurately map the relative spatial position relationship between the sound source and the microphone array, and guarantee the precision and practicality of the sound source positioning result. The sound source positioning parameter calculation mode combining the hardware parameters, the fitting result and the delay position statistical value can accurately convert the delay characteristics in the signal level into the angle and distance parameters in the spatial level, and guarantee the accuracy and interpretability of the target sound source positioning result.

[0089] It can be seen from the above that, by presetting a multi-channel microphone array to synchronously record an original sweep frequency signal emitted by a target sound source, a plurality of recording signals corresponding to the number of channels are obtained, each of the recording signals is subjected to cross-correlation operation with the original sweep frequency signal to generate a corresponding cross-correlation sequence, peak value detection is performed on the cross-correlation sequence, a target stable peak value position is determined by combining a preset iterative screening strategy, each target delay position corresponding to the number of channels is obtained, linear fitting is performed based on the target delay position, fitting error is calculated according to the fitting result, and it is judged whether the sound source positioning is successful or not according to the fitting error, if it is judged that the positioning is successful, the target angle and the target distance of the target sound source relative to the center point of the microphone array are calculated by combining the fitting result and the target delay position, so that the positioning of the target sound source is completed. In this way, by the above process of the embodiment of the application, the sound source positioning scheme combining the cross-correlation operation, the iterative peak value screening and the linear fitting method can effectively reduce the influence of environmental noise and signal interference on the delay position detection, verify the reliability of the positioning result through the fitting error, and realize accurate measurement of the angle and distance of the sound source, thereby improving the accuracy and stability of the sound source positioning, and further suppressing the false peak interference in the high reverberation indoor environment to improve the accuracy and robustness of the sound source positioning.

[0090] Correspondingly, referring to Figure 4 The embodiment of the application also provides an anti-reverberation sound source positioning device, which comprises:

[0091] A cross-correlation operation module 11 is configured to perform synchronous recording on an original sweep frequency signal by a preset multi-channel microphone array to obtain a plurality of recording signals corresponding to the number of channels, and perform cross-correlation operation on each of the recording signals and the original sweep frequency signal to obtain a corresponding cross-correlation sequence; the original sweep frequency signal is a sweep frequency signal emitted by a target sound source;

[0092] A position determination module 12 is configured to perform peak value detection on the cross-correlation sequence to obtain a corresponding peak value detection result, and determine a target stable peak value position based on the peak value detection result by using a preset iterative screening strategy, so as to obtain each target delay position corresponding to the number of channels;

[0093] A linear fitting module 13 is configured to perform linear fitting based on the target delay positions, calculate fitting error according to the fitting result, and judge whether the sound source positioning is successful or not by using the fitting error to obtain a corresponding judgment result;

[0094] The distance determination module 14 is configured to, if the judgment result indicates that the current sound source positioning is successful, determine a target angle and a target distance based on the fitting result and the target delay positions, so as to position the target sound source according to the target angle and the target distance. The target angle and the target distance are an angle and a distance of the target sound source relative to a center point of the multi-channel microphone array, respectively.

[0095] In some embodiments, the cross-correlation operation module 11 can specifically include:

[0096] The length determination unit is configured to determine a signal length of the original sweep signal.

[0097] The cross-correlation operation unit is configured to perform a discrete cross-correlation operation on the original sweep signal and the recorded signal with different time offsets based on the signal length and a preset discrete cross-correlation function, so as to generate a cross-correlation sequence based on an obtained operation result.

[0098] In some embodiments, the position determination module 12 can specifically include:

[0099] The local peak identification unit is configured to identify a local peak in the cross-correlation sequence by using a preset peak detection method, so as to determine a corresponding set of local peak position.

[0100] The interpolation operation unit is configured to, for each target peak in the set of local peak positions, perform an interpolation operation in the preset neighborhood, so as to obtain a reconstructed local continuous curve.

[0101] The peak replacement unit is configured to identify a maximum value in the local continuous curve, and replace the target peak with the maximum value, so as to obtain an optimized set of local peak positions.

[0102] The first position determination unit is configured to determine a global maximum peak position with the largest amplitude from the optimized set of local peak positions.

[0103] In some embodiments, the position determination module 12 can specifically include:

[0104] The second position determination unit is configured to determine the global maximum peak position as a current stable peak position.

[0105] The third position determination unit is configured to determine a left adjacent candidate peak position of the stable peak position from the optimized set of local peak positions.

[0106] An amplitude ratio comparison unit is configured to calculate an amplitude ratio between the stable peak position and the left neighboring candidate peak position, and compare the amplitude ratio with a preset amplitude ratio threshold to obtain a corresponding comparison result.

[0107] A step jumping unit is configured to, if the comparison result indicates that the amplitude ratio is not greater than the preset amplitude ratio threshold, determine the left neighboring candidate peak position as a new stable peak position, and jump to the step of determining the left neighboring candidate peak position of the stable peak position from the optimized local peak position set until the amplitude ratio is greater than the preset amplitude ratio threshold, and determine the current stable peak position as a target stable peak position.

[0108] In some embodiments, the straight line fitting module 13 can specifically include:

[0109] A position integration unit is configured to integrate the target delay positions to obtain a corresponding delay position point set.

[0110] An intercept calculation unit is configured to perform straight line fitting based on the delay position point set to obtain a corresponding fitting straight line, and calculate a slope and an intercept of the fitting straight line by least square method; the fitting result includes the slope and the intercept.

[0111] A first formula calculation unit is configured to calculate a preset fitting error formula for each point in the delay position point set by using the fitting result to obtain each fitting error corresponding to the delay position point set.

[0112] In some embodiments, the straight line fitting module 13 can specifically include:

[0113] A confidence degree determination unit is configured to, if there is a point in the delay position point set with a fitting error greater than a preset fitting error threshold, set the confidence degree of the current sound source positioning to zero; if there is no point in the delay position point set with a fitting error greater than the preset fitting error threshold, calculate a preset confidence degree formula by using the number of iterations to obtain a calculation result, and determine the calculation result as the confidence degree of the current sound source positioning; the number of iterations is the number of iterations for determining the target stable peak position.

[0114] A first result acquisition unit is configured to, if the confidence degree is not less than a preset confidence degree threshold, obtain a judgment result that the current sound source positioning is successful.

[0115] A second result acquisition unit is configured to, if the confidence degree is less than the preset confidence degree threshold, obtain a judgment result that the current sound source positioning fails.

[0116] In some embodiments, the distance determination module 14 can specifically include:

[0117] a distance determination unit configured to determine a microphone distance of adjacent microphones in the multi-channel microphone array;

[0118] a second formula calculation unit configured to calculate a preset angle formula by using the microphone distance and the fitting result to obtain a target angle;

[0119] a third formula calculation unit configured to average the target delay positions to obtain a corresponding position average value, and calculate a preset distance formula by using the position average value to obtain a target distance.

[0120] Further, the application also discloses an electronic device, Figure 5 is a structural diagram of an electronic device 20 according to an example embodiment, and the content in the figure cannot be considered as any limitation on the use range of the application. The electronic device 20 can specifically include at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25 and a communication bus 26. The memory 22 is used for storing a computer program, and the computer program is loaded and executed by the processor 21 to realize the related steps in the anti-reverberation sound source positioning method disclosed in any of the preceding embodiments. In addition, the electronic device 20 in the embodiment can be an electronic computer.

[0121] In the embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the application, which is not limited here; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to the specific application needs, which is not limited here.

[0122] In addition, the memory 22 as a carrier for resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage mode can be temporary storage or permanent storage.

[0123] The operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the anti-reverberation sound source positioning method executed by the electronic device 20 disclosed in any of the preceding embodiments, the computer program 222 can further include a computer program capable of completing other specific work.

[0124] Further, the application also discloses a computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to realize the anti-reverberation sound source positioning method disclosed above. For the specific steps of the method, refer to the corresponding content disclosed in the foregoing embodiments, which will not be repeated here.

[0125] The various embodiments are described in the specification by progressive stages, and each embodiment focuses on the difference from other embodiments. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts refer to the method part.

[0126] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0127] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.

[0128] Finally, it should be noted that in this document, relational terms such as first and second are used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0129] The technical solutions provided by the present application are described in detail above, and the principles and implementation manners of the present application are described by using specific examples. The above description of the examples is only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed, and the above description of the content of the specification should not be understood as a limitation on the present application.

Claims

1. A method for locating reverberation-resistant sound sources, characterized in that, include: The original sweep frequency signal is synchronously recorded using a preset multi-channel microphone array to obtain multiple recording signals corresponding to the number of channels. The recording signals are then cross-correlated with the original sweep frequency signal to obtain corresponding cross-correlation sequences. The original sweep frequency signal is the sweep frequency signal emitted by the target sound source. Peak detection is performed on the cross-correlation sequence to obtain the corresponding peak detection results. Based on the peak detection results, the target stable peak position is determined using a preset iterative filtering strategy to obtain the target delay position corresponding to the number of channels. Linear fitting is performed based on the delay positions of each target, and the fitting error is calculated based on the fitting result. The fitting error is then used to determine whether the sound source localization is successful, and the corresponding judgment result is obtained. If the judgment result indicates that the sound source localization is successful, then the target angle and target distance are determined based on the fitting result and the delay positions of each target, so as to locate the target sound source according to the target angle and the target distance; the target angle and the target distance are respectively the angle and distance of the target sound source relative to the center point of the multi-channel microphone array; The step of performing peak detection on the cross-correlation sequence to obtain corresponding peak detection results, and determining the target stable peak position based on the peak detection results using a preset iterative screening strategy, includes: A preset peak detection method is used to identify local peaks in the cross-correlation sequence to determine the corresponding set of local peak positions; the local peak is the sequence index with the largest amplitude in a preset neighborhood; for each target peak in the set of local peak positions, interpolation is performed in the preset neighborhood to obtain a reconstructed local continuous curve; Identify the maximum value in the local continuous curve and replace the target peak value with the maximum value to obtain an optimized set of local peak positions; determine the global maximum peak position with the largest amplitude from the optimized set of local peak positions; The global maximum peak position is determined as the current stable peak position; the left adjacent candidate peak position is determined from the optimized local peak position set; the amplitude ratio between the stable peak position and the left adjacent candidate peak position is calculated, and the amplitude ratio is compared with a preset amplitude ratio threshold to obtain the corresponding comparison result; If the comparison result indicates that the amplitude ratio is not greater than the preset amplitude ratio threshold, then the left adjacent candidate peak position is determined as the new stable peak position, and the process jumps to the step of determining the left adjacent candidate peak position of the stable peak position from the optimized local peak position set, until the amplitude ratio is greater than the preset amplitude ratio threshold, and the current stable peak position is determined as the target stable peak position.

2. The method for locating reverberation-resistant sound sources according to claim 1, characterized in that, The step of performing cross-correlation operations on the recorded signal and the original swept frequency signal to obtain corresponding cross-correlation sequences includes: Determine the signal length of the original sweep frequency signal; Based on the signal length and a preset discrete cross-correlation function, discrete cross-correlation operations are performed on the original swept frequency signal and the recording signals at different time offsets, respectively, to generate a cross-correlation sequence based on the obtained operation results.

3. The method for locating reverberation-resistant sound sources according to claim 1, characterized in that, The step of performing linear fitting based on the target delay positions, and calculating the fitting error based on the fitting results, includes: Integrate the aforementioned target delay locations to obtain the corresponding delay location point set; A straight line is fitted based on the set of delayed position points to obtain a corresponding fitted straight line, and the slope and intercept of the fitted straight line are calculated by the least squares method; the fitting result includes the slope and the intercept. Using the fitting results, a preset fitting error formula is calculated for each point in the set of delayed position points to obtain the fitting errors corresponding to the set of delayed position points.

4. The method for locating reverberation-resistant sound sources according to claim 3, characterized in that, The step of determining whether the sound source localization was successful based on the fitting error, and obtaining the corresponding judgment result, includes: If there are points in the set of delayed location points whose fitting error is greater than a preset fitting error threshold, then the confidence level of the current sound source localization is set to zero; if there are no points in the set of delayed location points whose fitting error is greater than a preset fitting error threshold, then a preset confidence formula is calculated using the number of iterations, and the calculated result is determined as the confidence level of the current sound source localization; the number of iterations is the number of iterations used to determine the stable peak position of the target. If the confidence level is not less than the preset confidence threshold, then the judgment result of successful sound source localization is obtained; If the confidence level is less than the preset confidence threshold, then the result of the judgment that the sound source localization failed is obtained.

5. The method for locating reverberation-resistant sound sources according to any one of claims 1 to 4, characterized in that, The step of determining the target angle and target distance based on the fitting results and the delay positions of each target includes: Determine the microphone spacing between adjacent microphones in the multi-channel microphone array; The target angle is obtained by calculating a preset angle formula using the microphone spacing and the fitting results; The average position of each target delay position is calculated to obtain the corresponding position average value, and the target distance is obtained by using the position average value to calculate the preset distance formula.

6. A reverberation-resistant sound source localization device, characterized in that, include: The cross-correlation operation module is used to synchronously record the original sweep frequency signal through a preset multi-channel microphone array to obtain multiple recording signals corresponding to the number of channels, and to perform cross-correlation operation on the recording signals and the original sweep frequency signal to obtain the corresponding cross-correlation sequence; the original sweep frequency signal is the sweep frequency signal emitted by the target sound source; The position determination module is used to perform peak detection on the cross-correlation sequence to obtain the corresponding peak detection results, and to determine the target stable peak position based on the peak detection results using a preset iterative filtering strategy, so as to obtain the target delay position corresponding to the number of channels. The linear fitting module is used to perform linear fitting based on the delay positions of each target, calculate the fitting error based on the fitting result, and determine whether the sound source localization is successful based on the fitting error, thereby obtaining the corresponding judgment result. The distance determination module is used to determine the target angle and target distance based on the fitting result and the delay positions of each target if the judgment result indicates that the sound source localization is successful, so as to locate the target sound source according to the target angle and the target distance; the target angle and the target distance are respectively the angle and distance of the target sound source relative to the center point of the multi-channel microphone array; The location determination module includes: A local peak identification unit is used to identify local peaks in the cross-correlation sequence using a preset peak detection method, so as to determine the corresponding set of local peak positions; the local peak is the sequence index with the largest amplitude in a preset neighborhood. An interpolation unit is used to perform interpolation operations on each target peak in the set of local peak locations within the preset neighborhood to obtain a reconstructed local continuous curve. A peak replacement unit is used to identify the maximum value in the local continuous curve and replace the target peak value with the maximum value to obtain an optimized set of local peak positions. The first position determination unit is used to determine the global maximum peak position with the largest amplitude from the optimized set of local peak positions; The second position determination unit is used to determine the global maximum peak position as the current stable peak position; The third position determination unit is used to determine the position of the left adjacent candidate peak of the stable peak position from the optimized set of local peak positions. An amplitude ratio comparison unit is used to calculate the amplitude ratio between the stable peak position and the left adjacent candidate peak position, and compare the amplitude ratio with a preset amplitude ratio threshold to obtain the corresponding comparison result; The step jump unit is used to determine the left adjacent candidate peak position as the new stable peak position if the comparison result indicates that the amplitude ratio is not greater than the preset amplitude ratio threshold, and jump to the step of determining the left adjacent candidate peak position of the stable peak position from the optimized local peak position set, until the amplitude ratio is greater than the preset amplitude ratio threshold, and the current stable peak position is determined as the target stable peak position.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the anti-reverberation sound source localization method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the anti-reverberation sound source localization method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Wire position identification method, system and equipment of electromagnetic cruise system and medium

    CN121274889A

  • Method for Encoding Multi-Channel Signal and Encoder

    US20190189134A1