Sound source localization and distance measurement method and device based on microphone array, equipment and storage medium

By using frequency sweep signal processing and iterative cross-correlation technology with microphone arrays, the accuracy and robustness issues of sound source localization and distance measurement in low signal-to-noise ratio environments are solved, achieving efficient sound source localization and distance measurement, which is suitable for complex sound fields and resource-constrained platforms.

CN121831683APending Publication Date: 2026-04-10MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In environments with low signal-to-noise ratio or strong reverberation, microphone array-based sound source localization and distance measurement methods suffer from reduced directional accuracy, multipath reflection and noise interference, and insufficient robustness. Furthermore, traditional methods have high requirements for array deployment accuracy and synchronization, making it difficult to operate stably on resource-constrained embedded platforms.

Method used

The frequency sweep signal is synchronously acquired by a microphone array, and then subjected to generalized cross-correlation and phase transformation weighting processing to construct a cross-correlation sequence set. Iterative cross-correlation operations are performed, and combined with upsampling and time delay alignment, a single-channel beamforming signal is constructed to determine the sound source distance.

Benefits of technology

It improves the efficiency of sound source localization and distance measurement, enhances the directional accuracy and stability in complex sound fields, and is suitable for resource-constrained embedded platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121831683A_ABST
    Figure CN121831683A_ABST
Patent Text Reader

Abstract

The invention discloses a sound source localization and distance measurement method, device and equipment based on a microphone array and a storage medium, and relates to the technical field of acoustics, and the method comprises the steps: collecting a frequency sweeping signal played by a sound source through the microphone array, obtaining the audio data of each channel, obtaining a cross-correlation sequence through generalized cross-correlation and phase transformation weighting processing, and obtaining a distance measurement result; constructing a current to-be-processed set, processing other sequences again by taking the first sequence as a reference to obtain a new sequence group, discarding the reference sequence, taking the new group as a to-be-processed set, repeating the iteration step until only one to-be-processed sequence is left in the set, and performing up-sampling on the sequences to obtain a to-be-processed set; extracting a first main peak position to obtain a time delay estimation value of a decimal sampling level so as to determine an incident angle of a sound source relative to an array normal, performing time delay alignment on each cross-correlation sequence, constructing a single-channel beam forming signal based on an alignment result, and determining a second main peak position so as to determine a sound source distance, the efficiency of sound source localization and distance measurement is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of acoustics, in particular to a sound source positioning and distance measurement method and device based on a microphone array, equipment and a storage medium. BACKGROUND

[0002] At present, the sound source positioning technology aims to determine the spatial position or direction of the sound source by using the sound wave signals received by the microphone array. The common implementation path includes methods based on time difference of arrival, angle of arrival and signal strength, etc. Among them, the method based on time difference of arrival is widely used in voice conference, intelligent robot and security monitoring scenes due to its mature implementation and moderate calculation amount. The traditional scheme usually estimates the time delay difference between channels through generalized cross-correlation, and then calculates the direction of the sound source combined with the array geometry, and the position of the sound source can be further obtained through geometric solution and other methods.

[0003] In a low signal-to-noise ratio or strong reverberation environment, multipath reflection and noise can cause multiple competing peaks in the cross-correlation function, and the peak value corresponding to the true time delay is easy to be deviated by interference, thereby causing the directional accuracy to decrease. At the same time, the traditional method has high requirements for array layout accuracy and synchronization, and the collaborative use of array multi-channel information is insufficient, which is easy to produce sidelobe interference and has limited robustness. Although existing researches improve the directional accuracy by improving the cross-power spectrum weighting (such as using GCC-PHAT) or introducing least squares optimization, there are still problems such as false peak and insufficient stability in complex sound fields. Moreover, part of the high-precision implementation relies heavily on iterative optimization or matrix operation, which is difficult to run stably and with low latency on resource-constrained embedded platforms for a long time.

[0004] From the above, how to improve the efficiency of sound source positioning and distance measurement in the process of sound source positioning and distance measurement based on a microphone array is a problem to be solved at present. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a sound source positioning and distance measurement method, device, equipment and storage medium based on a microphone array, which can improve the efficiency of sound source positioning and distance measurement in the process of sound source positioning and distance measurement based on a microphone array. The specific scheme is as follows: In a first aspect, the present application provides a sound source positioning and distance measurement method based on a microphone array, comprising: synchronously collecting, by a microphone array, a sweep signal played by a sound source to be positioned to obtain each audio data to be processed, performing generalized cross-correlation and phase transformation weighting processing on each audio data to be processed based on the sweep signal to obtain a cross-correlation sequence corresponding to each channel, and then constructing a current processing set based on each cross-correlation sequence; set the first sequence in the current to-be-processed set as a current reference sequence, and perform generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence to obtain a current cross-correlation sequence group, then discard the current reference sequence, and set the first sequence in the obtained new current cross-correlation sequence group as a new current reference sequence, and then set the current cross-correlation sequence group as a new current to-be-processed set to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set; perform an upsampling operation on the to-be-processed sequence to obtain an upsampled sequence, and extract a fractional sampling stage time delay estimation value based on a first main peak value position in the upsampled sequence, to determine an incident angle of the to-be-positioned sound source relative to an array normal line based on the time delay estimation value, an inter-element distance, a sampling rate, and a sound speed, and perform integer time delay alignment and fractional time delay alignment on each of the cross-correlation sequences based on the time delay estimation value and the incident angle to obtain corresponding alignment results, respectively; construct a target single-channel beam synthesis signal based on each of the alignment results, and determine a second main peak value position corresponding to the target single-channel beam synthesis signal, to determine a sound source distance based on the second main peak value position, the time delay estimation value, a preset system delay, and the sound speed.

[0006] Optionally, the sweep signal played by the to-be-positioned sound source is synchronously collected by the microphone array to obtain each to-be-processed audio data, and generalized cross-correlation and phase transform weighting processing is performed on each of the to-be-processed audio data based on the sweep signal to obtain a cross-correlation sequence corresponding to each channel, comprising: The sweep signal played by the to-be-positioned sound source is synchronously collected by the microphone array based on a preset starting frequency, a preset terminal frequency, and a preset duration; the sweep signal includes a linear sweep signal and a logarithmic sweep signal; The sweep signal corresponding to each channel is subjected to clock synchronization calibration and gain consistency calibration to obtain to-be-processed signals, and each of the to-be-processed signals is subjected to band-pass filtering and pre-emphasis filtering to obtain each to-be-processed audio data; The sweep signal is set as a reference signal, and each of the to-be-processed audio data is subjected to Fourier transform based on the reference signal to obtain a corresponding frequency domain representation, and then the cross-power spectrum is determined based on the conjugate multiplication result between each of the frequency domain representations and each of the to-be-processed audio data; Determine the modulus corresponding to the cross-power spectrum, and determine the to-be-transformed result based on the cross-power spectrum, the modulus, and a preset minimum positive number, then perform phase transformation weighting on the to-be-transformed result to obtain a weighted cross-power spectrum, and perform inverse Fourier transformation on the weighted cross-power spectrum to obtain the corresponding cross-correlation sequence.

[0007] Optionally, the first sequence in the current to-be-processed set is set as a current reference sequence, generalized cross-correlation and phase transformation weighting processing is performed on the current remaining sequences in the current to-be-processed set based on the current reference sequence to obtain a current cross-correlation sequence group, then the current reference sequence is discarded, and the first sequence in the obtained new current cross-correlation sequence group is set as a new current reference sequence, and the current cross-correlation sequence group is set as a new current to-be-processed set, so as to jump back to the step of performing generalized cross-correlation and phase transformation weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set, including: The first sequence in the current to-be-processed set is set as a current reference sequence, and generalized cross-correlation operations are respectively performed on the current reference sequence and each sequence in the current to-be-processed set to obtain corresponding cross-correlation processing results, and then phase transformation weighting operations are performed on the cross-correlation processing results to obtain corresponding cross-correlation sequences respectively; A current cross-correlation sequence group is constructed based on the cross-correlation sequences, then the current reference sequence in the current cross-correlation sequence group is discarded to obtain a new current cross-correlation sequence group, and the first sequence in the current cross-correlation sequence group is set as a new current reference sequence; The current cross-correlation sequence group is set as a new current to-be-processed set, and the step of performing generalized cross-correlation and phase transformation weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence is jumped back, until only one to-be-processed sequence is included in the current to-be-processed set; the to-be-processed sequence includes common time delay information corresponding to each channel.

[0008] Optionally, the up-sampling operation satisfying the preset high-multiple condition is performed on the to-be-processed sequence to obtain an up-sampled sequence, and a fractional sampling stage time delay estimation value is extracted based on a first main peak value position in the up-sampled sequence, including: The up-sampling operation satisfying the preset high-multiple condition is performed on the to-be-processed sequence to obtain an up-sampled sequence, and an index position corresponding to a global maximum value in the up-sampled sequence is searched, and the index position is set as a first main peak value position; the preset high-multiple condition is a condition determined based on a preset time delay estimation precision; determine a time delay estimation value of the fractional sampling stage based on the first main peak position and the up-sampling ratio; the time delay estimation value is used to represent a proportional relationship between an actual time delay of a received signal between adjacent array elements and a sampling interval.

[0009] Optionally, the integer time delay alignment and the fractional time delay alignment are respectively performed on each of the cross-correlation sequences based on the time delay estimation value and the incident angle, to obtain a corresponding alignment result respectively. determine an integer sampling delay amount corresponding to each channel based on the time delay estimation value and an array element number, and then perform an integer sampling point alignment operation on each of the cross-correlation sequences based on a preset alignment operation, the integer sampling delay amount, a preset delay direction, and the incident angle, to obtain an integer time delay alignment result; the delay direction is determined based on a sign of the time delay estimation value; the preset alignment operation includes a zero padding mode and a cyclic shift mode; determine a fractional delay amount corresponding to each channel based on the integer sampling delay amount, the time delay estimation value, and a channel number, and then perform a fractional time delay alignment on the integer time delay alignment result based on a preset first-order all-pass filter and a preset filter coefficient, to obtain a corresponding alignment result; the preset filter coefficient is determined based on a preset formula and the fractional delay amount corresponding to each channel.

[0010] Optionally, the target single-channel beamforming signal is constructed based on each of the alignment results, including: perform point-by-point and weighted addition on each sequence in each of the alignment results, to obtain an initial single-channel beamforming signal; the initial single-channel beamforming signal includes energy corresponding to the sound source to be positioned; perform a signal enhancement operation on the initial single-channel beamforming signal, to obtain a signal enhancement result, and determine interference and noise from other directions in the signal enhancement result, and then suppress the interference and the noise, to obtain a target single-channel beamforming signal.

[0011] Optionally, a second main peak position corresponding to the target single-channel beamforming signal is determined, and a sound source distance is determined based on the second main peak position, the time delay estimation value, a preset system delay, and the sound speed, including: search for an index corresponding to a maximum amplitude in a full sequence range of the target single-channel beamforming signal, and set the index as the second main peak position; perform peak position positioning in a target neighborhood of the second main peak position using a preset fitting algorithm, to obtain an index value satisfying a preset sub-sampling point level condition; the preset fitting algorithm includes an interpolation fitting algorithm and a parabolic fitting algorithm; the target neighborhood is a region corresponding to a preset nearby range; Determine an equivalent sampling point offset total based on the array geometry, the time delay relationship and the time delay estimation value, determine a time value based on the equivalent sampling point offset total, the index value and a preset sampling rate, determine a sound source distance between the sound source to be positioned and a preset reference point based on the time value, a preset system delay and the sound speed.

[0012] In a second aspect, the present application provides a microphone array based sound source positioning and distance measurement device, comprising: An audio data generation module is configured to synchronously collect a frequency sweeping signal played by a sound source to be positioned by using a microphone array, to obtain each audio data to be processed, and to perform generalized cross-correlation and phase transform weighting processing on each audio data to be processed based on the frequency sweeping signal, to obtain a cross-correlation sequence corresponding to each channel, and then to construct a current processing set based on each cross-correlation sequence. A sequence processing module is configured to set a first sequence in the current processing set as a current reference sequence, to perform generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current processing set based on the current reference sequence, to obtain a current cross-correlation sequence group, to discard the current reference sequence, to set a first sequence in the new current cross-correlation sequence group obtained as a new current reference sequence, and then to set the current cross-correlation sequence group as a new current processing set, so as to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current processing set based on the current reference sequence, until only one sequence to be processed is included in the current processing set. A sequence up-sampling module is configured to perform an up-sampling operation on the sequence to be processed to meet a preset high multiple condition, to obtain an up-sampled sequence, to extract a time delay estimation value of a decimal sampling stage based on a first main peak position in the up-sampled sequence, to determine an incident angle of the sound source to be positioned relative to an array normal line based on the time delay estimation value, an inter-element distance, a sampling rate and a sound speed, and to perform integer time delay alignment and decimal time delay alignment on each cross-correlation sequence based on the time delay estimation value and the incident angle, to obtain a corresponding alignment result. A sound source distance determination module is configured to construct a target single-channel beam synthesis signal based on each alignment result, to determine a second main peak position corresponding to the target single-channel beam synthesis signal, and to determine a sound source distance based on the second main peak position, the time delay estimation value, a preset system delay and the sound speed.

[0013] In a third aspect, the present application provides an electronic device, comprising: A memory is configured to save a computer program. A processor is configured to execute the computer program to implement the microphone array based sound source positioning and distance measurement method.

[0014] In a fourth aspect, the present application provides a computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the above-mentioned microphone array-based sound source positioning and distance measurement method.

[0015] As can be seen from the above, before the microphone array-based sound source positioning and distance measurement is performed, the present application needs to use the microphone array to synchronously collect the swept-frequency signals played by the sound source to be positioned to obtain each audio data to be processed, and perform generalized cross-correlation and phase transform weighting processing on each audio data to be processed based on the swept-frequency signals to obtain a cross-correlation sequence corresponding to each channel respectively, and then construct a current set to be processed based on each cross-correlation sequence; set the first sequence in the current set to be processed as a current reference sequence, and perform generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current set to be processed based on the current reference sequence to obtain a current cross-correlation sequence group, then discard the current reference sequence, and set the first sequence in the new current cross-correlation sequence group obtained as a new current reference sequence, then set the current cross-correlation sequence group as a new current set to be processed, so as to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current set to be processed based on the current reference sequence, until only one sequence to be processed is included in the current set to be processed; perform high-multiple up-sampling on the sequence to be processed to obtain an up-sampled sequence, and extract a time delay estimation value of a decimal sampling level based on the first main peak position in the up-sampled sequence, so as to determine the incident angle of the sound source to be positioned relative to the array normal based on the time delay estimation value, the inter-element distance, the sampling rate and the sound speed, and perform integer time delay alignment and decimal time delay alignment on each cross-correlation sequence based on the time delay estimation value and the incident angle to obtain a corresponding alignment result respectively; construct a target single-channel beam synthesis signal based on each alignment result, and determine the second main peak position corresponding to the target single-channel beam synthesis signal, so as to determine the distance of the sound source based on the second main peak position, the time delay estimation value, the preset system delay and the sound speed.

[0016] It can be seen that, first, the sweep signal played by the sound source to be positioned is synchronously collected by using the microphone array to obtain each audio data to be processed, and the generalized cross-correlation and phase transform weighting processing are performed on each audio data to be processed based on the sweep signal to obtain the cross-correlation sequence corresponding to each channel, and then the current processing set is constructed based on each cross-correlation sequence; secondly, the first sequence in the current processing set is set as the current reference sequence, and the generalized cross-correlation and phase transform weighting processing are performed on the current remaining sequences in the current processing set based on the current reference sequence to obtain the current cross-correlation sequence group, then the current reference sequence is discarded, and the first sequence in the new current cross-correlation sequence group obtained is set as the new current reference sequence, and then the current cross-correlation sequence group is set as the new current processing set to re-jump to the step of performing the generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current processing set based on the current reference sequence, until only one sequence to be processed is included in the current processing set; thirdly, the sequence to be processed is high-multiple up-sampled to obtain an up-sampled sequence, and a fractional sampling stage time delay estimation value is extracted based on the first main peak position in the up-sampled sequence, so as to determine the incidence angle of the sound source to be positioned relative to the array normal based on the time delay estimation value, the inter-element distance, the sampling rate and the sound speed, and the integer time delay alignment and the fractional time delay alignment are respectively performed on each cross-correlation sequence based on the time delay estimation value and the incidence angle to obtain the corresponding alignment results; finally, the target single-channel beam synthesis signal is constructed based on each alignment result, and the second main peak position corresponding to the target single-channel beam synthesis signal is determined, so as to determine the sound source distance based on the second main peak position, the time delay estimation value, the preset system delay and the sound speed. In this way, the efficiency of sound source positioning and distance measurement based on the microphone array is improved, and the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on the provided drawings.

[0018] Figure 1 A flow chart of a specific sound source positioning and distance measurement method based on a microphone array is disclosed in the present application. Figure 2 A flow chart of a specific sound source positioning and distance measurement method based on a microphone array is disclosed in the present application. Figure 3 A specific iteration process schematic diagram is disclosed in the present application. Figure 4A specific beam synthesis signal generation flowchart disclosed by the present application; Figure 5 A microphone array based sound source positioning and distance measurement device structure schematic diagram disclosed by the present application; Figure 6 An electronic device structure diagram disclosed by the present application. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0020] At present, the sound source positioning technology aims to determine the spatial position or direction of the sound source by using the sound wave signals received by the microphone array. The common implementation paths include methods based on time difference of arrival, angle of arrival and signal strength, etc. In a low signal-to-noise ratio or strong reverberation environment, multipath reflection and noise can cause multiple competing peaks in the cross-correlation function, and the peak value corresponding to the true time delay is easy to be pulled off by interference, thereby causing the directional accuracy to decrease. At the same time, the traditional method has high requirements for array layout accuracy and synchronization, and the collaborative use of array multi-channel information is insufficient, which is easy to produce sidelobe interference and has limited robustness. Therefore, the present application provides a microphone array based sound source positioning and distance measurement method, which can improve the efficiency of sound source positioning and distance measurement in the process of microphone array based sound source positioning and distance measurement.

[0021] Referring to Figure 1 The embodiments of the present application disclose a microphone array based sound source positioning and distance measurement method, which comprises: Step S11, synchronously collecting the swept frequency signals played by the sound source to be positioned by using the microphone array to obtain each audio data to be processed, and performing generalized cross-correlation and phase transformation weighting processing on each audio data to be processed based on the swept frequency signals to obtain a cross-correlation sequence corresponding to each channel, and then constructing a current processing set based on each cross-correlation sequence.

[0022] In the present embodiment, the embodiments of the present application perform sound source positioning and distance measurement based on linear array microphones as hardware basis, and the corresponding flowchart is as shown in Figure 2 In one specific embodiment, the typical parameters are: the number of array elements N≥3, the array element spacing d≤0.1m, the sampling rate ≥48 kHz), both in software on a general purpose processor and in hardware using DSP / FPGA / ASIC, further, the embodiment of the application implements high-precision orientation and ranging of the sound source through the following steps: First, sound source excitation and synchronous acquisition: a known linear or logarithmic sweep signal (for example, Chirp signal) is played, the starting frequency, the terminal frequency and the duration time of which are configurable; N-channel audio data is synchronously acquired and clock synchronization and gain calibration are performed. Further, to improve the signal-to-noise ratio, a band-pass filter or pre-emphasis processing can be added at the front end to suppress out-of-band noise and enhance the energy of the effective frequency band.

[0023] Subsequently, the generalized cross-correlation (GCC-PHAT) is calculated for each channel based on the played signal as the reference signal, thereby obtaining N cross-correlation sequences . The GCC-PHAT typical frequency domain expression is: ; Wherein, is the frequency domain representation of the reference signal, is the frequency domain representation of the recording of a certain channel, and the superscript represents conjugate, represents inverse Fourier transform, is a small positive number to prevent the denominator from being zero. In this way, the phase weighting suppresses environmental noise and amplitude fluctuations to highlight the phase consistency corresponding to the time delay.

[0024] Specifically, the sweep signal played by the sound source to be positioned is synchronously acquired by using the microphone array to obtain each audio data to be processed, and the generalized cross-correlation and phase transform weighting processing are performed on each audio data to be processed based on the sweep signal to obtain the cross-correlation sequence corresponding to each channel, which can include: the sweep signal played by the sound source to be positioned is synchronously acquired by using the microphone array based on the preset starting frequency, the preset terminal frequency and the preset duration; the sweep signal includes a linear sweep signal and a logarithmic sweep signal; the clock synchronization calibration and gain consistency calibration are performed on the sweep signal corresponding to each channel to obtain the processed signal, and the band-pass filtering operation and the pre-emphasis filtering operation are performed on each processed signal to obtain each audio data to be processed; the sweep signal is set as the reference signal, and the Fourier transform is performed on each audio data to be processed based on the reference signal to obtain the corresponding frequency domain representation, and then the cross-power spectrum is determined based on the conjugate multiplication result between each frequency domain representation and each audio data to be processed; the modulus value corresponding to the cross-power spectrum is determined, and the to-be-transformed result is determined based on the cross-power spectrum, the modulus value and the preset small positive number, then the phase transform weighting is performed on the to-be-transformed result to obtain the weighted cross-power spectrum, and the inverse Fourier transform is performed on the weighted cross-power spectrum to obtain the corresponding cross-correlation sequence.

[0025] Step S12, set the first sequence in the current to-be-processed set as the current reference sequence, and perform generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence to obtain a current cross-correlation sequence group, then discard the current reference sequence, set the first sequence in the obtained new current cross-correlation sequence group as a new current reference sequence, and set the current cross-correlation sequence group as a new current to-be-processed set to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set by the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set.

[0026] In this embodiment, the embodiment of the application aims to robustly and accurately extract a small value representing the time delay between adjacent elements from the obtained cross-correlation sequence The core is to gradually integrate multi-channel information through iterative cross-correlation operation, and then to strengthen the peak value corresponding to the true time delay, and the corresponding iterative flowchart is as shown in Figure 3 First, the obtained N sequences are set as the iteration input set of the first round of iteration. Subsequently, in each round of iteration, the first sequence in the current input set is taken as a reference, and GCC-PHAT calculation is performed with each of the remaining sequences in the set to obtain a new cross-correlation sequence group. Subsequently, the reference sequence is discarded, and the newly obtained sequence set is taken as the input of the next round of iteration. This process is repeated, and each round of iteration reduces the number of "channels" involved by 1. Finally, when the iteration is performed to the output set with only a single sequence left, the iteration is terminated. The final sequence is denoted as . It is worth mentioning that the above sequence condenses the commonality between all original channels and the main time delay, and has a higher peak signal-to-noise ratio.

[0027] Specifically, the first sequence in the current set to be processed is set as the current reference sequence. Based on the current reference sequence, generalized cross-correlation and phase transformation weighting are performed on the remaining sequences in the current set to be processed to obtain a current cross-correlation sequence group. Then, the current reference sequence is discarded, and the first sequence in the newly obtained current cross-correlation sequence group is set as the new current reference sequence. This current cross-correlation sequence group is then set as the new current set to be processed, and the process jumps back to the step of performing generalized cross-correlation and phase transformation weighting on the remaining sequences in the current set based on the current reference sequence, until the current set to be processed contains only one sequence to be processed. This may include: setting the first sequence in the current set to be processed as the current reference sequence, and then comparing the current reference sequence with the current sequence to be processed... Each sequence in the processing set is subjected to generalized cross-correlation to obtain the corresponding cross-correlation processing result. Then, a phase transformation weighting operation is performed on the cross-correlation processing result to obtain the corresponding cross-correlation sequence. A current cross-correlation sequence group is constructed based on each cross-correlation sequence. Then, the current reference sequence in the current cross-correlation sequence group is discarded to obtain a new current cross-correlation sequence group. The first sequence in the current cross-correlation sequence group is set as the new current current set to be processed. The process then jumps back to the current reference sequence to perform generalized cross-correlation and phase transformation weighting processing on the remaining sequences in the current set to be processed until the current set to be processed contains only one sequence to be processed. The sequence to be processed includes the common time delay information corresponding to each channel.

[0028] Step S13: Perform an upsampling operation on the sequence to be processed to meet the preset high multiple condition to obtain an upsampled sequence. Extract the time delay estimate of the fractional sampling level based on the first main peak position in the upsampled sequence. Determine the incident angle of the sound source to be located relative to the array normal based on the time delay estimate, the array element spacing, the sampling rate and the sound velocity. Perform integer time delay alignment and fractional time delay alignment on each of the cross-correlation sequences based on the time delay estimate and the incident angle to obtain the corresponding alignment results.

[0029] In this embodiment, the final sequence of this application embodiment Perform Q-fold upsampling (e.g., Q=8 or 16), and precisely search for the main peak position in the upsampled sequence. The final output fractional-sample delay estimate is... .

[0030] Specifically, the up-sampling operation satisfying the preset high-multiple condition is performed on the to-be-processed sequence to obtain an up-sampled sequence, and a time delay estimation value of the fractional sampling stage is extracted based on a first main peak position in the up-sampled sequence, which can include: performing the up-sampling operation satisfying the preset high-multiple condition on the to-be-processed sequence to obtain the up-sampled sequence, searching for an index position corresponding to a global maximum value in the up-sampled sequence, and setting the index position as the first main peak position; the preset high-multiple condition is a condition determined based on a preset time delay estimation accuracy; the time delay estimation value of the fractional sampling stage is determined based on the first main peak position and an up-sampling multiple; and the time delay estimation value is used to represent a proportional relationship between an actual time delay of received signals between adjacent array elements and a sampling interval.

[0031] Subsequently, the embodiment of the present application can determine the incident angle of the sound source relative to the array normal based on the obtained time delay estimation value . ; wherein c is the sound speed.

[0032] Then, the embodiment of the present application needs to perform time domain alignment on the cross-correlation result: first, integer delay alignment is performed: for the first cross-correlation result, the channel is subjected to an integer sampling delay of points to achieve coarse alignment between channels. In implementation, zero padding or cyclic shift method can be selected, wherein the delay direction is determined by the sign of . Subsequently, fractional delay alignment is performed: for channel i, the fractional delay is , wherein . Then, a first-order all-pass filter is applied to it to achieve fractional delay: ; wherein .

[0033] It is worth mentioning that the embodiment of the present application can use a higher-order (such as Thiran) all-pass structure to improve the group delay flatness.

[0034] ​Specifically, the integer time delay alignment and the decimal time delay alignment are respectively performed on each cross-correlation sequence based on the time delay estimation value and the incident angle to obtain a corresponding alignment result, which can include: determining an integer sampling delay amount corresponding to each channel based on the time delay estimation value and the array element number, then performing an integer sampling point alignment operation on each cross-correlation sequence based on a preset alignment operation and the integer sampling delay amount, a preset delay direction and the incident angle to obtain an integer time delay alignment result; the delay direction is a direction determined based on the sign of the time delay estimation value; the preset alignment operation includes a zero padding mode and a cyclic shift mode; determining a decimal delay amount corresponding to each channel based on the integer sampling delay amount, the time delay estimation value and the channel number, then performing a decimal time delay alignment on the integer time delay alignment result based on a preset first-order all-pass filter and a preset filter coefficient to obtain a corresponding alignment result; the preset filter coefficient is a coefficient determined based on a preset formula and the decimal delay amount corresponding to the channel.

[0035] Step S14, constructing a target single-channel beam synthesis signal based on each alignment result, and determining a second main peak value position corresponding to the target single-channel beam synthesis signal, to determine the sound source distance based on the second main peak value position, the time delay estimation value, a preset system delay and the sound speed.

[0036] In this embodiment, the N-channel signals compensated by the integer and decimal delay are added to obtain a single-channel beam synthesis signal, and the corresponding beam synthesis signal generation flowchart is as shown in Figure 4 Specifically, the target single-channel beam synthesis signal can be constructed based on each alignment result, which can include: point-by-point and weighted addition of each sequence in each alignment result to obtain an initial single-channel beam synthesis signal; the initial single-channel beam synthesis signal includes energy corresponding to the sound source to be positioned; performing a signal enhancement operation on the initial single-channel beam synthesis signal to obtain a signal enhancement result, and determining interference and noise from other directions in the signal enhancement result, then suppressing the interference and noise to obtain the target single-channel beam synthesis signal.

[0037] Subsequently, the main peak position index of the beam synthesis signal needs to be searched , and interpolation or parabolic fitting can be used in the peak neighborhood to improve accuracy. The determination formula of the sound source distance is as follows: ; Wherein, is a system inherent fixed delay (play, acquisition link), which can be obtained by measurement, and c is the sound speed.

[0038] Specifically, the determining the second main peak position corresponding to the target single-channel beamforming signal, and determining the sound source distance based on the second main peak position, the time delay estimation value, the preset system delay and the sound speed can include: searching an index corresponding to a maximum amplitude in a full sequence range of the target single-channel beamforming signal, and setting the index as the second main peak position; performing peak position positioning in a target neighborhood of the second main peak position by using a preset fitting algorithm to obtain an index value satisfying a preset sub-sampling point level condition; the preset fitting algorithm includes an interpolation fitting algorithm and a parabolic fitting algorithm; the target neighborhood is a region corresponding to a preset nearby range; determining an equivalent sampling point offset total amount based on an array geometry, a time delay relationship and the time delay estimation value, and determining a time value based on the equivalent sampling point offset total amount, the index value and a preset sampling rate, and determining the sound source distance between the to-be-positioned sound source and the preset reference point based on the time value, the preset system delay and the sound speed.

[0039] As can be seen from the above, the embodiments of the present application first need to use the microphone array to synchronously collect the sweep signal played by the to-be-positioned sound source, obtain each to-be-processed audio data, and perform generalized cross-correlation and phase transform weighting processing on each to-be-processed audio data based on the sweep signal to obtain a cross-correlation sequence corresponding to each channel respectively, and then construct a current to-be-processed set based on each cross-correlation sequence; secondly, the first sequence in the current to-be-processed set is set as a current reference sequence, and the current remaining sequences in the current to-be-processed set are processed by generalized cross-correlation and phase transform weighting based on the current reference sequence to obtain a current cross-correlation sequence group, then the current reference sequence is discarded, and the first sequence in the new current cross-correlation sequence group obtained is set as a new current reference sequence, then the current cross-correlation sequence group is set as a new current to-be-processed set, so as to jump back to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set by the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set; thirdly, the to-be-processed sequence is high-multiple up-sampled to obtain an up-sampled sequence, and a time delay estimation value of a decimal sampling level is extracted based on a first main peak position in the up-sampled sequence, the incident angle of the to-be-positioned sound source relative to the array normal line is determined based on the time delay estimation value, the inter-element distance, the sampling rate and the sound speed, and the integer time delay alignment and the decimal time delay alignment are performed on each cross-correlation sequence based on the time delay estimation value and the incident angle, to obtain a corresponding alignment result respectively; finally, a target single-channel beamforming signal is constructed based on each alignment result, and a second main peak position corresponding to the target single-channel beamforming signal is determined, and the sound source distance is determined based on the second main peak position, the time delay estimation value, the preset system delay and the sound speed. In this way, the efficiency of sound source positioning and distance measurement based on the microphone array is improved, and the user experience is improved.

[0040] Correspondingly, referring to Figure 5As shown, the application also provides a microphone array-based sound source positioning and distance measurement device, comprising: An audio data generation module 11 is configured to synchronously collect a sweep signal played by a sound source to be positioned by using a microphone array, to obtain each to-be-processed audio data, and to perform generalized cross-correlation and phase transform weighting processing on each to-be-processed audio data based on the sweep signal, to obtain a cross-correlation sequence corresponding to each channel respectively, and then to construct a current to-be-processed set based on each cross-correlation sequence. A sequence processing module 12 is configured to set a first sequence in the current to-be-processed set as a current reference sequence, to perform generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, to obtain a current cross-correlation sequence group, to discard the current reference sequence, and to set a first sequence in the obtained new current cross-correlation sequence group as a new current reference sequence, and then to set the current cross-correlation sequence group as a new current to-be-processed set, so as to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set. A sequence upsampling module 13 is configured to perform an upsampling operation on the to-be-processed sequence to meet a preset high-multiple condition, to obtain an upsampled sequence, to extract a time delay estimation value of a decimal sampling stage based on a first main peak value position in the upsampled sequence, to determine an incident angle of the sound source to be positioned relative to an array normal line based on the time delay estimation value, an inter-element distance, a sampling rate and a sound speed, and to perform integer time delay alignment and decimal time delay alignment on each cross-correlation sequence based on the time delay estimation value and the incident angle, to obtain a corresponding alignment result respectively. A sound source distance determination module 14 is configured to construct a target single-channel beam synthesis signal based on each alignment result, and to determine a second main peak value position corresponding to the target single-channel beam synthesis signal, to determine a sound source distance based on the second main peak value position, the time delay estimation value, a preset system delay and the sound speed.

[0041] In some embodiments, the audio data generation module 11 can specifically include: A sweep signal collection unit is configured to synchronously collect a sweep signal played by a sound source to be positioned by using a microphone array and based on a preset starting frequency, a preset terminal frequency and a preset duration period; the sweep signal includes a linear sweep signal and a logarithmic sweep signal. An audio data generation sub-unit is configured to perform clock synchronization calibration and gain consistency calibration on the sweep signal corresponding to each channel respectively, to obtain a to-be-processed signal, and to perform band-pass filtering and pre-emphasis filtering on each to-be-processed signal, to obtain each to-be-processed audio data. a cross-power spectrum determination unit, configured to set the sweep signal as a reference signal, and perform Fourier transform on each of the to-be-processed audio data based on the reference signal to obtain a corresponding frequency domain representation, and then determine a cross-power spectrum based on a conjugate multiplication result between each of the frequency domain representations and the to-be-processed audio data; a cross-correlation sequence generation unit, configured to determine a modulus value corresponding to the cross-power spectrum, and determine a to-be-transformed result based on the cross-power spectrum, the modulus value and a preset minimum positive number, then perform phase transform weighting on the to-be-transformed result to obtain a weighted cross-power spectrum, and perform inverse Fourier transform on the weighted cross-power spectrum to obtain a corresponding cross-correlation sequence.

[0042] In some specific embodiments, the sequence processing module 12 can specifically include: a cross-correlation processing result generation unit, configured to set a first sequence in a current to-be-processed set as a current reference sequence, and perform generalized cross-correlation operation on the current reference sequence and each sequence in the current to-be-processed set respectively to obtain a corresponding cross-correlation processing result, and then perform phase transform weighting operation on the cross-correlation processing result to obtain a corresponding cross-correlation sequence respectively; a current cross-correlation sequence group construction unit, configured to construct a current cross-correlation sequence group based on each of the cross-correlation sequences, then discard the current reference sequence in the current cross-correlation sequence group to obtain a new current cross-correlation sequence group, and set a first sequence in the current cross-correlation sequence group as a new current reference sequence; a step jump unit, configured to set the current cross-correlation sequence group as a new current to-be-processed set, and jump back to the step of performing generalized cross-correlation and phase transform weighting processing on the current reference sequence and the current remaining sequences in the current to-be-processed set until the current to-be-processed set includes only one to-be-processed sequence; the to-be-processed sequence includes common time delay information corresponding to each channel.

[0043] In some specific embodiments, the sequence up-sampling module 13 can specifically include: a first main peak position determination unit, configured to perform up-sampling operation on the to-be-processed sequence to obtain an up-sampled sequence, the up-sampling operation satisfying a preset high-multiplication condition, and search for an index position corresponding to a global maximum value in the up-sampled sequence, and set the index position as a first main peak position; the preset high-multiplication condition is a condition determined based on a preset time delay estimation accuracy; a time delay estimation value generation unit, configured to determine a time delay estimation value of a decimal sampling stage based on the first main peak position and an up-sampling multiple; the time delay estimation value is used to represent a proportional relationship between an actual time delay of received signals between adjacent array elements and a sampling interval.

[0044] In some embodiments, the sequence up-sampling module 13 can specifically include: The first alignment result generation unit is configured to determine an integer sampling delay amount corresponding to each channel based on the time delay estimation value and the array element serial number, and then perform an integer sampling point alignment operation on each of the cross-correlation sequences based on the integer sampling delay amount, a preset delay direction, and the incident angle by using a preset alignment operation, to obtain an integer time delay alignment result; the delay direction is determined based on the sign of the time delay estimation value; and the preset alignment operation includes a zero padding mode and a cyclic shift mode. The second alignment result generation unit is configured to determine a fractional delay amount corresponding to each channel based on the integer sampling delay amount, the time delay estimation value, and the channel serial number, and then perform a fractional time delay alignment on the integer time delay alignment result based on a preset filter coefficient and the integer sampling delay amount by using a preset first-order all-pass filter, to obtain a corresponding alignment result; the preset filter coefficient is determined based on a preset formula and the fractional delay amount corresponding to each channel.

[0045] In some embodiments, the sound source distance determination module 14 can specifically include: The sequence processing unit is configured to perform point-by-point and weighted addition on each sequence in each of the alignment results, to obtain an initial single-channel beamforming signal; the initial single-channel beamforming signal includes energy corresponding to the sound source to be positioned. The single-channel beamforming signal generation unit is configured to perform a signal enhancement operation on the initial single-channel beamforming signal, to obtain a signal enhancement result, and determine interference and noise from other directions in the signal enhancement result, and then suppress the interference and the noise, to obtain a target single-channel beamforming signal.

[0046] In some embodiments, the sound source distance determination module 14 can specifically include: The second main peak position determination unit is configured to search for an index corresponding to a maximum amplitude in a full sequence range of the target single-channel beamforming signal, and set the index as a second main peak position. The index value determination unit is configured to perform peak position positioning in a target neighborhood of the second main peak position by using a preset fitting algorithm, to obtain an index value satisfying a preset sub-sampling point level condition; the preset fitting algorithm includes an interpolation fitting algorithm and a parabolic fitting algorithm; and the target neighborhood is a region corresponding to a preset nearby range. The sound source distance determination sub-unit is configured to determine an equivalent sampling point offset total based on array geometry, time delay relationship and the time delay estimation value, determine a time value based on the equivalent sampling point offset total, the index value and a preset sampling rate, and determine a sound source distance between a to-be-positioned sound source and a preset reference point based on the time value, a preset system delay and the sound velocity.

[0047] Further, the application further discloses an electronic device, Figure 6 is an electronic device 20 structure diagram shown according to an exemplary embodiment, the contents in the figure cannot be considered as any limitation to the use range of the application. The electronic device 20, specifically can include: at least one processor 21, at least one memory 22, power supply 23, communication interface 24, input output interface 25 and communication bus 26. Wherein, the memory 22 is used for storing computer program, the computer program is loaded and executed by the processor 21, to realize the related steps in the foregoing any embodiment disclosed based on microphone array sound source positioning and distance measurement method. In addition, the electronic device 20 in the embodiment specifically can be electronic computer.

[0048] In the embodiment, the power supply 23 is used for providing working voltage for each hardware device on the electronic device 20; the communication interface 24 can create data transmission channel between the electronic device 20 and external device, and the communication protocol followed is any communication protocol applicable to the technical scheme of the application, which is not specifically limited here; the input output interface 25 is used for obtaining external input data or outputting data to the outside, and the specific interface type can be selected according to specific application needs, which is not specifically limited here.

[0049] In addition, the memory 22 as the carrier of resource storage can be read-only memory, random access memory, disk or optical disk, etc., and the resources stored thereon can include operating system 221, computer program 222, etc., and the storage mode can be temporary storage or permanent storage.

[0050] Wherein, the operating system 221 is used for managing and controlling each hardware device on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. The computer program 222 can further include computer programs capable of completing other specific work in addition to the computer programs capable of completing the microphone array based sound source positioning and distance measurement method executed by the electronic device 20 disclosed in any of the foregoing embodiments.

[0051] Further, the application also discloses a computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to realize the microphone array based sound source positioning and distance measuring method disclosed above. For the specific steps of the method, refer to the corresponding content disclosed in the foregoing embodiments, which will not be repeated here.

[0052] The various embodiments are described in the specification by progressive stages, and each embodiment focuses on the difference from other embodiments. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts refer to the method part.

[0053] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.

[0054] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0055] Finally, it should be noted that, in this document, relationship terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0056] The technical solutions provided by the present application are described in detail above, and the principles and implementation manners of the present application are described by using specific examples. The above description of the examples is only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed, and the above description of the content of the specification should not be understood as a limitation on the present application.

Claims

1. A method for sound source localization and distance measurement based on a microphone array, characterized in that, The method comprises the following steps: Synchronously collecting, by using a microphone array, a sweep signal played by a sound source to be positioned to obtain each to-be-processed audio data, and performing generalized cross-correlation and phase transform weighting processing on each to-be-processed audio data based on the sweep signal to obtain a cross-correlation sequence corresponding to each channel respectively, and then constructing a current to-be-processed set based on each cross-correlation sequence; Setting a first sequence in the current to-be-processed set as a current reference sequence, and performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence to obtain a current cross-correlation sequence group, then discarding the current reference sequence, setting a first sequence in the obtained new current cross-correlation sequence group as a new current reference sequence, and then setting the current cross-correlation sequence group as a new current to-be-processed set to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set; Performing an upsampling operation on the to-be-processed sequence to obtain an upsampled sequence, the upsampling operation satisfying a preset high-multiple condition, extracting a time delay estimation value of a decimal sampling stage based on a first main peak value position in the upsampled sequence, determining an incident angle of the sound source to be positioned relative to an array normal line based on the time delay estimation value, an inter-element distance, a sampling rate and a sound speed, and performing integer time delay alignment and decimal time delay alignment on each cross-correlation sequence based on the time delay estimation value and the incident angle to obtain a corresponding alignment result respectively; Constructing a target single-channel beamforming signal based on each alignment result, and determining a second main peak value position corresponding to the target single-channel beamforming signal, and determining a sound source distance based on the second main peak value position, the time delay estimation value, a preset system delay and the sound speed.

2. The microphone array based sound source positioning and distance measurement method according to claim 1, characterized in that, The method comprises the following steps: Synchronously collecting, by using a microphone array, a sweep signal played by a sound source to be positioned to obtain each to-be-processed audio data, and performing generalized cross-correlation and phase transform weighting processing on each to-be-processed audio data based on the sweep signal to obtain a cross-correlation sequence corresponding to each channel respectively, and then constructing a current to-be-processed set based on each cross-correlation sequence; Synchronously collecting, by using a microphone array, a sweep signal played by a sound source to be positioned based on a preset starting frequency, a preset ending frequency and a preset duration; the sweep signal comprises a linear sweep signal and a logarithmic sweep signal; Performing clock synchronization calibration and gain consistency calibration on the sweep signal corresponding to each channel to obtain a to-be-processed signal, and performing band-pass filtering and pre-emphasis filtering on each to-be-processed signal to obtain each to-be-processed audio data; Setting the sweep signal as a reference signal, and performing Fourier transform on each to-be-processed audio data based on the reference signal to obtain a corresponding frequency domain representation, and then determining a cross-power spectrum based on a conjugate multiplication result between each frequency domain representation and each to-be-processed audio data; Determine the modulus corresponding to the cross-power spectrum, and determine the to-be-transformed result based on the cross-power spectrum, the modulus, and a preset minimum positive number, then perform phase transformation weighting on the to-be-transformed result to obtain a weighted cross-power spectrum, and perform inverse Fourier transformation on the weighted cross-power spectrum to obtain the corresponding cross-correlation sequence.

3. The microphone array based sound source positioning and distance measurement method of claim 1, wherein, The first sequence in the current to-be-processed set is set as the current reference sequence, and the generalized cross-correlation and phase transformation weighting processing are performed on the current remaining sequences in the current to-be-processed set based on the current reference sequence to obtain a current cross-correlation sequence group, then the current reference sequence is discarded, and the first sequence in the obtained new current cross-correlation sequence group is set as a new current reference sequence, then the current cross-correlation sequence group is set as a new current to-be-processed set, so as to jump back to the step of performing the generalized cross-correlation and phase transformation weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set, including: The first sequence in the current to-be-processed set is set as the current reference sequence, and the generalized cross-correlation operation is performed on each sequence in the current to-be-processed set based on the current reference sequence to obtain corresponding cross-correlation processing results, then the phase transformation weighting operation is performed on the cross-correlation processing results to obtain corresponding cross-correlation sequences respectively; A current cross-correlation sequence group is constructed based on the cross-correlation sequences, then the current reference sequence in the current cross-correlation sequence group is discarded to obtain a new current cross-correlation sequence group, and the first sequence in the current cross-correlation sequence group is set as a new current reference sequence; The current cross-correlation sequence group is set as a new current to-be-processed set, and the step of performing the generalized cross-correlation and phase transformation weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence is jumped back, until only one to-be-processed sequence is included in the current to-be-processed set; the to-be-processed sequence includes the common time delay information corresponding to each channel.

4. The microphone array based sound source positioning and distance measurement method of claim 1, wherein, The upsampling operation satisfying the preset high-multiple condition is performed on the to-be-processed sequence to obtain an upsampled sequence, and the time delay estimation value of the fractional sampling stage is extracted based on the first main peak value position in the upsampled sequence, including: The upsampling operation satisfying the preset high-multiple condition is performed on the to-be-processed sequence to obtain an upsampled sequence, and the index position corresponding to the global maximum value in the upsampled sequence is searched, and the index position is set as the first main peak value position; the preset high-multiple condition is a condition determined based on a preset time delay estimation accuracy; The time delay estimation value of the fractional sampling stage is determined based on the first main peak value position and the upsampling multiple; the time delay estimation value is used to represent the proportional relationship between the actual time delay of the received signals between adjacent array elements and the sampling interval.

5. The microphone array based sound source positioning and distance measurement method of claim 1, wherein, The integer time delay alignment and the fractional time delay alignment are respectively performed on each cross-correlation sequence based on the time delay estimation value and the incident angle to obtain corresponding alignment results respectively, including: determining an integer sampling delay amount corresponding to each channel based on the time delay estimation value and the array element serial number, and then performing an integer sampling point alignment operation on each of the cross-correlation sequences based on the integer sampling delay amount, a preset delay direction, and the incident angle by using a preset alignment operation, to obtain an integer time delay alignment result; the delay direction is determined based on the sign of the time delay estimation value; the preset alignment operation includes a zero padding mode and a cyclic shift mode; determining a fractional delay amount corresponding to each channel based on the integer sampling delay amount, the time delay estimation value, and the channel serial number, and then performing a fractional time delay alignment on the integer time delay alignment result based on a preset first-order all-pass filter and a preset filter coefficient, to obtain a corresponding alignment result; the preset filter coefficient is determined based on a preset formula and the fractional delay amount corresponding to each channel.

6. The microphone array based sound source positioning and distance measurement method of claim 5, wherein, The target single-channel beamforming signal is constructed based on each of the alignment results, including: point-by-point and weighted addition of each sequence in each of the alignment results is performed to obtain an initial single-channel beamforming signal; the initial single-channel beamforming signal includes energy corresponding to the sound source to be positioned; signal enhancement is performed on the initial single-channel beamforming signal to obtain a signal enhancement result, and interference and noise from other directions in the signal enhancement result are determined, and then the interference and the noise are suppressed to obtain a target single-channel beamforming signal.

7. The microphone array based sound source positioning and distance measurement method according to any one of claims 1 to 6, characterized in that, The second main peak value position corresponding to the target single-channel beamforming signal is determined, and the sound source distance is determined based on the second main peak value position, the time delay estimation value, a preset system delay, and the sound speed, including: searching for an index corresponding to a maximum amplitude value in a full sequence range of the target single-channel beamforming signal, and setting the index as a second main peak value position; peak position positioning is performed in a target neighborhood of the second main peak value position by using a preset fitting algorithm to obtain an index value satisfying a preset sub-sampling point level condition; the preset fitting algorithm includes an interpolation fitting algorithm and a parabolic fitting algorithm; the target neighborhood is a region corresponding to a preset nearby range; an equivalent sampling point offset total amount is determined based on array geometry, time delay relationship, and the time delay estimation value, and a time value is determined based on the equivalent sampling point offset total amount, the index value, and a preset sampling rate, and a sound source distance between the sound source to be positioned and a preset reference point is determined based on the time value, a preset system delay, and the sound speed.

8. A microphone array based sound source localization and distance measurement apparatus, characterized in that, including: an audio data generation module configured to synchronously collect a sweep signal played by a sound source to be positioned by using a microphone array to obtain each of audio data to be processed, perform generalized cross-correlation and phase transformation weighting processing on each of the audio data to be processed based on the sweep signal to obtain a cross-correlation sequence corresponding to each channel, and then construct a current set to be processed based on each of the cross-correlation sequences; The sequence processing module is configured to set a first sequence in a current to-be-processed set as a current reference sequence, perform generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set based on the current reference sequence, obtain a current cross-correlation sequence group, discard the current reference sequence, set a first sequence in a new current cross-correlation sequence group obtained as a new current reference sequence, and set the current cross-correlation sequence group as a new current to-be-processed set to re-jump to the step of performing generalized cross-correlation and phase transform weighting processing on the current remaining sequences in the current to-be-processed set by using the current reference sequence, until only one to-be-processed sequence is included in the current to-be-processed set; The sequence up-sampling module is configured to perform an up-sampling operation on the to-be-processed sequence to obtain an up-sampled sequence, extract a time delay estimation value of a fractional sampling stage based on a first main peak position in the up-sampled sequence, determine an angle of incidence of the to-be-positioned sound source relative to an array normal line based on the time delay estimation value, an inter-element distance, a sampling rate and a sound speed, and perform integer time delay alignment and fractional time delay alignment on each of the cross-correlation sequences based on the time delay estimation value and the angle of incidence to obtain corresponding alignment results, respectively. The sound source distance determination module is configured to construct a target single-channel beamforming signal based on the alignment results, determine a second main peak position corresponding to the target single-channel beamforming signal, and determine a sound source distance based on the second main peak position, the time delay estimation value, a preset system delay and the sound speed.

9. An electronic device, comprising: The memory is configured to save a computer program. The processor is configured to execute the computer program to implement the microphone array-based sound source positioning and distance measurement method according to any one of claims 1 to 7. The memory is configured to save a computer program.

10. A computer-readable storage medium, characterized in that, The processor is configured to execute the computer program to implement the microphone array-based sound source positioning and distance measurement method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Positioning method and device for physical loudspeaker and virtual loudspeaker, equipment and storage medium

    CN122063541A