Sound source direction finding and positioning method and system for complex environments
By calculating the sound source synchronization sequence and sound intensity interference of the microphone array, combined with fixed beam formation, noise interference is suppressed, the accuracy and stability of sound source positioning in complex environments are solved, and accurate positioning is achieved in a multi-sound source environment.
Patent Information
- Application Number
- CN202510715166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In complex environments, background noise and reverb interference significantly reduce the accuracy and stability of sound source positioning. The traditional method's delay summing beamforming algorithm results in excessive weighting of noise data, resulting in sound source direction finding positioning errors.
By calculating the maximum reception interval between the microphones, short-term audio similarity, sound source synchronization sequence and sound intensity interference, combined with the fixed beam formation of the microphone array, noise interference is suppressed and pure sound source signal components are extracted to improve positioning accuracy.
The stability and accuracy of sound source positioning are significantly improved in complex noise environments, the adaptability and reliability of the system are enhanced, and the target sound source position can be accurately identified in a multi-sound source environment.
Smart Images

Figure CN120233305B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and in particular to a method and system for sound source direction finding and positioning in complex environments. Background Art
[0002] Sound source localization technology typically relies on an array of multiple microphones to receive sound signals. By precisely analyzing information such as the phase difference and time difference between signals from different sensors, the system can calculate the specific location of the sound source. However, in complex environments, interference from background noise and reverberation can significantly reduce positioning accuracy, making sound source localization a significant challenge. Core indicators of sound source localization include spatial resolution and noise immunity. Spatial resolution determines the minimum angle at which the system can distinguish adjacent sound sources. High spatial resolution means that the system can more accurately locate sound sources and accurately identify targets even in environments with multiple sound sources. Noise immunity is key to maintaining the stability and reliability of sound source localization systems in complex noisy environments. Improving the system's noise immunity can not only effectively suppress the interference of background noise, but also enhance the accuracy of sound source positioning, ensuring the system's stable operation in various complex scenarios.
[0003] Real-time sound source localization requires real-time location of the sound source. Therefore, simple algorithms are required to process speech signals. Traditional methods typically use Delay-and-Sum Beamforming (DSB) for real-time processing. However, because traditional DSB algorithms apply the same weight to the speech signal received by each microphone, this weighting can overweight noise data. This can cause residual incoherent noise in the processed signal, leading to errors in subsequent sound source direction finding and localization. Summary of the Invention
[0004] In view of the above, it is necessary to provide a sound source direction finding and positioning method and system for complex environments to solve the above problems.
[0005] In a first aspect, the present application provides a method for sound source direction finding and positioning in a complex environment, the method comprising:
[0006] The sound intensity of each microphone with a preset duration is combined into a sound source signal sequence;
[0007] Based on the distance between two adjacent microphones, combined with the speed of sound propagation and the frequency of collecting sound intensity, the maximum receiving interval between two adjacent microphones is obtained;
[0008] Based on the numerical value of the maximum receiving interval, the sound source signal sequences of the two microphones are shifted in sequence to obtain the short-time audio similarity between the two microphones after each shift;
[0009] Based on the distribution of short-term audio similarity between each microphone and the microphone after it, the sound source synchronization sequence and the sound source frequency sequence of each microphone are extracted, and the sound quality deviation of each microphone is obtained according to the difference between the two sequences;
[0010] Select one microphone as the reference microphone, analyze the sound source synchronization sequence of all remaining microphones based on the reference microphone, and obtain the sound source similarity sequence and delay data length of each microphone;
[0011] Obtaining a sound intensity distribution sequence for each microphone based on the arrangement values of elements at the same position in the sound source similarity sequences of all microphones, analyzing the discreteness of the elements in the sound intensity distribution sequence, and obtaining the sound intensity interference degree of each microphone in combination with the sound quality deviation degree;
[0012] Based on the sound intensity interference of each microphone and the length of the time delay data, a fixed beam of the microphone array is obtained, and the direction with the maximum output power is taken as the sound source direction. Combined with the far-field model, the distance between the reference microphone and the sound source is determined.
[0013] The maximum receiving interval between two adjacent microphones is obtained as follows:
[0014] Obtain the distance between two adjacent microphones; calculate the ratio of the distance to the speed of sound propagation in the air; and multiply the ratio by the frequency at which the microphones collect data as the maximum receiving interval between the two adjacent microphones.
[0015] The short-term audio similarity between the two microphones after each movement is obtained as follows:
[0016]
[0017]
[0018] Where, represents the short-term audio similarity between the i-th microphone sound source signal sequence at the j-th position and the i+1-th microphone sound source signal sequence; The data sequence representing the sound source signal sequence of the i-th microphone from the j-th position to the n-th position; represents the data sequence of the i+1th microphone sound source signal sequence from the 1st position to the n-j+1th position; n represents the length of the microphone sound source signal sequence; represents the similarity function; represents the short-term audio similarity between the i+1th microphone sound source signal sequence at the jth position and the i-th microphone sound source signal sequence; The data sequence representing the sound source signal sequence of the i+1th microphone from the jth position to the nth position; represents the data sequence of the i-th microphone sound source signal sequence from the 1st position to the n-j+1th position; j represents the position number, ; L represents the maximum receiving interval between the i-th microphone and the i+1-th microphone.
[0019] The extraction of the sound source synchronization sequence and the sound source frequency sequence of each microphone is specifically as follows:
[0020] The elements of the sound source signal sequence of each microphone and the next microphone corresponding to the position of each microphone when the short-time audio similarity is the largest are respectively used to form a sound source synchronization sequence of each microphone and the next microphone;
[0021] The same frequency with non-zero amplitude between the sound source synchronization sequence of each microphone and the next microphone is obtained, and the amplitude of the same frequency of the two microphones is processed by inverse Fourier transform to obtain the sound source frequency sequence of each microphone.
[0022] The sound quality deviation of each microphone is specifically a distance measure between a sound source synchronization sequence and a sound source same-frequency sequence corresponding to each microphone.
[0023] The sound source similarity sequence and time delay data length of each microphone are obtained as follows:
[0024] Extracting the intersection of the element serial numbers of the sound source synchronization sequence elements of each microphone based on the reference microphone in the sound source signal sequence, and combining all the elements with serial numbers in the sound source signal sequence of each microphone in the intersection into a sound source similarity sequence for each microphone;
[0025] The sequence number of the first element of the sound source similarity sequence in the sound source signal sequence is used as the time delay data length of the corresponding microphone.
[0026] The sound intensity distribution sequence of each microphone is obtained as follows:
[0027] The elements with the same sequence number in the sound source similarity sequence of each microphone are sorted in ascending order, and the corresponding elements in the sound source similarity sequence are replaced with the sorted values to obtain the sound intensity distribution sequence of each microphone.
[0028] The specific process of obtaining the sound intensity interference degree of each microphone is as follows:
[0029] The sound quality deviation of each microphone is forward fused with the discrete degree of the corresponding sound intensity distribution sequence to obtain the sound intensity interference of each microphone.
[0030] The process of obtaining the fixed beam of the microphone array is specifically as follows:
[0031] The negative correlation mapping results of the sound intensity interference of each microphone are normalized, and the normalized value is used as the weight of the sound source signal sequence after the delayed data length of the corresponding microphone. The delayed sound source signal sequences of all weighted microphones are summed to obtain the fixed beam of the microphone array.
[0032] In a second aspect, an embodiment of the present application also provides a sound source direction finding and positioning system for complex environments, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of any one of the above methods when executing the computer program.
[0033] This application has at least the following beneficial effects:
[0034] 1. By calculating the microphone's sound quality deviation and sound intensity interference, we identify microphones most susceptible to noise interference and address them accordingly. This method effectively suppresses background noise interference, enabling the sound source localization system to maintain high stability in complex noise environments, significantly improving the system's noise immunity and providing a reliable guarantee for subsequent, more accurate sound source localization.
[0035] 2. On the basis of suppressing noise interference, by analyzing the differences between sound source signals, calculating the similarities between different parts, and extracting pure sound source signal components, it can effectively reduce positioning errors and significantly improve the accuracy of sound source positioning. Even in a multi-sound source environment, the position of the target sound source can be more accurately identified.
[0036] 3. By taking into account multiple factors, including the geometric layout of the microphone array, signal latency, signal strength, and noise interference, the sound source localization system can better adapt to various complex environments. Whether in scenes with high noise levels, strong reverberation, or in the presence of multiple sound sources, it can stably output accurate sound source location information, enhancing the system's adaptability and reliability and broadening its application in various scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A flowchart of a method for sound source direction finding and positioning in a complex environment provided by one embodiment of the present application;
[0038] Figure 2 A schematic diagram of calculating short-term audio similarity provided by one embodiment of the present application;
[0039] Figure 3 A flowchart for obtaining a fixed beam is provided for one embodiment of the present application. DETAILED DESCRIPTION
[0040] In the description of the embodiments of this application, words such as "exemplary," "or," and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "or," and "for example" is intended to present the relevant concepts in a concrete manner.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in the art of this application. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0042] It should also be noted that the terms "first" and "second" in this application and the accompanying drawings are used to distinguish similar objects, rather than to describe a specific order or sequence. The methods disclosed in the embodiments of this application or the methods shown in the flowcharts include one or more steps for implementing the methods. Without departing from the scope of protection of this application, the order of executing multiple steps can be interchanged with each other, and some steps can also be deleted.
[0043] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0044] The specific solutions of the sound source direction finding and positioning method and system provided by the present application in complex environments are described in detail below with reference to the accompanying drawings.
[0045] See also Figure 1 , which shows a flowchart of a method for sound source direction finding and positioning in a complex environment provided by an embodiment of the present application, the method comprising the following steps:
[0046] The first step is to combine the sound intensity of each microphone with a preset duration into a sound source signal sequence.
[0047] First, a dual-microphone array is installed in the space where sound source localization is required. The dual-microphone array includes 6 microphones, with one microphone as the origin and the positions of the remaining microphones are 、 、 、 、 , is the distance between array elements, which is 0.06m in this embodiment. The six microphones are marked as Then, a dual-microphone array is used to collect the intensity of sound in complex environments, collecting data at each moment and the 30ms before it. The frequency of each data acquisition is 40kHz, and data is collected every 15ms. The data collected by each microphone is arranged in chronological order to obtain the sound source signal sequence of each microphone.
[0048] The second step: Based on the distance between two adjacent microphones, combined with the speed of sound propagation and the frequency of collecting sound intensity, the maximum receiving interval between two adjacent microphones is obtained.
[0049] For data collected by different microphones, the time it takes to collect data from the same sound source varies due to the different positions of the microphones from the sound source, so it is necessary to locate the time delay between the microphones. For the same sound source, the spatial distance with the maximum time delay between two adjacent microphones is the distance between the two microphones. Therefore, the maximum receiving interval between two adjacent microphones is calculated by: obtaining the distance between the two adjacent microphones; calculating the ratio of the distance to the speed of sound propagation in air; and multiplying the ratio by the frequency of the microphone data collected is used as the maximum receiving interval between the two adjacent microphones.
[0050] In this embodiment, the speed of sound propagation in air is 340 m / s; the frequency of the microphone collecting data is 40 kHz.
[0051] It should be understood that as the distance between two microphones increases, the time interval between recording the same audio signal from the same sound source also increases accordingly. This is because sound takes time to propagate through the air; the greater the distance, the later the sound arrives. Therefore, the maximum reception interval between two microphones also increases with distance.
[0052] The third step: based on the numerical value of the maximum receiving interval, the sound source signal sequences of the two microphones are moved in sequence to obtain the short-time audio similarity between the two microphones after each movement.
[0053] For the sound source signal sequences between different microphones, when performing data enhancement, it is necessary to determine the time delay position between the sound source signals. For the signals at the same time of the sound source, there is similarity between the data collected by the microphones, and in the absence of noise interference, the sound source signals collected by the two microphones are the same. Therefore, the similarity between the adjacent microphones can be determined by the similarity between the sound source signal sequence data between the two microphones. Thus, the short-term audio similarity between the partial data of the sound source signal sequence between the microphones is calculated:
[0054]
[0055]
[0056] Where, represents the short-term audio similarity between the i-th microphone sound source signal sequence at the j-th position and the i+1-th microphone sound source signal sequence; The data sequence representing the sound source signal sequence of the i-th microphone from the j-th position to the n-th position; represents the data sequence of the i+1th microphone sound source signal sequence from the 1st position to the n-j+1th position; n represents the length of the microphone sound source signal sequence; Represents the similarity function, and in this embodiment, cosine similarity is used. It should be noted that, due to It is less than the acquisition time of 30ms, so the maximum receiving interval is less than the length of the sound source signal sequence. represents the short-term audio similarity between the i+1th microphone sound source signal sequence at the jth position and the i-th microphone sound source signal sequence; The data sequence representing the sound source signal sequence of the i+1th microphone from the jth position to the nth position; represents the data sequence of the sound source signal sequence of the i-th microphone from the 1st position to the n-j+1th position; j represents the position number; L represents the maximum receiving interval between the i-th microphone and the i+1-th microphone.
[0057] Among them, the calculation diagram of short-term audio similarity is as follows Figure 2 shown; among them, 、 Represent the i-th and i+1-th microphone sound source signal sequences respectively.
[0058] It should be understood that the calculation process of the short-time audio similarity is equivalent to moving the sound source signal sequences of the i-th and i+1-th microphones in sequence. First, the sound source signal sequences of the two microphones are aligned. After alignment, the sound source signal sequence of the i-th microphone is moved to the left with a step size of 1 for each movement. At this time, the sound source signal sequence of the i+1-th microphone does not move. After each movement of j steps, the similarity between the overlapping parts of the sound source signal sequences of the i-th and i+1-th microphone sequences is calculated as the short-time audio similarity between the sound source signal sequence of the i-th microphone at the j-th position and the sound source signal sequence of the i+1-th microphone; correspondingly, the sound source signal sequence of the i+1-th microphone is moved to the left by j steps with a step size of 1 for each movement. At this time, the sound source signal sequence of the i-th microphone does not move, and the short-time audio similarity between the sound source signal sequence of the i+1-th microphone at the j-th position and the sound source signal sequence of the i-th microphone is obtained.
[0059] For example, the sound source signal sequence of the i-th microphone is {1,2,3}, and the sound source signal sequence of the i+1-th microphone is {4,5,6}. When the sound source signal sequence of the i-th microphone moves to the left by 1 step, the overlapping parts of the sound source signal sequences of the two microphones are {2,3} and {4,5}; when the sound source signal sequence of the i+1-th microphone moves to the left by 1 step, the overlapping parts of the sound source signal sequences of the two microphones are {1,2} and {5,6}.
[0060] It should be understood that the higher the similarity between the partial sequences of the sound source signal sequences collected by two microphones, the more likely it is that the two microphones at that location received the sound emitted by the sound source at the same time. This high similarity indicates that the time delay and spatial position relationship between the sound source signals, when reaching the two microphones via different paths, are consistent with expectations, allowing accurate determination of the direction and distance of the sound source.
[0061] The fourth step: Based on the distribution of short-term audio similarity between each microphone and the microphone after it, the sound source synchronization sequence and the sound source isofrequency sequence of each microphone are extracted, and the sound quality deviation of each microphone is obtained according to the difference between the two sequences.
[0062] Since the processing method for the sound source signal sequence of each microphone is consistent, this application takes the sound source signal of the i-th microphone as an example to illustrate: calculate the short-time audio similarity between different parts of the i-th and i+1-th microphone sound source signal sequences. In this embodiment, a short-time audio similarity is obtained each time the short-time audio similarity is moved, and a total of 2×L short-time audio similarities can be obtained; when the short-time audio similarity value is the largest, the elements at the corresponding positions of the i-th microphone sound source signal sequence are composed of the sound source synchronization sequence of the i-th microphone, and the elements at the corresponding positions of the i+1-th microphone sound source signal sequence are composed of the sound source synchronization sequence of the i+1-th microphone.
[0063] For the sound source signals collected by the microphones, since the frequency of the same sound is the same, noise can cause abnormal frequencies to appear in the sound source signals collected by the two microphones. Therefore, the sound source synchronization sequences of the i-th microphone and the (i+1)-th microphone are used as the output of the Fourier transform algorithm, and the frequencies of the two microphone sound source synchronization sequences are output. It should be noted that the output frequencies are frequencies with non-zero amplitudes.
[0064] Extract the same frequency in the synchronous sequence of the sound sources of the two microphones, and use the same frequency and corresponding amplitude in the i-th microphone and the i+1-th microphone as the output of the inverse Fourier transform, and output the same frequency sequence of the sound source of the i-th microphone. Note: and Input, get The same frequency sequence of the sound source; and Input, get The sound source has the same frequency sequence, and so on. and Input, get The calculation of Fourier transform and inverse Fourier transform is a well-known technology, and the specific calculation steps are not repeated here.
[0065] Because the same sound has the same frequency, the difference between the source synchronization sequence and the source frequency sequence indicates noise interference unique to that microphone. Therefore, the sound quality deviation of each microphone is calculated: the distance between the source synchronization sequence and the source frequency sequence corresponding to each microphone is used as the sound quality deviation of each microphone. This embodiment uses the Manhattan distance for this calculation.
[0066] It should be understood that a high sound quality deviation indicates that the microphone is subject to severe noise interference and the signal it collects contains a high level of noise. A low sound quality deviation indicates that the signal collected by the microphone is relatively pure and less affected by noise. If a microphone has a high sound quality deviation, its signal can be subjected to separate noise reduction processing or assigned a lower weight in subsequent signal processing to improve the signal quality of the entire array and the accuracy of sound source direction finding.
[0067] The fifth step: select any microphone as the reference microphone, analyze the sound source synchronization sequence of all remaining microphones based on the reference microphone, and obtain the sound source similarity sequence and delay data length of each microphone.
[0068] Any microphone is selected as the reference microphone. In this embodiment, the microphone at the origin is selected; the sound source synchronization sequence of other microphones based on the reference microphone is obtained. The intersection of the element serial numbers of the sound source synchronization sequence elements of each microphone based on the reference microphone in the sound source signal sequence is extracted, and the elements with all serial numbers in the intersection in the sound source signal sequence of each microphone are combined into a sound source similarity sequence for each microphone. For example, the element serial numbers of the sound source synchronization sequence elements of each microphone based on the reference microphone in the sound source signal sequence are: 2 to 7, 3 to 9, 2 to 8, of which the common part is 3 to 7. Therefore, the data at positions 3 to 7 are extracted to form a sound source similarity sequence for each microphone; then the serial number of the first element of the sound source similarity sequence in the sound source signal sequence is used as the time delay data length of the corresponding microphone. For example, the part extracted by the i-th microphone is from the j-th position to the n-th position, and j is the time delay data length of the i-th microphone.
[0069] The sixth step: according to the arrangement values of the elements at the same position in the sound source similarity sequence of all microphones, the sound intensity distribution sequence of each microphone is obtained, the discrete degree of the elements in the sound intensity distribution sequence is analyzed, and the sound intensity interference degree of each microphone is obtained in combination with the sound quality deviation degree.
[0070] Since there is a certain distance between the arrangement of the microphone array, and the sound will attenuate during the propagation process, the intensity received by different array elements in the microphone array for the same sound is different. In the absence of noise interference, the order of the audio intensity received by the microphone at different times should be in the same position. The elements with the same sequence number in the sound source similarity sequence of each microphone are sorted in ascending order, and the corresponding elements in the sound source similarity sequence are replaced with the sorted values to obtain the sound intensity distribution sequence of each microphone. Thus, the sound intensity interference degree of each microphone is calculated: the sound quality deviation degree of each microphone is forward fused with the discrete degree of the corresponding sound intensity distribution sequence to obtain the sound intensity interference degree of each microphone. In this embodiment, the discrete degree between sequence elements is calculated by the mean absolute deviation; the forward fusion between multiple variables is obtained by multiplication.
[0071] It should be understood that when a microphone's sound intensity interference is high, it indicates that the difference between the sound source signal received by that microphone and other microphones is significant, indicating that it is experiencing large intensity fluctuations and that the signal it collects may contain more abnormal intensity components. On the other hand, when it is low, the signal intensity collected by the microphone is relatively stable and less subject to interference. If a microphone's sound intensity interference is high, its signal can be individually intensity-corrected or assigned a lower weight in subsequent signal processing to improve the signal quality of the entire array and the accuracy of sound source direction finding.
[0072] The seventh step: Based on the sound intensity interference of each microphone and the length of the time delay data, a fixed beam of the microphone array is obtained, the direction with the maximum output power is taken as the direction of the sound source, and the distance between the reference microphone and the sound source is determined in combination with the far-field model.
[0073] Through the sound intensity interference of the microphone, the DSB algorithm is used to obtain the fixed beam of the microphone array for the sound source signal collected by the microphone array: the negative correlation mapping result of the sound intensity interference of each microphone is normalized, and the normalized value is used as the weight of the sound source signal sequence after the time-delay data length of the corresponding microphone. The delayed sound source signal sequences of all weighted microphones are summed to obtain the fixed beam of the microphone array.
[0074] The fixed beam acquisition flow chart is as follows: Figure 3 shown.
[0075] It should be understood that a higher value for the acoustic interference intensity of a microphone indicates poorer signal quality. Therefore, a smaller weight should be assigned to each microphone during signal fusion. By reducing the weight of microphones with high noise interference, noise can be effectively suppressed, reducing its impact on the final output signal, thereby improving the signal quality of the entire system and, consequently, enhancing the accuracy of sound source direction finding and positioning.
[0076] Within the predefined search range, the output power of the beam is calculated direction by direction. The search range in this embodiment is 0~360 degrees. In this embodiment, an integer angle is taken as a direction, and the output power of the beamformer in each direction is calculated, and the direction with the largest output power is taken as the direction of the sound source. Among them, the calculation of the output power of the beamformer is a well-known technology, and the specific calculation steps are not repeated here. Finally, the geometric layout of the microphone array and the direction between the sound source and the origin are used as the input of the far-field model, and the output is the distance between the sound source and the origin. At this point, the position and direction of the sound source from the microphone array are obtained to achieve direction finding and positioning of the sound source. Among them, the calculation of the far-field model is a well-known technology, and the specific calculation steps are not repeated here.
[0077] Based on the same inventive concept as the above method, an embodiment of the present application also provides a sound source direction finding and positioning system for complex environments, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-mentioned sound source direction finding and positioning methods for complex environments.
[0078] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to the embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.
[0079] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the basic features of the present application. Therefore, from any point of view, the above embodiments of the present application should be regarded as exemplary and non-restrictive; modifications to the technical solutions described in the above embodiments, or equivalent replacement of some of the technical features therein, do not deviate from the essence of the corresponding technical solutions within the scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application.
Claims
1. A sound source direction finding and positioning method in a complex environment, characterized by: The method comprises the following steps: The sound intensity of each microphone with a preset duration is combined into a sound source signal sequence; Based on the distance between two adjacent microphones, combined with the speed of sound propagation and the frequency of collecting sound intensity, the maximum receiving interval between two adjacent microphones is obtained; Based on the numerical value of the maximum receiving interval, the sound source signal sequences of the two microphones are moved in sequence to obtain the short-term audio similarity between the two microphones after each movement. The specific formula is: Where, represents the short-term audio similarity between the i-th microphone sound source signal sequence at the j-th position and the i+1-th microphone sound source signal sequence; The data sequence representing the sound source signal sequence of the i-th microphone from the j-th position to the n-th position; represents the data sequence of the i+1th microphone sound source signal sequence from the 1st position to the n-j+1th position; n represents the length of the microphone sound source signal sequence; represents the similarity function; represents the short-term audio similarity between the i+1th microphone sound source signal sequence at the jth position and the i-th microphone sound source signal sequence; The data sequence representing the sound source signal sequence of the i+1th microphone from the jth position to the nth position; represents the data sequence of the i-th microphone sound source signal sequence from the 1st position to the n-j+1th position; j represents the position number, ; L represents the maximum receiving interval between the i-th microphone and the i+1-th microphone; Based on the distribution of short-term audio similarity between each microphone and the microphone after it, the sound source synchronization sequence and the sound source isofrequency sequence of each microphone are extracted, and the distance between the sound source synchronization sequence and the sound source isofrequency sequence corresponding to each microphone is measured as the sound quality deviation of each microphone; Select one microphone as the reference microphone, analyze the sound source synchronization sequence of all remaining microphones based on the reference microphone, and obtain the sound source similarity sequence and delay data length of each microphone; The sound intensity distribution sequence of each microphone is obtained based on the arrangement values of the elements at the same position in the sound source similarity sequence of all microphones. The sound quality deviation of each microphone is forward fused with the discrete degree of the corresponding sound intensity distribution sequence to obtain the sound intensity interference degree of each microphone. Based on the sound intensity interference of each microphone and the length of the time delay data, a fixed beam of the microphone array is obtained, and the direction with the maximum output power is taken as the sound source direction. Combined with the far-field model, the distance between the reference microphone and the sound source is determined.
2. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein: The maximum receiving interval between two adjacent microphones is obtained as follows: Obtain the distance between two adjacent microphones; calculate the ratio of the distance to the speed of sound propagation in the air; and multiply the ratio by the frequency at which the microphones collect data as the maximum receiving interval between the two adjacent microphones.
3. The sound source direction finding and positioning method in a complex environment according to claim 1, wherein: The extraction of the sound source synchronization sequence and the sound source frequency sequence of each microphone is specifically as follows: The elements of the sound source signal sequence of each microphone and the next microphone corresponding to the position of each microphone when the short-time audio similarity is the largest are respectively used to form a sound source synchronization sequence of each microphone and the next microphone; The same frequency with non-zero amplitude between the sound source synchronization sequence of each microphone and the next microphone is obtained, and the amplitude of the same frequency of the two microphones is processed by inverse Fourier transform to obtain the sound source frequency sequence of each microphone.
4. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein: The sound source similarity sequence and time delay data length of each microphone are obtained as follows: Extracting the intersection of the element serial numbers of the sound source synchronization sequence elements of each microphone based on the reference microphone in the sound source signal sequence, and combining all the elements with serial numbers in the sound source signal sequence of each microphone in the intersection into a sound source similarity sequence for each microphone; The sequence number of the first element of the sound source similarity sequence in the sound source signal sequence is used as the time delay data length of the corresponding microphone.
5. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein: The sound intensity distribution sequence of each microphone is obtained as follows: The elements with the same sequence number in the sound source similarity sequence of each microphone are sorted in ascending order, and the corresponding elements in the sound source similarity sequence are replaced with the sorted values to obtain the sound intensity distribution sequence of each microphone.
6. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein: The process of obtaining the fixed beam of the microphone array is specifically as follows: The negative correlation mapping results of the sound intensity interference of each microphone are normalized, and the normalized value is used as the weight of the sound source signal sequence after the delayed data length of the corresponding microphone. The delayed sound source signal sequences of all weighted microphones are summed to obtain the fixed beam of the microphone array.
7. A sound source direction finding and positioning system for complex environments, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Audio processing method, processing system, medium and program product
CN118841022A
Pickup control method and device based on sound position recognition
CN119152878A