Sound source direction finding positioning method and system oriented to complex environment

By calculating the sound quality deviation and sound intensity interference of the microphone, combining short-term audio similarity and sound source sequence differences, pure sound source signal components are extracted, and the fixed beam of the microphone array is calculated to determine the sound source direction, which solves the problem of low sound source positioning accuracy in complex environments and achieves high noise resistance and precision sound source positioning.

CN120233305AActive Publication Date: 2025-07-01SUZHOU AUDITORYWORKS CO LTD

Patent Information

Application Number
CN202510715166.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In complex environments, the interference of background noise and reverb significantly reduces the accuracy of sound source positioning. Traditional DSB algorithms have excessive weights due to the same weight processing, resulting in signal errors.

Method used

By calculating the sound quality deviation and sound intensity interference between the microphones, we identify the microphones with greater noise interference and perform targeted processing. Combining short-term audio similarity, sound source synchronization and sequence differences in the same frequency, pure sound source signal components are extracted, and the fixed beam of the microphone array is calculated to determine the sound source direction.

Benefits of technology

Effectively suppress background noise interference, improve the noise immunity and accuracy of the sound source positioning system, ensure the stable operation of the system in complex environments, and adapt to multi-sound sources and high-noise scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233305A_ABST
    Figure CN120233305A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, in particular to a sound source direction-finding positioning method and system for a complex environment, and the method comprises the steps: serializing the sound intensity of each microphone, and calculating a maximum receiving interval according to the distance between adjacent microphones, the sound transmission speed and the sampling frequency; the method comprises the following steps: determining a short-time audio similarity through moving a sound source signal sequence, extracting a sound source synchronization sequence and a sound source same-frequency sequence, and further evaluating a tone quality deviation degree; optionally selecting a reference microphone, and comparing the reference microphone with a sound source synchronization sequence of other microphones to obtain a time delay data length and a sound source similar sequence; based on the arrangement values of the same position elements of the sound source similar sequence, a sound intensity distribution sequence is obtained, and the sound intensity interference degree of each microphone is obtained; and determining a fixed beam direction and calculating the distance between the reference microphone and the sound source. The invention aims to improve the accuracy of sound source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speech processing, and specifically to a sound source direction finding and positioning method and system for complex environments. Background Art

[0002] Sound source localization technology usually relies on an array composed of multiple microphones to receive sound signals. By precisely analyzing information such as the phase difference and time difference between signals of different sensors, the system can calculate the specific position of the sound source. However, in complex environments, the interference of background noise and reverberation will significantly reduce the positioning accuracy, posing a huge challenge to sound source localization. The core indicators of sound source localization include spatial resolution and noise resistance. Spatial resolution determines the minimum angle at which the system can distinguish adjacent sound sources. High spatial resolution means that the system can more accurately locate the sound source and can accurately identify the target even in a multi-source environment. Noise resistance is the key for a sound source localization system to maintain stability and reliability in a complex noise environment. Improving the noise resistance of the system can not only effectively suppress the interference of background noise, but also enhance the accuracy of sound source position localization, ensuring the stable operation of the system in various complex scenarios.

[0003] For the processing technology of real-time sound source localization, since it is necessary to locate the position of the sound source in real time, for the processing of speech signals, simple algorithms need to be used. The traditional method is generally to perform real-time processing through Delay-and-Sum Beamforming (DSB). However, since the traditional DSB algorithm uses the same weight for the speech signals received by each microphone, the same weight will cause the weight attached to the noise data to be too large, resulting in the presence of incoherent noise residues in the processed signal, leading to errors in subsequent sound source direction finding and positioning. Summary of the Invention

[0004] In view of the above, it is necessary to provide a sound source direction finding and positioning method and system for complex environments to solve the above problems.

[0005] In the first aspect of this application, a sound source direction finding and positioning method for complex environments is provided. The method includes: Composing a sound source signal sequence from the sound intensities of each microphone for a preset duration; Based on the distance between adjacent two microphones, combining the sound propagation speed and the frequency of collecting sound intensities, obtaining the maximum reception interval amount between adjacent two microphones; Based on the numerical magnitude of the maximum reception interval amount, sequentially moving the sound source signal sequences of two microphones to obtain the short-time audio similarity between the two microphones after each movement; Based on the distribution of the short-time audio similarity between each microphone and the next microphone, extract the sound source synchronization sequence and the sound source same-frequency sequence of each microphone, and obtain the sound quality deviation degree of each microphone according to the difference between the two sequences. Arbitrarily select one microphone as the reference microphone, analyze the sound source synchronization sequences of all the remaining microphones based on the reference microphone, and obtain the sound source similarity sequence and the time delay data length of each microphone. According to the arrangement values of the elements at the same positions in the sound source similarity sequences of all microphones, obtain the sound intensity distribution sequence of each microphone, analyze the dispersion degree of the elements in the sound intensity distribution sequence, and combine the sound quality deviation degree to obtain the sound intensity interference degree of each microphone. Based on the sound intensity interference degree of each microphone and the time delay data length, obtain the fixed beam of the microphone array, take the direction with the maximum output power as the sound source direction, and combine the far-field model to determine the distance between the reference microphone and the sound source.

[0006] Among them, the obtaining of the maximum reception interval amount between two adjacent microphones is specifically as follows: Obtain the distance between two adjacent microphones; calculate the ratio of the distance to the propagation speed of sound in the air; take the product of the ratio and the frequency of the data collected by the microphone as the maximum reception interval amount between two adjacent microphones.

[0007] Among them, the obtaining of the short-time audio similarity between two microphones after each movement is specifically as follows: In the formula, represents the short-time audio similarity between the sound source signal sequence of the i-th microphone at the j-th position and the sound source signal sequence of the (i + 1)-th microphone; represents the data sequence of the sound source signal sequence of the i-th microphone from the j-th position to the n-th position; represents the data sequence of the sound source signal sequence of the (i + 1)-th microphone from the 1st position to the (n - j + 1)-th position; n represents the length of the sound source signal sequence of the microphone; represents the similarity function; represents the short-time audio similarity between the sound source signal sequence of the (i + 1)-th microphone at the j-th position and the sound source signal sequence of the i-th microphone; represents the data sequence of the sound source signal sequence of the (i + 1)-th microphone from the j-th position to the n-th position; represents the data sequence of the sound source signal sequence of the i-th microphone from the 1st position to the (n - j + 1)-th position; j represents the position serial number. ; L represents the maximum reception interval amount between the i-th microphone and the (i + 1)-th microphone.

[0008] Among them, the extraction of the sound source synchronization sequence and the sound source same-frequency sequence of each microphone is specifically as follows: When the short-time audio similarity is the largest, the elements of the sound source signal sequences at the corresponding positions of each microphone and the microphone immediately following it are respectively combined to form the sound source synchronization sequence of each microphone and the microphone immediately following it; Obtain the same frequencies with non-zero amplitudes between the sound source synchronization sequences of each microphone and the microphone immediately following it, and use the inverse Fourier transform to process the amplitudes of the same frequencies of the two microphones to obtain the sound source same-frequency sequence of each microphone.

[0009] Among them, the sound quality deviation degree of each microphone is specifically the distance metric between the sound source synchronization sequence corresponding to each microphone and the sound source same-frequency sequence.

[0010] Among them, the obtaining of the sound source similarity sequence and the delay data length of each microphone is specifically as follows: Extract the intersection of the element numbers of the sound source synchronization sequence elements of each microphone based on the reference microphone in the sound source signal sequence, and form the sound source similarity sequence of each microphone with the elements in the sound source signal sequence of all the numbers in the intersection; Take the number of the first element of the sound source similarity sequence in the sound source signal sequence as the delay data length of the corresponding microphone.

[0011] Among them, the obtaining of the sound intensity distribution sequence of each microphone is specifically as follows: Sort the elements with the same number in the sound source similarity sequence of each microphone in ascending order, and replace the corresponding elements in the sound source similarity sequence with the sorting values to obtain the sound intensity distribution sequence of each microphone.

[0012] Among them, the specific process of obtaining the sound intensity interference degree of each microphone is as follows: Positively fuse the sound quality deviation degree of each microphone with the discreteness of the corresponding sound intensity distribution sequence to obtain the sound intensity interference degree of each microphone.

[0013] Among them, the process of obtaining the fixed beam of the microphone array is specifically as follows: Normalize the negative correlation mapping result of the sound intensity interference degree of each microphone, take the normalized value as the weight of the sound source signal sequence after the delay data length of the corresponding microphone, and sum the sound source signal sequences after the delay of all microphones after weighting to obtain the fixed beam of the microphone array.

[0014] In a second aspect, an embodiment of the present application further provides a sound source direction finding and positioning system for complex environments, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0015] The present application has at least the following beneficial effects: 1. By calculating the sound quality deviation degree and sound intensity interference degree of microphones, the microphones greatly interfered by noise are identified and targeted processing is performed on them. This method effectively suppresses the interference of background noise, enables the sound source positioning system to still maintain high stability in a complex noise environment, significantly improves the noise resistance of the system, and provides a reliable guarantee for more accurate sound source positioning subsequently.

[0016] 2. On the basis of suppressing noise interference, by analyzing the differences between sound source signals, calculating the similarity between different parts, and extracting pure sound source signal components, the positioning error can be effectively reduced, and the accuracy of sound source positioning can be significantly improved. Even in a multi-source environment, the position of the target sound source can be more accurately identified.

[0017] 3. Considering various factors such as the geometric layout of the microphone array, signal time delay, intensity, and noise interference, the sound source positioning system can better adapt to various complex environments. Whether in a scene with a high noise level and strong reverberation or in the case of multi-source interference, it can stably output accurate sound source position information, enhancing the adaptability and reliability of the system and broadening its application scope in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of the steps of a sound source direction finding and positioning method for complex environments provided by an embodiment of the present application; Figure 2 It is a schematic diagram for calculating the short-time audio similarity provided by an embodiment of the present application; Figure 3 It is a flowchart for obtaining a fixed beam provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example", etc. are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary", "or", "for example" is intended to present relevant concepts in a specific manner.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0021] In addition, it should be noted that the terms "first" and "second" in this application and the accompanying drawings are used to distinguish similar objects and are not used to describe a specific order or sequence. For the methods disclosed in the embodiments of this application or shown in the flowcharts, including one or more steps for implementing the methods, without departing from the scope of protection of this application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs.

[0023] The following specifically describes the specific solutions of the sound source direction finding and positioning method and system provided by this application in the complex environment with reference to the accompanying drawings.

[0024] Please refer to Figure 1 , which shows the flowchart of the steps of the sound source direction finding and positioning method in the complex environment provided by an embodiment of this application. The method includes the following steps: The first step: Compose the sound intensity of each microphone for a preset duration into a sound source signal sequence.

[0025] First, install a dual microphone array in the space where sound source positioning is required. The dual microphone array includes 6 microphones. Taking one microphone as the origin, the positions where the remaining microphones are installed are , , , , , is the distance between array elements, which takes a value of 0.06 m in this embodiment. The 6 microphones are respectively denoted as in the clockwise direction. Then, collect the sound intensity in the complex environment through the dual microphone array, collect the data of each moment and the previous 30 ms. The frequency of each data collection is 40 kHz, and the data is collected once every 15 ms. Arrange the data collected by each microphone in the order of time, and obtain the sound source signal sequence of each microphone.

[0026] The second step: Based on the distance between adjacent microphones, combined with the sound propagation speed and the frequency of collecting sound intensity, obtain the maximum reception interval amount between adjacent microphones.

[0027] For the data collected by different microphones, since the positions of the microphones relative to the sound source are different, the times at which the same sound source data is collected are also different. Therefore, it is necessary to locate the time delay between the microphones. For the same sound source, the spatial distance corresponding to the maximum time delay between two adjacent microphones is the distance between the two microphones. Thus, the maximum reception interval between two adjacent microphones is calculated as follows: obtain the distance between two adjacent microphones; calculate the ratio of the distance to the propagation speed of sound in air; multiply the ratio by the frequency of the data collected by the microphones, and take the product as the maximum reception interval between two adjacent microphones.

[0028] In this embodiment, the propagation speed of sound in air is taken as 340 m / s; the frequency of the data collected by the microphones is 40 kHz.

[0029] It should be understood that when the distance between two microphones increases, the time interval for collecting the same audio signal for the same sound source will also increase accordingly. This is because it takes time for sound to propagate in air, and the farther the distance, the later the sound arrives. Therefore, the maximum reception interval between two microphones will also increase with the increase in distance.

[0030] The third step: Based on the numerical value of the maximum reception interval, sequentially shift the sound source signal sequences of the two microphones to obtain the short-time audio similarity between the two microphones after each shift.

[0031] For the sound source signal sequences between different microphones, when enhancing the data, it is necessary to determine the time delay position between the sound source signals. For the signals at the same moment of the sound source, there is similarity between the data collected by the microphones, and in the absence of noise interference, the sound source signals collected by the two microphones are the same. Thus, the similar position between adjacent microphones can be determined by the similarity degree between the sound source signal sequences of the two microphones. Thus, the short-time audio similarity between partial data of the sound source signal sequences between the microphones is calculated as follows: In the formula, represents the short-time audio similarity between the sound source signal sequence of the i-th microphone at the j-th position and the sound source signal sequence of the (i + 1)-th microphone; represents the data sequence of the sound source signal sequence of the i-th microphone from the j-th position to the n-th position; represents the data sequence of the sound source signal sequence of the (i + 1)-th microphone from the 1st position to the (n - j + 1)-th position; n represents the length of the sound source signal sequence of the microphone; represents the similarity function, which adopts the cosine similarity in this embodiment. It should be noted that since It is less than the acquisition time of 30 ms, so the maximum reception interval is less than the length of the sound source signal sequence. represents the short-time audio similarity between the sound source signal sequence of the (i + 1)-th microphone at the j-th position and the sound source signal sequence of the i-th microphone; represents the data sequence of the sound source signal sequence of the (i + 1)-th microphone from the j-th position to the n-th position; represents the data sequence of the sound source signal sequence of the i-th microphone from the 1st position to the (n - j + 1)-th position; j represents the position serial number; L represents the maximum reception interval between the i-th microphone and the (i + 1)-th microphone.

[0032] Among them, the schematic diagram for calculating the short-time audio similarity is as Figure 2 shown; among them, and respectively represent the sound source signal sequences of the i-th and (i + 1)-th microphones.

[0033] It should be understood that the calculation process of the short-time audio similarity is equivalent to moving the sound source signal sequences of the i-th and (i + 1)-th microphones in turn. First, align the sound source signal sequences of the two microphones. After alignment, move the sound source signal sequence of the i-th microphone to the left, with a step size of 1 each time. At this time, the sound source signal sequence of the (i + 1)-th microphone does not move. After moving j step sizes each time, calculate the similarity between the overlapping parts of the sound source signal sequences of the i-th and (i + 1)-th microphone sequences as the short-time audio similarity between the sound source signal sequence of the i-th microphone at the j-th position and the sound source signal sequence of the (i + 1)-th microphone; correspondingly, move the sound source signal sequence of the (i + 1)-th microphone to the left by j step sizes, with a step size of 1 each time. At this time, the sound source signal sequence of the i-th microphone does not move, and the short-time audio similarity between the sound source signal sequence of the (i + 1)-th microphone at the j-th position and the sound source signal sequence of the i-th microphone is obtained.

[0034] For example, the sound source signal sequence of the i-th microphone is {1, 2, 3}, and the sound source signal sequence of the (i + 1)-th microphone is {4, 5, 6}. When the sound source signal sequence of the i-th microphone is moved to the left by 1 step size, the overlapping parts of the sound source signal sequences of the two microphones are {2, 3} and {4, 5}; when the sound source signal sequence of the (i + 1)-th microphone is moved to the left by 1 step size, the overlapping parts of the sound source signal sequences of the two microphones are {1, 2} and {5, 6}.

[0035] It should be understood that the higher the similarity between the partial sequences of the sound source signal sequences collected by two microphones, the more it indicates that the two microphones corresponding to this position receive the sound emitted by the sound source at the same moment. This high similarity indicates that when the sound source signal arrives at the two microphones through different paths, its time delay and spatial position relationship meet the expectations, so that the direction and distance of the sound source can be accurately judged.

[0036] The fourth step: Based on the distribution of the short-time audio similarity between each microphone and the next microphone, extract the sound source synchronization sequence and the sound source same-frequency sequence of each microphone, and obtain the sound quality deviation degree of each microphone according to the difference between the two sequences.

[0037] Since the processing method for the sound source signal sequence of each microphone is the same, this application takes the sound source signal of the i-th microphone as an example for illustration: Calculate the short-time audio similarity between the different parts of the sound source signal sequences of the i-th and the (i + 1)-th microphones. In this embodiment, one short-time audio similarity will be obtained each time it moves, and a total of 2×L short-time audio similarities can be obtained; when the value of the short-time audio similarity is the largest, the elements at the corresponding positions of the sound source signal sequence of the i-th microphone form the sound source synchronization sequence of the i-th microphone, and the elements at the corresponding positions of the sound source signal sequence of the (i + 1)-th microphone form the sound source synchronization sequence of the (i + 1)-th microphone.

[0038] For the sound source signal collected by the microphone, since the frequencies generated by the same sound are the same, while noise will cause abnormal frequencies in the frequencies of the sound source signals collected by the two microphones. Thus, take the sound source synchronization sequences of the i-th and the (i + 1)-th microphones as the outputs of the Fourier transform algorithm respectively, and output the frequencies of the sound source synchronization sequences of the two microphones. It should be noted that the output frequencies are the frequencies with non-zero amplitudes.

[0039] Extract the same frequencies in the sound source synchronization sequences of the two microphones, and take the same frequencies and the corresponding amplitudes in the i-th microphone and the (i + 1)-th microphone as the output of the inverse Fourier transform, and output the sound source same-frequency sequence of the i-th microphone. Note: and input, to obtain the sound source same-frequency sequence; and input, to obtain the sound source same-frequency sequence, and so on, and input, to obtain the sound source same-frequency sequence. Among them, the calculations of the Fourier transform and the inverse Fourier transform are well-known technologies, and the specific calculation steps will not be elaborated here.

[0040] Due to the same frequency of the same sound, the difference between the sound source synchronization sequence and the sound source same-frequency sequence indicates the unique noise interference of this microphone. Thus, calculate the sound quality deviation degree of the microphone: Take the distance metric between the sound source synchronization sequence corresponding to each microphone and the sound source same-frequency sequence as the sound quality deviation degree of each microphone. In this embodiment, the Manhattan distance is used for calculation.

[0041] It should be understood that when the sound quality deviation degree of the microphone is relatively high, it indicates that the noise interference received by this microphone is relatively serious, and there are more noise components in the collected signal; while when it is relatively low, it indicates that the signal collected by the microphone is relatively pure and less affected by noise. If the sound quality deviation degree of a certain microphone is relatively high, the signal of this microphone can be subjected to separate noise reduction processing, or a lower weight can be given to the signal of this microphone in subsequent signal processing to improve the signal quality of the entire array and the accuracy of sound source direction finding and positioning.

[0042] The fifth step: Arbitrarily select a microphone as the reference microphone, and analyze the sound source synchronization sequences of all the remaining microphones based on the reference microphone to obtain the sound source similarity sequence and the time delay data length of each microphone.

[0043] Arbitrarily select a microphone as the reference microphone. In this embodiment, the microphone at the origin position is selected; obtain the sound source synchronization sequences of other microphones based on the reference microphone. Extract the intersection of the element sequence numbers of the elements of the sound source synchronization sequence of each microphone based on the reference microphone in the sound source signal sequence, and form the sound source similarity sequence of each microphone with the elements of all the sequence numbers in the intersection in the sound source signal sequence of each microphone. For example, the element sequence numbers of the elements of the sound source synchronization sequence of each microphone based on the reference microphone in the sound source signal sequence are respectively: from 2 to 7, from 3 to 9, from 2 to 8, and the common part among them is from 3 to 7. Therefore, extract the data at positions from 3 to 7 to form the sound source similarity sequence of each microphone; then use the sequence number of the first element of the sound source similarity sequence in the sound source signal sequence as the time delay data length of the corresponding microphone. For example, if the part extracted by the i-th microphone is from the j-th position to the n-th position, then j is the time delay data length of the i-th microphone.

[0044] The sixth step: According to the arrangement values of the elements at the same positions in the sound source similarity sequences of all the microphones, obtain the sound intensity distribution sequence of each microphone, analyze the degree of dispersion of the elements in the sound intensity distribution sequence, and combine the sound quality deviation degree to obtain the sound intensity interference degree of each microphone.

[0045] Due to the certain distance between the arrangements of the microphone arrays, and the attenuation of sound during propagation, the intensities received by different elements in the microphone arrays for the same sound are different. In the case of no noise interference, the sorting of the audio intensities received by the microphones at different times should be at the same position. Sort the elements with the same serial number in the sound source similarity sequence of each microphone in ascending order, replace the corresponding elements in the sound source similarity sequence with the sorting values, and obtain the sound intensity distribution sequence of each microphone. Thus, calculate the sound intensity interference degree of each microphone: positively fuse the sound quality deviation degree of each microphone with the dispersion degree of the corresponding sound intensity distribution sequence to obtain the sound intensity interference degree of each microphone. In this embodiment, the dispersion degree between sequence elements is calculated by the mean absolute deviation; the positive fusion between multiple variables is obtained by multiplication.

[0046] It should be understood that when the sound intensity interference degree of a microphone is relatively high, it indicates that there is a relatively obvious difference between this microphone and other microphones when receiving the sound source signal, indicating that the intensity fluctuation it receives is relatively large, and the signal it collects may contain more abnormal intensity components; while when it is relatively low, it indicates that the signal intensity collected by the microphone is relatively stable and less interfered. If the sound intensity interference degree of a certain microphone is relatively high, the signal of this microphone can be subjected to separate intensity correction processing, or a lower weight can be assigned to the signal of this microphone in subsequent signal processing to improve the signal quality of the entire array and the accuracy of sound source direction finding and positioning.

[0047] The seventh step: Based on the sound intensity interference degree of each microphone and the length of the time delay data, obtain the fixed beam of the microphone array, take the direction with the maximum output power as the sound source direction, and combine the far-field model to determine the distance between the reference microphone and the sound source.

[0048] Using the sound intensity interference degree of the microphone, obtain the fixed beam of the microphone array from the sound source signal collected by the microphone array by the DSB algorithm: normalize the negative correlation mapping result of the sound intensity interference degree of each microphone, and use the normalized value as the weight of the sound source signal sequence after the time delay data length of the corresponding microphone, sum the time-delayed sound source signal sequences of all microphones after weighting to obtain the fixed beam of the microphone array.

[0049] Among them, the flow chart for obtaining the fixed beam is as Figure 3 shown.

[0050] It should be understood that the larger the value of the sound intensity interference degree of the microphone, the worse the signal quality of this microphone. Therefore, a smaller weight should be assigned during signal fusion. By reducing the weight of the microphone with large noise interference, noise can be effectively suppressed, the influence of noise on the final output signal can be reduced, and thus the signal quality of the entire system can be improved. Thereby improving the accuracy of sound source direction finding and positioning.

[0051] Within a predefined search range, the output power of the beam is calculated direction by direction. The search range is 0 to 360 degrees in this embodiment. In this embodiment, one integer angle is used as one direction, and the output power of the beamformer in each direction is calculated. The direction with the maximum output power is taken as the direction of the sound source. Among them, the calculation of the output power of the beamformer is a well-known technology, and the specific calculation steps are not elaborated here. Finally, the geometric layout of the microphone array and the direction between the sound source and the origin are used as the input of the far-field model, and the output is the distance between the sound source and the origin. Thus, the position and direction of the sound source relative to the microphone array are obtained, realizing the direction finding and positioning of the sound source. Among them, the calculation of the far-field model is a well-known technology, and the specific calculation steps are not elaborated here.

[0052] Based on the same inventive concept as the above method, an embodiment of the present application also provides a sound source direction finding and positioning system for complex environments, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for sound source direction finding and positioning in complex environments.

[0053] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the systems, methods, and computer program products according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order from that disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0054] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and the present application can be implemented in other specific forms without departing from the basic characteristics of the present application. Therefore, from any point of view, the above-described embodiments of the present application should be regarded as exemplary and non-limiting; modifications to the technical solutions described in the foregoing embodiments, or equivalent replacements of some of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A sound source direction finding and positioning method for complex environments, characterized in that, The method includes the following steps: Compose a sound source signal sequence from the sound intensity of each microphone over a preset duration; Based on the distance between two adjacent microphones, in combination with the sound propagation speed and the frequency of sound intensity acquisition, obtain the maximum reception interval amount between two adjacent microphones; Based on the numerical value of the maximum reception interval amount, sequentially shift the sound source signal sequences of two microphones to obtain the short-time audio similarity between the two microphones after each shift; Based on the distribution of the short-time audio similarity between each microphone and the next microphone, extract the sound source synchronization sequence and the sound source same-frequency sequence of each microphone, and obtain the sound quality deviation degree of each microphone according to the difference between the two sequences; Optionally select one microphone as the reference microphone, analyze the sound source synchronization sequences of all the remaining microphones based on the reference microphone, and obtain the sound source similarity sequence and the time delay data length of each microphone; According to the arrangement values of the elements at the same positions in the sound source similarity sequences of all microphones, obtain the sound intensity distribution sequence of each microphone, analyze the dispersion degree of the elements in the sound intensity distribution sequence, and in combination with the sound quality deviation degree, obtain the sound intensity interference degree of each microphone; Based on the sound intensity interference degree of each microphone and the time delay data length, obtain the fixed beam of the microphone array, take the direction with the maximum output power as the sound source direction, and in combination with the far-field model, determine the distance between the reference microphone and the sound source.

2. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein, The obtaining of the maximum reception interval amount between two adjacent microphones is specifically as follows: Obtain the distance between two adjacent microphones; calculate the ratio of the distance to the sound propagation speed in air; take the product of the ratio and the frequency of the data collected by the microphone as the maximum reception interval amount between two adjacent microphones.

3. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein The obtaining of the short-time audio similarity between the two microphones after each shift is specifically as follows: In the formula, represents the short-time audio similarity between the i-th microphone sound source signal sequence at the j-th position and the (i + 1)-th microphone sound source signal sequence; represents the data sequence of the i-th microphone sound source signal sequence from the j-th position to the n-th position; represents the data sequence of the (i + 1)-th microphone sound source signal sequence from the 1st position to the (n - j + 1)-th position; n represents the length of the microphone sound source signal sequence; represents the similarity function; represents the short-time audio similarity between the (i + 1)-th microphone sound source signal sequence at the j-th position and the i-th microphone sound source signal sequence; represents the data sequence of the (i + 1)-th microphone sound source signal sequence from the j-th position to the n-th position; represents the data sequence of the i-th microphone sound source signal sequence from the 1st position to the (n - j + 1)-th position; j represents the position serial number, ; L represents the maximum reception interval amount between the i-th microphone and the (i + 1)-th microphone.

4. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein, The extraction of the sound source synchronization sequence and the sound source same-frequency sequence of each microphone is specifically as follows: When the short-time audio similarity is the maximum, respectively compose the elements of the sound source signal sequences at the corresponding positions of each microphone and the next microphone to form the sound source synchronization sequence of each microphone and the next microphone; Obtain the same frequencies with non-zero amplitudes between the sound source synchronization sequences of each microphone and the next microphone, and perform processing on the amplitudes of the same frequencies of the two microphones using the inverse Fourier transform to obtain the sound source same-frequency sequence of each microphone.

5. The method for sound source direction finding and positioning in a complex environment according to claim 1, characterized in that, The sound quality deviation degree of each microphone is specifically the distance metric between the sound source synchronization sequence and the sound source same-frequency sequence corresponding to each microphone.

6. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein, The obtaining of the sound source similarity sequence and the time delay data length of each microphone is specifically as follows: Extract the intersection of the element numbers of the sound source synchronization sequence elements of each microphone based on the reference microphone in the sound source signal sequence, and compose the elements in the intersection in the sound source signal sequences of each microphone to form the sound source similarity sequence of each microphone; Take the number of the first element of the sound source similarity sequence in the sound source signal sequence as the time delay data length of the corresponding microphone.

7. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein, The obtaining of the sound intensity distribution sequence of each microphone is specifically as follows: Sort the elements with the same serial number in the sound source similarity sequence of each microphone in ascending order, replace the corresponding elements in the sound source similarity sequence with the sorting values, and obtain the sound intensity distribution sequence of each microphone.

8. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein The specific process of obtaining the sound intensity interference degree of each microphone is as follows: Positively fuse the sound quality deviation degree of each microphone with the discreteness of the corresponding sound intensity distribution sequence to obtain the sound intensity interference degree of each microphone.

9. The method for sound source direction finding and positioning in a complex environment according to claim 1, wherein The specific process of obtaining the fixed beam of the microphone array is as follows: Normalize the negative correlation mapping result of the sound intensity interference degree of each microphone, use the normalized value as the weight of the sound source signal sequence after the time delay data length of the corresponding microphone, and sum the sound source signal sequences after the time delay of all microphones after weighting to obtain the fixed beam of the microphone array.

10. A sound source direction finding and positioning system for complex environments, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Sound source positioning following system and method based on microphone array

    CN106782596A

  • Method for calibrating microphone arrays, sound localization method and related equipment

    CN110068797A

  • Audio processing method, processing system, medium and program product

    CN118841022A

  • Pickup control method and device based on sound position recognition

    CN119152878A

  • Audio data enhancement method and system suitable for complex environment

    CN119360876A

Cited By

  • Unmanned aerial vehicle detection method and system based on voiceprint recognition and microphone array

    CN121596351A

  • Far-field accidental sound source positioning method, system, medium and equipment

    CN121633985A