Acoustic sensor array-based sound source precise positioning method
By analyzing the characteristic waveforms of audio segments using an acoustic sensor array, the problem of inaccurate audio segment localization in existing technologies is solved, enabling fast and efficient audio segment localization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HUAIXIN IOT TECH CO LTD
- Filing Date
- 2024-12-17
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for audio file processing suffer from low accuracy and time-consuming audio segment localization, making it impossible to quickly and accurately determine the location of audio segments within an array of sound sources.
An acoustic sensor array is used to create audio images, acquire feature waveforms, segment them, analyze the feature information of audio segments, and compare them with the array's sound sources to determine the location of the audio segments.
It enables rapid and accurate location of audio segments within the array of sound sources, reducing the comparison range and improving judgment speed and accuracy.
Smart Images

Figure CN119724223B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to a method for precise sound source localization based on an acoustic sensor array. Background Technology
[0002] When processing audio files, sometimes the files are too large. In some scenarios, to determine whether a certain sound segment belongs to an array of sound sources (the entire audio file), the existing technology generally uses the following methods:
[0003] By manually listening to individual audio segments and then listening to the array audio to determine if the audio segments belong to the array audio file and their specific location within the file, a process that wastes a significant amount of time.
[0004] The method disclosed in the audio segment detection method and related equipment, such as application number CN201911399043.0, has low accuracy in practice.
[0005] Based on the above, this application proposes a method for precise sound source localization based on an acoustic sensor array. Summary of the Invention
[0006] To address this, the present invention provides a method for precise localization of sound sources based on an acoustic sensor array, in order to solve the problem of how to quickly and accurately determine the position of an audio segment within the array of sound sources.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] According to a first aspect of the present invention, a method for precise localization of sound sources based on an acoustic sensor array includes the following steps:
[0009] Step 1: Input an audio clip, which is then captured by an acoustic sensor.
[0010] Step two: Create an audio image, producing a waveform image of the input audio.
[0011] Step 3: Obtain the characteristic waveform and segment the waveform image to extract the changing sound waveform.
[0012] Step four, waveform analysis: Analyze the extracted waveform image and the input audio segment to obtain the feature information contained in the audio.
[0013] Step 5: Feature information analysis. The features obtained in Step 4 are analyzed, and the acquired feature information is compared with the array sound sources.
[0014] Step six: If there is no corresponding audio segment in the array sound source, the input audio segment is directly output. If the array sound source contains the input audio segment, the position of the audio segment in the array sound source is derived when outputting the input audio segment.
[0015] Step 7: Output the results.
[0016] Preferably, when creating the audio image, the audio image is created based on the audio decibel level, and the specific creation method is as follows:
[0017] Establish a two-dimensional coordinate system with time on the horizontal axis and decibels on the vertical axis. Based on the change of decibels in an audio segment over time, obtain a waveform image and mark the multiple audio time segments with the highest decibels in the audio segment, which are denoted as audio segments.
[0018] Preferably, in step three, the characteristic waveform is segmented from the starting point to the ending point of the characteristic waveform according to the already marked audio segments.
[0019] Preferably, waveform analysis is divided into content feature analysis and time feature analysis.
[0020] Content feature information analysis involves analyzing the audio content based on the segmented waveform image, determining the type of audio content, classifying the audio, and then analyzing and interpreting the classified audio content to determine the content feature information in the segmented waveform.
[0021] The temporal feature information analysis involves analyzing the input audio segment to determine whether it contains temporal feature information. If it does, the position of the input audio file in the array sound source is determined based on the temporal feature information. If it does not contain temporal feature information, the temporal feature information is no longer analyzed.
[0022] Preferably, when the content information feature is voice dialogue, then...
[0023] The audio dialogue content is extracted. If place names and / or personal names and / or events appear, the context containing the place names and / or personal names and / or events is determined to ascertain whether the time point of the input audio segment is before, at the present, or after the occurrence of the place names and / or personal names and / or events, thus obtaining the first content conclusion.
[0024] The array of sound sources is processed and segmented according to whether they appear before, at the present moment, or after the location name and / or person name and / or event.
[0025] Based on conclusion one, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
[0026] Preferably, when the content information feature is no voice dialogue, then there is
[0027] Content features are extracted from the segmented waveform to determine whether the extracted content is animal sound, object sound, or natural sound, leading to content conclusion two.
[0028] The array of sound sources is processed and segmented according to the time of occurrence of animal sounds, object sounds, or natural sounds.
[0029] According to conclusion two, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
[0030] Preferably, when the content information features include both speech dialogue information and animal sounds and / or object sounds and / or natural sounds, then...
[0031] Content features are extracted from the segmented waveform. The extracted content is determined to be animal sounds and / or object sounds and / or natural sounds and / or place names and / or person names and / or events, leading to content conclusion three.
[0032] The array of sound sources is processed and segmented according to animal sounds, object sounds, natural sounds, and / or place names and / or person names and / or the time of occurrence of events.
[0033] According to conclusion three, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
[0034] Preferably, when analyzing time feature information, the input audio segment is analyzed to determine whether it contains direct time information and / or indirect time information.
[0035] When direct temporal feature information is included, the array sound source is segmented according to the direct temporal information and compared with the input audio segment to determine the exact location of the input audio segment within the array sound source.
[0036] When indirect time feature information is included, the indirect time feature information is first analyzed to transform it into direct time feature information, and then the input audio segment is compared with the array sound source according to the direct time feature information.
[0037] The present invention has the following advantages:
[0038] When implementing this solution, by analyzing the features in the audio segments, anchoring the feature values, and then using the feature values to analyze the array sound sources, the feature values can be located quickly and accurately, thereby reducing the comparison range, speeding up the comparison judgment, and increasing the accuracy of the results. Attached Figure Description
[0039] Figure 1 The flowchart illustrates a method for precise sound source localization based on an acoustic sensor array, provided in some embodiments of the present invention.
[0040] Figure 2 The diagram shows the feature information analysis of the sound source precise localization method based on acoustic sensor array provided in some embodiments of the present invention. Detailed Implementation
[0041] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] like Figures 1 to 2 As shown, the method for precise sound source localization based on an acoustic sensor array in the first aspect embodiment of the present invention includes the following steps:
[0043] Step one: Input an audio clip. The audio clip is captured by an acoustic sensor. Of course, in this implementation, the audio clip can be user-inputted or captured by an acoustic sensor.
[0044] Step two: Create an audio image, producing a waveform image of the input audio.
[0045] Step 3: Obtain the characteristic waveform and segment the waveform image to extract the changing sound waveform.
[0046] Step four, waveform analysis: Analyze the captured waveform image and the input audio segment to obtain the feature information contained in the audio. The feature information can be the content of the dialogue, unusual sounds in the audio, or other sounds.
[0047] Step 5: Feature information analysis. The features obtained in Step 4 are analyzed, and the acquired feature information is compared with the array sound sources.
[0048] Step six: Result determination. If there is no corresponding audio segment in the array sound source, the input audio segment is directly output. If the array sound source contains the input audio segment, the position of the audio segment in the array sound source is derived when outputting the input audio segment.
[0049] Step 7: Output the results.
[0050] When implementing this solution, by analyzing the features in the audio segments, anchoring the feature values, and then using the feature values to analyze the array sound sources, the feature values can be located quickly and accurately, thereby reducing the comparison range, speeding up the comparison judgment, and increasing the accuracy of the results.
[0051] When creating audio images, the audio image is generated based on the audio decibel level. The specific creation method is as follows.
[0052] A two-dimensional coordinate system is established, with time on the horizontal axis and decibels on the vertical axis. Based on the change in decibels of sound within an audio segment over time, a waveform image is obtained. The multiple audio time intervals with the highest decibel levels in that segment are marked and denoted as audio segments. By creating audio spectrograms, the input audio can be compared and analyzed with the arrayed sound more intuitively and conveniently.
[0053] The highest decibel level among the audio segments can be a predefined number of segments. For example, if the goal is to collect n audio segments, where n is an integer, then when collecting the audio segments, the n highest frequencies are marked and collected. The duration of each collected audio segment is the total duration of the audio clip divided by m, where m is greater than n. m is any integer greater than n and is defined by the user.
[0054] In order to obtain the characteristic audio segments, the following technical solution is used: in step three, the characteristic waveform is segmented from the start point to the end point of the characteristic waveform according to the already marked audio segments.
[0055] When analyzing feature information, the following features can be analyzed: Waveform analysis is divided into content feature information analysis and time feature analysis.
[0056] Content feature information analysis involves analyzing the audio content based on the segmented waveform image, determining the type of audio content, classifying the audio, and then analyzing and interpreting the classified audio content to determine the content feature information in the segmented waveform.
[0057] The temporal feature information analysis involves analyzing the input audio segment to determine whether it contains temporal feature information. If it does, the position of the input audio file in the array sound source is determined based on the temporal feature information. If it does not contain temporal feature information, the temporal feature information is no longer analyzed.
[0058] Example 1
[0059] When the content information features are voice dialogue, then there are
[0060] Extract the content of the voice dialogue. If place names and / or personal names and / or events appear, determine the context containing place names and / or personal names and / or events to determine whether the time point of the input audio segment is before, at the present, or after the appearance of place names and / or personal names and / or events, and obtain the first content conclusion.
[0061] The array of sound sources is processed and segmented according to whether they appear before, at the present moment, or after the location name and / or person name and / or event.
[0062] Based on conclusion one, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
[0063] When analyzing time feature information, the input audio segment is analyzed to determine whether it contains direct and / or indirect time information.
[0064] When direct temporal feature information is included, the array sound source is segmented according to the direct temporal information and compared with the input audio segment to determine the exact location of the input audio segment within the array sound source.
[0065] When indirect time feature information is included, the indirect time feature information is first analyzed to transform it into direct time feature information, and then the input audio segment is compared with the array sound source according to the direct time feature information.
[0066] When analyzing input audio segments, obtaining the time information within the audio segments allows for a faster and more efficient location of the audio segments.
[0067] Example 2
[0068] When the content information features no voice dialogue, then there is
[0069] Content features are extracted from the segmented waveform to determine whether the extracted content is animal sound, object sound, or natural sound, leading to content conclusion two.
[0070] The array of sound sources is processed and segmented according to the time of occurrence of animal sounds, object sounds, or natural sounds.
[0071] According to conclusion two, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
[0072] When analyzing time feature information, the input audio segment is analyzed to determine whether it contains direct and / or indirect time information.
[0073] When direct temporal feature information is included, the array sound source is segmented according to the direct temporal information and compared with the input audio segment to determine the exact location of the input audio segment within the array sound source.
[0074] When indirect time feature information is included, the indirect time feature information is first analyzed to transform it into direct time feature information, and then the input audio segment is compared with the array sound source according to the direct time feature information.
[0075] When analyzing input audio segments, obtaining the time information within the audio segments allows for a faster and more efficient location of the audio segments.
[0076] Example 3
[0077] When the content information features include both speech dialogue information and animal sounds and / or object sounds and / or natural sounds, then there are
[0078] Content features are extracted from the segmented waveform. The extracted content is determined to be animal sounds and / or object sounds and / or natural sounds and / or place names and / or person names and / or events, leading to content conclusion three.
[0079] The array of sound sources is processed and segmented according to animal sounds, object sounds, natural sounds, and / or place names and / or person names and / or the time of occurrence of events.
[0080] According to conclusion three, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
[0081] When analyzing time feature information, the input audio segment is analyzed to determine whether it contains direct and / or indirect time information.
[0082] When direct temporal feature information is included, the array sound source is segmented according to the direct temporal information and compared with the input audio segment to determine the exact location of the input audio segment within the array sound source.
[0083] When indirect time feature information is included, the indirect time feature information is first analyzed to transform it into direct time feature information, and then the input audio segment is compared with the array sound source according to the direct time feature information.
[0084] When analyzing input audio segments, obtaining the time information within the audio segments allows for a faster and more efficient location of the audio segments.
Claims
1. A method for precise sound source localization based on an acoustic sensor array, characterized in that, Includes the following steps, Step 1, Input audio clip: Capture the audio clip using an acoustic sensor; Step 2, Create audio image: Create a waveform image of the input audio; Step 3, Obtain Feature Waveforms: Segment the waveform image to extract the changing sound waveforms; Step 4, Waveform Analysis: Analyze the captured waveform image and the input audio segment to obtain the feature information contained in the audio. Step 5, Feature Information Analysis: Analyze the features of the information obtained in Step 4, and compare the obtained feature information with the array sound source; Step 6, Result Judgment: If the array sound source does not contain the feature information of the current audio segment, the input audio segment is directly output; if the array sound source contains the input audio segment, the position of the audio segment in the array sound source is derived when outputting the input audio segment. Step 7: Output the results; When creating audio images, the audio image is created based on the audio decibel level. The specific creation method is as follows: Establish a two-dimensional coordinate system with time on the horizontal axis and decibels on the vertical axis. Based on the change of decibels in an audio segment over time, obtain a waveform image and mark the multiple audio time segments with the highest decibels in the audio segment, which are denoted as audio segments. In step three, based on the already marked audio segments, the characteristic waveform is segmented from the start point to the end point of the characteristic waveform; In step four, the waveform analysis is divided into content feature information analysis and time feature analysis; Content feature information analysis involves analyzing the audio content based on the segmented waveform image, determining the type of audio content, classifying the audio, and then analyzing and interpreting the classified audio content to determine the content feature information in the segmented waveform. The temporal feature information analysis involves analyzing the input audio segment to determine whether it contains temporal feature information. If it does, the position of the input audio file in the array sound source is determined based on the temporal feature information. If it does not contain temporal feature information, the temporal feature information is no longer analyzed.
2. The method for precise sound source localization based on an acoustic sensor array according to claim 1, characterized in that, When the content information features are voice dialogue, then there are The audio dialogue content is extracted. If place names and / or personal names and / or events appear, the context containing the place names and / or personal names and / or events is determined to ascertain whether the time point of the input audio segment is before, at the present, or after the occurrence of the place names and / or personal names and / or events, thus obtaining the first content conclusion. The array of sound sources is processed and segmented according to whether they appear before, at the present moment, or after the location name and / or person name and / or event. Based on conclusion one, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
3. The method for precise sound source localization based on an acoustic sensor array according to claim 1, characterized in that, When the content information features no voice dialogue, then there is Content features are extracted from the segmented waveform to determine whether the extracted content is animal sound, object sound, or natural sound, leading to content conclusion two. The array of sound sources is processed and segmented according to the time of occurrence of animal sounds, object sounds, or natural sounds. According to conclusion two, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
4. The method for precise sound source localization based on an acoustic sensor array according to claim 1, characterized in that, When the content information features include both speech dialogue information and animal sounds and / or object sounds and / or natural sounds, then there are Content features are extracted from the segmented waveform. The extracted content is determined to be animal sounds and / or object sounds and / or natural sounds and / or place names and / or person names and / or events, leading to content conclusion three. The array of sound sources is processed and segmented according to animal sounds, object sounds, natural sounds, and / or place names and / or person names and / or the time of occurrence of events. According to conclusion three, the input audio segment is compared with the segmented array sound source to determine the exact location of the input audio segment within the array sound source.
5. The method for precise sound source localization based on an acoustic sensor array according to claim 2, 3, or 4, characterized in that, When analyzing time feature information, the input audio segment is analyzed to determine whether it contains direct and / or indirect time information. When direct temporal feature information is included, the array sound source is segmented according to the direct temporal information and compared with the input audio segment to determine the exact location of the input audio segment within the array sound source. When indirect time feature information is included, the indirect time feature information is first analyzed to transform it into direct time feature information, and then the input audio segment is compared with the array sound source according to the direct time feature information.