Method, system, and computer-readable medium for comparing audio files and audio samples
By deconvolution processing the self-coherent sequence of audio files and deformed audio, the robustness problem of audio detection in noisy environments is solved, and accurate positioning under low signal-to-noise ratio conditions is achieved.
Patent Information
- Application Number
- CN202010854872.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-24
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-03-29
AI Technical Summary
In a complex acoustic environment, especially when there is a noise source and the signal-to-noise ratio is low, it is difficult to stably detect whether an audio playback device is playing a specific audio and locate its playback time.
The self-coherent sequence of the audio file and the deformed audio is obtained, and deconvolution processing is performed. The self-coherent sequence is used as a deconvolution kernel to process the coherent time series to identify and locate the audio file and audio sample.
In complex acoustic environments, it can more accurately locate the audio time position, has good robustness, and is suitable for low signal-to-noise ratio conditions.
Smart Images

Figure CN114090820B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of audio signal processing, and more particularly to methods, systems, and computer-readable media for comparing audio files and audio samples. Background Art
[0002] It is often necessary to search for a sample that matches a short audio clip in a database of sound samples. For example, some websites provide such a service: by inputting or uploading a short piece of music audio, it is possible to quickly search and match in a music database of millions of music samples to find the entire music sample including the short music clip uploaded by the user. To achieve this goal, some existing algorithms perform compact sound texture extraction on both the retrieved audio and the music in the database. Although the algorithm can quickly match the sound pattern in a relatively quiet environment, due to the compactness of the algorithm's feature extraction of the sound texture, it is sensitive to the influence of irrelevant ambient sounds around it. In the presence of noise sources and low signal-to-noise ratio conditions, the algorithm's robustness to the retrieved sound pattern is flawed. Therefore, there is a need in the prior art to provide a solution that can stably detect whether an audio playback device is playing the detected audio and locate its playback time in a complex acoustic environment, in the presence of noise sources and with a low signal-to-noise ratio.
[0003] The contents of the background technology section are merely the technologies known to the inventors and do not necessarily represent the existing technologies in this field. Summary of the Invention
[0004] In view of at least one drawback of the prior art, the present invention provides a method for identifying an audio sample, comprising:
[0005] S102: Obtain a self-coherent sequence of an audio file and a deformed audio, wherein the deformed audio is obtained based on the audio file;
[0006] S103: Obtaining a coherent time series between the audio sample and the audio file;
[0007] S104: using the self-coherent sequence as a deconvolution kernel to perform deconvolution processing on the coherent time series;
[0008] S105: Identify and / or locate the audio file and / or the audio sample according to the coherent time series after deconvolution.
[0009] According to one aspect of the present invention, the deformed audio includes inserting a silent segment at the front and / or back of the audio file, or the deformed audio includes inserting the audio file at the front and / or back of the audio file.
[0010] According to one aspect of the present invention, the method further comprises step S101: obtaining a complex frequency spectrum of the audio file.
[0011] According to one aspect of the present invention, the length of the audio sample is greater than the length of the audio file, and step S103 includes obtaining the coherence time series by a sliding window method, where the width of the sliding window is the same as the length of the audio file, and step S103 includes:
[0012] S103-1: Compare the portion of the audio sample within the sliding window with the audio file to obtain a coherence index;
[0013] S103 - 2 : Slide the sliding window over the audio samples and repeat step S103 - 1 to obtain the coherent time series.
[0014] According to one aspect of the present invention, the coherence index is a frequency-weighted average of coherence coefficients between the portion of the audio sample within the sliding window and the audio file at each frequency.
[0015] According to one aspect of the present invention, the length of the audio sample is less than the length of the audio file, and step S103 includes obtaining the coherence time series by a sliding window method, where the width of the sliding window is the same as the length of the audio sample, and step S103 includes:
[0016] S103-1: Compare the portion of the audio file within the sliding window with the audio sample to obtain a coherence index;
[0017] S103-2: Slide the sliding window over the audio file and repeat step S103-1 to obtain the coherent time series
[0018] According to one aspect of the present invention, the coherence index is a frequency-weighted average of coherence coefficients of the portion of the audio file within the sliding window and the audio sample at each frequency.
[0019] According to one aspect of the present invention, step S104 includes: using the LASSO method to perform fitting in the time domain to perform deconvolution processing.
[0020] According to one aspect of the present invention, step S105 includes: locating the audio file and / or the audio sample by peak detection according to the coherence time series after deconvolution.
[0021] The present invention also provides a system for comparing audio files and audio samples, comprising:
[0022] a unit for obtaining a self-coherent sequence of the audio file and a deformed audio, wherein the deformed audio is obtained based on the audio file;
[0023] A unit for obtaining a coherence time series of the audio sample and the audio file;
[0024] a unit for performing deconvolution processing on the coherent time series using the self-coherent sequence as a deconvolution kernel; and
[0025] The units of the audio file and / or the audio samples are located according to the deconvolved coherence time series.
[0026] The present invention also provides a computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, can implement the method described above.
[0027] In this embodiment of the present invention, the coherence time series between the audio sample and the audio file is deconvolved with the audio file's self-coherence time series, enabling more accurate localization of the retrieved audio's temporal position. This embodiment of the present invention has been proven to be highly robust in complex real-world scenarios, such as low signal-to-noise ratio environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0029] Figure 1 A method for comparing audio files and audio samples according to one embodiment of the present invention is shown;
[0030] Figure 2 An example of a coherence sequence obtained from a real audio signal is shown;
[0031] Figure 3 The autocorrelation sequence used as the deconvolution kernel is shown;
[0032] Figure 4 The final output similarity time series results are shown;
[0033] Figure 5 A method for identifying audio samples according to a preferred embodiment of the present invention is shown. DETAILED DESCRIPTION
[0034] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present invention. Therefore, the drawings and description are to be considered as illustrative in nature and not restrictive.
[0035] In the description of the present invention, it should be understood that terms such as "center," "longitudinal," "transverse," "length," "width," "thickness," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," "clockwise," and "counterclockwise" are used to indicate positions or relationships based on those shown in the accompanying drawings. These terms are intended solely to facilitate description and simplify the description of the present invention and are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed to indicate or imply relative importance or to implicitly specify the number of the technical features referred to. Thus, features designated "first" or "second" may explicitly or implicitly include one or more of the designated features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0036] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, removable, or integral connections; mechanical, electrical, or intercommunication connections; direct or indirect connections through an intermediary; and internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of these terms in the present invention based on specific circumstances.
[0037] In the present invention, unless otherwise expressly specified or limited, "above" or "below" a first feature may include direct contact between the first and second features, or may include contact between the first and second features not in direct contact but via another feature between them. Furthermore, "above," "above," and "above" a first feature may include both directly above and diagonally above the second feature, or simply indicate that the first feature is at a higher level than the second feature. "Below," "below," and "below" a first feature may include both directly above and diagonally above the second feature, or simply indicate that the first feature is at a lower level than the second feature.
[0038] The disclosure below provides many different embodiments or examples for realizing different structures of the present invention. In order to simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are merely examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numbers and / or reference letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed. In addition, the present invention provides examples of various specific processes and materials, but those skilled in the art will recognize the application of other processes and / or the use of other materials.
[0039] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0040] In the prior art, given a template signal and a detection signal, a common method is to directly perform cross-correlation counting on them for identification and location. This method is used to determine whether the template signal appears in the detection signal and where the template signal appears. However, in practice, this method is significantly affected by the limitations of the actual recording equipment and ambient noise, resulting in poor results. The present invention proposes a new method for identifying and / or locating the template signal in the detection signal. This is described below with reference to the accompanying drawings.
[0041] Figure 1 A method 100 for comparing an audio file s(t) and an audio sample q(t) according to one embodiment of the present invention is shown. The method can be used to identify whether the content of the audio file s(t) appears in the audio sample q(t), and if possible, further identify the specific location of the occurrence. A detailed description is provided below with reference to the accompanying drawings. The audio file s(t) is, for example, a template signal, and the audio sample q(t) is, for example, an object to be compared. The audio file s(t) can be, for example, a real-time audio signal, or it can be a sound wave, a radio wave, a digital audio pulse code modulation stream, a compressed digital audio stream, or an internet broadcast stream, all of which are within the scope of the present invention.
[0042] According to one embodiment of the present invention, for audio files, in actual application, if a piece of music is several minutes long, a 10-second segment containing features that significantly distinguish it from other music (other music included in the database) can be selected as the audio file.
[0043] In step S101: the complex frequency spectrum of the audio file s(t) is obtained.
[0044] A method for calculating the complex frequency spectrum of a signal according to an embodiment of the present invention is given below.
[0045] This calculation method can be expressed based on the complex time-frequency distribution (TFD) of an audio signal. Assume that an audio signal is y(t), where t represents time and y(t) represents the intensity of the audio signal at time t. Through the short-time Fourier transform (STFT), y(t) can be converted to Y(n, f), where n represents the number of frames in time and f represents the frequency. In the short-time Fourier transform, for example, a Hann window of length 256 can be used. Y(n, f) is the complex time-frequency distribution (TFD) of the sound signal y(t), which is a function of time and frequency. Unlike a general time-frequency spectrum (spectrogram) that is converted to a real number after arithmetic, the complex time-frequency distribution (TFD) is in complex form.
[0046] Based on the complex time-frequency distribution TFD, the complex frequency spectrum of the signal y(t) can be obtained through further processing. For example, through the Welch method, the complex frequency spectrum of an audio signal y(t) (or Y(n, f)) is:
[0047]
[0048] According to a preferred embodiment of the present invention, the above method is applied to an audio file s(t) to obtain the complex frequency spectrum of the audio file s(t). Of course, the present invention is not limited to the above method for calculating the complex frequency spectrum, and other methods can also be used for calculation, which are all within the scope of protection of the present invention.
[0049] In step S102 : a self-coherent sequence of the audio file and a deformed audio is obtained, wherein the deformed audio is obtained based on the audio file.
[0050] The deformed audio s'(t) is obtained by modifying the audio file s(t). For example, a silent segment can be inserted at the beginning and / or end of the audio file s(t) to obtain the deformed audio s'(t). The length of the silent segment is preferably equal to the length of the audio file. Alternatively, the audio file can be inserted at the beginning and / or end of the audio file, that is, the audio file is copied three times its original length to obtain the deformed audio. The deformed audio can also be obtained by other methods based on the audio file.
[0051] The self-coherence sequence of the audio file and the deformed audio can be obtained through various methods. The self-coherence sequence is obtained by calculating the similarity between the deformed audio and the audio file itself. The following describes methods for calculating the coherence index for signals of the same length and for signals of different lengths.
[0052] Since the coherence time series of the audio sample q(t) and the audio file s(t) is also calculated below, the audio sample q(t) and the audio file s(t) are used as an example here to illustrate the technical methods of the coherence coefficient, coherence index and coherence sequence.
[0053] Given signals of the same length at both ends: audio file (template signal) s(t), detected signal (audio sample) q(t), at each frequency f, their coherence coefficient is calculated using the following formula:
[0054]
[0055] Where W s (f) is the complex frequency spectrum of the template signal s(t), W q (f) is the complex frequency spectrum of the detection signal q(t), which can be calculated, for example, according to the method described in step S101. The coherence coefficient ranges from 0 to 1, where 0 indicates that the two are completely unrelated and 1 indicates that the two are completely coherent. According to a preferred embodiment of the present invention, a coherence score is obtained by weighted averaging in the frequency domain:
[0056]
[0057] The coherence index can indicate the degree of coherence between the template signal s(t) and the detected signal q(t) in the frequency domain.
[0058] The coherence timeseries of a long audio signal relative to a short audio template signal are calculated as follows. In the above description, it is assumed that the input detection signal q(t) is the same length as the template signal s(t), and the output is a real number index. If the length of the detection signal q(t) is greater than the template signal s(t), a coherence timeseries can be obtained by the square sliding window method. The length of the sliding window is the same as s(t). The q(t) signal in the sliding window is compared with s(t), and the coherence index C of the two is calculated according to the above method. sq After completion, the sliding window is continuously moved, and finally the similarity time series C is obtained. sq (t), the sequence includes multiple coherence indices. The step size of each movement of the sliding window can be set as needed. Generally, a smaller step size results in higher accuracy but a larger computational effort; a larger step size results in lower accuracy but a smaller computational effort. The step size can also be set by the user.
[0059] Since the complex frequency spectrum of the template signal is reused in each sliding window, it is preferable to pre-calculate the complex frequency spectrum W of the template signal. s (f) and stored.
[0060] Applying the above method to the audio file s(t) and the deformed audio s'(t) (the length of the deformed audio is longer than the audio file), that is, replacing the audio sample q(t) with the deformed audio s'(t), the self-coherent sequence of the audio file s(t) and the deformed audio s'(t) can be obtained, which will be used as the deconvolution kernel in the deconvolution operation later. Figure 3 is a self-coherent sequence obtained by a computational example (as a deconvolution kernel).
[0061] In the above step S102, the complex frequency spectrum calculated in step S101 is used to calculate the self-correlation sequence between the audio file and itself. The present invention is not limited to this, and other methods can also be used to calculate the self-correlation sequence.
[0062] In step S103 : obtaining a coherent time sequence between the audio sample and the audio file.
[0063] As mentioned above, during the detection process, given the signal to be detected q(t) and the pre-calculated frequency spectrum W of the template signal (audio file), s (f), the coherence sequence of the two can be obtained by the sliding window method. Figure 2 is an example of a coherence sequence obtained from a real audio signal.
[0064] In step S104: using the self-coherent sequence as a deconvolution kernel, deconvolution processing is performed on the coherent time series.
[0065] The coherent time series obtained in step S103 is deconvolved, and the deconvolution kernel is the auto-coherent sequence prepared in step S102. According to a preferred embodiment of the present invention, the LASSO method can be used to fit in the time domain to achieve the deconvolution effect. Figure 4 It is the final output similarity time series result.
[0066] In step S105 : the audio file and / or the audio sample is located according to the coherent time series after deconvolution.
[0067] The similarity time series obtained in the previous step can be simply tested for min / max peaks to locate the detected template audio. Given a discrete time series, such as Figure 4As shown, min / max peak detection checks whether the value at a given moment is higher than the value at the previous moment. If so, it is a potential peak. In practice, a threshold can be set: if the value at a given moment is higher than the previous moment and exceeds the set threshold, it is considered a signal peak.
[0068] Therefore, by peak detection, the audio file and / or the audio sample can be located according to the coherent time series after deconvolution, for example, the time coordinates at which the audio sample is aligned with the audio file can be found. Figure 4 Four peaks are identified in the audio sample, indicating that the template audio appears four times in the audio sample, thereby completing the positioning of the audio sample and / or audio file.
[0069] The above embodiments of the present invention are applicable to audio samples and audio files of various lengths. According to one embodiment of the present invention, the length of the audio sample is greater than the length of the audio file. In this case, step S103 includes obtaining the coherence time series using a sliding window method, where the width of the sliding window is the same as the length of the audio file. Step S103 includes:
[0070] S103-1: Compare the portion of the audio sample within the sliding window with the audio file to obtain a coherence index;
[0071] S103 - 2 : Slide the sliding window over the audio samples and repeat step S103 - 1 to obtain the coherent time series.
[0072] At this time, the coherence index is a frequency-weighted average of coherence coefficients between the portion of the audio sample within the sliding window and the audio file at each frequency.
[0073] According to another embodiment of the present invention, the length of the audio sample is less than the length of the audio file. In this case, step S103 includes obtaining the coherence time series using a sliding window method, where the width of the sliding window is the same as the length of the audio sample, and step S103 includes:
[0074] S103-1: Compare the portion of the audio file within the sliding window with the audio sample to obtain a coherence index;
[0075] S103-2: Slide the sliding window over the audio file and repeat step S103-1 to obtain the coherent time series
[0076] The coherence index is a frequency-weighted average of coherence coefficients between the portion of the audio file within the sliding window and the audio sample at each frequency.
[0077] In the above-described embodiment of the present invention, the coherence time series of the audio sample and the audio file is deconvolved using the audio file's self-coherence time series, enabling more accurate localization of the retrieved audio's temporal position. Practical verification has demonstrated that the embodiment of the present invention exhibits excellent robustness in complex real-world scenarios (e.g., low signal-to-noise ratio environments).
[0078] Figure 5 A method 200 for identifying an audio sample according to a preferred embodiment of the present invention is shown and will be described in detail below with reference to the accompanying drawings.
[0079] In step S201, a complex frequency spectrum of an audio file is obtained.
[0080] In step S202 , a self-coherent sequence of the audio file and a deformed audio is obtained, where the deformed audio is obtained based on the audio file.
[0081] According to a preferred embodiment of the present invention, the audio file is a pre-stored audio signal, so the audio file can be pre-processed to obtain its complex frequency spectrum and self-coherence sequence, which are stored for subsequent signal processing.
[0082] Assuming that the length of the audio sample is longer than the audio file (template), a sliding window approach is used to calculate the coherence coefficient. In step S203, the sliding window is first placed at the starting point of the audio sample, and the coherence coefficients of the audio sample and the audio file within the current sliding window are obtained.
[0083] In step S204, a weighted coherence index within the current sliding window is calculated, for example, by weighted averaging in the frequency domain to obtain a coherence index.
[0084] In step S205, it is determined whether the end of the audio sample has been reached. If the end of the audio sample has been reached, the process proceeds to step S206; otherwise, the process proceeds to step S207.
[0085] In step S207, the sliding window is slid forward, and then the process returns to step S203, and steps S203 and S204 are repeated.
[0086] In step S206, a deconvolution operation is performed on the coherent time series based on the audio file obtained in step S202 and its own self-coherent sequence.
[0087] In step S208, the audio sample and / or audio file is identified and / or located based on the deconvolved coherence time series. For example, it may be determined whether the audio file appears in the audio sample and, if so, the location of the audio file.
[0088] The present invention also relates to a system for comparing an audio file and an audio sample, comprising:
[0089] A unit for obtaining a complex frequency spectrum of the audio file;
[0090] a unit for obtaining a self-coherent sequence of the audio file and a deformed audio, wherein the deformed audio is obtained based on the audio file;
[0091] A unit for obtaining a coherence time series of the audio sample and the audio file;
[0092] a unit for performing deconvolution processing on the coherent time series using the self-coherent sequence as a deconvolution kernel;
[0093] The units of the audio file and / or the audio samples are located according to the deconvolved coherence time series.
[0094] The present invention also relates to a computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, can implement the method as described above.
[0095] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for identifying an audio sample, comprising: S102: Obtaining a self-coherent sequence of an audio file and a deformed audio, wherein the deformed audio is obtained based on the audio file, the deformed audio includes inserting a silent segment at the beginning and / or end of the audio file, or the deformed audio includes inserting the audio file at the beginning and / or end of the audio file, and obtaining the self-coherent sequence based on the length of the deformed audio and the length of the audio file; S103: Obtaining a coherent time series between the audio sample and the audio file based on the length of the audio sample and the length of the audio file; The step S103 includes: If the length of the audio sample is greater than the length of the audio file, and the width of the sliding window is the same as the length of the audio file, perform the following steps: S103-1: Compare the portion of the audio sample within the sliding window with the audio file to obtain a coherence index; S103-2: Slide the sliding window over the audio sample and repeat step S103-1 to obtain the coherent time series; The coherence index is a frequency-weighted average of coherence coefficients between the portion of the audio sample within the sliding window and the audio file at each frequency; If the length of the audio sample is less than the length of the audio file, the width of the sliding window is the same as the length of the audio sample, and the following steps are performed: S103-3: Compare the portion of the audio file within the sliding window with the audio sample to obtain a coherence index; S103-4: Slide the sliding window over the audio file and repeat step S103-3 to obtain the coherent time series; The coherence index is a frequency-weighted average of coherence coefficients between the portion of the audio file within the sliding window and the audio sample at each frequency; S104: performing deconvolution processing on the coherent time series using the self-coherent sequence as a deconvolution kernel; S105: Identify and / or locate the audio file and / or the audio sample according to the deconvolved coherence time series. 2 . The method according to claim 1 , further comprising step S101 : obtaining a complex frequency spectrum of the audio file.
3. The method according to claim 1, wherein the step S104 comprises: The LASSO method is used to fit in the time domain to perform deconvolution processing.
4. The method according to claim 1, wherein the step S105 comprises: The audio file and / or the audio sample is located by peak detection according to the coherence time series after deconvolution.
5. A system for comparing an audio file and an audio sample, comprising: a unit for obtaining a self-coherent sequence of the audio file and a deformed audio, wherein the deformed audio is obtained based on the audio file, the deformed audio includes inserting a silent segment at the beginning and / or end of the audio file, or the deformed audio includes inserting the audio file at the beginning and / or end of the audio file, and the self-coherent sequence is obtained based on the length of the deformed audio and the length of the audio file; A unit for obtaining a coherence time series between the audio sample and the audio file, wherein the coherence time series is obtained based on the length of the audio sample and the length of the audio file. If the length of the audio sample is greater than the length of the audio file, the width of the sliding window is the same as the length of the audio file, and the portion of the audio sample within the sliding window is compared with the audio file to obtain a coherence index; the sliding window is slid across the audio sample, and the step of obtaining the coherence index is repeated to obtain the coherence time series; wherein the coherence index is the coherence index between the portion of the audio sample within the sliding window and the audio file at each frequency. The method comprises the steps of: calculating a frequency-weighted average of the coherence coefficients of the audio sample; if the length of the audio sample is less than the length of the audio file, and the width of the sliding window is the same as the length of the audio sample, comparing the portion of the audio file within the sliding window with the audio sample to obtain a coherence index; sliding the sliding window over the audio file, and repeating the step of comparing the portion of the audio file within the sliding window with the audio sample to obtain a coherence index, thereby obtaining the coherence time series; wherein the coherence index is the frequency-weighted average of the coherence coefficients of the portion of the audio file within the sliding window and the audio sample at each frequency; a unit for performing deconvolution processing on the coherent time series using the self-coherent sequence as a deconvolution kernel; and Units of the audio file and / or the audio samples are identified and / or located based on the deconvolved coherence time series.
6. A computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, can implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Audio retrieval method and device
CN105893549A
Similarity calculation method based on multiple sound characteristics
CN107610715A
Method of, and apparatus for, full waveform inversion
US20160238729A1