Systems and methods for estimation of frequencies thereof
Patent Information
- Application Number
- TW113142522
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-05-02
- Filing Date
- 2024-11-06
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2044-11-05
AI Technical Summary
Existing audio systems struggle to accurately and efficiently detect and quantify harmonic distortions in audio signals, particularly in real-time, due to the reliance on lengthy measurement times and sensitivity to noise, which limits their application in high-quality audio processing.
A machine learning-based system using sparse input signals derived from audio signals, processed through phase correction and normalization, to identify frequency increases in audio signals, employing neural networks for rapid and accurate harmonic detection.
The system achieves high-accuracy harmonic detection within a short time frame, reducing measurement time by up to 600 times compared to conventional methods, enabling real-time correction of harmonic distortions in audio systems.
Smart Images

Figure TWG2TB001905458_001 
Figure TWG2TB001905458_002 
Figure TWG2TB001905458_003
Abstract
Description
[Technical Field]
[0001] This invention relates to the processing of audio signals, and more generally to a method and system for measuring the frequency of audio signals. [Previous Technology]
[0002] When an audio system adds an input signal, it is generally considered to be of high quality when the proportion of audio artifacts as a byproduct of the system itself is kept to a minimum. These artifacts can be classified as noise, non-harmonic distortion, and harmonic distortion. It is necessary to detect and quantify these artifacts in order to design better systems and provide real-time control of automatic tuning systems.
[0003] Previous patent documents have disclosed how to use machine learning (ML) to quantify the quality of signals. For example, U.S. Patent Publication No. 2023 / 0136698, which has been assigned to the applicant in this case, discloses a system comprising a memory and a processor, the memory being configured to store an ML model, the processor being configured to (1) acquire a set of training audio signals labeled with corresponding levels of distortion; (2) convert the training audio signals into individual images; (3) train the ML model based on the images to estimate the level of distortion; (4) receive an input audio signal; (5) convert the input audio signal to an image; and (6) estimate the level of distortion of the input audio signal by applying the trained ML model to the image.
[0004] Another U.S. Patent Publication No. 2023 / 0136220, also assigned to the applicant in this case, discloses a system comprising a memory and a processor. The memory is configured to store an ML model, and the processor is configured to (1) acquire a set of training audio signals existing in the form of a plurality of initial audio signals, the training audio signals having a first duration and being labeled with individual distortion levels, the first duration being within a first duration range; (2) train the ML model based on the training audio signals to estimate the distortion level; (3) receive an input audio signal having a second duration shorter than the first duration, the second duration being within a second duration range; and (4) estimate the distortion level of the input audio signal by applying the trained ML model to the input audio signal. [Summary of the Invention]
[0005] An embodiment of the present invention discloses a frequency estimation system comprising a memory and a processor. The memory is configured to store a machine learning model trained to estimate frequency increases in a plurality of sparse input signals derived individually from a plurality of input audio signals, the sparse input signals being used to identify one or more frequency increases in the corresponding input audio signals. The processor is configured to: (1) receive a first input audio signal; (2) derive a first sparse input signal from the first input audio signal, the first sparse input signal being used to identify frequency increases in the first input audio signal; and (3) estimate multiple frequency increases in the first input audio signal by applying the trained machine learning model to the first sparse input signal.
[0006] In one embodiment, the processor is configured to derive the first sparse input signal from the first input audio signal by retaining the portion of the first input audio signal that is adjacent to zero crossover and discarding the other portion of the first input audio signal.
[0007] In another embodiment, the processor is configured to derive the first sparse input signal from the first input audio signal by retaining portions of the first input audio signal that are near extreme values and discarding other portions of the first input audio signal.
[0008] In another embodiment, the processor is configured to derive the first sparse input signal from the first input audio signal by retaining the steepest portion of the first input audio signal and discarding the other portions of the first input audio signal.
[0009] In one embodiment, the processor is further configured to perform phase correction on the first input audio signal by applying an initialization step, and to derive the first sparse input signal from the first input audio signal.
[0010] In one embodiment, the processor is configured to detect the frequencies of one or more higher harmonics in the first input audio signal in order to estimate the increase in the value of those frequencies.
[0011] In some embodiments, the processor is configured to receive the first input audio signal to acquire the first input audio signal.
[0012] In some embodiments, the processor is further configured to filter out the DC component from the first input audio signal.
[0013] In one embodiment, the processor is further configured to normalize the first input audio signal.
[0014] In some embodiments, the machine learning model includes one of a convolutional neural network and a recursive neural network.
[0015] In some embodiments, the processor is further configured to control one of the audio systems that generates the first input audio signal using the estimated frequency increase values.
[0016] An embodiment of the present invention further discloses a frequency estimation system comprising a memory and a processor. The memory is configured to store a machine learning model. The processor is configured to: (1) acquire a plurality of audio signals individually labeled according to a plurality of frequency increase values; (2) derive a plurality of training sparse signals from the individual audio signals, each training sparse signal being used to identify one or more frequency increases in a corresponding audio signal; and (3) train the machine learning model using the training sparse signals to estimate the plurality of frequency increase values.
[0017] In some embodiments, the processor is configured to receive a plurality of initial audio signals having a first duration, and to segment the initial audio signals into a plurality of segments having a second duration, thereby acquiring a plurality of audio signals, the second duration being shorter than the first duration.
[0018] In one embodiment, the processor is further configured to filter out a DC component from each of the audio signals.
[0019] In another embodiment, the processor is further configured to normalize each of the audio signals.
[0020] In some embodiments, the machine learning model includes one of a convolutional neural network and a recursive neural network.
[0021] In some embodiments, the convolutional neural network classifies the frequency increase based on the numerical value of the frequency increase and marks it in the audio signal.
[0022] An embodiment of the present invention further discloses a frequency estimation method, comprising: storing a machine learning model in a memory, the machine learning model being trained to estimate multiple frequency increase values among multiple sparse input signals, the sparse input signals being derived individually from multiple input audio signals, the sparse input signals being used to identify one or more frequency increases among the corresponding input audio signals; receiving a first input audio signal; deriving a first sparse input signal from the first input audio signal, the first sparse input signal being used to identify frequency increases in the first input audio signal; and estimating multiple frequency increase values in the first input audio signal by applying the trained machine learning model to the first sparse input signal.
[0023] An embodiment of the present invention also discloses a frequency estimation method, comprising: storing a machine learning model in a memory; acquiring a plurality of audio signals, the audio signals being individually labeled according to a plurality of frequency increase values; individually deriving a plurality of training sparse signals from the audio signals, each training sparse signal being used to identify one or more frequency increases in a corresponding audio signal; and using the training sparse signals to train the machine learning model to estimate the values of the plurality of frequency increases.
Implementation Method
[0025] Overview
[0026] Audio signals (such as music or sound) exist primarily in the form of sound wave energy. In consumer technology products, this energy is typically converted to digital formats for various manipulations, such as storage, processing, or broadcasting. The distortions that may result from these manipulations are generally considered negative artifacts. These distortions can now be measured using high-accuracy analyzers, but achieving high accuracy requires relatively long measurement times, typically around 667 msec in the industry, thus limiting their application. Within one or more harmonics of a given fundamental frequency, even a single distortion can cause a small change (i.e., error). When the original signal can be encoded into multiple fundamental frequencies, one or more additional harmonics can cover a fairly wide frequency range. For simplicity, the disclosed content will consider the given fundamental frequency as 1 kHz. Furthermore, the disclosed technique is also applicable to a set of fundamental frequencies.
[0027] In some embodiments of the present invention, a machine learning (ML)-based technique is disclosed for detecting one or more increasing harmonic frequencies in a signal, such as an odd or even higher harmonic that increases to the fundamental frequency. This technique is also applicable over a wide bandwidth to detect any number of increasing harmonics. A processor using an ML algorithm can provide a faster method for identifying increasing frequencies while maintaining high accuracy and low or no signal noise sensitivity.
[0028] In one embodiment, the processor uses a trained artificial neural network (ANN) to detect and quantize an increased harmonic in a test input audio signal that rises to a fundamental frequency, and to obtain a representative accuracy error within 0.1% over a very short time span across multiple cycles. For example, the fundamental frequency is 1 kHz, and the duration is 5 msec. Using commercially available analyzers on the same test signal would require much longer (approximately 600 cycles) to provide similar results.
[0029] To effectively train an ML model, a preprocessing step is applied to a processor. This step includes deriving a set of sparse training signals from a set of labeled short audio signals. Each signal at an increment frequency in a harmonic has a known frequency error (FE). Generally, any increment harmonic signal will have a frequency significantly higher than the fundamental harmonic, for example, at least 10% higher. Smaller frequency differences at the fundamental frequency can be considered as jitter or timing errors.
[0030] In some examples, a given training sparse signal can identify frequencies in the corresponding audio signal derived from it. In one example, frequency identification is performed using a sparse signal that has a sequence of signal portions with nearest-zero crossovers in the corresponding audio signal. In another example, the sparse signal captures extrema in the signal to identify frequencies. In yet another example, the sparse signal captures the signal region with the steepest change to identify frequencies.
[0031] The training sparse signal retains information about the correlated increase frequency in the input audio signal, but the size of this information is quite small and easy to process. A set of labeled training sparse signals can be derived from the short audio signal using the various methods described above (e.g., the neighboring values of zero crossovers, the neighboring values of extrema, and the neighboring values of the steepest regions in the input signal).
[0032] In one example, the training sparse signal can be derived from the short audio signal by preserving the intervals adjacent to each zero-crossing interval and discarding the rest of the short audio signal. In another example, the training sparse signal can be derived from the short audio signal by generating a function that has a spike at zero crossings and discarding the rest of the short audio signal.
[0033] Another step that can reduce training workload is to apply a phase correction, such as a zero-crossed aligned preprocessing step, to a set of labeled short audio signals. A set of signals that has undergone this phase correction (e.g., zero-crossed aligned training) is more efficient than using the original short audio signals because it saves the processor from having to consider the least relevant data points (i.e., most of the signal). Other methods for phase correction of the input signal rely on extremum correction or correction of the steepest changes in the signal.
[0034] In one example, for training purposes, the set of raw data includes harmonic signal samples with a fundamental (i.e., basic) frequency of 1 kHz. (The terms "audio signal" and "audio sample" disclosed in this case are considered equivalent and interchangeable.) To simulate the higher harmonics in the audio sample, a harmonic signal with a higher frequency is randomly added to the audio sample in the range of 1.09 kHz to 9 kHz, wherein the minimum peak level of the added harmonic is at least 1% of the amplitude of the fundamental signal.
[0035] To further simulate real-world scenarios, random noise is added with an amplitude up to 1% of the fundamental signal amplitude. This set of audio samples is repeated ten times with different random phase values (the phase lies between the fundamental frequency and the increased frequency). After summing, a set of approximately 200,000 sparse signals is derived from individual short audio samples to train the example ML model. Increasing the database to millions or more will improve the detection accuracy.
[0036] In the inference phase, the trained system receives an input audio signal with a short duration (e.g., within 10 ms) and may contain one or more increasing harmonics. After deriving individual sparse signals, the trained system predicts the degree of distortion of the input audio signal (e.g., a list of one or more increasing harmonic frequencies) over the duration. Example simulation results show an accuracy of approximately 5%. Generally, the fundamental harmonic frequency signal is more accurate than distortions that may be non-harmonic, noise, etc. Therefore, the technique disclosed herein is suitable for detecting harmonics increasing to a fundamental frequency, rather than for quantifying jitter or timing-related problems caused by a system error that modifies the fundamental frequency.
[0037] The technology disclosed in this case can be applied to the analysis of offline or real-time audio signals. This example system helps in accurate system analysis (product design phase) and can also achieve real-time control to correct distortion caused by spurious harmonics added to the audio system.
[0038] System Description
[0039] According to an embodiment of the present invention, Figure 1 is a block diagram illustrating a frequency estimation system 101 used to estimate the frequencies of one or more added harmonics of a short audio sample emitted from an audio processing device 100. The output signal of the audio processing device 100 is directed to an output device 110, such as a speaker. The output of the frequency estimation system 101 is used to correct errors in the frequency domain of the audio processing device 100, enabling it to output a higher quality signal to the output device 110.
[0040] To train an ML model, namely an ANN 107 (e.g., an RNN or a 1D CNN), one of the processors 108 in the frequency estimation system 101 uses a set of labeled short training sparse signals 125, each short training signal having a known frequency of a higher harmonic, which is added to the fundamental harmonic. Initially, the frequency estimation system 101 receives short audio signals 121, or the system can receive long signals and segment the initial audio signals into several segments of shorter duration, and then generate a set of initial training audio signals.
[0041] In order to reduce irrelevant variables in the signal, the system applies a DC filter 102 to remove the DC offset in the signal (e.g., short audio signal 121) and then generates a signal similar to short audio signal 123.
[0042] A sparse signal generator 104 performs a range-based preprocessing step on the short audio signal 123 by converting each short audio signal 123 to an individual sparse signal 125, for example, only near the zero-crossing point. The sparse signal generator 104 can also remove different phases and delays between the training sparse signals. This phase correction preprocessing step eliminates an unnecessary search for a zero-crossing during a full cycle of signal training. When training an ML model to detect and predict one or more increasing frequencies, it is more efficient to use the sparse signal 125 than the short audio signal 123.
[0043] In one preferred embodiment, the sparse signal generator 104 first applies a zero-crossing correction preprocessing step (e.g., phase alignment) to a set of labeled short audio signals 123. When all the zero-crossing corrected short audio signals 123 start with a similar initial amplitude and phase, the short audio signals 123 can become sparser and more efficient through the sparse signal generator 104.
[0044] Subsequently, the sparse signal 125 is digitized by a circuit implemented by the processor 108, and the waveform of the sparse signal 125 is normalized when necessary. The circuit digitizes and normalizes the initial signals using a given minimum digital precision (e.g., 8 bits), and higher precision (e.g., 24 bits) will achieve better accuracy.
[0045] During a training phase, the processor 108 runs an algorithm to optimize the ANN 107 (e.g., determine the weights) and stores the optimized ANN in a memory 109.
[0046] During the inference phase, the frequency estimation system 101 is configured to perform a one-dimensional (1D) prediction of one or more increased frequencies in a short audio signal sample that is similar to one of the short audio signals 121. The processor 108 runs a trained ANN 107 and stores it in memory 109 to perform inference on the sparse signal 125, inferring the frequency domain error in the signal. In one embodiment, the ANN 107 is a long short-term memory (LSTM) RNN.
[0047] Finally, a feedback line 106 between the processor 108 and the audio processing device 100 can control the amount of frequency increase (FA) in the output audio signal in real time according to the estimated increase frequency.
[0048] The various components in the frequency estimation system 101 shown in Figure 1 and the audio processing device 100 can be implemented using suitable hardware, such as one or more discrete components, one or more application-specific integrated circuits (ASICs), and / or one or more field-programmable gate arrays (FPGAs). Some functions in the frequency estimation system 101, such as functions of the processor 108, can be implemented in one or more general-purpose processors, which are programmed in software to carry out the aforementioned functions. The software can be downloaded to the processor electronically via a network or from a host computer, for example, alternatively or otherwise provided and / or stored in non-transitory tangible media, such as magnetic memory, optical memory, or electronic memory.
[0049] For clarity, the embodiment shown in Figure 1 is described by way of example. Any other suitable configuration may be used in an alternative embodiment, for example, the circuitry for preprocessing may perform another type of preprocessing for the short audio signal 121 used as initial training samples.
[0050] Short audio signals use an artificial neural network (ANN) to determine the FA.
[0051] According to an embodiment of the present invention, Figure 2 is a graph 200, which discloses a short audio signal 206 having a frequency increase (FA). The short audio signal 206 is generated by adding a harmonic 204 to a fundamental frequency signal 202, the fundamental frequency signal being 1 kHz with an amplitude of -6 dB (peak level approximately 0.5). In Figure 2, the added harmonic 204 is -33 dB (peak level approximately 0.022, for enhanced visibility in the graph), and has a frequency of 7850 Hz and a randomly given phase value.
[0052] If the initial signal is very long (e.g., lasting hundreds of cycles), the system will truncate (e.g., segment) the training audio sample, leaving only a few (e.g., 5) cycles. Therefore, the training uses samples of short duration (e.g., 5 cycles of a 1 kHz wave), with each sample having a total duration of 5 msec. This duration is considered very short and not permitted for analysis, as emphasized above, such as meaningful Fast Fourier Transform (FFT) analysis of harmonic distortion.
[0053] According to an embodiment of the present invention, Figure 3 is a diagram 300 and discloses an example of an actual short audio signal 306 having one of the additional noise 304, which is derived from the short audio signal 206 in Figure 2.
[0054] The actual noise level (for better visibility in the figure) is at level 1, which is -82.1 dB (the peak level of short audio signal 206 is approximately 7.8E-5).
[0055] In one example, after converting a set of short audio signals 306 to the final training zero-crossed-aligned signals and making the signals sparse (as shown in Figure 1, sparse signal 125), the final audio sample is digitally sampled at a sampling rate of 48 kHz (the most common sampling rate for audio systems).
[0056] In order to train an artificial neural network (ANN) to detect sparse signals of FE.
[0057] According to an embodiment of the present invention, Figures 4A and 4B are graphs of two types of sparse signals for training, which show the short audio signal 306 from Figure 3.
[0058] Figure 4A discloses a training sparse signal 402, from which the signal portion 404 is derived by maintaining the intervals adjacent to each zero crossover in the short audio signal 306 and discarding the other parts of the short audio signal 306.
[0059] Figure 4B reveals a training sparse signal 406, which can be derived from the short audio signal 306 by generating a function 408 with a spike at zero crossover and discarding other parts of the short audio signal 306.
[0060] Performance analysis of FE estimation using artificial neural networks (ANN)
[0061] In the field of machine learning, particularly regarding statistical classification problems, a confusion matrix, also known as an error matrix, is a specific tabular layout that allows visualization of the performance of an algorithm, typically a supervised learning algorithm (that is, an algorithm that learns using labeled training data). Each column in the matrix represents an instance in an actual class, while each row represents an instance in a predicted class (or vice versa). Regardless of whether the system confuses the two classes (that is, often mislabeling one name as another), names derived from reality will make the name easy to see.
[0062] According to an embodiment of the present invention, Figure 5 is a graph 500 showing a comparison between the estimated detection frequency of an added harmonic using the system in Figure 1 and the ground-truth frequency of the added harmonic. This graph is generated as part of proof-of-concept testing in the disclosed technology, and the frequency of the added harmonic ranges from 1.1 kHz to 9 kHz. In this example, a dataset of 200,000 data points is used, each harmonic having an energy level more than 10% higher than the fundamental harmonic, and each sample having 5 cycles. A trend line 502 presents ideally correct results, indicating that a high-quality audio system will have small amplitude and noise when adding harmonics. When using the system in Figure 1 on a given dataset, the system detects the added frequency of the harmonic at a rate of approximately 0.1% 501.
[0063] When the energy level of an increased harmonic increases rapidly, for example, to the level of an increased harmonic in the fundamental harmonic under extreme distortion, zero crossover or other indicators (such as extreme values) may not be distinctive enough to distinguish one increased frequency from another. In this case, the detected FE will increase significantly, and this increase will be reflected by the vertical width of the increased frequency values in a set of selected data.
[0064] Method for estimating FA in a short audio signal
[0065] According to an embodiment of the present invention, Figure 6 is a flowchart illustrating a method for estimating the frequency increase (FA) in a sample of a short audio signal 121 using the system in Figure 1. According to this embodiment, the algorithm executes a procedure divided into a training phase 601 and an inference phase 603.
[0066] For clarity, the training and inference phases described above are presented as a single process executed by the frequency estimation system 101. In practical implementation, the training phase 601 may be executed by a system (similar to or different from the frequency estimation system 101) and the inference phase 603 may be executed by another system (e.g., the frequency estimation system 101 in Figure 1). In such an embodiment, the frequency estimation system 101 provides a preprocessing ML model for one of the ANN 107.
[0067] The training phase begins in an upload step 602, during which the processor 108 uploads a set of short audio samples (e.g., segmented) from the memory 109. These samples are similar to the samples of the 5-cycle short audio signal 306 used in Figure 3.
[0068] In a preprocessing step 604, the sparse signal generator 104 converts each of the short audio signals 123 into an individual sparse signal 125, each of the sparse signals 125 being only within the zero-crossing range.
[0069] In the training step 606 of an ANN, the processor 108 uses the sparse signal output from the preprocessing step 604 to train the ANN 107 and estimate the value of the FA for an audio signal.
[0070] The inference phase 603 begins when the frequency estimation system 101 receives an input, which is one of the short-duration audio samples (e.g., for a duration of several milliseconds) in the sample input step 608.
[0071] In a preprocessing step 610, the sparse signal generator 104 converts the audio sample into a sparse signal (e.g., as described in the sparse signal 125).
[0072] In an ANN inference phase 612, the processor 108 uses a trained ANN to estimate a value of FA in the audio signal.
[0073] Finally, in the FE output step 614, the processor 108 of the frequency estimation system 101 outputs the estimated FA to a user or a processor, for example, to adjust a parameter of the audio processing device 100 in order to reduce errors in the frequency domain.
[0074] The flowchart in Figure 6 is an example of a clear approach, for example, DC filtering can be used in other preprocessing steps.
[0075] Although the algorithms described in the embodiments primarily deal with audio processing for audio engineering suites and / or consumer devices, the methods and systems described herein can also be used in other applications, such as audio quality analysis, filter design, or the use of auto-self-control filters in still image or video processing, as well as encoding and decoding of data compression based on all or part of Fast Fourier Transform analysis.
[0076] The embodiments cited above by way of example and the present invention are not limited to the specific disclosures and descriptions herein. The scope of the invention includes the combinations and various features described herein, as well as variations and adjustments that may be made by those skilled in the art upon reading the foregoing description and content not disclosed in the prior art. The references included in this patent application will be considered an integral part of the patent application, unless any term defined in the references conflicts in some way with the explicit or implicit definitions in the specification of this patent, in which case only the definitions in the specification will be considered. [Simplified Explanation of the Diagram]
[0024] Figure 1 is a block diagram illustrating one embodiment of the present invention, showing a frequency estimation system used to estimate the frequency of one or more added harmonics of a short audio sample from an audio processing device. Figure 2 is a diagram illustrating one embodiment of the present invention, showing a short audio signal 206 having a frequency increase (FA). Figure 3 is a diagram illustrating one embodiment of the present invention and showing an example of an actual short audio signal with additional noise, derived from the short audio signal in Figure 2. Figures 4A and 4B are diagrams illustrating two types of sparse training signals according to an embodiment of the present invention, shown from the short audio signal in Figure 3. Figure 5 is a diagram illustrating one embodiment of the present invention, showing a comparison between the detection frequency of an added harmonic estimated using the system in Figure 1 and the reference truth frequency of the added harmonic. Figure 6 is a flowchart of one embodiment of the present invention, showing a method for estimating the frequency increase (FA) in a sample of a short audio signal 121 using the system in Figure 1.
Claims
1. A frequency estimation system, comprising: a memory configured to store a machine learning model trained to estimate frequency increases in a plurality of sparse input signals, the sparse input signals being individually derived from a plurality of input audio signals, the sparse input signals being used to identify frequency increases in corresponding input audio signals; and a processor configured to: receive a first input audio signal; derive a first sparse input signal from the first input audio signal, the first sparse input signal being used to identify frequency increases in the first input audio signal; estimate a plurality of frequency increases in the first input audio signal by performing one-dimensional estimation using the trained machine learning model applied to the first sparse input signal; and control an audio system generating the first input audio signal using the estimated frequency increases.
2. The system as claimed in claim 1, wherein the processor is configured to derive the first sparse input signal from the first input audio signal by retaining portions of the first input audio signal that are adjacent to zero crossovers and discarding other portions of the first input audio signal.
3. The system as claimed in claim 1, wherein the processor is configured to derive the first sparse input signal from the first input audio signal by retaining portions of the first input audio signal that are near extreme values and discarding other portions of the first input audio signal.
4. The system as claimed in claim 1, wherein the processor is configured to derive the first sparse input signal from the first input audio signal by retaining the steepest portion of the first input audio signal and discarding the other portions of the first input audio signal.
5. The system as claimed in claim 1, wherein the processor is further configured to perform phase correction on the first input audio signal by applying an initialization step, thereby deriving the first sparse input signal from the first input audio signal.
6. The system as claimed in claim 1, wherein the processor is configured to detect the frequencies of one or more higher harmonics in the first input audio signal in order to estimate the increase in value of those frequencies.
7. The system as claimed in claim 1, wherein the processor is configured to receive the first input audio signal to acquire the first input audio signal.
8. The system as described in claim 1, wherein the processor is further configured to filter out a DC component from the first input audio signal.
9. The system as described in claim 1, wherein the processor is further configured to normalize the first input audio signal.
10. The system as described in claim 1, wherein the machine learning model comprises one of a convolutional neural network and a recursive neural network.
11. A frequency estimation system comprising: a memory configured to store a machine learning model; and a processor configured to: acquire a plurality of audio signals, the audio signals being individually labeled according to a plurality of frequency increment values; derive a plurality of training sparse signals from the individual audio signals, each training sparse signal being used to identify one or more frequency increments in a corresponding audio signal; train the machine learning model using the training sparse signals to perform one-dimensional estimation to estimate the plurality of frequency increment values; and control an audio system that generates the audio signals using the estimated frequency increment values.
12. The system as claimed in claim 11, wherein the processor is configured to derive training sparse signals from the audio signals by retaining portions of the audio signals that are adjacent to zero crossovers and discarding other portions of the audio signals.
13. The system as claimed in claim 11, wherein the processor is configured to derive training sparse signals from the audio signals by retaining portions of the audio signals that are near extreme values and discarding other portions of the audio signals.
14. The system as claimed in claim 11, wherein the processor is configured to derive training sparse signals from the audio signals by retaining the steepest portions of the audio signals and discarding the other portions of the audio signals.
15. The system as described in claim 11, wherein the processor is further configured to first apply an initialization step to perform phase correction on the audio signals.
16. The system of claim 11, wherein the processor is configured to receive a plurality of initial audio signals having a first duration, and to segment the initial audio signals into a plurality of segments having a second duration, thereby acquiring a plurality of audio signals, the second duration being shorter than the first duration.
17. The system as claimed in claim 11, wherein the processor is further configured to filter out a DC component from each of the audio signals.
18. The system as described in claim 11, wherein the processor is further configured to normalize each of the audio signals.
19. The system as claimed in claim 11, wherein the machine learning model comprises one of a convolutional neural network and a recursive neural network.
20. The system as claimed in claim 19, wherein the convolutional neural network classifies the frequency increase according to the numerical value of the frequency increase and labels it in the audio signal.
21. A frequency estimation method, comprising: storing a machine learning model in a memory, the machine learning model being trained to estimate multiple frequency increments among multiple sparse input signals, the sparse input signals being derived individually from multiple input audio signals, the sparse input signals being used to identify one or more frequency increments among the corresponding input audio signals; receiving a first input audio signal; deriving a first sparse input signal from the first input audio signal, the first sparse input signal being used to identify frequency increments in the first input audio signal; estimating multiple frequency increments in the first input audio signal by performing a one-dimensional estimation using the trained machine learning model applied to the first sparse input signal; and controlling an audio system that generates the first input audio signal using the estimated frequency increments.
22. A frequency estimation method comprising: storing a machine learning model in a memory; acquiring a plurality of audio signals, the audio signals being individually labeled according to a plurality of frequency increment values; individually deriving a plurality of training sparse signals from the audio signals, each training sparse signal being used to identify one or more frequency increments in a corresponding audio signal; training the machine learning model using the training sparse signals to perform a one-dimensional estimation to estimate the plurality of frequency increment values; and controlling an audio system that generates the audio signals using the estimated frequency increment values.
Citation Information
Patent Citations
Unsupervised abnormal sound detection method and device based on dictionary learning
CN113327632A
Unsupervised machine abnormal sound detection method and device based on single classification algorithm
CN114462475A
Analog hardware realization of neural networks
TW202207093A
Biosignal analysis system
US20240000364A1