Stringed instrument fundamental frequency detection method, apparatus, device and medium

By performing preprocessing and windowing Fourier transform on the string instrument signal, combined with cepstral analysis and linear frequency modulation Z-transform algorithm, the fundamental frequency estimation is optimized, solving the problem of large fundamental frequency detection error in string instrument tuning and achieving high-precision fundamental frequency detection.

CN119889257BActive Publication Date: 2025-11-11ZHUHAI WEIKE TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510070685.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-11-11
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing technologies have significant errors in fundamental frequency detection during string instrument tuning, failing to meet high-precision requirements.

Method used

By preprocessing the string instrument signal, a preset detection strategy is used to distinguish between effective signals and white noise. After windowing, a fast Fourier transform and absolute value operation are performed. Combined with cepstral analysis and linear frequency modulation Z-transform algorithm, the fundamental frequency estimation is optimized using an iterative algorithm and loss function.

Benefits of technology

It reduces detection errors in the low-frequency band, improves the accuracy and robustness of fundamental frequency detection, and ensures the tuning precision of string instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889257B_ABST
    Figure CN119889257B_ABST
Patent Text Reader

Abstract

The application provides a stringed instrument fundamental frequency detection method, device, equipment and medium, the target frame signal obtained by preprocessing is updated to the preset cache; if the target frame signal in the updated cache is detected as a valid signal by using a preset detection strategy, the signal is windowed, and fast Fourier transform and absolute value operation are performed on the windowed signal; if it is determined that the first frequency with the maximum amplitude in the first spectrum information is less than the preset frequency, the first fundamental frequency is obtained by using cepstrum analysis; and the number of harmonics and the harmonic frequency of each harmonic are determined according to the first fundamental frequency and the preset harmonic frequency interval; the first fundamental frequency and the harmonic frequency of each harmonic are calculated by using a preset linear frequency modulation Z transform algorithm to obtain the maximum harmonic frequency of each harmonic; and the optimal fundamental frequency is determined according to the preset iteration algorithm, the preset loss function and the maximum harmonic frequency of each harmonic, thereby reducing the detection error of the low frequency band and improving the accuracy of the fundamental frequency detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio signal processing technology, and in particular to a method, apparatus, device and medium for detecting the fundamental frequency of a stringed instrument. Background Technology

[0002] In many natural signals (such as human voices and instrument sounds), the sound wave signal emitted by the sound source is usually a combination of a series of harmonic frequencies; among these, the lowest frequency is called the fundamental frequency. Fundamental Frequency Detection (F0 detection) is a key technology in audio signal processing, widely used in speech processing, music information retrieval, and instrument tuning. However, the fundamental frequency detection requirements in string instrument tuning differ significantly from those in speech signal tuning, mainly in terms of a wider frequency range, harmonic multiple relationships, high-precision frequency detection requirements, and harmonic shift phenomena. Therefore, fundamental frequency detection in string instrument tuning algorithms needs to meet higher requirements in terms of accuracy, harmonic processing, and frequency range to ensure accuracy and robustness.

[0003] Currently, most mature fundamental frequency detection algorithms are primarily designed for speech signals, such as time-domain-based autocorrelation algorithms and their improved versions (e.g., the YIN algorithm and pYIN algorithm), and frequency-domain-based Fast Fourier Transform (FFT) methods (e.g., peak detection, harmonic summation, harmonic product spectrum, etc.). However, these fundamental frequency detection methods have limitations in the application of string instrument tuning, especially in the low-frequency range where errors are relatively large, failing to meet the high-precision requirements of string instrument tuning. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for detecting the fundamental frequency of a stringed instrument, aiming to solve the problem that the existing technology has a large detection error in the low-frequency band, which cannot meet the high-precision requirements for tuning stringed instruments.

[0005] In a first aspect, embodiments of this application provide a method for detecting the fundamental frequency of a stringed instrument, the method comprising:

[0006] The current frame signal is acquired, and the current frame signal is preprocessed to obtain the target frame signal;

[0007] The target frame signal is updated in the preset buffer to obtain the updated buffer;

[0008] If the target frame signal in the updated cache is detected as a valid signal using a preset detection strategy, then the signal in the updated cache is windowed to obtain a windowed signal.

[0009] The windowed signal is subjected to a Fast Fourier Transform and an absolute value operation to obtain the first spectral information;

[0010] If it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than the preset frequency, then the first spectrum information is processed by cepstral analysis to obtain a coarse estimate of the first fundamental frequency.

[0011] The number of harmonics and the harmonic frequency of each harmonic are determined based on the first fundamental frequency and the preset harmonic frequency range.

[0012] The first fundamental frequency and the harmonic frequencies of each harmonic are analyzed using a preset linear frequency modulation Z-transform algorithm to obtain the maximum frequency value of each harmonic.

[0013] The optimal fundamental frequency is determined based on a preset iterative algorithm, a preset loss function, and the maximum frequency value of each harmonic.

[0014] Secondly, embodiments of this application provide a method for detecting the fundamental frequency of a stringed instrument, the apparatus comprising:

[0015] The current frame signal acquisition unit is used to acquire the current frame signal and preprocess the current frame signal to obtain the target frame signal;

[0016] A cache update unit is used to update the target frame signal into a preset cache to obtain an updated cache.

[0017] A windowing unit is used to perform windowing processing on the signal in the updated cache if the target frame signal in the updated cache is detected as a valid signal using a preset detection strategy, so as to obtain a windowed signal.

[0018] The frequency domain conversion unit is used to perform fast Fourier transform and absolute value operation on the windowed signal to obtain the first spectrum information;

[0019] The fundamental frequency estimation unit is used to process the first spectrum information by cepstral analysis to obtain a coarse estimate of the first fundamental frequency if it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than a preset frequency.

[0020] The harmonic determination unit is used to determine the number of harmonics and the harmonic frequency of each harmonic based on the first fundamental frequency and the preset harmonic frequency range.

[0021] The maximum frequency acquisition unit uses a preset linear frequency modulation Z-transform algorithm to analyze the first fundamental frequency and the harmonic frequencies of each harmonic to obtain the maximum frequency value of each harmonic.

[0022] The fundamental frequency acquisition unit is used to determine the optimal fundamental frequency based on a preset iterative algorithm, a preset loss function, and the maximum frequency value of each harmonic.

[0023] Thirdly, embodiments of this application also provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the fundamental frequency detection method for stringed instruments described in the first aspect.

[0024] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the fundamental frequency detection method for stringed instruments described in the first aspect.

[0025] This application provides a method, apparatus, device, and medium for detecting the fundamental frequency of a stringed instrument. The method involves updating a pre-processed target frame signal to a preset buffer. If the updated target frame signal is detected as valid using a preset detection strategy, it is windowed, and a fast Fourier transform and absolute value operation are performed on the windowed signal. If the first frequency with the maximum amplitude in the obtained first spectral information is determined to be less than a preset frequency, cepstral analysis is used to process it to obtain the first fundamental frequency. The number of harmonics and the harmonic frequency of each harmonic are determined based on the first fundamental frequency and a preset harmonic frequency range. A preset linear frequency modulation Z-transform algorithm is used to calculate the first fundamental frequency and the harmonic frequencies of each harmonic to obtain the maximum harmonic frequency of each harmonic. The optimal fundamental frequency is determined based on a preset iterative algorithm, a preset loss function, and the maximum harmonic frequency of each harmonic, thereby reducing detection errors in the low-frequency band and improving the accuracy of fundamental frequency detection. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A schematic flowchart illustrating the fundamental frequency detection method for stringed instruments provided in this application embodiment;

[0028] Figure 2 A waveform diagram of the input signal in the fundamental frequency detection method for stringed instruments provided in the embodiments of this application;

[0029] Figure 3 A waveform diagram of the output frequency in the fundamental frequency detection method for stringed instruments provided in the embodiments of this application;

[0030] Figure 4 A schematic block diagram of a fundamental frequency detection device for a stringed instrument provided in an embodiment of this application;

[0031] Figure 5 A schematic block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0034] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0035] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0036] Please see Figure 1-3 , Figure 1 A schematic flowchart illustrating the fundamental frequency detection method for stringed instruments provided in this application embodiment; Figure 2 A diagram illustrating the effect of the input signal in the fundamental frequency detection method for stringed instruments provided in this application embodiment; Figure 3 The diagram shows the effect of the output frequency in the fundamental frequency detection method for stringed instruments provided in this application embodiment; the fundamental frequency detection method for stringed instruments is applied to a server.

[0037] like Figure 1-3 As shown, the method includes steps S110-S180.

[0038] S110. Obtain the current frame signal and preprocess the current frame signal to obtain the target frame signal.

[0039] The step of preprocessing the current frame signal to obtain the target frame signal includes:

[0040] The current frame signal is subjected to noise filtering using a first filter to obtain a first signal;

[0041] The first signal is subjected to high-frequency filtering using a low-pass filter to obtain the second signal;

[0042] The second signal is downsampled to obtain the target frame signal.

[0043] In this embodiment, before acquiring the initial signal of each frame, the input signal to be processed, i.e., the audio stream signal (e.g., ...), can be acquired through an external device. Figure 2 As shown, the horizontal axis represents time (t), and the vertical axis represents the magnitude of the audio stream signal (amplitude m). For example, the recorded audio data is acquired using a microphone or directly input; and the audio stream signal to be processed is segmented into frames according to a preset framing algorithm to obtain an initial signal for each frame; the length of each initial signal is typically 20 to 50 milliseconds. After framing the audio stream signal, the current frame signal is obtained from each initial signal, and a first filter is used to filter out noise from the current frame signal to remove white noise, thereby obtaining a first signal; the first filter can be a Wiener filter or an adaptive filter. Furthermore, in this embodiment, the stringed instrument is preferably a piano. Since most pianos have an audio sampling rate of 44.1kHz or 48kHz, and the frequency range of a piano is 27.5Hz to 4186Hz, and according to the Nyquist sampling theorem, the sampling frequency should be greater than or equal to twice the highest frequency in the analog signal spectrum, the sampling rate must be at least 8372Hz. To reduce aliasing after the low-pass filter, the sampling rate is reduced by a factor of 5, decreasing the piano's audio sampling rate from 44.1kHz to 8820Hz and from 48kHz to 9600Hz. Specifically, a low-pass filter with a normalized sampling rate of 0.2 is first used to remove high-frequency noise from the first signal to obtain the second signal; then, the second signal is downsampled to prevent aliasing during downsampling, thereby obtaining a more reliable and efficient target frame signal.

[0044] S120. Update the target frame signal to the preset buffer to obtain the updated buffer.

[0045] In this embodiment, the preprocessed target frame signal is stored in a preset buffer for subsequent signal processing. Specifically, the target frame signal is added to the end of the preset buffer, and data of the same length as the target frame signal is removed from the beginning of the preset buffer. Subsequently, each input target frame signal is stored at the end of the preset buffer, and the same amount of data at the beginning of the preset buffer as the target frame signal is discarded, thereby obtaining an updated buffer to keep the signals in the preset buffer continuous and up-to-date.

[0046] S130. If the target frame signal in the updated cache is detected as a valid signal using a preset detection strategy, then the signal in the updated cache is windowed to obtain a windowed signal.

[0047] In this embodiment, after obtaining the updated cache, a preset detection strategy is used to detect the current target frame signal in the updated cache to distinguish between white noise and valid signals; wherein, the preset detection strategy in this embodiment is preferably a short-time average energy method; specifically, the number of signal frames detected by the preset detection is M and the preset average energy threshold is STE. ev Then, the short-time average energy value (STE) of the i-th frame signal (i.e., the target frame signal) is calculated to obtain the short-time average energy value corresponding to the i-th frame signal; the STE is then compared with the short-time average energy value of the i-th frame signal. i Compared with the short-time average energy value of each frame preceding the i-th frame signal, if STE i-M <STE i-(M-1) <…… <STE i-1 <STE i If the start of a valid signal is detected and a STE occurs, it is considered the beginning of a valid signal. i-M ≈STE i-(M-1) ≈……≈STE i-1 ≈STE i ≈STE ev If the signal in frame i is found to be valid, it is determined to be the end of a valid signal segment, and the signal in frame i is declared to be valid. At this point, the updated signal in the buffer is windowed, i.e., a window function is added to obtain the windowed signal; the window function can be a Hamming window, Hanning window, or Blackman window, etc., thereby reducing spectral leakage and improving the accuracy of spectral estimation through windowing.

[0048] In one embodiment, after obtaining the updated cache, the process further includes:

[0049] If the target frame signal in the updated cache is detected as white noise using a preset signal detection strategy, then the next frame signal of the current frame signal is obtained, the current frame signal is updated based on the next frame signal, and the process of preprocessing the current frame signal to obtain the target frame signal is returned.

[0050] In this embodiment, the preset detection strategy is preferably a short-time average energy method; specifically, the number of signal frames detected by the preset method is M, and the preset average energy threshold is STE. ev Specifically, a preset detection strategy is used to detect the i-th frame signal (i.e., the target frame signal) in the updated cache, and the short-time average energy value (STE) of the i-th frame signal is compared. iCompared with the short-time average energy value of each frame preceding the i-th frame signal, if STE i-M ≈STE i-(M-1) ≈……≈STE i-1 ≈STE i If the signal is white noise, it is determined to be white noise. At this time, the next frame signal of the current frame signal is obtained to continuously update the current frame signal, and the step of preprocessing the current frame signal to obtain the target frame signal is repeated. This allows for timely identification and elimination of white noise, avoiding wasting resources on useless signals.

[0051] S140. Perform a fast Fourier transform and absolute value operation on the windowed signal to obtain the first spectrum information.

[0052] In this embodiment, the windowed signal is first subjected to a Fast Fourier Transform (FFT) to convert it from the time domain to the frequency domain to obtain the Fourier Transform result. The absolute value of the Fourier Transform result is then taken to obtain the first spectral information of the windowed signal. The first spectral information includes the amplitude information of the spectrum. This facilitates the subsequent search for the spectral peak from the amplitude information in the first spectral information, thereby obtaining the first frequency corresponding to the spectral peak.

[0053] In one embodiment, after step S140, the following steps are included:

[0054] If it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is greater than the preset frequency, then the first frequency is taken as the optimal base frequency.

[0055] In this embodiment, in order to coarsely estimate the signal frequency, the windowed signal is subjected to a Fast Fourier Transform and absolute value operation to obtain the first spectrum information. Then, the first spectrum information is traversed to determine the maximum amplitude (i.e., the maximum peak value) in the first spectrum information, thereby determining the strongest frequency component. The first frequency corresponding to the maximum amplitude obtained by traversal is compared with a preset frequency. When it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is greater than the preset frequency, it indicates that there are no harmonics in the first spectrum information and only one peak value. The frequency corresponding to the peak value is the fundamental frequency of the signal. Then, the coarse frequency estimation detection ends, and the first frequency is taken as the optimal fundamental frequency.

[0056] S150. If it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than the preset frequency, then the first spectrum information is processed by cepstral analysis to obtain a coarse estimate of the first fundamental frequency.

[0057] In this embodiment, after obtaining the first spectrum information, the first spectrum information is traversed to determine the maximum amplitude (i.e., the maximum peak value) in the first spectrum information, thereby determining the strongest frequency component; and the first frequency corresponding to the maximum amplitude obtained by traversal is compared with a preset frequency. When it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than the preset frequency, it indicates that there is an overtone harmonic of the first frequency corresponding to the maximum amplitude in the first spectrum information. This means that in addition to the main frequency component corresponding to the first frequency, there are also components of integer multiples of its frequency. At this time, the first spectrum information is processed by cepstral analysis to identify periodic components, thereby obtaining a rough estimate of the first fundamental frequency.

[0058] In one embodiment, processing the first spectral information using cepstral analysis to obtain a coarse estimate of the first fundamental frequency includes:

[0059] Perform a Fast Fourier Transform operation on the first spectral information to obtain the transformation result;

[0060] The absolute value of the transformation result is then calculated to obtain the second spectral information;

[0061] The second spectrum information is traversed to determine the maximum amplitude, and the frequency corresponding to the maximum amplitude is taken as the first fundamental frequency.

[0062] In this embodiment, the process of processing the first spectral information using cepstral analysis to obtain a coarse estimate of the first fundamental frequency firstly involves performing a fast Fourier transform on the first spectral information to obtain a transform result including cepstral information; then, performing an absolute value operation on the transform result to calculate the second spectral information; furthermore, the second spectral information is traversed to find the maximum amplitude corresponding to the second spectral information, and the second frequency corresponding to the maximum amplitude is taken as the coarse estimate of the first fundamental frequency.

[0063] S160. Determine the number of harmonics and the harmonic frequency of each harmonic based on the first fundamental frequency and the preset harmonic frequency range.

[0064] In this embodiment, in order to facilitate a more accurate calculation of the optimal fundamental frequency, the number of harmonics is determined by the first fundamental frequency and the preset harmonic frequency range. Herein, a harmonic refers to a sine wave whose frequency is an integer multiple of the fundamental frequency. Thus, the number of existing harmonics can be calculated by using the known first fundamental frequency and the preset harmonic frequency range, and the approximate harmonic frequency of each harmonic can be obtained.

[0065] For example, if the preset harmonic frequency range is 0-4800Hz and the first fundamental frequency f1 = 1000Hz, then the other consecutive maximum peaks in the preset harmonic frequency range are f2 = 2000Hz, f3 = 3000Hz, and f4 = 4000Hz. The number of harmonics is n = 4, and the harmonic frequencies corresponding to each harmonic are f1 = 1000Hz, f2 = 2000Hz, f3 = 3000Hz, and f4 = 4000Hz, respectively.

[0066] S170. Analyze the first fundamental frequency and the harmonic frequencies of each harmonic using a preset linear frequency modulation Z-transform algorithm to obtain the maximum frequency value of each harmonic.

[0067] In this embodiment, in order to determine more accurate harmonic frequencies, after determining the number of harmonics and the harmonic frequency corresponding to each harmonic, the first fundamental frequency and the harmonic frequency corresponding to each harmonic are subjected to spectral analysis using the linear frequency modulated Z-transform (Cirp-Z, CZT) algorithm to determine the maximum frequency value corresponding to each harmonic position, that is, the actual frequency of each harmonic.

[0068] S180. Determine the optimal fundamental frequency based on the preset iterative algorithm, the preset loss function, and the maximum frequency value of each harmonic.

[0069] In this embodiment, after obtaining the maximum frequency value of each harmonic, the optimal fundamental frequency is determined specifically through a preset iterative algorithm, a preset loss function, and the maximum frequency value of each harmonic. First, it should be clarified that for stringed instruments, such as pianos, guitars, basses, and guzheng, the sound-producing unit is the vibration of a taut, rigid string under external force. Considering the bending stiffness and internal damping of the string, a string of length L can be represented as:

[0070]

[0071] Where μ represents the linear density of the string, R(ω) represents the damping coefficient, T represents the string tension, E represents Young's modulus, S represents the cross-sectional area, κ represents the radius of gyration, and d x (x,t) represents the external excitation force applied to the string, which is a function of space and time.

[0072] When d x When (x,t) is an impact signal of Fδ(t,t0)δ(x,x0) such as plucking, strumming, or striking, the characteristics of the audio signal generated by the vibration of the string over time are expressed as follows:

[0073]

[0074] Where t represents the vibration time of the string, and R k f represents the internal damping of the string.k y represents the k-th order vibration frequency. δ,k (t) represents the audio signal generated by the vibration of the string.

[0075] Meanwhile, the frequency of the nth harmonic can be expressed as:

[0076]

[0077] Where n is a positive integer, and the value of n is [1, ∞];

[0078] because and but

[0079] Where B represents the correction factor, L represents the string length, f0 represents the fundamental frequency of the string, and f n This represents the frequency of the nth harmonic of the string.

[0080] Furthermore, assume that the nth harmonic observed from the spectrum is f n ′ Then the error ∈ n The actual measured value f n ′ Compared with the theoretical value f n Differences between them:

[0081]

[0082] Furthermore, assuming the error ∈ n It follows a normal distribution with a mean of 0 and a variance of σ. 2 ,Right now

[0083] ∈ n ~N(0,σ 2 );

[0084] This means that the harmonic frequency f we observed n It also follows a normal distribution with a mean of 1 / 2. The variance is σ 2 .

[0085] Based on the assumption of error, the likelihood function L(f0,B|f1) ′ f2 ′ ,…f k ′ ), where k is the order of the harmonic:

[0086]

[0087] Taking the logarithm of the likelihood function, we obtain the log-likelihood function l(f0,B):

[0088]

[0089] Because the constant term and σ 2 Since L(f0,B) is independent of both f0 and B, maximizing L(f0,B) is equivalent to minimizing the error term. Therefore, the preset loss function can be expressed as:

[0090]

[0091] Therefore, by using the maximum frequency value of each harmonic and a preset iterative algorithm (e.g., gradient descent or Newton's method), the estimated value of the fundamental frequency is continuously adjusted to minimize the preset loss function, thereby achieving an accurate estimation of the fundamental frequency. This makes the estimated values ​​of f0 and B maximize the log-likelihood function, thus significantly improving the accuracy and robustness of the fundamental frequency estimation.

[0092] In one embodiment, step S180 includes:

[0093] Obtain an initial parameter set; wherein the initial parameter set includes a preset cutoff threshold, an initial function value, an initial fundamental frequency parameter, and an initial correction factor;

[0094] The preset loss function is calculated based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value;

[0095] If the absolute value of the difference between the current function value and the initial function value is less than the preset cutoff threshold, then the initial function value is updated based on the current function value, the initial fundamental frequency parameter and the initial correction factor are updated to obtain the updated correction factor and the updated fundamental frequency parameter. The updated fundamental frequency parameter is used as the initial fundamental frequency parameter, the updated correction factor is used as the initial correction factor, and the process returns to the step of calculating the preset loss function based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value.

[0096] The iteration continues until a preset number of iterations is reached, or the absolute value of the difference between the current function value and the initial function value is greater than the preset cutoff threshold, at which point the iteration stops and the optimal fundamental frequency is output.

[0097] In this embodiment, the preset iterative algorithm is a gradient descent method; the step of determining the optimal fundamental frequency based on the preset iterative algorithm, the preset loss function, and the maximum frequency value of each harmonic includes: firstly, obtaining an initial parameter set; wherein, the initial parameter set includes a preset cutoff threshold D, an initial function value loss, and an initial function value loss. pre The initial fundamental frequency parameter f0 and the initial correction factor B, including the iteration factor lr; the preset cutoff threshold D, and the initial function value loss. preThe values ​​of the initial fundamental frequency parameter f0, the initial correction factor B, and the iteration factor lr can be set according to actual needs. In this embodiment, the stringed instrument is a piano, and the range of the initial correction factor B is typically 2*10. -4 ~8*10 -4 Between these points, the initial fundamental frequency parameter f0 can be f0 ′ The initial correction factor B = 1 * 10 -4 Initial function value loss pre =0, iteration factor lr=0.1, preset cutoff threshold D=1*10 -2 Furthermore, in the first iteration, the maximum frequency value f0 of each harmonic was set. ′ f1 ′ ...f n ′ The initial fundamental frequency parameter f0 and the initial correction factor B are calculated using the preset loss function to obtain the current function value loss; and the current function value loss is then compared with the preset function threshold loss. pre The absolute value of the difference between them δ loss The current function value loss is compared with the preset function threshold loss. pre The absolute value of the difference between them δ loss The current function value (loss) is compared with a preset cutoff threshold; if the current function value (loss) is less than or equal to the preset function threshold (loss)... pre The absolute value of the difference between them δ loss If the value is less than the preset cutoff threshold D, it indicates that the optimal solution has been reached. At this point, the calculation ends, and the initial fundamental frequency parameter and initial correction factor corresponding to the current function value are output. The initial fundamental frequency parameter is then output as the optimal fundamental frequency. If the current function value is less than the preset function threshold loss... pre The absolute value of the difference between them δ loss If the value is greater than the preset cutoff threshold D, it indicates that the optimal solution has not been reached. The iteration continues, and the initial function value is updated based on the current function value. The initial fundamental frequency parameter and the initial correction factor are also updated to obtain the updated correction factor and the updated fundamental frequency parameter. The updated fundamental frequency parameter is used as the initial fundamental frequency parameter, and the updated correction factor is used as the initial correction factor. The process then repeats the step of calculating the preset loss function based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value. This continues until a preset number of iterations is reached, or the absolute value of the difference between the current function value and the initial function value is greater than the preset cutoff threshold. At this point, the iteration ends, and the optimal fundamental frequency (e.g., ...) is output. Figure 3 As shown, the horizontal axis represents time (t) and the vertical axis represents the fundamental frequency (Hz). This allows for a more accurate description of the string's vibration characteristics through the optimal fundamental frequency, ensuring the tuning accuracy of stringed instruments.

[0098] The initial fundamental frequency parameter f0 and the initial correction factor B are updated using the following formulas to obtain the updated correction factor and the updated fundamental frequency parameter:

[0099]

[0100] Where f0 represents the initial fundamental frequency parameter, B represents the initial correction factor, lr represents the iteration factor, loss represents the current function value, and n represents the number of harmonics, with n taking the value [1, k].

[0101] Furthermore, the step of calculating the current function value by using the preset loss function based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic is performed by calculating the current function value using the following formula:

[0102]

[0103] Where, f′ n Let f0 represent the maximum harmonic frequency of the nth harmonic, B represent the correction factor, f0 represent the initial fundamental frequency, and n represent the number of harmonics, with n taking values ​​from [1, k].

[0104] Therefore, this embodiment automatically calculates the optimal fundamental frequency parameter that minimizes the preset loss function through a preset iterative algorithm, thereby providing a more accurate and reliable optimal fundamental frequency.

[0105] As can be seen from the above technical solution, the target frame signal obtained through preprocessing is updated to a preset buffer. If the target frame signal in the updated buffer is detected as a valid signal using a preset detection strategy, it is windowed, and a fast Fourier transform and absolute value operation are performed on the windowed signal. If it is determined that the first frequency with the maximum amplitude in the first spectrum information is less than a preset frequency, it is processed using cepstral analysis to obtain the first fundamental frequency. The number of harmonics and the harmonic frequency of each harmonic are determined based on the first fundamental frequency and the preset harmonic frequency range. The first fundamental frequency and the harmonic frequency of each harmonic are calculated using a preset linear frequency modulation Z-transform algorithm to obtain the maximum harmonic frequency of each harmonic. The optimal fundamental frequency is determined based on a preset iterative algorithm, a preset loss function, and the maximum harmonic frequency of each harmonic, thereby reducing the detection error in the low-frequency band and improving the accuracy of fundamental frequency detection.

[0106] This application also provides a fundamental frequency detection device for a stringed instrument, which is used to perform any embodiment of the aforementioned fundamental frequency detection method for stringed instruments. Specifically, please refer to... Figure 2 , Figure 2 This is a schematic block diagram of the fundamental frequency detection device for stringed instruments provided in the embodiments of this application.

[0107] Among them, such as Figure 4 As shown, the fundamental frequency detection device 100 for the stringed instrument includes a current frame signal acquisition unit 110, a buffer update unit 120, a windowing unit 130, a frequency domain conversion unit 140, a fundamental frequency estimation unit 150, a harmonic determination unit 160, a maximum frequency acquisition unit 170, and a fundamental frequency acquisition unit 180.

[0108] The current frame signal acquisition unit 110 is used to acquire the current frame signal and preprocess the current frame signal to obtain the target frame signal;

[0109] The cache update unit 120 is used to update the target frame signal into a preset cache to obtain an updated cache.

[0110] The windowing unit 130 is used to perform windowing processing on the signal in the updated cache if the target frame signal in the updated cache is detected as a valid signal using a preset detection strategy, so as to obtain the windowed signal.

[0111] The frequency domain conversion unit 140 is used to perform fast Fourier transform and absolute value operation on the windowed signal to obtain the first spectrum information;

[0112] The fundamental frequency estimation unit 150 is used to process the first spectrum information by cepstral analysis if it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than a preset frequency, so as to obtain a coarse estimate of the first fundamental frequency.

[0113] The harmonic determination unit 160 is used to determine the number of harmonics and the harmonic frequency of each harmonic based on the first fundamental frequency and the preset harmonic frequency range.

[0114] The maximum frequency acquisition unit 170 uses a preset linear frequency modulation Z-transform algorithm to analyze the first fundamental frequency and the harmonic frequencies of each harmonic to obtain the maximum frequency value of each harmonic.

[0115] The fundamental frequency acquisition unit 180 is used to determine the optimal fundamental frequency based on a preset iterative algorithm, a preset loss function, and the maximum frequency value of each harmonic.

[0116] In some embodiments, the baseband acquisition unit 180 is specifically used for:

[0117] Obtain an initial parameter set; wherein the initial parameter set includes a preset cutoff threshold, an initial function value, an initial fundamental frequency parameter, and an initial correction factor;

[0118] The preset loss function is calculated based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value;

[0119] If the absolute value of the difference between the current function value and the initial function value is less than the preset cutoff threshold, then the initial function value is updated based on the current function value, the initial fundamental frequency parameter and the initial correction factor are updated to obtain the updated correction factor and the updated fundamental frequency parameter. The updated fundamental frequency parameter is used as the initial fundamental frequency parameter, the updated correction factor is used as the initial correction factor, and the process returns to the step of calculating the preset loss function based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value.

[0120] The iteration continues until a preset number of iterations is reached, or the absolute value of the difference between the current function value and the initial function value is greater than the preset cutoff threshold, at which point the iteration stops and the optimal fundamental frequency is output.

[0121] In some embodiments, the current function value is obtained in the step of obtaining the current function value using the following formula:

[0122]

[0123] Where n represents the number of harmonics, f′ n Let f0 represent the maximum harmonic frequency of the nth harmonic, B represent the correction factor, f0 represent the initial fundamental frequency, and n takes the value [1, k].

[0124] In some embodiments, when the fundamental frequency estimation unit 150 performs the step of processing the first spectral information using cepstral analysis to obtain a coarse estimate of the first fundamental frequency, it is specifically used for:

[0125] Perform a Fast Fourier Transform operation on the first spectral information to obtain the transformation result;

[0126] The absolute value of the transformation result is then calculated to obtain the second spectral information;

[0127] The second spectrum information is traversed to determine the maximum amplitude, and the frequency corresponding to the maximum amplitude is taken as the first fundamental frequency.

[0128] In some embodiments, after the frequency domain conversion unit 140, it is further configured to:

[0129] If it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is greater than the preset frequency, then the first frequency is taken as the optimal base frequency.

[0130] In some embodiments, when the current frame signal acquisition unit 110 performs the step of preprocessing the current frame signal to obtain the target frame signal, it is specifically used for:

[0131] The current frame signal is subjected to noise filtering using a first filter to obtain a first signal;

[0132] The first signal is subjected to high-frequency filtering using a low-pass filter to obtain the second signal;

[0133] The second signal is downsampled to obtain the target frame signal.

[0134] In some embodiments, after the cache update unit 120, it is further configured to:

[0135] If the target frame signal in the updated cache is detected as white noise using a preset signal detection strategy, then the next frame signal of the current frame signal is obtained, the current frame signal is updated based on the next frame signal, and the process of preprocessing the current frame signal to obtain the target frame signal is returned.

[0136] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the aforementioned string instrument fundamental frequency detection device and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.

[0137] The fundamental frequency detection device for the aforementioned stringed instrument can be implemented as a computer program, which can, for example... Figure 5 It runs on the computer device shown.

[0138] Please see Figure 5 , Figure 5 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 is a server, or it can be a server cluster. The server can be a standalone server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0139] See Figure 5 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a device bus 501, wherein the memory may include a storage medium 503 and internal memory 504.

[0140] The storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, it enables the processor 502 to execute a fundamental frequency detection method for stringed instruments.

[0141] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0142] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute the fundamental frequency detection method for stringed instruments.

[0143] This network interface 505 is used for network communication, such as providing data transmission. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0144] The processor 502 is used to run a computer program 5032 stored in a memory to implement the fundamental frequency detection method for stringed instruments disclosed in the embodiments of this application.

[0145] Those skilled in the art will understand that Figure 5 The embodiments of the computer device shown do not constitute a limitation on the specific configuration of the computer device. In other embodiments, the computer device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may include only memory and a processor. In such embodiments, the structure and function of the memory and processor are different from those shown. Figure 5 The embodiments shown are consistent and will not be repeated here.

[0146] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0147] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when executed by a processor, the computer program implements the fundamental frequency detection method for stringed instruments disclosed in the embodiments of this application.

[0148] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.

[0150] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0151] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0152] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a backend server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.

[0153] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting the fundamental frequency of a stringed instrument, characterized in that, The method includes: The current frame signal is acquired, and the current frame signal is preprocessed to obtain the target frame signal; The target frame signal is updated in the preset buffer to obtain the updated buffer; If the target frame signal in the updated cache is detected as a valid signal using a preset detection strategy, then the signal in the updated cache is windowed to obtain a windowed signal. The windowed signal is subjected to a Fast Fourier Transform and an absolute value operation to obtain the first spectral information; If it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than the preset frequency, then the first spectrum information is processed by cepstral analysis to obtain a coarse estimate of the first fundamental frequency. The number of harmonics and the harmonic frequency of each harmonic are determined based on the first fundamental frequency and the preset harmonic frequency range. The first fundamental frequency and the harmonic frequencies of each harmonic are analyzed using a preset linear frequency modulation Z-transform algorithm to obtain the maximum frequency value of each harmonic. The optimal fundamental frequency is determined based on a preset iterative algorithm, a preset loss function, and the maximum frequency value of each harmonic. The step of determining the optimal fundamental frequency based on a preset iterative algorithm, a preset loss function, and the maximum harmonic frequency at each harmonic includes: Obtain an initial parameter set; wherein the initial parameter set includes a preset cutoff threshold, an initial function value, an initial fundamental frequency parameter, and an initial correction factor; The preset loss function is calculated based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value; If the absolute value of the difference between the current function value and the initial function value is less than the preset cutoff threshold, then the initial function value is updated based on the current function value, the initial fundamental frequency parameter and the initial correction factor are updated to obtain the updated correction factor and the updated fundamental frequency parameter. The updated fundamental frequency parameter is used as the initial fundamental frequency parameter, the updated correction factor is used as the initial correction factor, and the process returns to the step of calculating the preset loss function based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value. The iteration continues until a preset number of iterations is reached, or the absolute value of the difference between the current function value and the initial function value is greater than the preset cutoff threshold, at which point the iteration stops and the optimal fundamental frequency is output. In the step of obtaining the current function value, the current function value is obtained using the following formula: ; Where n represents the number of harmonics, B represents the maximum harmonic frequency of the nth harmonic, and B represents the correction factor. This represents the initial fundamental frequency, where n takes the value [1, k], and k is the order of the harmonic.

2. The fundamental frequency detection method for stringed instruments according to claim 1, characterized in that, The step of processing the first spectral information using cepstral analysis to obtain a coarse estimate of the first fundamental frequency includes: Perform a Fast Fourier Transform operation on the first spectral information to obtain the transformation result; The absolute value of the transformation result is then calculated to obtain the second spectral information; The second spectrum information is traversed to determine the maximum amplitude, and the frequency corresponding to the maximum amplitude is taken as the first fundamental frequency.

3. The fundamental frequency detection method for stringed instruments according to claim 1, characterized in that, After obtaining the first spectrum information, the process further includes: If it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is greater than the preset frequency, then the first frequency is taken as the optimal base frequency.

4. The fundamental frequency detection method for stringed instruments according to claim 1, characterized in that, The step of preprocessing the current frame signal to obtain the target frame signal includes: The current frame signal is subjected to noise filtering using a first filter to obtain a first signal; The first signal is subjected to high-frequency filtering using a low-pass filter to obtain the second signal; The second signal is downsampled to obtain the target frame signal.

5. The fundamental frequency detection method for stringed instruments according to claim 1, characterized in that, After obtaining the updated cache, the process also includes: If the target frame signal in the updated cache is detected as white noise using a preset signal detection strategy, then the next frame signal of the current frame signal is obtained, the current frame signal is updated based on the next frame signal, and the process of preprocessing the current frame signal to obtain the target frame signal is returned.

6. A fundamental frequency detection device for a stringed instrument, characterized in that, The device includes: The current frame signal acquisition unit is used to acquire the current frame signal and preprocess the current frame signal to obtain the target frame signal; A cache update unit is used to update the target frame signal into a preset cache to obtain an updated cache. A windowing unit is used to perform windowing processing on the signal in the updated cache if the target frame signal in the updated cache is detected as a valid signal using a preset detection strategy, so as to obtain a windowed signal. The frequency domain conversion unit is used to perform fast Fourier transform and absolute value operation on the windowed signal to obtain the first spectrum information; The fundamental frequency estimation unit is used to process the first spectrum information by cepstral analysis to obtain a coarse estimate of the first fundamental frequency if it is determined that the first frequency corresponding to the maximum amplitude in the first spectrum information is less than a preset frequency. The harmonic determination unit is used to determine the number of harmonics and the harmonic frequency of each harmonic based on the first fundamental frequency and the preset harmonic frequency range. The maximum frequency acquisition unit uses a preset linear frequency modulation Z-transform algorithm to analyze the first fundamental frequency and the harmonic frequencies of each harmonic to obtain the maximum frequency value of each harmonic. The fundamental frequency acquisition unit is used to determine the optimal fundamental frequency based on a preset iterative algorithm, a preset loss function, and the maximum frequency value of each harmonic. The fundamental frequency acquisition unit is specifically used for: Obtain an initial parameter set; wherein the initial parameter set includes a preset cutoff threshold, an initial function value, an initial fundamental frequency parameter, and an initial correction factor; The preset loss function is calculated based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value; If the absolute value of the difference between the current function value and the initial function value is less than the preset cutoff threshold, then the initial function value is updated based on the current function value, the initial fundamental frequency parameter and the initial correction factor are updated to obtain the updated correction factor and the updated fundamental frequency parameter. The updated fundamental frequency parameter is used as the initial fundamental frequency parameter, the updated correction factor is used as the initial correction factor, and the process returns to the step of calculating the preset loss function based on the initial fundamental frequency parameter, the initial correction factor, and the maximum harmonic frequency at each harmonic to obtain the current function value. The iteration continues until a preset number of iterations is reached, or the absolute value of the difference between the current function value and the initial function value is greater than the preset cutoff threshold, at which point the iteration stops and the optimal fundamental frequency is output. In the step of obtaining the current function value, the current function value is obtained using the following formula: ; Where n represents the number of harmonics, B represents the maximum harmonic frequency of the nth harmonic, and B represents the correction factor. This represents the initial fundamental frequency, where n takes the value [1, k], and k is the order of the harmonic.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the fundamental frequency detection method for stringed instruments as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the fundamental frequency detection method for a stringed instrument as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • High-precision and high-stability string instrument fundamental frequency detection method

    CN111613241A

  • Musical instrument timbre conversion method based on musical sound signal harmonic energy

    CN118430485A