Fundamental frequency detection method, device, equipment and storage medium
By performing time domain preprocessing and frequency domain analysis on multi-frame splicing signals, combined with harmonic matching technology, the problem of difficulty in balancing stability and accuracy in real-time audio processing systems is solved, and high-precision fundamental frequency detection is achieved.
Patent Information
- Application Number
- CN202510651300.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-29
AI Technical Summary
Existing fundamental frequency detection algorithms are difficult to balance stability and accuracy in real-time audio processing systems. Frequency domain algorithms have limitations in frequency resolution. Time domain algorithms have poor stability when signal changes and are susceptible to noise and signal fluctuations.
By performing time-domain preprocessing (filtering, clipping and windowing processing) on multi-frame splicing signals, combining fast Fourier transform and peak detection, harmonic matching is performed based on the preset frequency deviation range, the support score of candidate fundamental frequency is determined, and fundamental frequency detection is realized.
While ensuring the real-time and efficient audio signal processing, high-precision fundamental frequency detection is achieved, balancing the requirements of stability and accuracy.
Smart Images

Figure CN120564754A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to a fundamental frequency detection method, apparatus, device and storage medium. Background Art
[0002] Fundamental frequency detection technology plays a vital role in numerous fields and is widely used across various technological sectors. For example, extracting fundamental frequency information helps identify speech characteristics such as pitch, rhythm, and emotion, enables data demodulation and signal recovery, and analyzes signals such as heartbeats and muscle tremors. Furthermore, in music theory, the combination of different fundamental frequencies determines the level of pitch and the composition of harmonics, directly impacting the expressiveness and aesthetic quality of music.
[0003] Currently, fundamental frequency detection algorithms are primarily categorized into two main categories: frequency domain and time domain. Frequency domain algorithms convert signals into the frequency domain for analysis using techniques such as Fourier transforms. This effectively filters out noise interference and offers good stability. However, frequency domain algorithms are limited in frequency resolution, resulting in poor accuracy. Time domain algorithms directly analyze signal characteristics, such as zero-crossing rate, along the time axis. While highly accurate, they are less stable in the face of signal fluctuations and are susceptible to noise and signal fluctuations. Therefore, balancing the stability and accuracy of fundamental frequency detection algorithms for real-time audio processing systems remains an unresolved issue.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a fundamental frequency detection method, device, equipment and storage medium, aiming to solve the technical problem of how to balance the stability and accuracy of the fundamental frequency detection algorithm applied to the real-time audio processing system.
[0006] To achieve the above objectives, the present application proposes a fundamental frequency detection method, which includes:
[0007] Performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, wherein the time domain preprocessing includes filtering processing, clipping processing, and windowing processing;
[0008] Performing a fast Fourier transform on the first time-domain preprocessed data to obtain preprocessed frequency-domain signal data;
[0009] Performing peak detection on the preprocessed frequency domain signal data to obtain a candidate fundamental frequency;
[0010] Harmonic matching is performed on the candidate fundamental frequencies based on a preset frequency deviation range to obtain a support score for the candidate fundamental frequencies, and a fundamental frequency detection result is determined according to the support score.
[0011] In one embodiment, the step of performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data includes:
[0012] Filtering the multi-frame splicing signal based on the filter coefficient to obtain filtered data;
[0013] performing clipping processing on the filtered data based on a clipping threshold by using a level clipping function to obtain clipping data;
[0014] The clipping data is subjected to windowing processing by means of an analysis window function to obtain first time-domain preprocessed data.
[0015] In one embodiment, the step of performing peak detection on the preprocessed frequency domain signal data to obtain a candidate fundamental frequency includes:
[0016] Traversing the pre-processed frequency domain signal data through a preset sliding window, obtaining a frequency point global subscript, a frequency point window local subscript, and a frequency point amplitude of the preset pre-processed frequency domain signal data;
[0017] Determining candidate frequency points from the preset pre-processed frequency domain signal data based on the frequency point amplitudes;
[0018] When it is detected based on the local subscript of the frequency window that the candidate frequency point is located at the center of the preset sliding window, obtaining the global subscript and frequency amplitude of the candidate frequency point to obtain a peak detection result;
[0019] A candidate fundamental frequency is determined based on the peak detection result.
[0020] In one embodiment, the step of performing harmonic matching on the candidate fundamental frequency based on the preset frequency deviation range to obtain a support score of the candidate fundamental frequency includes:
[0021] generating a theoretical harmonic sequence based on the candidate fundamental frequency;
[0022] Performing harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range and the theoretical harmonic sequence to obtain the number of matching harmonics and the matching harmonic amplitude;
[0023] A support score for the candidate fundamental frequency is determined based on the number of matching harmonics and the amplitude of the matching harmonics.
[0024] In one embodiment, the step of determining the fundamental frequency detection result according to the support score includes:
[0025] When the support score of the candidate fundamental frequency is greater than or equal to a preset harmonic support threshold, and the support score difference of the candidate fundamental frequency is greater than a preset difference, determining a target fundamental frequency from the candidate fundamental frequencies based on the support score, and obtaining a fundamental frequency detection result;
[0026] When the support score of the candidate fundamental frequencies is greater than or equal to a preset harmonic support threshold, and the difference in the support scores of the candidate fundamental frequencies is less than or equal to the preset difference, determining the target fundamental frequency from the candidate fundamental frequencies based on the frequency sorting result of the candidate fundamental frequencies to obtain a fundamental frequency detection result;
[0027] When the support score of the candidate fundamental frequency is less than a preset harmonic support threshold, it is determined that the fundamental frequency detection result is that the fundamental frequency does not exist.
[0028] In one embodiment, before the step of performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, the method further includes:
[0029] Get the frame length of the transmitted and received signals and the number of frames of the buffered spliced signals;
[0030] Determine the buffer length according to the frame length of the transmitted and received signals and the number of frames of the buffer spliced signal;
[0031] When the current frame data of the original audio signal arrives at the buffer, obtaining the buffer index of the current frame data;
[0032] The current frame data is frame-spliced and spliced based on the buffer length and the buffer index through a piecewise function to obtain a multi-frame splicing signal.
[0033] In one embodiment, after the steps of performing harmonic matching on the candidate fundamental frequencies based on the preset frequency deviation range to obtain a support score for the candidate fundamental frequencies, and determining a fundamental frequency detection result based on the support score, the method further includes:
[0034] Performing sound effect processing on the pre-processed frequency domain signal data based on the fundamental frequency detection result to obtain target frequency domain signal data;
[0035] Performing an inverse fast Fourier transform on the target frequency domain signal data to obtain second time domain signal data;
[0036] Performing windowing processing on the second time domain signal data by using a synthetic window function to obtain second time domain preprocessed data;
[0037] Acquire current frame synthesis window signal data and adjacent frame synthesis window signal data from the second time domain preprocessed data based on the frame length of the received and transmitted signal and the buffer index of the current frame data;
[0038] The current frame synthesis window signal data and the adjacent frame synthesis window signal data are superimposed to obtain a target audio signal.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a fundamental frequency detection device, which includes:
[0040] A data processing module, configured to perform time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, wherein the time domain preprocessing includes filtering, clipping, and windowing;
[0041] a data conversion module, configured to perform a fast Fourier transform on the first time-domain preprocessed data to obtain preprocessed frequency-domain signal data;
[0042] The data processing module is further configured to perform peak detection on the pre-processed frequency domain signal data to obtain a candidate fundamental frequency;
[0043] The fundamental frequency detection module is used to perform harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range, obtain a support score of the candidate fundamental frequency, and determine a fundamental frequency detection result according to the support score.
[0044] In addition, to achieve the above-mentioned purpose, the present application also proposes a baseband detection device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the baseband detection method described above.
[0045] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the steps of the baseband detection method described above are implemented.
[0046] In addition, to achieve the above-mentioned object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the baseband detection method described above are implemented.
[0047] One or more technical solutions proposed in this application have at least the following technical effects:
[0048] By performing time domain preprocessing (filtering, clipping and windowing) on the multi-frame spliced signal, noise and interference in the signal are removed, the signal quality is improved, and the accuracy of signal recognition is enhanced. The time domain signal is converted into a frequency domain signal using fast Fourier transform. The main frequency components in the spectrum are identified through peak detection, providing a preliminary estimate for the candidate fundamental frequency. The support score of the candidate fundamental frequency is obtained through harmonic matching, and the stability and matching degree of the candidate fundamental frequency under multiple harmonic relationships are determined. The fundamental frequency detection result is determined based on the support score. While ensuring the real-time and high efficiency of audio signal processing, high-precision fundamental frequency detection is achieved, balancing the requirements of stability and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0051] Figure 1 A flowchart of the first embodiment of the baseband detection method of the present application is provided;
[0052] Figure 2 A flowchart of the second embodiment of the baseband detection method of the present application is provided;
[0053] Figure 3 This is a schematic diagram of the module structure of the fundamental frequency detection device according to an embodiment of the present application;
[0054] Figure 4 Schematic diagram of the device structure of the hardware operating environment involved in the baseband detection method in the embodiment of the present application.
[0055] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0056] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0057] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0058] The main solution of the embodiment of the present application is: perform time domain preprocessing on the multi-frame splicing signal to obtain first time domain preprocessing data, and the time domain preprocessing includes filtering processing, clipping processing and windowing processing; perform fast Fourier transform on the first time domain preprocessing data to obtain preprocessed frequency domain signal data; perform peak detection on the preprocessed frequency domain signal data to obtain a candidate fundamental frequency; perform harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range to obtain a support score of the candidate fundamental frequency, and determine the fundamental frequency detection result according to the support score.
[0059] In this embodiment, for ease of description, the following description is made by taking the recognition real-time audio processing system as the execution subject.
[0060] Fundamental frequency detection algorithms are primarily categorized into two main categories: frequency domain and time domain. Frequency domain algorithms use techniques like Fourier transforms to convert signals into the frequency domain for analysis, effectively filtering out noise interference and offering good stability. However, frequency domain algorithms are limited in frequency resolution, resulting in poor accuracy. Time domain algorithms directly analyze signal characteristics, such as zero-crossing rate, along the time axis, offering high accuracy but poor stability in the face of signal changes and being susceptible to noise and signal fluctuations. Balancing the stability and accuracy of fundamental frequency detection algorithms for real-time audio processing systems remains an unresolved issue.
[0061] The present application provides a solution, which performs time domain preprocessing (filtering, clipping and windowing) on multi-frame spliced signals to remove noise and interference in the signal, improve signal quality, and enhance the accuracy of signal recognition. It uses fast Fourier transform to convert the time domain signal into a frequency domain signal, identifies the main frequency components in the spectrum through peak detection, provides a preliminary estimate for the candidate fundamental frequency, obtains the support score of the candidate fundamental frequency through harmonic matching, determines the stability and matching degree of the candidate fundamental frequency under multiple harmonic relationships, and determines the fundamental frequency detection result based on the support score. While ensuring the real-time and high efficiency of audio signal processing, it achieves high-precision fundamental frequency detection and balances the requirements of stability and accuracy.
[0062] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a real-time audio processing system, etc. The following uses the real-time audio processing system as an example to illustrate this embodiment and the following embodiments.
[0063] Based on this, the embodiment of the present application provides a baseband detection method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the baseband detection method of the present application.
[0064] In this embodiment, the baseband detection method includes steps S10 to S40:
[0065] Step S10, performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, wherein the time domain preprocessing includes filtering processing, clipping processing, and windowing processing;
[0066] It should be understood that during audio processing, the real-time audio processing system continuously receives audio input data to generate the original audio signal. Multiple consecutive audio frames from the original audio signal are concatenated to generate a multi-frame spliced signal. This multi-frame spliced signal contains information about the entire audio segment. Subsequent time and frequency domain processing of this multi-frame spliced signal facilitates analysis of the spectral characteristics of the entire audio signal.
[0067] It should be noted that the multi-frame spliced signal is time domain data. Preliminary signal cleaning and adjustment, namely time domain preprocessing, of the multi-frame spliced signal to obtain the first time domain preprocessed data can remove unnecessary noise and signal distortion to better perform frequency domain analysis and fundamental frequency detection. Time domain preprocessing includes filtering, clipping, and windowing. Filtering is used to remove high-frequency resonance peaks in the signal to reduce computational complexity. Clipping is used to eliminate the influence of low-energy noise and improve algorithm accuracy. Windowing is used to weight each frame of the signal to reduce spectral leakage.
[0068] In a feasible implementation manner, before step S10, steps S01 to S04 may be included:
[0069] Step S01, obtaining the frame length of the receiving and transmitting signal and the number of frames of the buffer spliced signal;
[0070] It's important to note that the frame length of the transmitted and received signals is the number of samples contained in each frame when the original audio signal is segmented into multiple frames. In real-time audio processing systems, the frame length of the transmitted and received signals is fixed and finite. This length is determined based on the processor speed and memory capacity of the real-time audio processing system to ensure that the audio processing task can be completed within each frame.
[0071] It should be understood that to improve the accuracy of the audio processing algorithm, multiple frames of the signal are concatenated. The number of concatenated signal frames in the buffer is the number of consecutive frames used when concatenating the original audio signal. The maximum supported total signal frame length can be determined based on the hardware resources (processor speed and memory capacity) of the real-time audio processing system. The signal frame number can be calculated by dividing the total signal frame length by the frame length of the transmitted and received signals.
[0072] Step S02, determining the buffer length according to the frame length of the receiving and transmitting signal and the number of frames of the buffer splicing signal;
[0073] It should be noted that the buffer length is the size of the memory area (i.e., the buffer) used to store and process audio data. It is calculated by multiplying the frame length of the transmitted and received signals by the number of signal frames spliced into the buffer. In real-time audio processing systems, the buffer is used to temporarily store received data frames for subsequent processing.
[0074] Step S03, when the current frame data of the original audio signal arrives at the buffer, obtaining the buffer index of the current frame data;
[0075] It should be understood that the original audio signal is the audio input data for baseband detection, which can be obtained from an input device (such as a microphone) or a data stream received over the network. The current frame data is the original audio signal that has currently arrived in the buffer. The buffer index is the data index of the current position in the buffer, indicating the position of the current frame data in the buffer.
[0076] Specifically, if the frame length of the transceiver signal is FrameLen and the number of buffer splicing signal frames is FrameNum, the starting index of the current frame data is FrameLen*(FrameNum-1), and the index range of the current frame data is FrameLen*(FrameNum-1) to FrameLen*FrameNum.
[0077] Step S04 , performing frame-by-frame splicing on the current frame data based on the buffer length and the buffer index by using a piecewise function to obtain a multi-frame splicing signal.
[0078] It should be noted that the piecewise function is used to process the current frame data according to the buffer length and buffer index. The current frame data can be spliced to the existing buffer data according to the specified positional relationship to obtain a complete multi-frame splicing signal, providing a continuous data basis for the subsequent baseband detection algorithm. After obtaining the current frame data, determining the starting position of the current frame data in the buffer, and obtaining the buffer index, the number of buffer splicing signal frames will be reduced by one to obtain the first preset number of frames. The frame data before the first preset number of frames in the buffer will be slid forward one frame to overwrite the old buffer data, and then the current frame data will be filled into the empty space at the end of the buffer. The length of the multi-frame splicing signal is equal to the frame length of the transceiver signal multiplied by the number of buffer splicing signal frames.
[0079] For example, the expression of the piecewise function is as follows:
[0080]
[0081] Where FrameLen represents the frame length of the transmitted and received signals, FrameNum represents the number of concatenated signal frames in the buffer, and FrameLen*FrameNum represents the buffer length. InputData represents the original audio signal, CurData represents the concatenated multi-frame signal, and i represents the sample index in the buffer.
[0082] In a feasible implementation, step S10 may include steps S11 to S13:
[0083] Step S11, filtering the multi-frame spliced signal based on the filter coefficient to obtain filtered data;
[0084] It should be noted that a filter is a signal processing tool used to filter a signal to remove unwanted frequency components or noise. It includes low-pass filters, high-pass filters, and band-pass filters. The filter coefficient is a parameter in the filter that defines the filter's characteristics and controls the filtering effect. The filter coefficient can be calculated based on the filter's design requirements. Filters are categorized as finite impulse response (FIR) filters and infinite impulse response (IIR) filters. FIR filters contain only forward coefficients, while IIR filters contain both forward and feedback coefficients.
[0085] It should be understood that the FIR filter depends on the current and past sample values of the input signal, and the FIR filter output is a linear combination of the input signal and the filter coefficients (i.e., forward coefficients). The IIR filter depends on the current and past sample values of the input signal and the current and past sample values of the output signal, and the IIR filter output is a linear combination of the input signal and the filter coefficients (i.e., forward coefficients) plus a linear combination of the past output signal and the filter coefficients (i.e., feedback coefficients). In filter design, the amplitude-frequency characteristic accuracy of the FIR filter is lower than that of the IIR filter, but the linear phase is high. The amplitude-frequency characteristic accuracy of the IIR filter is not linear phase.
[0086] For example, the expression of the IIR filter is as follows:
[0087]
[0088] Where x represents the filter input signal, and y represents the filter output signal. x(ni) represents the value of the input signal at time point ni, and y(nj) represents the value of the output signal at time point nj. a(i) represents the forward coefficient of the filter, where i is an integer from 0 to N. N is the order of the forward coefficient, which represents the number of forward coefficients used for filtering. b(j) represents the feedback coefficient of the filter, where j is an integer from 1 to M. M is the order of the feedback coefficient, which represents the number of feedback coefficients used for filtering. y(n) represents the output signal of the filter at time point n.
[0089] Step S12, performing clipping processing on the filtered data based on a clipping threshold using a level clipping function to obtain clipping data;
[0090] It should be noted that the level clipping function is a signal processing function used to eliminate the influence of small energy noise. The clipping threshold is a parameter in the level clipping function, which is used to define the lower limit of the signal amplitude. The appropriate clipping threshold can be selected according to the signal processing requirements. After obtaining the signal amplitude of the filtered data input at each time point, the signal amplitude is compared with the clipping threshold. When the signal amplitude is greater than or equal to the clipping threshold, the amplitude of the clipping data output at that time point remains unchanged; when the signal amplitude is less than the clipping threshold, the amplitude of the clipping data output at that time point is 0. That is, the output amplitude corresponding to the filtered data whose signal amplitude reaches the clipping threshold remains unchanged, and the output amplitude corresponding to the filtered data that does not reach the clipping threshold is 0.
[0091] Exemplarily, the expression of the level clipping function is as follows:
[0092]
[0093] Where x represents the clipped input signal and y represents the clipped output signal. x(n) represents the amplitude of the input signal at time point n, y(n) represents the amplitude of the output signal at time point n, and C L Indicates the clipping threshold.
[0094] Step S13: performing windowing processing on the clipping data by using an analysis window function to obtain first time-domain preprocessed data.
[0095] It should be noted that the analysis window function is used for signal windowing to reduce spectral leakage. During the windowing process, the clipped data is point-by-point multiplied by the analysis window function to produce the first time-domain preprocessed data, which modifies the signal shape. Windowing can smooth the transition between signal boundaries and prevent excessive abrupt changes from adversely affecting spectral analysis.
[0096] For example, the expression of the analysis window and windowing process is as follows:
[0097] CurDataAna(i)=CurData(i)*WindowAna(i),0≤i <FrameLen*FrameNum
[0098] Where CurData represents the clipping data, CurDataAna represents the first time domain preprocessed data, i represents the sample index of the buffer, and WindowAna represents the analysis window function.
[0099] For example, the Hanning window can be selected as the analysis window function. The expression of the Hanning window is as follows:
[0100]
[0101] Where Window(i) represents the value of the analysis window function at position i, and i represents the index of the current sample point, ranging from 0 to FrameLen*FrameNum-1.
[0102] It should be understood that the value range of the Hanning window is between 0 and 1. The shape of the Hanning window function is a symmetrical cosine curve. The value of the center point (i.e., i = (FrameLen*FrameNum-1) / 2) is 1, and the values of the two end points (i.e., i = 0 and i = FrameLen*FrameNum-1) are 0. It can smoothly attenuate the two ends of the signal, thereby reducing spectral leakage.
[0103] Step S20, performing fast Fourier transform on the first time domain preprocessed data to obtain preprocessed frequency domain signal data;
[0104] It should be noted that the Fast Fourier Transform (FFT) is an efficient signal conversion algorithm that can convert a time-domain signal into a frequency-domain signal. This algorithm processes the first time-domain preprocessed data to obtain preprocessed frequency-domain signal data. The Fast Fourier Transform (FFT) uses mathematical transformations to display the energy distribution of a signal at different frequencies.
[0105] For example, the fast Fourier transform process is expressed as follows:
[0106] FFTData=FFT(CurDataAna,FrameLen*FrameNum)
[0107] Where CurDataAna represents the time domain signal data obtained after time domain preprocessing, i.e., the first time domain preprocessed data. FrameLen*FrameNum represents the total length of the signal to be fast Fourier transformed, which is used to determine the resolution and range of the frequency domain signal after the fast Fourier transform. FFT represents the fast Fourier transform operation, and FFTData represents the frequency domain signal data obtained through the fast Fourier transform, i.e., the preprocessed frequency domain signal data.
[0108] It should be understood that during the fast Fourier transform process, the input first time domain preprocessed data will be recursively decomposed into smaller parts, the discrete Fourier transform (DFT) of each subsequence will be gradually calculated, and the calculated discrete Fourier transform of each subsequence will be merged to obtain complete frequency domain signal data, that is, preprocessed frequency domain signal data. Specifically, the first time domain preprocessed data will be decomposed into subsequences of even and odd positions, and DFT will be recursively applied to each subsequence. The frequency domain data of the subsequences will be merged into a complete frequency domain representation through rotation factors and merging operations. The preprocessed frequency domain signal data obtained by fast Fourier transform can determine the frequency component, amplitude and phase information of the original audio signal.
[0109] Step S30, performing peak detection on the pre-processed frequency domain signal data to obtain a candidate fundamental frequency;
[0110] It should be noted that peak detection is used to identify frequencies with large amplitudes in the preprocessed frequency-domain signal data. These frequencies may be fundamental frequencies or harmonics. A window can be slid across the preprocessed frequency-domain signal data to detect the frequency with the largest amplitude within the window. If the frequency is located in the center of the window, it is considered a peak point, and the candidate fundamental frequency is obtained.
[0111] It should be noted that in the field of signal processing, the fundamental frequency is the smallest positive periodic part in the Fourier series decomposition of a periodic function, called the fundamental wave or the first harmonic. The fundamental frequency is the basic frequency of the signal, which determines the pitch or basic oscillation frequency of the signal. The fundamental frequency usually has the largest amplitude and appears as a significant peak in the spectrum. Harmonics are frequency components that are integer multiples of the fundamental frequency, and are other frequency components in the periodic signal except the fundamental frequency. For example, in music, the fundamental frequency of a note determines the pitch of the note (such as the fundamental frequency of the A4 note is 440Hz).
[0112] In a feasible implementation, step S30 may include steps S31 to S34:
[0113] Step S31, traversing the pre-processed frequency domain signal data through a preset sliding window to obtain a frequency point global subscript, a frequency point window local subscript, and a frequency point amplitude of the preset pre-processed frequency domain signal data;
[0114] It should be understood that the preset sliding window is a fixed-length window used to slide on the preprocessed frequency domain signal data to detect peaks. The length of the preset sliding window can be set according to the length and resolution of the frequency domain signal. When the preset sliding window is used to traverse the preprocessed frequency domain signal data, two subscripts are obtained for each frequency point: one is the global subscript on the frequency domain signal, which is the global subscript of the frequency point, and the other is the local subscript within the sliding window, which is the local subscript of the frequency point window. At the same time, the amplitude corresponding to the frequency point is obtained to obtain the frequency point amplitude.
[0115] It should be understood that the global subscript of the frequency point is the position index of each frequency point in the entire preset preprocessed frequency domain signal data, the local subscript of the frequency point window is the relative position index of the frequency point in the preset sliding window, and the frequency point amplitude is the amplitude value of each frequency point in the preset preprocessed frequency domain signal data.
[0116] Step S32, determining a candidate frequency point from the preset pre-processed frequency domain signal data based on the frequency point amplitude;
[0117] It should be understood that the candidate frequency point is a possible peak frequency point detected in the sliding window. When the pre-processed frequency domain signal data is traversed using the preset sliding window, the frequency point with the largest frequency amplitude in the preset sliding window will be selected as the candidate frequency point.
[0118] Step S33, when it is detected based on the local subscript of the frequency window that the candidate frequency point is located at the center of the preset sliding window, obtaining the global subscript and frequency amplitude of the candidate frequency point to obtain a peak detection result;
[0119] It should be noted that after obtaining the candidate frequency point, the local subscript of the frequency window of the candidate frequency point will be obtained, and it will be detected whether the local subscript of the frequency window is half of the preset sliding window length. If so, the candidate frequency point is considered to be located at the center of the preset sliding window, and the global frequency subscript and frequency amplitude of the candidate frequency point are recorded. The candidate frequency point is used as the peak point until the preset sliding window traverses the preprocessed frequency domain signal data and all peak points of the current frame data are obtained, that is, the peak detection result.
[0120] Step S34: determining a candidate fundamental frequency based on the peak detection result.
[0121] It should be noted that by collecting and organizing all recorded peak points into a set, the candidate fundamental frequency can be obtained.
[0122] Exemplarily, for the above steps S31 to S34, a preset sliding window with a window length of 2*M+1 can be used to traverse the preprocessed frequency domain signal data FFTData. The global frequency subscripts, local frequency window subscripts, and frequency amplitudes of all frequency points in the preset sliding window are obtained, where the range of the global subscript is [0, FrameLen*FrameNum-1, and the range of the local frequency window subscript is [0, 2*M]. If the local frequency window subscript of the frequency point with the largest amplitude in the current preset sliding window is equal to M, then the frequency point is determined to be a peak point, the global subscript and amplitude of the frequency point are recorded, all peak points are found, and candidate fundamental frequencies are obtained, recorded as F=[f1, f2,…, fN], completing peak detection.
[0123] Step S40 , performing harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range, obtaining a support score of the candidate fundamental frequency, and determining a fundamental frequency detection result according to the support score.
[0124] It should be noted that for each candidate fundamental frequency, its theoretical harmonic sequence will be generated, and whether these harmonics exist in the actual frequency domain signal within the allowable preset frequency deviation range will be detected. The support score will be calculated based on the number and amplitude of the matched harmonics. The higher the score, the more likely the candidate fundamental frequency is to be the true fundamental frequency. Based on the support score, it can be determined whether there is a true fundamental frequency among the candidate fundamental frequencies, and the fundamental frequency detection result can be obtained.
[0125] In a feasible implementation manner, the step of performing harmonic matching on the candidate fundamental frequency based on the preset frequency deviation range to obtain the support score of the candidate fundamental frequency in step S40 may include steps S41 to S43:
[0126] Step S41, generating a theoretical harmonic sequence based on the candidate fundamental frequency;
[0127] It should be noted that for each candidate fundamental frequency of the current frame data, a plurality of integer multiple frequency components, ie, a plurality of theoretical harmonics, are generated to obtain a theoretical harmonic sequence of the candidate fundamental frequency.
[0128] Step S42, performing harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range and the theoretical harmonic sequence to obtain the number of matching harmonics and the matching harmonic amplitude;
[0129] It should be understood that in actual signals, due to various factors, the harmonic frequencies may not appear completely accurately on the theoretical values. Therefore, a frequency deviation range is required to allow for a certain error. This allows the theoretical harmonics to be matched with the candidate fundamental frequency within the allowable deviation range to accurately identify the true harmonic components.
[0130] It should be noted that the preset frequency deviation range is the allowable frequency deviation range between the candidate fundamental frequency and the theoretical harmonics. It can be set based on the frequency resolution or empirical value and is used to determine whether the candidate fundamental frequency matches the theoretical harmonics. The number of matching harmonics is the number of candidate fundamental frequencies that match the theoretical harmonics within the preset frequency deviation range. The matching harmonic amplitude is the amplitude of the candidate fundamental frequency that matches the theoretical harmonics within the preset frequency deviation range.
[0131] It should be understood that after obtaining the theoretical harmonic sequence, the allowable frequency deviation is determined based on the frequency resolution or empirical value, resulting in a preset frequency deviation range. Based on the frequencies of multiple integer multiple frequency components in the theoretical harmonic sequence and the preset frequency deviation range, multiple allowable frequency intervals can be determined. Candidate fundamental frequencies falling within the multiple allowable frequency intervals are considered to match the theoretical harmonics.
[0132] Step S43: determining the support score of the candidate fundamental frequency based on the number of matching harmonics and the matching harmonic amplitudes.
[0133] It should be noted that the support score is the credibility score of the candidate fundamental frequency, which can be obtained by weighted summation of the number of matching harmonics and the matching harmonic amplitudes.
[0134] For example, for the above steps S41 to S43, if the candidate fundamental frequency of the current frame data is recorded as F=[f1, f2, ..., f N ], for each candidate fundamental frequency f i ∈F(i=1,2,3,...,N), the theoretical harmonic position k*f will be calculated i (k=2,3,…,K,usually K=5),get the theoretical harmonic sequence. Take half of the frequency resolution or the empirical value as the allowable frequency deviation δ, then the preset frequency deviation range is [k*f i -δ,k*f i +δ]. For multiple frequency allowable intervals obtained based on the theoretical harmonic sequence and the preset frequency deviation range, it is detected whether there is a peak in the candidate fundamental frequency F of the current frame data whose frequency falls within the multiple frequency allowable intervals. If so, the number and amplitude of candidate fundamental frequencies matching the theoretical harmonics in F are obtained, and the number and amplitude of matching harmonics are obtained. The weighted sum of the harmonic amplitudes is used as the support score of the candidate fundamental frequency.
[0135] In a feasible embodiment, the step of determining the fundamental frequency detection result based on the support score in step S40 may include: when the support score of the candidate fundamental frequency is greater than or equal to the preset harmonic support threshold, and the support score difference of the candidate fundamental frequencies is greater than the preset difference, determining the target fundamental frequency from the candidate fundamental frequencies based on the support score to obtain the fundamental frequency detection result; when the support score of the candidate fundamental frequency is greater than or equal to the preset harmonic support threshold, and the support score difference of the candidate fundamental frequencies is less than or equal to the preset difference, determining the target fundamental frequency from the candidate fundamental frequencies based on the frequency sorting result of the candidate fundamental frequencies to obtain the fundamental frequency detection result; when the support score of the candidate fundamental frequency is less than the preset harmonic support threshold, determining the fundamental frequency detection result as the absence of the fundamental frequency.
[0136] It should be noted that the preset harmonic support threshold is the minimum support score threshold, which is used to determine whether a candidate fundamental frequency is valid. The preset difference is the minimum difference between support scores, which is used to determine whether the support scores of multiple candidate fundamental frequencies are similar. The target fundamental frequency is the fundamental frequency finally determined among the candidate fundamental frequencies.
[0137] It should be understood that the frequency sorting result is the result of sorting the candidate fundamental frequencies in ascending frequency order. Because the preprocessed frequency domain signal data is sorted in ascending frequency order, filtering and clipping are used during time domain preprocessing of the multi-frame spliced signal to eliminate some resonance peaks and false peaks caused by low-energy noise. After peak detection, no further peak preprocessing of sorting and screening is required to obtain the frequency sorting result of the candidate fundamental frequencies.
[0138] In addition, it should be understood that after obtaining the candidate fundamental frequency, all candidate fundamental frequencies of the current frame data will be traversed to obtain the support score corresponding to each candidate fundamental frequency, and the support score of the candidate fundamental frequency will be compared with the preset harmonic support threshold in turn.
[0139] If there are candidate fundamental frequencies in the current frame data whose support scores are greater than or equal to the preset harmonic support threshold, the support score difference between the candidate fundamental frequencies with support scores greater than or equal to the threshold will be calculated. If there is a candidate fundamental frequency with the highest score, and the difference between the support score of the highest-scoring candidate fundamental frequency and the score of the second-highest-scoring candidate fundamental frequency is greater than the preset difference, then this candidate fundamental frequency will be used as the target fundamental frequency. If there is a candidate fundamental frequency with the highest score, and the difference between the support score of the highest-scoring candidate fundamental frequency and the score of the second-highest-scoring candidate fundamental frequency is less than or equal to the preset difference, then the candidate fundamental frequency with the score difference less than or equal to the preset difference with the second-highest-scoring candidate fundamental frequency is obtained, and the candidate fundamental frequency with the lowest frequency is used as the target fundamental frequency based on the frequency sorting result. If the support scores of all candidate fundamental frequencies are less than the preset harmonic support threshold (indicating that there are not at least two harmonics supporting a candidate fundamental frequency), it is considered that there is no obvious fundamental frequency component in the current frame data.
[0140] This embodiment performs time domain preprocessing (filtering, clipping and windowing) on the multi-frame spliced signal to remove noise and interference in the signal, improve the signal quality, and enhance the accuracy of signal recognition. It uses fast Fourier transform to convert the time domain signal into a frequency domain signal, identifies the main frequency components in the spectrum through peak detection, provides a preliminary estimate for the candidate fundamental frequency, obtains the support score of the candidate fundamental frequency through harmonic matching, determines the stability and matching degree of the candidate fundamental frequency under multiple harmonic relationships, and determines the fundamental frequency detection result based on the support score. While ensuring the real-time and high efficiency of audio signal processing, it achieves high-precision fundamental frequency detection and balances the requirements of stability and accuracy.
[0141] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 After step S40, the baseband detection method further includes steps S50 to S90:
[0142] Step S50, performing sound effect processing on the pre-processed frequency domain signal data based on the fundamental frequency detection result to obtain target frequency domain signal data;
[0143] It should be noted that the fundamental frequency detection result obtained after harmonic matching and support scoring reflects the fundamental frequency components of the audio signal. Based on the fundamental frequency detection result, the specific frequency components of the preprocessed frequency domain signal data can be adjusted to enhance or change the audio effect, thereby obtaining the target frequency domain signal data. Based on the detected target fundamental frequency, the pitch of the preprocessed frequency domain signal data can be adjusted, specific frequency components can be enhanced, or other sound effect optimizations can be performed to obtain the target frequency domain signal data that meets the user's sound effect optimization needs.
[0144] For example, in voice communication, fundamental frequency detection can help restore the pitch and timbre of the voice signal, thereby improving call quality. After determining the basic period and structure of the signal through fundamental frequency detection, signal processing techniques such as filtering and interpolation can be used to remove noise and interference components, reconstruct the waveform of the original signal, and obtain the target frequency domain signal data.
[0145] Step S60, performing an inverse fast Fourier transform on the target frequency domain signal data to obtain second time domain signal data;
[0146] It should be noted that the Inverse Fast Fourier Transform (IFFT) is the inverse process of the Fast Fourier Transform, which is used to convert the frequency domain signal back to the time domain signal. Specifically, during the IFFT process, the complex data points in the target frequency domain signal data are converted into complex data points in the time domain, and the original signal waveform is reconstructed. The second time domain signal data is obtained by converting the amplitude and phase information of each frequency component in the frequency domain into time point samples in the time domain.
[0147] For example, the inverse fast Fourier transform process is expressed as follows:
[0148] IFFTData=IFFT(FFTData,FrameLen*FrameNum)
[0149] Where FFTData represents the frequency-domain signal data processed with the audio effect, i.e., the target frequency-domain signal data. FrameLen*FrameNum represents the total length of the signal to be inverse fast Fourier transformed (IFFT), which is used to determine the length of the time-domain signal after the IFFT. IFFT represents the inverse fast Fourier transform (IFFT), and IFFTData represents the time-domain signal data obtained through the IFFT, i.e., the second time-domain signal data.
[0150] Step S70, performing windowing processing on the second time domain signal data by using a synthetic window function to obtain second time domain preprocessed data;
[0151] It's important to note that the synthesis window function is used to compensate for energy attenuation and reduce reconstruction errors during time-domain signal reconstruction. Multiplying the time-domain signal by the synthesis window function point by point improves the quality of audio signal reconstruction, ensuring that the integrity and continuity of the audio signal are maintained during the reconstruction process. The synthesis window function complements the analysis window function, using the same window function to ensure smooth addition of overlapping signals.
[0152] Exemplarily, the expression of the windowing process of the synthetic window function is as follows:
[0153] CurDataSyn(i)=IFFTData(i)*WindowSyn(i),0≤i <FrameLen*FrameNum
[0154] Wherein, IFFTData represents the second time domain signal data, CurDataSyn represents the second time domain preprocessed data, i represents the sample index of the buffer, and WindowAna represents the synthesis window function.
[0155] Step S80, obtaining the current frame synthesis window signal data and the adjacent frame synthesis window signal data from the second time domain preprocessed data based on the frame length of the receiving and transmitting signal and the buffer index of the current frame data;
[0156] It should be noted that the synthetic window signal data of the current frame data, that is, the synthetic window signal data of the current frame, can be obtained based on the second time domain preprocessing data for subsequent signal reconstruction. According to the frame length of the transceiver signal and the buffer index of the current frame data, the synthetic window signal data of the adjacent frame can be extracted from the second time domain preprocessing data to obtain the adjacent frame synthetic window signal data. The adjacent frame synthetic window signal data is the data of the previous frame or the next frame adjacent to the current frame. The current frame synthetic window signal data of this round is superimposed on the adjacent frame synthetic window signal data of the previous round to obtain the target audio signal of this round; the adjacent frame synthetic window signal data of this round is superimposed on the current frame synthetic window signal data of the next round to obtain the target audio signal of the next round.
[0157] Exemplarily, the formula for extracting the adjacent frame synthesis window signal data is as follows:
[0158] TmpData(i)=CurDataSyn(i+FrameLen),0≤i <FrameLen
[0159] Where CurDataSyn represents the second time domain preprocessed data, TmpData represents the adjacent frame synthesis window signal data, i represents the sample index of the buffer, and FrameLen represents the frame length of the transmitted and received signal.
[0160] Step S90 : Superimposing the current frame synthesis window signal data and the adjacent frame synthesis window signal data to obtain a target audio signal.
[0161] It should be understood that when the signal data of the current frame synthesis window is superimposed with the signal data of the adjacent frame synthesis window, these signal data are synthesized along the time axis to produce a smooth and continuous audio signal, resulting in the target audio signal. This avoids breaks and discontinuities between frames and ensures a smooth transition of the audio signal. The target audio signal is the final audio data obtained through a series of sound effect processing and frame synthesis, and is available for subsequent playback or other processing by the user.
[0162] Exemplarily, the signal reconstruction formula is as follows:
[0163] OutputData(i)=CurDataSyn(i)+TmpData(i),0≤i <FrameLen
[0164] Where CurDataSyn represents the second time domain preprocessed data, TmpData represents the adjacent frame synthesis window signal data, OutputData represents the target audio signal, i represents the sample index of the buffer, and FrameLen represents the frame length of the transmitted and received signals.
[0165] This embodiment performs sound effect processing on the pre-processed frequency domain signal based on the fundamental frequency detection result, and can specifically adjust specific frequency components in the signal as needed, thereby improving the overall sound quality and clarity of the audio. The target frequency domain signal is converted back to the time domain through inverse fast Fourier transform, so that the signal after frequency domain processing can be restored to time domain waveform data in a suitable format for subsequent time domain processing. The signal is windowed by a synthetic window function to compensate for energy attenuation and reduce reconstruction errors. By extracting the current frame and adjacent frame data based on the frame length and buffer index, the current frame is superimposed with the adjacent frame signal data, which helps to improve the continuity and coherence of the audio signal and reduce the abruptness caused by the break between frames, and finally generates a high-quality target audio signal to ensure the continuity, stability and naturalness of the output audio.
[0166] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the baseband detection method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0167] This application also provides a fundamental frequency detection device, please refer to Figure 3 , the fundamental frequency detection device comprises:
[0168] The data processing module 10 is used to perform time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, wherein the time domain preprocessing includes filtering, clipping and windowing.
[0169] A data transformation module 20 is configured to perform a fast Fourier transform on the first time-domain preprocessed data to obtain preprocessed frequency-domain signal data;
[0170] The data processing module 10 is further configured to perform peak detection on the pre-processed frequency domain signal data to obtain a candidate fundamental frequency;
[0171] The fundamental frequency detection module 30 is configured to perform harmonic matching on the candidate fundamental frequencies based on a preset frequency deviation range, obtain a support score for the candidate fundamental frequencies, and determine a fundamental frequency detection result according to the support score.
[0172] In one embodiment, the data processing module 10 is further used to filter the multi-frame spliced signal based on the filter coefficient through a filter to obtain filtered data; to clip the filtered data based on the clipping threshold through a level clipping function to obtain clipping data; and to window the clipping data through an analysis window function to obtain first time domain preprocessed data.
[0173] In one embodiment, the data processing module 10 is further used to traverse the preprocessed frequency domain signal data through a preset sliding window to obtain the frequency global subscript, frequency window local subscript and frequency amplitude of the preset preprocessed frequency domain signal data; determine the candidate frequency from the preset preprocessed frequency domain signal data based on the frequency amplitude; when it is detected that the candidate frequency is located at the center of the preset sliding window based on the frequency window local subscript, obtain the frequency global subscript and frequency amplitude of the candidate frequency to obtain a peak detection result; and determine the candidate fundamental frequency based on the peak detection result.
[0174] In one embodiment, the fundamental frequency detection module 30 is further used to generate a theoretical harmonic sequence based on the candidate fundamental frequency; perform harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range and the theoretical harmonic sequence to obtain the number of matching harmonics and the matching harmonic amplitude; and determine the support score of the candidate fundamental frequency based on the number of matching harmonics and the matching harmonic amplitude.
[0175] In one embodiment, the fundamental frequency detection module 30 is further used to determine the target fundamental frequency from the candidate fundamental frequencies based on the support score to obtain a fundamental frequency detection result when the support score of the candidate fundamental frequency is greater than or equal to the preset harmonic support threshold and the difference in the support scores of the candidate fundamental frequencies is greater than the preset difference; determine the target fundamental frequency from the candidate fundamental frequencies based on the frequency sorting result of the candidate fundamental frequencies to obtain a fundamental frequency detection result when the support score of the candidate fundamental frequency is greater than or equal to the preset harmonic support threshold and the difference in the support scores of the candidate fundamental frequencies is less than or equal to the preset difference; and determine the fundamental frequency detection result as the absence of a fundamental frequency when the support score of the candidate fundamental frequency is less than the preset harmonic support threshold.
[0176] In one embodiment, the data processing module 10 is further used to obtain the frame length of the receiving and transmitting signal and the number of frames of the buffer spliced signal; determine the buffer length based on the frame length of the receiving and transmitting signal and the number of frames of the buffer spliced signal; when the current frame data of the original audio signal arrives at the buffer, obtain the buffer index of the current frame data; and perform frame splicing on the current frame data based on the buffer length and the buffer index through a piecewise function to obtain a multi-frame spliced signal.
[0177] In one embodiment, the data processing module 10 is further used to perform sound effect processing on the pre-processed frequency domain signal data based on the fundamental frequency detection result to obtain target frequency domain signal data; perform inverse fast Fourier transform on the target frequency domain signal data to obtain second time domain signal data; perform windowing processing on the second time domain signal data through a synthesis window function to obtain second time domain preprocessed data; obtain current frame synthesis window signal data and adjacent frame synthesis window signal data from the second time domain preprocessed data based on the frame length of the receiving and transmitting signal and the buffer index of the current frame data; and superimpose the current frame synthesis window signal data and the adjacent frame synthesis window signal data to obtain the target audio signal.
[0178] The fundamental frequency detection device provided in this application, employing the fundamental frequency detection method described in the aforementioned embodiments, can address the technical problem of balancing the stability and accuracy of fundamental frequency detection algorithms used in real-time audio processing systems. Compared to the prior art, the fundamental frequency detection device provided in this application achieves the same beneficial effects as the fundamental frequency detection method described in the aforementioned embodiments. Other technical features of the fundamental frequency detection device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0179] The present application provides a baseband detection device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the baseband detection method of the above-mentioned embodiment 1.
[0180] Reference below Figure 4 , which shows a schematic structural diagram of a baseband detection device suitable for implementing the embodiments of the present application. The baseband detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The fundamental frequency detection device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0181] like Figure 4As shown, the fundamental frequency detection device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the fundamental frequency detection device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output device 1008 including, for example, LCD (Liquid Crystal Display), speaker, vibrator, etc.; storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and communication device 1009. The communication device 1009 can allow the baseband detection device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows a baseband detection device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or provided instead.
[0182] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0183] The fundamental frequency detection device provided in this application, employing the fundamental frequency detection method described in the aforementioned embodiment, can address the technical problem of balancing the stability and accuracy of fundamental frequency detection algorithms used in real-time audio processing systems. Compared to the prior art, the fundamental frequency detection device provided in this application achieves the same beneficial effects as the fundamental frequency detection method described in the aforementioned embodiment. Other technical features of this fundamental frequency detection device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0184] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0185] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0186] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, wherein the computer-readable program instructions are used to execute the fundamental frequency detection method in the above embodiment.
[0187] The computer-readable storage medium provided in this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory) or flash memory, optical fiber, CD-ROM (CD-Read Only Memory, portable compact disk read-only memory), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0188] The computer-readable storage medium may be included in the baseband detection device, or may exist independently without being assembled into the baseband detection device.
[0189] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the baseband detection device, the baseband detection device: performs time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessing data, and the time domain preprocessing includes filtering processing, clipping processing and windowing processing; performs fast Fourier transform on the first time domain preprocessing data to obtain preprocessed frequency domain signal data; performs peak detection on the preprocessed frequency domain signal data to obtain a candidate baseband; performs harmonic matching on the candidate baseband based on a preset frequency deviation range to obtain a support score of the candidate baseband, and determines the baseband detection result according to the support score.
[0190] The computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).
[0191] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0192] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0193] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned fundamental frequency detection method. This computer-readable storage medium addresses the technical problem of balancing the stability and accuracy of fundamental frequency detection algorithms used in real-time audio processing systems. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the fundamental frequency detection method provided in the aforementioned embodiments and are not further elaborated here.
[0194] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned baseband detection method when executed by a processor.
[0195] The computer program product provided in this application can address the technical problem of balancing the stability and accuracy of fundamental frequency detection algorithms used in real-time audio processing systems. Compared to the prior art, the beneficial effects of the computer program product provided in this application are similar to those of the fundamental frequency detection methods provided in the aforementioned embodiments, and are not further elaborated here.
[0196] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A fundamental frequency detection method, characterized in that: The fundamental frequency detection method comprises: Performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, wherein the time domain preprocessing includes filtering processing, clipping processing, and windowing processing; Performing a fast Fourier transform on the first time-domain preprocessed data to obtain preprocessed frequency-domain signal data; Performing peak detection on the preprocessed frequency domain signal data to obtain a candidate fundamental frequency; Harmonic matching is performed on the candidate fundamental frequencies based on a preset frequency deviation range to obtain a support score for the candidate fundamental frequencies, and a fundamental frequency detection result is determined according to the support score.
2. The method according to claim 1, wherein The step of performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data includes: Filtering the multi-frame splicing signal based on the filter coefficient to obtain filtered data; performing clipping processing on the filtered data based on a clipping threshold by using a level clipping function to obtain clipping data; The clipping data is subjected to windowing processing by means of an analysis window function to obtain first time-domain preprocessed data.
3. The method according to claim 1, wherein The step of performing peak detection on the pre-processed frequency domain signal data to obtain a candidate fundamental frequency includes: Traversing the pre-processed frequency domain signal data through a preset sliding window, obtaining a frequency point global subscript, a frequency point window local subscript, and a frequency point amplitude of the preset pre-processed frequency domain signal data; Determining candidate frequency points from the preset pre-processed frequency domain signal data based on the frequency point amplitudes; When it is detected based on the local subscript of the frequency window that the candidate frequency point is located at the center of the preset sliding window, obtaining the global subscript and frequency amplitude of the candidate frequency point to obtain a peak detection result; A candidate fundamental frequency is determined based on the peak detection result.
4. The method according to claim 1, wherein The step of performing harmonic matching on the candidate fundamental frequency based on the preset frequency deviation range to obtain a support score of the candidate fundamental frequency includes: generating a theoretical harmonic sequence based on the candidate fundamental frequency; Performing harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range and the theoretical harmonic sequence to obtain the number of matching harmonics and the matching harmonic amplitude; A support score for the candidate fundamental frequency is determined based on the number of matching harmonics and the amplitude of the matching harmonics.
5. The method according to claim 1, wherein The step of determining the fundamental frequency detection result according to the support score includes: When the support score of the candidate fundamental frequency is greater than or equal to a preset harmonic support threshold, and the support score difference of the candidate fundamental frequency is greater than a preset difference, determining a target fundamental frequency from the candidate fundamental frequencies based on the support score, and obtaining a fundamental frequency detection result; When the support score of the candidate fundamental frequencies is greater than or equal to a preset harmonic support threshold, and the difference in the support scores of the candidate fundamental frequencies is less than or equal to the preset difference, determining the target fundamental frequency from the candidate fundamental frequencies based on the frequency sorting result of the candidate fundamental frequencies to obtain a fundamental frequency detection result; When the support score of the candidate fundamental frequency is less than a preset harmonic support threshold, it is determined that the fundamental frequency detection result is that the fundamental frequency does not exist.
6. The method according to claim 1, wherein Before the step of performing time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, the method further includes: Get the frame length of the transmitted and received signals and the number of frames of the buffered spliced signals; Determine the buffer length according to the frame length of the transmitted and received signals and the number of frames of the buffer spliced signal; When the current frame data of the original audio signal arrives at the buffer, obtaining the buffer index of the current frame data; The current frame data is frame-spliced and spliced based on the buffer length and the buffer index through a piecewise function to obtain a multi-frame splicing signal.
7. The method according to any one of claims 1 to 6, characterized in that After the steps of performing harmonic matching on the candidate fundamental frequency based on the preset frequency deviation range to obtain a support score of the candidate fundamental frequency, and determining a fundamental frequency detection result according to the support score, the method further includes: Performing sound effect processing on the pre-processed frequency domain signal data based on the fundamental frequency detection result to obtain target frequency domain signal data; Performing an inverse fast Fourier transform on the target frequency domain signal data to obtain second time domain signal data; Performing windowing processing on the second time domain signal data by using a synthetic window function to obtain second time domain preprocessed data; Acquire current frame synthesis window signal data and adjacent frame synthesis window signal data from the second time domain preprocessed data based on the frame length of the received and transmitted signal and the buffer index of the current frame data; The current frame synthesis window signal data and the adjacent frame synthesis window signal data are superimposed to obtain a target audio signal.
8. A fundamental frequency detection device, characterized in that: The device comprises: A data processing module, configured to perform time domain preprocessing on the multi-frame spliced signal to obtain first time domain preprocessed data, wherein the time domain preprocessing includes filtering, clipping, and windowing; a data conversion module, configured to perform a fast Fourier transform on the first time-domain preprocessed data to obtain preprocessed frequency-domain signal data; The data processing module is further configured to perform peak detection on the pre-processed frequency domain signal data to obtain a candidate fundamental frequency; The fundamental frequency detection module is used to perform harmonic matching on the candidate fundamental frequency based on a preset frequency deviation range, obtain a support score of the candidate fundamental frequency, and determine a fundamental frequency detection result according to the support score.
9. A fundamental frequency detection device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the baseband detection method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the baseband detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Method and system for determining fundamental frequency of Chinese lute strings
CN121725813A