Method, device, equipment, storage medium and computer program product for extracting audio fundamental frequency based on mel spectrum

CN122799883APending Publication Date: 2026-09-22FAW VOLKSWAGEN AUTOMOTIVE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326488.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

时域法为通过分析语音信号的时域特性来提取基频,这种方法对噪声敏感,要求较高的测量精度和较低的背景噪声水平导致提取基频的准确性差,且计算工作量较大、周期较长

Benefits of technology

[0015]本申请利用声纹具有一定的能量强度、时间上的连续性以及频谱中的显著峰值的这三个特征,可以从由音频转化成的梅尔谱中将声纹提取出来,然后再利用打分规则找到每个时间段中频率最低的声纹,再将每个时间段中频率最低的声纹连接起来即为基频,本申请的方法较传统的时域法和频域法准确性高、周期短,且相较机器学习法,因不需要准备数据集和采用复杂模型并进行训练,因此本申请的方法更准确且周期更短。因此本申请的方法可以快速且准确地识别出一段音频中的基频。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799883A_ABST
    Figure CN122799883A_ABST
Patent Text Reader

Abstract

The application discloses a method and device for extracting an audio fundamental frequency based on a mel spectrum, equipment, a storage medium and a computer program product, comprising receiving an audio to be processed, adjusting the volume to a preset volume and converting it into a corresponding mel spectrum; according to a preset energy intensity threshold, marking the voiceprints in the mel spectrum where the energy intensity exceeds the preset energy intensity threshold and extracting them into the same time-frequency graph; according to a preset scoring rule, scoring each voiceprint at each time point in the time-frequency graph and calculating the integral of each voiceprint; comparing the integral of each voiceprint with a preset minimum integral threshold and eliminating the voiceprints in the time-frequency graph whose integral is lower than the preset minimum integral threshold; in the remaining several voiceprints in the time-frequency graph, taking the point on the voiceprint with the highest integral at each time point and connecting them according to the corresponding voiceprint trend. The application has high accuracy and short cycle in extracting the fundamental frequency in the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound processing technology, and more specifically, to a method, apparatus, device, storage medium, and computer program product for extracting audio fundamental frequencies based on Mel spectrum. Background Technology

[0002] Fundamental frequency (FFM) has wide applications in sound processing, such as identity authentication and emotion recognition. The accuracy and speed of FFM extraction directly affect the accuracy and speed of sound processing, making it a crucial step. Existing methods for FFM extraction primarily include time-domain methods, frequency-domain methods, and machine learning methods. Time-domain methods extract FFM by analyzing the time-domain characteristics of the speech signal. This method is sensitive to noise, requiring high measurement accuracy and low background noise levels, resulting in poor accuracy. Furthermore, it involves significant computational workload and a long processing time. Frequency-domain methods convert the speech signal from the time domain to the frequency domain using Fourier transform and then analyze spectral features to extract the FFM. This method suffers from poor accuracy and lacks universality. Machine learning methods utilize deep learning models for feature extraction and FFM prediction. This method relies on large amounts of training data, requiring complex models and lengthy training times. Moreover, the trained model is affected by data quality and labeling accuracy. Therefore, all existing methods for FFM extraction suffer from poor accuracy, high computational workload, and long processing times. Summary of the Invention

[0003] To address at least one aspect of the above problems, the present invention provides a method, apparatus, device, storage medium, and computer program product for extracting audio fundamental frequencies based on Mel spectrum.

[0004] In the first aspect, this application provides a method for extracting audio fundamental frequencies based on Mel spectrum, comprising the following steps: a receiving step: receiving the audio to be processed; a preprocessing step: adjusting the volume of the audio to be processed to a preset volume; a first conversion step: converting the adjusted audio into a corresponding Mel spectrum; a first analysis step: marking the sound signatures in the Mel spectrum where the energy intensity exceeds a preset energy intensity threshold, based on a preset energy intensity threshold; a second conversion step: extracting several marked sound signatures and plotting them onto the same time-frequency spectrum; and a scoring step: scoring the time-frequency spectrum according to preset scoring rules. Each voiceprint at each time point in the time-frequency spectrum is scored, with the preset scoring rule being to score based on frequency at the same time point, giving higher scores to voiceprints at lower frequencies; Integration step: Calculate the integral for each voiceprint; Second analysis step: Compare the integral of each voiceprint with a preset minimum integration threshold. If the integral of a voiceprint is lower than the preset minimum integration threshold, that voiceprint is removed from the time-frequency spectrum; Fundamental frequency determination step: Among the remaining voiceprints in the time-frequency spectrum, select the point on the voiceprint with the highest integral at each time point and connect them according to the trend of the corresponding voiceprints.

[0005] Preferably, in the second conversion step, after extracting the marked voiceprints, the voiceprints with only single-point data are first removed, and then the remaining voiceprints are plotted onto the same time-frequency spectrum.

[0006] Preferably, the preset scoring rule is that at the same time point, the voiceprint at the lowest frequency scores 'a' points, the voiceprint at the second lowest frequency scores 'b' points, and the other voiceprints score 'c' points, where a is greater than b and b is greater than c.

[0007] More preferably, the preset scoring rule is that at the same time point, the voiceprint at the lowest frequency scores 2 points, the voiceprint at the second lowest frequency scores 1 point, and the voiceprint at other frequencies scores 0 points.

[0008] Preferably, the first analysis step further includes numbering the marked voiceprints.

[0009] Preferably, the preset energy intensity threshold is 50% to 90% of the highest energy intensity value in the Mel spectrum.

[0010] Secondly, this application provides an apparatus for extracting the fundamental frequency of audio based on Mel spectrum. The apparatus includes: a receiving module configured to receive audio to be processed; a preprocessing module configured to adjust the volume of the audio to be processed to a preset volume according to a preset volume; a first conversion module configured to convert the adjusted audio into a corresponding Mel spectrum; a first analysis module configured to mark the sound signatures in the Mel spectrum where the energy intensity exceeds a preset energy intensity threshold according to a preset energy intensity threshold; a second conversion module configured to extract several marked sound signatures and plot them on the same time-frequency spectrum; and a scoring module configured to score according to a preset scoring rule. Then, each voiceprint at each time point in the time-frequency spectrum is scored, with the preset scoring rule being to score based on frequency at the same time point, giving higher scores to voiceprints at lower frequencies; an integration module is configured to calculate the integral of each voiceprint; a second analysis module is configured to compare the integral of each voiceprint with a preset minimum integration threshold, and remove a voiceprint from the time-frequency spectrum when its integral is lower than the preset minimum integration threshold; a fundamental frequency determination module is configured to select the point on the voiceprint with the highest integral at each time point from the remaining voiceprints in the time-frequency spectrum, and connect them according to the trend of the corresponding voiceprints.

[0011] Thirdly, this application provides a device for extracting audio fundamental frequencies based on Mel spectrum, the device including a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method for extracting audio fundamental frequencies based on Mel spectrum as described above.

[0012] Fourthly, this application provides a storage medium storing computer-readable instructions that, when executed by a processor, perform the method according to any one of the preceding descriptions.

[0013] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0014] The present invention provides a method, apparatus, device, storage medium, and computer program product for extracting audio fundamental frequencies based on Mel spectrum, which has the following beneficial effects:

[0015] This application utilizes three characteristics of voiceprints: a certain energy intensity, temporal continuity, and significant peaks in the spectrum. Voiceprints can be extracted from the Mel spectrum converted from audio. Then, a scoring rule is used to find the lowest frequency voiceprint in each time segment. Connecting the lowest frequency voiceprints in each time segment yields the fundamental frequency. This method is more accurate and has a shorter cycle time than traditional time-domain and frequency-domain methods. Furthermore, compared to machine learning methods, it does not require preparing a dataset or using complex models for training. Therefore, this method is more accurate and has a shorter cycle time. Thus, this method can quickly and accurately identify the fundamental frequency in an audio segment. Attached Figure Description

[0016] To better understand the above and other objects, features, advantages, and functions of the present invention, reference can be made to the embodiments shown in the accompanying drawings. The same reference numerals in the drawings refer to the same parts. Those skilled in the art should understand that the drawings are intended to schematically illustrate preferred embodiments of the invention and are not intended to limit the scope of the invention; the parts in the drawings are not drawn to scale.

[0017] Figure 1 A flowchart of a method for extracting audio fundamental frequencies based on Mel spectrum according to an embodiment of the present invention is shown;

[0018] Figure 2 A schematic diagram of the first analysis step in a method for extracting audio fundamental frequencies based on Mel spectrum according to an embodiment of the present invention is shown;

[0019] Figure 3 A schematic diagram of the second conversion step in a method for extracting audio fundamental frequencies based on Mel spectrum according to an embodiment of the present invention is shown;

[0020] Figure 4 A schematic diagram of the second analysis step in a method for extracting audio fundamental frequencies based on Mel spectrum according to an embodiment of the present invention is shown;

[0021] Figure 5 A schematic diagram is shown of the fundamental frequency determination step in a method for extracting audio fundamental frequency based on Mel spectrum according to an embodiment of the present invention;

[0022] Figure 6 A block diagram of an apparatus for extracting audio fundamental frequencies based on Mel spectrum according to an embodiment of the present invention is shown. Detailed Implementation

[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0024] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] To at least partially address one or more of the aforementioned problems and other potential issues, embodiments of this disclosure propose a method for extracting audio fundamental frequencies based on Mel-spectrum analysis, such as... Figure 1 As shown, it includes the following steps:

[0026] Receiving steps: Receive the audio to be processed, specifically, receive the audio that needs to be extracted from the receiver, microphone, or sound acquisition device.

[0027] Preprocessing steps: Adjust the volume of the audio to be processed to the preset volume according to the preset volume. Specifically, adjust the volume of the audio to be processed according to the preset volume type. That is, when the preset volume type is the maximum preset volume, the maximum volume of the audio to be processed is adjusted to the maximum preset volume in this step; when the preset volume type is the average preset volume, the average volume of the audio to be processed is adjusted to the average preset volume in this step.

[0028] The first conversion step: convert the adjusted audio into the corresponding Mel spectrum, such as... Figure 2 As shown, the technique for converting audio into a me-spectrum is existing technology.

[0029] First analysis step: Based on a preset energy intensity threshold, mark the voiceprints in the Mel spectrum where the energy intensity exceeds the preset energy intensity threshold, such as... Figure 2 The red portion in the image is used, where the preset energy intensity threshold is preferably 50% to 90% of the highest energy intensity value in the Mel spectrum. Preferably, this step also includes numbering the marked voiceprints, and these voiceprints are always processed with the corresponding numbers, which helps to improve the accuracy of voiceprint processing.

[0030] The second conversion step: Extract the marked voiceprints and plot them onto the same time-frequency spectrum, such as... Figure 3 As shown; preferably, after extracting the marked voiceprints, the voiceprints with only single-point data are first removed, and then the remaining voiceprints are plotted on the same time-frequency spectrum. Since the voiceprints with only single-point data are noise, removing such voiceprints in this step helps to reduce the interference of noise on the extraction results and also helps to reduce the amount of subsequent calculations.

[0031] Scoring Steps: According to the preset scoring rules, each voiceprint at each time point in the time-frequency spectrum is scored. The preset scoring rules are to score according to frequency at the same time point, with lower frequency voiceprints receiving higher scores. That is, the scores are gradually increased from high to low frequency. Alternatively, since the goal of scoring is to find the lowest frequency voiceprint in each time period, only 2 to 6 voiceprints remain for comparison in the later stages of each time period. Therefore, the number of voiceprints to be compared in the later stages can be determined first. Then, the low-frequency voiceprints of that number of voiceprints are assigned scores gradually increasing from high to low frequency, and the remaining voiceprints are given the same score. For example, if the number of voiceprints to be compared in the later stages is 3, the preset scoring rules are to give the highest score to the voiceprint at the lowest frequency, the second highest score to the voiceprint at the second lowest frequency, the third highest score to the voiceprint at the third lowest frequency, and so on. The preset scoring rules are designed based on the number of voiceprints to be compared in the later stages. Preferably, the number of voiceprints required for later comparison is 2. The preset scoring rule is that at the same time point, the voiceprint at the lowest frequency scores 'a' points, the voiceprint at the second lowest frequency scores 'b' points, and the other voiceprints score 'c' points, where a is greater than b and b is greater than c. More preferably, the preset scoring rule is that at the same time point, the voiceprint at the lowest frequency scores 2 points, the voiceprint at the second lowest frequency scores 1 point, and the other voiceprints score 0 points.

[0032] Integration step: Calculate the integral for each voiceprint, which is the sum of the scores for each time point on each voiceprint in the previous step.

[0033] The second analysis step: Compare the integral of each voiceprint with a preset minimum integration threshold. When the integral of a voiceprint is lower than the preset minimum integration threshold, remove that voiceprint from the time-frequency spectrum. Figure 4 As shown, this is to avoid noise disturbance.

[0034] The steps for determining the fundamental frequency are as follows: In the remaining acoustic signatures in the time-frequency spectrum, select the point on the acoustic signature with the highest integral at each time point, and connect them according to the trend of the corresponding acoustic signatures, such as... Figure 5 As shown.

[0035] This application also provides a device for extracting the fundamental frequency of an audio signal based on Mel spectrum, such as... Figure 6 As shown, the device includes: a receiving module configured to receive audio to be processed; a preprocessing module configured to adjust the volume of the audio to be processed to a preset volume according to a preset volume; a first conversion module configured to convert the adjusted audio into a corresponding Mel spectrum; a first analysis module configured to mark the voiceprints in the Mel spectrum where the energy intensity exceeds a preset energy intensity threshold according to a preset energy intensity threshold; a second conversion module configured to extract several marked voiceprints and plot them onto the same time-frequency spectrum; and a scoring module configured to score each voiceprint in the time-frequency spectrum according to preset scoring rules. Each voiceprint at a given time point is scored, with the preset scoring rule being to score based on frequency at the same time point, giving higher scores to voiceprints at lower frequencies; an integration module is configured to calculate the integral of each voiceprint; a second analysis module is configured to compare the integral of each voiceprint with a preset minimum integration threshold, and remove a voiceprint from the time-frequency spectrum if its integral is lower than the preset minimum integration threshold; a fundamental frequency determination module is configured to select the point on the voiceprint with the highest integral at each time point from the remaining voiceprints in the time-frequency spectrum, and connect them according to the trend of the corresponding voiceprints.

[0036] This application also provides a device for extracting audio fundamental frequencies based on Mel spectrum, the device including a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method for extracting audio fundamental frequencies based on Mel spectrum as described above.

[0037] This application also provides a storage medium storing computer-readable instructions that, when executed by a processor, perform the method according to any one of the preceding descriptions.

[0038] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0039] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand this document.

Claims

1. A method for extracting audio fundamental frequencies based on Mel spectrum, characterized in that: Includes the following steps: Receiving steps: Receive the audio to be processed; Preprocessing steps: Adjust the volume of the audio to be processed to the preset volume. First conversion step: Convert the adjusted audio into the corresponding Mel spectrum; First analysis step: Based on the preset energy intensity threshold, mark the voiceprints in the Mel spectrum where the energy intensity exceeds the preset energy intensity threshold; The second conversion step is to extract the marked voiceprints and plot them onto the same time-frequency spectrum. Scoring steps: According to the preset scoring rules, each voiceprint at each time point in the time-frequency spectrum is scored. The preset scoring rules are to score according to the frequency at the same time point, with voiceprints at lower frequencies receiving higher scores. Integration steps: Calculate the integral over each voiceprint; The second analysis step is to compare the integral of each voiceprint with the preset minimum integration threshold. When the integral of a voiceprint is lower than the preset minimum integration threshold, the voiceprint is removed from the time-frequency spectrum. The fundamental frequency determination step is as follows: In the remaining several voiceprints in the time-frequency spectrum, take the point on the voiceprint with the highest integral at each time point, and connect them according to the trend of the corresponding voiceprint.

2. The method for extracting audio fundamental frequencies based on Mel spectrum according to claim 1, characterized in that: In the second conversion step, after extracting the marked voiceprints, the voiceprints with only single-point data are first removed, and then the remaining voiceprints are plotted on the same time-frequency spectrum.

3. The method for extracting audio fundamental frequencies based on Mel spectrum according to claim 1, characterized in that: The preset scoring rule is that at the same time point, the voiceprint at the lowest frequency scores 'a' points, the voiceprint at the second lowest frequency scores 'b' points, and the other voiceprints score 'c' points, where a is greater than b and b is greater than c.

4. The method for extracting audio fundamental frequencies based on Mel spectrum according to claim 3, characterized in that: The preset scoring rule is that at the same time point, the voiceprint at the lowest frequency scores 2 points, the voiceprint at the second lowest frequency scores 1 point, and the voiceprint at other frequencies scores 0 points.

5. The method for extracting audio fundamental frequencies based on Mel spectrum according to claim 1, characterized in that: The first analysis step also includes numbering the marked voiceprints.

6. The method for extracting audio fundamental frequencies based on Mel spectrum according to claim 1, characterized in that: The preset energy intensity threshold is 50% to 90% of the highest energy intensity value in the Mel spectrum.

7. A device for extracting the fundamental frequency of an audio signal based on Mel spectrum, characterized in that: The device includes: The receiving module is configured to receive audio to be processed. The preprocessing module is configured to adjust the volume of the audio to be processed to a preset volume based on the preset volume. The first conversion module is configured to convert the adjusted audio into the corresponding Mel spectrum. The first analysis module is configured to mark the voiceprints in the Mel spectrum where the energy intensity exceeds the preset energy intensity threshold, based on the preset energy intensity threshold. The second conversion module is configured to extract several marked voiceprints and plot them onto the same time-frequency spectrum. The scoring module is configured to score each voiceprint at each time point in the time-frequency spectrum according to a preset scoring rule. The preset scoring rule is to score according to the frequency at the same time point, with voiceprints at lower frequencies receiving higher scores. The integration module is configured to calculate the integral over each voiceprint. The second analysis module is configured to compare the integral of each voiceprint with a preset minimum integration threshold. When the integral of a voiceprint is lower than the preset minimum integration threshold, the voiceprint is removed from the time-frequency spectrum. Determine the fundamental frequency module and configure it to select the point on the voiceprint with the highest integral at each time point from the remaining voiceprints in the time-frequency spectrum, and connect them according to the trend of the corresponding voiceprints.

8. A device for extracting audio fundamental frequencies based on Mel spectrum, characterized in that: The device includes a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements a method for extracting audio fundamental frequencies based on Mel spectrum as described in any one of claims 1 to 6.

9. A storage medium, characterized in that: The device stores computer-readable instructions that, when executed by a processor, perform the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.