Audio data processing method, audio data processing apparatus, and storage medium

By identifying the characteristic parameters of human voice and accompaniment sound in the audio signal and dynamically adjusting the harmonic generation control parameters, the problem of reduced vocal clarity after bass enhancement is solved, and high-quality bass enhancement of audio data is achieved.

CN116095561BActive Publication Date: 2025-10-24NANJING GOERTEK ACOUSTICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310132337.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-17
Publication Date
2025-10-24
Estimated Expiration
2043-02-17

AI Technical Summary

Technical Problem

In the prior art, the harmonic control parameters of the low-frequency signal are fixed, resulting in reduced clarity of the human voice after the bass of the audio data is enhanced.

Method used

By identifying the characteristic parameters of human voice and accompaniment sound in the audio signal, the control parameters of harmonic generation are dynamically adjusted to generate enhanced harmonic signals to replace low-frequency signals.

Benefits of technology

Improves the clarity of human voices during bass enhancement, ensuring the quality of audio data playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116095561B_ABST
    Figure CN116095561B_ABST
Patent Text Reader

Abstract

The application discloses a processing method and device of audio data and a storage medium. The method comprises the following steps: acquiring an audio signal; extracting a low-frequency signal in the audio signal; identifying a first characteristic parameter of a human voice signal and a second characteristic parameter of an accompaniment sound signal in the audio signal, and determining a control parameter for controlling harmonic generation according to the first characteristic parameter and the second characteristic parameter; and generating the enhanced harmonic signal in the low-frequency signal according to the control parameter. The application aims to realize bass enhancement while improving the intelligibility of human voice in the played sound.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent audio, in particular to an audio data processing method, an audio data processing device and a storage medium. BACKGROUND

[0002] In the data processing process of a sound playing device (especially a device with a small speaker volume), a bass enhancement algorithm is generally used to process an audio signal, and low-frequency components in the audio signal are replaced by high-order harmonics generated by the low-frequency components.

[0003] However, the control parameters for forming harmonics in low-frequency signals are generally fixed parameters set in advance, which can cause the intelligibility of human voices in the sound played after bass enhancement of some audio data to be reduced. SUMMARY

[0004] The main purpose of the present application is to provide an audio data processing method, an audio data processing device and a storage medium, which can realize bass enhancement while improving the intelligibility of human voices in the played sound.

[0005] To achieve the above purpose, the present application provides an audio data processing method, which comprises the following steps:

[0006] obtaining an audio signal;

[0007] extracting a low-frequency signal in the audio signal;

[0008] identifying a first characteristic parameter of a human voice signal and a second characteristic parameter of an accompaniment sound signal in the audio signal, and determining a control parameter for controlling harmonic generation according to the first characteristic parameter and the second characteristic parameter;

[0009] generating the enhanced harmonic signal in the low-frequency signal according to the control parameter.

[0010] Optionally, the first characteristic parameter comprises a first proportion of the human voice signal in the audio signal, the second characteristic parameter comprises a second proportion of the accompaniment sound signal in the audio signal, and the step of determining the control parameter for controlling harmonic generation according to the first characteristic parameter and the second characteristic parameter comprises:

[0011] determining a harmonic generation amount and / or a harmonic generation proportion according to the first proportion and the second proportion, and the control parameter comprises the harmonic generation amount and / or the harmonic generation proportion.

[0012] Optionally, the harmonic generation amount is negatively correlated with the first proportion, and / or the harmonic generation proportion is negatively correlated with the first proportion.

[0013] The harmonic generation amount is positively correlated with the second proportion, and / or the harmonic generation proportion is positively correlated with the first proportion.

[0014] Optionally, the step of determining the harmonic generation amount and / or the harmonic generation proportion according to the first proportion and the second proportion comprises:

[0015] When the first proportion is greater than a first preset proportion and the second proportion is less than a second preset proportion, a first harmonic amount is determined as the harmonic generation amount, and / or a first proportion is determined as the harmonic generation proportion;

[0016] When the first proportion and the second proportion are both less than or equal to the first preset proportion and both greater than or equal to the second preset proportion, a second harmonic amount is determined as the harmonic generation amount, and / or a second proportion is determined as the harmonic generation proportion;

[0017] When the second proportion is greater than the first preset proportion and the first proportion is less than the second preset proportion, a third harmonic amount is determined as the harmonic generation amount, and / or a third proportion is determined as the harmonic generation proportion;

[0018] The first harmonic amount is less than the second harmonic amount, and the second harmonic amount is less than the third harmonic amount; the first proportion is less than the second proportion, and the second proportion is less than the third proportion.

[0019] Optionally, before the step of determining the second harmonic amount as the harmonic generation amount and / or determining the second proportion as the harmonic generation proportion, the method further comprises:

[0020] When the first proportion and the second proportion are both less than or equal to the first preset proportion and both greater than or equal to the second preset proportion, a relationship value of the first proportion and the second proportion is determined;

[0021] The third harmonic amount is adjusted according to the relationship value to obtain the second harmonic amount, and / or the third proportion is adjusted according to the relationship value to obtain the second proportion.

[0022] Optionally, before the step of determining the third harmonic amount as the harmonic generation amount and / or determining the third proportion as the harmonic generation proportion, the method further comprises:

[0023] The frequency amplitude of the accompaniment sound signal in a preset frequency band is identified;

[0024] The third harmonic amount and / or the third proportion are determined according to the frequency amplitude.

[0025] Optionally, the step of identifying the first characteristic parameter of the vocal sound signal and the second characteristic parameter of the accompaniment sound signal in the audio signal comprises:

[0026] extracting frequency domain features and time domain features in the audio signal;

[0027] determining the first feature parameter and the second feature parameter according to the frequency domain features and the time domain features.

[0028] Optionally, the time domain features include at least one of the following parameters: short-time energy, loudness, glottal excitation pulse;

[0029] The frequency domain features include at least one of the following parameters: cepstrum coefficient, spectral entropy, line spectrum pair.

[0030] In addition, in order to achieve the above object, the present application further provides an audio data processing device, which comprises a memory, a processor and an audio data processing program stored in the memory and executable on the processor, and the audio data processing program implements the steps of the audio data processing method according to any one of the above when executed by the processor.

[0031] In addition, in order to achieve the above object, the present application further provides a storage medium, which stores an audio data processing program, and the audio data processing program implements the steps of the audio data processing method according to any one of the above when executed by a processor.

[0032] The audio data processing method provided by the present application adjusts and controls the enhanced harmonic signal generated in the low-frequency signal extracted from the audio signal based on the first feature parameter of the vocal signal and the second feature parameter of the accompaniment signal identified from the audio signal. Therefore, the control parameter of the harmonic generation in the bass enhancement process is no longer determined by the fixed parameter set in advance, but is determined in combination with the features of the vocal and accompaniment in the audio signal, thereby effectively improving the vocal clarity after the audio data after bass enhancement is played. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 The hardware structure diagram related to the execution of an embodiment of the audio data processing device of the present application;

[0034] Figure 2 The flowchart of an embodiment of the audio data processing method of the present application;

[0035] Figure 3 The flowchart of another embodiment of the audio data processing method of the present application;

[0036] Figure 4 The flowchart of still another embodiment of the audio data processing method of the present application;

[0037] Figure 5FIG. 4 is a flow chart of another embodiment of the method for processing audio data according to the present invention.

[0038] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0039] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0040] An embodiment of the present invention provides an audio data processing device 1 .

[0041] In this embodiment, the audio data processing device 1 can be built into any device with an audio playback function. For example, the device may include, but is not limited to, mobile phones, head-mounted display devices (such as virtual reality devices or augmented reality devices), smart audio glasses, neck-mounted speakers, open-back headphones, speakers, tablet computers, televisions, and other products.

[0042] In this embodiment, the volume of the speaker in the device where the audio data processing apparatus 1 is located is smaller than a preset volume, and the cutoff frequency of the speaker of this volume is greater than a preset frequency.

[0043] In the embodiment of the present invention, referring to Figure 1 The audio data processing device 1 includes a processor 1001 (e.g., a CPU), a memory 1002, a timer 1003, and the like. The various components of the control device are connected via a communication bus. The memory 1002 can be a high-speed RAM memory or a non-volatile memory such as a disk drive. Alternatively, the memory 1002 can be a storage device independent of the processor 1001.

[0044] Those skilled in the art will understand that Figure 1 The device structure shown in the figure does not constitute a limitation of the device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0045] like Figure 1 As shown, the memory 1002 as a storage medium may include a processing program for audio data. Figure 1 In the device shown, the processor 1001 can be used to call the audio data processing program stored in the memory 1002 and execute the relevant steps of the audio data processing method in the following embodiments.

[0046] An embodiment of the present invention further provides an audio data processing method, which is applied to the above-mentioned audio data processing device.

[0047] Reference Figure 2In an embodiment of the method for processing audio data, the method comprises:

[0048] In step S10, an audio signal is acquired.

[0049] The audio signal can be read from a local memory, acquired in real time through a microphone, or acquired based on a network connection.

[0050] In step S20, a low-frequency signal in the audio signal is extracted.

[0051] Specifically, a cutoff frequency of a loudspeaker used to play the audio signal can be acquired, and a signal lower than the cutoff frequency in the audio signal can be taken as the low-frequency signal.

[0052] In step S30, a first characteristic parameter of a vocal signal and a second characteristic parameter of an accompaniment signal in the audio signal are identified, and a control parameter used to control harmonic generation is determined according to the first characteristic parameter and the second characteristic parameter.

[0053] The vocal signal can be any vocal signal, or a signal corresponding to a vocal signal that meets a preset condition (for example, a vocal signal corresponding to a person with a specified identity).

[0054] The accompaniment signal can be all signals other than the vocal signal in the audio signal. Alternatively, the accompaniment signal can be a signal that meets a preset melody condition in the audio signal.

[0055] Specifically, a vocal recognition model and an accompaniment recognition model can be preset, the audio signal can be input into the vocal recognition model, and an identification result output by the vocal recognition model can be taken as the first characteristic parameter; the audio signal can be input into the accompaniment recognition model, and an identification result output by the accompaniment recognition model can be taken as the second characteristic parameter. The vocal recognition model and the accompaniment recognition model can be machine learning models, wherein, in a model training process, the vocal recognition model can be adjusted according to a first error existing when the accompaniment recognition model identifies standard data including a vocal signal and an accompaniment signal, and the accompaniment recognition model can be adjusted according to a second error existing when the vocal recognition model identifies the standard data, thereby facilitating improvement of identification accuracy.

[0056] The first characteristic parameter can be any parameter (for example, a proportion, a number of vocal signals, and / or a frequency amplitude) representing a feature of the vocal signal in the audio signal. The second characteristic parameter can be any parameter (for example, a proportion, a type of sound, and / or a frequency amplitude) representing a feature of the accompaniment signal in the audio signal.

[0057] The control parameter is any parameter used to regulate the gain of the harmonic signal in the low-frequency signal. In this embodiment, the control parameter can include the harmonic generation amount and / or the harmonic generation ratio, etc.

[0058] The different first characteristic parameters and different second characteristic parameters correspond to different control parameters. Specifically, a corresponding relationship between the first characteristic parameters, the second characteristic parameters, and the control parameters can be established in advance. The corresponding relationship can include a mapping relationship, a calculation formula, etc. Based on the corresponding relationship, the control parameter can be obtained by looking up the table through the first characteristic parameter and the second characteristic parameter; the control parameter can also be calculated by substituting the first characteristic parameter and the second characteristic parameter into the formula, etc.

[0059] In step S40, the enhanced harmonic signal is generated in the low-frequency signal according to the control parameter.

[0060] In this embodiment, the enhanced harmonic signal is specifically a high-order harmonic signal.

[0061] Here, the enhanced harmonic signal is generated in the low-frequency signal according to the control parameter, specifically referring to generating the enhanced harmonic signal generated by the low-frequency signal according to the control parameter, and replacing the original low-frequency signal component in the audio signal with the enhanced harmonic signal.

[0062] Wherein, when the control parameter includes the harmonic generation amount, the initial harmonic is generated from the low-frequency signal, and the initial harmonic signal corresponding to the harmonic generation amount is generated as the enhanced harmonic signal, and the original low-frequency signal in the audio signal is replaced with the enhanced harmonic signal. When the control parameter includes the harmonic generation ratio, the initial harmonic is generated from the low-frequency signal, and the enhanced harmonic signal is obtained by amplifying the initial harmonic according to the harmonic generation ratio, and the original low-frequency signal in the audio signal is replaced with the enhanced harmonic signal.

[0063] Wherein, the original low-frequency signal in the audio signal can be output through a loudspeaker after being replaced with the enhanced harmonic signal.

[0064] The method for processing audio data proposed in the embodiment of the application regulates the enhanced harmonic signal generated in the low-frequency signal extracted from the audio signal based on the first characteristic parameter of the vocal signal and the second characteristic parameter of the accompaniment signal identified in the audio signal. Therefore, the control parameter of the harmonic generation in the bass enhancement process is no longer determined by the fixed parameter set in advance, but is determined in combination with the characteristics of the vocal and the accompaniment in the audio signal, thereby effectively improving the vocal clarity after the audio data after the bass enhancement is played.

[0065] Further, based on the above embodiment, another embodiment of the method for processing audio data of the application is proposed. In this embodiment, referring to Figure 3The first feature parameter comprises a first proportion of the vocal signal in the audio signal, and the second feature parameter comprises a second proportion of the accompaniment signal in the audio signal. The step of determining the control parameter for controlling the harmonic generation according to the first feature parameter and the second feature parameter comprises determining a harmonic generation amount and / or a harmonic generation proportion according to the first proportion and the second proportion. The control parameter comprises the harmonic generation amount and / or the harmonic generation proportion. Based on this, step S30 comprises: step S31, identifying the first feature parameter of the vocal signal and the second feature parameter of the accompaniment signal in the audio signal, and determining the harmonic generation amount and / or the harmonic generation proportion according to the first proportion and the second proportion. The control parameter comprises the harmonic generation amount and / or the harmonic generation proportion. Step S40 comprises: step S41, generating the enhanced harmonic signal in the low-frequency signal according to the harmonic generation amount and / or the harmonic generation proportion.

[0066] The harmonic generation amount is specifically a target number of harmonics to be generated. The harmonic generation proportion is specifically an amplification proportion of the harmonics generated from the low-frequency signal.

[0067] Specifically, a relationship value of the first proportion and the second proportion can be determined, and the harmonic generation amount and / or the harmonic generation proportion can be determined according to the relationship value. In addition, a first interval in which the first proportion is located and a second interval in which the second proportion is located can also be determined, and the harmonic generation amount and / or the harmonic generation proportion can be determined according to the first interval and the second interval.

[0068] Different first proportions and different second proportions correspond to different harmonic generation amounts. Different first proportions and different second proportions correspond to different harmonic generation proportions. Based on this, a first corresponding relationship (such as a mapping relationship, a formula, etc.) between the first proportion, the second proportion, and the harmonic generation amount can be established in advance, and the harmonic generation amount corresponding to the first proportion and the second proportion can be determined based on the first corresponding relationship. A second corresponding relationship (such as a mapping relationship, a formula, etc.) between the first proportion, the second proportion, and the harmonic generation proportion can be established in advance, and the harmonic generation proportion corresponding to the first proportion and the second proportion can be determined based on the second corresponding relationship. The harmonic generation amount is negatively correlated with the first proportion, and / or the harmonic generation proportion is negatively correlated with the first proportion. The harmonic generation amount is positively correlated with the second proportion, and / or the harmonic generation proportion is positively correlated with the first proportion. That is, the greater the proportion of the vocal signal, the smaller the harmonic generation amount and / or the harmonic generation proportion, and the greater the proportion of the accompaniment signal, the greater the harmonic generation amount and / or the harmonic generation proportion. Based on this, the vocal clarity of the sound played after the bass enhancement can be effectively improved.

[0069] Further, more than one second corresponding relationship can be pre-set, one of the more than one second corresponding relationship is determined as a target corresponding relationship according to the harmonic generation amount, different harmonic generation amounts correspond to different target corresponding relationships, and the first proportion and the second proportion correspond to the harmonic generation proportion are determined based on the target corresponding relationship.

[0070] In the embodiment, the first proportion and the second proportion can accurately reflect the influence of the bass enhancement on the vocal clarity of the audio signal, and the harmonic generation amount and the harmonic generation proportion can effectively adjust the quality of the vocal after the bass enhancement, so that the harmonic generation proportion and / or the harmonic generation amount are determined in combination with the first proportion and the second proportion, which can ensure that the bass enhancement is performed without excessively weakening the vocal clarity, thereby effectively improving the vocal clarity after the bass-enhanced audio data is played.

[0071] Further, based on any of the above embodiments, another embodiment of the audio data processing method of the application is provided. In this embodiment, referring to Figure 4 , the harmonic generation amount and / or the harmonic generation proportion are determined according to the first proportion and the second proportion, including:

[0072] In step S311, when the first proportion is greater than a first preset proportion and the second proportion is less than a second preset proportion, a first harmonic amount is determined as the harmonic generation amount and / or a first proportion is determined as the harmonic generation proportion.

[0073] In step S312, when the first proportion and the second proportion are both less than or equal to the first preset proportion and both greater than or equal to the second preset proportion, a second harmonic amount is determined as the harmonic generation amount and / or a second proportion is determined as the harmonic generation proportion.

[0074] In step S313, when the second proportion is greater than the first preset proportion and the first proportion is less than the second preset proportion, a third harmonic amount is determined as the harmonic generation amount and / or a third proportion is determined as the harmonic generation proportion.

[0075] The first harmonic amount is less than the second harmonic amount, and the second harmonic amount is less than the third harmonic amount. The first proportion is less than the second proportion, and the second proportion is less than the third proportion.

[0076] The first preset proportion and the second preset proportion can be pre-set fixed parameters, or can be determined according to the audio type corresponding to the audio signal and the relationship value between the first proportion and the second proportion. For example, the first preset proportion is 90%, and the second preset proportion is 10%. Alternatively, the first preset proportion and the second preset proportion can be set to other values according to actual conditions.

[0077] The first harmonic amount, the second harmonic amount, the third harmonic amount, the first ratio, the second ratio, and / or the third ratio can be preset fixed values, or can be values determined according to the audio signal, the first signal, and / or the second signal.

[0078] For example, when the first proportion is 100% and the second proportion is 0%, the first harmonic amount is determined to be the harmonic generation amount, and the first ratio is determined to be the harmonic generation ratio; when the first proportion is 0% and the second proportion is 100%, the third harmonic amount is determined to be the harmonic generation amount, and the third ratio is determined to be the harmonic generation ratio; and when the first proportion and the second proportion are between 0% and 100%, the second harmonic amount can be determined to be the harmonic generation amount, and the second ratio can be determined to be the harmonic generation ratio.

[0079] In this embodiment, different harmonic generation amounts and / or harmonic generation ratios are used as control parameters according to different proportion intervals of the first proportion and the second proportion, which can effectively reduce the influence of the identification error of the first proportion and the second proportion on the accuracy of the subsequent control parameters, thereby improving the accuracy of the control parameters, and further improving the vocal clarity after the bass-enhanced audio data is played.

[0080] Further, in this embodiment, before the step of determining the second harmonic amount to be the harmonic generation amount and / or determining the second ratio to be the harmonic generation ratio, the method further includes: when the first proportion and the second proportion are both less than or equal to a first preset proportion and both greater than or equal to a second preset proportion, determining a relationship value of the first proportion and the second proportion; adjusting the third harmonic amount according to the relationship value to obtain the second harmonic amount; and / or adjusting the third ratio according to the relationship value to obtain the second ratio.

[0081] The relationship value can include a difference value and / or a ratio value. Specifically, a first adjustment amplitude or a first adjustment ratio of the third harmonic amount can be determined according to the relationship value, and the second harmonic amount can be obtained by reducing the third harmonic amount according to the first adjustment amplitude or the first adjustment ratio; in addition, a second adjustment amplitude or a second adjustment ratio of the third ratio can be determined according to the relationship value, and the second ratio can be obtained by reducing the third ratio according to the second adjustment amplitude or the second adjustment ratio.

[0082] Specifically, in this embodiment, the relationship value is a ratio of the first proportion to the second proportion, the second harmonic amount is negatively correlated with the ratio, and the second ratio is negatively correlated with the ratio, that is, the smaller the ratio, the greater the proportion of the accompaniment sound relative to the vocal, and the corresponding second harmonic amount and / or second ratio can be greater.

[0083] Further, the first harmonic amount can be determined according to the second harmonic amount, and the first ratio can be determined according to the second ratio.

[0084] In the embodiment, the second harmonic quantity and / or the second ratio are adjusted based on the relationship value of the first ratio and the second ratio, so as to further ensure the accuracy of the second harmonic quantity and / or the second ratio determined when the deviation between the first ratio and the second ratio is not too large, so as to further realize the low-frequency enhancement and improve the clarity of the vocal.

[0085] Further, in the embodiment, before the step of determining the third harmonic quantity as the harmonic generation quantity and / or determining the third ratio as the harmonic generation ratio, the method further comprises: identifying a frequency amplitude of the accompaniment signal in a preset frequency band; and determining the third harmonic quantity and / or the third ratio according to the frequency amplitude.

[0086] The preset frequency band is a frequency band in a frequency range less than the cutoff frequency. The maximum value of the preset frequency band is less than the cutoff frequency.

[0087] Different frequency amplitudes correspond to different third harmonic quantities and / or third ratios. The third harmonic quantity and / or the third ratio are positively correlated with the frequency amplitude. That is, the greater the frequency amplitude, the greater the third harmonic quantity and / or the third ratio. Specifically, the third harmonic quantity and / or the third ratio can be calculated by substituting the frequency amplitude into a formula; or, an interval in which the frequency amplitude is located can be determined, and the corresponding third harmonic quantity and / or third ratio can be determined based on the interval.

[0088] It should be noted that the third harmonic quantity and / or the third ratio can be determined in the above manner to determine the second harmonic quantity and / or the second ratio.

[0089] In the embodiment, the above manner is used to improve the accuracy of the third harmonic quantity and / or the third ratio, so as to further realize the low-frequency enhancement and improve the clarity of the vocal.

[0090] Further, based on any of the above embodiments, another embodiment of a method for processing audio data is provided. In the embodiment, referring to Figure 5 , the step S20 comprises:

[0091] The step S21 comprises extracting frequency domain features and time domain features in the audio signal.

[0092] The time domain features comprise at least one of the following parameters: short-time energy, loudness, glottal excitation pulse. The frequency domain features comprise at least one of the following parameters: cepstrum coefficient, spectral entropy, line spectrum pair. The cepstrum coefficient can comprise one or more of linear prediction cepstrum coefficient, mel cepstrum coefficient, and first-order differential mel cepstrum coefficient.

[0093] Specifically, the audio signal can be divided into a plurality of data frames according to a preset rule, and time domain features of each data frame are extracted. In addition, after obtaining the plurality of data frames, each data frame can be subjected to windowing processing, the data frame subjected to the windowing processing is subjected to Fourier transform, frequency domain feature extraction is performed based on a result of the Fourier transform, and the frequency domain feature is obtained here.

[0094] In step S22, the first feature parameter and the second feature parameter are determined according to the frequency domain feature and the time domain feature.

[0095] The signal type in the audio signal is classified based on the frequency domain feature and the time domain feature to obtain a vocal signal and an accompaniment signal, the first feature parameter is determined according to a signal feature of the vocal signal and a signal feature of the audio signal, and the second feature parameter is determined according to a signal feature of the accompaniment signal and the signal feature of the audio signal.

[0096] In the embodiment, the first feature parameter and the second feature parameter are accurately analyzed through the frequency domain feature and the time domain feature extraction of the audio signal, so as to further improve the accuracy of the control parameter of the enhanced harmonic signal determined subsequently.

[0097] In addition, the embodiment of the present application further provides a storage medium, and the storage medium stores a processing program of audio data. The processing program of audio data is executed by a processor to realize the related steps of any one of the embodiments of the processing method of audio data.

[0098] It should be noted that, in this document, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or system. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of another identical element in the process, method, article or system including the element.

[0099] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0100] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, also can be through hardware, but in many cases the former is the better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art contribution can be embodied in the form of software product, the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disc, optical disc) as described above, including a number of instructions to make a terminal device (may be a mobile phone, computer, server, audio data processing device, or network equipment, etc.) executes the method described in various embodiments of the present application.

[0101] The above is only the preferred embodiment of the present application, not therefore limit the patent scope of the present application, any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An audio data processing method, characterized by, The audio data processing method comprises the following steps: obtaining an audio signal; extracting a low-frequency signal in the audio signal; identifying a first characteristic parameter of a human voice signal and a second characteristic parameter of an accompaniment sound signal in the audio signal, and determining a control parameter for controlling harmonic generation according to the first characteristic parameter and the second characteristic parameter; generating an enhanced harmonic signal in the low-frequency signal according to the control parameter, and replacing the original low-frequency signal component in the audio signal with the enhanced harmonic signal.

2. The audio data processing method of claim 1, wherein, The first characteristic parameter comprises a first proportion of the human voice signal in the audio signal, and the second characteristic parameter comprises a second proportion of the accompaniment sound signal in the audio signal, and the step of determining the control parameter for controlling harmonic generation according to the first characteristic parameter and the second characteristic parameter comprises: determining a harmonic generation amount and / or a harmonic generation proportion according to the first proportion and the second proportion, and the control parameter comprises the harmonic generation amount and / or the harmonic generation proportion.

3. The audio data processing method of claim 2, wherein, The harmonic generation amount is negatively correlated with the first proportion, and / or the harmonic generation proportion is negatively correlated with the first proportion; The harmonic generation amount is positively correlated with the second proportion, and / or the harmonic generation proportion is positively correlated with the first proportion.

4. The audio data processing method of claim 3, wherein, The step of determining the harmonic generation amount and / or the harmonic generation proportion according to the first proportion and the second proportion comprises: when the first proportion is greater than a first preset proportion and the second proportion is less than a second preset proportion, determining a first harmonic amount as the harmonic generation amount and / or determining a first proportion as the harmonic generation proportion; when the first proportion and the second proportion are both less than or equal to the first preset proportion and both greater than or equal to the second preset proportion, determining a second harmonic amount as the harmonic generation amount and / or determining a second proportion as the harmonic generation proportion; when the second proportion is greater than the first preset proportion and the first proportion is less than the second preset proportion, determining a third harmonic amount as the harmonic generation amount and / or determining a third proportion as the harmonic generation proportion; wherein the first harmonic amount is less than the second harmonic amount, and the second harmonic amount is less than the third harmonic amount; and the first proportion is less than the second proportion, and the second proportion is less than the third proportion.

5. The audio data processing method of claim 4, wherein, Before the step of determining the second harmonic amount as the harmonic generation amount and / or determining the second proportion as the harmonic generation proportion, the method further comprises: when the first proportion and the second proportion are both less than or equal to the first preset proportion and both greater than or equal to the second preset proportion, determining a relationship value of the first proportion and the second proportion; adjusting the third harmonic amount according to the relationship value to obtain the second harmonic amount; and / or adjusting the third proportion according to the relationship value to obtain the second proportion.

6. The audio data processing method of claim 4, wherein, Before the step of determining the third harmonic amount as the harmonic generation amount and / or determining the third proportion as the harmonic generation proportion, the method further comprises: identifying a frequency amplitude of the accompaniment sound signal in a preset frequency band; determining the third harmonic amount and / or the third proportion according to the frequency amplitude.

7. The audio data processing method of any one of claims 1 to 6, wherein, The step of identifying the first feature parameter of the human voice signal and the second feature parameter of the accompaniment sound signal in the audio signal comprises: extracting frequency domain features and time domain features in the audio signal; determining the first feature parameter and the second feature parameter according to the frequency domain features and the time domain features.

8. The audio data processing method of claim 7, wherein, The time domain features comprise at least one of the following parameters: short-time energy, loudness, glottal excitation pulse; The frequency domain features comprise at least one of the following parameters: cepstrum coefficient, spectral entropy, line spectrum pair.

9. An audio data processing apparatus, characterized by comprising: The audio data processing apparatus comprises a memory, a processor, and an audio data processing program stored in the memory and executable on the processor, and the audio data processing program, when executed by the processor, implements the steps of the audio data processing method according to any one of claims 1 to 8.

10. A storage medium, characterized by The storage medium stores an audio data processing program, and the audio data processing program, when executed by the processor, implements the steps of the audio data processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Mobile terminal and voice signal processing method thereof

    CN103594091A

  • Audio data processing method and electronic equipment

    CN114299976A