Method for obtaining audio adjustment strategy, computer device, and program product

The audio adjustment strategy enhances audio tuning accuracy by identifying key frequency points and applying gain coefficients based on amplitude ratios, addressing the inaccuracy of current vocal adjustment methods.

CN115862664BActive Publication Date: 2025-07-15TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211423674.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-07-15
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

Current audio adjustment strategies for singing vocals lack accuracy due to varying user musical understanding, leading to inadequate adjustment of song audio deficiencies.

Method used

An audio adjustment strategy that identifies key frequency points in vocal signals, determines adjustment frequencies based on amplitude ratios, and applies gain coefficients to enhance audio quality.

Benefits of technology

Improves the accuracy of audio adjustments by aligning frequency enhancements with user singing proficiency, resulting in more effective audio tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862664B_ABST
    Figure CN115862664B_ABST
Patent Text Reader

Abstract

The present application relates to a method for obtaining an audio adjustment strategy, a computer device, a storage medium, and a computer program product. After receiving an audio adjustment instruction, the fundamental frequency points and overtone frequency points of the human voice signal in the audio to be adjusted are obtained, the overtone frequency points to be adjusted are determined among the overtone frequency points, and according to the comparison result between the amplitude ratio between the fundamental frequency points and the overtone frequency points and a preset amplitude ratio, the amplitude gain coefficient of the frequency band to be adjusted composed of the overtone frequency points is determined, and the frequency band to be adjusted and the amplitude gain coefficient are determined as the audio amplitude adjustment strategy of the audio to be adjusted. Compared with the scheme of manually determining the strategy to adjust the audio, this scheme determines the amplitude gain coefficient based on the standard proportional relationship between the fundamental frequency and the overtone, and determines the frequency to be adjusted based on the comparison between the frequency range and the preset frequency, so as to determine the strategy for adjusting the audio, improving the accuracy of determining the audio adjustment strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and particularly to a method for obtaining an audio adjustment strategy, a computer device, a storage medium, and a computer program product. Background Art

[0002] With the development of computer technology, people can currently listen to songs through terminals such as mobile phones, and it has also become mainstream for users to record their singing using terminals. Since the singing levels of each user are different, in order to achieve better singing effects, it is necessary to adjust the corresponding equalization parameters of the user's singing audio. Currently, the strategy for adjusting the equalization parameters of singing audio is usually manually selected. However, the music understanding levels of each user are inconsistent, and it is impossible to discover the equalization defects in the song audio. By adjusting the equalization using the manually selected strategy, it is impossible to adjust the defective parts of the song audio, reducing the accuracy of song audio adjustment.

[0003] Therefore, the current audio adjustment strategy has the defect of low adjustment accuracy. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an audio adjustment strategy obtaining method, a computer device, a computer-readable storage medium, and a computer program product that can improve the adjustment accuracy.

[0005] In a first aspect, the present application provides an audio adjustment strategy obtaining method, and the method includes:

[0006] When an audio adjustment instruction for an audio to be adjusted is detected, obtain the fundamental frequency points and overtone frequency points corresponding to each frame of the human voice signal in the audio to be adjusted;

[0007] For each frame of the human voice signal, determine at least one overtone frequency point to be adjusted among the overtone frequency points corresponding to the human voice signal, and determine the frequency band to be adjusted according to the at least one overtone frequency point to be adjusted;

[0008] For each frame of the human voice signal, determine the amplitude gain coefficient corresponding to the frequency band to be adjusted according to the amplitude ratio between the at least one overtone frequency point to be adjusted and the fundamental frequency point;

[0009] Determine each frame of the human voice signal, the frequency band to be adjusted corresponding to each frame of the human voice signal, and the amplitude gain coefficient corresponding to the frequency band to be adjusted as the audio amplitude adjustment strategy corresponding to the audio to be adjusted.

[0010] In one embodiment, the determining at least one overtone frequency point to be adjusted among the overtone frequency points corresponding to the human voice signal includes:

[0011] Query the frequency to be adjusted in a preset frequency adjustment table according to the overtone frequency range of the overtone audio points of the human voice signal; the preset frequency adjustment table includes the frequencies of the overtone audio points that need to be frequency-adjusted.

[0012] Determine at least one overtone audio point corresponding to the frequency to be adjusted among the overtone audio points corresponding to the human voice signal as the overtone audio point to be adjusted.

[0013] In one embodiment, the querying the frequency to be adjusted in the preset frequency adjustment table according to the overtone frequency range of the overtone audio points of the human voice signal includes:

[0014] Determine the frequency range composed of the overtone audio points corresponding to the human voice signal as the overtone frequency range corresponding to the overtone audio points.

[0015] Obtain the first frequency to be adjusted within the overtone frequency range in the preset frequency adjustment table, and obtain the second frequency to be adjusted whose frequency difference from the minimum or maximum value of the overtone frequency range is within a preset frequency difference range.

[0016] Determine the first frequency to be adjusted and / or the second frequency to be adjusted as the frequency to be adjusted that needs to be adjusted.

[0017] In one embodiment, the determining the amplitude gain coefficient corresponding to the frequency band to be adjusted according to the amplitude ratio between the at least one overtone audio point to be adjusted and the fundamental frequency point includes:

[0018] Obtain the amplitude ratio between the average amplitude value of the at least one overtone audio point to be adjusted and the amplitude value of the fundamental frequency point.

[0019] Determine the amplitude gain coefficient of the frequency band to be adjusted according to the difference between the amplitude ratio and a preset amplitude ratio, where the amplitude gain coefficient is used to make the amplitude ratio between the frequency band to be adjusted and the fundamental frequency point conform to the preset amplitude ratio.

[0020] In one embodiment, after determining the audio amplitude adjustment strategy for the audio to be adjusted, it further includes:

[0021] Adjust the amplitude of the frequency band to be adjusted of each frame of the human voice signal according to the frequency band to be adjusted corresponding to each frame of the human voice signal in the audio to be adjusted and the amplitude gain coefficient corresponding to the frequency band to be adjusted, to obtain the target audio.

[0022] In one embodiment, after obtaining the fundamental frequency point and overtone audio points corresponding to each frame of the human voice signal in the audio to be adjusted, it further includes:

[0023] Among the overtone frequency points corresponding to each frame of the vocal signal, obtain a preset number of target overtone frequency points, and generate an envelope sequence corresponding to each frame of the vocal signal according to the fundamental frequency point and the preset number of target overtone frequency points; the preset number is determined based on the amplitude magnitude of the overtone frequency points;

[0024] Determine the overtone sufficiency level of the audio to be adjusted according to the amplitude decrease value of the envelope sequence corresponding to each frame of the vocal signal within a preset time; the overtone sufficiency level characterizes the singing level of the audio to be adjusted; the overtone sufficiency level is inversely proportional to the amplitude decrease value within the preset time;

[0025] Display the overtone sufficiency level.

[0026] In one embodiment, the obtaining a preset number of target overtone frequency points among the overtone frequency points corresponding to each frame of the vocal signal includes:

[0027] Among the overtone frequency points corresponding to each frame of the vocal signal, obtain the amplitude difference between each overtone frequency point and the fundamental frequency point;

[0028] Obtain a first overtone frequency point among the overtone frequency points whose amplitude difference from the fundamental frequency point is greater than or equal to a preset amplitude difference threshold for the first time, and use the first overtone frequency point and the overtone frequency points between the fundamental frequency point and the first overtone frequency point as the target overtone frequency points.

[0029] In one embodiment, the obtaining the fundamental frequency point and the overtone frequency points corresponding to each frame of the vocal signal in the audio to be adjusted includes:

[0030] Obtain the spectrogram corresponding to each frame of the vocal signal in the audio to be adjusted; wherein the spectrogram includes the amplitude and frequency of each frequency point of the vocal signal;

[0031] In the spectrogram corresponding to each frame of the vocal signal, determine the fundamental frequency point and the overtone frequency points corresponding to the fundamental frequency point.

[0032] In a second aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0033] In a third aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0034] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0035] The above method for obtaining an audio adjustment strategy, computer device, storage medium, and computer program product obtain the fundamental frequency points and harmonic frequency points of the human voice signal in the audio to be adjusted after receiving an audio adjustment instruction, determine the harmonic frequency points to be adjusted among the harmonic frequency points, and determine the amplitude gain coefficient of the frequency band to be adjusted composed of the harmonic frequency points according to the comparison result between the amplitude ratio between the fundamental frequency point and the harmonic frequency points and a preset amplitude ratio, and determine the frequency band to be adjusted and the amplitude gain coefficient as the audio amplitude adjustment strategy for the audio to be adjusted. Compared with the solution of manually determining the strategy to adjust the audio, this solution determines the amplitude gain coefficient based on the standard proportional relationship between the fundamental frequency and the harmonics, and determines the frequency to be adjusted based on the comparison between the frequency range and a preset frequency, thereby determining the strategy for adjusting the audio, improving the accuracy of determining the audio adjustment strategy. Description of the Drawings

[0036] Figure 1 It is a schematic flowchart of a method for obtaining an audio adjustment strategy in an embodiment;

[0037] Figure 2 It is a schematic diagram of a human voice signal in an embodiment;

[0038] Figure 3 It is a schematic diagram of a spectrogram in an embodiment;

[0039] Figure 4 It is a schematic diagram related to the step of obtaining target frequency points in an embodiment;

[0040] Figure 5 It is a schematic diagram related to the step of obtaining an envelope sequence in an embodiment;

[0041] Figure 6 It is a schematic structural diagram of a computer device in an embodiment. Detailed Embodiments

[0042] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0043] In one embodiment, as Figure 1 shown, a method for obtaining an audio adjustment strategy is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The method includes the following steps:

[0044] Step S202, when an audio adjustment instruction is detected, obtain the fundamental frequency points corresponding to each frame of the human voice signal in the adjusted audio and the harmonic frequency points corresponding to the fundamental frequency points.

[0045] Among them, the audio adjustment instruction can be an instruction triggered by the user on the terminal. For example, taking the audio as a song, when the terminal detects that the user has finished recording a song, it can display an audio adjustment instruction button for adjusting the equalization parameters of the song. After the user clicks the button, the audio adjustment instruction can be triggered. At this time, the terminal can obtain the audio to be adjusted, identify the human voice signal in the audio to be adjusted, and obtain the spectrogram corresponding to the human voice signal. Among them, the human voice signal can include multiple frames, each frame of the human voice signal includes multiple frequency points, and the spectrogram includes the amplitude and frequency of each frequency point in the human voice signal. In a specific implementation manner, the coordinate system can be used to represent the spectrogram, where the amplitude and frequency are used as an axis respectively, and each frequency point in the human voice signal is marked in the coordinate system by using the amplitude and frequency of each frequency point in the human voice signal.

[0046] Specifically, the human voice signal of the audio to be adjusted input by the user can be as Figure 2 shown. Figure 2 As shown in the schematic diagram of the human voice signal in an embodiment. Among them, the signal 301 can be the waveform diagram of the audio to be adjusted, and the signal 302 can be the human voice signal extracted from the audio to be adjusted, which can also be called the fundamental frequency signal. Among them, the terminal can obtain the spectrogram corresponding to each frame of the human voice signal through Fourier transform. For example, in an embodiment, obtaining the audio to be adjusted and the spectrogram corresponding to the human voice signal in the audio to be adjusted includes: obtaining the human voice signal of each frame of audio in the audio to be adjusted to obtain multiple frames of human voice signals; for each frame of the human voice signal, performing windowing processing on the frame of the human voice signal to obtain the windowed human voice signal; performing Fourier transform on the windowed human voice signal to obtain the spectrogram corresponding to the frame of the human voice signal.

[0047] In this embodiment, after the terminal obtains the audio to be adjusted, since the audio to be adjusted contains multiple frames of audio, the terminal can frame the audio to be adjusted and obtain the human voice signal of each frame of audio in the audio to be adjusted to obtain multiple frames of human voice signals. For each frame of the human voice signal, the terminal can perform windowing processing on the frame of the human voice signal to obtain the windowed human voice signal. Among them, the windowing processing means restricting the range of the human voice signal and performing Fourier transform within the range. The terminal can perform windowing on each frame of the human voice signal, so as to obtain multiple windowed human voice signals. The terminal can perform Fourier transform on the windowed human voice signal to obtain the spectrogram corresponding to the frame of the human voice signal. The terminal can perform the above-mentioned acquisition of the spectrogram on each frame of the human voice signal, so that the terminal can obtain the spectrogram corresponding to each frame of the human voice signal and obtain the spectrogram of the audio to be adjusted based on the human voice signals of all frames.

[0048] Among them, the above windowing can be a Hanning window, and the above Fourier transform can be an FFT (Fast Fourier Transform). The Hanning window is one of the window functions and is a special case of the raised cosine window. The Fast Fourier Transform is a fast algorithm for the discrete Fourier transform. It improves the algorithm of the discrete Fourier transform based on the odd, even, imaginary, real and other characteristics of the discrete Fourier transform. Specifically, the processing of the audio to be adjusted by the terminal is mainly the processing of the vocal part. When the terminal identifies the vocal signal, it can identify the fundamental frequency and harmonics in each frame of the audio to be adjusted, and obtain each frame of vocal signal, while the silent and soft sounds without vocal cord vibration will be discarded. After the terminal obtains multiple frames of vocal signals, it can perform windowing processing on each frame of vocal signal, such as adding a Hanning window, and perform Fourier transform on each window to obtain the amplitude and frequency of each frequency point, so as to obtain the above spectrogram. Among them, the formula of the above Hanning window can be shown as follows: w(i) = 0.54 - 0.46cos[(2πi) / (N - 1)], 0 ≤ i ≤ N - 1. Where i represents the sample point index, that is, the i-th sampling point in this frame, and N represents the window length, that is, the length of the above Hanning window. Here, the window length can be N = 512, and this window length can be equal to the length of the frame. After the terminal windows the above vocal signal, the windowed vocal signal can be shown as follows: xw n (Ln + i) = x(i)w(i), 0 ≤ i ≤ N - 1. Where n represents the n-th frame signal after windowing, L represents the frame shift. The frame shift represents a sliding window, and the offset of the start of the next frame relative to the start of the previous frame is the frame shift. For example, L = 256, and i represents the index starting from 0 of the N sample points in the n-th frame signal, that is, the i-th sample point among the N sample points. After the terminal performs Fourier transform on the n-th frame of vocal signal, the following results can be obtained:

[0049] Among them, xw n (Ln + i) represents the signal after windowing of the n-th frame, and (n, k) represents the k-th frequency point of the n-th frame, which can also be called a frequency point. After the terminal performs Fourier transform on each frame of the above vocal signal, the spectrogram corresponding to the vocal signal can be obtained. Among them, taking one frame of vocal signal as an example, after the terminal performs Fourier transform on one frame of vocal signal, it can obtain as Figure 3 shown in the spectrogram. Figure 3 It is a schematic diagram of the spectrogram of one frame of vocal signal in an embodiment. Among them, the spectrogram can form a curve as Figure 3 shown. Each point in the curve can be a frequency point. The abscissa of this spectrogram is the frequency, and the ordinate is the amplitude. Among them, both the frequency value in the abscissa and the amplitude in the ordinate are in a linearly increasing state.

[0050] After obtaining the spectrogram corresponding to the human voice signal, the terminal can obtain the fundamental frequency points in the spectrogram and the overtone frequency points corresponding to the fundamental frequency points. The fundamental frequency points represent the frequency points corresponding to the human voice, and the overtone frequency points represent the frequency points corresponding to the overtones of the human voice. The above human voice signal may include multiple frames of human voice signals. The terminal can, based on each frame of human voice signal, obtain the fundamental frequency points and overtone frequency points in the spectrogram of each frame of human voice signal, so that the terminal can obtain multiple fundamental frequency points and the overtone frequency points corresponding to each fundamental frequency point. Among them, the terminal can determine the target frequency points by searching for the maximum amplitude values in each frame of spectrogram, determine the fundamental frequency points from the multiple target frequency points, and determine the overtone frequency points among the other target frequency points except the fundamental frequency points based on the amplitude multiple relationship between the fundamental frequency points and the overtone frequency points, so as to obtain the fundamental frequency points and overtone frequency points in one frame.

[0051] Step S204: For each frame of human voice signal, determine at least one overtone frequency point to be adjusted among the overtone frequency points corresponding to the human voice signal, and determine the frequency band to be adjusted according to the at least one overtone frequency point to be adjusted.

[0052] Among them, in each frame of voice signal, there may be a fundamental frequency point and at least one harmonic frequency point. Since there can be multiple harmonic frequency points corresponding to a fundamental frequency point, the terminal can select one or more harmonic frequency points to be adjusted from them as the harmonic frequency points to be adjusted. For example, the terminal can obtain the frequencies of all harmonic frequency points from the above spectrogram, so as to obtain the harmonic frequency range corresponding to the fundamental frequency point in the spectrogram, that is, the harmonic frequency range represents the range corresponding to the frequencies of the above respective harmonic frequency points; the terminal can query a preset frequency adjustment table according to the above harmonic frequency range, where the preset frequency adjustment table includes multiple frequencies with different frequency values, and the terminal can determine multiple frequencies according to a preset multiple relationship and form the above preset frequency adjustment table. When there is a preset multiple relationship between the frequencies of the harmonic frequency points, the harmonics of the adjusted voice audio can be more sufficient, so that after adjusting the harmonic frequency points corresponding to the frequencies to be adjusted, the harmonics of the voice audio are more sufficient. Specifically, among the above multiple frequencies, each frequency represents the frequency corresponding to a harmonic, for example, the frequency of the fundamental frequency is 190 Hz, and there is a multiple relationship between the fundamental frequency and each harmonic, then the frequency of the first harmonic can be 380 Hz, etc., and 380 Hz can be a kind of frequency. The terminal can determine which frequencies in the preset frequency adjustment table are within the above harmonic frequency range by querying the preset frequency adjustment table, and use them as the frequencies to be adjusted, and use the harmonic frequency points corresponding to at least one determined frequency to be adjusted as the harmonic frequency points to be adjusted, so as to obtain at least one harmonic frequency point to be adjusted. In addition, in the above preset frequency adjustment table, the frequencies near the maximum or minimum value of the harmonic frequency range can also be used as the frequencies to be adjusted. After the terminal determines at least one harmonic frequency point to be adjusted, it can determine the frequency band to be adjusted based on the at least one harmonic frequency point to be adjusted. For example, when there is only one harmonic frequency point to be adjusted, the frequency band to be adjusted is determined as the value of the harmonic frequency point to be adjusted; when there are multiple harmonic frequency points to be adjusted, the terminal can calculate the frequency range where the respective harmonic frequency points to be adjusted are mainly concentrated, and determine the frequency band to be adjusted that needs to be gain-adjusted by equalization. For example, the terminal can determine the frequency range formed by the frequencies corresponding to the multiple harmonic frequency points to be adjusted as the frequency band to be adjusted. Among them, when displaying the above frequency band to be adjusted, the terminal can display the band name determined based on the above frequency range, such as low frequency, middle frequency, and high frequency, etc., or directly display the frequency range as the frequency band to be adjusted.

[0053] Step S206: For each frame of voice signal, determine the amplitude gain coefficient corresponding to the frequency band to be adjusted according to the amplitude ratio between at least one harmonic frequency point to be adjusted and the fundamental frequency point.

[0054] Among them, for each frame of vocal signal, there can be multiple overtone frequency points corresponding to the fundamental frequency point. The preset amplitude ratio is the amplitude ratio between the fundamental frequency point and at least one overtone frequency point to be adjusted. The closer the above amplitude ratio is to the preset amplitude ratio, the more it can indicate that the singer has a higher singing level. That is, when the amplitude ratio between at least one overtone frequency point to be adjusted and the fundamental frequency point can reach the above preset amplitude ratio, it indicates that the overtones of the audio are sufficient and the singing level is relatively high. When the amplitude ratio between the overtone frequency point to be adjusted and the fundamental frequency point cannot reach the above preset amplitude ratio, it is necessary to adjust at least one overtone frequency point to be adjusted through a gain coefficient. Therefore, the terminal can obtain the amplitude ratio between the above fundamental frequency point and the overtone frequency point to be adjusted corresponding thereto. For example, the terminal can obtain the average amplitude of all overtone frequency points to be adjusted corresponding to the above fundamental frequency point, and obtain the amplitude ratio between this average amplitude and the amplitude of the above fundamental frequency point. In addition, the terminal can also calculate the weighted sum of the amplitudes of the above respective overtone frequency points to be adjusted, and use the calculation result as a comparison parameter with the amplitude of the fundamental frequency point. For example, the terminal can pre-divide multiple amplitude ranges, and determine the weights of the respective overtone frequency points to be adjusted within each range according to the number of overtone frequency points to be adjusted belonging to the same range. For example, the more the number of overtone frequency points to be adjusted within the same range, the greater the weight can be. Thus, the terminal can obtain an amplitude value for amplitude comparison with the fundamental frequency point according to the weighted sum of multiple overtone frequency points to be adjusted. The terminal can compare the above amplitude ratio with the preset amplitude ratio to obtain a comparison result. Thus, the terminal can determine the amplitude gain coefficient of the above respective overtone frequency points based on this comparison result. For example, the amplitude gain coefficient of the frequency band composed of the respective overtone frequency points.

[0055] In addition, in some embodiments, the terminal can also calculate the ratio of the amplitude of each overtone frequency point to be adjusted to the amplitude of the fundamental frequency point respectively, so as to determine multiple amplitude ratios between multiple overtone frequency points to be adjusted and the fundamental frequency point. The terminal can calculate the difference between each amplitude ratio and the preset amplitude ratio, and determine the sub-amplitude gain coefficient corresponding to each overtone frequency point to be adjusted according to this difference. Thus, the terminal can determine the amplitude gain coefficient between the fundamental frequency point and the overtone frequency point to be adjusted according to the sub-amplitude gain coefficients of multiple overtone frequency points to be adjusted. For example, the terminal can calculate the average value of multiple sub-amplitude gain coefficients to obtain the above amplitude gain coefficient.

[0056] Step S208: Determine the audio amplitude adjustment strategy of the audio to be adjusted by using each frame of vocal signal, the frequency band to be adjusted corresponding to each frame of vocal signal, and the amplitude gain coefficient corresponding to the frequency band to be adjusted.

[0057] Among them, for each frame of voice signal, after the terminal determines the to-be-adjusted frequency and amplitude gain coefficient, it can obtain the corresponding pan audio points of the to-be-adjusted frequency as the to-be-adjusted pan audio points, and determine the audio amplitude adjustment strategy of the to-be-adjusted audio according to the to-be-adjusted frequency band formed by the frequencies of all the to-be-adjusted pan audio points and the corresponding amplitude gain coefficient. Among them, the above amplitude gain coefficient can be the amplitude gain coefficient of the to-be-adjusted frequency band in each frame of the voice signal, that is, the terminal can adjust the to-be-adjusted frequency band based on its corresponding amplitude gain coefficient.

[0058] Among them, the terminal can also generate corresponding equalization adjustment opinions based on the determined audio amplitude adjustment strategy and push them to the user. For example, based on the frequencies of all the to-be-adjusted pan audio points, a to-be-adjusted frequency band is determined, and the to-be-adjusted frequency band and the corresponding amplitude gain coefficient are pushed to the user. The user can adjust the amplitude of the to-be-adjusted frequency band based on the amplitude gain coefficient. After the terminal obtains the above audio amplitude adjustment strategy, it can be displayed to the user as a recommended solution, and only when the user determines to use it, the terminal adjusts the to-be-adjusted audio based on the audio amplitude adjustment strategy. For example, in one embodiment, after determining the audio amplitude adjustment strategy for the to-be-adjusted audio, it further includes: adjusting the amplitudes of the to-be-adjusted frequency bands of each frame of voice signal according to the to-be-adjusted frequency bands corresponding to each frame of voice signal in the to-be-adjusted audio and the amplitude gain coefficients corresponding to the to-be-adjusted frequency bands to obtain the target audio.

[0059] In this embodiment, the terminal can display the above audio adjustment strategy through a display device in the terminal. This audio adjustment strategy can also be regarded as an adjustment of the equalization parameters for each frame of voice signal. For example, the terminal can display "It is recommended to use this solution to adjust the equalization" on the display device. After the user determines to use the above audio adjustment strategy, the terminal can receive the audio adjustment strategy determination instruction triggered by the user. Then, the terminal can adjust the amplitude of the to-be-adjusted frequency band according to the to-be-adjusted frequency band and the amplitude gain coefficient corresponding to the to-be-adjusted frequency band within the frequency band corresponding to the to-be-adjusted frequency band in the to-be-adjusted audio. After the terminal adjusts the amplitudes of the to-be-adjusted frequency bands of each frame of voice signal based on the amplitude gain coefficient, the adjusted target audio can be obtained. Taking the audio as a song as an example, the terminal can also output the adjusted song so that the user can listen to the adjusted song.

[0060] In the above method for obtaining an audio adjustment strategy, after receiving an audio adjustment instruction, the fundamental frequency points and overtone frequency points of the human voice signal in the audio to be adjusted are obtained. The overtone frequency points to be adjusted are determined from the overtone frequency points, and according to the comparison result between the amplitude ratio between the fundamental frequency points and the overtone frequency points and a preset amplitude ratio, the amplitude gain coefficient of the frequency band to be adjusted composed of the overtone frequency points is determined. The frequency band to be adjusted and the amplitude gain coefficient are determined as the audio amplitude adjustment strategy of the audio to be adjusted. Compared with the solution of manually determining the strategy to adjust the audio, this solution determines the amplitude gain coefficient based on the standard proportional relationship between the fundamental frequency and the overtone, and determines the frequency to be adjusted based on the comparison between the frequency range and a preset frequency, so as to determine the strategy for adjusting the audio, improving the accuracy of determining the audio adjustment strategy.

[0061] In one embodiment, obtaining the fundamental frequency points and the overtone frequency points corresponding to the fundamental frequency points in the spectrogram corresponding to the human voice signal includes: determining the target frequency points in the spectrogram according to the comparison result between the amplitude magnitude of each frequency point in the spectrogram and the amplitude magnitudes of other frequency points; the other frequency points represent the frequency points within a preset range centered on each frequency point in the spectrogram; taking the target frequency point with the lowest frequency among the multiple target frequency points as the fundamental frequency point in the spectrogram; and determining the overtone frequency points among the other target frequency points except the fundamental frequency point among the multiple target frequency points according to the preset frequency multiple relationship between the fundamental frequency point and the overtone frequency points.

[0062] In this embodiment, the above spectrogram includes multiple frequency points, and there are fundamental frequency points and overtone frequency points corresponding to the fundamental frequency points among the multiple frequency points. Then the terminal can identify the fundamental frequency points and overtone frequency points among the multiple frequency points. Among them, one frame of spectrogram may include one fundamental frequency point and several overtone frequency points, and the terminal can identify the fundamental frequency points and overtone frequency points for the spectrogram of each frame of human voice signal. For each frequency point in the spectrogram corresponding to each frame of human voice signal, the terminal can compare the amplitude magnitude of this frequency point with the amplitude magnitudes of other frequency points corresponding to this frequency point to obtain a comparison result, and determine the target frequency point from the frequency points in the spectrogram according to this comparison result. Among them, the other frequency points represent the frequency points within a preset range centered on the above each frequency point in the spectrogram, that is, the terminal can take the above each frequency point as the center, obtain the frequency points within the preset range before and after each frequency point and compare the amplitude magnitudes to determine the target frequency point.

[0063] Among them, the target frequency points can be frequency points that may be pan audio frequency points. The terminal can perform the above comparison on each frequency point in the spectrogram to determine multiple target frequency points. For example, in one embodiment, multiple target frequency points in the spectrogram are determined according to the comparison result of the amplitude of each frequency point in the spectrogram and the amplitudes of other frequency points. The method includes: for each frequency point in the above spectrogram, if the amplitude of the frequency point is the maximum amplitude within a preset range, and the ratio of the amplitude of the frequency point to the second largest amplitude within the preset range is greater than or equal to a preset ratio, determine that the frequency point is a target frequency point; the second largest amplitude is the amplitude that is less than the maximum amplitude within the preset range and greater than the amplitudes of other frequencies within the preset range.

[0064] In this embodiment, the spectrogram of the above frame of human voice signal includes multiple frequency points. For each frequency point in the spectrogram of a frame of human voice signal, the terminal can obtain the amplitude of the frequency point and determine the preset range corresponding to the frequency point, where the preset range is a preset range centered on the frequency point in the spectrogram. The terminal can also obtain the frequency point corresponding to the second largest amplitude within the above preset range, where the second largest amplitude represents the amplitude of the frequency point that is less than the maximum amplitude and greater than the amplitudes of other frequencies within the preset range, that is, the second largest amplitude within the preset range. If the terminal detects that the amplitude of the frequency point is the maximum amplitude within the above preset range, and the ratio of the amplitude of the frequency point to the second largest amplitude within its preset range is greater than or equal to the preset ratio, the terminal can determine that the frequency point is a target frequency point.

[0065] Specifically, the spectrogram obtained after the terminal determines the target frequency points can be as Figure 4 shown, Figure 4 which is a schematic diagram of the target frequency point acquisition step in one embodiment. Figure 4Each point marked by a dot in it is the target frequency point. The terminal can traverse each frequency point f in this frame of spectrogram, delimit the comparison range n. If the terminal detects that the frequency point f is the frequency point with the largest amplitude within the range [f - n, f + n], and the amplitude value of this frequency point is greater than or equal to a preset multiple of the second largest frequency point, for example, 0.85 times, then the terminal can determine that this frequency point f is a target frequency point. Among them, the above [f - n, f + n] can be the above preset range, and the preset multiple can be set according to the actual situation. After the terminal determines multiple target frequency points, the above multiple target frequency points can be arranged in order according to the magnitude of the frequency. Then the terminal can use the target frequency point with the lowest frequency among the multiple target frequency points as the fundamental frequency point of this spectrogram, that is, the terminal can use the first target frequency point in the above spectrogram as the fundamental frequency point. Among them, there is a preset frequency multiple relationship between the fundamental frequency point and the harmonic frequency points. Taking the fundamental frequency point as f0 and the harmonic frequency points as f1…fn as an example, where the harmonic frequency points f1..fn include the first harmonic frequency point f1 to the nth harmonic frequency point fn. In the case where the harmonic frequency points are the harmonic frequency points corresponding to the fundamental frequency point, the frequency of f1 is twice the frequency of f0, the frequency of f2 is three times the frequency of f0, and the frequency of fn is n - 1 times the frequency of f0. The frequencies of the above f1…fn may not exactly be the accurate multiple relationship of f0. As long as the terminal detects that the frequency difference between the frequencies of f1…fn and the frequencies of the standard harmonic frequency points f'1…f'n corresponding to f0 is within the preset frequency difference range, these harmonic frequency points can also be used as the harmonic frequency points corresponding to the fundamental frequency point f0. After the terminal determines the preset frequency multiple relationship between the fundamental frequency point and the harmonic frequency points, it can determine the harmonic frequency points among the other target frequency points except the fundamental frequency point among the multiple target frequency points based on this relationship, that is, the terminal can determine each harmonic frequency point corresponding to the fundamental frequency point from the other target frequency points.

[0066] Specifically, as Figure 4 shown, Figure 4 the frequencies corresponding to each target frequency point in Figure 4In the first target frequency point f0 = 190 Hz, since f'1 is twice f'0 in the standard multiple relationship, the terminal can calculate that f1 is a target frequency point with a frequency of 380 Hz. The terminal can, based on this frequency, search backward from the fundamental frequency point in the spectrogram to find the first target frequency point that matches this frequency as the first harmonic frequency point f1. Additionally, the frequencies of f1...fn above may not be exactly multiples of f0. As long as the terminal detects that the frequency difference between f1...fn and the frequencies of the standard harmonic frequency points f'1...f'n corresponding to f0 is within the preset frequency difference range, these harmonic frequency points can also be regarded as the harmonic frequency points corresponding to the fundamental frequency point f0. For example, for the second harmonic frequency point f2, the terminal determines that the standard frequency of f2 is three times that of f0, i.e., 570 Hz. However, the frequency of the first closest target frequency point found by the terminal after the first harmonic frequency point is 569 Hz, and the difference from 570 Hz is within the preset frequency difference range. Then the terminal can also regard this target frequency point as the second harmonic frequency point f2. The terminal can perform the determination of the above harmonic frequency points for each target frequency point, so that the terminal can obtain multiple harmonic frequency points corresponding to each fundamental frequency point.

[0067] Through the above embodiments, the terminal can determine the target frequency points by performing range-based amplitude comparison on each frequency point in the spectrogram, and determine the fundamental frequency points and harmonic frequency points among multiple target frequency points based on the multiple relationship between the fundamental frequency points and the harmonic frequency points. Thus, the terminal can determine the audio adjustment strategy based on the fundamental frequency points and the harmonic frequency points, improving the accuracy of determining the audio adjustment strategy.

[0068] In one embodiment, determining the amplitude gain coefficient corresponding to the frequency band to be adjusted according to the amplitude ratio between at least one harmonic frequency point to be adjusted and the fundamental frequency point includes: obtaining the amplitude ratio between the average amplitude value of at least one harmonic frequency point to be adjusted and the amplitude value of the fundamental frequency point, and determining the amplitude gain coefficient of the frequency band of the harmonic frequency points to be adjusted according to the difference between the amplitude ratio and the preset amplitude ratio. For example, the amplitude gain coefficient of the frequency band composed of the frequencies of all harmonic frequency points to be adjusted, where the amplitude gain coefficient is used to make the amplitude ratio between the frequency band of the harmonic frequency points to be adjusted and the fundamental frequency point conform to the preset amplitude ratio.

[0069] In this embodiment, the spectrogram corresponding to each frame of voice signal includes a fundamental frequency point, and each fundamental frequency point corresponds to multiple harmonic frequency points. Taking a song as an example of the audio, there is a proportional relationship between the fundamental frequency point and the harmonic frequency points. When the fundamental frequency point and the harmonic frequency points conform to the corresponding proportional relationship, it represents a high singing level of the song. However, in actual singing, the ratios of the respective frequency points of the audio to be adjusted by the user often cannot reach the above ratio. Therefore, the terminal needs to determine the amplitude gain coefficients of the respective harmonic frequency points corresponding to the audio to be adjusted, so that after the terminal adjusts the harmonic frequency points based on the amplitude gain coefficients, the adjusted harmonic frequency points can conform to the above proportional relationship. For each fundamental frequency and its corresponding harmonic frequency points in the above spectrogram, the terminal can obtain the amplitude ratio between the fundamental frequency and at least one harmonic frequency point to be adjusted. For example, the terminal can first obtain the average amplitude of all the harmonic frequency points to be adjusted corresponding to the fundamental frequency point, and then obtain the amplitude ratio of the average amplitude to the amplitude of the above fundamental frequency point. Additionally, in some embodiments, the terminal can also obtain the weighted sum of the respective harmonic frequency points to be adjusted, and compare the weighted sum with the amplitude of the fundamental frequency point to obtain the amplitude ratio. The terminal can obtain the comparison result between the amplitude ratio and the preset amplitude ratio, such as how much the amplitude ratio differs from the preset amplitude ratio, and then further determine the amplitude gain coefficient corresponding to the frequency band where the at least one harmonic frequency point to be adjusted is located according to the comparison result, so that after the terminal adjusts the frequency band where the harmonic frequency point to be adjusted is located based on the amplitude gain coefficient, the amplitude ratio between the harmonic frequency point and the fundamental frequency point conforms to the above preset amplitude ratio.

[0070] Specifically, the terminal takes the ratio of the average value of the harmonics to the fundamental frequency as a reference, calculates the gain coefficient, and adjusts the amplitude of the harmonics to about a preset multiple of the corresponding fundamental frequency point, such as 0.8 times. Specifically, if the harmonic frequency points to be adjusted are the first harmonic frequency point f1 to the nth harmonic frequency point fn, the terminal can obtain the average amplitude of f1 - fn, obtain the amplitude ratio of the average amplitude to the amplitude of f0, so as to determine the corresponding amplitude gain coefficient, and adjust the amplitude of the frequency band where f1 - fn is located to about 0.8 times the fundamental frequency f0 based on the amplitude gain coefficient of the frequency band where f1 - fn is located.

[0071] Through this embodiment, the terminal can determine the amplitude gain coefficient of the harmonic frequency point based on the amplitude ratio between the harmonic frequency point and the fundamental frequency point, so that the terminal can determine the adjustment degree of the frequency band where the harmonic frequency point is located based on the amplitude gain coefficient, improving the accuracy of determining the audio adjustment strategy.

[0072] In one embodiment, querying the frequency to be adjusted in a preset frequency adjustment table according to the harmonic frequency range of the overtone points of the human voice signal includes: determining the frequency range formed by the overtone points corresponding to the human voice signal as the harmonic frequency range corresponding to the overtone points; obtaining the first frequency to be adjusted within the harmonic frequency range in the preset frequency adjustment table, and obtaining the second frequency to be adjusted whose frequency difference from the minimum value or the maximum value of the harmonic frequency range is within the preset frequency difference range; determining the first frequency to be adjusted and / or the second frequency to be adjusted as the frequency to be adjusted.

[0073] In this embodiment, the spectrogram corresponding to the above human voice signal includes multiple fundamental frequency points and multiple overtone points, and the frequencies of each overtone point are different. The terminal can frame the spectrogram corresponding to the human voice signal in the entire audio to be adjusted in units of one frame with a preset frame length, and determine the fundamental frequency points in each frame of the spectrogram corresponding to the human voice signal in the audio to be adjusted, so that multiple fundamental frequency points can be determined based on multiple frames of spectrograms. Therefore, the terminal can obtain the frequencies of the overtone points in each frame of the spectrogram corresponding to the human voice signal, and determine the harmonic frequency range corresponding to the overtone points. The terminal does not need to adjust the amplitude of each frequency, but needs to determine the frequency to be adjusted based on the harmonic frequency range. Therefore, the terminal can query the preset frequency adjustment table with the above harmonic frequency range to obtain the frequencies within the harmonic frequency range in the preset frequency adjustment table as the first frequencies to be adjusted, and the terminal can also obtain the frequencies near the maximum value or the minimum value of the harmonic frequency range as the second frequencies to be adjusted. For example, the terminal can obtain the second frequencies to be adjusted whose frequency differences from the minimum value or the maximum value of the harmonic frequency range are within the preset frequency difference range, so that the terminal can determine at least one of the obtained first frequencies to be adjusted and the second frequencies to be adjusted as the frequency to be adjusted.

[0074] Specifically, when the terminal determines the overtone points corresponding to the fundamental frequency points, it can determine the respective overtone points corresponding to the fundamental frequency points in each frame of the spectrogram based on the preset multiple relationship between the fundamental frequency points and the overtone points, including the above f1…fn, etc. The terminal can obtain the frequencies of all overtone points within the entire audio to be adjusted, so as to determine the harmonic frequency range, and determine the frequencies that need to be gain-adjusted by equalization based on the harmonic frequency range. Multiple adjustable frequencies are preset in the above preset frequency adjustment table, and the values of each frequency are different. These frequencies can be displayed on the display device of the terminal in different frequency bands, and the user can also adjust the frequency band where the frequency corresponding to the audio to be adjusted displayed in the terminal is located in the terminal. In order to determine the best audio adjustment strategy corresponding to the audio to be adjusted, the terminal can query the preset frequency adjustment table based on the above harmonic frequency range, and select the corresponding first frequency to be adjusted and / or the second frequency to be adjusted as the frequency to be adjusted.

[0075] Through this embodiment, the terminal can determine the frequency to be adjusted based on the overtone frequency range and the preset frequency adjustment table, so that the terminal can obtain the overtone frequency point with the frequency value being the frequency to be adjusted, and obtain the amplitude gain coefficient corresponding to the overtone frequency point to be adjusted, and adjust the amplitude of the overtone frequency point based on the amplitude gain coefficient, so as to improve the listening feeling of the audio to be adjusted and improve the accuracy of determining the audio adjustment strategy.

[0076] In one embodiment, after obtaining the fundamental frequency points and overtone frequency points corresponding to the voice signals of each frame in the audio to be adjusted, it further includes: among the overtone frequency points corresponding to the voice signals of each frame, obtaining a preset number of target overtone frequency points, and generating an envelope sequence corresponding to the voice signals of each frame according to the fundamental frequency points and the preset number of target overtone frequency points; the preset number is determined based on the amplitude of the overtone frequency points; determining the overtone sufficiency level of the audio to be adjusted according to the amplitude decrease value of the envelope sequence corresponding to the voice signals of each frame within a preset time; the overtone sufficiency level is inversely proportional to the amplitude decrease value within the preset time; the overtone sufficiency level characterizes the singing level of the audio to be adjusted; displaying the overtone sufficiency level.

[0077] In this embodiment, the terminal can also construct an envelope sequence based on the fundamental frequency point and its corresponding overtone frequency points. The above-mentioned spectrogram corresponding to the voice signal includes multiple fundamental frequency points, and each frame of the voice signal includes one fundamental frequency point. For each frame of the voice signal, there are multiple overtone frequency points in the spectrogram corresponding to the fundamental frequency point in this frame of the voice signal. The terminal can obtain a preset number of target overtone frequency points corresponding to the fundamental frequency point, that is, the terminal can select several target overtone frequency points from multiple overtone frequency points, and the selection rule can be determined based on the amplitude of the overtone frequency points, that is, the preset number can be determined based on the amplitude of the overtone frequency points. The terminal can generate an envelope sequence corresponding to the fundamental frequency point based on the above-mentioned fundamental frequency point and the preset number of target overtone frequency points.

[0078] Among them, the terminal can determine the target overtone frequency points to be selected based on the amplitude difference between the fundamental frequency point and its corresponding overtone frequency points. For example, in one embodiment, obtaining a preset number of target overtone frequency points among the overtone frequency points corresponding to the voice signals of each frame includes: obtaining the amplitude difference between each overtone frequency point and the fundamental frequency point among the overtone frequency points corresponding to the voice signals of each frame; obtaining the first overtone frequency point whose amplitude difference from the above-mentioned fundamental frequency point is greater than or equal to the preset amplitude difference threshold among each overtone frequency point, and taking the first overtone frequency point and the overtone frequency points between the fundamental frequency point and the first overtone frequency point as the target overtone frequency points.

[0079] In this embodiment, the spectrogram corresponding to the above-mentioned human voice signal contains multiple fundamental frequency points. Each frame of the human voice signal can correspond to one fundamental frequency point. For the fundamental frequency point and the harmonic frequency points in each frame of the human voice signal, the terminal can obtain a preset number of target harmonic frequency points after the fundamental frequency point in the human voice signal. For example, the terminal can obtain the amplitude difference between each harmonic frequency point in the spectrogram corresponding to the fundamental frequency point and the fundamental frequency point. The terminal can obtain the amplitude difference for each harmonic frequency point of the fundamental frequency point corresponding to this frame of the human voice signal in the order from left to right, and the terminal can obtain the first harmonic frequency point among the above-mentioned harmonic frequency points whose amplitude difference from the fundamental frequency point is greater than or equal to the preset amplitude difference threshold. This harmonic frequency point can be called the first harmonic frequency point. The terminal can use the first harmonic frequency point and the harmonic frequency points between the fundamental frequency point and the first harmonic frequency point as the target harmonic frequency points. Specifically, as Figure 5 shown, Figure 5 is a schematic diagram of the envelope sequence acquisition step in an embodiment. The point 501 in the figure is the fundamental frequency point. Usually, the amplitude of the fundamental frequency or the first harmonic f1 of a person is the strongest. The terminal can use the fundamental frequency point as a reference and compare the amplitude of each harmonic frequency point with that of the fundamental frequency point one by one to identify the harmonic frequency point at which the amplitude attenuation degree of the harmonic frequency point reaches the preset amplitude difference threshold, such as 45 dB. Specifically, as Figure 5 the point 502 in the figure, then the terminal can use this harmonic frequency point and the harmonic frequency points between this harmonic frequency point and the fundamental frequency point as the target harmonic frequency points, where the above-mentioned preset amplitude difference threshold can be set according to the actual situation.

[0080] After the terminal determines a preset number of target harmonic frequency points, it can also generate an envelope sequence based on the fundamental frequency point and the preset number of harmonic frequency points. For example, in one embodiment, generating the envelope sequence corresponding to each frame of the human voice signal according to the fundamental frequency point and the preset number of target harmonic frequency points includes: performing mean filtering on the fundamental frequency point of each frame of the human voice signal and the preset number of target harmonic frequency points of each frame of the human voice signal to obtain a preprocessed envelope sequence; normalizing the preprocessed envelope sequence to obtain the envelope sequence corresponding to each frame of the human voice signal.

[0081] In this embodiment, when generating an envelope sequence at the terminal, after obtaining the fundamental frequency points of each frame of vocal signals and a preset number of target overtone frequency points, the fundamental frequency points and the preset number of target overtone frequency points can be subjected to mean filtering to obtain a preprocessed envelope sequence. This preprocessed envelope sequence can be a rough envelope. The terminal can also normalize the preprocessed envelope sequence to obtain the envelope sequence corresponding to the fundamental frequency point. Among them, mean filtering means achieving a low-pass filtering effect by averaging the target point and the points in the nearby range and then outputting, so as to filter out some high-frequency components and make the signal smoother. Normalization means a processing process of retaining the original distribution ratio law of the data and compressing the range to between [0,1]. Specifically, after the terminal obtains the amplitude values of the above-mentioned fundamental frequency points and target overtone frequency points, it can obtain a rough envelope through mean filtering, normalize the rough envelope, and store it together with the fundamental frequency point of the current frame to obtain the envelope sequence corresponding to the fundamental frequency point. Among them, the above-mentioned preset number can have a maximum value. For example, the maximum is the seventh overtone frequency point f7. When the attenuation degree of the overtone frequency points corresponding to the above-mentioned fundamental frequency point still does not reach the preset amplitude difference threshold after reaching the maximum value, the terminal can directly use the first overtone frequency point f1 and the seventh overtone frequency point f7 as the target overtone frequency points.

[0082] After the terminal obtains the above envelope sequence, taking the audio as a song as an example, the envelope sequence can be used to evaluate the user's singing level. The terminal can obtain the amplitude decrease value of the above envelope sequence within a preset time, and determine the user's overtone sufficiency level according to the amplitude decrease value within the preset time. Among them, the overtone sufficiency level characterizes the singing level of the user when singing the above audio to be adjusted. The overtone sufficiency level is inversely proportional to the amplitude decrease value within the preset time, that is, the larger the amplitude decrease value within the preset time, the higher the overtone sufficiency level, which further indicates that the user's singing level is higher. The above preset time can be set according to the actual situation. After the terminal determines the overtone sufficiency level, it can display the overtone sufficiency level. For example, it can be displayed on the display device of the terminal. And when the terminal displays the overtone sufficiency level, it can also generate a certain piece of text for display based on the overtone sufficiency level. For example, when the overtone sufficiency level is the first level, it can generate a piece of text such as "Your voice is very infectious" for display to indicate that the user's singing level is relatively high; when the overtone sufficiency level is the second level, it can generate a piece of text such as "The mid-frequency is very full" for display for the advantages in the audio to be adjusted to indicate that the user's singing level is medium; when the overtone sufficiency level is the third level, it can generate a piece of text such as "Pay attention to holding your breath when singing and try the opinion of this teacher" for display to indicate that the user's singing level is relatively low.

[0083] Through the above embodiments, the terminal can generate an envelope sequence based on the fundamental frequency point and its corresponding pan audio points through methods such as mean filtering and normalization, determine the user's singing level based on the attenuation degree of the envelope sequence, and let the user know their level by displaying the copywriting, thereby improving the accuracy of determining the singing level.

[0084] It should be understood that although each step in the flowcharts involved in the above-described embodiments is sequentially shown as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0085] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 6 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an audio adjustment strategy acquisition method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0086] Those skilled in the art can understand that Figure 6 the structure shown in

[0087] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above-mentioned method for obtaining an audio adjustment strategy is implemented.

[0088] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for obtaining an audio adjustment strategy is implemented.

[0089] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above-mentioned method for obtaining an audio adjustment strategy is implemented.

[0090] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.

[0091] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0092] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0093] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for obtaining an audio adjustment strategy, characterized in that The method includes: When an audio adjustment instruction for the audio to be adjusted is detected, obtaining the fundamental frequency points and overtone frequency points corresponding to each frame of the vocal signal in the audio to be adjusted; For each frame of the vocal signal, determining at least one overtone frequency point to be adjusted among the overtone frequency points corresponding to the vocal signal, and determining the frequency band to be adjusted according to the at least one overtone frequency point to be adjusted; For each frame of the vocal signal, determining the amplitude gain coefficient corresponding to the frequency band to be adjusted according to the amplitude ratio between the at least one overtone frequency point to be adjusted and the fundamental frequency point; the amplitude gain coefficient is determined according to the difference between the amplitude ratio and a preset amplitude ratio; Determining each frame of the vocal signal, the frequency band to be adjusted corresponding to each frame of the vocal signal, and the amplitude gain coefficient corresponding to the frequency band to be adjusted as the audio amplitude adjustment strategy corresponding to the audio to be adjusted.

2. The method according to claim 1, wherein The determining at least one overtone frequency point to be adjusted among the overtone frequency points corresponding to the vocal signal includes: Querying the frequency to be adjusted in a preset frequency adjustment table according to the overtone frequency range of the overtone frequency points of the vocal signal; the preset frequency adjustment table includes the frequencies of the overtone frequency points that need to be frequency-adjusted; Determining at least one overtone frequency point corresponding to the frequency to be adjusted among the overtone frequency points corresponding to the vocal signal as the overtone frequency point to be adjusted.

3. The method according to claim 2, wherein The querying the frequency to be adjusted in a preset frequency adjustment table according to the overtone frequency range of the overtone frequency points of the vocal signal includes: Determining the frequency range formed by the overtone frequency points corresponding to the vocal signal as the overtone frequency range corresponding to the overtone frequency points; Obtaining a first frequency to be adjusted in the preset frequency adjustment table within the overtone frequency range, and obtaining a second frequency to be adjusted whose frequency difference from the minimum or maximum value of the overtone frequency range is within a preset frequency difference range; Determining the first frequency to be adjusted and / or the second frequency to be adjusted as the frequency to be adjusted.

4. The method according to claim 1, wherein The determining the amplitude gain coefficient corresponding to the frequency band to be adjusted according to the amplitude ratio between the at least one overtone frequency point to be adjusted and the fundamental frequency point includes: Obtaining the amplitude ratio between the average amplitude value of the at least one overtone frequency point to be adjusted and the amplitude value of the fundamental frequency point; Determining the amplitude gain coefficient of the frequency band to be adjusted according to the difference between the amplitude ratio and a preset amplitude ratio, where the amplitude gain coefficient is used to make the amplitude ratio between the frequency band to be adjusted and the fundamental frequency point conform to the preset amplitude ratio.

5. The method according to claim 1, wherein After determining the audio amplitude adjustment strategy for the audio to be adjusted, it further includes: Adjusting the amplitude of the frequency band to be adjusted of each frame of the vocal signal according to the frequency band to be adjusted corresponding to each frame of the vocal signal in the audio to be adjusted and the amplitude gain coefficient corresponding to the frequency band to be adjusted to obtain the target audio.

6. The method according to claim 1, wherein After obtaining the fundamental frequency points and overtone frequency points corresponding to each frame of the vocal signal in the audio to be adjusted, it further includes: Among the pan audio points corresponding to each frame of the vocal signal, obtain a preset number of target pan audio points, and generate an envelope sequence corresponding to each frame of the vocal signal according to the fundamental frequency point and the preset number of target pan audio points; the preset number is determined based on the amplitude size of the pan audio points; Determine the overtone sufficiency level of the audio to be adjusted according to the amplitude decrease value of the envelope sequence corresponding to each frame of the vocal signal within a preset time; the overtone sufficiency level characterizes the singing level of the audio to be adjusted; the overtone sufficiency level is inversely proportional to the amplitude decrease value within the preset time; Display the overtone sufficiency level.

7. The method according to claim 6, characterized in that, The obtaining a preset number of target pan audio points among the pan audio points corresponding to each frame of the vocal signal includes: Among the pan audio points corresponding to each frame of the vocal signal, obtain the amplitude difference between each pan audio point and the fundamental frequency point; Obtain the first pan audio point among each pan audio point whose amplitude difference from the fundamental frequency point is greater than or equal to a preset amplitude difference threshold, and use the first pan audio point and the pan audio points between the fundamental frequency point and the first pan audio point as target pan audio points.

8. The method according to claim 1, wherein The obtaining the fundamental frequency point and the pan audio points corresponding to each frame of the vocal signal in the audio to be adjusted includes: Obtain the spectrogram corresponding to each frame of the vocal signal in the audio to be adjusted; wherein the spectrogram includes the amplitude and frequency of each frequency point of the vocal signal; In the spectrogram corresponding to each frame of the vocal signal, determine the fundamental frequency point and the pan audio points corresponding to the fundamental frequency point.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for transforming voice signal characteristics

    CN103258539A

  • Audio signal processing method and device, and storage medium

    CN113362837A