Audio Processing Method, Apparatus, Device, and Storage Medium
By performing short-term Fourier transformation and fade-out window processing on the audio data, a spectrum diagram display effect that is more in line with the human ear hearing is generated, which solves the problem that the spectrum diagram is inconsistent with the human ear hearing in the existing technology, and improves the user's music viewing experience.
Patent Information
- Application Number
- CN202280003707.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-10
AI Technical Summary
In the prior art, the acquisition method of spectrum maps does not match the auditory effect of the human ear, resulting in inaccurate display effect.
By performing short-time Fourier transform on the audio data, the frequency domain loudness is calculated, and the frequency domain loudness is windowed by using the fade function to generate a display effect that is more in line with the hearing of the human ear.
Make the display effect of the spectrum diagram more suitable for the human ear and improve the user's music viewing experience.
Smart Images

Figure CN115956270B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing, and particularly to an audio processing method, apparatus, device, and storage medium. Background Art
[0002] When playing music, a spectrogram of the music will be displayed. The spectrogram of the music can be displayed as a bar chart, and users visually perceive the rhythm of the music by observing the ups and downs of the bar chart.
[0003] In the related art, the way to obtain the spectrogram is as follows: perform a short-time Fourier transform on the audio data to obtain the spectral data of each frame, and perform a smoothing process on the spectral data to obtain the spectrogram. While playing the audio data, the spectrogram of each frame of the audio data can be played synchronously, thereby displaying the rhythm of the audio data.
[0004] The spectrogram obtained by the method in the related art does not conform to the music effect heard by the human ear. Summary of the Invention
[0005] Embodiments of this application provide an audio processing method, apparatus, device, and storage medium, which can make the spectrogram more consistent with the human ear's hearing. The technical solutions are as follows.
[0006] According to one aspect of this application, an audio processing method is provided. The method includes:
[0007] Perform a short-time Fourier transform on the audio data to obtain a frequency-domain data set. Each time window in the frequency-domain data set corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency;
[0008] Calculate the loudness according to the amplitude in the frequency-domain data set to obtain a first frequency-domain loudness set. Each time window in the first frequency-domain loudness set corresponds to a set of frequency-domain loudness, and the frequency-domain loudness includes frequency and the loudness corresponding to the frequency;
[0009] Use a fade-in / fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set.
[0010] According to another aspect of this application, an audio processing apparatus is provided. The apparatus includes:
[0011] A processing module, configured to perform a short-time Fourier transform on the audio data to obtain a frequency-domain data set. Each time window in the frequency-domain data set corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency;
[0012] A loudness module, configured to calculate loudness based on the amplitudes in a frequency-domain dataset to obtain a first frequency-domain loudness set, where each time window in the first frequency-domain loudness set corresponds to a group of frequency-domain loudness, and the frequency-domain loudness includes a frequency and the loudness corresponding to the frequency;
[0013] A windowing module, configured to perform head and tail windowing processing on the frequency-domain loudness of each time window in the first frequency-domain loudness set by using a fade-in and fade-out function to obtain a second frequency-domain loudness set.
[0014] According to another aspect of the present application, there is provided a computer device, including: a processor and a memory, where at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the audio processing method described in the above aspect.
[0015] According to another aspect of the present application, there is provided a computer-readable storage medium, where at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the audio processing method described in the above aspect.
[0016] According to another aspect of an embodiment of the present disclosure, there is provided a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing method provided in the above optional implementation manner.
[0017] The beneficial effects brought by the technical solution provided in the embodiments of the present application at least include:
[0018] By performing short-time Fourier transform on audio data, the frequency-domain data of each time window can be obtained, and the frequency-domain data identifies the amplitude distribution of the audio data at each frequency within the time window. Then, the loudness is calculated based on the amplitudes in the frequency-domain data, and further, the frequency-domain loudness of each time window is obtained. The frequency-domain loudness can represent the loudness perception of the human ear for sound waves at each frequency within the current time window. Further, since the human ear has a weak perception of high-frequency and low-frequency sound waves, fade-in and fade-out windowing is performed on the head and tail of the audio loudness of each time window, so that the loudness value gradually decreases from the middle to both sides. The frequency-domain loudness of each time window calculated by the above method is more in line with the loudness distribution heard by the human ear during the playback of audio data. Based on this frequency-domain loudness, the relevant display effects during audio playback can be produced, so that the display effects are closer to the human ear auditory effects. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a block diagram of a computer device provided by an exemplary embodiment of the present application;
[0021] Figure 2 is a flowchart of an audio processing method provided by another exemplary embodiment of the present application;
[0022] Figure 3 is a flowchart of an audio processing method provided by another exemplary embodiment of the present application;
[0023] Figure 4 is a schematic diagram of an audio processing method provided by another exemplary embodiment of the present application;
[0024] Figure 5 is a schematic diagram of an audio processing method provided by another exemplary embodiment of the present application;
[0025] Figure 6 is a block diagram of an audio processing device provided by another exemplary embodiment of the present application;
[0026] Figure 7 is a schematic structural diagram of a server provided by another exemplary embodiment of the present application;
[0027] Figure 8 is a block diagram of a terminal provided by another exemplary embodiment of the present application. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the embodiments of the present application in detail with reference to the drawings.
[0029] Figure 1 shows a schematic diagram of a computer device 101 provided by an exemplary embodiment of the present application. The computer device 101 may be a terminal or a server.
[0030] The terminal may include at least one of a digital camera, a smart phone, a laptop computer, a desktop computer, a tablet computer, a smart speaker, and a smart robot. Optionally, the terminal may also be a device with a sound, such as an MP3 (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group 4), a speaker, a smart speaker, an in-vehicle computer, headphones, a smart home device, etc. In an alternative implementation, the audio processing method provided in this application may be applied to an application with audio processing capabilities, and the application may be: a music player, a video player, a short video player, an audio editor, a video editor, a social application, a life service application, a shopping application, a live broadcast application, a forum application, an information application, a life application, an office application, etc. Optionally, a client of the application is installed on the terminal.
[0031] Exemplarily, an audio processing algorithm is stored on the terminal. When the client needs to use the audio processing function provided in the embodiments of this application, the client may call the audio processing algorithm to complete the audio processing. Exemplarily, the audio processing process may be completed by the terminal or by the server.
[0032] The terminal and the server are interconnected through a wired or wireless network.
[0033] The terminal includes a first memory and a first processor. An audio processing algorithm is stored in the first memory; the above-mentioned audio processing algorithm is called and executed by the first processor to implement the audio processing method provided in this application. The first memory may include, but is not limited to, the following: Random Access Memory (RAM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM).
[0034] The first processor may be composed of one or more integrated circuit chips. Optionally, the first processor may be a general-purpose processor, such as a Central Processing Unit (CPU) or a Network Processor (NP). Optionally, the first processor may implement the audio processing method provided in this application by running a program or code.
[0035] The server includes a second memory and a second processor. An audio processing algorithm is stored in the second memory; the audio processing algorithm is called by the second processor to implement the audio processing method provided in this application. Optionally, the second memory may include, but is not limited to, the following: RAM, ROM, PROM, EPROM, EEPROM. Optionally, the second processor may be a general-purpose processor, such as a CPU or an NP.
[0036] Figure 2 The flowchart of the audio processing method provided by an exemplary embodiment of this application is shown. This method may be executed by a computer device, for example, by a terminal or a server as Figure 1 shown. The method includes the following steps.
[0037] Step 210, perform a short-time Fourier transform on the audio data to obtain a frequency-domain data set. Each time window in the frequency-domain data set corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency.
[0038] The audio data may be PCM (Pulse Code Modulation) audio data, or the audio data may be audio data in other audio formats.
[0039] The short-time Fourier transform divides the audio data into multiple time windows in time, performs a Fourier transform on the audio data within each time window, and obtains the frequency-domain data of each time window. The frequency-domain data of multiple time windows forms a frequency-domain data set. That is, the frequency-domain data set is a three-dimensional data composed of time windows, frequency, and amplitude. The frequency-domain data set includes the frequency-domain data of multiple time windows, and the frequency-domain data of each time window includes the amplitudes at each frequency. Exemplarily, the frequency-domain data may also include the real part and the imaginary part at each frequency, and the amplitude and phase are calculated based on the real part and the imaginary part. In the embodiments of this application, the loudness is calculated using the amplitude.
[0040] For example, each time window is one frame (30 ms, 60 ms, or 100 ms). Perform Fourier transform on each frame of audio signal of the audio data to obtain the frequency-domain data of each frame, and multiple frames of frequency-domain data form a frequency-domain data set. One frame of frequency-domain data includes the amplitude of the audio data of this frame at at least one frequency. Of course, the length of the time window can be set arbitrarily. For example, it can be set to two frames, one second, 1 ms, etc.
[0041] Step 220, calculate the loudness according to the amplitude in the frequency-domain data set to obtain the first frequency-domain loudness set. Each time window in the first frequency-domain loudness set corresponds to a group of frequency-domain loudness, and the frequency-domain loudness includes the frequency and the loudness corresponding to the frequency.
[0042] Calculate the amplitude in the frequency-domain data set as the loudness, and replace the amplitude in the frequency-domain data set with the loudness to obtain the first frequency-domain loudness set.
[0043] For example, loudness = 20lg(amplitude). When the amplitude of the first frame (the first time window) at 20 Hz (Hertz) is 10, the calculated loudness is 20. Then, the loudness of the first frame (the first time window) in the first frequency-domain loudness set at 20 Hz is 20.
[0044] The relationship between the first frequency-domain loudness set and the frequency-domain data set is: the time window remains unchanged, the frequency remains unchanged, and the amplitude is correspondingly replaced by the loudness. For example, if the frequency-domain data set includes the frequency-domain data of 30 time windows, then the first frequency-domain loudness set also includes the frequency-domain loudness of 30 time windows, and the time windows in the frequency-domain data set and the time windows in the first frequency-domain loudness set have a one-to-one correspondence.
[0045] Optionally, the sound pressure can also be calculated according to the amplitude in the frequency-domain data set to obtain the first frequency-domain sound pressure set. Each time window in the first frequency-domain sound pressure set corresponds to a group of frequency-domain sound pressure, and the frequency-domain sound pressure includes the frequency and the sound pressure corresponding to the frequency. Then, the "loudness" related terms in the subsequent steps can be correspondingly replaced by "sound pressure".
[0046] Step 230, use a fade-in / fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain the second frequency-domain loudness set.
[0047] The fade-in / fade-out function can be a function with starting and ending points of 0 / tending to 0, which first gradually increases and then gradually decreases. For example, the fade-in / fade-out function can be a broken-line function connected by (0,0), (1,1), (2,1), (3,0).
[0048] When the fade-in / fade-out function is a continuous function, the fade-in / fade-out function can be used to perform windowing processing on the full frequency band of the frequency-domain loudness of each time window. Windowing processing means multiplying the data to be windowed (frequency-domain loudness) by the windowing function (fade-in / fade-out function).
[0049] Optionally, the fade-in and fade-out function can also be two functions: a fade-in function and a fade-out function. The fade-in function is a gradually increasing function with a starting point of 0 / tending to 0. The fade-out function is a gradually decreasing function with an ending point of 0 / tending to 0.
[0050] The fade-in function can be used to perform windowing on the head part (a frequency band starting from the starting point) of the frequency-domain loudness of each time window, and the fade-out function can be used to perform windowing on the tail part (a frequency band ending at the ending point) of the frequency-domain loudness of each time window. The window length (frequency band length) of windowing can be set arbitrarily, and the window lengths of the head and tail can be the same or different. The window lengths for performing windowing on the heads of different time windows can be the same or different, and the window lengths for performing windowing on the tails of different time windows can be the same or different.
[0051] In summary, the method provided in this embodiment can obtain the frequency-domain data of each time window through short-time Fourier transform of the audio data. The frequency-domain data identifies the amplitude distribution of the audio data at each frequency within the time window. Then, the loudness is calculated based on the amplitude in the frequency-domain data, and thus the frequency-domain loudness of each time window is obtained. The frequency-domain loudness can represent the loudness perception of the human ear for sound waves at each frequency within the current time window. Further, since the human ear has a weak perception of high-frequency and low-frequency sound waves, fade-in and fade-out windowing are performed on the head and tail of the audio loudness of each time window, so that the loudness value gradually decreases from the middle to both sides. The frequency-domain loudness of each time window calculated by the above method is more in line with the loudness distribution heard by the human ear during audio data playback. Based on this frequency-domain loudness, the relevant display effects during audio playback can be produced, so that the display effects are closer to the human ear's auditory effects.
[0052] Figure 3 The flowchart of an audio processing method provided by an exemplary embodiment of the present application is shown. This method can be executed by a computer device, for example, executed by a terminal or a server as shown in Figure 1 The method includes the following steps.
[0053] Step 210, perform short-time Fourier transform on the audio data to obtain a frequency-domain data set. Each time window in the frequency-domain data set corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency.
[0054] For example, one time window is one frame, and if the audio data includes 1000 frames, then after performing short-time Fourier transform on the audio data, a frequency-domain data set composed of 1000 frames of frequency-domain data is obtained. One frame of frequency-domain data can form a dot plot / line graph / bar graph with the horizontal axis being the frequency (for example, the value range is 0 - 20000 Hz) and the vertical axis being the amplitude. Then the frequency-domain data set includes 1000 dot plots / line graphs / bar graphs of frequency-domain data.
[0055] Step 220: Calculate the loudness based on the amplitudes in the frequency-domain dataset to obtain a first frequency-domain loudness set. Each time window in the first frequency-domain loudness set corresponds to a group of frequency-domain loudness, and the frequency-domain loudness includes the frequency and the loudness corresponding to the frequency.
[0056] For example, the frequency-domain dataset of 1000 frames of frequency-domain data in the example of step 210 is converted into a first frequency-domain loudness set. The first frequency-domain loudness set includes 1000 frames of frequency-domain loudness. One frame of frequency-domain loudness can form a dot plot / line plot / bar plot with the horizontal axis being the frequency (for example, the value range is 0 - 20000 Hz) and the vertical axis being the amplitude. Then the first frequency-domain loudness set includes 1000 dot plots / line plots / bar plots of frequency-domain loudness.
[0057] Step 221: Perform A-weighting filtering processing and Mel-scale conversion on the first frequency-domain loudness set.
[0058] Optionally, after obtaining the first frequency-domain loudness set through step 220, further processing is also performed on the frequency-domain loudness of each time window in the first frequency-domain loudness set. The further processing includes A-weighting filtering processing and Mel-scale conversion.
[0059] A-weighting filtering processing simulates the loudness of a 40-phon pure tone by the human ear. When the signal passes through, there is a large attenuation in its low-frequency and mid-frequency (below 1000 Hz) bands. The characteristic curve of A-weighting filtering is close to the auditory perception characteristic of the human ear. For example, an A-weighting filter can be used to process each frequency-domain loudness in the first frequency-domain loudness set.
[0060] Mel-scale conversion is used to convert the frequency to the Mel scale. Most of the frequency range that the human ear can recognize is between 20 and 20000 Hz, but the recognition relationship of the human ear to the sound frequency unit Hz is not a simple linear relationship. For example, the human ear is most sensitive to sounds in the mid-low frequency range (around 1000 Hz). When the sound frequency increases from 1000 Hz to 2000 Hz, the human ear cannot feel that the frequency has doubled. Therefore, people often use the Mel scale to re-quantify the auditory perception characteristics of the human ear to frequency.
[0061] For example, if the horizontal axis of the frequency-domain loudness is the frequency with a value range of 0 - 20000 Hz, then through Mel-scale conversion, 0 - 20000 Hz is converted to the Mel scale. The loudness corresponding to the Mel scale is the sum of the loudness within the frequency loudness interval corresponding to the Mel scale. For example, Mel scale 1 corresponds to a certain frequency band, and the loudness of Mel scale 1 is the sum of the loudness corresponding to all frequencies within this frequency band. Thus, the first frequency-domain loudness set is made more in line with the human ear's perception of loudness.
[0062] The formula for mel-scale conversion is: Mel-scale = 2595 * lg(1 + frequency / 700). For example, 6300 Hz is converted to a mel-scale of 2595.
[0063] Mel-scale conversion can be accomplished using a mel-scale filter. The frequency-domain loudness concentrated in the first frequency-domain loudness after A-weighted filtering is input into the mel-scale filter to obtain the first frequency-domain loudness set after mel-scale conversion.
[0064] Optionally, the first frequency-domain loudness set used in subsequent steps can be the first frequency-domain loudness set obtained in step 220, or the first frequency-domain loudness set after A-weighted filtering, or the first frequency-domain loudness set after mel-scale conversion, or the first frequency-domain loudness set after A-weighted filtering and mel-scale conversion.
[0065] Through A-weighted filtering and mel-scale conversion, frequency-domain loudness that better conforms to human auditory changes can be obtained. However, this data is uneven and overly jittery. Therefore, in this embodiment, data smoothing, error correction to conform to the visual perception, and fade-in and fade-out are achieved through intra-frame and inter-frame data processing.
[0066] Step 222: When the distribution of the frequency-domain loudness satisfies the concentrated distribution condition, reduce the loudness values in the frequency-domain loudness that are lower than the average loudness.
[0067] The concentrated distribution condition is used to determine whether the distribution of the frequency-domain loudness is concentrated near the average loudness. For example, the average distribution condition is that the loudness variance is less than the first threshold, and the difference between the loudness expectation and the average loudness is less than the second threshold. The values of the first threshold and the second threshold are determined according to actual requirements.
[0068] Calculate the maximum loudness, minimum loudness, average loudness, loudness expectation, and loudness variance of the frequency-domain loudness for each time window, and then determine the distribution of the frequency-domain loudness in this time window. If the distribution is concentrated near the average loudness, then for the loudness lower than the average loudness, its value is pulled down to highlight the high-loudness part in the current time window.
[0069] Optionally, calculate the average loudness, loudness expectation, and loudness variance based on the frequency-domain loudness; when the loudness variance is less than the first threshold and the difference between the loudness expectation and the average loudness is less than the second threshold, reduce the values of the loudness in the frequency-domain loudness that are lower than the average loudness.
[0070] The way to reduce the values of the loudness in the frequency-domain loudness that are lower than the average loudness can be: subtracting a fixed value, multiplying by a coefficient, subtracting a gradient value, etc. For example, multiply all the loudness values lower than the average loudness by 0.5.
[0071] For example, as Figure 4As shown in (1) therein, it is the frequency-domain loudness of the first time window of the first frequency-domain loudness set, and its loudness is concentrated around the average loudness. Then, the loudness lower than the average loudness is reduced in value to obtain as shown in Figure 4 in (2) therein, so as to highlight the high-loudness part and make it conform to the human ear's auditory effect.
[0072] Step 230: Use a fade-in and fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set.
[0073] Optionally, use a fade-in function to perform windowing processing on the head of the frequency-domain loudness of each time window in the first frequency-domain loudness set, where the head is the part with a frequency lower than the first frequency threshold; use a fade-out function to perform windowing processing on the tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set, where the tail is the part with a frequency higher than the second frequency threshold.
[0074] The window sizes and positions of the head and tail can be set arbitrarily. For example, they can be set according to the value range of the horizontal axis. Starting from the minimum value of the value range to the first frequency threshold is the head, and from the second frequency threshold to the maximum value of the value range is the tail. It can also be determined according to the frequency range / Mel scale range where the loudness value in the frequency-domain loudness is greater than 0. For example, the horizontal axis value range of the frequency-domain loudness is 0-20000Hz, but the loudness from 0-10 is 0, the loudness at 10 is 1, the loudness at 90 is 1, and the loudness from 90-20000 is 0. Then the frequency / Mel scale range where the loudness value is greater than 0 is 10-90. Starting from the minimum value of the frequency / Mel scale range where the loudness value is greater than 0 to the first frequency threshold is the head, and from the second frequency threshold to the maximum value of the frequency / Mel scale range where the loudness value is greater than 0 is the tail. The window lengths of the head and tail can be the same or different.
[0075] For example, if the frequency value range of the frequency-domain loudness is 0-20000Hz, the head can be 0-100Hz and the tail can be 19900-20000Hz. Or, if the Mel scale range of the frequency-domain loudness is 0-100, the head can be 0-10 and the tail can be 90-100.
[0076] Optionally, the fade-in function is the part from 0 to π / 2 of the sine function; the fade-out function is the part from 0 to π / 2 of the cosine function.
[0077] Scale the fade-in function to be the same as the horizontal axis window length of the head, and then multiply the fade-in function by the frequency-domain loudness of the head to obtain the windowed head. Scale the fade-out function to be the same as the horizontal axis window length of the tail, and then multiply the fade-out function by the frequency-domain loudness of the tail to obtain the windowed tail.
[0078] For example, as shown in Figure 5As shown in (1) therein, it is the frequency-domain loudness of the first time window. The horizontal axis is the frequency / Mel scale and the vertical axis is the loudness. Then, a fade-in function is used to window the head part, and a fade-out function is used to window the tail part. After windowing, the frequency-domain loudness as shown in (2) in Figure 5 is obtained.
[0079] Step 231: Perform in-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness concentration.
[0080] Optionally, a polynomial smoothing algorithm is used to perform in-window data smoothing on the frequency-domain loudness of each time window. For example, if there are 10 loudness values from 0 - 100 hz in a window, then the j-th loudness is smoothed using the (j - 2)-th loudness, the (j - 1)-th loudness, the (j + 1)-th loudness, and the (j + 2)-th loudness. In this way, each loudness value in the time window is smoothed in turn. j is an integer greater than 2.
[0081] Step 232: Perform inter-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness concentration.
[0082] Optionally, a sliding smoothing weighted filtering algorithm is used to smooth the frequency-domain loudness of the i-th time window based on the frequency-domain loudness of the (i - 1)-th time window, the i-th time window, and the (i + 1)-th time window, where i is a positive integer.
[0083] If the loudness f3 of the (i + 1)-th time window at the x-th frequency is less than the loudness f2 of the i-th time window at the x-th frequency, then the smoothed loudness c3 of the i-th time window at the x-th frequency = (the loudness f3 of the (i + 1)-th time window at the x-th frequency) * decay coefficient a + (1 - decay coefficient a) * (1 - decay coefficient a) * (the loudness f2 of the i-th time window at the x-th frequency before smoothing) + (1 - decay coefficient a) * decay coefficient a * (the loudness f1 of the (i - 1)-th time window at the x-th frequency).
[0084] If the loudness f3 of the (i + 1)-th time window at the x-th frequency is greater than or equal to the loudness f2 of the i-th time window at the x-th frequency, then the smoothed loudness c3 of the i-th time window at the x-th frequency = (the loudness f3 of the (i + 1)-th time window at the x-th frequency) * rise coefficient b + (1 - rise coefficient b) * rise coefficient b * (the loudness f2 of the i-th time window at the x-th frequency before smoothing) + (1 - rise coefficient b) * (1 - rise coefficient b) * (the loudness f1 of the (i - 1)-th time window at the x-th frequency).
[0085] Among them, the decay coefficient a and the rise coefficient b can be set according to requirements. The value ranges of a and b are from 0 to 1.
[0086] Step 233: Generate the playback display effect of the audio data based on the second frequency-domain loudness set. The playback display effect includes at least one of the spectrogram of the audio data, the background image playback effect, the lyrics playback effect, the music fountain effect, and the lighting effect.
[0087] Optionally, for the second frequency-domain loudness set obtained after the above steps of processing, the frequency-domain loudness therein is the spectral data that conforms to the human ear's hearing. Then, based on the frequency-domain loudness in the second frequency-domain loudness set, the playback display effect during the playback of the audio data can be generated.
[0088] For example, a spectrogram that rhythmically changes with the music during music playback can be generated. Or, the background image playback effect can be controlled according to the audio loudness. For example, the display duration of the image, the scaling size of the image, the playback speed of the image, etc. can be controlled. Or, the lighting irradiation effect can also be controlled according to the frequency-domain loudness. Through these playback display effects, a music playback experience can be brought to the user visually, improving the consistency between the music the user hears and the playback display effect, and enhancing the user's music viewing experience.
[0089] Optionally, the above steps for processing the audio data can be arbitrarily deleted, the order can be adjusted, and new embodiments can be obtained by combination.
[0090] In summary, for the method provided in this embodiment, by performing a short-time Fourier transform on the audio data, the frequency-domain data of each time window can be obtained. This frequency-domain data identifies the amplitude distribution of the audio data at each frequency within the time window. Then, the loudness is calculated based on the amplitude in the frequency-domain data, and thus the frequency-domain loudness of each time window is obtained. The frequency-domain loudness can represent the loudness perception of the human ear for sound waves at each frequency within the current time window. Further, A-weighted filtering and Mel-scale conversion are performed on the frequency-domain loudness. Then, for the time window where the loudness distribution is concentrated around the average loudness, the loudness values below the average loudness are reduced to highlight the high-loudness part. Then, fade-in and fade-out windows are added to the beginning and end of the audio loudness of each time window, so that the loudness values gradually decrease from the middle to both sides. Finally, the data within the frame is smoothed and the data between frames is smoothed to make the visual effect of the frequency-domain loudness values more smooth. The frequency-domain loudness of each time window calculated by the above method is more in line with the loudness distribution when the audio data is played as heard by the human ear. Based on this frequency-domain loudness, the relevant display effects during audio playback can be produced, so that the display effect is closer to the human ear's auditory effect, improving the visual and auditory matching degree when the user views the playback display effect and enhancing the viewing experience.
[0091] The following is an embodiment of the apparatus of the present application. For the details not described in detail in the embodiment of the apparatus, reference can be made to the corresponding records in the above method embodiment, which will not be elaborated herein.
[0092] Figure 6The structural schematic diagram of an audio processing device provided by an exemplary embodiment of the present application is shown. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:
[0093] A processing module 401, configured to perform a short-time Fourier transform on audio data to obtain a frequency-domain data set. Each time window in the frequency-domain data set corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency;
[0094] A loudness module 402, configured to calculate loudness based on the amplitude in the frequency-domain data set to obtain a first frequency-domain loudness set. Each time window in the first frequency-domain loudness set corresponds to a set of frequency-domain loudness, and the frequency-domain loudness includes frequency and the loudness corresponding to the frequency;
[0095] A windowing module 403, configured to perform head and tail windowing processing on the frequency-domain loudness of each time window in the first frequency-domain loudness set by using a fade-in and fade-out function to obtain a second frequency-domain loudness set.
[0096] In an optional embodiment, the windowing module 403 is configured to perform windowing processing on the head of the frequency-domain loudness of each time window in the first frequency-domain loudness set by using a fade-in function, where the head is the part with a frequency lower than the first frequency threshold;
[0097] The windowing module 403 is configured to perform windowing processing on the tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set by using a fade-out function, where the tail is the part with a frequency higher than the second frequency threshold.
[0098] In an optional embodiment, the fade-in function is the part from 0 to π / 2 of the sine function;
[0099] The fade-out function is the part from 0 to π / 2 of the cosine function.
[0100] In an optional embodiment, the device further includes:
[0101] A reduction module 406, configured to reduce the loudness value lower than the average loudness in the frequency-domain loudness when the distribution of the frequency-domain loudness satisfies the concentrated distribution condition.
[0102] In an optional embodiment, the reduction module 406 is configured to calculate the average loudness, loudness expectation, and loudness variance based on the frequency-domain loudness;
[0103] The reduction module 406 is configured to reduce the value of the loudness lower than the average loudness in the frequency-domain loudness when the loudness variance is less than the first threshold and the difference between the loudness expectation and the average loudness is less than the second threshold.
[0104] In an optional embodiment, each time window in the second frequency-domain loudness concentration corresponds to a set of frequency-domain loudness; the apparatus further includes:
[0105] A smoothing module 405, configured to perform in-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness concentration.
[0106] In an optional embodiment, each time window in the second frequency-domain loudness concentration corresponds to a set of frequency-domain loudness; the apparatus further includes:
[0107] A smoothing module 405, configured to perform inter-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness concentration.
[0108] In an optional embodiment, the smoothing module 405 is configured to use a sliding smoothing weighted filtering algorithm to smooth the frequency-domain loudness of the i-th time window based on the frequency-domain loudness of the (i - 1)-th time window, the i-th time window, and the (i + 1)-th time window, where i is a positive integer.
[0109] In an optional embodiment, the processing module 401 is configured to perform A-weighted filtering processing and Mel scale conversion on the first frequency-domain loudness set.
[0110] In an optional embodiment, the apparatus further includes:
[0111] A display module 404, configured to generate a playback display effect of the audio data based on the second frequency-domain loudness set, where the playback display effect includes at least one of a spectrogram of the audio data, a background map playback effect, a lyrics playback effect, and a music fountain effect.
[0112] Figure 7 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specifically: The server 800 includes a central processing unit (Central Processing Unit, abbreviated as CPU) 801, a system memory 804 including a random access memory (Random Access Memory, abbreviated as RAM) 802 and a read-only memory (Read-Only Memory, abbreviated as ROM) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The server 800 further includes a basic input / output system (I / O system) 806 for facilitating information transmission between various devices in the computer, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0113] The basic input / output system 806 includes a display 808 for displaying information and input devices 809 such as a mouse, keyboard, etc. for user account input information. Both the display 808 and the input devices 809 are connected to the central processing unit 801 through an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include an input / output controller 810 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides outputs to a display screen, printer, or other types of output devices.
[0114] The mass storage device 807 is connected to the central processing unit 801 through a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable medium provide non-volatile storage for the server 800. That is to say, the mass storage device 807 may include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0115] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art know that computer storage media is not limited to the above several. The above system memory 804 and the mass storage device 807 can be collectively referred to as memory.
[0116] According to various embodiments of the present application, the server 800 may also run on a remote computer on the network connected through a network such as the Internet. That is, the server 800 may be connected to the network 812 through the network interface unit 811 connected to the system bus 805. Or rather, the network interface unit 811 may also be used to connect to other types of networks or remote computer systems (not shown).
[0117] The present application also provides a terminal, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the audio processing method provided by each of the above method embodiments. It should be noted that the terminal may be as follows Figure 8 the provided terminal.
[0118] Figure 8 FIG. shows a structural block diagram of a terminal 900 provided by an exemplary embodiment of the present application. The terminal 900 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 900 may also be referred to by other names such as a user account device, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0119] Generally, the terminal 900 includes: a processor 901 and a memory 902.
[0120] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0121] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 901 to implement the audio processing method or the audio processing method provided in the method embodiments of the present application.
[0122] In some embodiments, the terminal 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning component 908, and a power supply 909.
[0123] The peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0124] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Exemplarily, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user account identity module card, and so on. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0125] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905. The touch signals can be input to the processor 901 as control signals for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, which is provided on the front panel of the terminal 900; in other embodiments, there may be at least two display screens 905, which are respectively provided on different surfaces of the terminal 900 or are in a folding design; in still other embodiments, the display screen 905 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 900. Even further, the display screen 905 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 905 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0126] The camera module 906 is used to collect images or videos. Exemplarily, the camera module 906 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera respectively, to implement functions such as the combination of the main camera and the depth camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions, or other combined shooting functions. In some embodiments, the camera module 906 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0127] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves of the user account and the environment, and convert the sound waves into electrical signals for input to the processor 901 for processing, or input to the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 907 may further include a headphone jack.
[0128] The positioning component 908 is used to locate the current geographical location of the terminal 900 to achieve navigation or LBS (Location Based Service). The positioning component 908 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.
[0129] The power supply 909 is used to supply power to each component in the terminal 900. The power supply 909 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 909 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.
[0130] In some embodiments, the terminal 900 further includes one or more sensors 910. The one or more sensors 910 include but are not limited to: an acceleration sensor 911, a gyroscope sensor 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.
[0131] The acceleration sensor 911 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 900. For example, the acceleration sensor 911 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 901 can control the display screen 905 to display the user account interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 911. The acceleration sensor 911 can also be used for collecting motion data of games or user accounts.
[0132] The gyroscope sensor 912 can detect the body direction and rotation angle of the terminal 900. The gyroscope sensor 912 can cooperate with the acceleration sensor 911 to collect the 3D actions of the user account on the terminal 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can implement the following functions: motion sensing (such as changing the UI according to the tilt operation of the user account), image stabilization during shooting, game control, and inertial navigation.
[0133] The pressure sensor 913 can be disposed on the side frame of the terminal 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is disposed on the side frame of the terminal 900, it can detect the holding signal of the user account on the terminal 900, and the processor 901 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the display screen 905, the processor 901 can control the operable controls on the UI interface according to the pressure operation of the user account on the display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0134] The fingerprint sensor 914 is used to collect the fingerprint of the user account. The processor 901 can identify the identity of the user account according to the fingerprint collected by the fingerprint sensor 914, or the fingerprint sensor 914 can identify the identity of the user account according to the collected fingerprint. When the identity of the user account is identified as a trusted identity, the processor 901 authorizes the user account to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 914 can be disposed on the front, back, or side of the terminal 900. When there are physical buttons or manufacturer logos on the terminal 900, the fingerprint sensor 914 can be integrated with the physical buttons or manufacturer logos.
[0135] The optical sensor 915 is used to collect the ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 according to the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera module 906 according to the ambient light intensity collected by the optical sensor 915.
[0136] The proximity sensor 916, also known as a distance sensor, is usually disposed on the front panel of the terminal 900. The proximity sensor 916 is used to collect the distance between the user's account and the front of the terminal 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user's account and the front of the terminal 900 is gradually decreasing, the processor 901 controls the display screen 905 to switch from the lit state to the off state; when the proximity sensor 916 detects that the distance between the user's account and the front of the terminal 900 is gradually increasing, the processor 901 controls the display screen 905 to switch from the off state to the lit state.
[0137] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the terminal 900, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.
[0138] The memory further includes one or more programs, and the one or more programs are stored in the memory. The one or more programs include methods for performing the audio processing provided in the embodiments of the present application.
[0139] The present application further provides a computer device, which includes: a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the audio processing methods provided in the foregoing method embodiments.
[0140] The present application further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set, or an instruction set is stored. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the audio processing methods provided in the foregoing method embodiments.
[0141] The present application further provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing methods provided in the foregoing optional implementation manners.
[0142] It should be understood that as used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0143] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0144] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An audio processing method, characterized in that The method includes: Performing a short-time Fourier transform on the audio data to obtain a frequency-domain data set, where each time window in the frequency-domain data set corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency; Calculating the loudness based on the amplitudes in the frequency-domain data set to obtain a first frequency-domain loudness set, where each time window in the first frequency-domain loudness set corresponds to a set of frequency-domain loudness, and the frequency-domain loudness includes frequency and the loudness corresponding to the frequency; Using a fade-in and fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set.
2. The method according to claim 1, characterized in that The step of using a fade-in and fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set includes: Using a fade-in function to perform windowing processing on the head of the frequency-domain loudness of each time window in the first frequency-domain loudness set, where the head is the part with a frequency lower than the first frequency threshold; Using a fade-out function to perform windowing processing on the tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set, where the tail is the part with a frequency higher than the second frequency threshold.
3. The method according to claim 2, wherein The fade-in function is the part from 0 to π / 2 of the sine function; The fade-out function is the part from 0 to π / 2 of the cosine function.
4. The method according to any one of claims 1 to 3, characterized in that Before using the fade-in and fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set, it further includes: When the distribution of the frequency-domain loudness satisfies the concentrated distribution condition, reducing the loudness values in the frequency-domain loudness that are lower than the average loudness.
5. The method according to claim 4, characterized in that The step of reducing the loudness values in the frequency-domain loudness that are lower than the average loudness when the distribution of the frequency-domain loudness satisfies the concentrated distribution condition includes: Calculating the average loudness, loudness expectation, and loudness variance based on the frequency-domain loudness; When the loudness variance is less than the first threshold and the difference between the loudness expectation and the average loudness is less than the second threshold, reducing the values in the frequency-domain loudness that are lower than the average loudness.
6. The method according to any one of claims 1 to 3, characterized in that Each time window in the second frequency-domain loudness set corresponds to a set of frequency-domain loudness; After using the fade-in and fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set, it further includes: Performing in-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness set.
7. The method according to any one of claims 1 to 3, characterized in that Each time window in the second frequency-domain loudness set corresponds to a set of frequency-domain loudness; After using the fade-in and fade-out function to perform windowing processing on the head and tail of the frequency-domain loudness of each time window in the first frequency-domain loudness set to obtain a second frequency-domain loudness set, it further includes: Performing inter-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness set.
8. The method according to claim 7, wherein The step of performing inter-window data smoothing on the frequency-domain loudness of each time window in the second frequency-domain loudness set includes: Adopt a sliding smooth weighted filtering algorithm to smooth the frequency-domain loudness of the \(i\)-th time window based on the frequency-domain loudness of the \((i - 1)\)-th time window, the \(i\)-th time window, and the \((i + 1)\)-th time window, where \(i\) is a positive integer greater than 1.
9. The method according to claim 8, wherein After calculating the loudness according to the amplitude in the frequency-domain dataset to obtain the first frequency-domain loudness set, it further includes: Performing A-weighted filtering processing and Mel-scale conversion on the first frequency-domain loudness set.
10. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Generating a playback display effect of the audio data based on the second frequency-domain loudness set, where the playback display effect includes at least one of a spectrogram of the audio data, a background map playback effect, a lyrics playback effect, a music fountain effect, and a lighting effect.
11. An audio processing device, characterized in that, The device includes: A processing module for performing a short-time Fourier transform on the audio data to obtain a frequency-domain dataset, where each time window in the frequency-domain dataset corresponds to a set of frequency-domain data, and the frequency-domain data includes frequency and the amplitude corresponding to the frequency; A loudness module for calculating the loudness according to the amplitude in the frequency-domain dataset to obtain a first frequency-domain loudness set, where each time window in the first frequency-domain loudness set corresponds to a set of frequency-domain loudness, and the frequency-domain loudness includes frequency and the loudness corresponding to the frequency; A windowing module for performing head and tail windowing processing on the frequency-domain loudness of each time window in the first frequency-domain loudness set by using a fade-in and fade-out function to obtain a second frequency-domain loudness set.
12. A computer device, characterized in that, The computer device includes: a processor and a memory, where at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the audio processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the audio processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Audio identification method and device
CN103971689A
Audio coding method and device, computer readable storage medium and equipment
CN111462764A