Audio processing method and device, medium and product
By performing frame processing and style tag recognition on the songs, and selecting sound effects processing parameters based on the tags, the problem that a single sound effect cannot meet the user's diverse auditory needs is solved, and the personalized processing of song style is achieved.
Patent Information
- Application Number
- CN202510615658.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot meet the diverse auditory needs of users for different song segments with different styles, and a single sound effect processing cannot meet the personalized needs of users.
By performing frame-based processing on songs, the style tags of each audio frame are identified, and corresponding sound processing parameters are selected according to the tags, and the songs are personalized.
It realizes more different styles and more personalized sound effects processing for different song clips to meet the diverse auditory needs of users.
Smart Images

Figure CN120260526A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio processing, and particularly to an audio processing method, device, medium, and product. Background Art
[0002] With the development of the times, more and more music styles have flooded into the market. Currently, a smart phone can be installed with multiple music playing software, and each music playing software can provide multiple sound effects. During application, the smart phone can obtain the song to be played in the music playing software and the sound effect selected by the user. Furthermore, the smart phone can process the song to be played according to the sound effect selected by the user to change the style of the song to be played.
[0003] However, in some songs, different segments may have different styles. Thus, processing such a song with a single sound effect cannot meet the diverse auditory needs of users. Summary of the Invention
[0004] To solve the problem that processing a song with a single sound effect cannot meet the diverse auditory needs of users, embodiments of this application provide an audio processing method, device, medium, and product, including:
[0005] In a first aspect, embodiments of this application provide an audio processing method applied to an electronic device, including: obtaining first audio data of a song to be played; performing frame division processing on the first audio data to obtain a plurality of audio frame data; performing classification processing on each of the plurality of audio frame data to determine a style label for each of the audio frame data; and processing the first audio data based on audio processing parameters corresponding to the style labels of each of the audio frame data to obtain second audio data.
[0006] In some optional implementation manners of the first aspect, performing frame division processing on the first audio data to obtain a plurality of audio frame data includes: performing frame division processing on the first audio data based on a preset duration to obtain a plurality of audio frame data.
[0007] In some optional implementation manners of the first aspect, performing classification processing on each of the plurality of audio frame data to determine a style label for each of the audio frame data includes: performing feature extraction processing on each of the plurality of audio frame data to obtain a feature vector group for each of the audio frame data; and determining a style label for each of the audio frame data based on the feature vector group for each of the audio frame data.
[0008] In some optional implementation manners of the first aspect, the feature vector group for each of the audio frame data includes a waveform factor and a mel-frequency cepstral coefficient.
[0009] In some alternative implementations of the first aspect, determining the style labels of each audio frame data includes: inputting the feature vector groups of each audio frame data into a reverse neural network, and using the reverse neural network to classify each audio frame data to determine the style labels of each audio frame data.
[0010] In some alternative implementations of the first aspect, determining the style labels of each audio frame data includes: determining the style label with the largest proportion among the style labels of multiple audio frame data adjacent to each audio frame data as the style label of each audio frame data.
[0011] In some alternative implementations of the first aspect, processing the first audio data based on the audio processing parameters corresponding to the style labels of each audio frame data to obtain the second audio data includes: obtaining the audio processing parameters corresponding to the style labels of each audio frame data based on multiple frequency bands to be adjusted corresponding to the style labels of each audio frame data and the processing methods corresponding to each frequency band to be adjusted; adjusting the multiple frequency bands to be adjusted corresponding to the style labels of each audio frame data based on the processing methods corresponding to each frequency band to be adjusted to obtain the second audio data.
[0012] In a second aspect, the present application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the audio processing method mentioned in the first aspect or any one of the first aspects of the present application.
[0013] In a third aspect, the present application provides a readable storage medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device executes the audio processing method mentioned in the first aspect or any one of the first aspects of the present application.
[0014] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When executed by an electronic device, the electronic device executes the computer program code of the audio processing method mentioned in the first aspect or any one of the first aspects of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 According to some embodiments of the present application, a flowchart of an audio processing method is shown;
[0016] Figure 2 According to some embodiments of the present application, a flowchart block diagram of an audio processing method is shown;
[0017] Figure 3According to some embodiments of the present application, a schematic structural diagram of a backpropagation neural network is shown;
[0018] Figure 4 According to some embodiments of the present application, a flowchart of another audio processing method is shown;
[0019] Figure 5 According to some embodiments of the present application, a schematic hardware structure diagram of an electronic device is shown. Detailed implementation manners
[0020] Embodiments of the present application include but are not limited to an audio processing method, device, medium, and product.
[0021] It can be understood that the audio processing method mentioned in the embodiments of the present application can be applied to an electronic device. Among them, the electronic device can be called a terminal, user equipment (UE), mobile terminal (MT), etc. In some specific implementation manners, the electronic device can be a vehicle-mounted terminal, a smart phone, a portable music player, a smart watch, or other devices with an audio playback function.
[0022] Next, in combination with the scenario of playing a song on a smart phone, the audio processing method mentioned in the embodiments of the present application will be introduced.
[0023] It can be understood that a smart phone can install multiple music playback software, and each music playback software can provide multiple sound effects. During application, the smart phone can obtain the song to be played in the music playback software and the sound effect selected by the user. Furthermore, the smart phone can process the song to be played according to the sound effect selected by the user to change the style of the song to be played.
[0024] In some implementation manners, the music playback software can provide sound effects such as dynamic electro music, 3D surround, classic rock, heavy bass, and gentle ancient style for the user to freely select according to their own preferences. During application, the electronic device can obtain the sound effect selected by the user from these sound effects and process the original audio data of the song to be played based on the sound effect selected by the user. Specifically, the audio attributes of the original audio data of the song to be played are modified to the audio attributes corresponding to the sound effect selected by the user. Among them, the audio attributes can include attributes such as timbre, sense of space, and sense of surround.
[0025] For example, when the user selects dynamic electro music, regardless of the style of the song to be played itself, its audio attributes are modified to the audio attributes corresponding to the dynamic electro music sound effect. In this way, even some relatively gentle light music will become dynamic. However, in some songs, different segments may have different styles. Therefore, using a single sound effect to process such a song cannot meet the diverse auditory needs of users.
[0026] In some implementation manners, the music playing software can also provide intelligent sound effects for the user to select. When the user selects the intelligent sound effect, the electronic device can intelligently select the corresponding sound effect according to the style label of the song to be played. In this way, different styles of songs can be processed with different sound effects to meet the user's needs.
[0027] However, since the style labels of songs are pre-labeled and stored in the music software company's own information library and are not public information. Therefore, in the actual application process, the electronic device cannot obtain the style label of the song to be played, and thus the electronic device cannot select the corresponding sound effect for processing according to the style label of the song to be played. Moreover, the intelligent sound effect is also a single sound effect. For a song with different styles in different segments, using a single sound effect to process such a song cannot meet the diverse auditory needs of users.
[0028] In addition, in some implementation manners, the electronic device can intelligently select a corresponding sound effect for processing according to the album to which the song to be played belongs or the style of the singer who sings the song. Similarly, for a song with different styles in different segments, using a single sound effect to process such a song cannot meet the diverse auditory needs of users.
[0029] To solve the above problems, an embodiment of the present application provides an audio processing method. In this method, the electronic device obtains the original audio data (i.e., the first audio data) of the song to be played, and performs frame division processing on the first audio data to obtain a plurality of audio frame data. The electronic device can perform style prediction on each audio frame data to obtain the style label corresponding to each audio frame data. Furthermore, the electronic device can perform different sound effect processing on the audio frame data with different style labels in the first audio data to obtain the processed audio data (i.e., the second audio data).
[0030] In this way, by identifying different styles of different segments in the song to be played and performing different sound effect processing on segments with different styles, the song to be played can sound more prominent in style, and the differences between segments with different styles are greater and more personalized, so as to meet the diverse auditory needs of users.
[0031] As Figure 1 shown, a schematic flowchart of an audio processing method is shown. AsFigure 2 As shown, a flowchart of an audio processing method is presented. The following will combine Figure 1 and Figure 2 to introduce the audio processing method mentioned in the embodiments of this application in detail.
[0032] It can be understood that this audio processing method can be executed by an electronic device, such as the smartphone mentioned above. Specifically, this audio processing method may include:
[0033] S101: Obtain the first audio data of the song to be played.
[0034] It can be understood that the music playback software may have a playlist. In some implementation manners, the electronic device may identify the song currently being played in the playlist as the song to be played, or may also identify the next song of the song currently being played in the playlist as the song to be played.
[0035] If there is no currently playing song in the playlist, the electronic device may identify the first song in the playlist as the song to be played.
[0036] S102: Perform frame division processing on the first audio data to obtain multiple audio frame data.
[0037] In some optional implementation manners, after obtaining the first audio data of the song to be played, the electronic device may perform windowing processing on the first audio data to divide the first audio data into multiple audio frame data. Since the song to be played will not undergo a drastic style mutation within a small time window, that is, the song to be played will not change from one style to another within a very short time, therefore, the time window can be a short time window of 1s.
[0038] Suppose the total duration of the first audio data is 3 minutes. In some frame division processing manners, the electronic device may use the audio data within 0 - 1s as the first audio frame data, the audio data within 1 - 2s as the second audio frame data... and the audio data within 179 - 180s as the 180th audio frame data.
[0039] In some other frame division processing manners, the electronic device may use the audio data within 0 - 1s as the first audio frame data, the audio frame data within 0.5 - 1.5 as the second audio frame data... and the audio data within 179 - 180s as the 360th audio frame data.
[0040] It can be understood that the window length of the time window listed above is only an example listed in the embodiments of this application. In actual application processes, the time window may also be other window lengths, which are not limited herein.
[0041] S103: Classify each of the multiple audio frame data to determine the style label of each audio frame data.
[0042] In some alternative implementation manners, the electronic device may use a Back Propagation (BP) neural network to classify each of the audio frame data to determine the style label of each audio frame data. Specifically, the electronic device may extract features from each of the audio frame data, input the extracted features into the back propagation neural network, and use the back propagation neural network to classify each of the audio frame data based on these features to obtain the style label of each audio frame data.
[0043] It can be understood that features can represent certain prominent properties of audio frame data. In a classification task, the extracted features should be as independent of each other as possible. For this purpose, when extracting features from audio frame data, the time domain characteristics and frequency domain characteristics of the audio frame data can be analyzed, and the waveform factor and Mel-Frequency Cepstral Coefficients (MFCC) can be selected as the extracted features, and these values can be combined into a feature vector group as the input of the back propagation network.
[0044] Among them, the waveform factor is the ratio of the effective value to the absolute mean value of the audio frame data in one cycle, where the effective value is also called the root mean square. Specifically, the waveform factor can be calculated using the following formula (1):
[0045]
[0046] Among them, m can represent the waveform factor, RMS can represent the effective value, ARV can represent the absolute mean value, and x i (k) can represent the time domain signal corresponding to the audio frame data, and n can represent the number of sampling points in the audio frame data.
[0047] The human ear's perception of sound frequency is not linear. In the low-frequency range, the human ear can distinguish small frequency changes. In the high-frequency range, a larger frequency change is required for the human ear to perceive. The Mel scale is designed to simulate this non-linear perception of the human ear. When processing audio frame data, through means such as Fourier Transform (FFT), the time-domain signal corresponding to the audio frame data can be converted into a frequency-domain signal to obtain the energy distribution of each frequency, that is, the energy spectrum of the audio frame data. Since the energy range of the audio frame data is very wide, then, the energy spectrum of the audio frame data can be logarithmically calculated to obtain the logarithmic energy spectrum to compress the energy range of the audio frame data. Then, a linear transformation can be performed on the logarithmic energy spectrum to integrate the useful information in the logarithmic energy spectrum to obtain the Mel Frequency Cepstral Coefficient, making the audio frame data easier to be recognized and processed by electronic devices. Specifically, the Mel Frequency Cepstral Coefficient can be calculated using the following formula (2):
[0048]
[0049] Among them, mel(f) can represent the Mel Frequency Cepstral Coefficient, 2595 and 700 are parameters for frequency scale conversion, and f can represent frequency.
[0050] Moreover, it can be understood that the backpropagation neural network is a multi-layer feedforward network trained by error backpropagation. Its basic idea is the gradient descent method, and gradient search technology can be used to minimize the mean square error between the actual output value and the expected output value of the network. That is, the backpropagation neural network includes two processes: the forward propagation of signals and the backpropagation of errors. In other words, when calculating the error output, it is carried out in the direction from input to output, while when adjusting the weights (w) and thresholds (b), it is carried out in the direction from output to input.
[0051] As Figure 3 shown, a schematic structural diagram of a backpropagation neural network is shown. The backpropagation neural network includes an input layer, a hidden layer, and an output layer. The above-mentioned feature vector group can be used as the input of the input layer. The output of the input layer can be the input of the hidden layer. The output of the hidden layer can be used as the input of the output layer. The output of the output layer can include style labels.
[0052] Among them, the input layer can perform weighted calculations on the feature vector group. The hidden layer can continue to perform weighted calculations on the calculation results of the input layer and perform non-linear operations on the further weighted calculation results. Similarly, the output layer can continue to perform weighted calculations and non-linear operations on the calculation results of the hidden layer. Different from the hidden layer, the output layer will perform logistic regression on the calculation results and use the logistic regression results as the final output. If the mean square error of the actual output value and the expected output value of the network is less than the mean square error threshold, that is, the mean square error of the actual output value and the expected output value is not the smallest, then the error backpropagation process is entered. The error backpropagation process refers to: the error is propagated backward layer by layer from the hidden layer to the input layer, and the error is allocated to all units in each layer, so as to adjust the weights of each layer.
[0053] During the training process of the backpropagation neural network, the labeled style labels of the sample audio data can include dynamic and gentle. Similarly, the sample audio data can be first subjected to feature extraction, and the extracted features can be combined to obtain a feature vector group, which is input into the backpropagation neural network to be trained. The backpropagation neural network to be trained can output the predicted style labels of each sample audio data based on the feature vector group. Among them, "0" can indicate that the style label of the sample audio data is dynamic, and "1" can indicate that the style label of the sample audio data is gentle. Furthermore, the weights and thresholds of the backpropagation neural network to be trained can be adjusted according to the error between the labeled style label and the predicted style label of each sample audio data until the mean square error of the actual output value and the expected output value of the trained backpropagation neural network is less than the mean square error threshold, and the training of the backpropagation neural network is completed.
[0054] It can be understood that usually a song shows different styles in the form of fragments. For example, the style of the prelude part of the song is gentle, and the style of the chorus part is dynamic. And, for any one of the multiple audio frame data and several adjacent audio frame data to it, they are usually of the same style. Therefore, for any one of the multiple audio frame data, the electronic device can count the style labels of several adjacent audio frame data to it and use the style label with the largest proportion as the style label of this audio frame data.
[0055] For example, for the second audio frame data among the 180 audio frame data listed above, assume that the electronic device determines that the style label of the first audio frame data is gentle, the style label of the second audio frame data is dynamic, and the style label of the third audio frame data is gentle. In this way, the electronic device can count the style labels of the first audio frame data, the second audio frame data, and the third audio frame data, and determine that the style label of the second audio frame data is gentle.
[0056] Thus, in the process of determining the style tags of each audio frame data, by comprehensively considering the style tags of multiple audio frame data adjacent to each audio frame data, the accuracy and stability of determining the style tags of each audio frame data can be further improved on the basis of the recognition of the backpropagation model (generally with an accuracy of 60% - 80%).
[0057] S104: Process the first audio data based on the sound effect processing parameters corresponding to the style tags of each audio frame data to obtain the second audio data.
[0058] It can be understood that the sound effect processing parameters corresponding to the style tags may include the frequency bands to be adjusted and the processing methods to be performed on the frequency bands, such as boosting or attenuating. In some alternative implementation manners, as Figure 4 shown, the boosting or attenuating of different frequency bands in each audio frame data of the first audio data can be performed through multiple groups of parallel equalizers (EQs) to adjust the timbre, sense of space, sense of surround, etc. corresponding to the audio frame data.
[0059] It can be understood that the audio processing method provided in the embodiments of the present application can be applied to an electronic device. Next, an exemplary introduction to the hardware structure of the electronic device applicable to the audio processing method provided in the embodiments of the present application will be given.
[0060] As Figure 5 shown, the electronic device 500 may include a processor 510, an external memory interface 520, an internal memory 521, a universal serial bus (USB) interface 530, a charging management module 540, a power management module 541, a battery 542, an antenna, a wireless communication module 550, an audio module 560, a speaker 560A, a receiver 560B, a microphone 560C, a headphone interface 560D, a camera 570, a display screen 580, etc.
[0061] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 500. In other embodiments of the present application, the electronic device 500 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0062] The processor 510 may include one or more processing units. For example, the processor 510 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0063] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions. The processor 510 may control fetching and executing instructions through the controller to implement the audio processing method provided in the embodiments of the present application. For example, the processor 510 may control fetching and executing instructions through the controller to implement the Figure 1 、 Figure 2 or Figure 3 corresponding steps implemented in the processes shown.
[0064] A memory may also be provided in the processor 510 for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory may save the instructions or data that the processor 510 has just used or recycled. If the processor 510 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.
[0065] The wireless communication function of the electronic device may be implemented through an antenna, a wireless communication module 550, a modem processor, a baseband processor, etc.
[0066] The antenna is used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device may be used to cover a single or multiple communication frequency bands. Different antennas may also be multiplexed to improve the utilization rate of the antennas. For example, the antenna may be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.
[0067] The wireless communication module 550 may provide solutions for wireless communications applied to an electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 550 may be one or more devices integrating at least one communication processing module. The wireless communication module 550 receives electromagnetic waves via an antenna, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 510. The wireless communication module 550 may also receive signals to be sent from the processor 510, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna for radiation.
[0068] The electronic device realizes the display function through a GPU, a display screen 580, an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 580 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 510 may include one or more GPUs, which execute program instructions to generate or change display information.
[0069] The display screen 580 is used to display images, videos, etc. The display screen 580 includes a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N display screens 580, where N is a positive integer greater than 1.
[0070] In some cases, the embodiments disclosed in this application may be implemented in hardware, firmware, software, or any combination thereof.
[0071] The embodiments disclosed in this application can also be implemented as instructions carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, magneto-optical discs, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in the form of electrical, optical, acoustic, or other propagated signals using the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0072] Embodiments of this application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0073] The program code can be applied to the input instructions to perform the various functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit, or a microprocessor.
[0074] The program code can be implemented in a high-level procedural language or an object-oriented programming language in order to communicate with the processing system. When necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any particular programming language. In any case, the language can be a compiled language or an interpreted language.
[0075] The above introduced the possible hardware structure of the electronic device. It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device can include more or fewer components than shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0076] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0077] It should be noted that in the examples and description of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0078] Although this application has been illustrated and described by reference to certain embodiments thereof, those of ordinary skill in the art should understand that various changes may be made in form and detail without departing from the scope of this application.
Claims
1. An audio processing method, characterized in that, Applied to an electronic device, including: Obtain first audio data of a song to be played; Perform frame splitting on the first audio data to obtain a plurality of audio frame data; Perform classification processing on each of the plurality of audio frame data to determine style labels for each of the audio frame data; Process the first audio data based on audio processing parameters corresponding to the style labels of each of the audio frame data to obtain second audio data.
2. The method according to claim 1, characterized in that, The performing frame splitting on the first audio data to obtain a plurality of audio frame data includes: Perform frame splitting on the first audio data based on a preset duration to obtain a plurality of audio frame data.
3. The method according to claim 1, wherein The performing classification processing on each of the plurality of audio frame data to determine style labels for each of the audio frame data includes: Perform feature extraction processing on each of the plurality of audio frame data to obtain a feature vector group for each of the audio frame data; Determine style labels for each of the audio frame data based on the feature vector groups of each of the audio frame data.
4. The method according to claim 3, wherein The feature vector group of each of the audio frame data includes a waveform factor and a Mel-frequency cepstral coefficient.
5. The method according to claim 3, characterized in that, The determining style labels for each of the audio frame data based on the feature vector groups of each of the audio frame data includes: Input the feature vector groups of each of the audio frame data into a reverse neural network, and use the reverse neural network to perform classification processing on each of the audio frame data to determine style labels for each of the audio frame data.
6. The method according to claim 5, wherein The determining style labels for each of the audio frame data includes: Determine, as the style label for each of the audio frame data, the style label with the largest proportion among the style labels of a plurality of audio frame data adjacent to each of the audio frame data.
7. The method according to claim 1, wherein The processing the first audio data based on audio processing parameters corresponding to the style labels of each of the audio frame data to obtain second audio data includes: Based on a plurality of frequency bands to be adjusted corresponding to the style labels of each of the audio frame data and processing methods corresponding to each of the frequency bands to be adjusted, obtain audio processing parameters corresponding to the style labels of each of the audio frame data; Adjust the plurality of frequency bands to be adjusted corresponding to the style labels of each of the audio frame data based on the processing methods corresponding to each of the frequency bands to be adjusted to obtain the second audio data.
8. An electronic device, characterized in that, Including: A memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the audio processing method according to any one of claims 1-7.
9. A readable storage medium, characterized in that, Instructions are stored on the readable storage medium, and when the instructions are executed on the electronic device, the electronic device executes the audio processing method according to any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product includes computer instructions, and when executed by the electronic device, the electronic device executes computer program code of the audio processing method according to any one of claims 1-7.
Citation Information
Cited By
Audio processing method, audio processing device and electronic equipment
CN120600050A