Audio processing method and device, equipment and storage medium
By converting the audio signal from the time domain to the frequency domain, using a timbre recognition model to detect and apply timbre operators to enhance the signal, the problem of the inability to accurately identify and enhance timbre in existing technologies is solved, achieving efficient enhancement of target timbre and targeted and accurate audio processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-03-31
AI Technical Summary
Existing audio processing technologies are unable to precisely identify and enhance specific timbres, resulting in a lack of targeted and accurate audio processing, and a serious waste of computing resources and time.
通过将音频信号从时域转换为频域,利用音色识别模型检测目标音色,并应用相应的音色算子进行信号增强,仅在检测到目标音色时进行增强,避免对其他音色产生干扰。
It achieves efficient enhancement of the target timbre, saves computing resources and processing time, maintains the natural balance and overall quality of the audio, and provides a better listening experience.
Smart Images

Figure CN121768408A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an audio processing method, apparatus, device, and storage medium. Background Technology
[0002] With the development of computer and electronic device technologies, audio is being used in increasingly diverse business applications, leading to the advancement of audio processing technologies. In some audio processing techniques, audio developers can enhance sound effects to improve the immersive experience, enhance sound quality and clarity, and provide users with a more realistic auditory experience.
[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] The purpose of this disclosure is to provide an audio processing method, apparatus, device, and storage medium.
[0005] According to a first aspect of the present disclosure, an audio processing method is provided, comprising: acquiring an audio to be played; converting time-domain data of the audio to be played into frequency-domain data; detecting a timbre type contained in the audio to be played based on the frequency-domain data; in response to the presence of a target timbre in the timbre type, acquiring a timbre operator corresponding to the target timbre; and performing signal enhancement on the frequency-domain data for the target timbre according to the timbre operator to obtain an audio to be output.
[0006] In some implementations, detecting the timbre type contained in the audio to be played based on the frequency domain data includes: acquiring a trained timbre recognition model; processing the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played; and determining the timbre type contained in the audio to be played based on the timbre recognition result.
[0007] In some embodiments, the timbre recognition model includes a first fully connected layer, multiple convolutional layers, a gated recurrent layer, and a second fully connected layer. The process of processing the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played includes: processing the frequency domain data through the first fully connected layer to obtain a first frequency domain feature; processing the first frequency domain feature through the multiple convolutional layers to obtain a second frequency domain feature; processing the second frequency domain feature through multiple gated recurrent units in the gated recurrent layer to obtain multiple time-series features; and processing the multiple time-series features through the second fully connected layer to obtain the timbre recognition result.
[0008] In some implementations, the frequency domain data is enhanced with the timbre-specific signal according to the timbre operator to obtain the audio to be output, including: performing convolution processing on the timbre operator and the frequency domain data to obtain enhanced frequency domain data; and determining the audio to be output based on the enhanced frequency domain data.
[0009] In some implementations, determining the audio to be output based on the enhanced frequency domain data includes: smoothing the enhanced frequency domain data to obtain smoothed enhanced data; and converting the smoothed enhanced data into time domain data of the audio to be output.
[0010] In some embodiments, the audio processing method further includes: acquiring user preference selection information after the audio to be output is played; the user preference selection information indicates the operator identification information of the timbre operator; and establishing a correspondence between the operator identification information, the target timbre, and the user information.
[0011] In some implementations, obtaining the timbre operator corresponding to the target timbre includes: obtaining current user information using the audio to be played; and, in response to the existence of operator identifier information matching the target timbre and / or the current user information, determining the timbre operator corresponding to the target timbre based on the operator identifier information.
[0012] In some implementations, converting the time-domain data of the audio to be played into frequency-domain data includes: determining the business scenario of the audio to be played; and converting the time-domain data of the audio to be played into frequency-domain data in response to the business scenario meeting processing conditions; wherein the processing conditions include: the maximum allowed delay duration of the business scenario is greater than or equal to a delay threshold, or the business scenario belongs to a preset scenario set.
[0013] In some implementations, the target timbre includes at least one of the following: human voice, light music, mechanical sound, and ambient sound.
[0014] According to a second aspect of the present disclosure, an audio processing apparatus is provided, comprising: an audio conversion unit configured to acquire audio to be played and convert time-domain data of the audio to be played into frequency-domain data; a detection unit configured to detect a timbre type contained in the audio to be played based on the frequency-domain data; an operator acquisition unit configured to acquire a timbre operator corresponding to the target timbre in response to the presence of a target timbre in the timbre type; and an output unit configured to perform signal enhancement on the frequency-domain data for the target timbre according to the timbre operator to obtain audio to be output.
[0015] In some implementations, the detection unit detects the timbre type contained in the audio to be played based on the frequency domain data, including: acquiring a trained timbre recognition model; processing the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played; and determining the timbre type contained in the audio to be played based on the timbre recognition result.
[0016] In some implementations, the timbre recognition model includes a first fully connected layer, multiple convolutional layers, a gated recurrent layer, and a second fully connected layer. The detection unit processes the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played, including: processing the frequency domain data through the first fully connected layer to obtain a first frequency domain feature; processing the first frequency domain feature through the multiple convolutional layers to obtain a second frequency domain feature; processing the second frequency domain feature through multiple gated recurrent units in the gated recurrent layer to obtain multiple time-series features; and processing the multiple time-series features through the second fully connected layer to obtain the timbre recognition result.
[0017] In some implementations, the output unit enhances the frequency domain data for the target timbre based on the timbre operator to obtain the audio to be output, including: performing convolution processing on the timbre operator and the frequency domain data to obtain enhanced frequency domain data; and determining the audio to be output based on the enhanced frequency domain data.
[0018] In some implementations, the output unit determines the audio to be output based on the enhanced frequency domain data, including: smoothing the enhanced frequency domain data to obtain smoothed enhanced data; and converting the smoothed enhanced data into time domain data of the audio to be output.
[0019] In some embodiments, the audio processing apparatus further includes a preference acquisition unit and a preference establishment unit. The preference acquisition unit is used to acquire user preference selection information after the audio to be output is played. The user preference selection information indicates the operator identification information of the timbre operator. The preference establishment unit is used to establish the correspondence between the operator identification information, the target timbre, and the user information.
[0020] In some implementations, the operator acquisition unit acquires a timbre operator corresponding to the target timbre, including: acquiring current user information using the audio to be played; and in response to the existence of operator identifier information matching the target timbre and / or the current user information, determining a timbre operator corresponding to the target timbre based on the operator identifier information.
[0021] In some implementations, the audio conversion unit converts the time-domain data of the audio to be played into frequency-domain data, including: determining the business scenario of the audio to be played; and converting the time-domain data of the audio to be played into frequency-domain data in response to the business scenario meeting processing conditions; wherein the processing conditions include: the maximum allowed delay duration of the business scenario is greater than or equal to a delay threshold, or the business scenario belongs to a preset scenario set.
[0022] In some implementations, the target timbre includes at least one of the following: human voice, light music, mechanical sound, and ambient sound.
[0023] According to a third aspect of the present disclosure, an electronic device is provided, characterized in that it includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-described audio processing method.
[0024] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to execute an audio processing method, the method comprising: acquiring audio to be played; converting time-domain data of the audio to be played into frequency-domain data; detecting a timbre type contained in the audio to be played based on the frequency-domain data; in response to the presence of a target timbre in the timbre type, acquiring a timbre operator corresponding to the target timbre; and performing signal enhancement on the frequency-domain data for the target timbre according to the timbre operator to obtain audio to be output.
[0025] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the audio processing method described above.
[0026] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0027] This disclosure, on the one hand, can detect and identify the target timbre in the audio to be played based on frequency domain analysis technology, and then apply targeted timbre operators to enhance the signal, ultimately obtaining an audio output with superior sound quality for the target timbre. On the other hand, this selective enhancement, which applies timbre operators only when the target timbre is detected, avoids the indiscriminate use of timbre operators on all audio to be played. This not only saves computational resources and audio processing time, but also effectively avoids unnecessary interference with other timbres, achieving targeted and precise audio processing.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0030] Figure 1 This is a flowchart illustrating an audio processing method according to some embodiments of the present disclosure.
[0031] Figure 2 This is a flowchart illustrating another audio processing method according to some embodiments of the present disclosure.
[0032] Figure 3 This is a schematic diagram of a timbre recognition model in an audio processing method according to some embodiments of the present disclosure.
[0033] Figure 4 This is a schematic diagram illustrating the acquisition of timbre recognition results through a timbre recognition model in an audio processing method according to some embodiments of the present disclosure.
[0034] Figure 5 This is a flowchart illustrating an audio processing method according to some embodiments of the present disclosure.
[0035] Figure 6 This is a flowchart illustrating another audio processing method according to some embodiments of the present disclosure.
[0036] Figure 7 This is a block diagram illustrating an audio processing apparatus according to some embodiments of the present disclosure.
[0037] Figure 8 This is a block diagram illustrating an apparatus for audio processing according to some embodiments of the present disclosure. Detailed Implementation
[0038] Exemplary embodiments of this disclosure will be described in detail herein, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.
[0039] The embodiments described below, which are examples of some of the embodiments of this disclosure, do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0040] In related technologies, Fourier transform is commonly used to convert the time-domain amplitude spectrum of an audio signal into its frequency-domain amplitude spectrum, and then enhance the amplitude of a specific frequency band (such as the low-frequency band) to achieve signal enhancement. However, Fourier transform cannot perform more refined filtering operations, cannot identify and enhance specific timbres, and cannot highlight a particular timbre.
[0041] The specific implementation of the embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0042] Figure 1 This is a flowchart illustrating an audio processing method according to some embodiments of the present disclosure, such as... Figure 1 As shown, the audio processing method may include the following steps.
[0043] In step S110, the audio to be played is obtained, and the time-domain data of the audio to be played is converted into frequency-domain data.
[0044] In this embodiment of the disclosure, the original audio to be played can be obtained from an audio source (such as a locally stored file, a network streaming media, etc.). The original audio to be played is usually in the time domain (i.e., time domain). In the time domain, the audio signal is represented by the change in amplitude over time; that is, time domain data can represent the change of the audio signal over time. In order to analyze and process the audio signal more effectively, the time domain data can be converted into frequency domain data.
[0045] In an exemplary embodiment, the audio signal can be converted from time-domain data to frequency-domain data using either the STFT (short-time Fourier transform) algorithm or the FFT (fast Fourier transform). In the frequency domain, the audio signal is decomposed into components of different frequencies, each with a specific sound wave intensity. Therefore, the audio signal can be represented as a combination of different frequency components and their corresponding sound wave intensities, allowing for frequency-based analysis and processing in subsequent steps.
[0046] In step S120, the timbre type contained in the audio to be played is detected based on the frequency domain data.
[0047] In this embodiment of the disclosure, in the frequency domain representation of audio, the timbre characteristics of an audio signal can be presented through its spectral distribution (i.e., the energy distribution of the audio signal at different frequencies) and harmonic structure (referring to the distribution of the fundamental frequency and its integer multiples of frequency components in the audio signal). Machine learning models, pattern recognition algorithms, or preset timbre template matching methods can be used to detect the timbre type contained in the audio to be played. The timbre type can include human voices (such as male and female voices), instrument sounds (such as drum kits and guitars), etc.
[0048] In step S130, in response to the existence of a target timbre in the timbre type, a timbre operator corresponding to the target timbre is obtained.
[0049] In this embodiment of the disclosure, the target timbre can be a specific timbre that the user pre-sets to be enhanced. When the target timbre is detected in the audio, the timbre operator corresponding to that timbre can be found and obtained. The timbre operator can be used to enhance the signal of the specific timbre in the frequency domain.
[0050] In an exemplary embodiment, the timbre operator can be a filter (such as a low-pass, high-pass, or band-pass filter) or a gain function that is pre-customized according to the characteristics of the corresponding target timbre. The filter can enhance the timbre by adjusting the frequency components of the audio signal, while the gain function can enhance the timbre by controlling the amplification of different frequency components.
[0051] For example, if a bandpass filter is selected as the timbre operator, the center frequency, bandwidth and gain of the filter can be set according to the characteristics of the target timbre (such as frequency distribution, harmonic structure, dynamic range, etc.) so that signals within a specific frequency range of the target timbre can be passed through, while signals outside that range can be attenuated or blocked, thereby ensuring that the target timbre is highlighted.
[0052] For example, if a gain function is chosen as the timbre operator, different gain values can be set based on the spectral characteristics of the target timbre and other timbres. This allows different amplification or attenuation coefficients to be applied to signals of different frequencies, thereby achieving precise adjustment of the gain values of each frequency and thus precisely shaping and enhancing the specific frequency components of the target timbre.
[0053] In an exemplary embodiment, either the filter or the gain function can be used alone to customize the timbre operator, or the two can be used in combination to customize the timbre operator; this disclosure does not limit this.
[0054] In an exemplary embodiment, the target timbre can be a small-signal type timbre, characterized by a low frequency (e.g., below a preset frequency threshold) or low amplitude (e.g., below a preset amplitude threshold). Based on this, small-signal type timbres often present a soft and relatively muted tone. However, small-signal type timbres (e.g., whispers, wind sounds, light music sounds) can often carry rich detail. For example, in the context of music, these small signals can represent subtle changes in instruments, the breath of a performer, or slight sounds in the environment, adding depth and nuance to the overall timbre. Furthermore, small-signal type timbres are easily masked by other, stronger signals. Therefore, in audio processing or music production, identifying and enhancing small-signal type timbres can help them be better represented in the overall timbre.
[0055] For example, in some audio or video recordings, mechanical sounds that are far from the recording point or the inherently faint sound of machine operation are easily affected by ambient noise due to their low amplitude, thus exhibiting the characteristics of a small signal timbre in some scenarios. If the target timbre is set to a mechanical sound that conforms to the characteristics of a small signal, then such a timbre will be recognized.
[0056] For example, in some audio or video recordings, faint ambient sounds recorded in quiet environments (such as distant wind, drizzle, or traffic noise) can also be considered small-signal timbres because their amplitude is low and easily overlooked. If the target timbre is set to an ambient sound that matches the characteristics of a small signal, then such a timbre will also be recognized.
[0057] Furthermore, small-signal timbres, being low-frequency and low-amplitude sound signals in audio signals, can be distinguished using a spectrogram. Traditional Fourier algorithms can only differentiate sound signals by frequency band, not by specific sound type. Therefore, this solution, building upon the Fourier algorithm, utilizes deep learning models to effectively identify specific timbres in the spectrum and then enhances them, achieving the effect of strengthening specific sound types while maintaining the hierarchy and distinct layers of the enhanced audio.
[0058] In some embodiments of this disclosure, the target timbre includes at least one of the following: human voice, light music sound, mechanical sound, and ambient sound.
[0059] In this embodiment of the disclosure, human voice, light music, mechanical sound, and ambient sound are several specific timbre types among small signal type timbre.
[0060] By enhancing the human voice, you can ensure that the speaker's or singer's voice is clearly identifiable in mixed audio, such as in movies, music, or podcasts, thereby improving the listenability and comprehensibility of the audio.
[0061] In an exemplary embodiment, the target timbre can also be a subtle breathing sound, whisper, or low murmur in human voice. These subtle breathing sounds, whispers, and low murmurs are subcategories of small signals in human voice. Enhancing these timbres can make them more distinct in mixed audio, easier for listeners to recognize and understand, and can also help improve the signal-to-noise ratio of the audio, making these sounds purer and more prominent. For example, in situations where careful listening to whispers is required, detecting and enhancing whispers in the audio can provide listeners with a more natural and comfortable listening experience.
[0062] Soft music can include gentle instrument sounds, chords, and melodies. In certain scenarios, such as meditation apps, background music, or relaxing environments, enhancing soft music can create a more tranquil and pleasant atmosphere.
[0063] Mechanical sounds can be sounds produced by machines, equipment, or tools. In industrial, scientific, or educational videos, enhancing mechanical sounds can help viewers hear and understand more clearly how machines work or how they operate.
[0064] Ambient sounds can include sounds from natural environments (such as wind and flowing water) and urban environments (such as traffic noise and crowd noise). In documentaries, game sound effects, or virtual reality applications, enhanced ambient sounds can provide a more immersive experience, helping users better perceive the realism of a scene.
[0065] In step S140, the frequency domain data is enhanced with signal for the target timbre according to the timbre operator to obtain the audio to be output.
[0066] In this embodiment of the disclosure, based on the acquired timbre operator, targeted signal enhancement processing is performed on the frequency domain data. This can include increasing the amplitude of specific frequency components, adjusting the frequency response curve, and other similar processing. The enhanced frequency domain data can then be converted back to the time domain using an inverse Fourier transform to generate enhanced output audio, making the target timbre more prominent and clear during playback.
[0067] Furthermore, when using timbre operators to process audio, it is inevitable that other timbres will be affected. However, with this embodiment, timbre operators can be applied for enhancement only when the target timbre (such as whispers in human voices, soft music, mechanical sounds, or ambient sounds) is detected. This selective enhancement not only ensures that only the target timbre is highlighted and improves its texture and recognizability, but also avoids the indiscriminate use of timbre operators on all audio to be played. This saves computing resources and audio processing time, and avoids unnecessary enhancements that may have an undue impact on other timbres, thus achieving targeted and precise audio processing.
[0068] In an exemplary embodiment, when the target timbre is set to a small-signal type, since small signals are characterized by low frequency or low amplitude, their overlap with other timbres in the audio spectrum is relatively small. Therefore, when using timbre operators to perform signal enhancement processing on these specific small-signal timbres, the effect can be more precise on the target timbre, while relatively reducing the impact on other non-target timbres in the audio. This selective enhancement strategy not only ensures that the target timbre is significantly highlighted but also better maintains the overall natural balance of the audio, reducing unnecessary interference with other timbres.
[0069] As can be seen from the above steps, the audio processing method provided in this disclosure can perform frequency domain transformation on the audio to be played, detect whether a pre-set target timbre exists in the audio based on the frequency domain data, and if the target timbre exists, acquire and use the timbre operator corresponding to the target timbre to process the frequency domain data, thereby achieving signal enhancement of the target timbre in the audio to be played. Therefore, this solution can, on the one hand, detect and identify the target timbre in the audio to be played based on frequency domain analysis technology, and then apply targeted timbre operators to enhance the signal, ultimately obtaining an audio output with superior sound quality for the target timbre. On the other hand, this selective enhancement, which applies the timbre operator only when the target timbre is detected, avoids indiscriminately using the timbre operator on all audio to be played, saving computational resources and audio processing time, and effectively avoiding unnecessary interference to other timbres, thus achieving targeted and accurate audio processing.
[0070] Figure 2 This is a flowchart illustrating another audio processing method according to some embodiments of the present disclosure.
[0071] In this embodiment of the disclosure, Figure 2 In the audio processing method shown, steps S210, S250, and S260 are respectively related to... Figure 1 The steps S110, S130, and S140 in the audio processing method shown correspond to each other and will not be repeated here.
[0072] In some embodiments of this disclosure, in Figure 1 Based on the audio processing method shown, Figure 2 The audio processing method shown may also include the following steps.
[0073] Step S220: Obtain the trained timbre recognition model.
[0074] In this embodiment of the disclosure, the timbre recognition model can be a model trained by machine learning technology that can recognize and process timbre features in audio signals.
[0075] For example, an initial timbre recognition model can be built in advance based on a deep learning neural network or other classifiers, and a training set can be prepared for training. The training set can be used to train the initial timbre recognition model to recognize timbre types, resulting in a trained timbre recognition model.
[0076] In an exemplary embodiment, the timbre recognition model may be a CNN (Convolutional Neural Network) model, an RNN (Recurrent Neural Network) model, a CRNN (Convolutional Recurrent Neural Network) model, etc., and this disclosure does not limit it.
[0077] Step S230: Process the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played.
[0078] In this embodiment of the disclosure, after the audio is converted into frequency domain data, this frequency domain data can be input into a timbre recognition model. The model analyzes the characteristics of the frequency domain data, such as spectral shape and energy distribution, to identify the timbre information contained therein. After processing by the model, one or more timbre recognition results can be output.
[0079] In an exemplary embodiment, the timbre recognition result can be a feature vector of one or more specified timbres, a specific timbre type label, a string indicating the existence of one or more specified timbres, etc., and this disclosure does not limit this. The form and content of the model output can be determined according to actual needs, and the sample labels in the model's training set can be determined based on this. For example, if the desired output is a timbre type label, then each sample in the training set should be labeled with the corresponding timbre type. As another example, if the desired output is a feature vector or string representation, then the samples in the training set are also labeled in the corresponding form.
[0080] Step S240: Determine the timbre type contained in the audio to be played based on the timbre recognition result.
[0081] In this embodiment of the disclosure, the actual timbre type contained in the audio to be played can be determined based on the timbre recognition result.
[0082] For example, if the timbre recognition result is a feature vector of one or more specified timbres, the final timbre type can be determined further through feature comparison, threshold judgment, or cluster analysis. If the timbre recognition result is one or more specific timbre type labels, then these labels are the timbre types contained in the audio. If the timbre recognition result is a string indicating the existence of one or more specified timbres, then the existence of the specified timbre corresponding to each position can be determined based on the information of the values at each position in the string, thereby obtaining information about the timbre types contained in the audio to be played.
[0083] In an exemplary embodiment, the timbre recognition result in string form can use a JSON object as the output string, which can store the existence information of each specified timbre, where the key is the timbre type and the value is true or false (or 1 and 0) to indicate whether the timbre exists. This form of result is both flexible and easy to integrate with other systems.
[0084] Through the embodiments of this disclosure, an audio processing method based on frequency domain data and timbre recognition model can be used to automatically detect and identify the timbre type in audio, providing strong support for subsequent signal enhancement, audio editing and other processing.
[0085] Figure 3 This is a schematic diagram illustrating a timbre recognition model in an audio processing method according to some embodiments of the present disclosure. The timbre recognition model may be a model based on CRNN technology.
[0086] like Figure 3 As shown, the timbre recognition model includes a first fully connected layer 301, multiple convolutional layers (here, multiple convolutional layers including convolutional layer 302 and convolutional layer 303 are illustrated, but this disclosure is not limited to this), a gated recurrent layer (here, multiple convolutional layers including gated recurrent unit 304 and gated recurrent unit 305 are illustrated, but this disclosure is not limited to this), and a second fully connected layer 306.
[0087] Figure 4 This is a schematic diagram illustrating the acquisition of timbre recognition results through a timbre recognition model in an audio processing method according to some embodiments of the present disclosure.
[0088] Combination Figure 3 In some embodiments of this disclosure, processing the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played may include the following steps.
[0089] Step S410: The frequency domain data is processed by the first fully connected layer 301 to obtain the first frequency domain feature.
[0090] The first fully connected layer receives the input frequency domain data and maps it to a new feature space. This layer performs preliminary integration and transformation of the input data so that subsequent layers can extract features more effectively. The first frequency domain features may contain some global or preliminary spectral information of the audio signal.
[0091] Step S420: The first frequency domain features are processed by the multiple convolutional layers (including convolutional layer 302 and convolutional layer 303) to obtain the second frequency domain features.
[0092] Convolutional layers are an effective tool for processing local features in audio signals. They apply convolutional kernels to input features using a sliding window approach, thereby extracting local patterns (such as frequency components and texture) in the audio signal. In timbre recognition models, multiple stacked convolutional layers can extract higher-level frequency domain features layer by layer. Each convolutional layer extracts features based on the output of the previous layer, thus constructing a hierarchical representation of the audio signal. After processing by multiple convolutional layers, the first frequency domain features can be transformed into second frequency domain features, which can contain richer timbre and structural information.
[0093] Step S430: The second frequency domain features are processed by multiple gated loop units (including gated loop unit 304 and gated loop unit 305) in the gated loop layer to obtain multiple time series features.
[0094] Among them, the gated recurrent layer is a special recurrent neural network (RNN) structure that can solve the gradient vanishing or gradient explosion problems that traditional RNN models are prone to when processing long sequences by introducing gating mechanisms (such as forget gate, input gate and output gate).
[0095] In timbre recognition models, gated recurrent layers can process second-frequency domain features through multiple gated recurrent units (such as LSTM or GRU). Gated recurrent units are powerful tools for processing sequential data, capable of capturing temporal dependencies within the sequence. Since audio signals possess time-series characteristics, recurrent neural networks can effectively capture these temporal dependencies. Specifically, multiple gated recurrent units can process each time step in the second-frequency domain features sequentially, resulting in multiple time-series features. These time-series features can contain frequency domain information of the audio signal, temporal dependencies, and reflect the dynamic changes of the audio signal over time, making them a crucial component of timbre recognition tasks.
[0096] Step S440: The multiple time-series features are processed by the second fully connected layer 306 to obtain the timbre recognition result.
[0097] The second fully connected layer can receive multiple time-series features output from the gated recurrent layer, perform further nonlinear transformation and dimensionality reduction on these time-series features, and map them onto the final timbre recognition result, outputting the final timbre recognition result.
[0098] In an exemplary embodiment, the second fully connected layer may use a softmax function or other classifiers to convert the feature vector into a probability distribution or a specific timbre type label, and output the timbre recognition result.
[0099] Through the embodiments of this disclosure, the timbre recognition model, by combining fully connected layers, convolutional layers, and gated recurrent layers, can effectively process frequency domain data and extract feature representations useful for timbre recognition tasks. This structural design fully considers the characteristics of audio signals, including their frequency domain characteristics and time-series characteristics, thereby ensuring the accuracy and effectiveness of the output results.
[0100] In some embodiments of this disclosure, performing signal enhancement on the frequency domain data according to the timbre operator to target the timbre and obtain the audio to be output may include: performing convolution processing on the timbre operator and the frequency domain data to obtain enhanced frequency domain data; and determining the audio to be output based on the enhanced frequency domain data.
[0101] In this embodiment of the disclosure, convolution processing can simulate many effects in signal processing, such as filtering, smoothing, and sharpening. Convolution processing based on the timbre operator and the frequency domain data can involve applying a specific filter function (i.e., the timbre operator) to the frequency domain data. This filter function should be designed to identify and highlight the frequency domain characteristics of the target timbre while minimizing the impact on non-target timbres.
[0102] The convolution operation iterates through every point in the frequency domain data and performs a weighted summation of the timbre operator (filter) with that point and its surrounding data. This enhances the frequency domain features of the target timbre, while other features remain relatively unchanged or are minimally affected.
[0103] After convolution processing, the original frequency domain data is converted into enhanced frequency domain data. Then, through inverse Fourier transform (IFFT) or other similar mathematical operations, the enhanced frequency domain data can be converted back into time domain data to obtain the final output audio.
[0104] The present invention discloses a reasonable and effective method for enhancing the target timbre by using timbre operators and frequency domain data for convolution processing. This method can accurately target the target timbre while maintaining the overall natural harmony of the audio.
[0105] In some embodiments of this disclosure, determining the audio to be output based on the enhanced frequency domain data includes: smoothing the enhanced frequency domain data to obtain smoothed enhanced data; and converting the smoothed enhanced data into time domain data of the audio to be output.
[0106] Smoothing can reduce noise or unwanted fluctuations in data during signal processing, making the data smoother and easier to analyze or apply. In embodiments of this disclosure, smoothing can help improve the quality of audio signals, making them sound more natural and fluid.
[0107] In an exemplary embodiment, enhanced frequency domain data can be processed by pre-set low-pass filtering, smoothing algorithms (such as moving average, median filtering), etc., to reduce noise or irregular fluctuations in the enhanced frequency domain data, making the enhanced frequency domain data smoother, helping to eliminate sharp jumps or irregularities in the enhanced frequency domain data, thereby improving audio quality.
[0108] After smoothing, a new set of smoother frequency domain data, namely smoothed enhanced data, can be obtained. Then, the smoothed enhanced data can be converted from the frequency domain back to the time domain (e.g., through inverse Fourier transform) to obtain the audio to be output, so that the user can hear it in the form of sound waves.
[0109] Through the embodiments of this disclosure, the enhanced frequency domain data can be optimized through smoothing processing to eliminate sharp jumps or irregularities in the enhanced frequency domain data, thereby generating high-quality audio to be output and providing a clearer and more stable listening experience.
[0110] In some embodiments of this disclosure, the method further includes: obtaining user preference selection information after the audio to be output is played; the user preference selection information indicates the operator identification information of the timbre operator; and establishing a correspondence between the operator identification information, the target timbre, and the user information.
[0111] In this embodiment, a user feedback mechanism can be introduced to optimize and enrich the personalized effects in the audio processing method. By obtaining the user's preference selection information after the output audio is played, the user's satisfaction with the current sound processing effect can be understood.
[0112] In an exemplary embodiment, user preference selection information may be collected in the form of explicit user selections or implicit behaviors (such as playback duration, number of replays, etc.).
[0113] For example, controls can be provided to users to collect feedback, such as "like," "dislike," "favorite," and rating controls, allowing users to provide feedback through ratings, likes, and selecting "like" or "dislike." Another example is using behavioral data such as audio playback duration, number of replays, and number of skips to infer user satisfaction with the played audio.
[0114] In this embodiment, user information can be personalized information related to user identity, such as user ID, username, age, gender, and demographic tags. Operator identification information can be information used to uniquely identify a timbre operator (such as numbers, strings, or codes). By establishing a correspondence between operator identification information, target timbre, and user information, it is possible to record the timbre operators that a user or a group of users prefers to use when hearing a target timbre. This allows for the customization of unique timbre processing schemes for the user. In other words, timbre operators can be adjusted to match the user's usage habits based on their historical preferences, providing more accurate timbre recommendations in the future and improving user satisfaction and experience.
[0115] Furthermore, more user preference information can provide a data foundation for the general optimization of timbre operators, enabling continuous improvement and optimization of the design and performance of timbre operators.
[0116] This disclosure provides an audio processing method with user feedback and personalized adjustments. By introducing user feedback and establishing detailed correspondences, it is possible not only to better meet the personalized needs of users and provide specific users with an auditory experience that better meets their expectations, but also to continuously improve and optimize the design and performance of timbre operators, thereby providing a better auditory experience for the general public.
[0117] Figure 5 This is a flowchart illustrating an audio processing method according to some embodiments of the present disclosure.
[0118] In this embodiment of the disclosure, Figure 5 In the audio processing method shown, steps S510, S520, and S550 are respectively related to... Figure 1 Steps S110, S120, and S140 in the audio processing method shown correspond to each other and will not be repeated here.
[0119] In some embodiments of this disclosure, in Figure 1 Based on the audio processing method shown, Figure 5 The audio processing method shown may also include the following steps.
[0120] Step S530: In response to the existence of a target timbre in the timbre type, obtain the current user information using the audio to be played.
[0121] In this embodiment, more personalized and accurate matching can be achieved by combining current user information, thereby improving the intelligence level of audio processing and better adapting to the needs and preferences of different users. Current user information may include, for example, user ID, username, age, gender, and user category tags.
[0122] Step S540: In response to the existence of operator identifier information that matches the target timbre and / or the current user information, determine the timbre operator corresponding to the target timbre based on the operator identifier information.
[0123] The database can be searched to determine if operator identifiers match the target timbre and / or the current user information. Matching can be based on user preference selections previously provided by the user, or user preference selections provided by other users of the same type.
[0124] If a matching operator identifier is found, the timbre operator corresponding to the target timbre can be determined based on this identifier. This embodiment ensures that the selection of the timbre operator is based on the user's actual needs and preferences, thereby improving the personalization and satisfaction of the user with the audio processing results.
[0125] In addition, if no matching operator identifier information is found, the default timbre operator corresponding to the target timbre can be used for subsequent processing.
[0126] In some embodiments of this disclosure, converting the time-domain data of the audio to be played into frequency-domain data includes: determining the business scenario of the audio to be played; and converting the time-domain data of the audio to be played into frequency-domain data in response to the business scenario meeting processing conditions; wherein the processing conditions include: the maximum allowed delay duration of the business scenario is greater than or equal to a delay threshold, or the business scenario belongs to a preset scenario set.
[0127] In this embodiment of the disclosure, the business scenario can be a specific environment or purpose for audio usage, such as online music playback, real-time voice calls, video conferencing, game sound effects, movie playback, etc. Different business scenarios have different requirements and limitations for audio processing.
[0128] In an exemplary embodiment, the audio processing scheme provided in this disclosure, after determining that audio from an audio source or region can be obtained, can perform audio processing once after acquiring audio of a preset duration. That is, this scheme can output the final audio to be played in granular terms of the preset duration. Based on this, the delay threshold can be set to the preset duration. The preset duration can be, for example, 800 milliseconds, 1 second, 2 seconds, etc.
[0129] Some business scenarios (such as online music playback, non-real-time communication, etc.) have a high tolerance for latency. For example, the maximum allowed latency can reach 3 seconds or 5 seconds. When the maximum allowed latency of a business scenario exceeds the preset latency threshold, it is feasible to convert the audio from the time domain to the frequency domain and then perform timbre detection and timbre enhancement without causing a latency that is perceptible to the user.
[0130] In an exemplary embodiment, if the business scenario itself is a member of a preset scenario set, then a series of steps in this solution, such as time-domain to frequency-domain conversion, timbre detection, and timbre enhancement, can also be triggered. This preset scenario set can include scenarios that have been verified to be suitable for this solution and can significantly improve sound quality or meet specific needs.
[0131] In an exemplary embodiment, the scene set may include movie playback scene, voice call scene, and video conferencing scene.
[0132] In an exemplary embodiment, the audio processing method may further include: obtaining user preference selection information after the audio to be output is played; the user preference selection information indicates the operator identification information of the timbre operator; and establishing a correspondence between the operator identification information, the target timbre, the user information, and the scenario information of the business scenario.
[0133] Based on this, obtaining the timbre operator corresponding to the target timbre may further include: obtaining current user information using the audio to be played, and scene information of the business scenario; in response to the existence of operator identifier information matching the target timbre and / or the current user information and scene information, determining the timbre operator corresponding to the target timbre based on the operator identifier information.
[0134] Figure 6 This is a flowchart illustrating another audio processing method according to some embodiments of this disclosure. For example... Figure 6 As shown, the audio processing method may include the following steps.
[0135] Step S601: Obtain the audio to be played.
[0136] Step S603: Initialize the audio to be played.
[0137] For example, you can decode an MP4 format audio file to obtain a PCM format audio file.
[0138] Step S605: Determine whether the business scenario of the audio to be played meets the processing conditions. If yes, proceed to step S607; otherwise, proceed to step S619.
[0139] For example, this step can determine whether the business scenario for the audio to be played is a movie playback scenario.
[0140] Step S607: Convert the time-domain data of the audio to be played into frequency-domain data.
[0141] Step S609: Process the frequency domain data through the timbre recognition model to obtain the timbre type contained in the audio to be played.
[0142] Step S611: Determine if the target timbre exists in the timbre types. If it exists, proceed to step S613; otherwise, proceed to step S619.
[0143] Step S613: Perform signal enhancement on the frequency domain data according to the timbre operator to target the timbre, and obtain enhanced frequency domain data.
[0144] Step S615: Smooth the enhanced frequency domain data to obtain smoothed enhanced data.
[0145] Step S617: Use Fourier transform to convert the smoothed enhanced data into enhanced time-domain data, and determine the audio to be output based on the enhanced time-domain data.
[0146] Step S619: Perform normal sound effects processing to obtain the audio to be output.
[0147] Step S621: Play the audio to be output.
[0148] Figure 6 Other aspects of the embodiments can be referred to in the other embodiments described above, and will not be repeated here.
[0149] It should be noted that the above figures are merely illustrative representations of the processes included in methods according to some embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0150] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein. For details not disclosed in the apparatus embodiments of this disclosure, please refer to the embodiments of the method disclosed herein.
[0151] Figure 7 This is a block diagram illustrating an audio processing apparatus according to some embodiments of the present disclosure. (Refer to...) Figure 7 The device includes: an audio conversion unit 701, a detection unit 702, an operator acquisition unit 703, an output unit 704, a preference acquisition unit 705, and a preference establishment unit 706.
[0152] The audio conversion unit 701 is used to acquire the audio to be played and convert the time-domain data of the audio to be played into frequency-domain data; the detection unit 702 is used to detect the timbre type contained in the audio to be played based on the frequency-domain data; the operator acquisition unit is used to acquire the timbre operator corresponding to the target timbre in response to the existence of a target timbre in the timbre type; and the output unit is used to perform signal enhancement on the frequency-domain data for the target timbre according to the timbre operator to obtain the audio to be output.
[0153] In some embodiments of this disclosure, the detection unit 702 detects the timbre type contained in the audio to be played based on the frequency domain data, including: acquiring a trained timbre recognition model; processing the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played; and determining the timbre type contained in the audio to be played based on the timbre recognition result.
[0154] In some embodiments of this disclosure, the timbre recognition model includes a first fully connected layer, multiple convolutional layers, a gated recurrent layer, and a second fully connected layer. The detection unit 702 processes the frequency domain data through the timbre recognition model to obtain the timbre recognition result of the audio to be played, including: processing the frequency domain data through the first fully connected layer to obtain a first frequency domain feature; processing the first frequency domain feature through the multiple convolutional layers to obtain a second frequency domain feature; processing the second frequency domain feature through multiple gated recurrent units in the gated recurrent layer to obtain multiple time-series features; and processing the multiple time-series features through the second fully connected layer to obtain the timbre recognition result.
[0155] In some embodiments of this disclosure, the output unit 704 performs signal enhancement on the frequency domain data according to the timbre operator to target the timbre, thereby obtaining the audio to be output. This includes: performing convolution processing on the timbre operator and the frequency domain data to obtain enhanced frequency domain data; and determining the audio to be output based on the enhanced frequency domain data.
[0156] In some embodiments of this disclosure, the output unit 704 determines the audio to be output based on the enhanced frequency domain data, including: smoothing the enhanced frequency domain data to obtain smoothed enhanced data; and converting the smoothed enhanced data into time domain data of the audio to be output.
[0157] In some embodiments of this disclosure, the preference acquisition unit 705 is used to acquire user preference selection information after the audio to be output is played; the user preference selection information indicates the operator identification information of the timbre operator; the preference establishment unit 706 is used to establish the correspondence between the operator identification information, the target timbre, and the user information.
[0158] In some embodiments of this disclosure, the operator acquisition unit 703 acquires a timbre operator corresponding to the target timbre, including: acquiring current user information using the audio to be played; and in response to the existence of operator identifier information matching the target timbre and / or the current user information, determining a timbre operator corresponding to the target timbre based on the operator identifier information.
[0159] In some embodiments of this disclosure, the audio conversion unit 701 converts the time-domain data of the audio to be played into frequency-domain data, including: determining the business scenario of the audio to be played; and converting the time-domain data of the audio to be played into frequency-domain data in response to the business scenario meeting processing conditions; wherein the processing conditions include: the maximum allowed delay duration of the business scenario is greater than or equal to a delay threshold, or the business scenario belongs to a preset scenario set.
[0160] In some embodiments of this disclosure, the target timbre includes at least one of the following: human voice, light music sound, mechanical sound, and ambient sound.
[0161] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0162] Figure 8 This is a block diagram illustrating an audio processing apparatus 800 according to some embodiments of the present disclosure. For example, apparatus 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0163] Reference Figure 8 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 88, a sensor component 814, and a communication component 816.
[0164] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0165] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0166] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 800.
[0167] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0168] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0169] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0170] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0171] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 816 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0172] In some embodiments of this disclosure, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0173] In some embodiments of this disclosure, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions that can be executed by a processor 820 of device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0174] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform an audio processing method, the method comprising: acquiring audio to be played; converting time-domain data of the audio to be played into frequency-domain data; detecting a timbre type contained in the audio to be played based on the frequency-domain data; in response to the presence of a target timbre in the timbre type, acquiring a timbre operator corresponding to the target timbre; and performing signal enhancement on the frequency-domain data for the target timbre according to the timbre operator to obtain audio to be output.
[0175] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0176] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An audio processing method, characterized by, The method comprises: obtaining to-be-played audio, and converting time domain data of the to-be-played audio into frequency domain data; detecting a tone type contained in the to-be-played audio based on the frequency domain data; in response to the presence of a target tone in the tone type, obtaining a tone operator corresponding to the target tone; performing signal enhancement on the frequency domain data according to the tone operator for the target tone to obtain to-be-output audio.
2. The method of claim 1, wherein, The method comprises: obtaining a trained tone recognition model; processing the frequency domain data through the tone recognition model to obtain a tone recognition result of the to-be-played audio; determining the tone type contained in the to-be-played audio according to the tone recognition result.
3. The method of claim 2, wherein, The tone recognition model comprises a first full connection layer, a plurality of convolution layers, a gated recurrent layer, and a second full connection layer. The method comprises: processing the frequency domain data through the first full connection layer to obtain first frequency domain features; processing the first frequency domain features through the plurality of convolution layers to obtain second frequency domain features; processing the second frequency domain features through a plurality of gated recurrent units in the gated recurrent layer to obtain a plurality of time sequence features; processing the plurality of time sequence features through the second full connection layer to obtain the tone recognition result.
4. The method of claim 1, wherein, The method comprises: performing convolution processing on the tone operator and the frequency domain data to obtain enhanced frequency domain data; determining the to-be-output audio according to the enhanced frequency domain data.
5. The method of claim 4, wherein, The method comprises: performing smoothing processing on the enhanced frequency domain data to obtain smoothed enhanced data; converting the smoothed enhanced data into time domain data of the to-be-output audio.
6. The method of claim 1, wherein, The method further comprises: obtaining user preference selection information after the to-be-output audio is played; the user preference selection information indicates operator identification information of the tone operator; establishing a corresponding relationship among the operator identification information, the target tone, and user information.
7. The method according to claim 1 or 6, characterized in that, The method comprises: obtaining current user information using the to-be-played audio; in response to the presence of operator identification information matching the target tone and / or the current user information, determining a tone operator corresponding to the target tone according to the operator identification information.
8. The method of claim 1, wherein, The method comprises: determining a service scenario of to-be-played audio; in response to the service scenario satisfying a processing condition, converting time domain data of the to-be-played audio into frequency domain data; wherein the processing condition comprises that a maximum delay duration allowed by the service scenario is greater than or equal to a delay threshold, or the service scenario belongs to a preset scenario set.
9. The method of claim 1, wherein, The target tone comprises at least one of the following: human voice, light music sound, mechanical sound, and environmental sound.
10. An audio processing device, characterized by The method comprises: An audio converting unit is configured to obtain audio to be played and convert time domain data of the audio to be played into frequency domain data; A detecting unit is configured to detect a tone type contained in the audio to be played based on the frequency domain data; An operator obtaining unit is configured to obtain a tone operator corresponding to a target tone in the tone type in response to the target tone existing in the tone type; An output unit is configured to perform signal enhancement on the frequency domain data according to the tone operator for the target tone to obtain audio to be output.
11. An electronic device, comprising: It comprises: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the method of any one of claims 1-9.
12. A non-transitory computer readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform an audio processing method, the method comprising: obtaining audio to be played and converting time domain data of the audio to be played into frequency domain data; detecting a tone type contained in the audio to be played based on the frequency domain data; obtaining a tone operator corresponding to a target tone in the tone type in response to the target tone existing in the tone type; performing signal enhancement on the frequency domain data according to the tone operator for the target tone to obtain audio to be output.