Audio processing method, device, apparatus, and computer-readable storage medium
Through audio feature classification and directional noise reduction processing, the problem of poor audio processing effect in different recording scenarios is solved, and higher quality audio processing and user experience is achieved.
Patent Information
- Application Number
- CN202210155930.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-02-21
AI Technical Summary
In the prior art, the same noise reduction processing method has limited audio processing effects in different recording scenarios, which cannot meet the needs of different users and environments, resulting in a degradation of audio processing quality.
By classifying audio characteristics of the to be processed, multiple sound source types are identified, and directional noise reduction is performed according to the target sound source type selected by the user, including directional preservation or cancellation of specific noise.
It improves the audio processing effect, retains useful environmental background sound, enhances the user experience, and adapts to the noise reduction needs of different users and environments.
Smart Images

Figure CN114520005B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio processing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] Current terminals can record voice anytime and anywhere, obtain audio information, and improve user experience. However, there are various noise conditions in the recording environment. Therefore, after obtaining the audio information, it is necessary to perform noise reduction processing on the audio information to obtain the processed audio.
[0003] In the prior art, after obtaining audio information, a unified noise reduction process is used to reduce the noise of the audio information. However, the types of noise in different recording scenarios vary. In the prior art, using the same noise reduction process to reduce the noise of audio information in different recording scenarios can only achieve limited noise reduction effects, thereby reducing the audio processing quality. Summary of the Invention
[0004] Embodiments of the present application provide an audio processing method, apparatus, device, and computer-readable storage medium. By classifying and displaying multiple sound source types, noise reduction processing is performed on the processed audio according to the selected target sound source type, so that the noise reduction processing results are suitable for different users and different environments, thereby improving the audio processing effect.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] In a first aspect, an embodiment of the present application provides an audio processing method, the method comprising: obtaining audio to be processed; classifying audio features corresponding to the audio to be processed to obtain multiple sound source types, and displaying the multiple sound source types; for the multiple sound source types displayed, selecting a target sound source type in response to a sound source type selection operation; performing noise reduction processing on the audio to be processed according to the target sound source type to obtain target audio.
[0007] In a second aspect, an embodiment of the present application provides an audio processing device, comprising: an acquisition module for acquiring audio to be processed; a classification module for classifying audio features corresponding to the audio to be processed, obtaining a plurality of sound source types, and displaying the plurality of sound source types; a response module for selecting a target sound source type in response to a sound source type selection operation for the plurality of sound source types displayed; and a noise reduction module for performing noise reduction processing on the audio to be processed according to the target sound source type to obtain the target audio.
[0008] In a third aspect, an embodiment of the present application provides an audio processing device, which includes a memory for storing executable instructions and a processor for implementing the above-mentioned audio processing method when executing the executable instructions stored in the memory.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having executable instructions stored thereon for implementing the above-mentioned audio processing method when executed by a processor.
[0010] The embodiments of the present application provide an audio processing method, apparatus, device and computer-readable storage medium. According to the solution provided in the embodiments of the present application, the audio to be processed is obtained; the audio features corresponding to the audio to be processed are classified to obtain a variety of sound source types, and the multiple sound source types are displayed; by classifying the multiple sound source types, it is convenient for users to select a suitable noise reduction strategy according to their actual recording environment. For the multiple sound source types displayed, in response to the sound source type selection operation, the target sound source type is selected, and through the user interface interaction, the noise reduction strategy is made more suitable for different users and different environments. The audio to be processed is subjected to noise reduction processing according to the target sound source type to obtain the target audio, thereby improving the audio processing effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 An optional flowchart of an audio processing method provided in an embodiment of the present application;
[0012] Figure 2 An optional flowchart of another audio processing method provided in an embodiment of the present application;
[0013] Figure 3 An optional flowchart of another audio processing method provided in an embodiment of the present application;
[0014] Figure 4 An optional flowchart of another audio processing method provided in an embodiment of the present application;
[0015] Figure 5 An optional flowchart of another audio processing method provided in an embodiment of the present application;
[0016] Figure 6 An optional flowchart of another audio processing method provided in an embodiment of the present application;
[0017] Figure 7 A schematic diagram of the structure of an audio processing device provided in an embodiment of the present application;
[0018] Figure 8 A schematic diagram of the structure of an audio processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that some of the embodiments described here are only used to explain the technical solutions of the present application and are not used to limit the technical scope of the present application.
[0020] To facilitate understanding of this solution, before describing the embodiments of the present application, the application background (related technology) in the embodiments of the present application is described.
[0021] To facilitate understanding of this solution, before describing the embodiments of this application, the relevant technologies in the embodiments of this application are described.
[0022] Related art recording noise reduction technology considers all audio except voice messages as noise, classifying them into speech and non-speech scenarios. The same noise reduction process is then applied to all recorded audio in different scenarios. For example, certain neural network algorithms or other artificial intelligence detection algorithms are used to distinguish between speech and non-speech signals, and the amplitude of the non-speech signal is weakened to achieve the desired noise reduction effect.
[0023] The above-mentioned recording noise reduction solution can achieve the function of overall noise reduction of recorded audio. However, because the distinction dimension is considered from the perspective of whether it is speech, it is relatively simple, which reduces the audio processing effect. Different users' judgment of noise in different scenarios is not static. Whether it is noise requires users to judge based on the actual recording environment. For example, in a scene where users are talking, the sound of wind is noise in the recorded audio. However, in a scene where wind sound is collected outdoors, the wind sound is the audio to be collected. In this scenario, the wind sound can no longer be eliminated as noise.
[0024] The audio processing method provided in the embodiments of the present application can be applied to a terminal, which can be a vehicle-mounted device, a wearable device, a personal computer (PC), a smart phone, a tablet computer, a portable computer, or other device with a display function.
[0025] The audio processing method provided in the embodiment of the present application can be applied to a terminal. For example, the audio processing method is carried in an application (Application, APP), installed on the terminal, and the terminal obtains the audio to be processed; the audio features corresponding to the audio to be processed are classified to obtain a plurality of sound source types. The terminal has a display function, and the display function is used to display a plurality of sound source types and receive sound source type selection operations for the displayed plurality of sound source types. The terminal is also used to perform noise reduction processing on the audio to be processed according to the selected target sound source type to obtain the target audio.
[0026] The audio processing method provided in the embodiment of the present application can also be applied between two devices, and the communication connection between the two devices. The first device is used to obtain the audio to be processed, classify the audio features corresponding to the audio to be processed, obtain multiple sound source types, and transmit multiple sound source types to the second device. The second device is used to receive and display multiple sound source types, and the second device is also used to receive the sound source type selection operation performed on the displayed multiple sound source types, and transmit the selected target sound source type to the first device. The first device is also used to perform noise reduction processing on the audio to be processed according to the target sound source type to obtain the target audio.
[0027] The present application embodiment provides an audio processing method, such as Figure 1 As shown, Figure 1 This is an optional flowchart of an audio processing method provided in an embodiment of the present application, the audio processing method comprising the following steps:
[0028] S101: Obtain audio to be processed.
[0029] S102: Classify audio features corresponding to the audio to be processed to obtain multiple sound source types, and display the multiple sound source types.
[0030] In the embodiment of the present application, the audio to be processed includes a variety of sounds and noises, and the audio features are the features of the audio to be processed after quantization. The audio features can be understood as data features and can be used for classification, differentiation and other calculation processes.
[0031] Exemplarily, audio features of multiple preset sound source types are obtained, and the similarity between the audio features and the audio features of the multiple preset sound source types is calculated to obtain multiple similarities, and the sound source type with a greater than preset similarity among the multiple similarities is determined as multiple sound source types.
[0032] It should be noted that the preset similarity can be appropriately set by those skilled in the art according to actual circumstances, as long as the types of various sound sources can be distinguished, and this embodiment of the present application does not limit this.
[0033] In this embodiment of the present application, the aforementioned preset sound source types include, but are not limited to, whistles, birdsong, water, wind, music, rain, road noise, snoring, crying, baby sounds, equipment sounds, and renovation noise. Equipment sounds include, but are not limited to, audio echoes, interference sounds, and off-frequency sounds. This embodiment of the present application does not impose any limitations on this.
[0034] In some embodiments, the above S102 can be implemented in the following manner: Classifying audio features corresponding to the audio to be processed based on a preset classification model to obtain multiple sound source types, wherein the preset classification model is trained based on the audio features of multiple preset sound source types.
[0035] In an embodiment of the present application, audio of multiple preset sound source types is collected, and feature extraction is performed on the audio of the multiple preset sound source types to obtain audio features of the multiple preset sound source types. The audio features of the multiple preset sound source types are input into an initial classification model to obtain a predicted sound source type. A loss value is obtained based on the predicted sound source type and a preset loss function. The initial classification model is trained and optimized based on the loss value until a training termination condition is met, such as when the number of training cycles reaches a preset number or when the loss value reaches a preset threshold, thereby obtaining a preset classification model.
[0036] In the embodiment of the present application, the preset loss function can be set by those skilled in the art according to actual needs, including but not limited to the intersection-over-union ratio (DiceLoss), smooth L1 loss function and cross entropy loss function.
[0037] In the embodiment of the present application, the preset classification model can be understood as a machine learning model, which can be any appropriate neural network (NN) model that can be used to classify audio features, including but not limited to: Convolutional Recurrent Neural Network (CRNN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), etc. The above-mentioned NN model is a multi-classification model with multiple inputs and multiple outputs. The multi-classification model can be understood as a model that outputs more than two types of sound sources. The preset classification model can also be other machine learning methods, for example, binary classification algorithms, such as Logistic Regression, k-Nearest Neighbors, Decision Trees, Naive Bayes, and Support Vector Machines (SVM). The above-mentioned binary classification algorithm outputs two types of sound sources. The embodiment of the present application does not limit the specific structure of the classification model.
[0038] S103 : For the multiple audio source types displayed, in response to an audio source type selection operation, a target audio source type is selected.
[0039] In an embodiment of the present application, an example is given in which a plurality of sound source types are displayed to a user through a user interface (UI). The user performs a selection operation based on the displayed plurality of sound source types and selects a target sound source type. The target sound source type includes one or more sound source types, and this embodiment of the present application does not impose any limitation on this.
[0040] In the embodiments of the present application, in certain scenarios, the bird calls and water sounds in the background noise are the environmental details required by the user. The recording noise reduction technology in the related art cannot retain the environmental details required by the user, and all the non-speech signals obtained by distinction are eliminated, thereby reducing the audio processing effect.
[0041] In an embodiment of the present application, different types of sound sources are distinguished according to the type of sound source, so that users can eliminate directional noise according to their personal needs. While choosing to eliminate noise, useful recording environment background sounds and voices are retained, thereby improving the user experience.
[0042] S104: Perform noise reduction processing on the audio to be processed according to the target sound source type to obtain the target audio.
[0043] In an embodiment of the present application, an audio processing method is applied to a terminal as an example for explanation, and the terminal receives and responds to a selected target sound source type. The selected target sound source type can be a sound source type that needs to be eliminated, or it can be a sound source type that needs to be retained. In either case, after obtaining the selected target sound source type, a noise reduction strategy can be determined based on the target sound source type. The noise reduction strategy is to directional retain noise or directional filter noise, and directional noise reduction processing is performed on the audio to be processed according to the noise reduction strategy.
[0044] For example, the directional noise reduction effect can be achieved by superimposing reverse amplitude signals or performing blind source separation on a specific noise frequency, wherein blind source separation can also be understood as blind signal separation (BSS).
[0045] Exemplarily, when a noise reduction method of superimposing reverse amplitude signals is adopted, if the selected target sound source type is the sound source type that needs to be eliminated, the output power of the reverse sound wave spectrum is adapted according to the selected target sound source type, and the two opposite waveforms are offset against each other. The selected target sound source type overlaps with the reverse sound wave spectrum, thereby achieving a directional noise reduction effect of offsetting the specified sound source type.
[0046] Exemplarily, when using the blind source separation noise reduction method, if the selected target sound source type is the sound source type that needs to be eliminated, the selected target sound source type is used as the specific noise, and blind source separation is used to separate or restore the original source signal (i.e., the target audio) from the mixed signal, thereby achieving the effect of eliminating the specific noise.
[0047] In the embodiment of the present application, by selectively selecting the target sound source type that needs to be retained or eliminated and performing noise reduction processing on the target sound source type, the noise reduction result is more targeted, the audio noise reduction effect is improved, and at the same time, the environmental background sound other than the noise is retained, avoiding missing useful environmental background sound. The retained environmental background sound serves as a bridge connecting the recorded audio with the outside world, reducing the audio distortion caused by the lack of environmental background sound, and improving the audio processing quality.
[0048] According to the solution provided in the embodiment of the present application, the audio to be processed is obtained; the audio features corresponding to the audio to be processed are classified to obtain a variety of sound source types, and the multiple sound source types are displayed; by classifying the multiple sound source types, it is convenient for users to select a suitable noise reduction strategy according to their actual recording environment. For the multiple sound source types displayed, in response to the sound source type selection operation, the target sound source type is selected, and through UI interaction, the noise reduction strategy is made more suitable for different users and different environments. The audio to be processed is subjected to noise reduction processing according to the target sound source type to obtain the target audio, thereby improving the audio processing effect.
[0049] In some embodiments, the audio features corresponding to the audio to be processed can be obtained through S201-S203. Figure 2 As shown, Figure 2 An optional flowchart of another audio processing method provided in an embodiment of the present application.
[0050] S201 , performing multi-dimensional feature extraction on the audio to be processed to obtain multi-dimensional features, where the multi-dimensional features include at least one of time domain, frequency domain, spatial domain, and amplitude.
[0051] In the embodiment of the present application, the multi-dimensionality may also include the highest frequency, the lowest frequency, timbre, decibels, and the frequency variation range, etc., which are not limited by the embodiment of the present application. The time domain represents the change of sound over time. For example, the sound of a whistle is a very short pulse, and the sound of water flow is a continuous pulse. The amplitude represents the intensity of the sound. For example, the amplitude of the whistle is relatively concentrated and relatively high, while the amplitude of the sound of water flow is relatively low. The frequency domain represents the frequency band in which the sound is distributed. For example, the frequency band in which the whistle is distributed is relatively concentrated and usually distributed in the high-frequency band; the sound of water flow is distributed in the low-frequency band. The spatial domain represents the distribution of sound in space. The sound of a whistle and the sound of water flow can be understood as white noise. White noise is relatively uniformly distributed in space and has a relatively low autocorrelation. It can be understood that the amplitude of white noise at different times is unrelated. White noise is different from the human voice, which has a relatively high autocorrelation.
[0052] In the embodiment of the present application, by performing multi-dimensional feature extraction on the audio to be processed, the richness of the features is improved, so that the audio features can be obtained by subsequent fusion based on the multiple dimensional features, thereby improving the accuracy of the audio features.
[0053] S202: Normalize each dimensional feature separately to obtain multiple normalized dimensional features.
[0054] In the embodiment of the present application, since the magnitudes (which can also be understood as dimensions, ranges or units) corresponding to various dimensional features are different and some features are discrete, it is also necessary to normalize the multiple dimensional features to achieve quantitative unification of the multiple dimensional features, so as to facilitate the subsequent fusion processing of the multiple normalized dimensional features.
[0055] S203: fusing multiple normalized dimensional features to obtain audio features.
[0056] In the embodiments of the present application, dimensional features can be represented in the form of vectors. Multiple normalized dimensional features are fused to obtain audio features that contain multi-dimensional features of the processed audio, thereby improving the accuracy of the processed audio. Fusion methods include, but are not limited to, merging and concatenation.
[0057] In an embodiment of the present application, multi-dimensional feature extraction is performed on the audio to be processed, thereby improving the richness and diversity of the features, normalizing the extracted multiple dimensional features, and fusing the multiple normalized dimensional features to obtain audio features, thereby improving the accuracy of the audio features.
[0058] It should be noted that when obtaining audio features, the feature extraction of the audio to be processed can be performed through an algorithm or a neural network model to obtain audio features. The embodiments of this application do not limit the specific implementation of the specific algorithm and neural network model method.
[0059] In some embodiments, the above S201 can be implemented in the following manner: Based on a preset feature extraction model, multi-dimensional feature extraction is performed on the audio to be processed to obtain multi-dimensional features; the preset feature extraction model is used for feature extraction.
[0060] In an embodiment of the present application, a preset feature extraction model is used for feature extraction of different dimensions. The preset feature extraction model has the function of taking the audio to be processed as input and outputting corresponding multiple dimensional features. There is no restriction on the specific implementation form and processing process of the feature extraction model, as long as it can output multiple dimensional features based on the audio to be processed. In practical applications, such as neural networks (NN), etc., can be applied, and this embodiment of the present application does not limit this.
[0061] It should be noted that the preset feature extraction model can also perform the above steps S201-S203. In other words, feature extraction, normalization, and fusion are all performed by the feature extraction model. Here, the example of the preset feature extraction model performing the above step S201 is used for explanation. The embodiment of the present application does not limit the execution steps of the feature extraction model.
[0062] In some embodiments, the above-mentioned preset feature extraction model can be obtained by the following method. Including S301-S305. Figure 3 As shown, Figure 3 An optional flowchart of another audio processing method provided in an embodiment of the present application.
[0063] S301: Acquire an audio sample of a preset sound source type, where the preset sound source type includes at least one of a whistle, a bird's cry, a water flow, a wind, a music sound, and a device sound.
[0064] In the embodiment of the present application, the audio samples of the preset sound source types can be understood as audio samples of known sound source types. The preset sound source types can be supplemented and added in the subsequent application process. For example, the sound source types can also include rain sound, road noise, snoring sound, baby sound, crying sound, and renovation noise. This embodiment of the present application is not limited to this.
[0065] S302 : Based on the initial feature extraction model, perform multi-dimensional feature extraction on audio samples of various preset sound source types to obtain multi-dimensional feature samples of various preset sound source types.
[0066] S303: Determine audio feature samples of various preset sound source types according to the multiple dimensional feature samples of various preset sound source types.
[0067] In the embodiment of the present application, S302 and S303 are the process of extracting features from audio samples during the training process. Figure 2 The feature extraction steps in S201-S203 are the same and will not be repeated here.
[0068] S304: Calculate the discrimination between each pair of audio feature samples of various preset sound source types to obtain a plurality of feature discriminations.
[0069] S305: If the discrimination of multiple features is greater than a preset threshold, a preset feature extraction model is obtained.
[0070] In the embodiments of the present application, after obtaining audio feature samples of various preset sound source types, it is necessary to determine whether the audio feature samples extracted by the initial feature extraction model meet the requirements. In other words, whether the discrimination between each pair of audio feature samples meets the requirement of being greater than a preset threshold. If the requirements are met, it indicates that the parameters of the feature extraction model and the dimensions corresponding to the features can be used in subsequent feature extraction applications, which can also be understood as the completion of feature extraction model training.
[0071] It should be noted that the preset threshold can be set by those skilled in the art according to actual conditions, as long as the audio features can be distinguished. For example, it can be determined based on a large number of threshold analysis adopted in the training process, and this embodiment of the application does not limit this.
[0072] In an embodiment of the present application, the initial feature extraction model is trained using audio samples of preset sound source types until the discrimination between each pair of audio feature samples of various preset sound source types meets the requirements, thereby obtaining a feature extraction model. This feature extraction model can be used for subsequent feature extraction, and multi-dimensional features are extracted and fused to obtain audio features, thereby improving the accuracy of the audio features.
[0073] In some embodiments, in the above Figure 3 After S305, the audio processing method further includes S306 and S307.
[0074] S306: If there is a feature discrimination degree among the multiple feature discrimination degrees that is less than or equal to a preset threshold, parameters of the initial feature extraction model are adjusted to obtain a feature extraction model after parameter adjustment.
[0075] S307. Based on the feature extraction model after parameter adjustment, continuously extract features of target dimensions for audio samples of various preset sound source types until multiple feature discriminations are greater than a preset threshold, thereby obtaining a preset feature extraction model; wherein the target dimension is at least one of the multiple dimensions.
[0076] In the embodiment of the present application, if the discrimination between audio feature samples does not meet the requirements, it means that the dimension of the currently extracted features is inappropriate, or the parameters of the feature extraction model are inappropriate. It is necessary to reselect the target dimension. The target dimension can be a single dimension or multiple dimensions. The target dimension can be one of the multiple dimensions or a newly added dimension. This embodiment of the present application does not limit this.
[0077] In an embodiment of the present application, the parameters of the initial feature extraction model are adjusted, and the dimensions of the extracted features are screened, and the initial feature extraction model is continuously trained until the requirements are met to obtain a feature extraction model, thereby improving the accuracy of the feature extraction model, thereby improving the discrimination between the audio features of different sound source types extracted by it, facilitating the discrimination of audio features of different sound source types, and improving the accuracy of noise classification results.
[0078] Below, we will explain the exemplary application of the embodiment of the present application in a practical application scenario. Before the processed audio is subjected to noise reduction processing, the noise classification model is used as an example to represent the classification model. Figure 4 As shown, Figure 4 An optional flowchart of another audio processing method provided in an embodiment of the present application includes S401-S405.
[0079] S401: Noise refinement.
[0080] The recorded ambient noise is divided into various preset sound source types according to the sound source type. The preset sound source types include but are not limited to whistles, birdsong, water flow, wind, music, and equipment sounds.
[0081] S402: Optimal feature extraction.
[0082] For each preset sound source type, features of different preset sound source types are extracted in the time domain, frequency domain, spatial domain and amplitude to obtain multi-dimensional features.
[0083] It should be noted that since the audio features determined by multiple dimensional features need to be evaluated to see if they meet the conditions for strong discrimination, if they do, the audio features can be output as audio feature samples; if they do not, the dimensional features need to be re-extracted. Therefore, dimensional features are used here to represent dimensional feature samples, and audio features are used to represent audio feature samples.
[0084] S403: Feature quantification.
[0085] The multi-dimensional features are normalized and then fused to achieve quantization of the multi-dimensional features and obtain audio features.
[0086] S404. Feature evaluation.
[0087] Evaluate the audio characteristics of various preset source types.
[0088] S405: If the feature distinction between the audio features of the plurality of preset sound source types is strong, output the noise classification database.
[0089] In an embodiment of the present application, the noise classification database includes audio features of multiple preset sound source types and is used to train a preset noise classification model.
[0090] If the feature discrimination is weak, S402-S404 are executed again to select the best feature until the feature discrimination between the audio features of the multiple preset sound source types is strong, and the noise classification database is output.
[0091] In the embodiments of the present application, the recorded audio is classified into preset sound source types and features are extracted to obtain audio feature samples. After establishing a noise classification database, it is necessary to select an appropriate noise classification model to learn the audio feature samples so that the noise classification model can classify the audio feature samples and determine the sound source type. The noise classification model can be a multi-output NN model or other machine learning method, such as a binary classification model SVM.
[0092] In an embodiment of the present application, the recorded environmental noise is refined by sound source classification, and a complete noise classification database is established, providing users with a personalized and selectable noise reduction solution so that users can choose different sound source types for elimination according to different scenarios and conditions, thereby improving the sound quality and detail processing of the recorded audio.
[0093] When performing noise reduction on the audio to be processed, Figure 5 As shown, Figure 5 An optional flowchart of another audio processing method provided in an embodiment of the present application includes S501-S506.
[0094] S501, recording audio source.
[0095] The recording source indicates the audio to be processed
[0096] S502: Feature extraction.
[0097] Perform feature extraction on the recorded sound source to obtain audio features.
[0098] S503: Noise classification.
[0099] Classify the audio features and obtain the noise classification results.
[0100] S504: Present the noise classification result.
[0101] The audio features of the audio to be processed are input into the noise classification model, which then outputs a noise classification result, which includes multiple sound source types. The noise classification model converts the audio features into multiple sound source types and displays them to the user through the UI, allowing the user to select the sound source type to be eliminated based on the recording environment.
[0102] S505: User selectively eliminates noise.
[0103] S506: Noise elimination.
[0104] According to the target sound source type corresponding to the user's selective noise elimination, noise elimination is performed on the processed audio to obtain the target audio.
[0105] In the embodiment of the present application, after the user selects directional noise elimination for a certain source type, directional noise reduction processing is performed on the audio to be processed based on the selected source type, for example, superimposing an inverse amplitude signal or performing blind source separation on a specific noise frequency, thereby achieving the purpose of specific noise reduction and improving the noise reduction effect.
[0106] In the embodiment of the present application, by selecting a technical solution for eliminating background noise, users can choose the type of sound source they want to retain or eliminate, thereby improving user experience and audio processing effects.
[0107] In some embodiments, the present application provides an audio processing method, including S601-S605. Figure 6 As shown, Figure 6 An optional flowchart of another audio processing method provided in an embodiment of the present application.
[0108] S601: Obtain audio to be processed.
[0109] S602: Classify audio features corresponding to the audio to be processed to obtain multiple audio source types.
[0110] It should be noted that S601 and S602 are Figure 1 The specific implementations of S101 and S102 are the same and will not be repeated here.
[0111] S603: Perform noise reduction prediction for multiple sound source types based on a preset recommendation model, generate at least one noise reduction solution, and display at least one noise reduction solution; wherein the preset recommendation model is used to predict the user's noise reduction preference.
[0112] In the embodiment of the present application, the preset recommendation model can also be understood as a machine learning model, which can be any appropriate neural network (NN) model that can be used to predict user preference information. The preset recommendation model predicts the user's noise reduction preference based on the target sound source type selected historically, and the user's noise reduction preference can also be understood as the user's noise reduction habits. Based on the preset recommendation model, noise reduction predictions can be made based on a variety of sound source types and user noise reduction preferences to generate a noise reduction solution. The noise reduction solution can be one or more, and the noise reduction solution can represent the elimination of one or more sound source types, or it can represent the retention and elimination of one or more sound source types, and this embodiment of the present application does not limit this.
[0113] For example, in one scenario, by recommending a noise reduction solution to the user, the user's operation steps can be reduced. When the user wants to eliminate multiple sound source types, the user does not need to select multiple sound source types, but only needs to select the corresponding noise reduction solution, thereby improving the noise reduction efficiency and user experience. In another scenario, the user has no or little knowledge of the field of audio processing, and when faced with multiple sound source types, it is impossible to select the most preferred noise reduction solution. The at least one recommended noise reduction solution can serve as a prompt. The user can select the noise reduction solution with a high probability of being selected in the past from the at least one recommended noise reduction solution, thereby ensuring the quality of audio processing.
[0114] In some embodiments, the preset recommendation model in S603 is implemented as follows: obtaining user historical behavior information, wherein the user historical behavior information includes historically selected target audio source types and / or historically selected solutions corresponding to historical noise reduction solutions; and training an initial recommendation model based on the user historical behavior information to obtain a preset recommendation model.
[0115] In an embodiment of the present application, in a scenario where multiple sound source types are presented to the user, the selected historical target sound source type is selected by the user from among the multiple sound source types, reflecting the user's noise reduction habits. Taking the recommendation of multiple noise reduction solutions as an example, in a scenario where multiple noise reduction solutions are recommended to the user, the user can select a specific noise reduction solution from among the multiple noise reduction solutions, or abandon the multiple noise reduction solutions and reselect the target sound source type on their own. The historical selection solution represents the user's selection results for the multiple noise reduction solutions. The historical selection solution includes the target noise reduction solution and the target sound source type, and the historical selection solution reflects the user's noise reduction habits.
[0116] In this embodiment of the present application, the recommendation model can be trained based on the user's historical behavior information corresponding to either of the two scenarios above, or a combination of the user's historical behavior information corresponding to the two scenarios above. The preset recommendation model is used to predict the user's noise reduction preference based on the user's historical behavior information and output at least one noise reduction solution.
[0117] It should be noted that when training the recommendation model, the user historical behavior information can be the user historical behavior information of multiple different users, or the user historical behavior information of a specific user. For a certain user, when initially processing the audio, there is no user historical behavior information of the current user, so it is impossible to predict the noise reduction preference information of the current user. At this time, the noise reduction solution can be recommended based on the universal recommendation model trained with the user historical behavior information of other users. As the number of recommendations increases, the user historical behavior information of the current user is obtained, and it can be trained based on the user historical behavior information of the current user to obtain a recommendation model that is suitable for the current user, thereby recommending a personalized noise reduction solution for the current user.
[0118] S604: In response to the selection operation of at least one noise reduction scheme, determine a target noise reduction scheme.
[0119] In the embodiments of this application, a noise reduction solution is presented to the user through a UI as an example, and the user selects the noise reduction solution presented. For example, if multiple noise reduction solutions are presented to the user, the user can select the target noise reduction solution by confirming it with one click based on the recording environment or their own needs, thereby improving the user experience.
[0120] In the embodiment of the present application, the target noise reduction scheme represents a selected noise reduction scheme, and the target noise reduction scheme includes one or more sound source types.
[0121] S605: Perform noise reduction processing on the audio to be processed according to the target noise reduction scheme to obtain the target audio.
[0122] In the embodiment of the present application, the application of the audio processing method to the terminal is used as an example for explanation, and the terminal receives and responds to the selected target noise reduction scheme. The target noise reduction scheme can be one or more sound source types that need to be eliminated, or one or more sound source types that need to be retained. In either case, after obtaining the target noise reduction scheme, a noise reduction strategy can be determined according to the target noise reduction scheme, and the processed audio can be subjected to directional noise reduction processing according to the noise reduction strategy. The noise reduction strategy can include superimposing reverse amplitude signals and blind source separation, which can be referred to in Figure 1 The description of S104 is omitted here.
[0123] In an embodiment of the present application, noise reduction predictions are made for various audio source types based on a preset recommendation model, generating at least one noise reduction solution. This solution is then recommended through UI interaction, allowing users to select an appropriate target noise reduction solution based on their actual recording environment with a single click, thereby improving noise reduction efficiency. The selected noise reduction solution is more suitable for different users and environments. Then, noise reduction processing is performed on the audio being processed based on the target noise reduction solution to obtain the target audio, improving the audio processing effect.
[0124] In some embodiments, after the above S605, the audio processing method further includes: for the same sound source type, determining the recommendation deviation based on the historical selection scheme and the target noise reduction scheme; adjusting the preset recommendation model based on the recommendation deviation to obtain the adjusted recommendation model, and the adjusted recommendation model is used for the next noise reduction scheme recommendation process.
[0125] In this embodiment of the present application, the preset recommendation model is trained based on historical user behavior information and is used to predict user noise reduction preferences. For a particular user, during initial audio processing, no historical user behavior information is available, and the user noise reduction preferences have not yet been adjusted. Therefore, when recommending at least one noise reduction solution to the current user, the recommendation is based on the default, universal user noise reduction preferences. In other words, it is not possible to recommend a personalized noise reduction solution for the current user.
[0126] In an embodiment of the present application, as the number of recommendations increases, multiple sound source types or at least one noise reduction scheme are continuously recommended to the current user, and the target sound source type and / or target noise reduction scheme selected by the current user are received as the user history behavior information of the current user. Adjusting the preset recommendation model based on the user history behavior information of the current user can also be understood as adjusting the user preference information of the current user. Moreover, as the number of recommendations increases, the current user's user noise reduction preferences in different environments are different, that is, as the recording environment changes, the user's noise reduction habits in different recording environments are different. For example, in a user conversation scenario, the user will choose to filter out the wind as noise, and in a scenario where the wind is collected outdoors, the user will choose to keep the wind as background sound.
[0127] In an embodiment of the present application, for a certain user, after obtaining the target noise reduction solution, it is also necessary to determine the recommendation deviation based on the target noise reduction solution and the final selected solution, so as to continuously iteratively update the preset recommendation model based on the recommendation deviation, so that the recommendation model is more suitable for the current user, thereby recommending a personalized noise reduction solution.
[0128] For example, the noise reduction solution generated by the preset recommendation model is not static. As the recording environment changes and the user's actual preferred noise reduction habits are different, the recommendation model needs to be continuously learned and adjusted. In other words, there will be a deviation between the noise reduction solution output by the recommendation model and the solution actually selected by the user, which will reduce the output accuracy of the recommendation model. Therefore, it is also necessary to input the deviation information into the recommendation model and automatically update the recommendation model to adapt to the current user's noise reduction habits. It is conceivable that the noise reduction solution generated for the audio recorded by the user at a concert is completely different from the noise reduction solution generated for the audio recorded by the user during an outdoor outing.
[0129] In the embodiment of the present application, as the number of sound source types gradually increases and the machine learning model is continuously iterated and updated, the classification of noise types will gradually improve, and the distinction between audio features will become more refined. Due to the continuous updating of feature extraction models, classification models, and recommendation models, a mapping bridge is formed to convert the user's recorded audio into audio containing multi-dimensional information as input and output a noise reduction solution suitable for the user. The technical solution for selective noise elimination provided in the embodiment of the present application can provide a suitable noise reduction solution based on the user's personal preferences and the recording environment in which the user is located, thereby improving the noise reduction efficiency.
[0130] Based on the audio processing method of the embodiment of the present application, the embodiment of the present application also provides an audio processing device, such as Figure 7 As shown, Figure 7 A structural schematic diagram of an audio processing device provided in an embodiment of the present application, the audio processing device 70 includes: an acquisition module 701, used to obtain audio to be processed; a classification module 702, used to classify audio features corresponding to the audio to be processed, obtain multiple sound source types, and display the multiple sound source types; a response module 703, used to select a target sound source type in response to a sound source type selection operation for the multiple sound source types displayed; a noise reduction module 704, used to perform noise reduction processing on the audio to be processed according to the target sound source type to obtain the target audio.
[0131] In some embodiments, the audio processing device 70 further includes a feature extraction module;
[0132] A feature extraction module is used to perform multi-dimensional feature extraction on the audio to be processed to obtain multiple dimensional features, where the multiple dimensions include at least one of the time domain, frequency domain, spatial domain and amplitude; normalize each dimensional feature separately to obtain multiple normalized dimensional features; and fuse the multiple normalized dimensional features to obtain the audio feature.
[0133] In some embodiments, the feature extraction module is further used to perform the multi-dimensional feature extraction on the audio to be processed according to a preset feature extraction model to obtain the multi-dimensional features; the preset feature extraction model is used for feature extraction.
[0134] In some embodiments, the audio processing device 70 further includes a training module;
[0135] The acquisition module 701 is further configured to acquire an audio sample of a preset sound source type, wherein the preset sound source type includes at least one of a whistle, a bird's cry, a water flow, a wind, a music sound, and a device sound;
[0136] The feature extraction module is further configured to perform the multi-dimensional feature extraction on the audio samples of each of the preset sound source types based on the initial feature extraction model to obtain multi-dimensional feature samples of each of the preset sound source types;
[0137] The training module is used to determine the audio feature samples of various preset sound source types based on multiple dimensional feature samples of various preset sound source types; calculate the discrimination between each pair of audio feature samples of various preset sound source types to obtain multiple feature discriminations; if the multiple feature discriminations are all greater than a preset threshold, the preset feature extraction model is obtained.
[0138] In some embodiments, the training module is also used to adjust the parameters of the initial feature extraction model if there is a feature discrimination less than or equal to the preset threshold value among the multiple feature discriminations, so as to obtain a feature extraction model after parameter adjustment; according to the feature extraction model after parameter adjustment, continuously perform feature extraction of target dimensions on audio samples of various preset sound source types, until the multiple feature discriminations are all greater than the preset threshold value, thereby obtaining the preset feature extraction model; wherein, the target dimension is at least one item among the multiple dimensions.
[0139] In some embodiments, the classification module 702 is further used to classify the audio features corresponding to the audio to be processed based on a preset classification model to obtain the multiple sound source types, wherein the preset classification model is trained based on the audio features of multiple preset sound source types.
[0140] In some embodiments, the audio processing device 70 further includes a generation module;
[0141] a generation module, configured to perform noise reduction prediction for the multiple sound source types based on a preset recommendation model, generate at least one noise reduction scheme, and display the at least one noise reduction scheme; wherein the preset recommendation model is used to predict the user's noise reduction preference;
[0142] The response module 703 is further configured to determine a target noise reduction scheme in response to the selection operation of the at least one noise reduction scheme;
[0143] The noise reduction module 704 is further configured to perform noise reduction processing on the audio to be processed according to the target noise reduction scheme to obtain the target audio.
[0144] In some embodiments, the acquisition module 701 is configured to acquire user historical behavior information, wherein the user historical behavior information includes selected historical target sound source types and / or historical selection schemes corresponding to historical noise reduction schemes;
[0145] The training module is used to train the initial recommendation model according to the user's historical behavior information to obtain the preset recommendation model.
[0146] In some embodiments, the training module is also used to determine the recommendation deviation for the same sound source type based on the historical selection scheme and the target noise reduction scheme; based on the recommendation deviation, the preset recommendation model is adjusted to obtain an adjusted recommendation model, and the adjusted recommendation model is used in the process of recommending the noise reduction scheme next time.
[0147] It should be noted that the audio processing device provided in the above embodiment only uses the division of the above program modules as an example to illustrate when performing audio processing. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the audio processing device provided in the above embodiment and the audio processing method embodiment belong to the same concept. The specific implementation process and beneficial effects are detailed in the method embodiment and will not be repeated here. For technical details not disclosed in the embodiment of this device, please refer to the description of the method embodiment of this application for understanding.
[0148] In the embodiments of this application, Figure 8 This is a schematic diagram of the structure of the audio processing device proposed in the embodiment of the present application, as shown in FIG. Figure 8 As shown, the device 80 proposed in the embodiment of the present application may also include a processor 801 and a memory 802 storing executable instructions of the processor 801. In some embodiments, the audio processing device 80 may also include a communication interface 803 and a bus 804 for connecting the processor 801, the memory 802 and the communication interface 803.
[0149] In the embodiment of the present application, the processor 801 may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor function may also be other, and the embodiment of the present application does not specifically limit this.
[0150] In the embodiment of the present application, the bus 804 is used to connect the communication interface 803, the processor 801 and the memory 802, as well as the mutual communication between these devices.
[0151] In an embodiment of the present application, the processor 801 is configured to execute the audio processing method described in any one of the above embodiments.
[0152] The memory 802 in the audio processing device 80 can be connected to the processor 801. The memory 802 is used to store executable program code and data. The program code includes computer operating instructions. The memory 802 may include high-speed RAM memory, and may also include non-volatile memory, for example, at least two disk memories. In actual applications, the memory 802 can be a volatile memory (volatile memory), such as random-access memory (RAM); or a non-volatile memory (non-volatile memory), such as read-only memory (ROM), flash memory, hard disk drive (HDD) or solid-state drive (SSD); or a combination of the above types of memory, and provides instructions and data to the processor 801.
[0153] In addition, the functional modules in the embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional modules.
[0154] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0155] An embodiment of the present application provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the audio processing method as described in any of the above embodiments is implemented.
[0156] Exemplarily, the program instructions corresponding to an audio processing method in this embodiment can be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When the program instructions corresponding to an audio processing method in the storage medium are read or executed by an electronic device, the audio processing method described in any of the above embodiments can be implemented.
[0157] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0158] The present application is described with reference to the implementation flow charts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flow charts and / or block diagrams, as well as the combination of processes and / or boxes in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the implementation flow charts. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0159] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which is implemented in the implementation flow diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process described in the flowchart. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0161] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. An audio processing method, characterized in that: The method comprises: Get the audio to be processed; Classifying the audio features corresponding to the audio to be processed to obtain multiple sound source types; Performing noise reduction predictions on the multiple sound source types based on a preset recommendation model, generating at least one noise reduction scheme, and displaying the at least one noise reduction scheme; wherein the preset recommendation model is used to predict the user's noise reduction preference; In response to the selection operation of the at least one noise reduction scheme, determining a target noise reduction scheme; performing noise reduction processing on the audio to be processed according to the target noise reduction scheme to obtain target audio; In the absence of responding to the selection operation of the at least one noise reduction scheme, the multiple sound source types are displayed; for the multiple sound source types displayed, a target sound source type is selected in response to the sound source type selection operation; and according to the target sound source type, noise reduction processing is performed on the audio to be processed to obtain the target audio.
2. The method according to claim 1, characterized in that After obtaining the audio to be processed, the method further includes: Performing multi-dimensional feature extraction on the audio to be processed to obtain multi-dimensional features, where the multi-dimensional features include at least one of time domain, frequency domain, spatial domain, and amplitude; Normalize each dimensional feature separately to obtain multiple normalized dimensional features; The audio feature is obtained by fusing the multiple normalized dimensional features.
3. The method according to claim 2, characterized in that The multi-dimensional feature extraction is performed on the audio to be processed to obtain multiple dimensional features, including: According to a preset feature extraction model, the multi-dimensional feature extraction is performed on the audio to be processed to obtain the multi-dimensional features; the preset feature extraction model is used for feature extraction.
4. The method according to claim 3, characterized in that Before performing the multi-dimensional feature extraction on the audio to be processed according to the preset feature extraction model to obtain the multi-dimensional features, the method further includes: Obtaining an audio sample of a preset sound source type, where the preset sound source type includes at least one of a whistle, a bird call, a water flow, a wind, a music sound, and a device sound; Based on the initial feature extraction model, performing the multi-dimensional feature extraction on the audio samples of each of the preset sound source types to obtain multi-dimensional feature samples of each of the preset sound source types; Determining audio feature samples of various preset sound source types according to the multiple dimensional feature samples of various preset sound source types; Calculating the discrimination between each pair of audio feature samples of the preset sound source types to obtain a plurality of feature discriminations; If the multiple feature discriminations are all greater than a preset threshold, the preset feature extraction model is obtained.
5. The method according to claim 4, characterized in that After calculating the discrimination between each pair of audio feature samples of the various preset sound source types to obtain a plurality of feature discriminations, the method further includes: If there is a feature discrimination degree among the multiple feature discrimination degrees that is less than or equal to the preset threshold, adjusting the parameters of the initial feature extraction model to obtain a feature extraction model after parameter adjustment; According to the feature extraction model after parameter adjustment, feature extraction of target dimensions is continuously performed on audio samples of various preset sound source types until the multiple feature discriminations are all greater than a preset threshold, thereby obtaining the preset feature extraction model; wherein, the target dimension is at least one item of the multiple dimensions.
6. The method according to any one of claims 1 to 5, characterized in that The audio features corresponding to the audio to be processed are classified to obtain multiple sound source types, including: The audio features corresponding to the audio to be processed are classified based on a preset classification model to obtain the multiple sound source types, wherein the preset classification model is trained based on the audio features of the multiple preset sound source types.
7. The method according to claim 1, characterized in that Before performing noise reduction prediction on the multiple sound source types based on the preset recommendation model and generating at least one noise reduction solution, the method further includes: Acquiring user historical behavior information, wherein the user historical behavior information includes historical target sound source types selected and / or historical selection schemes corresponding to historical noise reduction schemes; The initial recommendation model is trained according to the user historical behavior information to obtain the preset recommendation model.
8. The method according to claim 1, characterized in that After performing noise reduction processing on the audio to be processed according to the target noise reduction scheme to obtain the target audio, the method further includes: For the same sound source type, determining a recommended deviation based on historical selection schemes and the target noise reduction scheme; The preset recommendation model is adjusted according to the recommendation deviation to obtain an adjusted recommendation model, and the adjusted recommendation model is used in the process of recommending a noise reduction solution next time.
9. An audio processing device, characterized in that: The device comprises: An acquisition module is used to obtain the audio to be processed; A classification module, configured to classify the audio features corresponding to the audio to be processed to obtain multiple sound source types; a generation module, configured to perform noise reduction prediction for the multiple sound source types based on a preset recommendation model, generate at least one noise reduction scheme, and display the at least one noise reduction scheme; wherein the preset recommendation model is used to predict the user's noise reduction preference; a response module, configured to determine a target noise reduction scheme in response to a selection operation of the at least one noise reduction scheme; a noise reduction module, configured to perform noise reduction processing on the audio to be processed according to the target noise reduction scheme to obtain target audio; The response module is further configured to, if no response is received to the selection operation of the at least one noise reduction scheme, display the multiple sound source types, and select a target sound source type in response to the sound source type selection operation from among the multiple displayed sound source types; The noise reduction module is further configured to perform noise reduction on the audio to be processed according to the target sound source type to obtain the target audio.
10. An audio processing device, characterized in that: The device includes a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium, characterized in that Executable instructions are stored thereon, which are used to implement the method described in any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Audio processing method and device, computer equipment and computer readable storage medium
CN111540370A
Feature space determining method and device
CN112183653A
Play content recommendation method and device, electronic equipment and storage medium
CN113609387A
Voice recognition method and device, electronic equipment and storage medium
CN113889077A