Audio processing method and device, equipment and storage medium
Through time-frequency conversion, sound source recognition model and adaptive time-frequency filtering algorithm, the target sound source is accurately separated in the audio data, and combined with the time-frequency compensation algorithm to reconstruct the audio data, the problem of difficulty in accurately separating the sound source by fixed-band filters is solved, and the reliability and quality of audio data processing is improved.
Patent Information
- Application Number
- CN202510500505.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing audio data processing methods, it is difficult to accurately separate the target sound source by relying on fixed-band filters, resulting in damage to the sound quality of non-target sound sources and degradation of the overall audio quality.
The time-frequency feature set of audio data is obtained through time-frequency conversion processing, the sound source identification model is used to identify each sound source object, and the target sound source is suppressed through an adaptive time-frequency filtering algorithm, and the audio data is reconstructed in combination with the time-frequency compensation algorithm to ensure the integrity and audio quality of other sound source objects.
It realizes efficient and accurate identification and separation of designated sound sources in audio data, maintains the integrity of other sound source objects, and improves the reliability and applicability of audio data processing.
Smart Images

Figure CN120279933A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio data processing, and particularly to a method, device, equipment and storage medium for processing audio. Background Art
[0002] In the technical field of audio data processing, it involves eliminating the frequency components of a specified sound source in audio data to achieve special customization of the audio data.
[0003] In related audio data processing methods, band-pass, low-pass or high-pass filters with fixed frequency bands are relied on to attenuate the frequency components of the target sound source. However, it is difficult to accurately separate the target sound source by relying on fixed filtering, which damages the sound quality of non-target sound sources and thus leads to a decline in the overall audio quality. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, equipment and storage medium for processing audio to efficiently and accurately identify and separate a specified sound source object in audio data, thereby improving the reliability and applicability of the results of processing audio data.
[0005] In a first aspect, the present application provides a method for processing audio, including: Obtaining the audio data to be processed, performing time-frequency transformation processing on the audio data to obtain a time-frequency feature set of the audio data, and inputting the time-frequency feature set into a preset sound source recognition model for data processing to obtain each sound source object in the audio data; Determining a target sound source object to be eliminated among each sound source object, and suppressing the time-frequency features of the target sound source object in the time-frequency feature set based on a preset adaptive time-frequency filtering algorithm to obtain target audio data from which the target sound source object has been eliminated; Reconstructing the time-frequency features of the target audio data based on a preset time-frequency compensation algorithm to obtain reconstructed audio data corresponding to the target audio data, and performing inverse time-frequency transformation processing on the reconstructed audio data to obtain target reconstructed audio data in a time domain expression corresponding to the reconstructed audio data.
[0006] In a second aspect, the present application further provides a device for processing audio, including: An identification module, configured to obtain the audio data to be processed, perform time-frequency transformation processing on the audio data to obtain a time-frequency feature set of the audio data, and input the time-frequency feature set into a preset sound source recognition model for data processing to obtain each sound source object in the audio data; A separation module, configured to determine a target sound source object to be eliminated among each sound source object, and suppress the time-frequency feature of the target sound source object in the time-frequency feature set based on a preset adaptive time-frequency filtering algorithm, so as to obtain target audio data with the target sound source object eliminated; A reconstruction module, configured to reconstruct the time-frequency feature of the target audio data based on a preset time-frequency compensation algorithm to obtain reconstructed audio data corresponding to the target audio data, and perform an inverse time-frequency transformation on the reconstructed audio data to obtain target reconstructed audio data in the time-domain representation corresponding to the reconstructed audio data.
[0007] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above steps are implemented.
[0008] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above steps are implemented.
[0009] For the above audio processing method, device, equipment and storage medium, first, a time-frequency transformation is performed on the audio data to be processed to obtain a time-frequency feature set, and then the time-frequency feature set is analyzed through a sound source recognition model, so that each independent sound source object in the audio data can be accurately and effectively extracted; furthermore, based on the adaptive time-frequency filtering algorithm, the time-frequency feature of the target sound source object is suppressed in the time-frequency feature set to obtain target audio data with the target sound source object eliminated, so that the target sound source object can be effectively eliminated while maintaining the integrity of other sound source objects, and the customization processing ability of the audio data is improved; furthermore, based on the time-frequency compensation algorithm, the time-frequency feature of the target audio data is reconstructed to obtain reconstructed audio data, so as to repair the information loss problem caused by eliminating the target sound source object to ensure the audio quality; based on this, the specified sound source object is efficiently and accurately identified and separated in the audio data, and the integrity of other sound source objects and the data quality of the overall audio are ensured, thereby improving the reliability and applicability of the result of processing the audio data. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of the audio processing method in an embodiment; Figure 2 It is a structural block diagram of an audio processing device in an embodiment. Specific implementation manners
[0012] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0013] In one embodiment, as Figure 1 shown, a method for processing audio is provided. In this embodiment, it is exemplified that the method is applied to a server. It can be understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S101 to S103.
[0014] Step S101: Obtain the audio data to be processed, perform time-frequency transformation processing on the audio data to obtain a set of time-frequency features of the audio data, and input the set of time-frequency features into a preset sound source recognition model for data processing to obtain each sound source object in the audio data.
[0015] Among them, the audio data to be processed represents the original audio signal that needs to be identified and separated for sound source objects. For example, an audio file containing background noise and human voices collected by a microphone device; a sound source object represents an independent sound component separated from the audio data. For example, in the above-mentioned audio file containing background noise and human voices, the two independent sound components of "background noise" and "human voice" respectively identified are used as two independent sound source objects.
[0016] Among them, the set of time-frequency features represents a set of distribution features of the audio data in the time and frequency dimensions. For example, a time-frequency spectrum matrix obtained by short-time Fourier transform, where the horizontal axis represents time, the vertical axis represents frequency, and different colors or amplitude values represent the energy magnitude of the signal.
[0017] Among them, the sound source recognition model represents an algorithm or neural network model for analyzing the set of time-frequency features to identify the sound source objects in the audio data. For example, a model trained based on a convolutional neural network or an independent component analysis method.
[0018] Exemplarily, first, the audio data to be processed can be obtained through a microphone, a recording device or other audio input source; secondly, the obtained audio data is subjected to time-frequency transformation processing, for example, the audio data is divided into multiple short time periods through short-time Fourier transform, discrete wavelet transform, continuous wavelet transform and other methods, and frequency domain transformation is performed in each short time period, so as to obtain the characteristic distribution of the audio data in the time and frequency dimensions, and form a time-frequency feature set, wherein the obtained time-frequency feature set can represent a set of data matrices containing time dimension, frequency dimension and corresponding amplitude information, for subsequent sound source analysis; thirdly, the obtained time-frequency feature set is input into a preset sound source recognition model for processing, and the time-frequency feature set is subjected to feature extraction and classification through the sound source recognition model, so as to identify each sound source object contained in the audio data, wherein the sound source recognition model can learn the characteristic distribution of large-scale audio data in advance through convolutional neural network, recurrent neural network, independent component analysis-based method, etc., so as to realize the recognition of different sound source objects.
[0019] Step S102, determining a target sound source object to be eliminated from each sound source object, suppressing the time-frequency features of the target sound source object in the time-frequency feature set based on a preset adaptive time-frequency filtering algorithm, and obtaining target audio data from which the target sound source object has been eliminated.
[0020] Among them, the target sound source object represents the sound component selected from each sound source object and needs to be eliminated; the adaptive time-frequency filtering algorithm represents an algorithm for dynamically adjusting the signal strength of different frequencies and different time periods of audio data, for example, a method based on Wiener Filter or Mask-based Filtering, so as to adaptively reduce the energy of the target sound source object while maintaining the energy integrity of other sound source objects as much as possible.
[0021] The target audio data refers to the audio data after the target sound source object has been eliminated after being processed by the adaptive time-frequency filtering algorithm. For example, in an audio file containing background noise and human voice, the audio data containing only human voice is obtained after the background noise is eliminated.
[0022] Exemplarily, first, the target sound source object to be eliminated can be determined among each sound source object by a predefined selection rule for the sound source object or by a selection instruction input by the user; furthermore, the gain of the audio data in different frequency ranges is adjusted in the time-frequency feature set through a preset adaptive time-frequency filtering algorithm to achieve the suppression of the specified target sound source object, and the target audio data with the target sound source object eliminated is obtained. Among them, it can be achieved by methods such as dynamic masking, Wiener filtering, and adaptive filtering. For example: in the dynamic masking method, by analyzing the energy distribution of the target sound source object, the energy of the target sound source object in the time-frequency feature set is reduced or set to zero, so as to achieve the suppression effect; in the Wiener filtering method, the optimal filtering parameters are calculated based on statistical characteristics, and the contribution of the target sound source object is minimized through the optimal filtering parameters, so as to achieve the suppression effect; in the adaptive filtering method, the filtering parameters are continuously adjusted through a preset feedback mechanism to achieve adaptive adjustment of the suppression of the target sound source object, so as to adapt to different scenarios and changes in the input audio data.
[0023] Step S103, based on a preset time-frequency compensation algorithm, perform reconstruction processing on the time-frequency features of the target audio data to obtain the reconstructed audio data corresponding to the target audio data, and perform inverse time-frequency transformation on the reconstructed audio data to obtain the target reconstructed audio data in the time domain corresponding to the reconstructed audio data.
[0024] Among them, the time-frequency compensation algorithm refers to an algorithm used to repair the time-frequency feature loss or distortion caused by the elimination of the target sound source object, and is used to optimize the audio quality and maintain naturalness. For example, through interpolation compensation, residual prediction, or a deep learning model (such as a generative adversarial network), the time-frequency features affected by the filtering suppression are inferred and restored to ensure that the audio listening experience after removing the target sound source object is more natural.
[0025] Among them, the reconstructed audio data refers to an audio signal with a complete time-frequency structure obtained after being processed by the time-frequency compensation algorithm; the target reconstructed audio data refers to the final audio signal after converting the reconstructed audio data back to the time domain through inverse time-frequency transformation.
[0026] Exemplarily, first, the time-frequency characteristics of the target audio data are reconstructed based on a preset time-frequency compensation algorithm. In the time-frequency compensation algorithm, the target audio data can be reconstructed to generate corresponding reconstructed audio data through interpolation compensation, residual prediction, or a deep learning-based feature generation method. For example, in the interpolation compensation method, interpolation calculation is performed based on known time-frequency characteristics to fill in the missing parts during the filtering process according to the interpolation calculation results; in the residual prediction method, the missing time-frequency characteristics are inferred through a trained model and compensated within a possible range; in the deep learning-based feature generation method, models such as generative adversarial networks or variational autoencoders can be used to learn the statistical characteristics of the complete audio data and perform compensatory reconstruction after eliminating the target sound source object. Furthermore, the reconstructed audio data expressed in time-frequency can be converted into the target reconstructed audio data expressed in the time domain through methods such as inverse short-time Fourier transform or inverse wavelet transform.
[0027] In the above audio processing method, first, the time-frequency characteristics set is obtained by performing time-frequency transformation on the audio data to be processed, and then the time-frequency characteristics set is analyzed through a sound source recognition model, so that each independent sound source object in the audio data can be accurately and effectively extracted; furthermore, according to the adaptive time-frequency filtering algorithm, the time-frequency characteristics of the target sound source object in the time-frequency characteristics set are suppressed to obtain the target audio data with the target sound source object eliminated, so that the target sound source object can be effectively eliminated while maintaining the integrity of other sound source objects, improving the customization processing ability of the audio data; furthermore, according to the time-frequency compensation algorithm, the time-frequency characteristics of the target audio data are reconstructed to obtain the reconstructed audio data, thereby repairing the information loss problem caused by eliminating the target sound source object to ensure the audio quality; based on this, the specified sound source object is efficiently and accurately identified and separated in the audio data, and the integrity of other sound source objects and the data quality of the overall audio are ensured, thereby improving the reliability and applicability of the processing result of the audio data.
[0028] In an exemplary embodiment, the time-frequency characteristics set is input into a preset sound source recognition model for data processing to obtain each sound source object in the audio data, including steps S201 to S203.
[0029] In step S201, the time-frequency characteristics set is input into a preset sound source recognition model for feature extraction processing to obtain the energy distribution characteristics of the time-frequency characteristics set at different time periods.
[0030] Among them, the energy distribution characteristics of the time-frequency characteristics set at different time periods represent the frequency energy distribution of the time-frequency characteristics set at each time period. For example, in the data matrix obtained through short-time Fourier transform, each column corresponds to a different time period, and the longitudinal data in each column respectively represent the energy intensities of different frequency components within the corresponding time period.
[0031] Exemplarily, the time-frequency feature set is input into the sound source recognition model for feature extraction processing. During the feature extraction process, the sound source recognition model respectively converts the time-frequency features corresponding to each time period of the time-frequency feature set into a high-dimensional feature vector. Each high-dimensional feature vector represents the energy status of different frequency components within the corresponding time period, that is, the energy distribution characteristics of the time-frequency feature set in different time periods are obtained.
[0032] Step S202: Based on the temporal correlation between the energy distribution characteristics of the time-frequency feature set in adjacent time periods, perform sound source division processing on the time-frequency feature set to obtain each initial sound source object.
[0033] Among them, the initial sound source object represents the sound source object preliminarily divided based on the temporal correlation between the energy distribution characteristics of the time-frequency feature set in adjacent time periods.
[0034] Exemplarily, in the sound source recognition model, based on the temporal correlation between the energy distribution characteristics of the time-frequency feature set in adjacent time periods, and in combination with a feature analysis method based on clustering, perform sound source division processing on the time-frequency feature set to obtain each initial sound source object. That is, by calculating the feature similarity of the energy distribution characteristics of adjacent time periods, similar energy distribution characteristics are classified into the same sound source object, thereby forming the initial sound source object. Furthermore, each initial sound source object contains audio data with similar energy distribution characteristics within a certain time range, and these initial sound source objects still need to be further refined to obtain sound source objects to ensure clear boundaries between different sound source objects.
[0035] Step S203: Based on the multi-dimensional audio components of each initial sound source object in different frequency intervals, calculate the distribution probability of each initial sound source object in different frequency intervals. According to the distribution probability of each initial sound source object in different frequency intervals, screen out each sound source object in the audio data from each initial sound source object.
[0036] Among them, the multi-dimensional audio components of each initial sound source object in different frequency intervals represent the energy distribution of the corresponding initial sound source object in different frequency ranges. For example, in a mixed audio segment, a certain initial sound source object may have higher energy in the low-frequency interval, while another initial sound source object may be mainly concentrated in the mid-high frequency interval.
[0037] Among them, the distribution probability of each initial sound source object in different frequency intervals represents quantifying the contribution degree of the corresponding initial sound source object in different frequency intervals to determine the main frequency components of each initial sound source object.
[0038] Exemplarily, since the audio data may be composed of the sound components of multiple sound source objects in different frequency ranges, a simple time-domain division cannot guarantee the accurate identification of the sound source objects. It is necessary to further analyze the characteristics of each initial sound source object in the frequency dimension, so as to screen out each sound source object in the audio data from each initial sound source object. For example, in the analysis method based on the Gaussian mixture model, under the assumption that the frequency distribution of the audio data follows a mixture of multiple Gaussian distributions, based on the multi-dimensional audio components of each initial sound source object in different frequency ranges, the distribution probability of each initial sound source object in different frequency ranges is calculated by combining the expectation maximization algorithm to obtain the energy contribution of each initial sound source object in different frequency ranges, and then the main composition components of each initial sound source object are determined to refine and adjust the boundaries of each initial sound source object, so as to screen out each sound source object in the audio data from each initial sound source object.
[0039] In this embodiment, first, according to the feature extraction process of the time-frequency feature set, the energy distribution characteristics of the time-frequency feature set in different time periods are obtained, so as to accurately extract the changes of the audio data in time and frequency. Furthermore, according to the time correlation between the energy distribution characteristics of the time-frequency feature set in adjacent time periods, the time-frequency feature set is subjected to sound source division processing to obtain each initial sound source object, so as to ensure that the sound source division can maintain temporal continuity, avoid sound source confusion caused by short-term changes, and improve the stability and accuracy of identifying the sound source object. Furthermore, according to the multi-dimensional audio components of each initial sound source object in different frequency ranges, the distribution probability in different frequency ranges is calculated, and thus the final sound source object is screened out from each initial sound source object, so that the feature expression of each sound source object in the frequency dimension is more explicit, thereby further improving the stability and accuracy of identifying the sound source object.
[0040] In an exemplary embodiment, based on the time correlation between the energy distribution characteristics of the time-frequency feature set in adjacent time periods, the time-frequency feature set is subjected to sound source division processing to obtain each initial sound source object, including steps S301 to S302.
[0041] Step S301, based on the time correlation between the energy distribution characteristics of the time-frequency feature set in adjacent time periods, divide the time-frequency feature set into multiple time-frequency feature subsets, and each time-frequency feature subset corresponds to the energy distribution characteristics of adjacent time periods whose time correlation satisfies the preset threshold condition.
[0042] Among them, the time-frequency feature subset represents multiple subsets divided from the time-frequency feature set, and each subset contains time-frequency feature data that is continuous in time and has similar energy distribution.
[0043] Among them, the preset threshold condition represents a judgment criterion for measuring the temporal correlation between adjacent time periods, and is used to determine which time-frequency feature data of adjacent time periods should be divided into the same time-frequency feature subset.
[0044] Exemplarily, through calculation methods such as Euclidean distance and cosine similarity, the feature similarity degree of the energy distribution features of the time-frequency feature set in the time correlation level between adjacent time periods can be measured. If the feature similarity degree corresponding to the energy distribution features of two adjacent time periods is greater than a preset threshold, it indicates that the energy distribution features of the adjacent time periods have small differences. Further, it is determined that the temporal correlation of the energy distribution features of the adjacent time periods satisfies the preset threshold condition, so as to classify the energy distribution features of the adjacent time periods whose temporal correlation satisfies the preset threshold condition into the same time-frequency feature subset.
[0045] Step S302: Perform statistical analysis processing on the spectral shape features of each time-frequency feature subset to obtain audio components that maintain similar spectral shape features within consecutive time-frequency feature subsets. Based on the audio components that maintain similar spectral shape features within consecutive time-frequency feature subsets, each initial sound source object is identified.
[0046] Among them, the spectral shape feature represents the frequency distribution pattern in the time-frequency feature subset, including features such as energy distribution, harmonic structure, and amplitude change. For example, human voices have a relatively stable harmonic structure, while environmental noise is manifested as randomly distributed broadband noise; the audio component represents a sound unit extracted from consecutive time-frequency feature subsets with similar spectral shape features.
[0047] Exemplarily, first, the principal component analysis method can be used to perform dimensionality reduction analysis on the spectral shape features of each time-frequency feature subset to extract the audio components of each time-frequency feature subset and the spectral shape features to which the audio components belong. Furthermore, the clustering analysis method can be used to perform clustering statistical analysis processing on the spectral shape features within consecutive time-frequency feature subsets to obtain audio components that maintain similar spectral shape features within consecutive time-frequency feature subsets, so as to realize that the audio components that show high consistency in both the time and frequency dimensions are used as the audio components corresponding to each initial sound source object.
[0048] In this embodiment, first, based on the temporal correlation between the energy distribution features of the time-frequency feature set in adjacent time periods, the time-frequency feature set is divided into multiple time-frequency feature subsets, so as to ensure a complete feature division process at the level of temporal correlation. Furthermore, statistical analysis processing is performed according to the spectral shape features of each time-frequency feature subset to identify audio components that maintain similar spectral shape features within consecutive time-frequency feature subsets, so as to accurately distinguish each initial sound source object in the time-frequency dimension.
[0049] In an exemplary embodiment, based on the multi-dimensional audio components of each initial sound source object in different frequency ranges, the distribution probabilities of each initial sound source object in different frequency ranges are calculated, including steps S401 to S403.
[0050] Step S401: Based on a preset frequency resolution and time resolution, perform multi-dimensional audio decomposition processing on the time-frequency features of each initial sound source object to obtain the multi-dimensional audio components of each initial sound source object in different frequency ranges.
[0051] Among them, the frequency resolution represents the division precision on the frequency axis during the decomposition of the time-frequency features, and is used to determine the minimum frequency interval that can be distinguished when analyzing the time-frequency features; the time resolution represents the division precision on the time axis during the decomposition of the time-frequency features, and is used to determine the minimum time interval that can be distinguished when analyzing the time-frequency features.
[0052] Exemplarily, first, reasonable frequency resolution and time resolution need to be determined. A higher frequency resolution can make the spectral structure of the initial sound source object clearer, and a higher time resolution helps to capture the sound features of short-term changes. Furthermore, based on the determined frequency resolution and time resolution, perform multi-dimensional audio decomposition processing on the time-frequency features of each initial sound source object, so that the time-frequency features of each initial sound source object are parsed into audio components in multiple frequency ranges. For example: decompose the time-frequency features of each initial sound source object through methods such as short-time Fourier transform and non-negative matrix factorization to obtain the audio components of each initial sound source object in different frequency ranges respectively, which are used as the multi-dimensional audio components of each initial sound source object in different frequency ranges.
[0053] Step S402: Obtain a preset probability distribution function, and based on the distribution density characteristics corresponding to the multi-dimensional audio components of each initial sound source object in different frequency ranges, obtain the cumulative probability values of each initial sound source object in different frequency ranges of the probability distribution function respectively.
[0054] Among them, the probability distribution function represents a function used to describe the statistical characteristics of each initial sound source object in different frequency ranges, and is used to calculate the distribution probability of the corresponding initial sound source object in different frequency ranges; the cumulative probability value represents the proportion of the cumulative energy of each initial sound source object in different frequency ranges, and is used to measure the cumulative occurrence probability of the corresponding initial sound source object in the corresponding frequency range.
[0055] Among them, the distribution density feature represents the energy density distribution of each initial sound source object in different frequency intervals, and is used to measure the contribution degree of the corresponding initial sound source object at each frequency. For example, the high-density feature of a certain initial sound source object in a certain frequency interval can indicate that the main frequency components of the initial sound source object are concentrated in this frequency range.
[0056] Exemplarily, different data models such as normal distribution, Poisson distribution or mixture Gaussian distribution can be used in advance to fit the distribution density feature of the audio signal with the distribution probability value, so as to obtain the probability distribution function corresponding to the initial sound source object, and this probability distribution function is used to quantitatively describe the statistical characteristics of the audio signal in different frequency ranges; furthermore, the distribution density features corresponding to the multi-dimensional audio components of each initial sound source object in different frequency intervals are respectively input into the corresponding probability distribution function, and then the probability distribution function uses the numerical integration method to evaluate and obtain the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function, so as to quantitatively represent the cumulative energy proportion of each initial sound source object in different frequency intervals.
[0057] Step S403: Use the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function as the distribution probabilities of each initial sound source object in different frequency intervals.
[0058] Exemplarily, the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function can be normalized to obtain the normalized distribution probabilities of each initial sound source object in different frequency intervals, so that the frequency contribution degrees of each initial sound source object in different frequency intervals can be directly measured by the distribution probabilities.
[0059] In this embodiment, first, the time-frequency features of each initial sound source object are subjected to multi-dimensional audio decomposition processing to obtain the multi-dimensional audio components of each initial sound source object in different frequency intervals, so as to accurately extract the fine-grained audio components of each initial sound source object in different frequency intervals; furthermore, according to the distribution density features corresponding to the multi-dimensional audio components of each initial sound source object in different frequency intervals, the cumulative probability values in different frequency intervals of the probability distribution function are calculated to obtain the distribution probabilities of each initial sound source object in different frequency intervals, so as to effectively quantify the cumulative energy proportion and frequency contribution degree of each initial sound source object in different frequency ranges through statistical methods.
[0060] In an exemplary embodiment, according to the distribution probabilities of each initial sound source object in different frequency intervals, each sound source object in the audio data is screened out from each initial sound source object, including steps S501 to S503.
[0061] Step S501: Based on the distribution probabilities of each initial sound source object in different frequency intervals and in combination with the distribution probability threshold conditions for different frequency intervals, determine the characteristic stability indicators corresponding to each initial sound source object in all frequency intervals respectively.
[0062] Among them, the distribution probability threshold condition refers to the threshold standard used to determine whether the distribution probabilities of each initial sound source object in different frequency intervals meet the stability requirements. For example, when the distribution probability of a certain initial sound source object in a certain frequency interval is lower than the set threshold, it may be considered that this frequency interval does not belong to the main component of this initial sound source object, so as to adjust the boundary of this initial sound source object.
[0063] Among them, the characteristic stability indicator refers to the stability measurement parameter obtained by statistically analyzing the distribution probabilities of each initial sound source object in different frequency intervals, and is used to measure the energy distribution consistency of the corresponding initial sound source object in different frequency ranges. For example, if the distribution of a certain initial sound source object in the entire frequency interval is relatively stable, it indicates that its characteristic stability indicator is relatively high, while if the energy distribution fluctuates greatly, its characteristic stability indicator is relatively low.
[0064] Exemplarily, the time-frequency characteristics of the initial sound source object will be affected by different factors in different frequency ranges. For example, the frequency spectra of some initial sound source objects may be stably distributed in some frequency ranges, while significant fluctuations occur in other frequency ranges. Therefore, it is necessary to evaluate the characteristic stability indicators of each initial sound source object in different frequency intervals, that is, to examine the energy change situation in different frequency ranges and whether it meets certain stability requirements. Among them: if the distribution probabilities of a certain initial sound source object in different frequency ranges all meet the corresponding distribution probability threshold conditions, it indicates that this initial sound source object has high stability in each frequency range, and a high characteristic stability indicator needs to be assigned to this initial sound source object; while if the distribution probabilities of a certain initial sound source object in some frequency ranges do not meet the corresponding distribution probability threshold conditions, it indicates that this initial sound source object has low stability in the above-mentioned some frequency ranges, and a low characteristic stability indicator needs to be assigned to this initial sound source object.
[0065] Step S502: Based on the distribution probabilities of each initial sound source object in the edge frequency intervals and in combination with the distribution probability threshold conditions for the corresponding edge frequency intervals, adjust the boundary parameters of each initial sound source object to obtain the adjusted boundary parameters corresponding to each initial sound source object respectively.
[0066] Among them, the boundary parameter refers to the boundary value used when adjusting the frequency range of each initial sound source object, and is used to determine the frequency spectrum range of the corresponding initial sound source object in the time-frequency domain.
[0067] Exemplarily, it is necessary to reasonably adjust the boundary parameters of the initial sound source object to ensure that the frequency range of the initial sound source object neither contains too many irrelevant components nor misses key frequency information, where: it is possible to determine whether to expand or narrow the frequency boundary range of the initial sound source object by judging whether the distribution probability of the frequency interval of the initial sound source object in the edge region meets the corresponding distribution probability threshold condition. For example, if the distribution probability of a certain initial sound source object in the edge frequency interval does not meet the corresponding distribution probability threshold condition, the frequency boundary range of the initial sound source object can be appropriately narrowed to reduce the interference of other sound sources; if the distribution probability of a certain initial sound source object in the edge frequency interval meets the corresponding distribution probability threshold condition, the frequency boundary range of the initial sound source object can be appropriately expanded to ensure the integrity of the initial sound source object.
[0068] Step S503, combining the characteristic stability indexes of each initial sound source object and the adjusted boundary parameters, screen out each sound source object in the audio data from each initial sound source object.
[0069] Exemplarily, comprehensively consider the characteristic stability indexes and the adjusted boundary parameters of each initial sound source object to determine the final sound source category among each initial sound source object. Among them, through the weighted scoring method, the characteristic stability indexes and boundary parameters of each initial sound source object can be weighted and comprehensively calculated, so that multiple initial sound source objects with higher weighted scores are used as the screened sound source objects. For example, for an initial sound source object with a higher characteristic stability index and a clear frequency boundary, it can be preferentially selected, while for an initial sound source object with a lower characteristic stability index or still having great uncertainty after boundary adjustment, it can be considered for elimination.
[0070] In this embodiment, first, according to the distribution probability of each initial sound source object in different frequency intervals and in combination with the distribution probability threshold conditions of different frequency intervals, determine the characteristic stability indexes corresponding to each initial sound source object in all frequency intervals, so as to effectively quantify the characteristic stability of each initial sound source object; furthermore, according to the distribution probability of each initial sound source object in the edge frequency interval and in combination with the corresponding distribution probability threshold conditions of the edge frequency interval, adjust the boundary parameters of each initial sound source object, so as to ensure that the frequency spectrum range of each initial sound source object can be more accurate, neither missing its main frequency components nor including irrelevant frequency information; furthermore, combining the characteristic stability indexes of each initial sound source object and the adjusted boundary parameters, screen out each sound source object in the audio data from each initial sound source object, so as to ensure that the finally screened sound source objects have a stable energy distribution and accurate boundary discrimination in the time-frequency domain.
[0071] In an exemplary embodiment, suppressing the time-frequency features of a target sound source object in a time-frequency feature set based on a preset adaptive time-frequency filtering algorithm to obtain target audio data with the target sound source object eliminated, including steps S601 to S603.
[0072] Step S601: Combining the distribution characteristics of the time-frequency features of the target sound source object in the time dimension and the frequency dimension in the time-frequency feature set respectively, to determine the energy weight distribution of the target sound source object at different time-frequencies in the time-frequency feature set.
[0073] Among them, the distribution characteristics in the time dimension and the frequency dimension represent the energy distribution law of the time-frequency features of the corresponding target sound source object changing with time and frequency in the time-frequency feature set.
[0074] Among them, the energy weight distribution at different time-frequencies represents the energy contribution of the time-frequency features of the corresponding target sound source object at different time points and different frequency points in the time-frequency feature set.
[0075] Exemplarily, in the time dimension, the energy distribution of the target sound source object may show high energy concentration in some time periods, while the energy is lower or completely disappears in other time periods. In the frequency dimension, the energy distribution of some target sound source objects may be concentrated in the low-frequency range, while the energy distribution of some target sound source objects may be concentrated in the medium-high frequency range. Based on this, it is necessary to combine the distribution characteristics of the time-frequency features of the target sound source object in the time dimension and the frequency dimension in the time-frequency feature set respectively, and perform statistical calculations on the time-frequency features of the target sound source object by calculating the energy distribution matrix, so as to clarify the energy contribution ratio of the target sound source object at different time-frequencies, and then convert to obtain the energy weight distribution of the target sound source object at different time-frequencies in the time-frequency feature set.
[0076] Step S602: Based on the energy weight distribution of the target sound source object at different time-frequencies, combining the preset adaptive time-frequency filtering algorithm, to determine the filtering weights of the target sound source object at different time-frequencies.
[0077] Among them, the filtering weights at different time-frequencies represent the filtering intensity set at the corresponding time point and the corresponding frequency point according to the energy contribution of the corresponding target sound source object at different time points and different frequency points.
[0078] Exemplarily, since the energy distribution of the target sound source object is non-uniform at different times and frequencies, during the filtering process based on the adaptive time-frequency filtering algorithm, it is necessary to allocate appropriate filtering weights to the target sound source object at different time-frequency points, so that the filtering process can effectively suppress the time-frequency characteristics of the target sound source object without significantly interfering with the components of non-target sound source objects. Among them, in the process of calculating the filtering weights, in combination with the energy weight distribution of the target sound source object at different time-frequency points, the filtering weights are allocated to each time-frequency point according to the preset filtering rules in the adaptive time-frequency filtering algorithm, so that the filtering process can adapt to different audio environments through a dynamic adjustment mechanism. For example, for the target sound source object in the high-energy region, the filtering weight needs to be appropriately enhanced to ensure that this part of the energy can be fully suppressed, while for the target sound source object in the low-energy region, the filtering weight needs to be reduced to reduce the impact on other non-target sound source objects. Step S603: Based on the filtering weights of the target sound source object at different time-frequencies, suppress the time-frequency characteristics of the target sound source object in the time-frequency feature set through the adaptive time-frequency filtering algorithm to obtain the target audio data from which the target sound source object has been eliminated.
[0079] Exemplarily, the adaptive time-frequency filtering algorithm processes the time-frequency characteristics of the target sound source object point by point according to the filtering weights of the target sound source object at different time-frequencies, so that the energy of the target sound source object at different time-frequencies is gradually reduced or completely suppressed. Furthermore, the adaptive time-frequency filtering algorithm usually dynamically adjusts the filtering weights of the target sound source object at different time-frequencies in the time-frequency domain to match the energy characteristics of the target sound source object at different time-frequency points, so that the energy of the target sound source object can be reasonably attenuated without introducing additional spectral distortion or artifacts. After the filtering process is completed, the energy of the target sound source object will be effectively weakened, so that the finally obtained target audio data no longer contains the main characteristics of the target sound source object, while the time-frequency characteristics of other non-target sound source objects remain intact.
[0080] In this embodiment, first, according to the distribution characteristics of the target sound source object in the time dimension and frequency dimension of the time-frequency feature set, the energy weight distribution of the target sound source object at different time-frequencies is determined, so as to accurately depict the energy distribution law of the target sound source object in the multi-dimensional time-frequency; furthermore, according to the energy weight distribution of the target sound source object at different time-frequencies and in combination with a preset adaptive time-frequency filtering algorithm, the filtering weights of the target sound source object at different time-frequencies are determined, so as to dynamically adjust the filtering intensity of different time-frequency points, making the filtering process effectively adapt to the energy change of the target sound source object and improving the accuracy and self-adaptability of sound source suppression; furthermore, according to the filtering weights of the target sound source object at different time-frequencies, the time-frequency features of the target sound source object are suppressed by the adaptive time-frequency filtering algorithm, so as to effectively achieve the refined removal of the target sound source object, improve the accuracy of sound source separation, and at the same time reduce the interference to other sound sources.
[0081] In an exemplary embodiment, the time-frequency features of the target audio data are reconstructed based on a preset time-frequency compensation algorithm to obtain the reconstructed audio data corresponding to the target audio data, including steps S701 to S703.
[0082] Step S701, based on the difference between the time-frequency features of the target audio data and the time-frequency feature set, determine the time-frequency feature missing information corresponding to the target audio data.
[0083] Wherein, the time-frequency feature missing information represents the time-frequency component information lost or weakened by the target audio data after sound source suppression or filtering processing compared with the original time-frequency feature set, and is used to identify and quantify the eliminated or affected part of the original audio data.
[0084] Exemplarily, compare the time-frequency features of the target audio data with the complete time-frequency feature set to statistically obtain the deviation of the time-frequency features of the target audio data relative to the time-frequency feature set. For example, relative to the complete time-frequency feature set, if the energy attenuation in some frequency regions of the time-frequency features of the target audio data is too large or the spectral information in some time regions shows abnormal changes, it can be considered that there are feature losses in these specified regions, and thus the time-frequency features in the time-frequency feature set related to the above specified regions are used as the time-frequency feature missing information; furthermore, the above specified regions can represent the time-frequency features corresponding to the target sound source object eliminated in the time-frequency feature set, or can also represent the affected part of the time-frequency features of other non-target sound source objects during the process of eliminating the time-frequency features of the target sound source object, such as the time-frequency features missing or distorted due to being affected.
[0085] Step S702: Based on the continuity of the time-frequency feature missing information with the time-frequency features of adjacent time-frequencies in the time-frequency feature set, and in combination with a preset time-frequency compensation algorithm, determine the compensated time-frequency features corresponding to the time-frequency feature missing information.
[0086] Among them, the compensated time-frequency features represent the time-frequency component information used to reconstruct the missing part caused by the time-frequency feature missing information in the target audio data.
[0087] Exemplarily, based on the distribution characteristics of the time-frequency feature missing information in the time-frequency feature set, and in combination with the continuity of the time-frequency features of adjacent time-frequencies, determine the compensated time-frequency features corresponding to the time-frequency feature missing information, so that the reconstructed audio data based on the compensated time-frequency features maintains a normal time-frequency structure as a whole on the premise of not reflecting the time-frequency features of the target sound source object, so that the reconstructed audio data can maintain a smooth transition in both the time and frequency dimensions, without introducing abrupt breakpoints or abnormal frequency gaps, that is, when compensating, it is necessary to ensure that the newly generated time-frequency features can fill the eliminated part without reintroducing the time-frequency features related to the target sound source object. Among them, in the process of calculating the compensated time-frequency features, it is necessary to combine the continuity of the time-frequency features of adjacent time-frequencies and analyze the energy change trend of adjacent time-frequency points to infer a suitable compensation value. For example: if the missing area of a certain time-frequency point is within a stable background signal range, the time-frequency features of the surrounding time-frequencies can be directly used to calculate a reasonable compensation value for the missing area, and then filled by interpolation to maintain the smooth transition of the signal; for the missing area in a complex audio scene, such as a situation accompanied by sudden changes or dynamic changes, a prediction model or statistical method needs to be used to calculate a reasonable compensation value for the missing area.
[0088] Step S703: Based on the compensated time-frequency features corresponding to the time-frequency feature missing information, perform a reconstruction process on the time-frequency features of the target audio data through the time-frequency compensation algorithm to obtain the reconstructed audio data corresponding to the target audio data.
[0089] Exemplarily, in the reconstruction process, in the time-frequency compensation algorithm, through methods such as interpolation, regression analysis, and statistical modeling, fuse the compensated time-frequency features with the target audio data to generate the reconstructed audio data, that is, reconstruct the missing area of the target audio data through the compensated time-frequency features to generate the reconstructed audio data, so that the reconstructed audio data maintains overall coherence in the time-frequency domain and has reasonable spectral characteristics, ensuring that no abrupt energy changes or additional spectral distortions are introduced, and not showing the time-frequency features of the target sound source object.
[0090] In this embodiment, first, based on the difference between the time-frequency features of the target audio data and the time-frequency feature set, the time-frequency feature missing information is efficiently determined, so that the loss of time-frequency components caused by sound source suppression or filtering processing can be accurately identified. Furthermore, based on the continuity of the time-frequency features of the time-frequency feature missing information and the adjacent time-frequency in the time-frequency feature set, combined with a preset time-frequency compensation algorithm, the compensated time-frequency features are determined, so that the missing part caused by the time-frequency feature missing information in the target audio data can be effectively inferred at the associated level of continuous time-frequency. Furthermore, the time-frequency compensation algorithm reconstructs the time-frequency features of the target audio data according to the compensated time-frequency features to obtain the reconstructed audio data, so as to effectively restore the integrity of the target audio data, maintain a smooth frequency spectrum structure after removing the target sound source object, and avoid the deterioration of sound quality or information loss caused by filtering processing.
[0091] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0092] Based on the same inventive concept, an embodiment of the present application also provides a processing device for audio for implementing the above-mentioned audio processing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the audio processing device provided below can refer to the limitations on the audio processing method in the above text, and will not be repeated here.
[0093] In an exemplary embodiment, as Figure 2 shown, a processing device for audio is provided, including: an identification module 201, a separation module 202, and a reconstruction module 203, where: The identification module 201 is configured to obtain the audio data to be processed, perform time-frequency transformation processing on the audio data to obtain a time-frequency feature set of the audio data, and input the time-frequency feature set into a preset sound source identification model for data processing to obtain each sound source object in the audio data; A separation module 202 is configured to determine a target sound source object to be eliminated among each sound source object, and suppress the time-frequency feature of the target sound source object in the time-frequency feature set based on a preset adaptive time-frequency filtering algorithm, so as to obtain target audio data with the target sound source object eliminated. A reconstruction module 203 is configured to reconstruct the time-frequency feature of the target audio data based on a preset time-frequency compensation algorithm, so as to obtain reconstructed audio data corresponding to the target audio data, and perform an inverse time-frequency transformation on the reconstructed audio data to obtain target reconstructed audio data in the time domain corresponding to the reconstructed audio data.
[0094] In an exemplary embodiment, the recognition module 201 is further configured to: input the time-frequency feature set into a preset sound source recognition model for feature extraction processing to obtain the energy distribution feature of the time-frequency feature set in different time periods; based on the time correlation between the energy distribution features of the time-frequency feature set in adjacent time periods, perform sound source division processing on the time-frequency feature set to obtain each initial sound source object; calculate the distribution probability of each initial sound source object in different frequency intervals based on the multi-dimensional audio components of each initial sound source object in different frequency intervals, and screen each sound source object in the audio data from each initial sound source object according to the distribution probability of each initial sound source object in different frequency intervals.
[0095] In an exemplary embodiment, the recognition module 201 is further configured to: based on the time correlation between the energy distribution features of the time-frequency feature set in adjacent time periods, divide the time-frequency feature set into multiple time-frequency feature subsets, and each time-frequency feature subset respectively corresponds to the energy distribution feature of adjacent time periods whose time correlation satisfies a preset threshold condition; perform statistical analysis processing on the spectral shape features of each time-frequency feature subset to obtain audio components that maintain similar spectral shape features within consecutive time-frequency feature subsets, and recognize each initial sound source object based on the audio components that maintain similar spectral shape features within consecutive time-frequency feature subsets.
[0096] In an exemplary embodiment, the recognition module 201 is further configured to: perform multi-dimensional audio decomposition processing on the time-frequency features of each initial sound source object based on a preset frequency resolution and time resolution to obtain multi-dimensional audio components of each initial sound source object in different frequency intervals; obtain a preset probability distribution function, and obtain the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function based on the distribution density features corresponding to the multi-dimensional audio components of each initial sound source object in different frequency intervals; use the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function as the distribution probabilities of each initial sound source object in different frequency intervals.
[0097] In an exemplary embodiment, the recognition module 201 is further configured to: determine the feature stability indicators corresponding to each initial sound source object in all frequency ranges according to the distribution probabilities of each initial sound source object in different frequency ranges and in combination with the distribution probability threshold conditions of different frequency ranges; adjust the boundary parameters of each initial sound source object according to the distribution probabilities of each initial sound source object in the marginal frequency ranges and in combination with the corresponding distribution probability threshold conditions of the marginal frequency ranges, so as to obtain the adjusted boundary parameters corresponding to each initial sound source object; and screen out each sound source object in the audio data from each initial sound source object by combining the feature stability indicators and the adjusted boundary parameters of each initial sound source object.
[0098] In an exemplary embodiment, the separation module 202 is further configured to: determine the energy weight distribution of the target sound source object in different time-frequencies in the time-frequency feature set according to the distribution characteristics of the time-frequency features of the target sound source object in the time dimension and the frequency dimension in the time-frequency feature set respectively; determine the filtering weights of the target sound source object in different time-frequencies based on the energy weight distribution of the target sound source object in different time-frequencies and in combination with a preset adaptive time-frequency filtering algorithm; and suppress the time-frequency features of the target sound source object in the time-frequency feature set through the adaptive time-frequency filtering algorithm based on the filtering weights of the target sound source object in different time-frequencies, so as to obtain the target audio data with the target sound source object eliminated.
[0099] In an exemplary embodiment, the reconstruction module 203 is further configured to: determine the missing time-frequency feature information corresponding to the target audio data based on the difference between the time-frequency features of the target audio data and the time-frequency feature set; determine the compensated time-frequency features corresponding to the missing time-frequency feature information based on the continuity between the missing time-frequency feature information and the time-frequency features of adjacent time-frequencies in the time-frequency feature set and in combination with a preset time-frequency compensation algorithm; and reconstruct the time-frequency features of the target audio data through the time-frequency compensation algorithm based on the compensated time-frequency features corresponding to the missing time-frequency feature information, so as to obtain the reconstructed audio data corresponding to the target audio data.
[0100] Each module in the above audio processing device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0101] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in any of the above embodiments are implemented.
[0102] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any of the above embodiments are implemented.
[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0104] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0105] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for processing audio, characterized in that, The method includes: Obtain the audio data to be processed, perform time-frequency transformation processing on the audio data to obtain the time-frequency feature set of the audio data, and input the time-frequency feature set into a preset sound source recognition model for data processing to obtain each sound source object in the audio data; Determine the target sound source object to be eliminated among each sound source object, and based on a preset adaptive time-frequency filtering algorithm, suppress the time-frequency features of the target sound source object in the time-frequency feature set to obtain the target audio data with the target sound source object eliminated; Based on a preset time-frequency compensation algorithm, reconstruct the time-frequency features of the target audio data to obtain the reconstructed audio data corresponding to the target audio data, and perform inverse time-frequency transformation processing on the reconstructed audio data to obtain the target reconstructed audio data with a time-domain expression corresponding to the reconstructed audio data.
2. The method according to claim 1, wherein The step of inputting the time-frequency feature set into a preset sound source recognition model for data processing to obtain each sound source object in the audio data includes: Input the time-frequency feature set into a preset sound source recognition model for feature extraction processing to obtain the energy distribution features of the time-frequency feature set in different time periods; Based on the time correlation between the energy distribution features of the time-frequency feature set in adjacent time periods, perform sound source division processing on the time-frequency feature set to obtain each initial sound source object; Based on the multi-dimensional audio components of each initial sound source object in different frequency intervals, calculate the distribution probabilities of each initial sound source object in different frequency intervals, and according to the distribution probabilities of each initial sound source object in different frequency intervals, screen and obtain each sound source object in the audio data among each initial sound source object.
3. The method according to claim 2, wherein The step of based on the time correlation between the energy distribution features of the time-frequency feature set in adjacent time periods, perform sound source division processing on the time-frequency feature set to obtain each initial sound source object includes: Based on the time correlation between the energy distribution features of the time-frequency feature set in adjacent time periods, divide the time-frequency feature set into multiple time-frequency feature subsets, and each time-frequency feature subset respectively corresponds to the energy distribution features of adjacent time periods whose time correlation satisfies a preset threshold condition; Perform statistical analysis processing on the spectral shape features of each time-frequency feature subset to obtain the audio components that maintain similar spectral shape features within continuous time-frequency feature subsets, and based on the audio components that maintain similar spectral shape features within continuous time-frequency feature subsets, identify each initial sound source object.
4. The method according to claim 2, wherein The step of based on the multi-dimensional audio components of each initial sound source object in different frequency intervals, calculate the distribution probabilities of each initial sound source object in different frequency intervals includes: Based on a preset frequency resolution and time resolution, perform multi-dimensional audio decomposition processing on the time-frequency features of each initial sound source object to obtain the multi-dimensional audio components of each initial sound source object in different frequency intervals; Obtain a preset probability distribution function, and based on the distribution density characteristics corresponding to the multi-dimensional audio components of each initial sound source object in different frequency intervals, obtain the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function respectively; Use the cumulative probability values of each initial sound source object in different frequency intervals of the probability distribution function as the distribution probabilities of each initial sound source object in different frequency intervals.
5. The method according to claim 2, wherein The screening of each sound source object in the audio data from each initial sound source object according to the distribution probabilities of each initial sound source object in different frequency intervals includes: According to the distribution probabilities of each initial sound source object in different frequency intervals, and in combination with the distribution probability threshold conditions of different frequency intervals, determine the characteristic stability indexes of each initial sound source object corresponding to all frequency intervals respectively; According to the distribution probabilities of each initial sound source object in the marginal frequency intervals, and in combination with the distribution probability threshold conditions of the corresponding marginal frequency intervals, adjust the boundary parameters of each initial sound source object to obtain the adjusted boundary parameters corresponding to each initial sound source object respectively; In combination with the characteristic stability indexes and the adjusted boundary parameters of each initial sound source object, screen each sound source object in the audio data from each initial sound source object.
6. The method according to claim 1, wherein The suppression processing of the time-frequency characteristics of the target sound source object in the time-frequency feature set based on a preset adaptive time-frequency filtering algorithm to obtain the target audio data from which the target sound source object has been eliminated includes: In combination with the distribution characteristics of the time-frequency characteristics of the target sound source object in the time dimension and the frequency dimension in the time-frequency feature set respectively, determine the energy weight distribution status of the target sound source object at different time-frequencies in the time-frequency feature set; Based on the energy weight distribution status of the target sound source object at different time-frequencies, and in combination with a preset adaptive time-frequency filtering algorithm, determine the filtering weights of the target sound source object at different time-frequencies; Based on the filtering weights of the target sound source object at different time-frequencies, perform suppression processing on the time-frequency characteristics of the target sound source object in the time-frequency feature set through the adaptive time-frequency filtering algorithm to obtain the target audio data from which the target sound source object has been eliminated.
7. The method according to claim 1, wherein The reconstruction processing of the time-frequency characteristics of the target audio data based on a preset time-frequency compensation algorithm to obtain the reconstructed audio data corresponding to the target audio data includes: Based on the difference between the time-frequency characteristics of the target audio data and the time-frequency feature set, determine the time-frequency feature missing information corresponding to the target audio data; Based on the continuity of the time-frequency characteristics of the time-frequency feature missing information in the time-frequency feature set with the time-frequency characteristics of adjacent time-frequencies, and in combination with a preset time-frequency compensation algorithm, determine the compensated time-frequency characteristics corresponding to the time-frequency feature missing information; Based on the compensated time-frequency characteristics corresponding to the time-frequency feature missing information, perform reconstruction processing on the time-frequency characteristics of the target audio data through the time-frequency compensation algorithm to obtain the reconstructed audio data corresponding to the target audio data.
8. An audio processing device, characterized in that, The device includes: An identification module, configured to obtain audio data to be processed, perform time-frequency transformation processing on the audio data to obtain a time-frequency feature set of the audio data, and input the time-frequency feature set into a preset sound source identification model for data processing to obtain each sound source object in the audio data; A separation module, configured to determine a target sound source object to be eliminated among each sound source object, and perform suppression processing on the time-frequency features of the target sound source object in the time-frequency feature set based on a preset adaptive time-frequency filtering algorithm to obtain target audio data with the target sound source object eliminated; A reconstruction module, configured to perform reconstruction processing on the time-frequency features of the target audio data based on a preset time-frequency compensation algorithm to obtain reconstructed audio data corresponding to the target audio data, and perform inverse time-frequency transformation processing on the reconstructed audio data to obtain target reconstructed audio data in a time-domain expression corresponding to the reconstructed audio data.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.