Audio data processing method, computer equipment and computer storage medium
By identifying common features in audio data, combining audio tracks with similar functions into one group, and using music information retrieval technology MIR for processing, the problems of large amount of computing and low efficiency in existing audio track division processing are solved, and efficient audio processing is achieved.
Patent Information
- Application Number
- CN202510516677.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-01
AI Technical Summary
The existing audio track division processing methods require individual processing of each track, resulting in large amounts of computing, low efficiency and increased manual workload, making it difficult to meet the needs of real-time rendering and efficient processing.
By identifying common features in audio data, combining tracks with similar functions into one group, reducing the processing calculation amount, and using music information retrieval technology MIR to identify track signal characteristics and perform superimposition processing.
It significantly reduces the amount of calculation of track processing, improves the efficiency of audio processing, reduces the workload of manually setting parameters, and improves the speed and quality of audio processing.
Smart Images

Figure CN120236603A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio processing, and specifically to an audio data processing method, a computer device, and a computer storage medium. Background Art
[0002] With the maturity of audio coding and decoding technologies and the increase in data transmission bandwidth, some audio processing methods and applications based on audio tracks have emerged. Especially in the fields of music production, movie sound effect design, and virtual reality, there are many applications of audio track processing. These methods often process the original music master tape or the tracks after sound source separation.
[0003] However, there are some potential performance problems with this processing method, that is, when processing each audio, for example, when processing each song, it is necessary to process multiple tracks of the song separately. Therefore, the amount of computation for processing each song increases sharply, resulting in the inability to improve the audio processing efficiency. Summary of the Invention
[0004] The embodiments of the present application disclose an audio data processing method, a computer device, and a computer storage medium. By combining audio tracks with similar functions into a group, the amount of processing computation is significantly reduced, while maintaining the quality of audio processing.
[0005] The first aspect of the embodiments of the present application provides an audio data processing method, and the method includes:
[0006] Obtain N track signals of target audio data; where N is a positive integer greater than 1;
[0007] Extract the signal features of each track signal of the target audio data;
[0008] Determine the common features of the signal features of at least two of the N track signals;
[0009] Superimpose the at least two track signals including the common features to obtain a superimposed track signal;
[0010] Perform signal processing on the superimposed track signal to obtain processed audio data.
[0011] The second aspect of the embodiments of the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method of the first aspect is implemented.
[0012] The third aspect of the embodiments of the present application provides a computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer is caused to execute the method of the first aspect.
[0013] In the fourth aspect of the embodiments of the present application, a computer program product is provided. When the computer program product runs on a computer device, the computer device is caused to execute the method in the foregoing first aspect.
[0014] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0015] By identifying the common features of multiple track signals of the target audio data, multiple track signals including the common features are superimposed into the same functional signal to obtain a superimposed track signal. The processed audio data can be directly obtained by processing the superimposed track signal. Compared with processing each track signal separately, the number of tracks to be processed is greatly reduced, the computation amount of track signal processing can be significantly reduced, and thus the audio processing efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of a network framework in an embodiment of the present application;
[0017] Figure 2 It is a schematic flowchart of a method for processing audio data in an embodiment of the present application;
[0018] Figure 2A Based on Figure 2 the shown embodiment, it is an exemplary effect schematic diagram of superimposing at least two track signals including common features to obtain a superimposed track signal;
[0019] Figure 2B Based on Figure 2 the shown embodiment, it is an exemplary flowchart of processing multiple track signals of target audio data based on the parameter values of the track signals;
[0020] Figure 3 It is a schematic diagram of various grouping methods of multiple exemplary track signals in an embodiment of the present application;
[0021] Figure 4 It is a schematic diagram of the structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Please refer to Figure 1 , the network framework in the embodiments of the present application includes:
[0023] A service server 100 and a terminal cluster; the terminal cluster may include terminal devices such as terminal device 200a, terminal device 200b, terminal device 200c,..., terminal device 200n.
[0024] Among them, the above-mentioned service server 100 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices (including terminal devices 200a, 200b, 200c, ……, 200n) can be smart phones, tablet computers, laptop computers, desktop computers, palmtop computers, mobile internet devices (MID), wearable devices (such as smart watches, smart bracelets, etc.), smart computers, smart vehicles and other intelligent terminals.
[0025] Among them, the service server 100 can establish communication connections with each terminal device in the terminal cluster, and communication connections can also be established among the terminal devices in the terminal cluster. In other words, the service server 100 can establish communication connections with each of the terminal devices 200a, 200b, 200c, ……, 200n. For example, a communication connection can be established between the terminal device 200a and the service server 100. A communication connection can be established between the terminal device 200a and the terminal device 200b, and a communication connection can also be established between the terminal device 200a and the terminal device 200c. Among them, the above-mentioned communication connection does not limit the connection method, and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, etc., which can be specifically determined according to the actual application scenario, and this application does not make any restrictions here.
[0026] It should be understood that, as Figure 1 shown, each terminal device in the terminal cluster can be installed with an application client. When the application client runs on each terminal device, data interaction can be carried out with the service server 100 respectively, so that the service server 100 can receive service data from each terminal device (such as user identity data uploaded by the user through the terminal device). Among them, the application client can be a music player application, a KTV software application, a browser application, a social application, an instant messaging application, a live broadcast application, a game application, a short video application, a video application, a shopping application, a novel application, a payment application and other application clients with functions of displaying data information such as text, images, audio and video, which can be specifically determined according to the actual application scenario requirements, and no restrictions are made here. Among them, the application client can be an independent client, or an embedded sub-client integrated in a certain client (such as a music application, a KTV software application, etc.), which can be specifically determined according to the actual application scenario, and no limitations are made here.
[0027] With the maturity of audio coding and decoding technologies and the increase in data transmission bandwidth, some audio processing methods and applications based on audio tracks have emerged. Especially in the fields of music production, film sound design, and virtual reality, there are many applications for audio track processing. For example, in the scenario of spatial audio rendering, by encoding different audio tracks into the Ambisonics format, sounds can be accurately positioned and moved in three-dimensional space. This method is applicable to 360-degree videos and virtual reality applications. Or, using multi-track technology, different audio track objects (such as vocals, musical instruments, environmental sound effects) are processed independently and precisely positioned in three-dimensional space to create Dolby Atmos. This technology is widely used in cinemas and high-end home theater systems.
[0028] Also, in the scenario of interactive audio processing, during game development, some audio processing middleware allows developers to divide audio files into multiple tracks and adjust the parameters of the tracks in real time according to the game state and player actions to achieve dynamic audio effects.
[0029] The above-mentioned audio track processing technologies have covered all aspects of music and entertainment, reflecting the application space of audio processing based on tracks.
[0030] However, due to the need for separate signal processing for each audio track, the above-mentioned audio processing methods based on tracks generally have the problem of requiring a large amount of computing. Especially for tracks at the master level, often dozens of tracks need to be processed simultaneously. For scenarios that require real-time rendering, this will bring a large computing burden and prevent the improvement of audio processing efficiency.
[0031] On the other hand, due to the strong diversity between different songs, most song tracks are different and require different processing means. Therefore, whether it is the aforementioned spatial audio rendering or other applications, most of the time, the processing parameters of audio track processing are set manually. This requires manual setting of processing parameters for each audio track, which will greatly increase the workload of manual labor and also delay the start and execution of audio processing, resulting in low audio processing efficiency.
[0032] To solve the above technical problems, the embodiments of this application propose an audio data processing method. By combining audio tracks with similar functions into a group (bus), the processing computing amount is significantly reduced while maintaining the quality of audio processing.
[0033] The following will be combined with Figure 1 the network framework and the application scenarios of the embodiments of this application to describe the audio data processing method in the embodiments of this application:
[0034] Please refer to Figure 2 , an embodiment of the audio data processing method in the embodiments of this application includes:
[0035] 201. Obtain N track signals of the target audio data; where N is a positive integer greater than 1;
[0036] The method of this embodiment can be applied to a computer device, which can be Figure 1 the service server 100 or each terminal device in the network framework shown. In some embodiments, the computer device can implement the audio data processing method provided in this embodiment by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can also be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module or plug-in.
[0037] The computer device can obtain N track signals of the target audio data to be processed. For example, these N track signals can include track signals of multiple tracks such as drum track, bass track, guitar sound track, piano sound track, vocal track, string track, etc. The target audio data can be any audio data with multiple track signals. For example, a music producer, as a user of a music platform, can separate multiple track signals of a target song on the music platform and process some of the track signals so that the processed track signals meet the user's listening requirements or meet the usage requirements of a specific scenario.
[0038] Among them, to separate multiple track signals from the target audio data, for example, a spectrum-based method can be used, such as using methods like short-time Fourier transform analysis, principal component analysis, independent component analysis, etc. to separate multiple track signals; or a spatial separation technique can be used to separate the tracks of the target audio data. This method is applicable to multi-channel audio (such as stereo), which uses the phase and amplitude differences between channels to separate multiple track signals.
[0039] Separating multiple track signals from the target audio data can be performed locally by the computer device, or other devices can separate multiple track signals from the target audio data and then send them to the computer device. Therefore, the method for the computer device to obtain multiple track signals of the target audio data is not limited.
[0040] 202. Extract the signal features of each track signal of the target audio data;
[0041] For each track signal of the target audio data, the computer device can extract the signal features of each track signal. The signal features can characterize the characteristics of the track signal. For example, they can be information such as the signal category and signal parameters of the track signal. For instance, the harmonic richness index, as one of the signal parameters, can be determined by analyzing the energy ratio of the fundamental frequency (f0) and its integer multiples of harmonics (2f0, 3f0, …) in the spectrum; the spectral flatness, as one of the signal parameters, can be calculated by framing the track signal, calculating the short-time Fourier transform, and calculating the spectral flatness based on the short-time Fourier transform result.
[0042] 203. Determine the common features of the signal features of at least two of the N track signals;
[0043] After extracting the signal features of each track signal, the common features of the signal features among multiple track signals in the N track signals of the target audio data can be determined. The so-called common features refer to that multiple track signals have the same or similar index results in one or some indicators. For example, it can be summarized that multiple track signals all belong to the same signal category, or it can be determined that the parameter values of a certain signal parameter of multiple track signals all fall within the same numerical range, and so on.
[0044] 204. Superimpose the at least two track signals including the common features to obtain a superimposed track signal;
[0045] 205. Perform signal processing on the superimposed track signal to obtain processed audio data;
[0046] The existence of common features in the signal features of multiple track signals indicates that the functions of these multiple track signals are similar and can be combined for processing during track signal processing. As Figure 2A shown, the 3 track signals shown on the left side of the arrow in the figure have common features, so these 3 track signals can be superimposed into one track signal to obtain the superimposed track signal pointed to by the arrow on the right side of the figure. This superimposed track signal can be subjected to signal processing to obtain processed audio data.
[0047] Among them, the processing of the superimposed track signal, for example, can be processing of spectral and spatial effects such as spectral repair, harmonic enhancement, surround sound processing, or signal processing such as sound effect processing, equalization control, reverberation, etc., to make the music work sound more professional and full, and more in line with the atmosphere of a specific scene. For example, the processed audio data can be applied to various scenarios such as game production, audio-visual production, or virtual reality fields, cinemas, and high-end home theater systems to enhance the listening experience in a specific scene.
[0048] For example, in related solutions, it is necessary to separately process the 8 track signals of the target audio data, which requires a huge amount of computation. However, through the method of this embodiment, assuming that these 8 track signals are divided into 2 functional groups, when processing the track signals, multiple track signals in each functional group are superimposed into one track signal for combined processing, which is equivalent to only processing 2 track signals. In contrast, the amount of computation for track processing can be greatly reduced, and the processing performance and efficiency of the audio can be improved.
[0049] Therefore, multiple track signals containing this common feature can be superimposed into the same functional signal to obtain a superimposed track signal, and the processed audio data can be directly obtained by processing this superimposed track signal. Compared with separately processing each track signal, the number of tracks to be processed is greatly reduced, thereby reducing the amount of computation for track signal processing and improving the efficiency of audio processing.
[0050] Moreover, by merging multiple sub-tracks into the same functional signal, the user only needs to set processing parameters for one functional signal, without separately setting processing parameters for multiple track signals of this functional signal, thus reducing the workload of manually setting processing parameters and enabling the audio signal processing to be quickly started and executed, so the efficiency of audio signal processing can also be improved.
[0051] When identifying the common features of track signals, the music information retrieval technology MIR (Music Information Retrieval) can be used to identify the signal features of track signals and the common features between multiple track signals. Among them, the music information retrieval technology MIR uses computational methods to understand and analyze the content and characteristics of digital music. Therefore, based on Figure 2 In the embodiment shown, in an optional implementation manner, the signal features of the track signals may include the signal parameter values of the track signals. The common features between multiple track signals can be determined according to these signal parameter values and grouped based on this common feature. Therefore, when determining the common features of the signal features of multiple track signals, it can be determined that the signal parameter values of multiple track signals among the N track signals of the target audio data all fall within a preset parameter range. Among them, this preset parameter value range corresponds to a preset functional signal. Furthermore, when grouping, multiple track signals whose signal parameter values fall within this preset parameter range can be divided into this preset functional signal.
[0052] The signal parameter value can be a parameter characterizing any feature of the audio track signal. For example, the signal parameter value can be the Harmonic Richness Factor (HRF). The Harmonic Richness Factor is used to characterize the distribution density and intensity of harmonic components in the signal, and can be judged by analyzing the energy proportion of the fundamental frequency (f0) and its integer multiple harmonics (2f0, 3f0, …) in the spectrum. According to the Harmonic Richness Factor, it can be identified whether the audio track signal is a melody signal, and thus a corresponding judgment interval can be set. For example, a threshold of HRF is set. Multiple audio track signals greater than the threshold can be determined as melody signals (such as piano tracks, string tracks, etc.), and can be divided into the same functional group (such as the melody group), and then superimposed into the same functional signal (melody signal); multiple audio track signals less than the threshold are determined as non-melody signals (such as percussion, noise, etc.), and can be divided into the same functional group (such as the non-melody group), and then superimposed into the same functional signal (non-melody signal).
[0053] Therefore, by using the signal parameter values of the audio track signals to determine the common features of multiple audio tracks and grouping based on the common features, the common functions among multiple audio track signals can be summarized from the perspective of the features of the audio track signals, making the grouping of multiple audio track signals more accurate, capable of dividing audio track signals with similar functions into the same group, and contributing to the later superposition and merging processing of multiple audio track signals in the same functional group.
[0054] Among them, when grouping multiple audio track signals according to the signal parameter value, an optional implementation manner is that multiple audio track signals can be divided into multiple different functional groups based on the common features of the signal parameter values of the multiple audio track signals. Specifically, it can be determined that the signal parameter values of some audio track signals among the N audio track signals of the target audio data all fall within the first numerical range, and the signal parameter values of some audio track signals among the N audio track signals of the target audio data all fall within the second numerical range. Among them, the first numerical range corresponds to the first shallow functional signal, and the second numerical range corresponds to the second shallow functional signal.
[0055] Furthermore, multiple audio track signals whose signal parameter values fall within the first numerical range can be divided into the first shallow functional signal; multiple audio track signals whose signal parameter values fall within the second numerical range can be divided into the second shallow functional signal.
[0056] For example, the target audio data may be song audio, the signal parameter value may be a harmonic richness index, the first shallow functional signal may include a melody signal, and the second shallow functional signal may include a harmony signal. A numerical range of the melody signal on the harmonic richness index may be set, and a numerical range of the harmony signal on the harmonic richness index may be set. Further, a plurality of track signals whose harmonic richness index falls within the numerical range corresponding to the melody signal may be divided into a melody group, and a plurality of track signals whose harmonic richness index falls within the numerical range corresponding to the harmony signal may be divided into a harmony group, thereby achieving the division of a plurality of track signals of the target audio data into a plurality of groups respectively.
[0057] Therefore, by dividing a plurality of track signals into a plurality of functional groups according to the signal parameter value, a plurality of track signals with similar characteristics or functions can be divided into the same functional group, and then superimposed into the same functional signal, making the grouping and superposition of the track signals more accurate, ensuring the rationality and accuracy of the grouping, facilitating the superposition and merging processing of a plurality of track signals within the same functional group, and capable of accelerating the speed of audio production and post-processing.
[0058] The signal parameter value may be the parameter value of a single signal parameter of the track signal, or may be a comprehensive index value obtained by inductive calculation of the parameter values of a plurality of signal parameters of the track signal. Therefore, in some other alternative embodiments, when extracting the signal features of each track signal, the parameter values of a plurality of signal parameters may be extracted for each track signal, and then the signal parameter comprehensive scores corresponding to the plurality of signal parameters may be calculated according to the parameter values of the plurality of signal parameters. The signal parameter comprehensive score may characterize the comprehensive characteristics of the track signal on the plurality of signal parameters.
[0059] Further, the functional commonality among a plurality of track signals may be determined according to the signal parameter comprehensive score, that is, it is determined that the signal parameter comprehensive scores of some of the N track signals of the target audio data all fall within the first score range, and the signal parameter comprehensive scores of some of the N track signals of the target audio data all fall within the second score range. Among them, the first score range corresponds to the first deep functional signal, and the second score range corresponds to the second deep functional signal. Therefore, a plurality of track signals whose signal parameter comprehensive scores fall within the first score range may be divided into the first deep functional signal; a plurality of track signals whose signal parameter comprehensive scores fall within the second score range may be divided into the second deep functional signal.
[0060] For example, the target audio data may be song audio. In song production, multiple tracks of a song can be divided into tracks of main functional components and tracks of secondary functional components according to their functions and importance. Therefore, in this embodiment, multiple track signals in the song audio can be divided into track signals that undertake the main functions of the song and track signals that undertake the secondary functions of the song. Therefore, a score range corresponding to the track signals that undertake the main functions of the song can be set. When it is determined that the comprehensive scores of the signal parameters of multiple track signals all fall within this score range, these multiple track signals are divided into main functional signals. A score range corresponding to the track signals that undertake the secondary functions of the song can also be set. When it is determined that the comprehensive scores of the signal parameters of multiple track signals all fall within this score range, these multiple track signals are divided into secondary functional signals.
[0061] Among them, the track signals in the song audio that undertake the main functions of the song are the core of the music, directly carrying the leading parts of the melody, rhythm or emotional expression. For example, they include the lead vocals (the core of pop music), and include the main instrumental sounds, such as piano solos, electric guitar main melodies (such as the guitar solos in rock music), violin main melodies, etc.
[0062] The track signals in the song audio that undertake the secondary functions of the song are used to supplement, decorate or enhance the main components, creating an atmosphere or enriching the details. For example, the track signals that undertake the secondary functions of the song may include: background choruses, such as the gospel choir; chorus overlays, such as the harmony tracks outside the lead vocals (such as "Ahh" or "Ooh"); environmental sound effects, such as rain sounds, wind sounds; percussion fills, such as tambourines, maracas, etc...
[0063] Therefore, by dividing multiple tracks of the target audio data into main functional groups and secondary functional groups and stacking them into main functional signals and secondary functional signals, more track signals can be divided into the same functional group and stacked into the same functional signal, further increasing the intensity of track grouping, making the tracks of the target audio data more concentrated, reducing the number of functional groups, and being more conducive to the combined processing of multiple track signals, further improving the audio processing efficiency.
[0064] In some alternative embodiments, multiple signal parameters of the audio track signal may include the autocorrelation function, spectral flatness, and zero crossing rate. Among them, the autocorrelation function (ACF) can be used to evaluate the similarity of a signal with itself at different time delays, and it helps to identify periodic and repetitive patterns. The spectral flatness (SF) is an index that measures the smoothness of the spectrum. A low spectral flatness means that the signal has significant peaks and flat regions in the frequency domain, while a high spectral flatness indicates that the signal is more uniform in the frequency domain. The zero crossing rate (ZCR) refers to the number of times the signal crosses zero within a unit time, and it can be used to evaluate the stationarity of the signal. Therefore, when calculating the comprehensive score of the signal parameters of the audio track signal, the weighted sum of multiple parameter values of the audio track signal in terms of the autocorrelation function, spectral flatness, and zero crossing rate can be calculated, and the obtained sum value is used as the comprehensive score of the signal parameters.
[0065] For example, based on the above signal parameters, namely the autocorrelation function (ACF), spectral flatness (SF), and zero crossing rate (ZCR), the comprehensive score of the signal parameters corresponding to the above multiple signal parameters can be calculated according to the following formula:
[0066] Score = w1 × ACF + w2 × SF + w3 × ZCR;
[0067] where w1, w2, w3 ≥ 0 are weights, representing the importance and influence degree of each signal parameter on the final score, and satisfying w1 + w2 + w3 = 1.
[0068] After calculating the comprehensive score of the signal parameters for each audio track signal, the multiple audio track signals can be divided into a main function group (main function) and a secondary function group (Sectiondary function) based on the following method:
[0069]
[0070] where T is a preset threshold.
[0071] Therefore, by calculating the comprehensive score of multiple signal parameters, the overall functional characteristics of the audio track signal can be characterized by the comprehensive score of the signal parameters, and the audio track signals can be grouped more accurately based on the comprehensive score of the signal parameters.
[0072] For example, when performing audio track superposition according to the signal parameters of the audio track signal, an alternative embodiment can be as Figure 2BAs shown, corresponding track layer stratification parameters can be extracted for each track signal of the target audio data. For example, multiple track signals can be divided into tracks such as drum tracks and bass tracks according to the stratification parameters; then, the multiple track signals can be merged according to the stratification parameters of the bus layer, where the bus refers to a mixing channel that merges multiple tracks (such as merging all rhythm tracks into a "rhythm bus" for unified processing); further, the track signals obtained by superimposing the bus layer can be merged according to the stratification parameters of the master layer, and the merged track signals can be mastered to globally process the multiple track signals and finally output the processed audio data.
[0073] In addition to grouping multiple track signals of the target audio data according to the signal parameter values of the track signals, the multiple track signals can also be grouped according to other signal characteristics of the track signals. Therefore, in some alternative embodiments, the signal characteristics of the track signals may further include the track category confidence of the track signals, which is used to represent the confidence of the track signal corresponding to the target track category. Furthermore, when determining the common characteristics among multiple track signals and performing track grouping based on the common characteristics, it can be determined that the track category confidences of multiple track signals among the N track signals of the target audio data all fall within the confidence interval corresponding to the target track category, and the multiple track signals whose track category confidences fall within this confidence interval are superimposed to obtain the functional signal corresponding to the target track category.
[0074] For example, the track category confidence can be obtained by a deep learning model identifying the input track signal and outputting an identification result, and the identification result may include the confidence that the input track signal is the target track category. The deep learning model can be trained based on deep learning algorithms for multiple track signals, such as an unsupervised learning model training method or a supervised learning model training method. During the model training process, the deep learning model can learn the characteristics of multiple track signals and establish a mapping relationship between the characteristics of the track signals and the track signal categories, so that when an input track signal to be identified is input, it can output a judgment result on whether the track signal is the target track category and give the corresponding confidence.
[0075] Taking melody recognition as an example below, for the track signal to be identified, melody detection can be performed through a deep neural network, and the confidence p of the melody detection is output. If the confidence > threshold, it is determined that the track signal to be identified is a melody, otherwise it is harmony or other categories. The formula can be expressed as follows:
[0076]
[0077] where p = f DNN(target track), where T is the threshold threshold.
[0078] After determining that the audio track signal to be recognized belongs to the corresponding category, it can be classified into the function group corresponding to that category, and grouping of multiple audio track signals can also be achieved.
[0079] Therefore, the deep learning-based method can provide more accurate audio track classification results, can achieve audio track grouping based on audio track categories, so that multiple audio track signals with the same audio track category are classified into the same function group and superimposed into the same function signal, which can improve the accuracy of audio track signal grouping.
[0080] In addition, in addition to grouping multiple audio track signals of the target audio data according to the signal parameter values of the audio track signals or the confidence of the audio track category, it is also possible to judge whether the audio track signal is a silent track or a non-silent track according to the energy characteristics of the audio track signal (such as total energy, average energy, root mean square energy, peak energy, etc.), and grouping of multiple audio tracks of the target audio data can also be achieved.
[0081] The following further illustrates the above-mentioned various optional implementation manners by way of examples. Taking an eight-track example, the target audio data can be separated to obtain Figure 3 8 audio track signals as shown, including Drum, Bass, Acousticguitar, Electricguitar, Piano, String, Vocals, and Others. Among them, these 8 audio track signals can determine common characteristics according to the signal parameter values of the audio track signals and be grouped according to the common characteristics. Such as Figure 3 The grouping method indicated by the first arrow divides multiple audio track signals into multiple function groups such as melody and harmony according to the harmonic richness index or other signal parameter values, realizing shallow grouping of audio track signals. For example, in the first layer of grouping, multiple audio track signals can be first divided into five categories based on function, namely Drum, Bass, Melody, Harmony, and others. Among them, Melody is an audio track with obvious melody characteristics in music, Harmony is an audio track with stronger functional harmony attributes than melody attributes in music, others are tracks with relatively unclear functional divisions, and Drum and Bass are the same as the original track divisions. When classified into the corresponding groups, multiple audio track signals within the group can be superimposed into the same function signal.
[0082] In addition, it is also possible to group multiple audio track signals in a more centralized manner, dividing more audio track signals into the same function group. Correspondingly, the number of function groups will be less. For example, asFigure 3 For the grouping method indicated by the second arrow, the comprehensive scores of the parameter values of multiple signal parameters of each track signal can be calculated, and the common features of the multiple track signals can be determined based on the comprehensive scores and grouped according to the common features. Then, the multiple track signals can be divided into the Mainfunction group or the Secondaryfunction group. Alternatively, the above first-level grouping can be further integrated. For example, the Melody, Harmony, and others tracks can be judged functionally. The multiple track signals in the Melody group can be divided into the Main function group, while the multiple track signals in the Harmony and others groups can be divided into the Secondary function group to achieve the second-level grouping. When divided into the corresponding groups, the multiple track signals within the group can be superimposed into the same functional signal.
[0083] In this embodiment, a track function grouping scheme based on split tracks is adopted, which can group a large number of split tracks. By combining tracks with similar functions into one group, the processing operation amount is significantly reduced, and the speed of audio production and post-processing is accelerated. At the same time, combining signal parameters and MIR music information retrieval technology for track grouping ensures the rationality and accuracy of the grouping, and improves the overall coordination and listening experience of the audio. This technology not only optimizes the music production process, but also provides technical support for subsequent audio processing solutions such as real-time audio processing and spatial audio rendering.
[0084] The computer device in the embodiment of the present application will be described below. Please refer to Figure 4 , an embodiment of the computer device in the embodiment of the present application includes:
[0085] The computer device 400 may include one or more central processing units (CPUs) 401 and a memory 405, and one or more application programs or data are stored in the memory 405.
[0086] Among them, the memory 405 may be volatile storage or persistent storage. The program stored in the memory 405 may include one or more modules, and each module may include a series of instruction operations on the computer device. Further, the central processing unit 401 may be set to communicate with the memory 405 and execute a series of instruction operations in the memory 405 on the computer device 400.
[0087] The computer device 400 may further include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or, one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0088] The central processing unit 401 may perform the operations executed by the computer device in the foregoing Figure 2 illustrated embodiments and their various alternative embodiments. Details are not elaborated herein.
[0089] An embodiment of the present application also provides a computer storage medium. In one embodiment, the computer storage medium stores instructions that, when executed on a computer, cause the computer to perform the operations executed by the computer device in the foregoing Figure 2 illustrated embodiments and their various alternative embodiments.
[0090] An embodiment of the present application also provides a computer program product. In one embodiment, when the computer program product runs on a computer device, it causes the computer device to perform the operations executed by the computer device in the foregoing Figure 2 illustrated embodiments and their various alternative embodiments.
[0091] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and details are not elaborated herein.
[0092] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0093] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0095] When the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
Claims
1. An audio data processing method, characterized in that: The method comprises: Obtain N audio track signals of target audio data; wherein N is a positive integer greater than 1; Extracting signal features of each audio track signal of the target audio data; Determine common features of signal features of at least two audio track signals among the N audio track signals; Superimposing the at least two audio track signals containing the common feature to obtain a superimposed audio track signal; The superimposed audio track signal is subjected to signal processing to obtain processed audio data.
2. The method according to claim 1, characterized in that The signal characteristics of the audio track signal include signal parameter values of the audio track signal; The determining of the common features of the signal features of at least two audio track signals among the N audio track signals comprises: Determine that the signal parameter values of at least two of the N track signals fall within a preset parameter range; wherein the preset parameter value range corresponds to a preset function signal; The step of superimposing the at least two audio track signals containing the common feature to obtain a superimposed audio track signal includes: The at least two audio track signals whose signal parameter values fall within the preset parameter range are superimposed to obtain the preset function signal.
3. The method according to claim 2, characterized in that The determining that the signal parameter values of at least two of the N track signals fall within a preset parameter range includes: Determine that the signal parameter values of some of the N track signals all fall within a first value range, and the signal parameter values of some of the N track signals all fall within a second value range; wherein the first value range corresponds to a first shallow function signal, and the second value range corresponds to a second shallow function signal; The step of superimposing the at least two audio track signals whose signal parameter values fall within the preset parameter range to obtain the preset function signal includes: Superimposing a plurality of audio track signals whose signal parameter values fall within the first numerical range to obtain the first shallow function signal; Multiple audio track signals whose signal parameter values fall within the second numerical range are superimposed to obtain the second shallow functional signal.
4. The method according to claim 3, characterized in that The target audio data is song audio; the signal parameter value includes a harmonic richness index; The first shallow function signal includes a melody signal, and the second shallow function signal includes a harmony signal.
5. The method according to claim 2, characterized in that: The signal parameter value is a signal parameter comprehensive score calculated from parameter values of multiple signal parameters of the audio track signal; The determining that the signal parameter values of at least two of the N track signals fall within a preset parameter range includes: Determine that the signal parameter comprehensive scores of some of the N track signals all fall within a first score range, and that the signal parameter comprehensive scores of some of the N track signals all fall within a second score range; Wherein, the first score range corresponds to a first deep function signal, and the second score range corresponds to a second deep function signal; The step of superimposing the at least two audio track signals whose signal parameter values fall within the preset parameter range to obtain the superimposed audio track signal includes: Superimposing a plurality of audio track signals whose signal parameter comprehensive scores fall within the first score range to obtain the first deep function signal; Multiple audio track signals whose comprehensive scores of the signal parameters fall within the second score range are superimposed to obtain the second deep function signal.
6. The method according to claim 5, characterized in that The target audio data is song audio; The first deep function signal comprises a primary function signal, and the second deep function signal comprises a secondary function signal; Among them, the main function signal is the audio track signal in the song audio that bears the main function of the song, and the secondary function signal is the audio track signal in the song audio that bears the secondary function of the song.
7. The method according to claim 5, characterized in that The plurality of signal parameters include autocorrelation function, spectral flatness and zero crossing rate; The steps of calculating the comprehensive score of the signal parameters include: For each audio track signal of the target audio data, a weighted sum of multiple parameter values of the audio track signal in terms of autocorrelation function, spectral flatness and zero crossing rate is calculated, and the obtained sum value is used as the comprehensive score of the signal parameters.
8. The method according to claim 1, characterized in that The signal feature of the audio track signal includes an audio track category confidence of the audio track signal, wherein the audio track category confidence is used to indicate the confidence that the audio track signal corresponds to the target audio track category; The determining of the common features of the signal features of at least two audio track signals among the N audio track signals comprises: Determine that the track category confidences of at least two of the N track signals fall within the confidence interval corresponding to the target track category; The step of superimposing the at least two audio track signals containing the common feature to obtain a superimposed audio track signal includes: The at least two audio track signals whose audio track category confidences fall within the confidence interval are superimposed to obtain a functional signal corresponding to the target audio track category.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer storage medium, characterized in that: The computer storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to perform the method according to any one of claims 1 to 8.