Music master tape recovery model processing method, music master tape recovery method and equipment

Through the combination of the subband model and the U-net model, the problem of poor master restoration quality in the prior art is solved, and effective generation of high-frequency spectrum and high-quality restoration of master signals are achieved.

CN120279955APending Publication Date: 2025-07-08TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318433.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing music mastering model is difficult to generate high-frequency spectrum from low frequency spectrum, especially the 44.1kHz high frequency spectrum from 192kHz, resulting in poor mastering quality.

Method used

The combination of subband model and U-net model is adopted to generate high-frequency spectrum through band splitting, band sequence modeling and reconstruction, and the model is trained using time domain and frequency domain loss values to improve the restoration quality.

Benefits of technology

Through the fusion of the subband model and the U-net model, high-quality restoration of the music master signal is achieved, and the frequency and time domain accuracy of the master signal is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279955A_ABST
    Figure CN120279955A_ABST
Patent Text Reader

Abstract

The invention relates to a music master tape recovery model processing method, computer equipment and a storage medium. The method comprises the following steps: acquiring a sample music master band signal and a first sample input signal associated with the sample music master band signal; the first sample input signal is a music signal obtained by removing a high-frequency spectrum from a sample music master band signal; inputting the first sample input signal into a sub-band model included in a music master band restoration model to be trained to obtain a sub-band signal corresponding to the first sample input signal, and generating a second sample input signal by using the sub-band signal; superposing the first sample input signal and the second sample input signal, and inputting a U-net model included in a music master band restoration model to obtain a predicted music master band signal; and training a music master band recovery model by using the difference between the predicted music master band signal and the sample music master band signal to obtain a trained music master band recovery model. By adopting the method, the quality of the recovered master band signal can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to a method for processing a music master restoration model, a music master restoration method, a computer device, a storage medium, and a computer program product. Background Art

[0002] With the development of audio processing technology, a technology for restoring music masters has emerged. A music master refers to a file that has not undergone any lossy compression at the initial stage of song production, usually in the format of 192kHz 24bit. Through the master restoration technology, a compressed music signal can be restored to a master audio signal. For example, a music signal in the format of 44.1kHz 16bit can be restored to a master signal in the format of 192kHz 24bit, thereby better improving the sound quality and restoring music details.

[0003] In traditional technologies, the way to restore a music master signal can be achieved through neural network models, such as generative network models like GAN, Diffusion, etc. However, the above-mentioned generative network models are usually only applicable to generating high-frequency spectra from low-frequency spectra. For example, they are only applicable to generating a 44.1kHz high-frequency spectrum from a 20kHz spectrum, but it is difficult to generate an extremely high-frequency spectrum of 192kHz based on a 44.1kHz high-frequency spectrum, and it is difficult to meet the requirements of master restoration. Therefore, the quality of the existing music master restoration model for restoring master signals is poor. Summary of the Invention

[0004] Based on this, it is necessary to provide a method for processing a music master restoration model, a music master restoration method, a device, a computer device, a computer-readable storage medium, and a computer program product that can improve the quality of the music master restoration model for restoring master signals in view of the above technical problems.

[0005] In a first aspect, the present application provides a method for processing a music master restoration model, including:

[0006] Obtain a sample music master signal and a first sample input signal associated with the sample music master signal; the first sample input signal is a music signal obtained by removing the high-frequency spectrum from the sample music master signal;

[0007] Input the first sample input signal into a sub-band model included in the music master restoration model to be trained, obtain a sub-band signal corresponding to the first sample input signal, and generate a second sample input signal by using the sub-band signal; the second sample input signal is a music signal restored by using the sub-band signal;

[0008] Superimpose the first sample input signal and the second sample input signal, and input them into the U-net model included in the music master restoration model to obtain a predicted music master signal;

[0009] Utilize the difference between the predicted music master signal and the sample music master signal to train the music master restoration model to obtain a trained music master restoration model.

[0010] In one embodiment, the sub-band model includes: a frequency band splitting module; the step of inputting the first sample input signal into the sub-band model included in the music master restoration model to be trained to obtain the sub-band signal corresponding to the first sample input signal includes: inputting the first sample input signal into the sub-band model, performing Fourier transform on the first sample input signal through the sub-band model to obtain the first sample frequency domain signal corresponding to the first sample input signal; inputting the first sample frequency domain signal into the frequency band splitting module, and dividing the first sample frequency domain signal into sub-bands according to frequency points through the frequency band splitting module to obtain the sub-band signal.

[0011] In one embodiment, the sub-band model further includes: a frequency band sequence modeling module and a frequency band reconstruction module; the step of generating the second sample input signal by using the sub-band signal includes: inputting the sub-band signal into the frequency band sequence modeling module, performing bidirectional modeling processing on the sub-band signal through the frequency band sequence modeling module to obtain the sub-band signal after bidirectional modeling; inputting the sub-band signal after bidirectional modeling into the frequency band reconstruction module, generating the second sample frequency domain signal through the frequency band reconstruction module; and performing inverse Fourier transform on the second sample frequency domain signal to obtain the second sample input signal.

[0012] In one embodiment, the step of training the music master restoration model by using the difference between the predicted music master signal and the sample music master signal includes: obtaining a first loss value according to the time domain difference between the predicted music master signal and the sample music master signal, and obtaining a second loss value according to the frequency domain difference between the predicted music master signal and the sample music master signal; obtaining the total loss value of the music master restoration model according to the first loss value, the second loss value and a preset proportional factor between the time domain loss and the frequency domain loss, and training the music master restoration model by using the total loss value.

[0013] In one embodiment, obtaining the second loss value according to the frequency domain difference between the predicted music master signal and the sample music master signal includes: obtaining a plurality of preset resolution parameters, and performing short-time Fourier transform on the predicted music master signal and the sample music master signal by using the plurality of resolution parameters to obtain a predicted master frequency domain signal and a sample master frequency domain signal corresponding to each of the resolution parameters; for each current resolution parameter, obtaining a current second loss value corresponding to the current resolution parameter according to the difference between the current predicted master frequency domain signal corresponding to the current resolution parameter and the current sample master frequency domain signal corresponding to the current resolution parameter; the current resolution parameter being any one of the resolution parameters; and obtaining the second loss value according to each of the current second loss values.

[0014] In one embodiment, obtaining the current second loss value corresponding to the current resolution parameter according to the difference between the current predicted master frequency domain signal corresponding to the current resolution parameter and the current sample master frequency domain signal corresponding to the current resolution parameter includes: obtaining a first sub-loss value according to the complex spectrum difference between the current predicted master frequency domain signal and the current sample master frequency domain signal, and obtaining a second sub-loss value according to the spectral amplitude difference between the current predicted master frequency domain signal and the current sample master frequency domain signal; and obtaining the current second loss value according to the first sub-loss value and the second sub-loss value.

[0015] In one embodiment, the step of obtaining the first sample input signal includes: performing downsampling processing on the sample music master signal to obtain a downsampled signal; where the sampling frequency of the sample music master signal is a first sampling frequency, the sampling frequency of the downsampled signal is a second sampling frequency, and the first sampling frequency is higher than the second sampling frequency; and performing upsampling processing on the downsampled signal to obtain a first sample input signal with a sampling frequency of the first sampling frequency.

[0016] Second, the present application also provides a method for restoring a music master, including:

[0017] Obtaining a target music signal to be restored to a music master;

[0018] Inputting the target music signal into the trained music master restoration model, and generating a music master signal corresponding to the target music signal through the music master restoration model; the music master restoration model is obtained by processing through the music master restoration model processing method according to any one of the embodiments in the first aspect.

[0019] Third, the present application also provides a music master restoration model processing device, including:

[0020] A first signal acquisition module, configured to acquire a sample music master tape signal and a first sample input signal associated with the sample music master tape signal; the first sample input signal is a music signal obtained by removing the high-frequency spectrum from the sample music master tape signal.

[0021] A second signal acquisition module, configured to input the first sample input signal into a sub-band model included in a music master tape restoration model to be trained, obtain a sub-band signal corresponding to the first sample input signal, and generate a second sample input signal by using the sub-band signal; the second sample input signal is a music signal restored by using the sub-band signal.

[0022] A predicted master tape acquisition module, configured to superimpose the first sample input signal and the second sample input signal, input the superimposed signals into a U-net model included in the music master tape restoration model, and obtain a predicted music master tape signal.

[0023] A restoration model training module, configured to train the music master tape restoration model by using the difference between the predicted music master tape signal and the sample music master tape signal, so as to obtain a trained music master tape restoration model.

[0024] In a fourth aspect, the present application further provides a music master tape restoration device, including:

[0025] A target music acquisition module, configured to acquire a target music signal for which music master tape restoration is to be performed.

[0026] A master tape signal restoration module, configured to input the target music signal into the trained music master tape restoration model, and generate a music master tape signal corresponding to the target music signal through the music master tape restoration model; the music master tape restoration model is obtained by processing according to the music master tape restoration model processing method described in any embodiment of the first aspect.

[0027] In a fifth aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any embodiment of the first aspect or the second aspect are implemented.

[0028] In a sixth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any embodiment of the first aspect or the second aspect are implemented.

[0029] In a seventh aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in any embodiment of the first aspect or the second aspect are implemented.

[0030] The above-mentioned method for processing a music master tape restoration model, music master tape restoration method, device, computer device, storage medium, and computer program product obtain a sample music master tape signal and a first sample input signal associated with the sample music master tape signal; the first sample input signal is a music signal obtained by removing the high-frequency spectrum from the sample music master tape signal; input the first sample input signal into a sub-band model included in the music master tape restoration model to be trained to obtain a sub-band signal corresponding to the first sample input signal, and generate a second sample input signal using the sub-band signal; the second sample input signal is a music signal restored using the sub-band signal; superimpose the first sample input signal and the second sample input signal, input them into the U-net model included in the music master tape restoration model to obtain a predicted music master tape signal; use the difference between the predicted music master tape signal and the sample music master tape signal to train the music master tape restoration model to obtain a trained music master tape restoration model. In this embodiment, when training the music master tape restoration model, a sample music master tape signal and a first sample input signal obtained by removing the high-frequency spectrum from the sample music master tape signal can be collected, the first sample input signal is input into the sub-band model included in the music master tape restoration model, the sub-band model outputs a sub-band signal and generates a corresponding second sample input signal, and then the first sample input signal and the second sample input signal can be superimposed and input into the U-net model in the music master tape restoration model to obtain a predicted music master tape signal, so as to train the sample music master tape signal according to the difference between the predicted music master tape signal and the sample music master tape signal. Compared with using a GAN model to restore the master tape, the music master tape restoration model trained in this application can restore the master tape based on the time-frequency domain structure fused by the frequency-domain sub-band model and the time-domain U-net model, thereby improving the quality of the restored master tape signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a schematic flowchart of a method for processing a music master tape restoration model in an embodiment;

[0033] Figure 2 It is a schematic flowchart of obtaining a sub-band signal in an embodiment;

[0034] Figure 3 It is a schematic flowchart of obtaining a second sample input signal in an embodiment;

[0035] Figure 4 Schematic diagram of the process for obtaining the second loss value in an embodiment;

[0036] Figure 5 Schematic diagram of the process for the music master tape restoration method in an embodiment;

[0037] Figure 6 Schematic diagram of the training and inference process for the music master tape restoration in an embodiment;

[0038] Figure 7 Schematic diagram of the model input signal in an embodiment;

[0039] Figure 8 Schematic diagram of the real master tape signal in an embodiment;

[0040] Figure 9 Schematic diagram of the structure of the sub-band model Apollo in an embodiment;

[0041] Figure 10 Schematic diagram of the structure of the time-domain U-net model in an embodiment;

[0042] Figure 11 Schematic block diagram of the structure of the music master tape restoration model processing device in an embodiment;

[0043] Figure 12 Schematic block diagram of the structure of the music master tape restoration device in an embodiment;

[0044] Figure 13 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0045] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0046] In one embodiment, as Figure 1 shown, a music master tape restoration model processing method is provided. In this embodiment, it is exemplified that this method is applied to a server. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0047] Step S101, obtaining a sample music master tape signal and a first sample input signal associated with the sample music master tape signal; the first sample input signal is a music signal obtained by removing the high-frequency spectrum from the sample music master tape signal.

[0048] Among them, the sample music master tape signal can be the real master tape signal corresponding to a certain music signal, and the first sample input signal is the music signal obtained by removing the high-frequency spectrum from the sample music master tape signal. Since master tape restoration usually needs to restore a 44.1kHz standard audio signal to a 192kHz master tape signal, the sample music master tape signal used can be a real 192kHz master tape signal, and the first sample input signal is the music signal obtained by removing the high-frequency spectrum part that needs to be restored from the real master tape signal.

[0049] Specifically, when the server trains the music master tape restoration model, it first needs to obtain the real music master tape signal, that is, the sample music master tape signal as the learning target, and obtain the music signal obtained by removing the high-frequency spectrum from the sample music master tape signal, that is, the first sample input signal as the input signal of the model.

[0050] Step S102: Input the first sample input signal into the sub-band model included in the music master tape restoration model to be trained, obtain the sub-band signal corresponding to the first sample input signal, and generate a second sample input signal using the sub-band signal; the second sample input signal is the music signal restored using the sub-band signal.

[0051] The music master tape restoration model to be trained refers to the music master tape restoration model that needs to be trained. This model can be composed of two parts, namely the sub-band model and the U-net model. Among them, the sub-band model is mainly used to predict the missing high frequency, and the U-net model is mainly used to smooth the original low-frequency part and the predicted high-frequency part. The sub-band signal refers to the sub-band signal obtained by the sub-band model according to the input first sample input signal, and the second sample input signal is the music signal restored using the sub-band signal, and this signal can be used as part of the input signal of the U-net model.

[0052] Specifically, when training the music master tape restoration model, the first sample input signal associated with the sample music master tape signal can be input into the sub-band model included in the music master tape restoration model to be trained, and the corresponding sub-band signal is output by the sub-band model. Then, the sub-band model can further use the sub-band signal to generate the second sample input signal.

[0053] Step S103: Superimpose the first sample input signal and the second sample input signal, and input them into the U-net model included in the music master tape restoration model to obtain the predicted music master tape signal.

[0054] The predicted master music signal refers to the master music signal finally output by the master music restoration model. After obtaining the second sample input signal, the first sample input signal and the second sample input signal can be superimposed and processed, and the superimposed signal is input into the U-net model included in the master music restoration model to output the predicted master music signal.

[0055] Step S104: Use the difference between the predicted master music signal and the sample master music signal to train the master music restoration model to obtain a trained master music restoration model.

[0056] The trained master music restoration model refers to the neural network model that finally completes training for realizing master music restoration. After obtaining the predicted master music signal, the difference between the predicted master music signal and the sample master music signal can be used to obtain the model loss of the master music restoration model, and then the master music restoration model can be trained using this loss. It can be to fine-tune the sub-band model and the U-net model of the master music restoration model using this loss, so as to obtain a trained master music restoration model.

[0057] In the above master music restoration model processing method, when training the master music restoration model, the sample master music signal and the first sample input signal obtained by removing the high-frequency spectrum from the sample master music signal can be collected. The first sample input signal is input into the sub-band model included in the master music restoration model. The sub-band model outputs a sub-band signal and generates a corresponding second sample input signal. Then, the first sample input signal and the second sample input signal can be superimposed and input into the U-net model in the master music restoration model to obtain the predicted master music signal, and then the sample master music signal can be trained according to the difference between the predicted master music signal and the sample master music signal. Compared with using the GAN model to restore the master, the master music restoration model trained in this application can restore the master based on the time-frequency domain structure that combines the frequency domain model (sub-band model) and the time domain model (U-net model), thereby improving the quality of the restored master signal.

[0058] In one embodiment, the sub-band model includes: a frequency band splitting module; as Figure 2 shown, step S102 can further include:

[0059] Step S201: Input the first sample input signal into the sub-band model, and perform Fourier transform on the first sample input signal through the sub-band model to obtain the first sample frequency domain signal corresponding to the first sample input signal.

[0060] The first sample frequency-domain signal refers to the frequency-domain signal corresponding to the first sample input signal. Since the first sample input signal is a time-domain signal, and in order to generate sub-band signals corresponding to different frequency bands, it is necessary to first convert the time-domain signal into a frequency-domain signal, which can be achieved through Fourier transform. Specifically, the server can input the first sample input signal into the sub-band model, and the sub-band model first performs Fourier transform processing on the first sample input signal to obtain the first sample frequency-domain signal.

[0061] Step S202: Input the first sample frequency-domain signal into the frequency band splitting module, and obtain sub-band signals through the frequency band splitting module.

[0062] The frequency band splitting module is a module in the sub-band model used to separate the frequency range, that is, to separate sub-bands. After obtaining the first sample frequency-domain signal, the first sample frequency-domain signal can be input into the frequency band splitting module in the sub-band model, and the frequency band splitting module divides the frequency-domain signal into sub-bands according to frequency points, thereby obtaining sub-band signals.

[0063] In this embodiment, the server can input the first sample input signal into the sub-band model included in the music master restoration model. The sub-band model first converts the first sample input signal into a frequency-domain signal, and then the frequency band splitting module included in the sub-band model is used to obtain sub-band signals. By this method, the efficiency of obtaining sub-band signals can be improved.

[0064] In addition, the sub-band model further includes: a frequency band sequence modeling module and a frequency band reconstruction module; as Figure 3 shown, step S102 can further include:

[0065] Step S301: Input the sub-band signals into the frequency band sequence modeling module, and perform bidirectional modeling processing on the sub-band signals through the frequency band sequence modeling module to obtain the sub-band signals after bidirectional modeling.

[0066] The frequency band sequence modeling module is a module used to perform bidirectional modeling processing on sub-band signals. Bidirectional modeling can include sequence modeling in the frequency domain dimension and sequence modeling in the time domain dimension. Among them, sequence modeling in the frequency domain dimension can be implemented using Ropeformer, and sequence modeling in the time domain dimension is implemented through the TCN model.

[0067] Specifically, after obtaining sub-band signals through the frequency band splitting module, the sub-band signals can also be input into the frequency band sequence modeling module. First, the Ropeformer in the frequency band sequence modeling module is used to complete sequence modeling in the frequency domain dimension, and then the sequence modeling result in the frequency domain dimension is input into the TCN model to implement sequence modeling in the time domain dimension, thereby completing bidirectional modeling of the sub-band signals and obtaining the sub-band signals after bidirectional modeling.

[0068] Step S302: Input the sub-band signals after bidirectional modeling into the frequency band reconstruction module, and generate the second sample frequency domain signal through the frequency band reconstruction module.

[0069] Step S303: Perform inverse Fourier transform on the second sample frequency domain signal to obtain the second sample input signal.

[0070] The frequency band reconstruction module is a module used to merge and restore sub-band signals. After completing the bidirectional modeling of sub-band signals, the sub-band signals can be merged and restored, and the restored frequency domain signal can be used as the second sample frequency domain signal. The server can input the modeled sub-band signals into the frequency band reconstruction module included in the sub-band model to obtain the second sample frequency domain signal, and can also perform inverse Fourier transform on the second sample frequency domain signal to restore it to the time domain signal form, thereby obtaining the second sample input signal.

[0071] In this embodiment, after obtaining the sub-band signals, the frequency band sequence modeling module can also be used to perform bidirectional modeling on the sub-band signals, then merge and restore the bidirectionally modeled sub-band signals to obtain the second sample frequency domain signal and perform inverse Fourier transform to obtain the second sample input signal. Through this method, the smooth high-frequency part can be better predicted, thereby further improving the accuracy of restoring high-frequency signals through the sub-band model.

[0072] In one embodiment, step S104 can further include: obtaining a first loss value according to the time domain difference between the predicted music master signal and the sample music master signal, and obtaining a second loss value according to the frequency domain difference between the predicted music master signal and the sample music master signal; obtaining the total loss value of the music master restoration model according to the first loss value, the second loss value, and the preset proportional factor between the time domain loss and the frequency domain loss, and training the music master restoration model using the total loss value.

[0073] The first loss value refers to the loss value obtained according to the time domain difference between the predicted music master signal and the sample music master signal, that is, the time domain loss, and the second loss value is the loss value obtained according to the frequency domain difference between the predicted music master signal and the sample music master signal, that is, the frequency domain loss. The total loss value is the total loss value of the music master restoration model. In this embodiment, the loss of the model consists of two parts, namely the time domain loss and the frequency domain loss, and the proportional factor is a weighting factor used to perform weighted processing on the time domain loss and the frequency domain loss.

[0074] Specifically, after obtaining the predicted music master tape signal, the server can calculate the first loss value based on the time-domain difference between the predicted music master tape signal and the sample music master tape signal, or calculate the second loss value based on the frequency-domain difference between the predicted music master tape signal and the sample music master tape signal. Then, according to the first loss value, the second loss value, and the scaling factor, the total loss value of the music master tape restoration model is calculated, and the training of the music master tape restoration model is completed using this total loss value.

[0075] For example, the total loss value of the model can be calculated by the following formula:

[0076]

[0077] where, represents the time-domain loss, that is, the first loss value, represents the frequency-domain loss, that is, the second loss value, and is the scaling factor of the time-domain loss and the frequency-domain loss, and J is 1.

[0078] And, the first loss value can be calculated by the following formula:

[0079]

[0080] where, represents the sample music master tape signal, represents the predicted music master tape signal, represents the first-order norm.

[0081] In this embodiment, the loss value for training the music master tape restoration model can include two parts, namely, the first loss value corresponding to the time-domain difference between the predicted music master tape signal and the sample music master tape signal, and the second loss value corresponding to the frequency-domain difference between the predicted music master tape signal and the sample music master tape signal. The music master tape restoration model trained in this way can further improve the accuracy of master tape restoration.

[0082] Furthermore, as Figure 4 shown, obtaining the second loss value according to the frequency-domain difference between the predicted music master tape signal and the sample music master tape signal can further include:

[0083] Step S401, obtain a plurality of preset resolution parameters, and perform short-time Fourier transform on the predicted music master tape signal and the sample music master tape signal using the plurality of resolution parameters to obtain the predicted master tape frequency-domain signal and the sample master tape frequency-domain signal corresponding to each resolution parameter respectively.

[0084] In this embodiment, the frequency-domain difference may refer to the difference between the predicted master music signal and the sample master music signal after performing short-time Fourier transform (STFT), that is, the difference between the predicted master frequency-domain signal and the sample master frequency-domain signal obtained after STFT. The multiple resolution parameters are the resolution parameters for STFT. For example, the window lengths of STFT can be set to 4096, 2048, and 1024 respectively, and the frame shifts are 1024, 512, and 256 respectively. In this way, 9 resolution parameters can be obtained.

[0085] Specifically, after the server obtains the preset multiple resolution parameters, it can use each resolution parameter to perform short-time Fourier transform on the predicted master music signal and the sample master music signal respectively, so as to obtain the predicted master frequency-domain signal and the sample master frequency-domain signal corresponding to each resolution parameter.

[0086] For example, the resolution parameters include resolution parameter a, resolution parameter b, and resolution parameter c. The server can use resolution parameter a, resolution parameter b, and resolution parameter c to perform short-time Fourier transform on the predicted master music signal and the sample master music signal respectively, so as to obtain predicted master frequency-domain signal a, predicted master frequency-domain signal b, and predicted master frequency-domain signal c, and also obtain sample master frequency-domain signal a, sample master frequency-domain signal b, and sample master frequency-domain signal c.

[0087] Step S402: For each current resolution parameter, obtain the current second loss value according to the difference between the current predicted master frequency-domain signal corresponding to the current resolution parameter and the current sample master frequency-domain signal corresponding to the current resolution parameter; the current resolution parameter is any one of the resolution parameters.

[0088] Step S403: Obtain the second loss value according to each current second loss value.

[0089] The current resolution parameter refers to any one of the resolution parameters. The current predicted master frequency-domain signal refers to the predicted master frequency-domain signal corresponding to the current resolution parameter. The current sample master frequency-domain signal is the sample master frequency-domain signal corresponding to the current resolution parameter. The current second loss value is the second loss value corresponding to the current resolution parameter. In this embodiment, for each resolution parameter, a loss value can be calculated respectively. Then, the second loss value representing the frequency-domain difference can be obtained by comprehensively considering the loss values obtained for each resolution parameter. For example, the sum of each current second loss value can be used as the final second loss value, or the average of each current second loss value can be used as the final second loss value.

[0090] For example, the resolution parameters include resolution parameter a, resolution parameter b, and resolution parameter c. The server can obtain the second loss value a corresponding to resolution parameter a based on the difference between the predicted master band frequency domain signal a and the sample master band frequency domain signal a. It can also obtain the second loss value b corresponding to resolution parameter b based on the difference between the predicted master band frequency domain signal b and the sample master band frequency domain signal b, and obtain the second loss value c corresponding to resolution parameter c based on the difference between the predicted master band frequency domain signal c and the sample master band frequency domain signal c. Thus, by combining the second loss value a, the second loss value b, and the second loss value c, the final second loss value can be obtained.

[0091] For example, the second loss value can be calculated through the following formula:

[0092]

[0093] where represents the current second loss value corresponding to a certain current resolution parameter, and M represents the number of different resolution parameters.

[0094] In this embodiment, the second loss value can be calculated from the respective current second loss values corresponding to different resolution parameters. In this way, the accuracy of calculating the second loss value can be further improved.

[0095] Further, step S402 can further include: obtaining a first sub-loss value based on the complex spectrum difference between the current predicted master band frequency domain signal and the current sample master band frequency domain signal, and obtaining a second sub-loss value based on the spectral amplitude difference between the current predicted master band frequency domain signal and the current sample master band frequency domain signal; obtaining the current second loss value based on the first sub-loss value and the second sub-loss value.

[0096] In this embodiment, each current second loss value can also be composed of two parts, namely the first sub-loss value and the second sub-loss value. The first sub-loss value corresponds to the complex spectrum difference between the current predicted master band frequency domain signal and the current sample master band frequency domain signal, while the second sub-loss value corresponds to the spectral amplitude difference between the current predicted master band frequency domain signal and the current sample master band frequency domain signal.

[0097] For example, the current second loss value can be represented by the following formula:

[0098]

[0099] where represents the first sub-loss value, and represents the second sub-loss value.

[0100] And, the first sub-loss value can be calculated by the following formula:

[0101]

[0102] While the second sub-loss value can be calculated by the following formula:

[0103]

[0104] Wherein, represents the current sample master tape frequency domain signal, represents the current predicted master tape frequency domain signal, represents the 2-norm, while represents the 1-norm.

[0105] In this embodiment, the complex spectral difference between the current predicted master tape frequency domain signal and the current sample master tape frequency domain signal can also be used as the first sub-loss value, and the spectral amplitude difference between the current predicted master tape frequency domain signal and the current sample master tape frequency domain signal can be used as the second sub-loss value. Thus, the first sub-loss value and the second sub-loss value are used to obtain the current second loss value, and in this way, the accuracy of obtaining the current second loss value can be improved.

[0106] In one embodiment, step S101 may further include: performing downsampling on the sample music master tape signal to obtain a downsampled signal; wherein the sampling frequency of the sample music master tape signal is the first sampling frequency, and the sampling frequency of the downsampled signal is the second sampling frequency, and the first sampling frequency is higher than the second sampling frequency; performing upsampling on the downsampled signal to obtain a first sample input signal with a sampling frequency of the first sampling frequency.

[0107] In this embodiment, the sampling frequency of the first sample input signal may be the same as the sampling frequency of the sample music master tape signal, both being the first sampling frequency. And the way to remove the high-frequency spectrum may be to first perform downsampling on the sample music master tape signal to obtain a downsampled signal corresponding to the second sampling frequency, and then perform upsampling on the downsampled signal to obtain a first sample input signal with a sampling frequency of the first sampling frequency.

[0108] For example, the first sampling frequency may be 192 kHz, and the second sampling frequency may be 44.1 kHz. In this embodiment, the 192 kHz sample music master tape signal can be first downsampled to 44.1 kHz to obtain a downsampled signal, and then the signal can be upsampled to 192 kHz to obtain the first sample input signal, and in this way, the high-frequency components in the sample music master tape signal are removed.

[0109] In this embodiment, the first sample input signal can be obtained by first downsampling the first sample input signal and then upsampling it. By this method, the high-frequency components in the sample music master signal can be removed, enabling the music master restoration model to more easily learn to restore the high-frequency components, that is, the process of restoring the master tape, thereby further improving the accuracy of the music master restoration model in restoring the master tape.

[0110] In one embodiment, as Figure 5 shown, a music master restoration method is also provided. In this embodiment, taking the application of this method to a server as an example, it can be understood that this method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0111] Step S501: Obtain a target music signal to be restored to its music master.

[0112] The target music signal refers to the music signal that needs to be restored to its music master. For example, it can be a signal obtained by upsampling the music signal in an audio file of a certain standard quality. In this embodiment, when performing music master restoration, the server can first obtain the target music signal that needs to be restored to its music master.

[0113] Step S502: Input the target music signal into the trained music master restoration model to generate a music master signal corresponding to the target music signal through the music master restoration model. The music master restoration model is obtained by processing through the music master restoration model processing method described in any of the above embodiments.

[0114] The trained music master restoration model refers to the music master restoration model that has completed training. This model can be obtained by processing through the music master restoration model processing method provided in the previous embodiments. The server can input the target music signal into the trained music master restoration model, thereby generating a music master signal corresponding to the target music signal through the music master restoration model. For example, the target music signal can be first input as an input signal into the sub-band model in the music master restoration model. After the sub-band model restores the high-frequency features of the target music signal to obtain an output signal, the output signal of the sub-band model and the target music signal can be superimposed and input into the U-net model in the music master restoration model, so that the U-net model outputs the music master signal corresponding to the target music signal.

[0115] In the above music master tape restoration method, a target music signal to be restored is obtained; the target music signal is input into a trained music master tape restoration model, and a music master tape signal corresponding to the target music signal is generated by the music master tape restoration model; the music master tape restoration model is obtained by the music master tape restoration model processing method described in any of the above embodiments. In the music master tape restoration method provided by the present application, music master tape restoration can be performed through a trained music master tape restoration model, and the model can restore the master tape based on the time-frequency domain structure of the sub-band and U-net fusion during the processing. Therefore, the restoration of the master tape signal achieved by this model can improve the quality of the restored master tape signal.

[0116] In one embodiment, a method for music master tape restoration is also provided. This method first constructs effective training data. Then, by leveraging the advantages of different models, a new model architecture is designed, greatly improving the quality of master tape restoration. This method can be divided into the following two steps: including model design and loss function.

[0117] 1. Model design:

[0118] The training and inference processes adopted in this embodiment are as Figure 6 shown. First, training data is constructed. The model input is the true master tape signal at 192 kHz downsampled to 44.1 kHz and then upsampled to 192 kHz, as Figure 7 shown. The training target is the true master tape signal at 192 kHz, as Figure 8 shown.

[0119] This embodiment finds that master tape restoration requires restoring a 192 kHz master tape based on a 44.1 kHz standard-quality audio file SQ. Nearly 80% of the frequency band needs to be restored, far exceeding the 20% frequency band that does not need to be restored. Therefore, the challenge lies in how to restore a master tape with a natural spectrum. This embodiment first uses a sub-band model to predict the missing high frequencies. The bidirectional modeling of the time series and frequency-domain sub-band sequences in the sub-band model can better predict the smooth high-frequency part. However, relying solely on the sub-band model will result in excessive unevenness between the existing low-frequency spectrum and the predicted high-frequency spectrum. Therefore, in this embodiment, a time-domain Unet model is cascaded on the basis of the sub-band model to further fine-tune the full-frequency band spectrum, making the overall spectrum smoother. This embodiment realizes the integration of time-domain and frequency-domain models, and at the same time combines the rough estimation of the missing part by the sub-band model and the further fine-tuning of details by the Unet model. Therefore, the quality of master tape restoration can be improved.

[0120] Specifically, a sub-band model, such as the Apollo model, is first used, as Figure 9As shown in the figure, the process includes three parts: (1) The time-domain song signal x (i.e., the first sample input signal) is converted into a frequency-domain signal through Fourier transform, denoted as X (i.e., the first sample frequency-domain signal), which includes a time-dimension sequence divided by frames. At the same time, the frequency-domain signal is divided into subbands by frequency points to form a sequence in the frequency-domain dimension, denoted as Z (i.e., the subband signal); (2) Use Ropeformer to first model the sequence in the frequency-domain dimension to obtain Z' (i.e., the subband signal after frequency-domain dimension modeling), and then use the TCN model to model the sequence in the time dimension to obtain Q (i.e., the subband signal after bidirectional modeling); (3) The subband signal after bidirectional modeling is merged and restored to a frequency-domain signal to obtain Y (i.e., the second sample frequency-domain signal), and then the time-domain signal y1 (i.e., the second sample input signal) is obtained through inverse Fourier transform.

[0121] Then, x + y1 = y2 is used as the input of the time-domain U-net model. As Figure 10 shown, after the encoding and decoding of the Encoder and Decoder, the output is the final master tape, denoted as , where the Encoder and Decoder are composed of 1D convolutions, and GLU and GELU are activation functions.

[0122] 2. Loss function:

[0123] The time-frequency domain loss function in this embodiment is defined as:

[0124]

[0125] Among them, the time-domain loss function is as follows, which is used to constrain the consistency of the energy and distribution of the generated master tape and the original master tape.

[0126]

[0127] The frequency-domain loss function is as follows, which is used to characterize the high-frequency spectrum and enable the model to learn in the direction of restoring the master tape.

[0128]

[0129]

[0130]

[0131]

[0132] , are the second-order norm and the first-order norm respectively. Each contains a set of STFTs with a resolution. M is the number of STFTs with different resolutions. For example, the window lengths can be 4096, 2048, 1024 respectively; the frame shifts are 1024, 512, 256 respectively. is the proportional factor between time domain loss and frequency domain loss. J is 1. and They are the real master signal and the generated master signal respectively.

[0133] This embodiment is based on the time-frequency domain structure of the fusion of sub-band and U-net, and implements a robust master restoration algorithm by adding an effective loss function, which greatly improves the quality of master restoration.

[0134] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0135] Based on the same inventive concept, the embodiment of the present application also provides a music master restoration model processing device for implementing the music master restoration model processing method involved above, and a music master restoration device for implementing the music master restoration method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of one or more music master restoration model processing devices or music master restoration devices provided below can refer to the limitations of the music master restoration model processing method or music master restoration method in the above text, and will not be repeated here.

[0136] In one embodiment, Figure 11 As shown, a music master restoration model processing device is provided, comprising: a first signal acquisition module 1101, a second signal acquisition module 1102, a predicted master acquisition module 1103 and a restoration model training module 1104, wherein:

[0137] The first signal acquisition module 1101 is used to acquire a sample music master signal and a first sample input signal associated with the sample music master signal; the first sample input signal is a music signal obtained by removing a high-frequency spectrum from the sample music master signal;

[0138] The second signal acquisition module 1102 is configured to input the first sample input signal into the sub-band model included in the music master restoration model to be trained, obtain the sub-band signal corresponding to the first sample input signal, and generate a second sample input signal by using the sub-band signal;

[0139] The predicted master acquisition module 1103 is configured to superimpose the first sample input signal and the second sample input signal, input them into the U-net model included in the music master restoration model, and obtain the predicted music master signal;

[0140] The restoration model training module 1104 is configured to train the music master restoration model by using the difference between the predicted music master signal and the sample music master signal, so as to obtain the trained music master restoration model.

[0141] In one embodiment, as Figure 12 shown, a music master restoration device is provided, including: a target music acquisition module 1201 and a master signal restoration module 1202, wherein:

[0142] The target music acquisition module 1201 is configured to acquire a target music signal to be restored to the music master;

[0143] The master signal restoration module 1202 is configured to input the target music signal into the trained music master restoration model, and generate a music master signal corresponding to the target music signal through the music master restoration model; the music master restoration model is obtained by processing according to the music master restoration model processing method described in any one of the above embodiments.

[0144] Each module in the above music master restoration model processing device and the music master restoration device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0145] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 13As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample music master tape signal data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for processing a music master restoration model or a method for restoring a music master.

[0146] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0147] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0148] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0149] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0150] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0151] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiments of the method can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0152] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0153] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for processing a music master tape restoration model, characterized in that, The method includes: Obtaining a sample music master signal and a first sample input signal associated with the sample music master signal; the first sample input signal is a music signal obtained by removing the high-frequency spectrum from the sample music master signal; Inputting the first sample input signal into a sub-band model included in a music master restoration model to be trained, obtaining a sub-band signal corresponding to the first sample input signal, and generating a second sample input signal by using the sub-band signal; the second sample input signal is a music signal restored by using the sub-band signal; Superimposing the first sample input signal and the second sample input signal, and inputting them into a U-net model included in the music master restoration model to obtain a predicted music master signal; Training the music master restoration model by using the difference between the predicted music master signal and the sample music master signal to obtain a trained music master restoration model.

2. The method according to claim 1, wherein The sub-band model includes: a frequency band splitting module; the step of inputting the first sample input signal into the sub-band model included in the music master restoration model to be trained and obtaining a sub-band signal corresponding to the first sample input signal includes: Inputting the first sample input signal into the sub-band model, performing a Fourier transform on the first sample input signal through the sub-band model to obtain a first sample frequency domain signal corresponding to the first sample input signal; Inputting the first sample frequency domain signal into the frequency band splitting module, and dividing the first sample frequency domain signal into sub-bands by frequency points through the frequency band splitting module to obtain the sub-band signal.

3. The method according to claim 2, wherein The sub-band model further includes: a frequency band sequence modeling module and a frequency band reconstruction module; the step of generating a second sample input signal by using the sub-band signal includes: Inputting the sub-band signal into the frequency band sequence modeling module, performing a bidirectional modeling process on the sub-band signal through the frequency band sequence modeling module to obtain a bidirectional modeled sub-band signal; Inputting the bidirectional modeled sub-band signal into the frequency band reconstruction module, and generating a second sample frequency domain signal through the frequency band reconstruction module; Performing an inverse Fourier transform on the second sample frequency domain signal to obtain the second sample input signal.

4. The method according to claim 1, characterized in that The step of training the music master restoration model by using the difference between the predicted music master signal and the sample music master signal includes: Obtaining a first loss value according to the time domain difference between the predicted music master signal and the sample music master signal, and obtaining a second loss value according to the frequency domain difference between the predicted music master signal and the sample music master signal; Obtaining a total loss value of the music master restoration model according to the first loss value, the second loss value, and a proportionality factor between the preset time domain loss and the frequency domain loss, and training the music master restoration model by using the total loss value.

5. The method according to claim 4, wherein The step of obtaining a second loss value according to the frequency domain difference between the predicted music master signal and the sample music master signal includes: Obtain a plurality of preset resolution parameters, and perform short-time Fourier transform on the predicted music master signal and the sample music master signal respectively by using the plurality of resolution parameters to obtain a predicted master frequency-domain signal and a sample master frequency-domain signal corresponding to each resolution parameter; For each current resolution parameter, obtain a current second loss value corresponding to the current resolution parameter according to the difference between the current predicted master frequency-domain signal corresponding to the current resolution parameter and the current sample master frequency-domain signal corresponding to the current resolution parameter; the current resolution parameter is any one of the resolution parameters; Obtain the second loss value according to each of the current second loss values.

6. The method according to claim 5, characterized in that, The obtaining the current second loss value corresponding to the current resolution parameter according to the difference between the current predicted master frequency-domain signal corresponding to the current resolution parameter and the current sample master frequency-domain signal corresponding to the current resolution parameter includes: Obtain a first sub-loss value according to the complex spectrum difference between the current predicted master frequency-domain signal and the current sample master frequency-domain signal, and obtain a second sub-loss value according to the spectral amplitude difference between the current predicted master frequency-domain signal and the current sample master frequency-domain signal; Obtain the current second loss value according to the first sub-loss value and the second sub-loss value.

7. The method according to any one of claims 1 to 6, characterized in that The step of obtaining the first sample input signal includes: Perform downsampling processing on the sample music master signal to obtain a downsampled signal; wherein the sampling frequency of the sample music master signal is a first sampling frequency, the sampling frequency of the downsampled signal is a second sampling frequency, and the first sampling frequency is higher than the second sampling frequency; Perform upsampling processing on the downsampled signal to obtain a first sample input signal with a sampling frequency of the first sampling frequency.

8. A method for restoring a music master tape, characterized in that, The method includes: Obtain a target music signal to be restored to a music master; Input the target music signal into the trained music master restoration model, and generate a music master signal corresponding to the target music signal through the music master restoration model; the music master restoration model is obtained by processing with the music master restoration model processing method according to any one of claims 1 to 7.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.