Information processing system, information processing method, and information processing program

The demastering process using machine learning models reverses the effects of CD mastering, enabling flexible conversion of audio content to new formats by restoring pre-mastered data, thus facilitating easier remixing and distribution.

WO2025173586A1PCT designated stage Publication Date: 2025-08-21SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/003499
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-15
Filing Date
2025-02-04
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing audio content mastered for CDs is not suitable for conversion into new formats like spatial audio, as the mastering process is often irreversible and not adaptable, limiting the remixing and distribution of such content.

Method used

A demastering process using machine learning models, specifically a diffusion model, to reverse the effects of mastering processes like compression and limiting, restoring audio data to its pre-mastered state, enabling easier conversion to new formats.

Benefits of technology

The demastering process effectively removes the effects of mastering, allowing for clearer sound source separation and easier remixing into new formats, such as spatial audio, enhancing the flexibility and quality of audio content conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025003499_21082025_PF_FP_ABST
    Figure JP2025003499_21082025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing system according to the present disclosure comprises a processing unit that generates second audio data by performing demastering processing on first audio data that has been subjected to mastering processing.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and information processing program

[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program.

[0002] Commercially available audio content, such as CDs (CD-DA: Compact Disk Digital Audio), is basically mastered to suit the format of the audio data it contains. In the mastering process, stereo audio data, for example, consisting of left (L) and right (R) channels, is typically subjected to various processes, such as volume level compression and limiting, and sound quality adjustment, using a compressor, limiter, equalizer, etc.

[0003] Meanwhile, spatial audio has been commercialized in recent years, and several formats relating to audio data for realizing spatial audio have been developed.

[0004] Japanese Patent Application Laid-Open No. 2002-186100

[0005] For example, audio data that has been mastered for recording on a CD could be remade into audio data in a new format, such as the spatial audio described above, and then commercialized.

[0006] However, in many cases, the mastering process that has already been applied to audio data is not suitable for converting it into a new format. Since almost all existing audio content with high reusability has been mastered, the only way to convert it into a new format is to use the mastering process as a prerequisite.

[0007] An object of the present disclosure is to provide an information processing system, an information processing method, and an information processing program that can easily convert existing audio content into audio content in a different format.

[0008] The information processing system according to the present disclosure includes a processing unit that generates second audio data by performing a demastering process on mastered first audio data.

[0009] FIG. 1 is a schematic diagram for explaining a mastering process. FIG. 1 is a schematic diagram for explaining the operation of a compressor. FIG. 1 is a schematic diagram for generally explaining the demastering process according to the embodiments. FIG. 2 is a block diagram showing an example of a configuration of an information processing system for performing the demastering process, applicable to each embodiment. FIG. 2 is a block diagram showing an example of a configuration of an information processing system for performing the demastering process, applicable to each embodiment. FIG. 3 is a block diagram showing an example of a configuration of an information processing device according to the embodiment. FIG. 4 is a functional block diagram showing an example of a function of the information processing device according to the embodiment. FIG. 5 is a schematic diagram showing a method for training a machine learning model according to the embodiment. FIG. 6 is a schematic diagram showing an example of using non-mastered data in training a machine learning model according to the embodiments. FIG. 7 is a schematic diagram showing a method for training a diffusion model, applicable to the embodiments. FIG. 8 is a block diagram showing an example of a configuration for performing processing according to the first embodiment. FIG. 9 is a block diagram showing processing in a guide processing unit according to the first embodiment. FIG. 10 is a block diagram showing an example of a configuration for performing processing according to the second embodiment. FIG. 11 is a schematic diagram showing an example of a parameter adjustment screen according to the second embodiment. FIG. 12 is a block diagram showing an example of a configuration for performing processing according to the third embodiment. FIG. 13 is a block diagram showing an example of a configuration for performing processing according to the fourth embodiment. FIG. 14 is a block diagram showing an example of a configuration for performing processing according to the fifth embodiment. FIG. 15 is a schematic diagram showing an example of a configuration for performing processing according to the sixth embodiment.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are denoted by the same reference numerals, and redundant description will be omitted.

[0011] Hereinafter, embodiments of the present disclosure will be described in the following order: 1. Overview of embodiments of the present disclosure 1-1. Overview of mastering process 1-2. Regarding existing technology 1-3. General explanation of embodiments of the present disclosure 2. Configurations applicable to each embodiment 3. Learning method according to each embodiment 4. First embodiment of the present disclosure 5. Second embodiment of the present disclosure 6. Third embodiment of the present disclosure 7. Fourth embodiment of the present disclosure 7-1. Modified example of the fourth embodiment 8. Fifth embodiment of the present disclosure 9. Sixth embodiment of the present disclosure

[0012] (1. Overview of Embodiments of the Present Disclosure) Commercially available audio content, such as audio data recorded on a CD-DA (Compact Disk-Digital Audio), is generally subjected to a so-called mastering process to adjust sound quality and sound pressure, etc. In the present disclosure, a demastering process is performed on the mastered audio data to remove the mastering process, thereby restoring the audio data to its pre-mastering state.

[0013] (1-1. Mastering Process Overview) Here, for ease of understanding, the mastering process for audio data will be briefly described. Fig. 1 is a schematic diagram for explaining the mastering process.

[0014] The mastering process is basically performed using an equalizer (EQ) 500, a compressor 501, and a limiter 502. Here, the audio data to be mastered is audio data obtained by mixing audio data from multiple tracks into L (left) and R (right) channels (hereinafter referred to as 2MIX data).

[0015] The EQ 500 adjusts the level of a specified frequency band of input audio data. The compressor 501 and limiter 502 process the input audio data in relation to the dynamic range. More specifically, the compressor 501 compresses the level of the input audio data that exceeds a specified threshold Th at a specified ratio. The limiter 502 limits the level of the input audio data to a specified level.

[0016] FIG. 2 is a schematic diagram illustrating the operation of compressor 501. In FIG. 2, the horizontal axis represents the level of input audio data (dB: decibels), and the vertical axis represents the level of output audio data relative to the input audio data (dB). In the diagram, "1:1," "2:1," "4:1," and "∞:1" represent compression ratios relative to the input level, with the ratio "1:1" indicating a state in which no compression processing is performed by compressor 501. Meanwhile, the ratios "2:1" and "4:1" indicate that the output level for an input level exceeding threshold Th is compressed by a ratio of 1 / 2 and 1 / 4, respectively. The ratio "∞:1" limits the output level for an input level exceeding threshold Th to a level corresponding to threshold Th, which corresponds to the operation of limiter 502.

[0017] Compressor 501 compresses the level of input audio data, and in the case of a ratio of 2:1, for example, information is lost in the shaded area in Fig. 2. The processing by compressor 501 and limiter 502 is nonlinear and generally irreversible. On the other hand, EQ 500 is linear, increasing or decreasing the level of a specified frequency band, and therefore generally reversible.

[0018] The EQ 500 adjusts the level of a specified frequency band for the input 2MIX data in accordance with instructions from a mastering engineer who performs the mastering process, and outputs the result to the compressor 501. The compressor 501 compresses the input audio data in accordance with a specified threshold Th and ratio in accordance with instructions from the mastering engineer, and outputs the result to the limiter 502. The limiter 502 limits the level of the input audio data in accordance with instructions from the mastering engineer, and outputs the result as mastered data. The mastered data is used, for example, as data to be recorded on a commercially available CD-DA.

[0019] The mastering process described above is a basic example, and may be modified in various ways depending on the mastering engineer, the work environment in which the mastering process is performed, the intended use of the mastered data, etc. For example, in addition to the EQ 500, compressor 501, and limiter 502 described above, the mastering process may also include reverb processing to add reverberation. Furthermore, the order of the processes in the mastering process is not limited to the above, and the EQ process, compressor process, and limiter process may also be performed multiple times.

[0020] Hereafter, unless otherwise specified, CD-DA will be abbreviated as "CD." Furthermore, the mastering process suitable for recording on a CD is sometimes called "CD mastering," and audio data that has undergone CD mastering is sometimes called "mastered data."

[0021] (1-2. Existing Technology) As mentioned above, commercially available audio content (e.g., CDs) is basically mastered in its format (e.g., a stereo format with two channels, an L (left) channel and an R (right) channel). In some cases, this mastered audio data is remade into audio data in a different format for commercialization.

[0022] In such cases, the mastering process that has already been performed is often not suitable for converting to a new format. On the other hand, almost all existing content with high reusability has been mastered. Therefore, in the past, the only way to convert to a new format was to assume that the mastering process had already been performed.

[0023] As an example, various spatial audio formats have been developed in recent years that are becoming commercially available, such as 360RA (360 Reality Audio) (registered trademark), Dolby Atmos (registered trademark), and Auro 3D (registered trademark). However, new content incorporating spatial audio has not been distributed as widely as expected. This is thought to be largely due to the cost of producing music in a new format and the lack of clarity regarding its appeal over products in standard stereo formats.

[0024] Therefore, there is a demand for commercializing existing content by remixing it into the spatial audio format described above. However, most of the existing content with great reusability is resale of audio sources recorded on CDs, and in many cases, only the audio data (CD master) as the master material recorded on the CD remains. Therefore, it is necessary to use audio source separation to decompose this master audio data into elements for each part before remixing.

[0025] However, a CD master is mastered data that has undergone mastering processes for CDs, such as compressor processing, limiter processing, and in some cases even reverb processing, as explained with reference to Figure 1, and is therefore often not suitable for creating a new remix using that sound source. For this reason, there is a demand to remove the effects of the mastering processes applied to the CD master, such as the effects of the mastering process.

[0026] In response to such a request, the present disclosure performs a demastering process on the mastered data to remove the mastering process, thereby restoring the audio data before the mastering process.

[0027] In this specification, the demastering process refers to a process of restoring information that has been reduced to fit a specific format. As a specific example, the demastering process according to the embodiment may be a process of removing the effects of mastering processes, such as compressor processing and limiter processing, that have been applied to audio data to be recorded on a CD (an example of a specific format), and restoring the signal characteristics before the mastering process, as described above.

[0028] Furthermore, although the present specification describes the demastering process as being performed on audio data, the demastering process is not limited to audio data. In other words, the present disclosure can be applied to any process for the purpose of demastering, even if the demastering process is performed on data other than audio data. As an example, the demastering technology of the present disclosure may be applied when converting the format of data that has been mastered for recording on a recording medium such as a DVD (Digital Versatile Disk) or a Blu-Ray Disc (registered trademark) into a different format (for example, a format compatible with a video distribution service via the Internet).

[0029] (1-3. General Description of Embodiments of the Present Disclosure) Next, an embodiment of the present disclosure will be described in brief. Fig. 3 is a schematic diagram for explaining the demastering process according to the embodiment.

[0030] 3, the mastered data 20 and the demastered data 21 are each shown as waveforms with the vertical axis representing level (amplitude) with 0 dB at the center and the horizontal axis representing time. This applies to similar figures described below.

[0031] In an embodiment of the present disclosure, a demastering process 10 is performed on mastered data (first audio data) that has been subjected to a mastering process, and demastered data (second audio data) in which the mastering process in the mastered data has been removed is generated and output.

[0032] The demastering process 10 according to the embodiment may be performed using a machine learning model trained on unmastered audio data. As will be described in detail later, the demastering process 10 according to the embodiment may be performed using a diffusion model, which is one of the machine learning models.

[0033] The machine learning model applicable to the demastering process 10 according to the embodiment is not limited to the diffusion model. For example, a generative AI model based on a generative adversarial network (GAN), a variational autoencoder (VAE), or a transformer, which is one of deep learning models, may be applied as the machine learning model for executing the demastering process 10 according to the embodiment.

[0034] 3 is processed by compressor 501 in the mastering process described with reference to Fig. 1, whereby the level is increased and the level (amplitude) is compressed in areas where the absolute value of the input level is large, and further the level is limited to a predetermined level by limiter 502. As a result, the mastered data 20 has a waveform whose amplitude is clipped at a predetermined level.

[0035] On the other hand, the demastered data 21 has the effects of the processing by the compressor 501 and limiter 502 removed from the mastered data 20. More specifically, in the example of Fig. 3, the demastered data 21 has eliminated the amplitude clipping in the mastered data 20, and the level has been reduced even in low-level areas, making the amplitude clearer.

[0036] By subjecting this demastered data 21 to, for example, sound source separation processing, the audio data of each part can be obtained in a clearer state, making it possible to execute the remix processing more easily.

[0037] (2. Configurations Applicable to Each Embodiment) Next, configurations applicable to each embodiment of the present disclosure will be described.

[0038] 4 is a block diagram showing an example of the configuration of an information processing system for performing demastering processing that can be applied to each embodiment. As shown in FIG. 4, the information processing system 1a that can be applied to the embodiment can be configured as a standalone system without being connected to a communication network such as the Internet.

[0039] In FIG. 4, the information processing system 1 a includes a demastering processing device 1000 , a storage device 1100 , and an audio I / F 1200 .

[0040] The demastering processing device 1000 receives mastered data 20. The mastered data 20 may be input as a file or as stream data. The demastering processing device 1000 performs the above-described demastering process 10 on the input mastered data 20 and outputs demastered data 21.

[0041] The demastered data 21 output from the demastering processing device 1000 is stored in a storage device 1100 including a storage medium such as a hard disk drive or flash memory.

[0042] The audio I / F (interface) 1200 converts the audio data output from the demastering processing device 1000 into an analog audio signal and outputs the converted signal. The audio signal output from the audio I / F 1200 may be converted into sound by a speaker 1210 via, for example, an amplifier (not shown).

[0043] The demastering processing device 1000 may output the mastered data 20 and the demastered data 21 to the audio I / F 1200 and convert them into sound via a speaker 1210. This allows an engineer performing the demastering process to perform the demastering process 10 while, for example, comparing the sound based on the mastered data 20 with the sound based on the demastered data 21.

[0044] FIG. 5 is a block diagram showing an example of the configuration of a demastering processing device 1000 according to an embodiment.

[0045] The demastering processing device 1000 includes a CPU (Central Processing Unit) 1010, a ROM (Read Only Memory) 1011, a RAM (Random Access Memory) 1012, a display control unit 1013, a storage device 1014, a data I / F 1015, a communication I / F 1016, and a GPU (Graphics Processing Unit) 1017, and these units are connected to each other so as to be able to communicate with each other via a bus 1020. In this way, a general computer can be applied to the demastering processing device 1000.

[0046] The storage device 1014 is a non-volatile storage medium such as a hard disk drive or flash memory. A part of the storage area of ​​the storage device 1014 may be applied to the storage device 1100 described with reference to FIG.

[0047] The CPU 1010 controls the overall operation of the demastering processing device 1000 in accordance with programs stored in the storage device 1014 and the ROM 1011, using the RAM 1012 as a work memory.

[0048] The display control unit 1013 generates a display signal that can be displayed by the display device 1040, based on a display control signal generated by the CPU 1010 according to a program. The display device 1040 includes a display device such as an LCD (Liquid Crystal Display) and a drive circuit that drives the display device, and displays a screen according to the display signal output from the display control unit 1013.

[0049] The data I / F 1015 is an interface for transmitting and receiving data to and from an external device. In the example shown in the figure, an input device 1030 for receiving user operations and an audio I / F 1200 are connected to the data I / F 1015. In addition, mastered data 20 is input to the data I / F 1015.

[0050] The communication I / F 1016 controls communication with a communication network such as the Internet or a LAN (Local Area Network).

[0051] The GPU 1017 is a processor with high parallel processing capabilities and suitable for processing large amounts of data. The GPU 1017 is suitable for use in processing machine learning models using neural networks such as DNN (Deep Neural Network). In the figure, the GPU 1017 is shown as being built into the demastering processing device 1000, but this is not limited to this example. For example, the GPU 1017 may be configured as an external device connected to the demastering processing device 1000 via a data I / F 1015 or the like. Furthermore, if the CPU 1010 has sufficient processing capabilities, the GPU 1017 may be omitted.

[0052] FIG. 6 is a functional block diagram illustrating an example of the functions of the demastering processing device 1000 according to the embodiment.

[0053] 6 , the demastering processing device 1000 includes a processing unit 100, an overall control unit 101, a UI unit 102, an output unit 103, and a storage unit 104. The overall control unit 101 controls the overall operation of the demastering processing device 1000.

[0054] The processing unit 100 includes an acquisition unit 110 and a demastering processing unit 120. The acquisition unit 110 acquires mastered data 20. The demastering processing unit 120 performs a demastering process 10 on the mastered data 20 acquired by the acquisition unit 110 to generate demastered data 21. The demastering processing unit 120 performs the demastering process 10 using a machine learning model that has been trained in advance. More specifically, the demastering processing unit 120 according to the embodiment executes the demastering process 10 using a diffusion model. A method for training the diffusion model in the demastering processing unit 120 will be described later.

[0055] The UI unit 102 provides a UI (User Interface) related to the demastering process 10. For example, under the control of the overall control unit 101, the UI unit 102 displays a UI screen including a display of parameters related to the demastering process 10 on the display device 1040, and accepts user operations performed on the input device 1030 in accordance with the UI screen.

[0056] The output unit controls the output to the outside of the demastered data 21 generated by the processing unit 100. Furthermore, the storage unit 14 stores the demastered data 21 generated by the processing unit 100, for example, in the storage device 1100. The storage unit 104 may further store the mastered data 20 acquired by the acquisition unit 110 in the storage device 1100.

[0057] The processing unit 100, overall control unit 101, UI unit 102, output unit 103, and storage unit 104 may be configured by running an information processing program according to the embodiment on the CPU 1010. Alternatively, some or all of the processing unit 100, overall control unit 101, UI unit 102, output unit 103, and storage unit 104 may be configured by hardware circuits that operate in cooperation with each other. Furthermore, the processing unit 100 may be configured on the GPU 1017.

[0058] In the demastering processing device 1000, for example, the CPU 1010 executes the information processing program relating to the embodiment, thereby configuring the above-mentioned processing unit 100, overall control unit 101, UI unit 102, output unit 103 and memory unit 104, for example as modules, in the main memory area of ​​RAM 1012.

[0059] The information processing program can be obtained from outside via a communication network, for example, by communication via the communication I / F 1016, or from a storage medium connected to the data I / F 1015, and installed on the demastering processing device 1000.

[0060] (3. Learning Method According to Each Embodiment) Next, a learning method of a machine learning model used by the demastering processing unit 120 according to the embodiment will be described. In the embodiment, the demastering process 10 is performed using a diffusion model, which is one of the deep neural networks (DNNs) of generative AI (artificial intelligence).

[0061] 7 is a schematic diagram illustrating a learning method for a machine learning model according to an embodiment. Non-mastered data 31, which is audio data that has not been subjected to mastering processing, is input to a diffusion model 30. The diffusion model 30 is trained using the input non-mastered data 31.

[0062] The non-mastered data 31 used for training the diffusion model 30 is preferably audio data (e.g., 2MIX data) of the mastered data 20 recorded on a commercially available CD before mastering. Such audio data is generally difficult to obtain because it is not commercially available, but it can be obtained from production companies that produce content to be recorded on CDs or from some musicians who produce commercial content.

[0063] 8 is a schematic diagram showing an example of using non-mastered data 31 in training a machine learning model according to an embodiment. In FIG. 8, section (a) shows, for example, the entire non-mastered data 31. As shown in section (b), non-mastered data 31, 31, 31, ... of a predetermined length are extracted from this non-mastered data 31, and training of the diffusion model 30 is performed for each of these non-mastered data 31, 31, 31, ....

[0064] In the example shown in the figure, the non-mastered data 311, 312, 313, . . . are obtained by dividing the original non-mastered data 31 into 131,072 samples (=2 17The audio data is extracted in units of samples and used as is. As mentioned above, the mastering process involves nonlinear processing, and the rise of the signal is important for restoring this nonlinearly processed audio data. DNNs often use spectral data as input signals, but it is difficult to improve the accuracy of spectral data in the time axis direction due to the averaging effect of FFT (Fast Fourier Transform) and so waveform data in the time axis direction is used instead.

[0065] Also, as shown in section (b) of FIG. 8, the non-mastered data 31 1 , 31 2 , 31 3 , . . . may be overlapped in the time axis direction, and the data in the central portion where they do not overlap each other may be used for learning.

[0066] FIG. 9 is a schematic diagram illustrating a training method for the diffusion model 30 applicable to the embodiment. In the embodiment, a U-Net-based diffusion model 30 is used. As will be described later, the diffusion model 30 adds noise to training data, estimates the inverse process of the process leading to complete noise, and trains to minimize the error between the original training data and data from which noise has been removed by the inverse process. Due to these characteristics, the diffusion model 30 does not need to use paired data and can train using only the target audio data.

[0067] Section (a) of Figure 9 shows the forward process (diffusion process) in the learning process of the diffusion model 30. In the forward process, noise is gradually added to the original data. Looking at each stage of the forward process as a time series, noise is gradually added to the non-mastered data 31 at the initial time point X0, as shown as data 321, 322, 323, and 324. The added noise is, for example, Gaussian noise. From time point X0 to time point X t , time X t-1 Then noise is gradually added until the final point X T In this case, the data 33 is completely noise.

[0068] Section (b) of Figure 9 shows the reverse process in the learning process of the diffusion model 30. In the reverse process, estimation is performed in the reverse direction of the forward process based on the data 33 generated by the processing of the forward process. At this time, the parameters of the diffusion model 30 are optimized so that data close to the original data (non-mastered data 31) is obtained.

[0069] In the example shown in the figure, time X T From time X0 to time X0, data from which noise has been gradually removed is obtained, as shown by data 341, 342, 343, and 344. t-1 , time X t The noise is gradually removed through T In this case, data 35 is obtained by restoring the original non-mastered data 31.

[0070] Here, when estimation is performed from data 33 which is completely noise, data that matches the original non-mastered data 31 is not necessarily obtained. For this reason, mastered data 20, which is obtained by performing a mastering process on the non-mastered data 31, is convolved with the diffusion model 30 using a predetermined weight at each stage of noise removal. This makes it possible to obtain data 35 that is correlated with the original non-mastered data 31.

[0071] In addition, when amplitude is represented on the X axis and time on the Y axis, audio data can be considered as data in which each sampling point is represented by coordinates (X, Y). Therefore, a method similar to the diffusion model learning method for image data can also be applied to audio data.

[0072] In the above description, the non-mastered data 31 is used as is, i.e., as data on amplitude changes along the time axis, for training the diffusion model 30. However, this is not limited to this example. For example, spectral data based on the non-mastered data 31 may also be used for training the diffusion model 30.

[0073] In this embodiment, by performing the demastering process 10 on the mastered data 20 using the diffusion model 30 trained based on the non-mastered data 31, it is possible to remove the nonlinear processing (compressor processing, limiter processing, etc.) related to the dynamic range applied to the mastered data 20 and restore the original dynamic range.

[0074] As described above, mastering may involve EQ and reverb processing in addition to compressor and limiter processing. EQ processing is generally linear, so the original characteristics can be restored by performing EQ processing with the opposite characteristics to those of the mastering processing. On the other hand, reverb processing adds reverberation, so generally no information is lost and it can be removed using a known technique called WPE (Weighted Prediction Error).

[0075] 4. First Embodiment of the Present Disclosure Next, a first embodiment of the present disclosure will be described. Fig. 10 is a block diagram showing an example configuration for performing processing according to the first embodiment.

[0076] 10, the demastering processing unit 120 includes a guide processing unit 130 and a demastering model 140. The mastered data 20 is input to the guide processing unit 130 and the demastering model 140.

[0077] The demastering model 140 is the diffusion model 30 trained according to the training method described with reference to FIGS.

[0078] The guide process unit 130 passes information that serves as a guide for the demastering process 10 in the demastering model 140 to the demastering model 140. For example, the guide process unit 130 uses a simple reverse logic of the mastering process to guide the estimation process by the demastering model 140. As an example, the guide process unit 130 applies an existing dereverberation method such as the above-mentioned WPE to the reverb processing.

[0079] The demastering model 140 has a demastering function configured using a guide that is tailored to the mastering processing (compressor processing, limiter processing, reverb processing, etc.) performed by the guide processing unit 130 .

[0080] 11 is a block diagram for explaining the processing in the guide processing unit 130 according to the first embodiment. The guide processing unit 130 executes the processing of parameter estimation 131, coefficient extraction 132, and weighting 133 based on input mastered data.

[0081] The parameter estimation 131 is a process for estimating parameters of the processing executed in the mastering process for the input mastered data. For example, for the compressor processing in the mastering process, parameters such as the threshold Th and ratio, which are parameters related to the dynamic range in the compressor processing, and the attack time and release time, which are parameters in the time axis direction in the compressor processing, may be estimated.

[0082] Coefficient extraction 132 is a process of extracting coefficients to be applied to the demastering model 140 for each parameter estimated by parameter estimation 131. Weighting 133 is a process of weighting the coefficients extracted by coefficient extraction 132. The weighting value may be a preset value, or may be a value specified by an engineer using a UI (User Interface) described later.

[0083] The guide processor 130 applies the weighted coefficients to the demastering model 140 as parameters of the demastering model 140 .

[0084] The guide processing unit 130 can apply a trained machine learning model. For example, the guide processing unit 130 may apply a machine learning model trained based on the mastered data 20, non-mastered data corresponding to the mastered data 20, and parameters used in each process in the mastering process for the mastered data 20.

[0085] 10 , in the demastering processing unit 120, the demastering model 140 performs the demastering process 10 on the input mastered data 20 based on the parameters applied by the guide processing unit 130, and removes the mastering process from the mastered data 20. The demastering model 140 outputs demastered data 21 in which the mastering process has been removed from the mastered data 20.

[0086] As described above, in the first embodiment, the demastering process 10 is performed on the mastered data 20 using the demastering model 140, which is the trained diffusion model 30, to remove the mastering process from the mastered data 20. In the first embodiment, this makes it possible to restore the state of the mastered data 20 before the mastering process, and by using the demastered data 21 in which the state before the mastering process has been restored, it becomes possible to easily remake existing audio content into audio content in a different format.

[0087] (5. Second Embodiment of the Present Disclosure) Next, a second embodiment of the present disclosure will be described. The second embodiment is an example of a case where a mastering process (remastering process) is performed again on demastered data 21, from which the mastering process of the mastered data 20 has been removed as described above.

[0088] FIG. 12 is a block diagram showing an example of a configuration for performing processing according to the second embodiment.

[0089] 12, an adjustment unit 150 is added to the configuration shown in Fig. 10 above, and remastering 40, which is a second mastering process, is performed on the demastered data 21 output from the demastering processing unit 120. The remastering 40 is a process of adjusting the sound quality and sound pressure using an EQ 500, a compressor 501, a limiter 502, etc., as outlined above with reference to Fig. 1, for example.

[0090] In this way, in the second embodiment, the remastering 40 is performed on the demastered data 21, which is the data obtained by removing the mastering process from the mastered data 20. Therefore, the remastering 40 can be performed with a higher degree of freedom.

[0091] The adjustment unit 150 controls the guide process unit 130 to adjust the characteristics of the demastered data 21 output from the demastering model 140. The adjustment unit 150 may control the guide process unit 130 based on predetermined parameters, or may allow an engineer to make adjustments using a UI (described later) by the UI unit 102.

[0092] The guide process unit 130 according to the second embodiment may output, as demastering information, information including the coefficients extracted by the coefficient extraction 132. As an example, when compressor processing parameters (e.g., threshold Th, ratio, attack time, and release time) are acquired by the parameter estimation 131 and the coefficient extraction 132, these parameters may be included in the demastering information.

[0093] The demastering information is provided to the remastering unit 40. The engineer performing the remastering unit 40 (remastering engineer) can perform the remastering unit 40 by referring to this demastering information. As an example, consider a case where the demastering information includes parameters for the compressor processing performed on the mastered data 20. In this case, the remastering engineer can avoid performing the same mastering processing on the mastered data 20 in the remastering unit 40, and can estimate how to adjust each parameter in the remastering unit 40.

[0094] 13 is a schematic diagram showing an example of a parameter adjustment screen 200 according to the second embodiment. The parameter adjustment screen 200 is presented by the UI unit 102 based on, for example, information acquired from the adjustment unit 150. The parameter adjustment screen 200 is also shown as a parameter adjustment UI in the drawing. In the example of FIG. 13, the parameter adjustment screen 200 has display areas 210, 211, and 212, and buttons 213 and 214 arranged thereon.

[0095] The display area 210 displays identification information (e.g., name) that identifies the remastering engineer who will perform the remastering 40. The display area 210 may be configured so that the remastering engineer can input information. Furthermore, the display area 211 displays identification information (e.g., name) that identifies the mastering engineer who performed the mastering process on the mastered data 20 that is the target of the remastering 40. For example, the adjustment unit 150 may acquire the identification information of the mastering engineer that is displayed in the display area 212 from, for example, metadata of the mastered data 20 or a database that links the mastered data 20 with mastering engineers.

[0096] The display area 212 displays adjustable parameters for the demastering processing unit 120. In the illustrated example, each parameter of the compressor processing (threshold Th, ratio, attack, release) and its respective value are displayed in display sections 220a, 220b, 220c, and 220d within the display area 212. For example, the name and value of each parameter may be each parameter estimated by parameter estimation 131 based on the mastered data 20 in the guide processing unit 130, and a value extracted for each parameter by coefficient extraction 132.

[0097] In the display area 212, the values ​​displayed in the display sections 220a, 220b, 220c, and 220d can be edited. Editing these values ​​makes it possible to adjust the sound quality of the demastered data 21. For example, a remastering engineer may edit these values ​​while checking the sound of the demastered data 21 output from the demastering processing section 120, for example, by outputting it from the speaker 1210.

[0098] Button 213 is a parameter update button that instructs updating of parameters. For example, in response to operation of button 213, the demastering processing unit 120 updates the parameters of the demastering model 140 based on the values ​​of each parameter displayed in each of the display units 220a to 220d in the display area 212. Button 214 is a process start button for starting the demastering process using the demastering model 140. For example, in response to operation of button 214, the demastering processing unit 120 starts the demastering process 10 using the demastering model 140.

[0099] Here, the identification information of the mastering engineer who performed the mastering process on the mastered data 20 may be linked to the parameters and values ​​estimated and extracted from the mastered data 20 and stored in a database or the like. Similarly, the identification information of the remastering engineer who performs remastering 40 based on the mastered data 20 may be linked to each parameter and value in the remastering 40 and further stored in the database.

[0100] In this way, the control of the guide process unit 130 may be stored and learned based on information from the mastering engineer who performed the mastering process on the mastered data 20 and / or the remastering engineer who performs the remastering 40 based on the demastered data 21 that has been demastered 10 from the mastered data 20. This makes it possible to develop the demastering processing unit 120 into a state that is more customized for each individual, for example.

[0101] (6. Third Embodiment of the Present Disclosure) Next, a third embodiment of the present disclosure will be described. The third embodiment is an example in which data in a format different from that of the mastered data 20 is generated based on demastered data 21 obtained by performing a demastering process 10 on mastered data 20.

[0102] FIG. 14 is a block diagram showing an example of a configuration for performing processing according to the third embodiment.

[0103] In Figure 14, an adjustment unit 150 is added to the configuration shown in Figure 10 above, and sound source separation 41, remixing 42 and remastering 40 are performed on the demastered data 21 output from the demastering processing unit 120.

[0104] The demastered data 21 output from the demastering processing unit 120 is separated into sound sources, i.e., audio data, for each part (vocals, guitar, bass, drums, etc.) by a sound source separation 41. Each piece of audio data separated from the demastered data 21 by the sound source separation 41 is converted by a remix 42 into data in a format suitable for another format. The other format may be, for example, a spatial audio format such as 360RA, Dolby Atmos, or Auro 3D. However, the other format may also be a surround audio format such as 5.1ch or 7.1ch.

[0105] The audio data whose format has been converted by the remix 42 is remastered by the remastering 40 into data conforming to the other format and output.

[0106] In the third embodiment, sound source separation 41 is performed on demastered data 21, which is the mastering process removed from mastered data 20, to generate audio data in another format. This allows for more flexible remixing 42 and remastering 40.

[0107] (7. Fourth Embodiment of the Present Disclosure) Next, a fourth embodiment of the present disclosure will be described. The fourth embodiment is an example in which sound source separation is performed on mastered data 20, and demastering processing 10 is performed on audio data of each part resulting from the sound source separation.

[0108] FIG. 15 is a block diagram showing an example of a configuration for performing processing according to the fourth embodiment.

[0109] 15, the mastered data 20 is separated into audio data of each part by the sound source separation 41. For example, the mastered data 20 is separated into n pieces of audio data by the sound source separation 41. As shown in FIG. 15, in the fourth embodiment, there are as many pairs of an adjustment unit 150 and a demastering processing unit 120 as there are pieces of audio data separated by the sound source separation 41 (n pieces in this example). Each piece of audio data separated from the sound source is demastered by the demastering processing units 120, 120, ..., 120. n The demastering processing units 1201, 1202, ..., 120 n Each piece of audio data input to has been subjected to mastering processing.

[0110] Demastering processing units 1201, 1202, ..., 120 n are the adjustment units 1501, 1502, ..., 150 n Based on the parameters passed from the , the input audio data is subjected to a demastering process 10 and output as demastered data.

[0111] Each demastering processing unit 1201, 1202, ..., 120 n Each demastered data output from is processed by remix 42 to become audio data in a predetermined format. The audio data that has been remixed 42 is then remastered 40 in accordance with the predetermined format and output as remastered data.

[0112] For example, the audio data obtained by performing the sound source separation 41 on the mastered data 20 may have different characteristics. n By adjusting the audio data, each demastering processing unit 1201, 1202, ..., 120 n It is possible to perform a demastering process 10 optimized for each part.

[0113] In addition, each of the demastering processing units 1201, 1202, ..., 120 n The parts that each of the demastering processing units 120 performs the demastering process 10 on can be fixed. For example, the demastering processing unit 120 performs the demastering process 10 on the vocal audio data, and the demastering processing unit 120 performs the demastering process 10 on the guitar audio data. n The type of audio data to be subjected to the demastering process 10 is fixed.

[0114] In this case, each of the demastering processing units 1201, 1202, ..., 120 n The demastering model 140 in each of the demastering processing units 120 is trained according to the type of audio data. In the above example, the demastering model 140 in the demastering processing unit 120 is trained using vocal audio data that has not been mastered. Similarly, the demastering model 140 in the demastering processing unit 120 is trained using vocal audio data that has not been mastered.

[0115] (7-1. Modification of the Fourth Embodiment) A modification of the fourth embodiment will be described. The modification of the fourth embodiment is an example in which the third embodiment and the fourth embodiment described above are combined.

[0116] For example, one of the multiple demastering processing units 120 performs demastering processing 10 on audio data of a specific part (e.g., vocal audio data) among the audio data that has undergone sound source separation on the mastered data 20. On the other hand, for the other parts, another of the multiple demastering processing units 120 performs demastering processing 10, and then performs sound source separation.

[0117] According to a modified example of the fourth embodiment, an optimized demastering process 10 can be applied to a specific part of audio data contained in the mastered data 20, and the number of required demastering processing units 120 can be reduced.

[0118] 8. Fifth Embodiment of the Present Disclosure Next, a fifth embodiment of the present disclosure will be described. The fifth embodiment is an example in which the technology of the present disclosure is applied to audio data recorded live.

[0119] Fig. 16 is a block diagram showing an example of a configuration for performing processing according to the fifth embodiment. In the fifth embodiment, the configuration described in the first embodiment using Fig. 10 can be applied almost as is.

[0120] Here, we will give a brief overview of live recording. For example, in a live performance at a concert hall or outdoor stage, a microphone is installed for each part (vocals, guitar amp, percussion, etc.), and the individual audio signals output from each microphone for each part are sent to a mixing console. In some cases, the audio signals for some parts are sent directly to the mixing console without using microphones.

[0121] A mixing console mixes each audio signal for each output channel and sends the output signal of each output channel to a speaker system. At this time, the mixing console applies compressor and / or limiter processing as necessary to each channel being mixed in order to prevent damage to the speaker system from sudden excessive output (such as a sudden vocal shout or a percussion attack) and to maximize the signal-to-noise (SNR) ratio.

[0122] When recording live, the signals of each channel are typically split and extracted at a mixing console for multi-channel recording. Therefore, the recorded audio signals of each channel are often subjected to compression or limiting. Therefore, even if the recorded audio signals of each channel are mixed into, for example, stereo audio data of the left and right channels, it may not be possible to obtain audio data with the same high sound quality as in a studio recording.

[0123] 16 , in the fifth embodiment, a demastering processing unit 120 performs demastering processing 10 on signals of each channel output from a mixing console (shown as recorded sound source data 23 in the figure) using a demastering model 140. The demastering model 140 removes nonlinear characteristics caused by compressor processing and / or limiter processing from the recorded sound source data 23, and outputs the nonlinear characteristic-removed data 24.

[0124] By mixing the non-linear characteristic removed data 24 for each channel, it is possible to obtain high quality audio data that is close to that of a studio recording.

[0125] (9. Sixth Embodiment of the Present Disclosure) Next, a sixth embodiment of the present disclosure will be described. In the above, the information processing system 1a applicable to the embodiment has been described as a standalone system configured without being connected to a communication network. In contrast, the sixth embodiment is configured as an information processing system using a communication network.

[0126] FIG. 17 is a schematic diagram illustrating an example of the configuration of an information processing system according to the sixth embodiment.

[0127] 17, an information processing system 1b includes a server 60 and a terminal device 50 connected via a communication network 2. The communication network 2 may be the Internet or a local area network (LAN) established in a closed environment such as an in-house network.

[0128] The server 60 includes functions equivalent to those of the demastering processing device 1000 shown in Fig. 6, for example. In the example of Fig. 17, the server 60 is shown configured in a cloud network 3 connected to a communication network 2. However, the server 60 is not limited to this, and may be configured as a single computer or may be configured as a distributed server across multiple computers.

[0129] A general personal computer can be applied as the terminal device 50. However, the terminal device 50 is not limited to this, and may also be a smartphone or a tablet computer.

[0130] The terminal device 50 transmits the mastered data to the server 60 via the communication network 2. The server 60 performs a demastering process 10 on the mastered data transmitted from the terminal device 50 to generate demastered data. The server 60 transmits the generated demastered data to the terminal device 50.

[0131] Here, the server 60 may transmit, via the UI unit 102, display control information to the terminal device 50 for displaying, for example, the parameter adjustment screen 200 described with reference to FIG. 13 . Based on the display control information transmitted from the server 60, the terminal device 50 may display the parameter adjustment screen 200 on its own display unit using, for example, a browser application installed on the terminal device 50. The terminal device 50 transmits control information corresponding to a user operation on the parameter adjustment screen 200 to the server 60. The server 60 may generate demastered data based on the control signal transmitted from the terminal device 50, using a function equivalent to that of the demastering processing device 1000.

[0132] As described above, in the sixth embodiment, the server 60 has the same functions as the demastering processing device 1000, and the terminal device 50 can obtain demastered data from which the mastering process has been removed by transmitting mastered data to the server 60. This reduces the processing load on the terminal device 50, making it possible to reduce costs.

[0133] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0134] The present technology may also be configured as follows: (1) An information processing system including a processing unit that generates second audio data by performing a demastering process on first audio data that has been mastered. (2) The information processing system according to (1), wherein the processing unit generates the second audio data by inputting the first audio data to a trained machine learning model. (3) The information processing system according to (2), wherein the machine learning model is a diffusion model. (4) The information processing system according to (2) or (3), wherein the machine learning model is trained based on audio data that has not been mastered. (5) The information processing system according to any of (2) to (4), wherein the machine learning model is trained using a forward process that gradually adds noise to non-mastered audio data that is audio data that has not been mastered, and a reverse process that gradually removes noise from the non-mastered audio data to which noise has been added by the forward process, wherein in the reverse process, the mastered audio data is convolved with a predetermined weight at each stage of the noise removal. (6) The information processing system according to any one of (1) to (5), wherein the mastering process includes nonlinear processing of a dynamic range of audio data, and the processing unit generates the second audio data by estimating a state of the first audio data before the nonlinear processing through the demastering process. (7) The information processing system according to (6), wherein the nonlinear processing includes at least one of a compression process that compresses a dynamic range of audio data and a limiting process that limits the dynamic range. (8) The information processing system according to any one of (1) to (7), wherein the processing unit includes an adjustment unit that adjusts characteristics of the second audio data obtained through the demastering process.(9) The information processing system according to (8), wherein the adjustment unit includes: an estimation unit that estimates parameters for the mastering process based on the first audio data; an extraction unit that extracts coefficients for the demastering process based on the parameters estimated by the estimation unit; and a weighting unit that weights the coefficients extracted by the extraction unit. (10) The information processing system according to (9), wherein the adjustment unit displays the parameters and presents a user interface that accepts user operations on the parameters. (11) The information processing system according to (9) or (10), further including: a subsequent-stage system that performs predetermined processing on the second audio data, wherein the adjustment unit outputs demastering information including the parameters estimated by the estimation unit to the subsequent-stage system. (12) The information processing system according to (11), wherein the subsequent-stage system performs remastering processing on the second audio data as a mastering process different from the mastering process. (13) The information processing system according to (12), wherein the subsequent system further performs a sound source separation process on the second audio data and a remix process in which the second audio data is mixed with each piece of audio data separated by the sound source separation process, and the remastering process is performed on the audio data mixed by the remix process. (14) The information processing system according to (13), wherein the subsequent system converts a format of audio data obtained by mixing each piece of audio data separated by the remix process into a format different from that of the first audio data. (15) The information processing system according to (14), wherein the format different from that of the first audio data is a format related to spatial audio. (16) The information processing system according to any of (1) to (15), wherein the processing unit is provided for each of a plurality of audio data obtained by sound source separation of the first audio data.(17) The information processing system according to any one of (1) to (17), further comprising a subsequent-stage system that performs predetermined processing on the second audio data, wherein the subsequent-stage system performs a remix process that mixes each of the second audio data output from the processing units corresponding to each of the plurality of audio data, and a remastering process that is a mastering process different from the mastering process, on the audio data mixed by the remix process. (18) The information processing system according to any one of (1) to (17), wherein the first audio data is audio data that has been subjected to at least one of a compression process that compresses a dynamic range of audio data containing individual sound sources and a limiting process that limits the dynamic range, and the second audio data is audio data obtained by removing nonlinear components from the first audio data by the demastering process. (19) The information processing system according to (10), wherein the adjustment unit causes the user interface to display information identifying a first user who performed the mastering process on the first audio data and information identifying a second user who will perform the demastering process. (20) The information processing system according to (19), wherein the adjustment unit associates identification information identifying the first user with first parameters estimated for the mastering process performed by the first user, and associates identification information identifying the second user with second parameters obtained when the second user performed the demastering process on the first audio data that has been mastered by the first user, and stores the associations in a storage unit. (21) The information processing system according to any of (1) to (20), comprising: an information processing device including the processing unit. (22) The information processing system according to any one of (1) to (20), including: a server including the processing unit; and a terminal device connected to the server via a communication network, the terminal device transmitting the first audio data to the server and instructing the server to perform the demastering process.(23) An information processing method including: a generation process of generating second audio data by performing a demastering process on first audio data that has been mastered. (24) An information processing program that causes a computer to function as a processing unit that generates second audio data by performing a demastering process on first audio data that has been mastered.

[0135] 1a, 1b Information processing system 10 Demastering process 20 Mastered data 21 Demastered data 23 Recorded sound source data 24 Nonlinear characteristic removed data 30 Diffusion model 31, 311, 312, 313 Unmastered data 40 Remastering 41 Sound source separation 42 Remix 50 Terminal device 60 Server 100 Processing unit 102 UI unit 104 Storage unit 110 Acquisition unit 120, 1201, 1202, 120 n Demastering processing unit 130 Guide processing unit 131 Parameter estimation 132 Coefficient extraction 133 Weighting 140 Demastering model 150, 1501, 1502, 150 n Adjustment section 200 Parameter adjustment screen 1000 Demastering processing device

Claims

1. An information processing system comprising: a processing unit that generates second audio data by performing a demastering process on mastered first audio data.

2. The information processing system according to claim 1, wherein the processing unit generates the second audio data by inputting the first audio data into a trained machine learning model.

3. The information processing system according to claim 2, wherein the machine learning model is a diffusion model.

4. The information processing system according to claim 2, wherein the machine learning model is trained based on audio data that has not been subjected to mastering processing.

5. The information processing system of claim 2, wherein the machine learning model is trained using a forward process that gradually adds noise to non-mastered audio data, which is audio data that has not been mastered, and a reverse process that gradually removes noise from the non-mastered audio data to which noise has been added by the forward process, and wherein in the reverse process, the mastered audio data is convolved with a predetermined weight at each stage of the noise removal.

6. The information processing system of claim 1, wherein the mastering process includes nonlinear processing of the dynamic range of the audio data, and the processing unit generates the second audio data by estimating the state of the first audio data before the nonlinear processing through the demastering process.

7. The information processing system according to claim 6, wherein the nonlinear processing includes at least one of a compression process for compressing the dynamic range of audio data and a limiting process for limiting the dynamic range.

8. The information processing system according to claim 1, wherein the processing unit includes an adjustment unit that adjusts characteristics of the second audio data obtained by the demastering process.

9. The information processing system of claim 8, wherein the adjustment unit includes: an estimation unit that estimates parameters for the mastering process based on the first audio data; an extraction unit that extracts coefficients for the demastering process based on the parameters estimated by the estimation unit; and a weighting unit that weights the coefficients extracted by the extraction unit.

10. The information processing system according to claim 9, wherein the adjustment unit displays the parameters and presents a user interface that accepts user operations on the parameters.

11. An information processing system as described in claim 9, further comprising a subsequent system that performs predetermined processing on the second audio data, wherein the adjustment unit outputs demastering information including the parameters estimated by the estimation unit to the subsequent system.

12. The information processing system according to claim 11, wherein the latter system performs a remastering process on the second audio data as a mastering process different from the mastering process.

13. The information processing system of claim 12, wherein the latter system further performs a sound source separation process on the second audio data and a remix process in which the second audio data mixes each of the audio data separated by the sound source separation process, and the remastering process is performed on the audio data mixed by the remix process.

14. The information processing system according to claim 13, wherein the subsequent system converts the format of the audio data obtained by mixing the respective pieces of audio data separated by the remix process into a format different from that of the first audio data.

15. The information processing system according to claim 14, wherein the format different from the first audio data is a format related to spatial audio.

16. The information processing system according to claim 1, wherein the processing unit is provided for each of a plurality of audio data obtained by sound source separation of the first audio data.

17. An information processing system as described in claim 16, further comprising a subsequent system that performs predetermined processing on the second audio data, wherein the subsequent system performs a remix process that mixes each of the second audio data output from the processing units corresponding to each of the plurality of audio data, and a remastering process that is a mastering process different from the mastering process on the audio data mixed by the remix process.

18. The information processing system of claim 1, wherein the first audio data is audio data recorded from an individual sound source that has been subjected to at least one of a compression process that compresses the dynamic range and a limitation process that limits the dynamic range, and the second audio data is audio data from which nonlinear components have been removed by the demastering process.

19. The information processing system of claim 10, wherein the adjustment unit causes the user interface to display information identifying a first user who performed the mastering process on the first audio data and information identifying a second user who will perform the demastering process.

20. The information processing system described in claim 19, wherein the adjustment unit associates identification information identifying the first user with a first parameter estimated for the mastering process performed by the first user, and associates identification information identifying the second user with a second parameter when the second user performed the demastering process on the first audio data that was mastered by the first user, and stores the associations in a memory unit.

21. The information processing system according to claim 1, comprising: an information processing device including the processing unit.

22. The information processing system according to claim 1, comprising: a server including the processing unit; and a terminal device connected to the server via a communication network, for transmitting the first audio data to the server and instructing the server to perform the demastering process.

23. An information processing method including a generation process of generating second audio data by performing a demastering process on first audio data that has been mastered.

24. An information processing program for causing a computer to function as a processing unit that generates second audio data by performing a demastering process on mastered first audio data.

Citation Information

Patent Citations

  • Information processing device, method, and program

    WO2018131513A1

  • Signal processing device, signal processing method, and program

    WO2021059718A1

  • Signal processing device and method, and program

    WO2021172053A1

  • Diffusion models having improved accuracy and reduced consumption of computational resources

    WO2022265992A1

  • Local cross-attention operations in neural networks

    WO2023144385A1