Information processing system, information processing method, and computer program

The system generates trained models for music production tasks using a music-based model and training datasets, addressing resource and cost inefficiencies by adapting to common features and user preferences, achieving high-performance and efficient music production.

WO2025204213A1PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/JP2025/004474
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-02-12
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Developing music production applications using machine learning models requires significant development resources and costs due to the need for large amounts of training data and retraining for specific musical genres or trends, leading to inefficiencies in model training for various music-related tasks.

Method used

An information processing system that generates trained models for downstream music production tasks using a music-based model and corresponding training datasets, allowing adaptation to common intermediate features, reducing the need for extensive training data and resources.

Benefits of technology

Enables high-performance trained models for music production tasks with reduced resource and cost requirements, while maintaining consistency across tasks and allowing adaptation to specific musical genres or user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025004474_02102025_PF_FP_ABST
    Figure JP2025004474_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides an information processing system for performing processing related to music using AI technology. An information processing system according to the present invention includes an acquisition unit that acquires a trained music foundation model, and a generating unit that generates a trained model adapted to a downstream task on the basis of a training dataset related to the music foundation model and the downstream task. The model generating unit generates a plurality of trained models for each downstream task, on the basis of common intermediate features with the music foundation model, and a plurality of training datasets respectively corresponding to the plurality of the downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and computer program

[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing system, an information processing method, and a computer program that perform music-related processing.

[0002] In recent years, with the advancement of AI (Artificial Intelligence) technology, machine learning has come to be applied to content production. For example, a system that generates music from text using a machine learning model has been proposed (see Patent Document 1).

[0003] Many applications using machine learning models have been developed to assist in parts of the music production process. When developing applications for various music production tasks (e.g., music generation, audio source separation, automatic transcription, automatic arrangement, automatic mixing, song tagging, etc.), model developers must design and train models with appropriate architectures and prepare large amounts of training data for each task, resulting in enormous development resources and costs. Furthermore, if you want to incorporate specific musical genres or musical trends into applications for a series of music-related tasks, you must retrain each model for each task using the desired musical information.

[0004] US Patent Application Publication No. 2018 / 0190249

[0005] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, "LoRA: Low-Rank Adaptation of Large Language Models" (arXiv:2106.09685)

[0006] An object of the present disclosure is to provide an information processing system, an information processing method, and a computer program that perform music-related processing using AI technology.

[0007] The present disclosure has been made in consideration of the above-mentioned problems, and a first aspect thereof is an information processing system including: an acquisition unit that acquires a trained music-based model; and a generation unit that generates a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

[0008] However, the term "system" used here refers to a logical collection of multiple devices (or functional modules that realize specific functions), regardless of whether each device or functional module is contained within a single housing. In other words, both a single device consisting of multiple parts or functional modules and a collection of multiple devices are considered "systems."

[0009] The generation unit generates a plurality of trained models for each downstream task based on the music-based model and a plurality of training datasets respectively corresponding to the downstream tasks, and generates the plurality of trained models for each downstream task using common intermediate features of the music-based models.

[0010] For example, the generation unit generates a trained model for each downstream task performed in each process of music production. The music production process includes at least one of composing, arranging, and mixing, and the generation unit generates a trained model adapted to a downstream task performed in at least one of the processes of composing, arranging, and mixing. Specifically, the downstream tasks include at least one of automatic composition, sound source separation, automatic transcription, and music tagging, and the trained model generated by the generation unit for each downstream task is used in at least the composition process.

[0011] The information processing system according to the first aspect may further include an adjustment unit that adjusts the music-based model based on at least one of the specified song and music genre, the music genre, the era, the tempo of the song, and the mood.

[0012] A second aspect of the present disclosure is an information processing method including: an acquisition step of acquiring a trained music-based model; and a generation step of generating a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

[0013] Furthermore, a third aspect of the present disclosure is a computer program written in a computer-readable format to cause a computer to function as: an acquisition unit that acquires a trained music-based model; and a generation unit that generates a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

[0014] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or via a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative effect is exerted on the computer, and the same effects as those of the information processing system according to the first aspect of the present disclosure can be obtained.

[0015] FIG. 1 is a diagram illustrating a learning method proposed in the present disclosure. FIG. 2 is a diagram illustrating a method for adapting each downstream task to a music infrastructure model 101. FIG. 3 is a diagram illustrating an example in which the present disclosure is applied to a music production process. FIG. 4 is a diagram illustrating a music generation process 401 using a machine learning model. FIG. 5 is a diagram illustrating a music generation process 500 for a professional user. FIG. 6 is a diagram illustrating an example of a UI when a user controls the music generation process 401 using a machine learning model. FIG. 7 is a diagram illustrating an automatic music arrangement process 701 using a machine learning model. FIG. 8 is a diagram illustrating an automatic music arrangement process 800 for a professional user. FIG. 9 is a diagram illustrating an example of a UI when a user controls the automatic music arrangement process 701 using a machine learning model. FIG. 10 is a diagram illustrating an automatic mixing process 1000 using a machine learning model. FIG. 11 is a diagram illustrating an example of a UI when a user controls the automatic mixing process 1001 using a machine learning model. FIG. 12 is a diagram showing an example UI when a user reflects a specific music genre or the tendencies of music created by the user in the music-based model itself. FIG. 13 is a flowchart showing the process of editing background music for content using AI technology. FIG. 14 is a diagram showing a learning method applied in a second embodiment of the present disclosure. FIG. 15 is a diagram showing the operation of adapting a training dataset 1421 of data amount for music generation and a music generation model 1411 to a music-based model 1401. FIG. 16 is a diagram showing the operation of adapting a training dataset 1422 of data amount for music search and a music search model 1412 to the music-based model 1401. FIG. 17 is a diagram showing the operation of adapting a training dataset 1423 of data amount for automatic remixing and an automatic remix model 1413 to the music-based model 1401. FIG. 18 is a diagram showing the operation of adapting a training dataset 1424 of data amount for automatic editing and an automatic editing model 1414 to the music-based model 1401. Fig. 19 is a diagram showing a process for automatically generating background music to be added to content. Fig. 20 is a diagram showing an example of a UI when a user uploads content when background music is automatically generated. Fig. 21 is a diagram showing a process for automatically searching for background music to be added to content.Fig. 22 is a diagram showing an example of a UI when a user uploads content in the case of automatically searching for background music. Fig. 23 is a diagram showing a process of automatically searching for background music to be added to content and then automatically remixing it. Fig. 24 is a diagram showing an example of a UI when a user uploads content in the case of automatically searching for background music and then automatically remixing it. Fig. 25 is a diagram showing an example of a UI when the functions of automatically generating, automatically searching, and automatically editing background music are used in an integrated manner. Fig. 26 is a diagram showing an example of the hardware configuration of the information processing device 2000.

[0016] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.

[0017] A. Overview B. First Example B-1. Overall Configuration B-2. Automatic Composition Process B-3. Automatic Arrangement Process B-4. Automatic Mixing Process B-5. Control of the Entire Music Production Process B-6. Summary C. Second Example C-1. Overview C-2. Basic Configuration C-3. Background Music Editing Process Using Machine Learning C-4. Integrated UI Screen C-5. Summary D. Configuration of Information Processing Device

[0018] A. Overview Numerous applications have been developed that use machine learning models to assist in parts of the music production process. When developing applications for various music production tasks (e.g., music generation, audio source separation, automatic transcription, automatic arrangement, automatic mixing, song tagging, etc.), model developers must collect large amounts of training data for each individual task, design models with appropriate architectures, and train them, which poses challenges in terms of development resources and costs. Furthermore, if you want to reflect a specific music genre or musical trends across all music-related task applications, you must retrain each individual model for each task using the desired musical information.

[0019] Therefore, this disclosure proposes a technology for generating trained models for downstream tasks using a music-based model and training data corresponding to each task, instead of performing model training individually for multiple tasks related to music production. According to this disclosure, a high-performance trained model can be generated with a smaller amount of training data than when performing model training individually. In other words, according to this disclosure, a wide variety of models can be realized as downstream tasks with fewer resources and at fewer costs.

[0020] The base model is generated by learning various types of data, such as text, images, audio, and structured data, rather than data specialized for individual tasks. As a result, the base model also reflects the "combination of rules and principles / features" regarding the interrelationships between these data. By providing the base model with a small amount of data required to perform a specific task and learning from it, it becomes possible to adapt to new tasks.

[0021] FIG. 1 schematically illustrates the learning method proposed in this disclosure. In FIG. 1, downstream tasks include a wide variety of tasks related to music production, such as music generation, sound source separation, automatic music transcription, automatic music arrangement, automatic mixing, and music tagging. For each downstream task, a music generation model 111, a sound source separation model 112, an automatic music transcription model 113, an automatic music arrangement model 114, an automatic mixing model 115, a music tagging model 116, and so on are provided. The models for each downstream task may have an architecture appropriate for that task. In other words, the architecture of the models for each downstream task may differ.

[0022] In the present disclosure, the music-based model 101 is adapted to the training data and models corresponding to the downstream tasks described above, thereby generating trained models 111, 112, .... Adaptation makes it possible to generate high-performance trained models 111, 112, ... with a smaller amount of training data than when training models individually. In other words, it is possible to reduce the resources and costs required for training. As used herein, "adaptation" refers to the process of adapting a foundation model to a specific task. Adaptation allows the model for the downstream task to acquire knowledge appropriate for a specific domain or task.

[0023] Specifically, a music generation training dataset 121 and a music generation model 111 are adapted to the music-based model 101. The music generation model 111 is a model that generates music from input data of any modality, such as text, image, video, or audio (or a combination of two or more modalities). Therefore, the music generation training dataset 121 is a collection of data consisting of pairs of input data of a desired modality and music data (teacher data). By adapting the music-based model 101, a high-performance music generation model 111 can be generated with a smaller amount of music generation training dataset 121 compared to training the music generation model individually. Note that while the music-based model 101 itself is a model that generates music, the music generation model 111 generated by adaptation can generate music with further fine-tuned elements such as rhythm, tempo, the musical scales and instruments used, and so on.

[0024] Similarly, the training dataset 122 for sound source separation and the sound source separation model 112 are adapted. The sound source separation model 112 is a model that separates a piece of music into individual instruments. Therefore, the training dataset 122 for sound source separation is a collection of data consisting of pairs of music data in which the sound sources of all instruments are mixed before sound source separation and sound source data for each instrument after sound source separation. By using adaptation of the music-based model 101, it is possible to generate a high-performance sound source separation model 112 with a smaller amount of training dataset 122 for sound source separation compared to training sound source separation models individually.

[0025] In addition, a training dataset 123 for automatic music transcription and an automatic music transcription model 113 are adapted. The automatic music transcription model 113 is a model that converts Wave-format music data into MIDI data that allows elements such as pitch, dynamics, and timbre to be edited. Therefore, the training dataset 123 for automatic music transcription is a collection of data consisting of pairs of the original Wave data and MIDI data. By using adaptation of the music infrastructure model 101, a high-performance automatic music transcription model 113 can be generated with a smaller amount of training dataset 123 for automatic music transcription compared to training sound source separation models individually.

[0026] In addition, the automatic arrangement training dataset 124 and the automatic arrangement model 114 are adapted. The automatic arrangement training dataset 124 is a collection of data consisting of pairs of music data before and after arrangement. By using adaptation of the music foundation model 101, a high-performance automatic arrangement model 114 can be generated with a smaller amount of the automatic arrangement training dataset 124 compared to training sound source separation models individually.

[0027] In addition, the automatic-mixmodel training dataset 125 and the automatic-mixmodel model 115 are adapted. The automatic-mixmodel training dataset 125 is a collection of data consisting of pairs of music data before mixing (STEM) and music data after mixing. By using the adaptation of the music-based model 101, it is possible to generate a high-performance automatic-mixmodel model 115 with a smaller amount of the automatic-mixmodel training dataset 125 compared to training the automatic-mixmodel individually.

[0028] In addition, the song tagging training dataset 126 and the song tagging model 116 are adapted. The song tagging training dataset 126 is a collection of data consisting of pairs of music data and tag information associated with the music data. The tag information is, for example, text information describing the mood or genre of the music. By using adaptation of the music-based model 101, a high-performance song tagging model 116 can be generated with a smaller song tagging training dataset 126 than when training individual sound source separation models.

[0029] FIG. 2 shows a schematic diagram of a method for adapting each downstream task to the music infrastructure model 101.

[0030] 2, the music-based model 101 is constructed based on, for example, a Transformer. The music-based model 101 is composed of an encoder 201 and a decoder 202, and is a trained model that has been trained to restore input data.

[0031] As shown in Figure 2, when music data is input to the trained music-based model 101, common intermediate features extracted from the encoder 201 are applied (input) to models 111, 112, ... for each downstream task, and then each model 111, 112, ... is trained using the training data for each downstream task, thereby achieving adaptation. The music data input to the music-based model 101 may be any music. In the example shown in Figure 2, intermediate features 211 of a specific layer in the encoder 201 are input to the input layers of the models 111, 112, ... for all downstream tasks.

[0032] The method of inputting music data into the trained music-based model 101 is not particularly limited. For example, music data may be divided into multiple tokens and input to the encoder 201. Furthermore, the process of inputting intermediate features extracted from the music-based model 101 (encoder 201) to the input layers of the models 111, 112, ... of each downstream task may be performed only once, or may be performed multiple times with different music data input to the music-based model 101. Furthermore, in the example shown in FIG. 2 , only intermediate features of one specific layer are used, but in practice, features of any layer and their combinations may be used. Furthermore, in addition to the intermediate features, the music data itself may also be input to the input layers of the models 111, 112, ... of each downstream task.

[0033] The intermediate features extracted by the music-based model 101 include features that express the tendencies of the music. Therefore, in the present disclosure, by performing learning using common intermediate features of the music-based model 101 in each downstream task, it is expected that consistency in the behavior of the models 111, 112, ... for each downstream task will be maintained.

[0034] Furthermore, the learning method according to the present disclosure makes it easy to control the downstream tasks as a whole. Specifically, by reflecting the tendencies of a particular music genre or music created by the user in the music-based model 101 itself, it is not necessary to re-train each downstream task model 111, 112, ... using data on a particular genre or music, and it is possible to control the output of the downstream tasks as a whole to match the tendencies of a particular music genre or music created by the user.

[0035] In summary, the present disclosure provides an information processing system and an information processing method that use AI technology to perform various music-related processes. Note that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto. Furthermore, the present disclosure may also provide additional effects in addition to the above-described effects. Further, other objects, features, and advantages of the present disclosure will become apparent from the following detailed description of the embodiments and the accompanying drawings.

[0036] B. First Example In this section B, as a first example of the present disclosure, we introduce an example in which a downstream task trained using a music-based model is used in an actual music production process.

[0037] B-1. Overall Configuration In this embodiment, the music production process is roughly divided into four processes: "Composition," "Arrangement," "Mixing," and "Rendering." The upper part of Figure 3 shows a conventional example in which the user himself or herself performs all of the processes of composition 301, arrangement 302, mixing 303, and rendering 304. In the composition 301 process and arrangement 302 process, sound source data for each instrument is output in MIDI (Musical Instrument Digital Interface) format and Wave format, respectively. From the mixing process 303 onwards, music data obtained by mixing the sound source data for each piece of music is output in Wave format. In the following description of the embodiment, the case where MIDI format data is used will be explained as an example, but it is not limited to MIDI format data, and data in any format can be used as long as it contains musical score information.

[0038] The bottom section of Figure 3 shows an example of an automatic music generation process in which the composition, arrangement, and mixing stages of the music production process have been changed to processes in which ideas are output from a machine learning model and edited by the user. In Figure 3, processes entirely performed by the user are indicated by open blocks, while processes in which at least some processing is automated using a machine learning model are indicated by filled-in gray blocks. Comparing the top and bottom sections of Figure 3, the composition process 301 has been replaced by an automatic composition process 311 using a machine learning model and a user edit process 312, the arrangement process 302 has been replaced by an automatic arrangement process 321 using a machine learning model and a user edit process 322, and the mixing process 303 has been replaced by an automatic mixing process 331 using a machine learning model and a user edit process 332. However, in this embodiment, the rendering process 304 is not subject to automation using a machine learning model.

[0039] The automatic composition process 311 inputs multimodal data including text data, sound source data such as humming, and image data and video data such as illustrations, sketches, and photographs, and outputs sound source data for each instrument of the generated music in MIDI format and Wave format. The automatic arrangement process 321 inputs text data added by the user editing process 312 and MIDI format and Wave format sound source data for each instrument edited by the user editing process 312, and outputs MIDI format and Wave format sound source data for each instrument after automatic arrangement. The automatic mixing process 331 inputs text data added by the user editing process 322 and MIDI format and Wave format sound source data for each instrument after arrangement by the automatic arrangement process 321, and outputs automatically mixed Wave format music data together with effect parameters added during automatic mixing. The rendering process 304 inputs automatically mixed Wave format music data after editing it in the user editing process 332, and outputs rendered Wave format music data.

[0040] The automatic composition process 311 will be described in detail in Section B-2 below with reference to Figures 4 and 5. The automatic arrangement process 321 will be described in detail in Section B-3 below with reference to Figures 7 and 8. The automatic mixing process 331 will be described in detail in Section B-4 below with reference to Figure 10.

[0041] B-2. Automatic Composition Process Figure 4 shows a music generation (Music Generation from Multimodal) process 401 using a machine learning model. In this music generation process 401, the input is multimodal input including text data, sound source data such as humming, image data such as illustrations, sketches, and photographs, and video data, and the output is music data that mixes the sound source data of each instrument. The music generation process 401 uses a machine learning model to support the process itself of turning a subject that inspires the user into sound.

[0042] The music generation process 401 outputs a piece of music in which all the various parts are mixed together. Professional users with experience and knowledge in music production have a need not only to simply generate music but also to make detailed adjustments to the sound source, such as making fine adjustments to each part of the generated music. Therefore, when the music generation process 401 outputs completed music data for a piece of music, there may be cases where a sound source separation process is required to separate the generated music into individual instruments, or an automatic transcription process is required to convert the music into a MIDI file that allows editing of elements such as pitch, dynamics, and timbre.

[0043] 5 shows a music generation process 500 for professional users, which adds a sound source separation process 402 and an automatic music transcription process 403 to a music generation process 401. The music generation process 401 inputs multimodal data including text data, sound source data such as humming, and image and video data such as illustrations, sketches, and photographs, and outputs the generated music data in Wave format. The sound source separation process 402 separates the Wave format data (music file) output from the music generation process 401, which is a mix of sound source data for all instruments, into data for each instrument, and outputs the sound source data for each instrument in Wave format. The automatic music transcription process 403 converts the Wave format sound source data for each instrument output from the sound source separation process 402 into MIDI format.

[0044] By applying the music generation process 500 shown in Figure 5 to the automatic composition process 311, the automatic composition process 311, which is one of the music production processes that uses the machine learning model shown in the lower part of Figure 3, can be made for users who have experience and knowledge about music production, i.e., professional users.

[0045] 4 and 5 , processes using models for downstream tasks trained based on the music-based model (i.e., models generated using the present disclosure) are shown in gray. The models used in the music generation process 401, sound source separation process 402, and automatic music transcription process 403 can be generated with reduced resources and costs by adapting small amounts of training data and models corresponding to each downstream task (music generation, sound source separation, and automatic music transcription) to a single music-based model. Furthermore, by reflecting the trends of a specific music genre or music created by the user in the music-based model itself, there is no need to retrain each downstream task model 401-403 using data from a specific genre or music. Instead, the output for the entire downstream task can be controlled to match the trends of a specific music genre or music created by the user.

[0046] FIG. 6 shows an example of a UI (User Interface) when a user controls a music generation process 401 using a machine learning model.

[0047] The UI screen shown in FIG. 6 includes an add photo / video button 601, and by clicking this button, the user can add (upload) photos and videos to be input to the music generation process 401. As long as the image and video files are in a predetermined format, illustrations, sketches, and the like can also be added using the add photo / video button 601. A photo / video display section 611 displays the photos and videos that have been added (uploaded). The user can check the photos and videos that they have added in the photo / video display section 611.

[0048] 6 also includes an add sound / music button 602, and the user can click this button to add sound or music to be input to the music generation process 401. For example, various sound source data such as humming can be specified as input data to the music generation process 401. A sound / music list display section 612 shows a list of sound source data that has already been added. Clicking the play button for each sound source data listed in the sound / music list display section 612 plays the corresponding sound source, allowing the user to actually listen to it and select the sound or music to be input to the music generation process 401.

[0049] The UI screen shown in Fig. 6 also includes an add text button 603, and clicking this button allows the user to add text to be input to the music generation process 401. A text input section 613 is an area where the user can input and edit text specifying the content of the composition. In the example shown in Fig. 6, the following sentence is entered in the text input section 613: "Dance music mainly featuring vocals and synthesizers, BPM = 120, song length approximately 3 minutes." After entering text in the text input section 613, the user can click the add text button 603 to input the text to the music generation process 401.

[0050] 6 also includes a check box 604 for selecting whether or not to perform sound source separation (Export generated sound separately for each instrument), and a check box 605 for selecting whether or not to perform automatic music transcription (Export generated sound in MIDI file). By checking the check box 604, the user can select to perform sound source separation in which the music generated in the music generation process 401 is separated into individual instruments in the sound source separation process 402, and by checking the check box 605, the user can select to convert the WAVE file of the sound source data separated into individual instruments in the sound source separation process 402 into a MIDI file that can be edited element by element in the automatic music transcription process 403.

[0051] A Generate Music button 606 is located at the bottom of the UI screen shown in Fig. 6. By clicking the Generate Music button 606, the user can instruct the music generation process 401 using a machine learning model to start generating music based on the settings made on this UI screen.

[0052] The user editing using the UI screen shown in Fig. 6 is placed immediately before the automatic composition process 311 in the music production process shown in the lower part of Fig. 3 (not shown in Fig. 3). Via the UI screen shown in Fig. 6, the user can input multimodal information, including text data, sound source data such as humming, image data such as illustrations, sketches, and photographs, and video data, into the music generation process 401, and can also select whether to use the sound source separation and automatic transcription processes.

[0053] B-3. ​​Automatic Arrangement Process Fig. 7 shows an automatic arrangement process 701 using a machine learning model. In this automatic arrangement process 701, the input is unfinished demo sound source data that has been roughly composed and text data that explains the instructions for the desired arrangement, and the output is music data obtained by automatically arranging the unfinished demo sound source data that has been roughly composed based on the instructions in the text data, and is assumed to be music data such as MIDI or Wave.

[0054] In the automatic arrangement process 701, as with the music generation process 401, professional users with experience and knowledge in music production may have a need to make detailed adjustments to the sound source, such as making fine adjustments to each part of the arranged music, and there may be cases where a sound source separation process or automatic transcription process is required after the automatic arrangement process 701.

[0055] 8 shows an automatic arrangement process 800 for professional users, which adds a sound source separation process 702 and an automatic music transcription process 703 to an automatic arrangement process 701. The automatic arrangement process 701 inputs text data describing the desired arrangement and MIDI and Wave format sound source data for each instrument of a roughly composed, unfinished demo sound source, automatically arranges the music, and outputs music data in Wave format that mixes the sound source data for all instruments. The sound source separation process 702 separates the Wave format data (music file) output from the automatic arrangement process 701, which is a mix of the sound source data for all instruments, for each instrument, and outputs the sound source data for each instrument in Wave format. The automatic music transcription process 703 converts the Wave format sound source data for each instrument output from the sound source separation process 702 into MIDI format.

[0056] By applying the automatic arrangement process 800 shown in Figure 8 to the automatic arrangement process 321, the automatic arrangement process 321, which is one of the music production processes that uses the machine learning model shown in the lower part of Figure 3, can be made for users who have experience and knowledge about music production, i.e., professional users.

[0057] 7 and 8 , processes using models for downstream tasks trained based on the music-based model (i.e., models generated using the present disclosure) are shown in gray. The models used in the automatic music arrangement process 701, the sound source separation process 702, and the automatic music transcription process 703 can be generated with reduced resources and costs by adapting small amounts of training data and models corresponding to each downstream task (automatic music arrangement, sound source separation, and automatic music transcription) to a single music-based model. Furthermore, by reflecting the trends of a specific music genre or music created by the user in the music-based model itself, there is no need to retrain each downstream task model 701-703 using data from a specific genre or music. Instead, the output for the entire downstream task can be controlled to match the trends of a specific music genre or music created by the user.

[0058] FIG. 9 shows an example of a UI when a user controls an automatic musical arrangement process 701 using a machine learning model.

[0059] 9 includes a track selection button 901, and by clicking this button, the user can select tracks (voice tracks, tracks for individual instruments such as guitar or drums, etc.) to be processed by the automatic arrangement process 701. A track list display section 911 shows a list of tracks. By clicking the play button for each track listed in the track list display section 911, the corresponding track is played, allowing the user to actually listen to the track and select which tracks to input to the automatic arrangement process 701.

[0060] The UI screen shown in Fig. 9 also includes an add text button 902, and clicking this button allows the user to add text to be input to the automatic arrangement process 701. A text input section 912 is an area where the user can input and edit text specifying the arrangement content. In the example shown in Fig. 9, the following sentence is entered in the text input section 912: "To create an overall pop impression, I added a vocal chorus to the chorus and a guitar solo to the impression section. I also want to make the drum tone lighter overall." After entering text in the text input section 912, the user can input the text to the automatic arrangement process 701 by clicking the add text button 902.

[0061] 9 also includes a check box 903 for selecting whether to perform sound source separation, and a check box 904 for selecting whether to perform automatic transcription. By checking the check box 903, the user can select to perform sound source separation, in which the music arranged by the automatic arrangement process 701 is separated into individual instruments by the sound source separation process 702, and by checking the check box, the user can select to convert the WAVE files of the sound source data separated into individual instruments by the sound source separation process 702 into MIDI files that can be edited element by element by the automatic transcription process 703.

[0062] An Arrange button 905 is located at the bottom of the UI screen shown in Fig. 9. By clicking the Start Music Generation button 905, the user can instruct the automatic arrangement process 701 using a machine learning model to start automatic arrangement based on the settings made on this UI screen.

[0063] User editing using the UI screen shown in Fig. 9 can be used in the user editing process 312 located immediately before the automatic arrangement process 321 in the music production process shown in the lower part of Fig. 3. Via the UI screen as shown in Fig. 9, the user can input the sound source data to be arranged and text information indicating the arrangement content to the automatic arrangement process 321, and can also select whether or not to use the sound source separation and automatic transcription processes.

[0064] B-4. Automatic Mixing Process Fig. 10 shows an automatic mixing process 1000 using a machine learning model. The automatic mixing process 1000 corresponds to the automatic mixing process 331 in the music production process that uses a machine learning model shown in the lower part of Fig. 3.

[0065] 10 includes an equalizer 1001 that controls sound frequency, a compressor 1002 that controls sound pressure, a reverb 1003 that controls tone color and reverberation, and a volume 1004 that controls volume. Each of the processes 1001 to 1004 is MIDI-format and Wave-format sound source data for each instrument after automatic arrangement, and text data that explains instructions on how mixing is desired in each process, and the output is MIDI-format and Wave-format sound source data for each instrument after each process has been executed based on the instructions in the text data, and parameters such as the equalizer, compressor, reverb, and volume that control sound frequency, sound pressure, tone color, reverberation, etc.

[0066] In Figure 10, processes using models for downstream tasks trained based on the music-based model (i.e., models generated using the present disclosure) are shown in gray. The models used in the equalizer process 1001, compressor process 1002, reverb process 1003, and volume process 1004 can be generated with reduced resources and costs by adapting small amounts of training data and models corresponding to each downstream task (equalizer, compressor, reverb, and volume) to a single music-based model. Furthermore, by reflecting the trends of a specific music genre or music created by the user in the music-based model itself, there is no need to retrain each downstream task model 1001-1004 using data from a specific genre or music. Instead, the output for the entire downstream task can be controlled to match the trends of a specific music genre or music created by the user.

[0067] It is also possible that the order of some of the processes 1001-1004 in the automatic mixing process 1000 may be reversed. It is also possible that the automatic mixing process 1000 does not include some of the processes 1001-1004, or that it includes processes other than those shown. For processes not shown in FIG. 10 , models can be generated with reduced resources and costs by applying a small amount of training data and model corresponding to the downstream task to the same music-based model. Furthermore, by reflecting the trends of a specific music genre or music created by the user in the music-based model itself, the output can be controlled to match the trends of a specific music genre or music created by the user, just like models for other downstream tasks.

[0068] FIG. 11 shows an example of a UI for a user to control an automatic mixing process 1001 using a machine learning model.

[0069] 11 includes a track selection button 1101, and by clicking this button, the user can select tracks (such as audio tracks, tracks for individual instruments such as guitar or drums, etc.) to be processed by the automatic mixing process 1001. A track list display section 1111 shows a list of tracks. By clicking the play button for each track listed in the track list display section 1111, the corresponding track is played, allowing the user to actually listen to the tracks and select which tracks to input to the automatic mixing process 1001.

[0070] 11 also includes an add text button 1102, which can be clicked to add text to be input to the automatic mixing process 1001. A text input section 1112 is an area where the user can input and edit text specifying the mixing content. In the example shown in FIG. 11 , the following sentence is entered in the text input section 1112: "Make sure the vocals in the chorus aren't buried. Adjust the guitar interlude so that it stands out. Make sure the volume of the intro and outro isn't too quiet. I want the sound to sound like it's being performed in a slightly larger live music venue." After entering text in the text input section 1112, the user can click the add text button 1102 to input the text to the automatic mixing process 1001.

[0071] An automatic mixing start (Mix) button 1103 is located at the bottom of the UI screen shown in Fig. 11. By clicking the automatic mixing start button 1103, the user can instruct the automatic mixing process 1001 using a machine learning model to start automatic arrangement based on the settings made on this UI screen.

[0072] User editing using the UI screen shown in Fig. 11 can be used in the user editing process 322 located immediately before the automatic mixing process 331 in the music production process shown in the lower part of Fig. 3. The user can input the instrument data for each instrument to be mixed and text information specifying the mixing content to the automatic mixing process 331 via the UI screen shown in Fig. 11.

[0073] B-5. Controlling the Entire Music Production Process The learning method disclosed herein facilitates control of the entire downstream task. In other words, by incorporating the characteristics of a specific music genre or the music created by the user into the music-based model itself, it is possible to control the output of the entire downstream task to match the characteristics of the specific music genre or the music created by the user.

[0074] FIG. 12 shows an example of a UI when a user reflects a specific music genre or the tendencies of music that the user has created in the music base model itself.

[0075] The UI screen shown in FIG. 12 includes a song selection button 1201, and the user can click this button to select a song whose tendency is to be reflected in the music infrastructure model itself. A song list display section 1211 shows a list of songs to be selected. In the example shown in FIG. 12, the song list display section 1211 lists the titles and artist names of three candidate songs. Clicking the play button for each song listed in the song list display section 1211 plays the corresponding song, allowing the user to actually listen to the song and select the song whose tendency is to be reflected in the music infrastructure model itself. The song list display section 1211 can be used to instruct the music infrastructure model to tune files in the style of a specific artist.

[0076] The UI screen shown in FIG. 12 also includes a music genre list display section 1212 that displays a list of music genres whose tendencies can be reflected in the music infrastructure model itself. In the example shown in FIG. 12 , the music genre list display section 1212 displays a total of nine music genres: “Classic,” “Folk,” “Electric / Dance,” “Jazz,” “Rock,” “Hiphop,” “Pop,” “Funk / Soul / R&B,” and “Country.” The user can click and select one or more music genres from the music genre list display section 1212 whose tendencies the user wants to reflect in the music infrastructure model itself. The example shown in FIG. 12 shows two music genres, “Funk / Soul / R&B” and “Electric / Dance,” selected. In addition to music genre, the user may also be able to specify other tendencies, such as era, song tempo, and mood, as tendencies to be reflected in the music infrastructure model.

[0077] At the bottom of the UI screen shown in Fig. 12 is a start fine-tune model button 1202. By clicking the start fine-tune model button 1202, the user can instruct the start of fine-tuning based on the settings made on this UI screen for the music-based model used for training the entire downstream task.

[0078] Fine-tuning of the music-based model using the UI screen shown in FIG. 12 can be performed at the start of the music production process, or at any time during the music production process. Fine-tuning of the music-based model can be performed, for example, using Low-Rank Adaptation (LoRA) (see Non-Patent Document 1). LoRA is a technique that fixes the weights of an existing music-based model and injects a trainable rank decomposition matrix into each transformer layer. The rank decomposition matrix serves to approximate the calculations performed in each layer of the original model, and this approximation significantly reduces the number of trainable parameters in downstream tasks. However, fine-tuning of the music-based model is not necessarily limited to LoRA.

[0079] B-6. Summary According to the first embodiment of the present disclosure, it is possible to build an integrated music production platform using a music-based model, as shown in the lower part of Fig. 3. This disclosure makes it possible to train models for various downstream tasks using a small amount of training data, without imposing a significant burden on resources and costs by collecting large amounts of training data and building and training models for each individual function, such as automatic composition, automatic arrangement, and automatic mixing.

[0080] Furthermore, according to the first embodiment of the present disclosure, the models of each downstream task for realizing an individual application use common intermediate features in a single music-based model, so it is expected that consistency in behavior will be maintained.

[0081] Furthermore, according to the first embodiment of the present disclosure, by adapting the music-based model itself to a specific music genre or the tendencies of the music created by the user himself, it is possible to control the output of the entire downstream task collectively to be tailored to the tendencies of a specific music genre or the music created by the user himself.

[0082] C. Second Example In this section C, as a second example of the present disclosure, an example of automatically generating music such as background music (BGM) to be attached to user generated content (UGC) using AI technology will be introduced.

[0083] C-1. Overview When users upload content they have created themselves, such as text, photos, or videos, many social media platforms have their own media-specific contexts, such as their own specifications for image quality, video length, field of view, file size, etc., and when users upload content, it is automatically edited and converted to fit the specifications of each platform.

[0084] Additionally, some social media platforms have implemented a function that allows users to add background music to content they upload, allowing users to search for existing song titles or artist names, or to use music recommended by the platform. However, because social media platforms do not have a music editing function, there are many cases where the music selected as background music is adjusted to fit the length of the content uploaded by the user, resulting in the music being cut off in an unnatural way or the music being played at an unnaturally fast (or slow) speed.

[0085] Therefore, in the present disclosure, as a second embodiment, a technology is proposed that enables even a user with no experience in music editing to automatically edit background music to match the specifications of the social media platform to which the content is uploaded and the length of the content, without performing any special editing work.

[0086] C-2. Basic Configuration Fig. 13 is a flowchart showing an outline of the process of editing background music for content using AI technology.

[0087] The background music for the content can be broadly divided into two cases: when a new piece of music is used and when an existing piece of music is used (step S1301). When a new piece of music is used (Yes in step S1301), a new piece of music is automatically composed using a machine learning model (step S1302). On the other hand, when an existing piece of music is used (No in step S1301), a piece of music is searched from among the existing pieces using the machine learning model (step S1303).

[0088] Furthermore, when using an existing song, there is a case where a remix is ​​performed to arbitrarily arrange the song (Yes in step S1304), and an automatic remix is ​​performed on the existing song using a machine learning model (step S1305).

[0089] Then, in either the case of automatically generating a new song or automatically selecting an existing song, the song is automatically edited using a machine learning model so that it conforms to the specifications of the social media platform to which it is uploaded (step S1306).

[0090] In the flowchart shown in Figure 13, tasks to which machine learning models are applied are indicated in gray. When developing applications for tasks related to background music editing of content (i.e., music generation, music search, remixing, editing), model developers need to collect large amounts of training data for each task, design a model with an appropriate architecture, and perform training. This presents a challenge, resulting in enormous development resources and costs. Furthermore, if it is desired to reflect the preferences of users of the social media platform to which the content is uploaded in applications for each task related to background music editing of content, the individual models corresponding to each task need to be retrained using the music information to be reflected.

[0091] Therefore, in a second embodiment of the present disclosure, instead of performing model training individually for each task related to background music editing of content, a music-based model and smaller amounts of training data corresponding to each task are used to realize many types of models as downstream tasks with fewer resources and at fewer costs.

[0092] FIG. 14 shows a learning method applied in a second embodiment of the present disclosure. In FIG. 14, downstream tasks are music generation, music search, automatic remixing, and automatic editing, which are tasks related to background music editing. For each downstream task, a music generation model 1411, a music search model 1412, an automatic remixing model 1413, and an automatic editing model 1414 are prepared. The models for each downstream task may have an architecture appropriate for the respective task. In other words, the architecture of the models for each downstream task may be different.

[0093] In a second embodiment of the present disclosure, the music foundation model 1401 is adapted to small amounts of training data and models corresponding to the above downstream tasks, thereby reducing the resources and costs required for training the music generation model 1411, music search model 1412, automatic remix model 1413, and automatic editing model 1414. Adaptation is the process of adapting the foundation model to a specific task. Adaptation allows the models for the downstream tasks to acquire knowledge appropriate for a specific domain or task.

[0094] The method for adapting each downstream task to the music-based model 1401 is as described in Section A above with reference to Figure 2. Adapting each downstream task to the music-based model 1401 corresponds to applying (inputting) the common intermediate features extracted by the music-based model 1401 to the music generation model 1411, music search model 1412, automatic remix model 1413, and automatic editing model 1414 for each downstream task. LoRA (see Non-Patent Document 1) can be cited as an example of a method for inputting the intermediate features of the foundation model to the downstream tasks, but the second embodiment of the present disclosure is not limited to this method.

[0095] The intermediate features extracted by the music-based model 1401 include features that express the tendencies of the music. Therefore, in the present disclosure, by performing learning using common intermediate features of the music-based model 1401 in each downstream task, it is expected that consistency in the behavior of the music generation model 1411, music search model 1412, automatic remix model 1413, and automatic editing model 1414 for each downstream task can be maintained.

[0096] Furthermore, the learning method according to the present disclosure makes it easy to control downstream tasks as a whole. For example, by reflecting the taste trends of the user demographic of the social media platform to which content is uploaded in the music-based model 1401 itself, it is possible to eliminate the need to individually retrain the music generation model 1411, music search model 1412, automatic remix model 1413, and automatic editing model 1414 for each downstream task, and instead control the output of all downstream tasks collectively to match the taste trends of a specific user demographic.

[0097] FIG. 15 shows the operation of adapting a small training dataset 1421 for music generation and a music generation model 1411 to a music-based model 1401. The small training dataset 1421 includes music data and music information (such as era, genre, tempo, and mood) of the social media platform to which the music is uploaded.

[0098] FIG. 16 also shows the operation of adapting a small training dataset 1422 for music search and a music search model 1412 to the music-based model 1401. The small training dataset 1422 includes music data and music information (such as era, genre, tempo, and mood) of the social media platform to which the data is uploaded.

[0099] 17 also shows the operation of adapting a small amount of training data set 1423 for automatic remixing and an automatic remix model 1413 to the music-based model 1401. The small training data set 1423 includes music before remixing, music after remixing, and music genre information of the social media platform to which the music is uploaded. The remixing process may include, but is not limited to, multiple processes such as adding new musical elements, adding effects such as reverb, and restructuring the structure of the music.

[0100] 18 also shows an operation of adapting a small amount of training dataset 1424 for automatic editing and an automatic editing model 1414 to the music-based model 1401. The small training dataset 1424 includes analytical data such as hit song information and user generation information for each social media platform, and specifications or context of background music (such as the length of the song). The automatic editing model 1414 generated by the operation shown in FIG. 18 can adapt automatically generated or automatically searched music to the context of the social media platform to which it is uploaded, or edit it to suit the popularity of hit songs, user demographics (age groups), etc.

[0101] C-3. Background Music Editing Process Using Machine Learning C-3-1. Automatic Generation of Background Music Figure 19 shows a process for automatically generating background music to be added to content uploaded by a user to a social media platform. The illustrated automatic background music generation process uses a music generation model 1411 and an automatic editing model 1414, which were generated using common intermediate features of the same music-based model 1401 through the operations shown in Figures 15 and 18, respectively.

[0102] The music generation model 1411 automatically generates music suitable for background music for content by inputting content created by the user, such as text (captions), photos, and videos to be uploaded to a social media platform, and tags containing instructions from the user regarding the generation of background music. Note that whether or not tags are used for the automatic generation of background music (in other words, whether or not tags are input into the music generation model 1411) is optional.

[0103] The automatic editing model 1414 adapts the background music automatically generated by the music generation model 1411 to the context of the social media platform to which the music is uploaded, and edits it to match the hit songs or user demographics (age groups) of the social media platform to which the music is uploaded.

[0104] The background music edited by the automatic editing model 1414 is then finalized and uploaded to a social media platform together with the content, for example, to a streaming service with background music.

[0105] 20 shows an example of a UI for automatically generating background music when a user uploads content. This UI screen is displayed assuming that the user has already specified the social media platform to which the content will be uploaded and the content to be uploaded.

[0106] The UI screen shown in the figure is, for example, the screen of a smartphone, and is arranged with a content display section 2201 that displays content such as images and videos to be uploaded, a slide bar 2202 that displays the playback position of the content (only if the content is a video), and a caption display area 2203 that displays text information such as captions that will be uploaded along with the content.

[0107] This UI screen also has an option specification section 2204 for specifying option information related to the automatic generation of background music for the content displayed in the content display section 2201. The option specification section 2204 displays buttons for specifying the genre and tempo of the music when the background music is automatically generated, as well as text information such as tags containing instructions from the user related to the generation of the background music (in the example shown in FIG. 20, the text displayed is "relaxing, soothing music").

[0108] The UI screen shown in Fig. 20 may be, for example, a screen of a smartphone on which a user performs an operation to upload content, a UI screen provided by a specific social media platform, or a screen of an editing app for editing background music for content across multiple social media platforms.

[0109] The automatic background music generation process shown in Fig. 19 may be performed entirely on the smartphone of the user who uploads the content (or on an information terminal that performs the content upload operation). Alternatively, only the UI screen shown in Fig. 20 is operated on the user's smartphone (or on an information terminal), and the automatic background music generation process may be performed on the social media platform to which the content is uploaded, or on a server other than the upload destination.

[0110] C-3-2. Automatic Search for Background Music Figure 21 shows a process for automatically searching for background music to be added to content uploaded by a user to a social media platform. The illustrated automatic background music generation process uses a music search model 1412 and an automatic editing model 1414 that were generated using common intermediate features of the same music-based model 1401 through the operations shown in Figures 16 and 18, respectively.

[0111] The music search model 1412 receives input of content created by the user, such as text (captions), photos, and videos to be uploaded to a social media platform, and tags including instructions from the user regarding the generation of background music, and automatically searches for music suitable as background music for the content from among the existing songs stored in the music database 2101. Note that whether or not to use tags for the automatic search for background music (in other words, whether or not to input tags into the music search model 1412) is optional.

[0112] The automatic editing model 1414 adapts the background music automatically searched by the music search model 1412 to the context of the social media platform to which the music is uploaded, or edits it to match the hit songs or user demographic (age group) of the social media platform to which the music is uploaded.

[0113] The background music edited by the automatic editing model 1414 is then finalized and uploaded to a social media platform together with the content, for example, to a streaming service with background music.

[0114] 22 shows an example of a UI for uploading content when background music is automatically searched for. This UI screen is displayed assuming that the user has already specified the social media platform to which the content is to be uploaded and the content to be uploaded.

[0115] The UI screen shown in the figure is, for example, the screen of a smartphone, and is arranged with a content display section 2201 that displays content such as images and videos to be uploaded, a slide bar 2202 that displays the playback position of the content (only if the content is a video), and a caption display area 2203 that displays text information such as captions that will be uploaded along with the content.

[0116] This UI screen also includes an option designation section 2205 for designating option information related to an automatic search for background music for the content displayed in the content display section 2201. The option designation section 2205 displays a list of background music candidates (recommended music for you) automatically searched from the music database 2101 and recent hit songs. In the example shown in FIG. 22, two background music candidates and one recent hit song are displayed. The user can play and listen to the song by clicking the play button for each candidate song displayed in the option designation section 2205 (in the example shown in FIG. 22, the first background music candidate is being played). Checking the checkbox for each candidate song allows the song to be designated as background music for the content.

[0117] The UI screen shown in Fig. 22 may be, for example, a screen of a smartphone on which a user performs an operation to upload content, a UI screen provided by a specific social media platform, or a screen of an editing app for editing background music for content across multiple social media platforms.

[0118] The automatic background music search process shown in Fig. 21 may be performed entirely on the smartphone of the user who uploads the content (or on an information terminal that performs the content upload operation). Alternatively, only the UI screen shown in Fig. 22 is operated on the user's smartphone (or on an information terminal), and the automatic background music search process may be performed on the social media platform to which the content is uploaded, or on a server other than the upload destination.

[0119] C-3-3. Automatic Search for Background Music and Automatic Remixing Fig. 23 shows a process for automatically searching for and automatically remixing background music to be added to content uploaded by a user to a social media platform. The illustrated automatic background music generation process uses a music search model 1412, an automatic remix model 1413, and an automatic editing model 1414, which are generated using common intermediate features of the same music-based model 1401 through the operations shown in Figs. 16, 17, and 18, respectively.

[0120] The music search model 1412 receives input of content created by the user, such as text (captions), photos, and videos to be uploaded to a social media platform, and tags including instructions from the user regarding the generation of background music, and automatically searches for music suitable as background music for the content from among the existing songs stored in the music database 2101. Note that whether or not to use tags for the automatic search for background music (in other words, whether or not to input tags into the music search model 1412) is optional.

[0121] The automatic remix model 1413 automatically remixes the background music automatically searched for by the music search model 1412. The remix process may include, but is not limited to, a plurality of processes such as adding new musical elements, adding effects such as reverb, and restructuring the structure of the music.

[0122] The automatic editing model 1414 adapts the background music after automatic remixing by the automatic remix model 1413 to the context of the social media platform to which the music is uploaded, or edits it to suit the hit songs or user demographic (age group) of the social media platform to which the music is uploaded.

[0123] The background music edited by the automatic editing model 1414 is then finalized and uploaded to a social media platform together with the content, for example, to a streaming service with background music.

[0124] 24 shows an example of a UI for uploading content when background music is automatically searched and automatically remixed. This UI screen is displayed assuming that the user has already specified the social media platform to which the content will be uploaded and the content to be uploaded.

[0125] The UI screen shown in the figure is, for example, the screen of a smartphone, and is arranged with a content display section 2201 that displays content such as images and videos to be uploaded, a slide bar 2202 that displays the playback position of the content (only if the content is a video), and a caption display area 2203 that displays text information such as captions that will be uploaded along with the content.

[0126] This UI screen also includes an option designation section 2206 for designating option information related to the automatic search and automatic remixing of background music. The option designation section 2206 displays a list of background music candidates (recommended music) automatically searched from the music database 2101 and remix options to be applied to the background music candidates. In the example shown in FIG. 24 , three remix options are displayed: "80's Remix," "Funk Remix," and "Acoustic Remix." By clicking the play button for each of the background music candidates and remix option candidates displayed in the option designation section 2206, the user can play and listen to the background music candidate songs and the remixed candidate songs.

[0127] The UI screen shown in Fig. 24 may be, for example, a screen of a smartphone on which a user performs an operation to upload content, a UI screen provided by a specific social media platform, or a screen of an editing app for editing background music for content across multiple social media platforms.

[0128] The automatic background music search and automatic remix process shown in Fig. 23 may be performed entirely on the smartphone of the user uploading the content (or on an information terminal performing the content upload operation). Alternatively, only the UI screen shown in Fig. 24 is operated on the user's smartphone (or information terminal), and the automatic background music search and automatic remix process may be performed on the social media platform to which the content is uploaded, or on a server other than the upload destination.

[0129] C-4. Integrated UI Screen Figure 25 shows an example of a UI for integrating the background music automatic generation, automatic search, and automatic editing functions. This UI screen is displayed assuming that the user has already specified the social media platform to which the content will be uploaded and the content to be uploaded.

[0130] The UI screen shown in the figure is, for example, the screen of a smartphone, and is arranged with a content display section 2501 that displays content such as images and videos to be uploaded, a slide bar 2502 that displays the playback position of the content (only if the content is a video), and a caption display area 2503 that displays text information such as captions that will be uploaded along with the content.

[0131] This UI screen also includes an option designation section 2504 for designating option information related to the automatic generation of background music.

[0132] The option designation section 2504 includes an option information designation section 2511 for automatically generating background music. The option information designation section 2511 displays buttons for designating the genre and tempo of the music to be used when automatically generating background music, as well as text information such as tags containing user instructions regarding the generation of background music (in the example shown in FIG. 25 , the text displayed is "relaxing, soothing music"). When a music generation start button 2505 in the option information designation section 2511 is clicked, background music is automatically generated based on the content designated in the option information designation section 2511.

[0133] The option designation section 2504 also includes an option information designation section 2512 for automatically searching and remixing background music. The option information designation section 2512 displays a list of background music candidates (recommended music) automatically searched from the music database 2101 and recent hit songs. In the example shown in Fig. 25, two background music candidates and one recent hit song are displayed. The user can play and listen to the music by clicking the play button for each candidate song displayed in the option information designation section 2512.

[0134] A list of remix options to be applied to background music is displayed as option information related to automatic remixing of background music candidates in the option information designation section 2512. In the example shown in Fig. 25, three types of remix options are displayed for the first background music candidate: "80's Remix," "Funk Remix," and "Acoustic Remix."

[0135] By clicking the play button for each of the background music candidates and remix option candidates in the option information designation section 2512, the user can play and listen to the background music candidate songs and the remixed candidate songs.

[0136] The UI screen shown in Fig. 25 may be, for example, a screen of a smartphone on which a user performs an operation to upload content, a UI screen provided by a specific social media platform, or a screen of an editing app for editing background music for content across multiple social media platforms.

[0137] The processes of automatic generation, automatic search, and automatic remixing of background music may all be performed on the smartphone of the user who uploads the content (or on an information terminal that performs the content upload operation). Alternatively, only the UI screen shown in Fig. 25 is operated on the user's smartphone (or information terminal), and the processes of automatic generation, automatic search, and automatic remixing of background music may be performed on the social media platform to which the content is uploaded, or on a server other than the upload destination.

[0138] 20, 22, 24, and 25 have the advantage that even users who do not have sufficient musical knowledge, such as composition, can generate, arrange, and edit background music to be added to content they have created, such as text, photos, and videos. Therefore, even general users who do not have specialized knowledge of BPM (Beats Per Minute) or chords can create music (background music to be added to content, etc.) by selecting their own uploaded content, text, simple tags, and providing emotional instructions (comforting, exciting, calming, emotional, etc.).

[0139] C-5. Summary According to the second embodiment of the present disclosure, a music-based model adapted to media-specific contexts, such as specifications for uploading background music for each social media platform, and trends and age groups of users of each social media platform, can be incorporated into models for each downstream task, such as music generation, music search, automatic remixing, and automatic editing. Therefore, even users who do not have sufficient musical knowledge, such as composition, can add background music appropriate to the content they have created, such as text, photos, or videos, and upload them to a social media platform without having to perform detailed editing work.

[0140] Furthermore, according to a second embodiment of the present disclosure, a model is generated for each downstream task, using common intermediate features in a single music-based model to automatically generate, or automatically search for and remix, music from content (text, photos, videos, etc.) uploaded by a user, and edit the music to fit the context of the upload destination. Therefore, a model with sufficient performance can be trained with a smaller training dataset and a smaller model size than when models for each downstream task are trained individually.

[0141] D. Configuration of Information Processing Device FIG. 26 shows an example hardware configuration of an information processing device 2000 applied to the present disclosure. The information processing device 2000 can be used, for example, in a process of generating a model for an entire downstream task using common intermediate features of a music-based model. Furthermore, in a first embodiment, the information processing device 2000 can be used to realize at least some or all of the tasks of automatic composition, automatic arrangement, and automatic mixing, each of which uses a machine learning model. Furthermore, in a second embodiment, the information processing device 2000 can be used to realize at least some or all of the tasks of music generation, music search, automatic remixing, and automatic editing included in the editing process of background music for content.

[0142] This information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by a personal computer, but some of its functions may be configured by an information terminal such as a tablet or a smartphone.

[0143] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which is composed of a CPU bus and other components. The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. The ROM 2002 stores programs (such as a basic input / output system) and calculation parameters used by the CPU 2001 in a non-volatile manner. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as work data that change as appropriate during program execution. The CPU 2001 can execute various application programs in an execution environment provided by an operating system (OS) through the cooperative operation of the ROM 2002 and RAM 2003, thereby realizing a variety of functions and services.

[0144] If the information processing device 2000 is a PC, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. The programs loaded into the RAM 2003 and executed by the CPU 2001 include the OS and various application programs. Note that the application programs or some modules in the application programs may use existing libraries that are stored, shared, or made public through, for example, a source code management service. For example, at least one of the following programs (1) to (3) is executed on the information processing device 2000:

[0145] (1) A processing program for generating a model for the entire downstream task by utilizing common intermediate features of a music-based model. (2) In a first embodiment of the present disclosure, a processing program for realizing at least some or all of the tasks of automatic composition, automatic arrangement, and automatic mixing, each using a machine learning model. (3) In a second embodiment of the present disclosure, a processing program for realizing at least some or all of the tasks of music generation, music search, automatic remixing, and automatic editing, which are included in the editing process of background music for content.

[0146] When performing computationally intensive processing such as learning an AI (Artificial Intelligence) model on the information processing device 2000, it is desirable that the CPU 2001 be a multi-core CPU (for example, Apple M1 Max, etc.), and that the information processing device 2000 further be equipped with a multi-core processor such as a GPU or GPGPU (General-purpose computing on graphics processing units) (for example, NVIDIA's "Quadro A6000"). However, for convenience, these will be collectively referred to as the CPU 2001 below.

[0147] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).

[0148] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 26 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.

[0149] The input unit 2008 is composed of an input control circuit that generates an input signal based on an input from a user and outputs the signal to the CPU 2001. When the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, a mouse, a touch panel, a camera, and a microphone. The output unit 2009 includes display devices such as a liquid crystal display (LCD) device, an organic electroluminescence (EL) display device, and an LED (light emitting diode), as well as an audio output device such as a speaker. The input unit 2008 is used to input a video to be processed, and the output unit 2009 is used to display a GUI screen, etc.

[0150] For example, when the information processing device 2000 is applied to the first embodiment of the present disclosure, control of each process by the user via a UI in at least one of the user edits 312, 322, and 332 included in the music production process can be realized using the input unit 2008 and the output unit 2009. Furthermore, when the information processing device 2000 is applied to the second embodiment of the present disclosure, UI operations shown in at least one of Figures 20, 22, 24, and 25 can be realized using the input unit 2008 and the output unit 2009, and background music to be added to content such as text, photos, and videos created by the user can be generated, arranged, edited, and the like.

[0151] The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.

[0152] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 2012. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.

[0153] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark), and cellular communication networks such as 4G and 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) and a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, etc. Programs executed on the information processing device 2000 are installed from the outside, for example, via the communication unit 2013.

[0154] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.

[0155] Although the present specification has mainly described embodiments in which the present disclosure is applied to music-related processing, the gist of the present disclosure is not limited thereto. For example, the present disclosure can also be applied to processes such as the production and editing of various content, such as images, videos, and text. According to the present disclosure, an integrated content production platform can be built using a single platform model. According to the present disclosure, it is possible to avoid the significant resource and cost burden of collecting large amounts of training data individually for multiple functions related to content production to build and train models.

[0156] According to the present disclosure, models of downstream tasks for realizing individual functions use common intermediate features in a single base model, so consistency in behavior can be expected. Furthermore, according to the present disclosure, by adapting the base model itself to a specific genre or the tendencies of content created by the user, it is possible to collectively control the output of all downstream tasks to match the tendencies of a specific genre or content created by the user.

[0157] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.

[0158] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.

[0159] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc (registered trademark)), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to the realization of the present disclosure can be provided as so-called package software.

[0160] The program may also be transferred wirelessly or via a wire from a download site to a computer via a network such as a wide area network (WAN) typified by cellular, a local area network (LAN), the Internet, etc. The computer can receive the program transferred in this manner and install it in a large-capacity storage device such as an HDD or SSD within the computer.

[0161] The present disclosure may also be configured as follows.

[0162] (1) An information processing system comprising: an acquisition unit that acquires a trained music-based model; and a generation unit that generates a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

[0163] (2) The information processing system according to any one of (1) or (2), wherein the generation unit generates a plurality of trained models for each downstream task based on the music-based model and a plurality of training datasets corresponding to the plurality of downstream tasks.

[0164] (3) The information processing system according to (2), wherein the generation unit applies common intermediate features obtained from the music-based model to models for each downstream task to generate multiple trained models for each downstream task.

[0165] (4) The information processing system according to any one of (2) or (3), wherein the generation unit generates a trained model for each downstream task performed in each process of music production.

[0166] (5) The information processing system described in (4) above, wherein the music production process includes at least one of composing, arranging, and mixing, and the generation unit generates a learning model adapted to downstream tasks performed in at least one of the processes of composing, arranging, and mixing.

[0167] (6) The information processing system according to (5), wherein the downstream tasks include at least one of automatic composition, sound source separation, automatic music transcription, and music tagging, and the trained model generated by the generation unit for each downstream task is used at least in the composition process.

[0168] (7) The information processing system according to (5), wherein the downstream tasks include at least one of automatic music arrangement, sound source separation, automatic music transcription, automatic mixing, and music tagging, and the trained model generated for each downstream task by the generation unit is used at least in the music arrangement process.

[0169] (8) The information processing system according to (5), wherein the downstream task includes at least one of an equalizer, a compressor, a reverb, and a volume, and the trained model generated for each downstream task by the generation unit is used at least in the mixing process.

[0170] (9) The information processing system according to any one of (1) to (8) above, further comprising an adjustment unit that adjusts the music-based model based on at least one of a specified song and music genre, a music genre, an era, a tempo of the song, and a mood.

[0171] (10) The information processing system described in any one of (1) to (3) above, wherein the downstream tasks include at least one of music generation for generating music to be assigned to data, music search for searching for music to be assigned to data, music remixing, and music editing for adapting music to a predetermined context, and further includes a model acquisition unit for acquiring the trained model generated by the generation unit.

[0172] (11) The information processing system according to (10) above, wherein the model acquisition unit acquires a trained music generation model; the information processing system includes a data acquisition unit that acquires data to be uploaded to external media; and the information processing system generates a piece of music to be assigned to the data using the music generation model.

[0173] (12) The information processing system according to (11), wherein the music based on the specified option information is generated using the music generation model.

[0174] (13) The information processing system according to (10) above, wherein the model acquisition unit acquires a trained music search model, and the information processing system includes a data acquisition unit that acquires data to be uploaded to external media, and searches for songs to be assigned to the data using the music search model.

[0175] (14) The information processing system according to (13) above, wherein the music based on the specified option information is searched for using the music search model.

[0176] (15) The information processing system according to any one of (13) or (14), wherein the model acquisition unit further acquires a trained remix model generated by the generation unit, and remixes the searched music piece using the remix model.

[0177] (16) The information processing system according to (15) above, wherein the remix model is used to remix the searched music piece based on the specified option information.

[0178] (17) The information processing system according to any one of (11) to (16), wherein the model acquisition unit further acquires a trained music editing model, and uses the music editing model to edit the music to be assigned to the data so as to fit the context of the external media.

[0179] (18) The information processing system according to any one of (10) to (17), wherein the data includes at least one of text data, image data, and video data.

[0180] (19) An information processing method comprising: an acquisition step of acquiring a trained music-based model; and a generation step of generating a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

[0181] (20) A computer program written in a computer-readable format to cause a computer to function as: an acquisition unit that acquires a trained music-based model; and a generation unit that generates a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

[0182] (31) An information processing system comprising: a data acquisition unit that acquires data to be uploaded to external media; and a music generation unit that generates music from the data that fits the context of the external media using a music generation model generated based on a trained music-based model.

[0183] (32) The information processing system according to (31) above, further comprising a reception unit that receives specifications regarding the external media, wherein the music generation unit generates music that matches the context of the external media received by the reception unit.

[0184] (33) The information processing system described in any one of (31) or (32) above, wherein the music generation model is a model generated based on the music-based model and a training dataset for a downstream task of generating music from data.

[0185] (34) The information processing system according to any one of (31) to (33), wherein the data includes at least one of text data, image data, and video data.

[0186] (35) The information processing system according to any one of (31) to (34), wherein the data acquisition unit further acquires tag information, and the music generation model generates the music from the data and the tag information.

[0187] 100...Music-based model, 111...Music generation model, 112...Sound source separation model, 113...Automatic music transcription model, 114...Automatic arrangement model, 115...Automatic mix model, 116...Music tagging model, 201...Encoder, 202...Decoder, 211...Intermediate features, 1401...Music-based model, 1411...Music generation model, 1412...Music search model, 1413...Automatic remix model, 1414...Automatic editing model, 1421...Learning dataset (for music generation), 1422...Learning dataset (for music search), 1423...Learning dataset (for automatic remix), 1424...Learning dataset (for automatic editing), 2000...Information processing device, 2001...CPU, 2002...ROM, 2003...RAM, 2004...Host bus, 2005...Bridge 2006... expansion bus, 2007... interface section, 2008... input section, 2009... output section, 2010... storage section, 2011... drive, 2012... removable recording medium, 2013... communication section

Claims

1. An information processing system comprising: an acquisition unit that acquires a trained music-based model; and a generation unit that generates a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

2. The information processing system according to claim 1, wherein the generation unit generates a plurality of trained models for each downstream task based on the music-based model and a plurality of training datasets corresponding to each downstream task.

3. The information processing system according to claim 2, wherein the generation unit applies common intermediate features obtained from the music-based model to models for each downstream task to generate multiple trained models for each downstream task.

4. The information processing system according to claim 2, wherein the generation unit generates a trained model for each downstream task performed in each process of music production.

5. The information processing system according to claim 4, wherein the music production process includes at least one of composing, arranging, and mixing, and the generation unit generates a learning model adapted to downstream tasks performed in at least one of the processes of composing, arranging, and mixing.

6. The information processing system according to claim 5, wherein the downstream tasks include at least one of automatic composition, sound source separation, automatic music transcription, and music tagging, and the trained model generated for each downstream task by the generation unit is used at least in the composition process.

7. The information processing system of claim 5, wherein the downstream tasks include at least one of automatic music arrangement, sound source separation, automatic music transcription, automatic mixing, and music tagging, and the trained model generated for each downstream task by the generation unit is used at least in the music arrangement process.

8. The information processing system according to claim 5, wherein the downstream task includes at least one of an equalizer, a compressor, a reverb, and a volume, and the trained model generated for each downstream task by the generation unit is used at least in the mixing process.

9. The information processing system according to claim 1, further comprising an adjustment unit that adjusts the music-based model based on at least one of a specified song and music genre, and in addition to the music genre, an era, a tempo of the song, and a mood.

10. The information processing system of claim 1, wherein the downstream tasks include at least one of music generation for generating music to be assigned to data, music search for searching for music to be assigned to data, music remixing, and music editing for adapting music to a predetermined context, and further comprising a model acquisition unit for acquiring the trained model generated by the generation unit.

11. The information processing system according to claim 10, wherein the model acquisition unit acquires a trained music generation model; the information processing system includes a data acquisition unit that acquires data to be uploaded to external media; and the information processing system generates music to be assigned to the data using the music generation model.

12. The information processing system according to claim 11, wherein the music based on the specified option information is generated using the music generation model.

13. The information processing system according to claim 10, wherein the model acquisition unit acquires a trained music search model; the information processing system includes a data acquisition unit that acquires data to be uploaded to external media; and the information processing system uses the music search model to search for songs to be assigned to the data.

14. The information processing system according to claim 13, wherein the music based on the specified option information is searched for using the music search model.

15. The information processing system according to claim 13, wherein the model acquisition unit further acquires a trained remix model generated by the generation unit, and remixes the searched music piece using the remix model.

16. The information processing system according to claim 15, wherein the remix model is used to remix the searched music piece based on specified option information.

17. The information processing system according to claim 11, wherein the model acquisition unit further acquires a trained music editing model, and uses the music editing model to edit the music to be assigned to the data so as to fit the context of the external media.

18. The information processing system according to claim 10, wherein the data includes at least one of text data, image data, and video data.

19. An information processing method comprising: an acquisition step of acquiring a trained music-based model; and a generation step of generating a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

20. A computer program written in a computer-readable format to cause a computer to function as: an acquisition unit that acquires a trained music-based model; and a generation unit that generates a trained model adapted to the downstream task based on the music-based model and a training dataset related to the downstream task.

Citation Information

Cited By

  • Information processing method, information processing system, and program

    WO2026155052A1