Information processing system, information processing method, and storage medium
The information processing system addresses the lack of video generation control for music pieces by using a machine learning model to generate explanatory text data from multimodal inputs, enabling the creation of music videos with integrated visual and musical elements.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2025-11-10
- Publication Date
- 2026-07-30
AI Technical Summary
There is a need for a technology that can automatically generate a video corresponding to a music piece, as conventional methods lack a control mechanism for video generation based on music pieces.
An information processing system utilizing a machine learning model that receives multimodal data including music data, type of song, and lyric understanding to generate explanatory text data for creating a music video.
Enables the automatic generation of music videos by processing multimodal data to produce descriptive text data that can be used to create music videos, enhancing the integration of visual elements with musical features.
Smart Images

Figure JP2025039323_30072026_PF_FP_ABST
Abstract
Description
Information Processing System, Information Processing Method, and Storage Medium
[0001] The present disclosure relates to an information processing system, an information processing method, and a storage medium.
[0002] In recent years, especially in the music field such as pop music, it has been common to present a music piece together with a video corresponding to the music piece (referred to as a music video).
[0003] Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, Tat-Seng Chua, 'NExT-GPT: Any-to-Any Multimodal LLM', [ONLINE], September 11, 2023, [September 26, 2025], Internet <URL: https: / / arxiv.org / abs / 2309.05519>, English
[0004] There is a need for a technology that can automatically generate a video corresponding to a music piece, that is, a music video. However, since video generation corresponding to a music piece is not video generation by direct description of visual information, a control method for such video generation has not been established conventionally.
[0005] Therefore, an object of the present disclosure is to provide an information processing system, an information processing method, and a storage medium that can control the automatic generation of a video corresponding to a music piece.
[0006] The information processing system according to the present disclosure includes a processing circuit. The processing circuit receives first multimodal data related to first music data, and by inputting the first multimodal data into a machine learning model, outputs explanatory text data related to the first music data. The explanatory text data is used for generating first music video data.
[0007] This is a schematic diagram showing an example of an information processing system applicable to embodiments of this disclosure. This is a block diagram showing the configuration of an example of a server applicable to embodiments. This is an example of a functional block diagram for explaining the functions of a server applicable to embodiments of this disclosure. This is a schematic diagram for explaining the usage of a model according to an embodiment. This is a schematic diagram for explaining an example of a training dataset used for training a multimodal LLM according to an embodiment. This is a schematic diagram for explaining the training method of a multimodal LLM according to an embodiment. This is a schematic diagram for explaining the process of generating text data from non-text data according to an embodiment. This is a schematic diagram for explaining the method of creating ground truth data according to an embodiment. This is a schematic diagram showing another example of the method of creating integrated music captions according to an embodiment. This is a schematic diagram for explaining a model according to an embodiment. This is a schematic diagram showing an example of input and output to a trained LLM (multimodal LLM) according to an embodiment. This is a schematic diagram showing a more specific example of an MV description output from an LLM (multimodal LLM) according to an embodiment. This is a schematic diagram showing an example of an MV description as ground truth data according to an embodiment. This is a schematic diagram showing an example of evaluating processing results according to an embodiment. This is a schematic diagram showing an example of an original music video. This is a schematic diagram showing an example of video captions 7 generated based on an original MV, applicable to the embodiment. This is a schematic diagram showing an example of integrated music captions generated based on music data extracted from an original MV, applicable to the embodiment. This is a schematic diagram showing an example of lyric understanding, interpreting the lyrics corresponding to the original MV, applicable to the embodiment. This is a schematic diagram showing an example of an MV description generated using the model according to the embodiment, based on video captions, integrated music captions, and lyric understanding. This is a schematic diagram showing an example of a music video generated by the model according to the embodiment. This is a schematic diagram showing an example of a dataset consisting of four types of data: interrelated image data, video data, music data (audio data), and text data. This is a schematic diagram showing an example of a dataset consisting of four types of data: interrelated image data, video data, music data (audio data), and text data. This is a schematic diagram for explaining the generation of integrated music captions according to the embodiment.This is a schematic diagram illustrating, in more detail, the input data and instructions input to the model according to the embodiment, as well as an example of the MV description text generated by the model. This is a schematic diagram illustrating an example of video generation using the MV description text according to the embodiment. This is a schematic diagram illustrating an example of video generation using the MV description text according to the embodiment. This is a schematic diagram illustrating an example of video generation using the MV description text according to the embodiment.
[0008] The embodiments of this disclosure will be described in detail below with reference to the drawings. In the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.
[0009] The embodiments of this disclosure will be described below in the following order: 1. Overview of this disclosure 1-1. Information processing system applicable to the embodiments of this disclosure 2. Embodiments of this disclosure 2-1. Outline of the usage of the model according to the embodiment 2-2. Method for creating correct answer data according to the embodiment 2-3. Specific example of the model according to the embodiment 2-4. Example of evaluation of processing results according to the embodiment 2-5. Specific example of training data generation according to the embodiment 2-6. Specific example of a dataset applicable to the embodiment 2-7. Video generation using MV description text according to the embodiment
[0010] (1. Overview of this Disclosure) In this disclosure, a machine learning model is trained using multimodal data that includes the music data of the song, the type of the song, the type of music video to be generated, and lyric understanding that interprets the lyrics of the song. For example, a user inputs multimodal data related to the music data of the song for which they want to generate a music video into the machine learning model, and the model outputs descriptive text data for generating the music video. The multimodal data input into the machine learning model includes at least music data and text data that define the type of music video.
[0011] Users can input this generated descriptive text data into any video generation application to obtain a music video related to their desired song.
[0012] (1-1. Information Processing System Applicable to Embodiments of the Present Disclosure) Now, an information processing system applicable to embodiments of the present disclosure will be described.
[0013] Figure 1 is a schematic diagram showing an example of an information processing system applicable to embodiments of the present disclosure. In Figure 1, the information processing system 1 includes a server 10, which is connected to a terminal device 20 via a communication network 2 such as the Internet.
[0014] Server 10 includes a machine learning model trained using multimodal data, which includes the music data of the song, the type of the song, the type of music video to be generated, and lyric understanding that interprets the lyrics of the song. Terminal device 20 may be a personal computer, a smartphone, or a tablet computer. However, it is not limited to these, as long as terminal device 20 includes data input means, communication means for communicating via the communication network 2, and a processor for processing data.
[0015] Furthermore, the configuration of the information processing system 1 is not limited to a configuration using the terminal device 20; for example, it may be a configuration in which data is directly input to the server 10. In Figure 1, the information processing device 21 has a video generation function that generates video data corresponding to the input text data. The information processing device 21 is connected to the communication network 2 and is capable of communicating with the server 10 and the terminal device 20.
[0016] Figure 2 is a block diagram showing an example configuration of a server 10 applicable to the embodiment. In Figure 2, the server 10 includes a CPU (Central Processing Unit) 1000, a ROM (Read Only Memory) 1001, a RAM (Random Access Memory) 1002, a storage device 1003, a data interface 1004, and a communication interface 1005, and each of these parts is connected to each other via a bus 1010 so that they can communicate with one another.
[0017] The storage device 1003 is a storage device that uses a hard disk drive or flash memory as a storage medium. The CPU 1000 controls the operation of the server 10 by using the RAM 1002 as work memory, for example, according to a program stored in the storage device 1003 or ROM 1001.
[0018] The data interface 1004 is an interface for sending and receiving data with external devices, and may use USB (Universal Serial Bus) or similar technologies. The communication interface 1005 controls communication via the communication network 2 according to instructions from the CPU 1000.
[0019] In addition to the configuration shown in Figure 2, Server 10 may also have input devices such as a keyboard and output devices such as a display. Furthermore, although Figures 1 and 2 show Server 10 as being composed of a single computer, this is not an example; Server 10 may be configured in a distributed manner using multiple computers that are connected to each other in a communicative manner, or it may be configured using cloud computing.
[0020] Figure 3 is an example functional block diagram illustrating the functions of a server 10 applicable to embodiments of this disclosure.
[0021] In Figure 3, the server 10 includes a processing unit 100, an overall control unit 101, and a communication unit 102. These processing unit 100, overall control unit 101, and communication unit 102 are configured to run on a CPU 1000 with information processing programs according to each embodiment of this disclosure. However, these processing unit 100, overall control unit 101, and communication unit 102 may also be configured with hardware circuits that work together.
[0022] The overall control unit 101 controls the overall operation of the server 10. The overall control unit 101 may be, for example, an OS (Operating System) running on the CPU 1000. The communication unit 102 controls the communication interface 1005 to communicate with the communication network 2.
[0023] The processing unit 100 includes a model 110, which is a machine learning model according to each embodiment. It inputs music data and text data transmitted from the terminal device 20 into the model 110 to generate MV description text data corresponding to the music data. The processing unit 100 also trains the model 110 using a dataset that includes music data of a song, the type of song, the type of music video to be generated, and lyric understanding obtained by interpreting the lyrics of the song.
[0024] In the following, music videos may be abbreviated as "MV".
[0025] The information processing program according to each embodiment may be stored in a non-volatile or volatile storage medium and provided to the server 10, or it may be provided to the server 10 via the communication network 2. For example, in the server 10, when the information processing program according to each embodiment is executed, the CPU 1000 configures the above-described processing unit 100 in the main memory area of the RAM 1002, for example, as a module.
[0026] In such an information processing system 1, if a user wants to create a music video related to a certain song, for example, they input the music data of the song, text data indicating the type of song, text data indicating the type of music video they want to create, and text data for understanding the lyrics of the song into the terminal device 20. The terminal device 20 transmits the input music data and each text data to the server 10 via the communication network 2.
[0027] Server 10 passes the music data and text data transmitted from terminal device 20 to processing unit 100. Processing unit 100 inputs the received music data and text data into a pre-trained model 110. Based on the input music data and text data, processing unit 100 uses model 110 to generate MV description text data for generating a music video corresponding to the input music data.
[0028] Server 10 transmits the generated MV description text data to the terminal device 20 via the communication network 2. Based on the MV description text data transmitted from Server 10 to the terminal device 20, the user can use a predetermined video generation function to generate a music video corresponding to the transmitted music data.
[0029] For example, a user may send MV description text data transmitted from server 10 to information processing device 21, and generate a music video using the video generation function of information processing device 21. Alternatively, server 10 may directly transmit the generated MV description text data to information processing device 21, generate a music video using the video generation function of information processing device 21, and transmit the generated music video to terminal device 20. The server may have a video generation function that generates a video according to the text data, and generate the music video on server 10, or terminal device 20 may have such a video generation function, and generate the music video on terminal device 20.
[0030] Since the configuration of the terminal device 20 and the information processing device 21 can be adapted to that of a general computer, a detailed explanation is omitted here.
[0031] (2. Embodiments of the Disclosure) Next, embodiments of the disclosure will be described.
[0032] (2-1. Outline of the usage of the model according to the embodiment) The usage of the model according to the first embodiment will be described in general terms. Figure 4 is a schematic diagram for explaining the usage of the model according to the embodiment. More specifically, Figure 4 schematically shows the data input / output relationship in the multimodal LLM600 after training is complete.
[0033] In Figure 4, the multimodal LLM (Large Language Model) 600 corresponds to the model 110 shown in Figure 3, and performs inference using multimodal data 500 containing data in multiple formats as input to generate a music video (MV) description 700 using text data. Hereinafter, music video may be abbreviated as MV. The multimodal data 500 uses multiple types of music information that may influence the description of the MV description 700. More specifically, in the embodiment, the multimodal data 500 includes at least (1) music data and (3) MV type from among (1) music data, (2) music (song) type, (3) MV type, and (4) lyric understanding. The multimodal data 500 may further include (2) music data and (4).
[0034] Here, (1) music data may be waveform data, for example, and (2) music type, (3) MV type, and lyric understanding may each be text data. (2) Music type may be the genre of the music, for example. MV type may be inferred using a model such as GPT (Generative Pre-trained Transformer) based on the genre of the music video or the music video data corresponding to the music data. (4) Lyric understanding may be a summary or interpretation of the lyrics using an LLM, or it may be the lyrics themselves.
[0035] (2-2. Learning of the Machine Learning Model According to the Embodiment) Figure 5 is a schematic diagram illustrating an example of a training dataset used to train the multimodal LLM 600 according to the embodiment. As shown in Figure 5, the training dataset includes multiple pairs of data 800 (pairs #1, #2, ..., #n) of multimodal data including (1) music data, (2) music type, (3) MV type, and (4) lyric understanding, and MV description text 710 corresponding to the multimodal data 510. Here, the MV description text 710 is the ground truth data that is expected to be generated by the multimodal LLM 600 based on the corresponding multimodal data 510.
[0036] Note that the training dataset according to the embodiment is not limited to the example in Figure 5. For example, in the example in Figure 5, the multimodal data 510 includes (1) music data, (2) music type, (3) MV type, and (4) lyrics understanding. However, the multimodal data 510 may selectively use (1) music data, (2) music type, (3) MV type, and (4) lyrics understanding, or it may also include other data.
[0037] Figure 6 is a schematic diagram illustrating the learning mode of the multimodal LLM 600 according to the embodiment. As shown in Figure 6, in this embodiment, the multimodal LLM 600 is trained to output MV description text 710 as ground truth data based on the input multimodal data 510 by learning paired data 800 in the learning dataset described using Figure 5.
[0038] Next, a method for creating correct answer data according to the embodiment will be described using Figures 7 to 9. In this embodiment, text data is generated based on non-text data such as MV data and music data so that LLM can be used to process multimodal data. Using the text data thus generated based on the non-text data, LLM generates an MV description 710 as correct answer data.
[0039] Figure 7 is a schematic diagram illustrating a process for generating text data from non-text data according to an embodiment. As shown in section (a) of Figure 7, for example, pre-existing content, such as MV data 501, is input to a trained AI model 601a. The AI model 601a performs inference based on the input MV data 501 and generates an MV type corresponding to the MV data 501 using text data. Note that if the MV type 720 is attached to the MV data 501, for example as metadata, that data may be used directly or indirectly. Furthermore, any means of determining the type of video data based on the video data may be used, not limited to the AI model 601a. For example, the MV type 720 may be manually entered by the user.
[0040] Similarly, as shown in section (b) of Figure 7, the MV data 501 is input to the trained AI model 601b. The AI model 601b performs inference based on the input MV data 501 and generates a video caption 721 corresponding to the MV data 501 as text data. Note that any means of summarizing the video content based on the video data and generating text data may be used, not limited to the AI model 601b. For example, the video caption 721 may be manually entered by the user, or if metadata is attached to the MV data 501, it may be created and attached using that metadata.
[0041] In section (c) of FIG. 7, the music data 502 is data corresponding to the MV data 501, and may be data extracted from the MV data 501, or may be data obtained separately from the MV data 501 in relation to the MV data 501. The music data 502 is input into the trained AI model 601c. The AI model 601c makes an inference based on the input music data 502 and generates a music caption and low-level music features related to the music data. The low-level music features may include elements of the music by the music data, such as information on tempo, chords (harmony), downbeat, key, etc. The low-level music features may be extracted by an open-source tool (Bock et al., 2016) according to a predetermined tool, such as the method of L Lark (Gardner et al., 2024).
[0042] The AI model 601c integrates the generated music caption and low-level music features to generate an integrated music caption 722. Note that, not limited to the AI model 601c, any means for summarizing the music by the music data based on the music data 502 may be used. For example, the integrated music caption 722 may be manually input by the user, or when metadata is added to the music data 502, the metadata may be used to create and assign it.
[0043] Note that the above-described MV type 720, video caption 721, and integrated music caption 722 may be manually input by the user.
[0044] FIG. 8 is a schematic diagram for explaining a method of creating correct data according to an embodiment. The video caption 721 and integrated music caption 722 generated as described using FIG. 7, and the lyric understanding 723 are input into the trained AI model 602. Note that the text data of the lyric understanding 723 may be generated by a pre-trained model (e.g., OpenMU music understanding model). Also, the text data of the lyric understanding 723 may be the lyrics themselves.
[0045] The AI model 602 may be an LLM. The AI model 602 makes inferences based on the input video caption 721, integrated music caption 722, and lyric understanding 723, and generates the MV description text 710 as the correct data in text data. The AI model 602 may further use the MV type 720 to generate the MV description text 710.
[0046] In FIG. 7 described above, the integrated music caption 722 was described as being generated based on the music data 502, but this is not limited to this example. FIG. 9 is a schematic diagram showing another example of a method for creating an integrated music caption according to an embodiment. As shown in FIG. 9, based on the image data 350a, video data 350b, and music data 350c, an image caption 351a, a video caption 351b, and a music caption 351c are generated. Here, the image data 350a, video data 350b, and music data 350c are related to each other. For example, the image data 350a and music data 350c may be data extracted from the video data 350b. Also, the image caption 351a, video caption 351b, and music caption 351c may be created by the user based on the image data 350a, video data 350b, and music data 350c.
[0047] These image caption 351a, video caption 351b, and music caption 351c, and the attribute information 351d (tempo information, chord information, downbeat information, key information, etc.) of the music data 350c are integrated by an integration unit 352, for example, GPT, to generate the integrated music caption 722.
[0048] Not limited to this, for example, bone data such as a person included in the MV data 501 may be estimated, and the estimated bone data may be further used to generate the integrated music caption 722.
[0049] The MV description 710 output from the learned multimodal LLM 600 needs to be able to provide rich content regarding the visual elements of the music video while maintaining a close relationship with advanced features such as the tempo of the song, the downbeat, and the mood conveyed by the music. For this reason, in this embodiment, a video caption 721 for the MV data is generated based on the MV data 501 corresponding to the music data 502, and the visual context related to the MV data 501 is extracted. Next, an integration unit 352 using a model such as GPT narrows down these captions and generates an integrated music caption 722.
[0050] (2-3. Specific Examples of Models and Data According to the Embodiment) Next, an example of the model 110 according to the embodiment will be described in more detail. Note that the model 110 is not particularly limited as long as it is a model that operates with multimodal data 500 as input and MV description text 700 as output. Here, as an example of the model 110 according to the embodiment, its specific components will be described.
[0051] Figure 10 is a schematic diagram illustrating a model according to an embodiment. Model 110 includes an encoder 120, a projection unit 121, and an LLM 122. Model 110 is configured, for example, by a DNN (Deep Neural Network), and the encoder 120, projection unit 121, and LLM 122 are each configured as layers of the DNN. Model 110 may utilize an existing model, NExT-GPT: Any-to-Any Multimodal LLM (hereinafter, NExT-GPT).
[0052] Of the multimodal data 500, the music data, which is waveform data, is input to the encoder 120 as input data 500a and converted into, for example, feature data. The feature data may be, for example, vector data. The encoder 120 may apply, for example, image binding, a method developed by Meta Corporation in the United States that enables the integrated handling of multiple data of different modalities. The feature data output from the encoder 120 is converted into feature data suitable for the subsequent LLM 122 by the projection unit 121. The feature data converted by the projection unit 121 is input to the LLM 122.
[0053] Of the multimodal data 500, the music type, MV type, and lyrics understanding, which are text data respectively, and the instructions for the task are input to the LLM 122 as input data 500b. The LLM 122 may be the multimodal LLM 600 described above, and for example, Vicuna-7B can be applied.
[0054] The LLM 122 performs inference based on the feature data input from the projection unit 121 and each data input as input data 500b, and generates generated output data 730. The LLM 122 may include an adaptation unit 123 that performs fine tuning. The adaptation unit 123 may perform fine tuning using a method called LoRA (Low-Rank Adaptation), which is one of the efficient additional learning methods. Note that the LLM 122 is not limited to the configuration shown in Figure 10, as long as it is configured to perform learning that takes multimodal data 500 into consideration.
[0055] In Model 110, the projection unit 121 and the adaptation unit 123, indicated by the "★ (star)" symbol, are trained so that the generated output data 730 produced by the LLM 122 based on the input multimodal data 500 approaches the correct data, which is the MV description text 710. On the other hand, in the example in Figure 10, the encoder 120 and LLM 122, which are not indicated by the "★ (star)" symbol in Model 110, are already trained and are shown here as examples of what is not trained. Note that the way in Model 110 is trained is not limited to the example shown in Figure 10.
[0056] Figure 11 is a schematic diagram showing an example of input and output to the LLM 122 (multimodal LLM 600) learned as described above, according to the embodiment. The input to the LLM 122 is multimodal data 500, which in this example includes (1) music data, (2) music type (music genre), (3) MV type, and (4) lyric understanding. (1) Music data is input as waveform data, as described above. (2) Music type, (3) MV type, and (4) lyric understanding are each input as text data.
[0057] Instruction 60 is text data used to instruct the LLM 122 on the content of the music video to be generated. The content instructed by Instruction 60 may be the scenario of the music video to be generated, or it may be an overview of the music video. Furthermore, Instruction 60 may be a simplified summary of the gist of the music video. Instruction 60 is input to the LLM 122, for example, as text data.
[0058] The output of LLM122 is an MV description 700 generated by LLM122 based on the multimodal data 500 and instructions 60 input to LLM122. The MV description 700 may be output from LLM122 as text data. The MV description 700 includes an overview of the music video generated based on this MV description 700, and an overview of each divided period (frame-by-frame breakdown) when the generated music video is divided into frame units. The overview of each divided period may be, for example, an overview of each period when the generated music video is divided into predetermined periods (for example, every 2 seconds). In the example in Figure 5, the overview of each divided period includes the title of each period and an overview of the video content of each period.
[0059] The period for outputting a music video summary is not limited to predetermined periods into which the music video is divided. For example, a summary of the entire music video may be output without including time information. Furthermore, scenario and video summaries may be output for specific periods within the music video, or scenario and video summaries may be output for any specified period within the music video.
[0060] Figure 12 is a schematic diagram showing a more specific example of the MV description 700 output from the LLM 122 (multimodal LLM 600) according to the embodiment. In Figure 12, the MV description 700 includes an overview of the generated music video shown in the upper section and an overview of each divided period of the music video shown in the lower section. In the example of Figure 12, the overview of the music video includes the style and visual and musical characteristics of the generated music video. The overview of each divided period includes the title and video content overview of each period obtained by dividing the generated music video into, for example, 13 periods of predetermined duration.
[0061] Figure 13 is a schematic diagram showing an example of an MV description 710 as correct answer data according to the embodiment. In Figure 13, the MV description 710 corresponds to the MV description 700 shown in Figure 12 and includes the correct summary of the music video shown in the upper section and the summaries of each correct segmented period for the music video shown in the lower section. As shown in Figure 13 and Figure 12 described above, the MV description 700 generated by the LLM 122 based on the input multimodal data 500 and the MV description 710 prepared as correct answer data may be different.
[0062] The MV description 710 for the correct answer data may be created by the user through manual input or by creating it in advance using a model such as GPT.
[0063] (2-4. Evaluation Examples of Processing Results According to the Embodiment) Next, an example of evaluation of processing results according to the embodiment will be described.
[0064] Figure 14 is a schematic diagram showing an example of evaluating the processing results according to the embodiment. Figure 14 shows an example of comparing the results of fine-tuning using the dataset according to the embodiment with Model #A (e.g., NExT-GPT) as the baseline. As described above, the dataset includes (1) music data, (2) music type, (3) MV type, and (4) lyric comprehension. In the example in Figure 9, each combination of data (1) to (4) included in this dataset is evaluated using multiple evaluation indices Eva#A to Eva#H.
[0065] In Figure 14, the top row shows the evaluation results relative to the baseline, and below that, the evaluation results for the primary results (2 sets), sanity checks (2 sets), and ablation studies (6 sets) are shown. The primary results evaluate the combinations (1) + (4) and (1) + (2) + (3) + (4). The sanity checks evaluate the results with inputs (2) and (3) removed from (1) to (4), and the results with all of (1) to (4) included. The ablation studies evaluate four results with one input removed from (1) to (4), and two results with inputs (1) and (4) and (2) and (4) removed.
[0066] Figure 14 shows that (1) music data and (3) MV type are important factors that greatly influence the quality of the MV description. On the other hand, (2) music type and (4) lyric comprehension also influence the quality of the MV description, but their effects are interchangeable.
[0067] Furthermore, ablation studies exploring different data source combinations reveal that the sets (1) + (2) + (3) and (1) + (3) + (4) achieve performance comparable to the total data combination (1) + (2) + (3) + (4). This indicates that when (2) music type and (4) lyric comprehension text are used together, their contributions are the same without adding any other advantages.
[0068] The results for (1) + (3) show that music type (2) and (4) lyric comprehension text have a positive impact on the results and are not redundant inputs. When comparing the top three best-performing sets (1) + (2) + (3), (1) + (3) + (4), and (1) + (2) + (3) + (4) with each combination of (2) + (3) + (4) and (1) + (2) + (4), a significant performance degradation is observed. This makes it clear that including both (1) music data and (3) MV type is important and effective. Furthermore, the results for (2) + (3) show that even with simple inputs of music type and MV type, the model can generate reasonable MV description data.
[0069] In other words, this suggests the possibility of improving model performance by utilizing fine-grained features such as lyric sequences and low-level musical characteristics.
[0070] (2-5. Specific Examples of Training Data Generation According to the Embodiment) Next, specific examples of training data generation according to the embodiment will be described.
[0071] Figure 15 is a schematic diagram showing an example of an original music video (hereinafter referred to as the original MV). Figure 15 may, for example, show an example of the MV data 501 in Figure 7 described above. In the example of Figure 15, the original MV is divided into predetermined time intervals (for example, 2 seconds), and examples of the first frame images of each divided period, Frame #01 to Frame #15, are shown. Music data is extracted from this original MV, and based on the extracted music data, music captions are generated using, for example, a GPT model. Similarly, video captions are generated based on the original MV. Furthermore, if lyrics corresponding to the original MV can be obtained, lyric comprehension text is generated.
[0072] Figure 16 is a schematic diagram showing an example of a video caption 721 generated based on the original MV that can be applied to the embodiment. Figure 16 may, for example, show the example of the video caption 721 in Figure 7 described above. In Figure 16, the video caption 721 includes an overview of the entire original MV and an overview of each of the 15 divided segments, Frame #01 to Frame #15, from which the music video is divided.
[0073] Figure 17 is a schematic diagram showing an example of an integrated music caption 722 generated based on music data extracted from an original MV, applicable to the embodiment. Section (a) of Figure 17 shows an example of input data for generating the integrated music caption 722, and section (b) shows an example of an integrated music caption 722 generated based on the input data shown in section (a). Section (b) of Figure 17 may show, for example, the integrated music caption 722 examples of Figure 7 and Figure 9 described above.
[0074] As shown in section (a) of Figure 17, music captions and low-level music characteristics based on music data are input into Model #X, for example, shown in section (b). Music captions are input as text data, describing the musical features based on the music data. On the other hand, low-level music characteristics are input as text data, describing the musical elements based on the music data, such as tempo, chords, downbeat, key, and other information.
[0075] Model #X generates an integrated music caption 722, as shown in section (b), based on the input data. The integrated music caption 722 may include, for example, an overview of the input music data, the progression of the music during playback, and the impression of the music data.
[0076] Figure 18 is a schematic diagram showing an example of a lyric interpretation 723 that interprets lyrics corresponding to an original music video, applicable to the embodiment. Figure 18 may show, for example, the lyric interpretation of Figure 6(4) described above or the lyric interpretation 723 of Figure 8. Section (a) of Figure 18 shows an example of input data for generating the lyric interpretation 82, and section (b) shows an example of the lyric interpretation 723 generated based on the input data shown in section (a).
[0077] As shown in section (a) of Figure 18, music data and lyrics data are input from the dataset to model #Y, for example, shown in section (b). Based on this input data, model #Y generates a lyrics understanding 82, as shown in section (b). The lyrics understanding 723 may be, for example, an interpretation of the lyrics. For example, the lyrics understanding 82 may be the intent of the lyrics, the scene or background expressed by the lyrics, etc.
[0078] Figure 19 is a schematic diagram showing an example of an MV description 700 generated using the model 110 according to the embodiment, based on video captions 721, integrated music captions 722, and lyric understanding 723. Figure 19 may show an example of the MV description 700 in Figure 4 described above. Also, the MV description 700 shown in Figure 19 may correspond to the generated output 730 in Figure 10 described above. Section (a) of Figure 19 shows an example of input data 500b for generating the MV description 700, and section (b) shows an example of an MV description 700 generated by the model 110 (shown as Model#Z in the figure) based on the input data 500b shown in section (a).
[0079] As shown in section (a) of Figure 19, the input data 500b includes a video caption 721 (labeled "(1) Video caption" in the figure), an integrated music caption 722 (labeled "(2) Unified Music Caption" in the figure), and a lyrics comprehension 723 (labeled "(3) Lyrics understanding" in the figure). The MV description 700 includes an overview of the music video based on the input data 500b, and overviews of each of the 15 segments into which the music video is divided, from Frame #01 to Frame #15.
[0080] For example, the server 10 transmits the generated MV description 700 to the terminal device 20. Based on the MV description 700 received by the terminal device 20, the user can generate a music video, for example, using the video generation function of the information processing device 21.
[0081] The server 10 may also have a video generation function for generating a music video (composite music video) based on the MV description 700. Furthermore, the server 10 may have a function for generating image data (composite image data) based on the MV description 700, and a function for generating a music video based on said composite image data. These functions may also be provided by the terminal device 20.
[0082] Figure 20 is a schematic diagram showing an example of a music video generated by the model 110 according to this embodiment. Figure 20 may be an example of a music video generated based on the generation output 730 of Figure 10 described above. In the example of Figure 20, similar to the example of Figure 15 described above, the original MV is divided into predetermined time intervals (for example, 2 seconds), and the first frame image of each divided period, Frame #01 to Frame #15, is shown. The moving images of each divided section correspond to the outline of each divided period, Frame #01 to Frame #15, contained in the MV description text data 73 shown in Figure 19.
[0083] Thus, according to the embodiments of this disclosure, it is possible to generate an MV description 700 for generating a music video based on at least the music data 502 of the music for which a music video is to be created and the MV type 720.
[0084] Furthermore, in this embodiment, by further using the music type of the music data 502 and the lyrics understanding 723 corresponding to the music data 502, it is possible to generate an MV description 700 that can produce a more accurate music video. Here, the music type may be the integrated music caption 722 described with reference to Figures 7 and 9.
[0085] (2-6. Specific Examples of Datasets Applicable to Embodiments) Next, we will describe in more detail examples of datasets applicable to embodiments.
[0086] Figures 21A and 21B are schematic diagrams illustrating examples of datasets consisting of four types of data, each related to the other: image data, video data, music data (audio data), and text data.
[0087] In Figure 21A, section (a) shows an example of image data 90a and an image caption 83a for the image data 90a, and section (b) shows an example of video data 91a and a video caption 84a for the video data 91a. Section (c) shows an example of music data (audio data) 92a and a music caption 85a for the music data 92a.
[0088] Image data 90a, video data 91a, and music data 92a are related data, and for example, image data 90a and music data 92a can be extracted from video data 91a. In addition, image captions 83a, video captions 84a, and music captions 85a may be generated by applying models such as Image to Text, Video to Text, and Music to Text to image data 90a, video data 91a, and music data 92a, respectively.
[0089] The example in Figure 21B is similar to the example in Figure 22A. That is, in Figure 21B, section (a) shows an example of image data 90b and an image caption 83b for the image data 90b, and section (b) shows an example of video data 91b and a video caption 84b for the video data 91b. Section (c) shows an example of music data (audio data) 92b and a music caption 85b for the music data 92b. In the example in Figure 21B, for example, the music caption 85b is more detailed than the music caption 85a shown in section (c) of Figure 21A.
[0090] Model 305 may be trained using a dataset that includes image data, video data, and music data, along with their caption data. Alternatively, this dataset and the instructions 343 may be input to model 305 to generate output data 354 (MV description text data).
[0091] Figure 22 is a schematic diagram illustrating the generation of an integrated music caption 722 according to an embodiment. In Figure 22, section (a) shows examples of image data 350a, video data 350b, and music data 350c as described using Figure 9. The image data 350a and music data 350c may be extracted from the video data 350b, as described above, or they may be acquired individually. Based on these image data 350a, video data 350b, and music data 350c, image captions 351a, video captions 351b, and music captions 351c are generated, respectively.
[0092] In Figure 22, section (b) shows an example of an integrated music caption 722 obtained by integrating an image caption 351a, a video caption 351b, a music caption 351c, and attribute information 351d (not shown) by an integration unit 352. The integrated music caption 353 may be text data obtained by reconstructing and integrating summaries of each of the image caption 351a, video caption 351b, music caption 351c, and attribute information 351d.
[0093] Figure 23 is a schematic diagram that more specifically shows an example of input data 500 and instructions 60 input to the model 110 according to the embodiment, as well as an example of an MV description 700 generated by the model 110.
[0094] In the example in Figure 23, the input data 500 is tag " <input> Text data is embedded in the portion enclosed by " and ". Additionally, for example, access information (such as a URL (Uniform Resource Locator)) for music data (audio data), image data, and video data is embedded in the text data using the tag " <audio>" " and " <video>They are embedded in the positions indicated by ".
[0095] Similarly, Instruction 60 has the tag " <instruction>" and "< / instruction> Instructional text data is embedded in the section between the quotation marks. Instruction 60 is expected to include anything related to music data, image data, or video data captions, or integrated captions. Model 110 generates a response MV description 700 in response to such instructions 60.
[0096] MV description 700 is tag " <output> " and "< / output> The text data of the MV description 700 generated by Model 110 is embedded in the section enclosed by the quotation marks.
[0097] Various datasets may be used to train Model 110. In the simple Stage #1 training, the adapter (e.g., projection unit 121) is fine-tuned.
[0098] In the following equations (1) to (10), the meaning of each character is as follows. Also, the left side of the "→ (arrow)" indicates the input modal, and the right side indicates the output modal. I: Image data M: Music data (audio data) V: Video data T: Text data T(1): Image caption T(2): Music caption T(3): Video caption TU: Integrated caption
[0099] Note that the integrated caption (TU) referred to here is a combination of music data and multimodal data consisting of video data and / or image data, and is different from the integrated music caption 722 described in Figures 7 and 9.
[0100] <Stage #1> Dataset A: I → T (1) ... (1) Dataset B: M → T (2) ... (2) Dataset B: V → T (3) ... (3)
[0101] Furthermore, as a variation of the learning process, in Stage #2, fine-tuning of the adapter (e.g., projection unit 121) and LoRA (adaptation unit 321c) is performed.
[0102] <Stage #2> Dataset C: T → T …(4) Dataset D: M + T → T …(5) Dataset E: M → T (2) …(6) Dataset B: I → T (1) …(7) Dataset B: M → T (2) …(8) Dataset B: V → T (3) …(9) Dataset B: M + V / I → TU …(10)
[0103] Of these equations (1) to (10), datasets A and C to E of equations (1) and (4) to (6) are existing datasets. On the other hand, dataset B of equations (2) and (3), and equations (7) to (10) are datasets according to the embodiment, created using the method described with reference to Figures 21A, 21B, and 22. In particular, equation (10) corresponds to Figure 22 and takes multimodal data consisting of music data, video data, and / or image data as input and outputs integrated captions.
[0104] Existing music understanding models are trained using only data from two modalities. In contrast, the music understanding model according to the embodiment of this disclosure improves the accuracy of music understanding by incorporating visual information obtained from multi-way data. Furthermore, in the embodiment, for example, as shown in equations (7) to (10) above, the model 305 can be trained using pairs of inputs from different modalities and appropriate outputs for those inputs, thereby generating outputs that understand music in a general way for various inputs.
[0105] Furthermore, the integrated caption may be added to the video caption 721, integrated music caption 722, and lyrics comprehension 723 used to generate the MV description 710 as correct data, as explained using Figure 8. However, the integrated caption may also be used to replace one of these video captions 721, integrated music caption 722, and lyrics comprehension 723, or it may be used as the MV description 710 as correct data.
[0106] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.
[0107] (2-7. Video generation using MV description text according to the embodiment) Next, an example of video generation using MV description text 700 generated by model 110 (multimodal LLM600) according to the embodiment will be explained with reference to Figures 24A, 24B, and 25.
[0108] Figure 24A schematically shows an example of video generation using a single-modal AI model 805 that processes only one type of data. In Figure 24A, the MV description text 700 output from the multimodal LLM 600 (see Figure 4) is input to the AI model 800, which is to which a model such as Text-to-Video that generates video data based on text data is applied. The AI model 805 to which the MV description text 700 is applied is not limited to Text-to-Video; any model capable of generating video data based on text data may be applied. The AI model 805 generates MV data 900, which is video data, based on the input MV description text 700.
[0109] In the configuration shown in Figure 24A, instructions 60 may also be input to the AI model 805. In this case, the AI model 805 can generate MV data 900 that reflects the description of the instructions 60 in addition to the MV description 700.
[0110] Figure 24B schematically shows an example of video generation using a multimodal AI model 810 capable of processing multiple types of data. In Figure 24B, music data 510 is input to the AI model 810, which generates a video based on multimodal data, along with the MV description text 700 output from the multimodal LLM 600 (see Figure 4). The AI model 800 generates MV data 910, which is video data, based on the input MV description text 700 and music data 510. The AI model 810 generates MV data 910 that reflects the content of the music data 510 in addition to the MV description text 700.
[0111] In the configuration shown in Figure 24B, instructions 60 may also be input to the AI model 810. In this case, the AI model 810 can generate MV data 910 that reflects the contents of the instructions 60 in addition to the MV description 700 and music data 510.
[0112] Figure 25 schematically shows examples of MV explanatory text 700 and instructions 60 that can be applied to the examples in Figure 24A or Figure 24B described above.
[0113] In Figure 25, section (a) schematically shows an example of an MV description 700. In the example in Figure 25, the MV description 700, similar to the example in Figure 12 described above, includes an overview of the generated music video shown in the upper section and an overview of each divided period of the music video shown in the lower section. In the example in Figure 25, the overview of the music video also includes the style and visual and musical characteristics of the generated music video. The overview of each divided period includes the title and video content overview of each period, obtained by dividing the generated music video into, for example, five periods of predetermined duration.
[0114] In Figure 25, section (b) schematically shows an example of instruction 60. In this example, instruction 60 instructs the creation of a 10-second music video based on the music video outline shown in the MV description 700 and the outlines of each segment described at 2-second intervals. Instruction 60 may be entered entirely manually by the user, or it may be selected from pre-prepared instructions.
[0115] The AI model 805 or 810 may, for example, combine multiple short videos of about 10 seconds each to create a longer music video, or it may create a music video with a predetermined length from the outset.
[0116] Furthermore, AI model 805 or 810 may be mounted on the information processing device 21 in Figure 1. However, it is not limited to this; AI model 805 or 810 may be mounted on the server 10 or on the terminal device 20. Moreover, AI model 805 or 810 may be mounted on a standalone information processing device that is not connected to the communication network 2.
[0117] Furthermore, this technology can also take the following configurations: (1) An information processing system including a processing circuit, the processing circuit receiving first multimodal data related to first music data, inputting the first multimodal data into a machine learning model to output descriptive text data relating to the first music data, and the descriptive text data being used to generate first music video data. (2) The information processing system according to (1), wherein the first multimodal data includes the first music data and first text data for specifying the type of first music video data. (3) The information processing system according to (2), wherein the first text data is specified by a user. (4) The information processing system according to (2) or (3), wherein the first multimodal data further includes at least one of the following: second text data indicating the type of music by the first music data, third text data which is a description of lyrics relating to the first music data, and fourth text data corresponding to instructions for a task. (5) The information processing system according to any one of (1) to (4), wherein the machine learning model is trained on a training dataset created based on a second music video data which is existing content. (6) The information processing system according to (5), wherein the training dataset is paired data of a second multimodal data created based on the second music video data and a second descriptive text data relating to the second music video data. (7) The information processing system according to (5) or (6), wherein the machine learning model is trained on the second music video data which corresponds to the second music data. (8) The information processing system according to (7), wherein the machine learning model is trained on a first training text data relating to the second music video data. (9) The information processing system according to any one of (5) to (8), wherein the machine learning model is trained on a second training text data which specifies the type of the second music video data.(10) The information processing system according to (8), wherein the first learning text data is generated based on video caption data based on the second music video data. (11) The information processing system according to (8), wherein the first learning text data is generated based on integrated music caption data and a fifth text data which is an explanatory text relating to lyrics associated with the second music data. (12) The information processing system according to any one of (1) to (11), wherein the processing circuit is further configured to generate the first music video data based on the explanatory text data. (13) The information processing system according to (12), wherein the processing circuit is further configured to generate the first music video data based on a third music data. (14) The information processing system according to any one of (1) to (13), wherein the processing circuit is further configured to generate composite image data based on the explanatory text data. (15) The information processing system according to (14), wherein the processing circuit is further configured to generate the composite image data based on the fourth music data. (16) The information processing system according to (14) or (15), wherein the processing circuit is further configured to generate a second composite music video data based on the composite image data. (17) The information processing system according to any one of (1) to (16), wherein the first multimodal data includes the first music data, a first text data specifying the type of the first music video data, a second text data specifying the type of music by the first music data, and a third text data which is a description of lyrics related to the first music data. (18) The information processing system according to (17), wherein the first multimodal data includes at least the first music data and the first text data from among the first music data, the first text data, the second text data and the third text data.(19) An information processing method comprising: receiving multimodal data related to first music data; inputting the multimodal data into a machine learning model to output descriptive text data relating to the first music data; and the descriptive text data being used to generate music video data. (20) The information processing method according to (19), further comprising: generating music video data based on the descriptive text data. (21) The information processing method according to (19), further comprising: generating synthesized image data based on the descriptive text data. (22) The information processing system according to (19), wherein the processing circuit further generates the synthesized image data based on a fourth music data. (23) The information processing method according to (21) or (22), further comprising: generating second synthesized music video data based on the synthesized image data. (24) A computer-readable storage medium containing a program for causing a computer to perform a process in which it receives multimodal data related to music data, inputs the multimodal data into a machine learning model to output descriptive text data relating to the music data, and the descriptive text data is used to generate music video data.
[0118] 1. Information Processing System 10. Server 20. Terminal Device 21. Information Processing Device 60. Instruction 100. Processing Unit 110. Model 120. Encoder 121. Projection Unit 122. LLM 123. Adaptive Unit 350a. Image Data 350b. Video Data 350c, 502. Music Data 351a. Image Caption 351b, 721. Video Caption 351c. Music Caption 351d. Attribute Information 352. Integration Unit 500, 510. Multimodal Data 500a, 500b. Input Data 501, 900, 910. MV Data 600. Multimodal LLM 601a, 601b, 601c, 602. AI Model 700, 710. MV Description 720. MV Type 722. Integrated music captions: 723; Lyrics comprehension: 730; Generated output: 800; Paired data: 805,810; AI model< / video> < / audio>
Claims
1. An information processing system comprising a processing circuit, the processing circuit receiving first multimodal data related to first music data, inputting the first multimodal data into a machine learning model to output descriptive text data relating to the first music data, and the descriptive text data being used to generate first music video data.
2. The information processing system according to claim 1, wherein the first multimodal data includes the first music data and first text data for specifying the type of the first music video data.
3. The information processing system according to claim 2, wherein the first text data is specified by the user.
4. The information processing system according to claim 2, wherein the first multimodal data further comprises at least one of the following: second text data indicating the type of music based on the first music data; third text data which is a description of lyrics related to the first music data; and fourth text data which corresponds to instructions for a task.
5. The information processing system according to claim 1, wherein the machine learning model is trained on a training dataset created based on a second music video data which is existing content.
6. The information processing system according to claim 5, wherein the training dataset is paired data of a second multimodal data created based on the second music video data and a second descriptive text data relating to the second music video data.
7. The information processing system according to claim 5, wherein the machine learning model is trained on the second music video data corresponding to the second music data.
8. The information processing system according to claim 7, wherein the machine learning model is trained based on first training text data relating to the second music video data.
9. The information processing system according to claim 7, wherein the machine learning model is trained based on second training text data specifying the type of second music video data.
10. The information processing system according to claim 8, wherein the first learning text data is generated based on video caption data based on the second music video data.
11. The information processing system according to claim 8, wherein the first learning text data is generated based on integrated music caption data and a fifth text data which is an explanatory text relating to lyrics associated with the second music data.
12. The information processing system according to claim 1, wherein the processing circuit is further configured to generate the first music video data based on the descriptive text data.
13. The information processing system according to claim 12, wherein the processing circuit is further configured to generate the first music video data based on the third music data.
14. The information processing system according to claim 1, wherein the processing circuit is further configured to generate synthesized image data based on the descriptive text data.
15. The information processing system according to claim 14, wherein the processing circuit is further configured to generate the composite image data based on the fourth music data.
16. The information processing system according to claim 14, wherein the processing circuit is further configured to generate a second synthesized music video data based on the synthesized image data.
17. The information processing system according to claim 1, wherein the first multimodal data includes the first music data, the first text data specifying the type of the first music video data, the second text data specifying the type of music by the first music data, and the third text data which is a description of lyrics related to the first music data.
18. The information processing system according to claim 17, wherein the first multimodal data includes at least the first music data and the first text data from among the first music data, the first text data, the second text data and the third text data.
19. An information processing method comprising: receiving multimodal data related to first music data; inputting the multimodal data into a machine learning model to output descriptive text data relating to the first music data; and using the descriptive text data to generate music video data.
20. The information processing method according to claim 19, further comprising generating the music video data based on the explanatory text data.
21. The information processing method according to claim 19, further comprising generating the music video data based on the second music data.
22. A computer-readable storage medium containing a program for performing a process in which a computer receives multimodal data related to music data, inputs the multimodal data into a machine learning model to output descriptive text data relating to the music data, and the descriptive text data is used to generate music video data.