Enabling conversation-based video music recommendations
By using a dialogue-based music recommendation system that combines video and text input for an interactive recommendation process, the system addresses the shortcomings of personalized recommendations in existing systems, achieving more accurate music selection and enhanced user trust.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-04-10
AI Technical Summary
Existing music recommendation systems fail to effectively meet users' personalized needs, especially in terms of insufficient accuracy in recommendations for new users and a lack of consideration for subtle differences in user preferences, resulting in a poor user experience.
A dialogue-based music recommendation system is adopted. The first sub-model processes video and text input, and the second sub-model uses a large language model to generate natural language explanations, realizing an interactive recommendation process between users and the system, providing music clips and their recommendation reasons.
It improves the personalization of music recommendations, meets user preferences, enhances users' trust and understanding of the recommendations, and achieves more accurate music selection.
Smart Images

Figure CN121844306A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims priority to U.S. Application 18 / 368,253, filed September 14, 2023, and entitled “Implementing Dialog-Based Video Music Recommendations,” the disclosure of which is incorporated by reference herein in its entirety. BACKGROUND
[0002] Techniques for music recommendation are widely used in the music and entertainment industry. Recently, the demand for music recommendation has become more urgent. However, traditional music recommendation techniques can not be able to meet the needs of users due to various limitations. Therefore, there is a need to improve music recommendation techniques. BRIEF DESCRIPTION OF DRAWINGS
[0003] The following detailed description can be better understood when read in conjunction with the accompanying drawings. For the purpose of illustrating the various aspects of the disclosure, example embodiments of the various aspects are shown in the drawings; however, the application is not limited to the precise arrangements and instrumentalities shown. The detailed description set forth below, in connection with the appended drawings and appended examples, is intended as a description of various configurations and is not intended to limit the scope of the disclosure. Rather, the intent is to convey the true scope of the application to those skilled in the art.
[0004] Figure 1 An example system for implementing dialog-based music recommendations for an input video is shown.
[0005] Figure 2 An example system for implementing dialog-based music recommendations for an input video is shown.
[0006] Figure 3 An example system for generating a conversational music recommendation dataset is shown.
[0007] Figure 4 An example first submodel is shown.
[0008] Figure 5 An example first submodel is shown.
[0009] Figure 6 An example system for implementing dialog-based music recommendations for an input video is shown.
[0010] Figure 7 An example process for implementing dialog-based music recommendations for an input video is shown.
[0011] Figure 8 An example process for generating a simulated music recommendation conversation is shown.
[0012] Figure 9 An example process for training a machine learning model that implements dialog-based music recommendations for videos is shown.
[0013] Figure 10 An example dialog when the target music is retrieved in two turns is shown.
[0014] Figure 11 An example table showing the music retrieval results is provided.
[0015] Figure 12 An example table is shown illustrating a comparison of semantic similarity between outputs and simulated dialogues using various metrics.
[0016] Figure 13 An example computing device is shown that can be used to perform any of the techniques disclosed herein. Detailed Implementation
[0017] Music serves as a complementary mode in videos, enriching the viewing experience and aiding in understanding. Therefore, selecting appropriate music for a video is crucial. Current music recommendation systems can effectively curate tracks that harmonize with the video content. For example, a current system might select horror music for a horror movie or high-energy tracks for a dance video. While this focus on content compatibility is important, user preferences are equally crucial. For instance, individuals born in the 1980s might prefer synth-pop music for nostalgic-themed videos, while teenagers might prefer contemporary pop music for videos they create. Although both genres fall under the "pop music" category, the choice between them can significantly impact user engagement with the video.
[0018] Generating personalized music recommendations is challenging. While many systems utilize user profiles and activity data to generate recommendations, limitations remain. One limitation of existing personalized music recommendation systems is their inability to consistently meet user preferences. A second limitation is their inability to generate personalized music recommendations for new users (e.g., users not associated with previous data). A third limitation is that they provide lists of recommended songs based on user history, which may not always align with a user's needs for specific videos. Consequently, existing personalized music recommendation systems result in poor user experiences because they fail to account for the complexities of predicting preferences. Improvements in personalized music recommendation technology are desired.
[0019] This paper describes an improved technique for personalized music recommendation. It describes an innovative dialogue-based music recommendation system (e.g., MuseChat). Unlike existing systems that primarily emphasize content compatibility (often ignoring nuances of individual user preferences), the system described in this paper provides interactive user engagement and also suggests music tailored to the input video, allowing users to refine and personalize their music choices. The dialogue-based music recommendation system described in this paper may include a first sub-model (e.g., a multimodal recommendation engine). The first sub-model can match music by aligning it with visual cues from the video and / or by coordinating visual information, feedback from previously recommended music, and user text input. The dialogue-based music recommendation system described in this paper may include a second sub-model. The second sub-model can bridge music representations and text data with a Large Language Model (LLM) (such as Vicuna-7B and / or any other suitable LLM).
[0020] The dialogue-based music recommendation system described in this paper can be configured to perform a conversational synthesis method. This method simulates a two-round interaction between the user and the recommendation system. It utilizes pre-trained music tags and artist information. The user can submit videos to the system. In response, the system can suggest suitable music clips and explain the rationale behind the suggestion. The user can then provide feedback on the suggested music clips. In response to the user feedback, improved music recommendations can be provided, along with the rationale behind the improved recommendations. Therefore, the dialogue-based music recommendation system described in this paper can provide music recommendations and reasoning on why certain music clips are recommended in a manner similar to human communication. This dialogue-based music recommendation system surpasses existing state-of-the-art models in music retrieval tasks and pioneers the integration of the recommendation process within a natural language framework.
[0021] Figure 1An example comprehensive conversational music recommendation system 100 is illustrated. System 100 may include a machine learning model. The machine learning model may include two main components: a first sub-model 104 (e.g., a music recommendation sub-model) and a second sub-model 106 (e.g., a sentence generator sub-model). The first sub-model 104 may be configured to operate in two modes. In a first mode, the first sub-model 104 may only process the input video 102 to select one or more music tracks (e.g., clips) from a music pool 103 for recommendation. The first sub-model 104 may send the music embedding(s) and titles(s) associated with the selected music tracks to the second sub-model 106. The second sub-model 106 may receive the music embedding(s) and titles(s) associated with the selected music tracks as input. The second sub-model 106 may generate natural language (e.g., words and / or sentences) associated with the music embedding(s) and titles(s). For example, the second sub-model 106 may generate a first set of sentences that explain which music tracks(s) are recommended and / or explain why those music tracks are recommended. The first set of statements can be displayed on the interface of the computing device associated with the user.
[0022] At point 108, the user can provide input indicating whether they prefer music different from at least one music track. If the user provides input indicating they do not prefer music different from at least one music track (e.g., the user is satisfied) and / or the user provides no input, the music recommendation process can be terminated. Conversely, if the user provides input indicating they prefer music different from at least one music track (e.g., the user is dissatisfied), the first sub-model 104 can operate in the second mode. In the second mode, the first sub-model can process the input video 102, (multiple) previously recommended music tracks, and user input to select one or more different music tracks from the music pool 103 for recommendation.
[0023] The first sub-model 104 can send multiple music embeddings and multiple titles associated with the selected different music tracks to the second sub-model 106. The second sub-model 106 can receive the multiple music embeddings and multiple titles associated with the selected different music tracks as input. The second sub-model 106 can generate natural language (e.g., words and / or sentences) associated with the multiple music embeddings and multiple titles. For example, the second sub-model 106 can generate a second set of sentences explaining the reasoning behind recommending the multiple different music tracks and / or why those multiple different music tracks are recommended. The second set of sentences can be displayed on an interface of a computing device associated with the user. This process can continue until the multiple recommended music tracks satisfy the user.
[0024] Users may find it difficult to explain the internal workings of existing music recommendation models because these models typically operate as black boxes. Therefore, users often lack confidence in the recommendations made by these existing systems. System 100 addresses this problem by providing a rationale for its music track recommendations. System 100 not only provides users with the reasons for its music track recommendations, but it also helps them create their own personal narratives through music.
[0025] Figure 2 Showing more details Figure 1 A comprehensive conversational music recommendation system 100 is provided. Users can upload videos 102 to the system 100 and receive one or more recommended music tracks. Users can interact with the system 100 in a conversational manner. In each conversation round, users can improve these recommendations by specifying criteria (e.g., mood, genre, instrument, theme, artist details, etc.) using natural language until they identify the desired tracks.
[0026] A user inputs video 102 into system 100. The user can request music tracks corresponding to the input video 102 using natural language. In response to this request, a first sub-model 104 can process only the input video 102 to select one or more music tracks (e.g., clips) from a music pool 103 for recommendation. The first sub-model 104 can send multiple music embeddings and multiple titles associated with the selected music tracks to a second sub-model 106. The second sub-model 106 can receive the multiple music embeddings and multiple titles associated with the selected music tracks as input. The second sub-model 106 can generate natural language (e.g., words and / or sentences) associated with the multiple music embeddings and multiple titles. The second sub-model 106 can generate a first set of sentences explaining which music tracks(s) are recommended and / or explaining why those music tracks are recommended. Figure 2 In the example, the first set of statements recommends the music track "Save Some". The first set of statements also explains why the music track "Save Some" is recommended. This first set of statements can be displayed on the interface of the computing device associated with the user.
[0027] Users can view recommendations for the music track "Save Some". Users can decide which different music tracks they want to use for video 102. Users can provide input indicating their preference for music different from at least one music track. For example, users can provide input (e.g., in natural language) indicating that they want to combine music tracks from electronic and rock genres. The first sub-model 104 can process the input video 102, the previously recommended music track "Save Some", and the user input to select one or more different music tracks from the music pool 103 for recommendation.
[0028] The first sub-model 104 can send multiple music embeddings and multiple titles associated with selected different music tracks to the second sub-model 106. The second sub-model 106 can receive the multiple music embeddings and multiple titles associated with the selected different music tracks as input. The second sub-model 106 can generate natural language (e.g., words and / or sentences) associated with the multiple music embeddings and multiple titles. For example, the second sub-model 106 can generate a second set of sentences explaining the reasoning behind recommending the multiple different music tracks and / or why those multiple different music tracks are recommended. Figure 2 In the example, the second set of statements recommends the music track "I Like Not Knowing". The second set of statements also explains why the music track "I Like Not Knowing" is recommended. This second set of statements can be displayed on the interface of the computing device associated with the user. The process can continue until (multiple) recommended music tracks satisfy the user.
[0029] Building System 100 presents three core challenges. First, existing datasets primarily consist of music-video pairs, music-text pairs, or music-text-video triples. These datasets are not well-suited for training System 100. For example, such datasets only include single-turn interactions, lacking the multi-turn dialogues crucial for more interactive and dynamic recommendation systems. Furthermore, these datasets omit the explanation of recommendations, a key feature for enhancing user understanding and trust. To address these challenges, System 100 is trained on a novel dataset specifically designed for dialogue-driven music recommendation and inference within video contexts. The data includes 98,206 quartets: videos, original music, candidate music, and two-turn dialogues. This setup mimics user interaction with a recommendation system. (See below for reference.) Figure 3 The generation of this new dataset will be described in more detail.
[0030] A second challenge associated with building System 100 involves joint multimodal learning. Creating a joint embedding space for video, music, and text is a complex task. Each of these modes has its unique sequential features, making it challenging to combine them into a unified representation. System 100 effectively integrates spatiotemporal information from these different modes, resulting in a more comprehensive representation. Specifically, the first sub-model 104 includes a three-modal architecture designed for music-video matching with text input. Thus, the first sub-model 104 not only processes previously recommended music and video content but also integrates user-provided textual prompts to fine-tune its music recommendations.
[0031] The third challenge associated with building System 100 involves predictive reasoning. While current research on Multimodal Large Language Models (MLLMs) demonstrates the ability to process and understand different modalities, such as video and audio, significant gaps remain. Specifically, these models are not built for the nuanced tasks of music interpretation and recommendation. The second subsystem 106 is able to elucidate the reasoning behind its music recommendation by leveraging the capabilities of LLMs. Utilizing music representations from upstream modules, the second subsystem 106 gains a deep understanding of music features, subsequently producing coherent inference outputs that ensure harmonious alignment between music and text descriptors.
[0032] As described above, System 100 is trained on a novel dataset specifically designed for dialogue-driven music recommendation and reasoning in video context. Figure 3 System 300 for generating datasets (e.g., conversational music recommendation datasets) is illustrated. System 300 can be generated by simulating a two-round dialogue to create a data sample. In the first round, only a video is provided, and candidate music tracks can be recommended by an underlying music recommendation system. In the second round, based on the recommended music, simulated user dialogue (e.g., in natural language) can prompt changes to the target music, along with the video and the recommended music. System 100 can then output another recommended music track that best matches the video.
[0033] Music Video Dataset 302 (e.g., the YouTube-8M dataset) can be used to construct a conversational music recommendation dataset. Music Video Dataset 302 can include a large-scale collection of videos. It can contain hundreds of thousands or millions of music video IDs and associated tags distributed across thousands of categories, including genres such as music, sports, and documentaries. Music Video Dataset 302 can serve as a valuable dataset for video understanding, particularly for broader applications such as music identification and classification within the broader scope of video research. Videos labeled "music video" can be filtered out from Music Video Dataset 302. Any unusable videos can be removed. The resulting dataset can include 98,206 music videos. Clips can be extracted from each video (e.g., 120-second clips). Clips can focus on the central segments of the video. A portion of these music videos (e.g., 88,000) can be randomly assigned to the training set. The remaining portion (e.g., the remaining 10,206 videos) can be assigned to the test set. Each video and its corresponding music can be set as the baseline truth (e.g., video and target music).
[0034] The music video pre-trained (MVP) model can be used to generate a conversational music recommendation dataset. The MVP model can include a video branch (306) and a music branch (305). The video branch (306) can utilize a pre-trained CLIP image encoder for video feature extraction. The music branch (305) can utilize a pre-trained audio spectrogram transformer (AST) for music feature extraction. The MVP model can be trained on a dataset consisting of millions of music-video pairs.
[0035] The MVP model receives candidate music and videos as input. The MVP model outputs a similarity score. This similarity score can indicate the similarity between each candidate music item and the input video (e.g., a cosine similarity score). The music from the music pool can then be ranked in descending order of similarity (e.g., the music most similar to the video is ranked highest). In an embodiment, the candidate music pool can be limited for training and testing (e.g., limited to 2000 candidate music items and 500 for testing). Original music can be excluded from both the training and testing sets. Limiting the candidate pool ensures that recommendations are not influenced by low-quality music. The MVP model is not designed to identify the most similar tracks to the original tracks. Instead, the MVP model is configured to identify tracks that represent a significant deviation from previous recommendations that the user might not be entirely satisfied with.
[0036] Given a triple consisting of a video, its original music track, and recommended candidate music tracks, system 300 can be used to construct a two-round dialogue. Specifically, during each user round, cue word constructor 316 can provide cue words to bridge the gap between the original music and the currently recommended candidate music. During the robot round, a description of the returned music (e.g., the recommended candidate music in the first round and the original music in the second round) is essential.
[0037] Music tags 312 can be assigned to each music track. Music tags 312 can effectively summarize a song by providing descriptive keywords covering various elements such as mood, genre, and theme. Music tags 312 can be generated using a model that extracts acoustic features using shallow convolutional layers. The acoustic features can then be processed by stacked self-attention layers in a semi-supervised setting. Music tags 312 can be generated using one or more separate systems. The (multiple) systems can have 50 tag words. Using more than one system to generate music tags 312 can enhance tag robustness. Music metadata 314 can be collected for each music video. Music metadata 314 can indicate the title and / or video description of the music video. Music metadata 314 can indicate the official artist name, album details, and / or release date.
[0038] Music tags 312 and music metadata 314 can be fed into cue word builder 316. Cue word builder 316 can utilize music tags 312 and music metadata 314 to generate cue words for guiding chat generative pre-trained transformer 318 to generate a two-round dialogue (e.g., simulated dialogue 320) between the user and the music recommendation system. GPT 318 can receive cue words from cue word builder 316. GPT 318 can utilize one or more of the cue words, music tags 312, and music metadata 314 to generate simulated dialogue 320. Simulated dialogue 320 can be used to generate a conversational music recommendation dataset 322. Each instance (e.g., entry) in the conversational music recommendation dataset 322 includes a video. v Original target music tracks m t Candidate music tracks m c and simulated dialogue text t More specifically, t i Indicates the sequence order of each dialogue round: t 1 and t 3 From users, and t 2 and t4 From the recommendation system.
[0039] Figure 4 An example of a first sub-model 104 is shown. The first sub-model 104 can be trained on the conversational music recommendation dataset 322. The first sub-model 104 includes three types of input: video, music, and text. The first sub-model 104 may include a video encoder 402. The video encoder 402 can be configured to extract basic embeddings from the video. The video encoder 402 can utilize a multimodal visual and language model (e.g., CLIP) to extract basic embeddings from the video. The basic embeddings extracted from the video can be fed into a video self-attention layer 408. The video self-attention layer 408 enables the first sub-model 104 to weigh the importance of different basic embeddings extracted from the video and dynamically adjust their impact on the output. The output of the video self-attention layer 408 can be fed into a cross-attention model 412.
[0040] The first sub-model 104 may include a music encoder 404. The music encoder 404 may be configured to extract representations from music. The music encoder 404 may utilize an audio spectrogram transform (AST) to extract representations from music. The first sub-model 104 may include a text encoder 406. The text encoder 406 may be configured to extract basic embeddings from text. The text encoder 406 may utilize a multimodal visual and language model (e.g., CLIP) to extract basic embeddings from text. The representations extracted from music and the basic embeddings extracted from text may be combined. This combination may be fed into a music self-attention model 410. The music self-attention model 410 enables the first sub-model 104 to weigh the importance of different representations extracted from music and basic embeddings extracted from text, and dynamically adjust their impact on the output.
[0041] The output of the music self-attention layer 410 can be fed into the cross-attention model 412. The cross-attention model 412 can receive the outputs of the video self-attention layer 408 and the music self-attention layer 410. The cross-attention model 412 can fuse the outputs of the video self-attention layer 408 and the music self-attention layer 410. The context-rich fusion features can be combined with the video embedding 414 and the music embedding 416, resulting in significant improvements.
[0042] Figure 5 The first sub-model 104 is shown in more detail. The first sub-model 104 can be trained on the conversational music recommendation dataset 322. Each training sample in the conversational music recommendation dataset 322 is defined as a quartet (…). v, m c ,m t ,t 3 ),in v It's a video. m c Let t1 represent the candidate music track, mt be the target original music track, and t3 be the text indicating user preferences. The first sub-model 104 is trained to recommend music tracks from the previous track. m c Transition to the target music track m t .
[0043] The first sub-model 104 can be trained to select the most relevant music from a music pool using various inputs, such as video, music, and text. The first sub-model 104 can include three types of input: video, music, and text. Each training sample can be transformed into basic features: features specific to visual input. For text input For audio input and . g v and g t It can be frozen during training, and and It can be fine-tuned. To transform each training sample into basic features, the first sub-model 104 can be configured to extract basic embeddings from the video. The first sub-model 104 can utilize a multimodal visual and language model (e.g., a CLIP encoder) to extract basic embeddings from the video. The first sub-model 104 can be configured to extract basic embeddings from the text. The first sub-model 104 can utilize a multimodal visual and language model (e.g., a CLIP encoder) to extract basic embeddings from the text. The first sub-model 104 can be configured to extract representations from candidate music and / or from the original music. The first sub-model 104 can be configured to extract representations from candidate music using a first AST. The first sub-model 104 can be configured to extract representations from the original music using a second AST.
[0044] In this embodiment, since these features come from different backbone models, a trainable linear projection layer can be used to map them into a common embedding space. This produces... and Given the target music track m t The target, aligned with the overall video content, can be applied to each x v Average the sequence dimensions to generate To summarize the target music, you can use music from... The first cls tag, thus generating .
[0045] To better capture information from audio and text, a transformer layer can be used to encode latent features from candidate music and text. For example, a transformer encoder can be applied to audio and text (represented as follows). and Both are used to capture long-range dependencies and complex relationships in sequence data. This results in:
[0046] In this embodiment, encoded latent features from candidate music and text can be fused using a multi-head cross-attention layer. The transformed features... and It can be represented as the following sequence: in cls As a summary of the corresponding sequence, and with other elements capturing detailed features, the multi-head crossover attention layer can be defined as: in d k It is the dimension of the key vector. Q and K, V come from two different patterns.
[0047] In an embodiment, and They can be fused to generate the following final fused embedding: Context-rich fusion features can be combined with extracted video embeddings, resulting in significant improvements.
[0048] In this embodiment, a contrastive multi-view encoding loss function can be used during training. For each batch B, the following ranking loss can be used: in and These are the i-th fusion vector and the target music representation in the batch, respectively. Let be the discriminant function, and let τ be the temperature hyperparameter, which is also a trainable hyperparameter. Larger batch sizes may be beneficial in contrastive learning.
[0049] Figure 6 The second sub-model 106 is shown in more detail. During the training of the second sub-model 106, only the linear projection layers and the additional low-rank adaptation (LoRA) weights of the large language model can be trained. LoRA is a training method that accelerates the training of large models while consuming less memory. It adds pairs of rank decomposition weight matrices (called update matrices) to existing weights and trains only those newly added weights. The parameters of Vicuna-7B can be frozen during training. A Vicuna-7B-based multimodal LLM can be constructed by fine-tuning the Llama2-7B weights. Each training instance can include a music representation from the music encoder of the first sub-model 104. and corresponding recommendation reasoning statements from simulated dialogues in the conversational music recommendation dataset 322 t 4 In order to express the music Aligning with the text embedding space allows for the training of linear projections. f l To express music Connect to Vicuna. To reduce the number of trainable parameters, LoRA can be used to fine-tune the attention structure of Vicuna.
[0050] In this embodiment, a loss function can be used during training: in y i It is a response y The first in i There are 10 labels, and θ is a trainable parameter in the linear projection layer and LoRA weights.
[0051] Figure 7 The illustration shows an example process 700 for implementing dialogue-based music recommendation on an input video. Although in Figure 7 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.
[0052] At 702, the video can be input into the machine learning model. The video can be received (e.g., by) a user (e.g., uploaded). The machine learning model can be configured to perform dialogue-based music recommendation on the input video. The machine learning model can include a first sub-model and a second sub-model. The first sub-model can be configured to operate in two modes. In a first mode, the first sub-model can process only the input video. The first sub-model can process only the input video to identify one or more music tracks corresponding to the input video. At 704, at least one music track can be identified. At least one music track can be identified by the first sub-model based on the input video. The first sub-model can send an indication of at least one music track (e.g., multiple music embeddings and multiple titles associated with at least one music track) to the second sub-model.
[0053] The second sub-model can receive instructions for at least one musical track. At 706, a first set of statements can be generated. The first set of statements can be generated by the second sub-model. The first set of statements can indicate or describe at least one musical track. The first set of statements can be written in natural language (e.g., a language that naturally develops in use, in contrast to artificial language or computer code). Additionally, a recommendation rationale can be generated. The recommendation rationale can (e.g., in natural language) explain why at least one musical track is recommended. At 708, the first set of statements can be displayed. The first set of statements can be displayed on a computing device. The computing device can be associated with a user (e.g., a user who uploaded a video).
[0054] Users can prefer music that differs from at least one music track. Users can generate (e.g., in natural language) input indicating their preference for music that differs from at least one music track. Users can generate input (e.g., written or audio) using a keyboard, microphone, touchscreen device, etc. The input can indicate one or more characteristics of the different music preferred by the user. At 710, at least one other music track can be identified. At least one other music track can be identified in response to receiving input indicating that the user prefers music that differs from at least one music track. At least one other music track can be identified based on the input video, previously recommended music tracks, and the input indicating the user's preference. At least one other music track can be identified by a first sub-model. The first sub-model can send an indication of at least one other music track (e.g., multiple music embeddings and multiple titles associated with at least one other music track) to a second sub-model.
[0055] The second sub-model can receive instructions for at least one other musical piece. At point 712, a second set of statements can be generated. This second set of statements can describe at least one other musical piece. The second set of statements can be written in natural language (e.g., a language that naturally evolves in use, in contrast to artificial languages or computer code). Additionally, reasons for recommendation can be generated. These reasons can (e.g., in natural language) explain why at least one other musical piece is recommended. The second set of statements can be generated for display on a computing device.
[0056] Figure 8 The illustration depicts an example process 800 for generating a simulated music recommendation session. Although in Figure 8 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.
[0057] A music video pre-trained (MVP) model can be used to generate a conversational music recommendation dataset. The MVP model can include video branches and music branches. The MVP model can be trained on a dataset consisting of millions of music-video pairs. The MVP model can receive multiple sets of videos and music tracks. At 802, the similarity (e.g., cosine similarity) between each video in the multiple videos and each music track in the music track set can be calculated. The similarity can indicate the similarity between a music track and each video for each candidate music track item. The music from the music track set can then be ranked for each video in descending order of similarity (e.g., the music most similar to the video ranks highest for that video). At 804, candidate music tracks can be selected. Candidate music tracks can correspond to each video in the multiple videos. Candidate music tracks can be selected based on the cosine similarity between each video in the multiple videos and the music track set. For example, a candidate music track corresponding to a specific video could be the music track with the highest similarity to the video.
[0058] At point 806, music tags and music metadata can be determined. Music tracks and music metadata can be associated with the original music track associated with each of the multiple videos, as well as candidate music tracks corresponding to each of the multiple videos. Music tags can effectively summarize a song by providing descriptive keywords covering various elements such as mood, genre, and theme. Music metadata can indicate the title and / or video description of the music video. Music metadata can indicate the official artist name, album details, and / or release date. At point 808, a simulated music recommendation session can be generated. The simulated music recommendation session can be generated based on music tags and music metadata. The simulated music recommendation session can be generated using a generative pre-trained transformer (GPT). For example, a cue word builder can utilize music tags and music metadata to generate cue words to guide GPT in generating a two-round dialogue (e.g., a simulated dialogue) between the user and the music recommendation system. GPT can receive cue words from the cue word builder. GPT can utilize one or more of the cue words, music tags, and music metadata to generate the simulated dialogue. The simulated dialogue can be used to generate a session music recommendation dataset.
[0059] Figure 9 The illustration depicts an example process 900 for training a machine learning model to perform dialogue-based music recommendations on videos. Although in Figure 9 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.
[0060] At point 902, simulated music recommendation data can be generated. Each sample of the simulated music recommendation data (e.g., training sample, instance, etc.) may include information indicating the following: a video, the original music track associated with the video, at least one candidate music track corresponding to the video, and simulated dialogue text. At point 904, a machine learning model can be trained. The machine learning model can be trained on the simulated music recommendation data. The machine learning model may include a first sub-model and a second sub-model. The first sub-model may be trained to pair at least one music track from a music pool to fit the entire video. The first sub-model may be trained to change the music recommendation from a previously recommended music track to the currently recommended music track. The second sub-model may be trained to express the music recommendation and the rationale for the recommendation in natural language (e.g., language that has naturally evolved in use, in contrast to artificial language or computer code).
[0061] Figure 10 Example dialogue 1000 is shown when the target music is retrieved in two rounds. Example dialogue 1000 highlights the versatility and effectiveness of the system described herein. Example dialogue 100 illustrates how the system described herein interacts seamlessly with the user, dynamically adjusting its recommendations based on video content, user preferences, and contextual information about the music.
[0062] In the first round, the first sub-model can process the input video to select one or more music tracks (e.g., clips) from a music pool for recommendation. The first sub-model can send multiple music embeddings and multiple titles associated with the selected music tracks to the second sub-model. The second sub-model can receive the multiple music embeddings and multiple titles associated with the selected music tracks as input. The second sub-model can generate natural language (e.g., words and / or sentences) associated with the multiple music embeddings and multiple titles. For example, the second sub-model can generate a first set of sentences explaining which music tracks(s) are recommended and / or why those tracks are recommended. The first set of sentences can be displayed on the interface of a computing device associated with the user. Figure 10 In the example, the first statement set recommends the music track "Down on My Luck" and explains why the music track "Down on My Luck" is recommended.
[0063] Users may want music tracks different from "Down on My Luck". Users can provide input (e.g., in natural language) that explains their preference for music different from "Down on My Luck". User input can explain one or more features the user wants different music tracks to have. In the second round, the first sub-model can process the input video, the previously recommended music track "Down on My Luck", and the user input to select one or more different music tracks from the music pool for recommendation.
[0064] The first sub-model can send multiple music embeddings and multiple titles associated with the selected different music tracks to the second sub-model. The second sub-model can receive the multiple music embeddings and multiple titles associated with the selected different music tracks as input. The second sub-model can generate natural language (e.g., words and / or sentences) associated with the multiple music embeddings and multiple titles. For example, the second sub-model can generate a second set of sentences explaining the reasoning behind recommending the multiple different music tracks and / or why those multiple different music tracks are recommended. The second set of sentences can be displayed on the interface of a computing device associated with the user. Figure 10 In the example, the second set of statements recommends the music track "Ladder Song" and explains why "Ladder Song" is recommended.
[0065] The comprehensive conversational music recommendation system described in this paper outperforms existing music recommendation systems. Experiments are conducted to evaluate the performance of the comprehensive conversational music recommendation system described in this paper. Each music video clip in multiple 120-second music video clips is divided into twelve 10-second segments, and 5 frames per second are captured from each segment. During the training of the first sub-model 104, each training sample includes a 10-second video clip, the corresponding 10-second original music clip, a 10-second candidate music clip, and a user cue. A CLIP model is used to extract video and text features. An AST model is used to extract audio features. These basic features are transformed into embeddings of size 256 using linear projection for each input type. These embeddings are then processed by applying four transformer encoder layers and multi-head cross-attention layers (each with 16 heads). In the second sub-model 106, the maximum sequence length is limited to 128 and the temperature hyperparameter is set to 0.1.
[0066] The ranking capability of the first sub-model 104 was evaluated. The test set comprised a total of 10,206 music tracks. Each of these tracks was randomly divided into 20 distinct music pools, each containing over 500 tracks. Importantly, each music pool had only one correct track for each video. For track-level testing, embeddings for all 12 segments of each 120-second video and music track were computed. The average of these 12 embeddings was used to create a single representative embedding for each video and each music track. The performance of the first sub-model 104 was evaluated using these average embeddings. In the first round, music was suggested based solely on video features, as it could be assumed that the user had not provided any specific requests at this point. In the second round, user text prompts and candidate music were included along with video features. This setting was used to evaluate the system's ability to modify its initial recommendations based on new information. For both rounds, music tracks were ranked by calculating the cosine similarity between the features of the music in the music pools and the input features. Various metrics were then calculated. These metrics include, for example, Recall@K (K=1, 5, 10), median rank, and "Success Rate at 10" (abbreviated as SR@10). Success Rate at 10 measures the percentage of videos whose correct music tracks appear in the top 10 recommended lists within two rounds. The average performance of each of these metrics is evaluated across the entire test music pool.
[0067] To evaluate the effectiveness of the conversational recommendation system described in this paper (e.g., MuseChat), a robust baseline model with a dual-tower architecture was developed. This baseline model shares the same encoder model with MuseChat for processing video and raw music tracks, but lacks the ability to process text data. Both the baseline and MuseChat are trained using the same dataset and loss function. Figure 11 Example Table 1100 is shown. Table 1100 summarizes the music retrieval results of the baseline multi-round MuseChat. The performance of different models was evaluated under various input conditions. When only visual information is given in the first round, MuseChat, trained on fused features from three different modes, performs comparably to the baseline. However, when additional modes are introduced in the second round, improvements of over 10% are observed across the metrics. As shown in Table 1100, in the second round, MuseChat significantly outperforms the baseline model on all metrics.
[0068] The second sub-model 106 was evaluated. To emphasize the importance of training the second sub-model 106 with both music embeddings and music titles as input, two baseline models were introduced for comparison. The first baseline used a frozen Vicuna-7B model, which is based on the Llama2-7B architecture. Since this model cannot handle music embeddings, only recommended music titles were presented to it. The second baseline utilized the same architecture as the second sub-model 106, but only used music embeddings as input. Various common metrics were used to evaluate the performance of these baseline models and the second sub-model 106 on simulated dialogue.
[0069] Figure 12 Example Table 1200 is shown. Table 1200 illustrates a comparison of semantic similarity between the output and the simulated dialogue using various metrics. BERTScore evaluates lemma-level similarity, while AB divergence, L2 distance, and Fisher-Rao distance are derived based on InfoLM. Figure 12 As shown in Table 1200, the Vicuna-7B model performed the worst. This is mainly because it failed to extract the music title and artist name from the given music video title, thus lacking a comprehensive understanding of the recommended tracks. Even when this information is explicitly provided, the model struggles to grasp the musicality of a given track because it was trained only in text mode. As for the second baseline, while it successfully captured the musical essence of the recommended tracks due to its training on both music and text modes, it still has shortcomings. The model cannot accurately identify the correct music title and artist name based solely on audio information. In contrast, the second sub-model 106, which uses both audio information and music title input, outperforms the baseline, demonstrating the effectiveness of the techniques described in this paper.
[0070] In summary, traditional music recommendation systems primarily focus on providing personalized suggestions through implicit methods, which may not always capture the user's true preferences. The technique described in this paper enables the generation of more accurate, user-tailored recommendation outputs.
[0071] Figure 13 The illustration shows that it can be used in, for example Figure 1 The computing devices used in various aspects of the described services, networks, modules, and / or devices. About Figure 1 In this example architecture, the cloud network (and any of its components), client devices, and / or networks can be independent. Figure 13 This is achieved through one or more instances of the computing device 1300. Figure 13 The computer architecture shown illustrates conventional server computers, workstations, desktop computers, laptop computers, tablet computers, network devices, PDAs, e-readers, digital cellular phones, or other computing nodes, and can be used to perform any aspect of the computer described herein, such as implementing the methods described herein.
[0072] The computing device 1300 may include a substrate or “motherboard,” which is a printed circuit board on which multiple components or devices can be connected via a system bus or other electrical communication path. One or more central processing units (CPUs) 1304 may operate in conjunction with chipset 1306. The CPUs 1304 may be standard programmable processors that perform the arithmetic and logic operations necessary to perform the operation of the computing device 1300.
[0073] Multiple CPUs 1304 can perform necessary operations by manipulating switching elements that distinguish and change these states, transitioning from one discrete physical state to the next. Switching elements typically include electronic circuitry, such as flip-flops, that maintains one of two binary states, and electronic circuitry that provides an output state based on a logical combination of the states of one or more other switching elements, such as logic gates. These basic switching elements can be combined to create more complex logic circuits, including registers, adder-subtractor units, arithmetic logic units, floating-point units, etc.
[0074] The (multiple) CPUs 1304 can be expanded or replaced by other processing units, such as (multiple) GPUs 1305. The (multiple) GPUs 1305 may include processing units specifically designed for, but not necessarily limited to, highly parallel computing, such as graphics and other visualization-related processing.
[0075] Chipset 1306 can provide an interface between CPU(s) 1304 and the remaining components and devices on the substrate. Chipset 1306 can provide an interface for random access memory (RAM) 1308, which serves as the main memory in computing device 1300. Chipset 1306 can also provide an interface for computer-readable storage media such as read-only memory (ROM) 1320 or non-volatile RAM (NVRAM) (not shown) for storing basic routines that can help boot computing device 1300 and transfer information between various components and devices. ROM 1320 or NVRAM can also store other software components required for the operation of computing device 1300 according to the aspects described herein.
[0076] Computing device 1300 can operate in a networked environment using a logical connection to remote computing nodes and computer systems via a local area network (LAN). Chipset 1306 may include functionality for providing network connectivity via a network interface controller (NIC) 1322, such as a Gigabit Ethernet adapter. NIC 1322 may be able to connect computing device 1300 to other computing nodes via network 1316. It should be understood that multiple NICs 1322 may exist in computing device 1300, connecting the computing device to other types of networks and remote computer systems.
[0077] Computing device 1300 can be connected to mass storage device 1328, which provides non-volatile storage for the computer. Mass storage device 1328 can store system programs, application programs, other program modules, and data, as described in more detail herein. Mass storage device 1328 can be connected to computing device 1300 via storage controller 1324 connected via chipset 1306. Mass storage device 1328 may include one or more physical storage units. Mass storage device 1328 may include management component 1313. Storage controller 1324 can interface with physical storage units via Serial Attached SCSI (SAS) interface, Serial Advanced Technology Attached (SATA) interface, Fibre Channel (FC) interface, or other types of interfaces used for physically connecting and transferring data between the computer and physical storage units.
[0078] The computing device 1300 can store data on the mass storage device 1328 by changing the physical state of the physical storage units to reflect that information is being stored. The specific changes in physical state can depend on various factors and the different implementations described herein. Examples of such factors may include, but are not limited to, the technology used to implement the physical storage units and whether the mass storage device 1328 is characterized as a primary storage device or a secondary storage device.
[0079] For example, computing device 1300 can store information in mass storage device 1328 by issuing instructions via storage controller 1324 to change the magnetic properties of a specific location within a disk drive unit, the reflection or refraction properties of a specific location in an optical storage unit, or the electrical properties of a specific capacitor, transistor, or other discrete component in a solid-state storage unit. Other transformations of the physical medium are also possible without departing from the scope and spirit of this specification; the foregoing examples are provided merely for ease of description. Computing device 1300 can also read information from mass storage device 1328 by detecting the physical state or characteristics of one or more specific locations within a physical storage unit.
[0080] In addition to the aforementioned high-capacity storage device 1328, the computing device 1300 can access other computer-readable storage media to store and retrieve information, such as program modules, data structures, or other data. Those skilled in the art will understand that a computer-readable storage medium can be any available medium that provides storage for non-transitory data and can be accessed by the computing device 1300.
[0081] By way of example and not limitation, computer-readable storage media can include volatile and non-volatile, transient and non-transitory computer-readable storage media, as well as removable and non-removable media, implemented in any method or technology. Computer-readable storage media include, but are not limited to, RAM, ROM, erasable programmable ROM (“EPROM”), electrically erasable programmable ROM (“EEPROM”), flash memory or other solid-state memory technologies, compact disc ROM (“CD-ROM”), digital versatile disc (“DVD”), high-definition DVD (“HD-DVD”), Blu-ray or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices, other magnetic storage devices, or any other medium that can be used to store desired information in a non-transitory manner.
[0082] Such as Figure 13 The mass storage device 1328 depicted can store an operating system used to control the operation of the computing device 1300. The operating system may include a version of the Linux operating system. The operating system may include a version of the Windows Server operating system from Microsoft Corporation. According to a further aspect, the operating system may include a version of the UNIX operating system. Various mobile phone operating systems, such as iOS and Android, may also be used. It should be understood that other operating systems may also be used. The mass storage device 1328 can store other systems, applications, and data used by the computing device 1300.
[0083] Mass storage device 1328 or other computer-readable storage medium may also be encoded with computer-executable instructions that, when loaded into computing device 1300, transform the computing device from a general-purpose computing system into a special-purpose computer capable of implementing the aspects described herein. As described above, these computer-executable instructions transform computing device 1300 by specifying how CPU(s) 1304 transition between states. Computing device 1300 can access the computer-readable storage medium storing the computer-executable instructions, which, when executed by computing device 1300, can perform the methods described herein.
[0084] Such as Figure 13The computing device 1300 depicted may further include an input / output controller 1332 for receiving and processing input from multiple input devices such as a keyboard, mouse, touchpad, touchscreen, electronic pen, or other types of input devices. Similarly, the input / output controller 1332 may provide output to a display such as a computer monitor, flat panel display, digital projector, printer, plotter, or other types of output devices. It should be understood that the computing device 1300 may not include... Figure 13 All components shown may include Figure 13 Other components not explicitly shown in the document, or those that can be utilized with Figure 13 The architecture shown is completely different.
[0085] As described in this article, a computing device can be a physical computing device, such as... The computing device 1300. A computing node may also include virtual machine host processes and one or more virtual machine instances. Computer-executable instructions may be indirectly executed by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed in the context of the virtual machine.
[0086] It should be understood that the methods and systems are not limited to any particular method, component, or implementation. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0087] As used in the specification and appended claims, the singular forms “a,” “an,” and “the” include plural indicators unless the context clearly indicates otherwise. A range may be expressed herein as from “about” a particular value, and / or to “about” another particular value. When expressing such ranges, another embodiment includes from one particular value and / or to another particular value. Similarly, when a value is expressed as an approximation using the antecedent “about,” it should be understood that the particular value forms another embodiment. It should also be understood that the endpoints of each range are important both relative to and independent of the other endpoint.
[0088] "Optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes instances where the event or situation occurs and instances where it does not occur.
[0089] Throughout the description and claims of this specification, the word “comprising” and variations thereof, such as “comprising” and “including,” mean “including, but not limited to,” and are not intended to exclude, for example, other components, integers, or steps. “Exemplary” means “an example of…” and is not intended to convey indications of preferred or ideal embodiments. “Like” is not used in a limiting sense but for illustrative purposes.
[0090] Components that can be used to perform the described methods and systems are described. When describing combinations, subsets, interactions, groups, etc., of these components, it should be understood that although specific references to each of the various individual and collective combinations and arrangements of these components may not be explicitly described, each is specifically considered and described herein for all methods and systems. This applies to all aspects of this application, including but not limited to operations in the described methods. Therefore, if various additional operations exist that can be performed, it should be understood that each of these additional operations can be performed using any particular embodiment or combination of embodiments of the described methods.
[0091] The methods and systems of the present invention can be more readily understood by referring to the following detailed description of preferred embodiments and examples included therein, as well as the accompanying drawings.
[0092] Those skilled in the art will understand that the methods and systems may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the methods and systems may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) embodied within the storage medium. More specifically, the methods and systems may take the form of computer software implemented on the web. Any suitable computer-readable storage medium may be used, including hard disks, CD-ROMs, optical storage devices, or magnetic storage devices.
[0093] The following description of embodiments of methods and systems is based on block diagrams and flowcharts of methods, systems, apparatuses, and computer program products. It should be understood that each block in the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, can be implemented by computer program instructions. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute on the computer or other programmable data processing apparatus, create apparatus for implementing the functions specified in one or more flowchart blocks.
[0094] These computer program instructions may also be stored in a computer-readable storage medium that can instruct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including computer-readable instructions for implementing the functions specified in one or more flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus, thereby producing a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowchart blocks.
[0095] The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Furthermore, certain methods or processing blocks may be omitted in some implementations. The methods and processes described herein are not limited to any particular order and may be executed in other suitable orders with respect to their associated blocks or states. For example, described blocks or states may be executed in a different order than specifically described, or multiple blocks or states may be combined in a single block or state. Example blocks or states may be executed serially, in parallel, or in some other manner. Blocks or states may be added to or removed from the described example embodiments. The example systems and components described herein may be configured differently from those described. For example, elements may be added to, removed from, or rearranged from the described example embodiments compared to the described exemplary embodiments.
[0096] It should also be understood that the items are illustrated as being stored in memory or on a storage device when in use, and these items or portions thereof may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, some or all of the software modules and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Furthermore, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as at least in part as firmware and / or hardware, including but not limited to one or more application-specific integrated circuits (“ASICs”), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field-programmable gate arrays (“FPGAs”), complex programmable logic devices (“CPLDs”), etc. Some or all of the modules, systems, and data structures may also be stored (e.g., as software instructions or structured data) on computer-readable media, such as hard disks, memory, networks, or portable media articles, for retrieval by appropriate devices or via appropriate connections. The systems, modules, and data structures can also be transmitted as generated data signals (e.g., as part of a carrier or other analog or digital propagation signal) over various computer-readable transmission media, including wireless and wired / cable-based media, and can take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). In other embodiments, such computer program products can also take other forms. Therefore, the present invention can be implemented using other computer system configurations.
[0097] While methods and systems have been described in conjunction with preferred embodiments and specific examples, they are not intended to limit the scope to the particular embodiments illustrated, as the embodiments herein are intended in all respects to be illustrative rather than restrictive.
[0098] Unless otherwise expressly stated, it is not intended that any method described herein require its operations to be performed in a particular order. Therefore, no order is intended to be inferred in any way where the method claims do not actually describe the order of their operations or where the claims or description do not otherwise specify that the operations will be limited to a particular order. This applies to any possible non-expressive basis of interpretation, including: logical questions concerning the arrangement of steps or flow of operations; general meanings derived from grammatical organization or punctuation; and the number or type of embodiments described in the description.
[0099] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit of this disclosure. Other embodiments will be apparent to those skilled in the art in light of the description and practice described herein. The description and example figures are to be considered exemplary only, and their true scope and spirit are indicated by the appended claims.
Claims
1. A method for implementing dialogue-based music recommendation on input videos, comprising: The video is input into a machine learning model, wherein the machine learning model includes a first sub-model and a second sub-model, and the machine learning model is configured to perform the dialogue-based music recommendation on the input video; The first sub-model identifies at least one music track based on the input video; The second sub-model generates a first set of statements, which describes the at least one music track and the reasons for its recommendation in natural language. The first set of statements is displayed on the computing device associated with the user; In response to receiving input indicating that the user's preference is different from the at least one music track, at least one other music track is identified based on the input video, previously recommended music tracks, and the input indicating the user's preference; as well as Generate a second set of statements describing the at least one other musical piece in natural language for display on the computing device.
2. The method of claim 1, wherein the machine learning model is trained on conversational music recommendation data, and wherein the conversational music recommendation data is generated by: Calculate the cosine similarity between each video in a set of multiple videos and each music track in a set of music tracks; Candidate music tracks corresponding to each of the plurality of videos are selected based on the cosine similarity between each video in the plurality of videos and the music track set; Determine the music tags and music metadata of the original music tracks associated with each of the plurality of videos, as well as the candidate music tracks corresponding to each of the plurality of videos; as well as A simulated music recommendation session is generated using a generative pre-trained transformer (GPT) based on the music tags and the music metadata.
3. The method of claim 2, wherein each of the simulated music recommendation sessions in the simulated music recommendation session begins by recommending one of the candidate music tracks with associated reasons, and then recommends the original music track with reasoning based on identifying musical differences between the one candidate music track and the original music track and in response to simulated input indicating a preference for the original music track.
4. The method according to claim 2, further comprising: Generate simulated music recommendation training data, wherein each sample includes information indicating the following: a video, an original music track associated with the video, at least one candidate music track corresponding to the video, and simulated dialogue text.
5. The method of claim 1, wherein the first sub-model is trained to pair at least one music track from the music pool to fit the entire video, and wherein the first sub-model is trained to change the music recommendation from the previously recommended music track to the currently recommended music track.
6. The method of claim 1, wherein the first sub-model includes a trainable linear projection layer to project video patterns, music patterns, and text patterns into the same embedding space.
7. The method of claim 1, wherein the first sub-model is trained on video data, music data, and text data including simulated music recommendation conversation text.
8. The method of claim 1, wherein the second sub-model is trained to express music recommendations and reasons for recommendations in natural language.
9. The method of claim 1, wherein the second sub-model comprises a trainable linear projection layer, and wherein the second sub-model is trained on data, each instance of the data comprising a music representation from the first sub-model and a corresponding recommendation inference statement from a simulated conversational music recommendation dataset.
10. A system comprising: At least one processor; as well as At least one memory includes computer-readable instructions that, when executed by the at least one processor, cause the system to perform operations including: The video is input into a machine learning model, wherein the machine learning model includes a first sub-model and a second sub-model, and the machine learning model is configured to perform the dialogue-based music recommendation on the input video; The first sub-model identifies at least one music track based on the input video; The second sub-model generates a first set of statements, which describes the at least one music track and the reasons for its recommendation in natural language. The first set of statements is displayed on the computing device associated with the user; In response to receiving input indicating that the user's preference differs from the at least one music track, at least one other music track is identified based on the input video, previously recommended music tracks, and the input indicating the user's preference; and Generate a second set of statements describing the at least one other musical piece in natural language for display on the computing device.
11. The system of claim 10, wherein the machine learning model is trained on conversational music recommendation data, and wherein the conversational music recommendation data is generated by: Calculate the cosine similarity between each video in a set of multiple videos and each music track in a set of music tracks; Candidate music tracks corresponding to each of the plurality of videos are selected based on the cosine similarity between each video in the plurality of videos and the music track set; Determine the music tags and music metadata of the original music tracks associated with each of the plurality of videos, as well as the candidate music tracks corresponding to each of the plurality of videos; as well as A simulated music recommendation session is generated using a generative pre-trained transformer (GPT) based on the music tags and the music metadata.
12. The system of claim 11, wherein each of the simulated music recommendation sessions in the simulated music recommendation session begins by recommending one of the candidate music tracks with associated reasons, and then recommends the original music track with reasoning based on identifying musical differences between the one candidate music track and the original music track and in response to a simulated input indicating a preference for the original music track.
13. The system according to claim 11, further comprising: Generate simulated music recommendation training data, wherein each sample includes information indicating the following: a video, an original music track associated with the video, at least one candidate music track corresponding to the video, and simulated dialogue text.
14. The system of claim 10, wherein the first sub-model is trained to pair at least one music track from the music pool to fit the entire video, and wherein the first sub-model is trained to change the music recommendation from a previously recommended music track to a currently recommended music track.
15. The system of claim 10, wherein the second sub-model is trained to express music recommendations and reasons for recommendations in natural language.
16. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a processor, cause the processor to perform operations, the operations including: The video is input into a machine learning model, wherein the machine learning model includes a first sub-model and a second sub-model, and the machine learning model is configured to perform the dialogue-based music recommendation on the input video; Based on the input video, at least one music track is identified by the first sub-model; The second sub-model generates a first set of statements, which describes the at least one music track and the reasons for its recommendation in natural language. The first set of statements is displayed on the computing device associated with the user; In response to receiving input indicating that the user's preference is different from the at least one music track, at least one other music track is identified based on the input video, previously recommended music tracks, and the input indicating the user's preference; as well as Generate a second set of statements describing the at least one other musical piece in natural language for display on the computing device.
17. The non-transitory computer-readable storage medium of claim 16, wherein the machine learning model is trained on conversational music recommendation data, and wherein the conversational music recommendation data is generated by: Calculate the cosine similarity between each video in a set of multiple videos and each music track in a set of music tracks; Candidate music tracks corresponding to each of the plurality of videos are selected based on the cosine similarity between each video in the plurality of videos and the music track set; Determine the music tags and music metadata of the original music tracks associated with each of the plurality of videos, as well as the candidate music tracks corresponding to each of the plurality of videos; as well as A simulated music recommendation session is generated using a generative pre-trained transformer (GPT) based on the music tags and the music metadata.
18. The non-transitory computer-readable storage medium of claim 17, wherein each of the simulated music recommendation sessions in the simulated music recommendation session begins by recommending one of the candidate music tracks with associated reasons, and then recommends the original music track with reasoning based on identifying musical differences between the one candidate music track and the original music track and in response to a simulated input indicating a preference for the original music track.
19. The non-transitory computer-readable storage medium of claim 16, wherein the first sub-model is trained to pair at least one music track from the music pool to fit the entire video, and wherein the first sub-model is trained to change music recommendations from previously recommended music tracks to currently recommended music tracks.
20. The non-transitory computer-readable storage medium of claim 16, wherein the second sub-model is trained to express music recommendations and reasons for recommendations in natural language.