Media retrieval with semi-supervised contrastive learning
A semi-supervised contrastive learning system addresses the challenge of matching music and video retrieval by jointly training encoders with modality-symmetric loss functions, enabling efficient and customizable retrieval through user-controlled emphasis on self-supervised or label-supervised information.
Patent Information
- Application Number
- PCT/US2025/023928
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
Existing music and video retrieval systems struggle to efficiently match music and video based on nuanced descriptors, often requiring manual search and lacking support for highly specific queries, and prioritize metadata over auditory attributes.
A semi-supervised contrastive learning system that jointly trains audio and video encoders using modality-symmetric contrastive loss functions, enabling automatic matching of music and video by leveraging self-supervised and label-supervised learning, with user-controlled emphasis on self-supervised or label-supervised information.
Facilitates efficient and customizable retrieval of matching music or video based on nuanced descriptors, reducing manual effort and improving accuracy by balancing self-supervised and label-supervised learning.
Smart Images

Figure US2025023928_16102025_PF_FP_ABST
Abstract
Description
Docket No.: D24032WO01 MEDIA RETRIEVAL WITH SEMI-SUPERVISED CONTRASTIVE LEARNING 1. Cross-Reference to Related Applications
[0001] This application claims the benefit of Indian Provisional Patent Application No. 202411029608 filed on April 12, 2024, and entitled “A SYSTEM FOR CONTROLLABLE MUSIC-VIDEO RETRIEVAL USING SEMI-SUPERVISED CONTRASTIVE LEARNING.” 2. Field of the Disclosure
[0002] Various example embodiments relate to media retrieval and, more specifically but not exclusively, to finding a suitable piece of music for a given video or finding a suitable video for a given piece of music using machine learning methods. 3. Background
[0003] Synergy between the visuals and music is important for impactful storytelling. However, finding a piece of music that matches the style / genre / emotion of a video or finding a video that matches the style / genre / emotion of a piece of music can be challenging and time- consuming. Accordingly, a system that can automatically match music to video and vice versa is desirable. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS
[0004] Various embodiments provide a media retrieval system that can automatically match music to video and vice versa. In some examples, the media retrieval system includes an audio encoder and a video encoder that are jointly trained using semi-supervised contrastive learning and further using a modality-symmetric contrastive loss function. In operation, the media retrieval system projects an input media dataset (e.g., a video) into a latent space, identifies database latents that are most similar to the projection, and retrieves output media datasets (e.g., audio) corresponding to the identified latents from a media library. In some examples, users of the media retrieval system have the ability to customize the media retrieval process, ensuring that the latter aligns well with their specific creative goals and optimizes the outcome based on the selected balance between self-supervised and label-supervised learning components of the information obtained via the semi-supervised contrastive learning.
[0005] In one example, an apparatus for media retrieval comprises: at least one processor; and at least one memory including program code, wherein the at least one memory and theDocket No.: D24032WO01 program code are configured to, with the at least one processor, cause the apparatus at least to: receive an input media dataset of a first modality; obtain, with a first-modality encoder, a first embedding representing the input media dataset in a latent space; select a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space, each of the second embeddings representing, in the latent space, a respective candidate media dataset of a different second modality, wherein the second embeddings are obtained using a second-modality encoder; and identify one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings, wherein the first-modality encoder and the second-modality encoder are jointly trained using semi-supervised contrastive learning.
[0006] In another example, a method of media retrieval comprises: receiving an input media dataset of a first modality; obtaining, with a first-modality encoder, a first embedding representing the input media dataset in a latent space; selecting a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space, each of the second embeddings representing, in the latent space, a respective candidate media dataset of a different second modality, wherein the second embeddings are obtained using a second-modality encoder; and identifying one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings, wherein the first-modality encoder and the second-modality encoder are jointly trained using semi-supervised contrastive learning.
[0007] According to yet another example, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:
[0009] FIGS.1A-1B are block diagrams illustrating a music-video retrieval functionality according to some examples.Docket No.: D24032WO01
[0010] FIG.2 is a block diagram illustrating a training process for a music-video retrieval system according to some examples.
[0011] FIG.3 is a block diagram illustrating a process of generating a music embeddings database for video-to-music retrieval according to some examples.
[0012] FIG.4 is a block diagram illustrating a video-to-music retrieval process configured to use the music embeddings database generated using the process of FIG.3 according to some examples.
[0013] FIG.5 is a block diagram illustrating a process of generating a video embeddings database for music-to-video retrieval according to some examples.
[0014] FIG.6 is a block diagram illustrating a music-to-video retrieval process configured to use the video embeddings database generated using the process of FIG.5 according to some examples.
[0015] FIG.7 is a flowchart illustrating a method of media retrieval according to some examples.
[0016] FIG.8 is a block diagram illustrating a computing device used to implement at least some of the above-indicated systems, methods, and processes according to some examples. DETAILED DESCRIPTION
[0017] Content creators may use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. In some examples, a process of selecting a piece of music for a given video can be conceptualized as including the following steps: (1) the content creator decides what kind of music best matches their video, and (2) the content creator searches for a piece of music that matches their decided preferences and expectations. For step (1), the content creator typically needs to have a deep understanding of music (musical genres, styles, etc.) as well as some expertise and / or talent pertaining to recognition of what kinds of music might fit well with various types of visual content. Step (2) involves a search for music that matches the decision(s) made at step (1).
[0018] In some examples, a music search engine may operate by matching language queries with music metadata. The metadata may typically include some pertinent details, such as the artist’s name, album title, and / or track title. In some additional examples, a music search engineDocket No.: D24032WO01 may allow for searches that are based on certain musical characteristics, such as the genre or mood. However, such an engine may lack support for highly specific and / or rather general queries. Music search-engine users may typically be limited to a predetermined set of descriptors, such as “jazz” (genre) and “happy” (mood), rather than more nuanced musical descriptors such as, “a cheerful, lively Latin jazz piece featuring saxophone and bass.” In some additional examples, music retrieval systems may tend to prioritize metadata over the auditory attributes of the music itself. Accordingly, even if the content creator knows rather precisely what kind of music is desired for the given video, it can still be challenging and / or time- consuming to find an appropriate piece of music via the above-indicated search tools. Furthermore, when the search results are returned, the content creator may still need to carefully listen through a relatively large number of the retrieved pieces of music to identify the best match.
[0019] At least some of the above-indicated problems in the state of the art can beneficially be addressed using various embodiments disclosed herein. For example, one embodiment provides a system that can automatically recommend a matching piece of music for a given video or a matching video for a given piece of music. Herein, we refer to these processes collectively by the term “music-video retrieval,” which should be construed to encompass both video-to-music and music-to-video retrieval use cases. In some examples, a music-video retrieval system can be viewed as effectively combining the above-mentioned steps (1) and (2) into a single step, as perceived by the user. That is, the music-video retrieval system internally decides what kind of music best matches the given video and then finds and presents to the user one or more recommended music tracks, in some cases listed in a ranked order.
[0020] FIGS.1A-1B are block diagrams illustrating a music-video retrieval functionality according to some examples. More specifically, FIG.1A illustrates a video-to-music retrieval function 110 of the music-video retrieval functionality. FIG.1B similarly illustrates a music-to- video retrieval function 120 of the music-video retrieval functionality. In some examples, both of the retrieval functions 110 and 120 are implemented as different modalities of the same music-video retrieval unit.
[0021] As indicated in FIG.1A, the video-to-music retrieval function 110 operates to retrieve a set 112 of music pieces from a collection 104 of music tracks in response to a video query 102. In some examples, the set 112 includes a plurality of music pieces (tracks) ranked in the order of estimated relevance to the video inputted, as part of the video query 102, into the video-to-music retrieval function 110. In some examples, the collection 104 includes a relativelyDocket No.: D24032WO01 large number of music pieces pre-mapped to a feature space used by the video-to-music retrieval function 110.
[0022] As indicated in FIG.1B, the music-to-video retrieval function 120 operates to retrieve a set 122 of videos from a collection 108 of video tracks in response to a music query 106. In some examples, the set 122 includes a plurality of videos ranked in the order of estimated relevance to the music track inputted, as part of the music query 106, into the music-to-video retrieval function 120. In some examples, the collection 108 includes a relatively large number of video clips pre-mapped to a feature space used by the music-to-video retrieval function 120.
[0023] In some examples, the retrieval functions 110, 120 are implemented using a music- video retrieval framework that is based on both the artistic correspondence and the temporal alignment between music and video tracks. In one example, such music-video retrieval framework is based on a large-scale self-supervised learning approach that does not rely on human annotations and substantially relies only on the inherent correspondence between music and visual elements in music videos. However, in some examples, the music-video retrieval framework may benefit from human annotations, which can sometimes provide useful high-level information that may not be available from pure self-supervised objectives. The music-video retrieval framework may also benefit from the ability to impose explicit user-defined control of the retrieval process.
[0024] At least some of the above-indicated problems can beneficially be addresses using various embodiments disclosed herein. At least some of the disclosed embodiments have some or all of the following features: 1. Semi-supervised contrasting learning for music videos. 1.1. In some examples, a music-video retrieval system combines both self-supervised and supervised training objectives, using a multi-task learning approach. 1.1.1. For self-supervised learning, natural co-occurrence between music audio and video in music videos is leveraged. 1.1.2. For supervised learning, human-annotated music labels are used. 1.2. Self-supervised contrastive learning and supervised contrastive learning are used to learn a joint embedding space between music audio and video. 1.2.1. Contrastive loss functions are modality-symmetric, wherein directional audio-to-video and video-to-audio loss components are summed together with equal weights. 2. A method to enable explicit user control over the retrieval process.Docket No.: D24032WO01 2.1. A model architecture that enables fine-grained control over how much the retrieval process focuses on self-supervised versus label information at inference time. 2.2. This method is believed to be advantageous in that the model only needs to be trained once but supports variously configured retrieval processes, with a user- defined input.
[0025] FIG.2 is a block diagram illustrating a training process for a music-video retrieval system 200 according to some examples. As indicated in FIG.2, the system 200 has a dual- branch architecture for separate processing of audio and video. The audio branch of the system 200 includes an audio encoder 210. The video branch of the system 200 includes a video encoder 250. The audio encoder 210 and the video encoder 250 are jointly trained using a plurality of training video clips 202, each of which includes a respective audio track 204 and a respective video track 206. After being extracted from the video clip 202, the audio track 204 and the video track 206 are fed into the audio encoder 210 and the video encoder 250, respectively, as indicated in FIG.2. A plurality of loss functions computed based on various corresponding embeddings generated in the audio encoder 210 and the video encoder 250 are used to adjust the parameters of various trainable elements of the system 200, e.g., as described in more detail below, until the training stoppage criteria are met.
[0026] Given the audio track 204, ^^, and the video track 206, ^^, of the selected video clip 202, a pretrained audio representation model 212, ^^(∙), and a pretrained video representation model 252, ^^(∙), operate to extract an audio base feature vector 214, ^^^, and a video base feature vector 254, ^^^, respectively. Each audio track 204 is represented by a single respective base feature vector 214 obtained via temporal aggregation of corresponding sequential features. Each video track 206 is similarly represented by a single respective base feature vector 254 obtained via temporal aggregation of corresponding sequential features.
[0027] In some examples, the pretrained audio representation model 212 is the MERT model. The abbreviation “MERT” stands for “Music undERstanding model with large-scale self-supervised Training.” The MERT model is an open-source music audio representation model with a transformer architecture, which is described in more detail, e.g., in the following publication: Yizhi Li, Ruibin Yuan, Ge Zhang, et al., “MERT: ACOUSTIC MUSIC UNDERSTANDING MODEL WITH LARGE-SCALE SELF-SUPERVISED TRAINING,” arXiv:2306.00107v5, which is incorporated herein by reference in its entirety. MERT isDocket No.: D24032WO01 pretrained through large-scale self-supervised learning and produces music representations that show state-of-the-art performance on a wide range of music understanding tasks. In other examples, other pretrained audio representation models can similarly be used to implement the model 212.
[0028] In some examples, the model 212 is configured to ingest 24-kHz raw audio and output seventy-five 1024-dimensional feature vectors per second. To generate the audio base feature vector 214, the model 212 operates to extract the outputs of twenty-five transformer encoder layers, which results in features of the shape (8, 1024). In this manner, the model 212 produces a feature sequence of the shape (75, 8, 1024) for each second of audio. The model 212 then operates to take a global temporal average of the entire feature sequence, resulting in a tensor of the shape (8, 1024) for each audio track 204. The model 212 further operates to compute a learnable weighted average over the transformer layers dimension, which produces the 1024-dimensional feature representation 214. During the training process illustrated in FIG. 2, the model 212 is frozen (remains locked / unchanged) except for the learnable weighted average component thereof.
[0029] In some examples, the pretrained video representation model 252 is the vision component (ViT-B / 32 architecture) of the Contrastive Language-Image Pre-training (CLIP) model. The CLIP model is described, e.g., in A. Radford, J. W. Kim, C. Hallacy, et al., “Learning transferable visual models from natural language supervision,” 38th International Conference on Machine Learning, ICML, 2021, which is incorporated herein by reference in its entirety. The model 252 operates to extract a CLIP image feature for each video frame and then to temporally average the extracted features, resulting in the corresponding single 512- dimensional feature vector 254 for each video. In other examples, other pretrained video representation models can similarly be used to implement the model 252. During the training process illustrated in FIG.2, the model 252 is frozen.
[0030] In the following description, we use the notation ^ to represent both the audio (^) and video (^) branches of the system 200. A person of ordinary skill in the art will readily understand how to apply the provided description to the pertinent components of the audio encoder 210 and the video encoder 250.
[0031] Each of the extracted base features 214, 254 is passed through a respective one of base networks 216 and 256, ^^(∙). One purpose of having the base networks 216, 256 in the corresponding processing paths is to learn general audio / video representations that are sharedDocket No.: D24032WO01 across the downstream self-supervised and supervised tasks. Subsequently, outputs 218, 258 of the base networks 216, 256 are passed through respective task-specific head networks ℎ^^^^(∙) and ℎ^^^^(∙). More specifically, the audio encoder 210 includes the head networks 220 (ℎ^^^^(∙)) and 230 (ℎ^(∙)) co^^^nnected to receive the output 218 of the base network 216. The 250 similarly includes the head networks 260 (ℎ^^^^(∙)) and 270 (ℎ^^^^(∙)) connected to receive the output 258 of the base network 256. The head networks 220, 260, 270 are used to learn ^^embeddings ^^^^and ^^^^that are specific to the self- and supervised tasks, respectively. A self-supervised cross-modal contrastive loss, ^^^^^, operates on the embeddings {^^^^^and ^^^^^}. A supervised cross-modal contrastive loss, ^^^^^, operates on the embeddings {^^^^^and ^^^^^}. In this manner, ^^^^^is encouraged to represent self-supervised information, and ^^^^^is encouraged to represent supervised information. Pertinent details of the contrastive loss functions ^^^^^and ^^^^^will be described below. Based on the respective ones of the inputs 218 and 258, the head networks 220, 230, 260, and 270 generate task-specific embeddings 222 (ℎ^^^^), 232 (ℎ^^^^), 262 (ℎ^^^^), 272 (ℎ^^^^), respectively.
[0032] specific projection networks ^^^^^(∙) and ^^^^^(∙) are used to project the task- specific embeddings 222, 232, 262, and 272 (ℎ^^^^and ℎ^^^^) to a different embedding space where the projections can be combined.^the projection network 224 (^^^^(∙)) operates on the task-specific embedding 222 (ℎ^^^^) to generate a corresponding(^^^^^). The projection network 234 (^^^^^^on the task-specific embedding 232 (ℎ^^^) to generate a corresponding projection 236 (^^^^^). The projection network 264 (^^^^^on the task-specific embedding 262 (ℎ^^^^) to generate a corresponding projection 266 (^^^^^). The projection network 274 (^^^^^on the task-specific embedding 272 (ℎ^^^^) to generate a corresponding projection 276 (^^^^^).
[0033] The projections 226, 236, 266, and 276 are weighted and pairwise linearly combined to generate corresponding combined projections 240 (^^) and 280 (^^), e.g., as follows: ^^ = (1 − ^) ^^^^^ (^^^^^ ) + ^ ^^ ^^^^ ^^^^^ ^ (1)where ^ is thesame in both the audio and video branches of the system 200. In some other examples, the audio branch is configured to use a first combination weight, ^ , and the video branch is configured to use a different second combination weight, ^!. During the training process, ^ is a hyperparameter value that can be set to any selected value in the range between 0 and 1. In some examples, ^ =Docket No.: D24032WO01 0.5 can be used. The value(s) of ^ used during inference (and retrieval) may not necessarily be the same as that (those) used during the training, as will be explained in more detail below.
[0034] In some examples, the system 200 is trained using cross-modal semi-supervised contrastive learning between audio and video. As used herein, the term “semi-supervised contrastive learning” refers to a combination of self-supervised contrastive learning and supervised (e.g., label-supervised) contrastive learning. Self-supervised contrastive learning (SSCL) is a machine learning technique that leverages unlabeled data to learn meaningful representations by training a model to distinguish between similar and dissimilar data instances, without relying on explicit labels. SSCL falls under the umbrella of self-supervised learning, where the model learns from the data itself by creating its own “supervisory signals” or “pretext tasks.” Contrastive learning is generally directed to training a machine-learning (ML) model such as to bring similar data instances closer together in a latent space while pushing dissimilar data instances further apart. Supervised contrastive (SupCon) learning extends self-supervised contrastive learning to a fully supervised setting, using labels to guide the model in learning representations, pulling similar data instances closer and pushing dissimilar data instances further apart in the latent space.
[0035] Herein below, we adopt the following notations: ^"^denotes an embedding of sample# of modality ^; $^" denotes the label of sample # of modality ^; % = {1, … , )} denotes allindices in a training batch; and + denotes the temperature hyperparameter.
[0036] For SSCL, we use a cross-modal version of self-supervised InfoNCE loss. In general, the Information Noise Contrastive Estimation (InfoNCE) loss maximizes the agreement between positive samples and minimizes the agreement between negative samples in the learned representation space. The cross-modal version of the self-supervised InfoNCE loss can be described as follows. Given a batch of ) music videos (each split into the corresponding audio(^) track 204 and video (^) track 206) {(^^" , ^^" )} -", , we compute two different cross-modalself-supervised InfoNCE losses: an^^^^→^^and a video-to-audio loss ^^^^→^^. In one example, the audio-to-video loss ^^^^→^^is expressed as follows: ^^→^ ^ ^ / -345 (678∙ 679)^^^ (^ , ^ ) = ∑ 12^(2) A similar exto-video loss^^^and the video-to-audio loss ^^^^→^^operate to “pull together” audio and video embeddings from the same music video, and “push apart” audio and video embeddings from different musicDocket No.: D24032WO01 videos. In order to create modality-symmetric loss functions, we combine ^^^^→^^and ^^^^→^^with equal weights: ^^^^(^^, ^^) = 0.5 B^^→^^^^ (^^, ^^) + ^^→^^^^ (^^, ^^)C (3)
[0037] Forto the cross-modalsetting, as follows. Given a batch of ) music audio tracks with their labels {(^^ ^ -" , $" )}", and )videos with their labels {(^^ ^ -" , $" )}", , we compute two different cross-modal SupCon losses: anaudio-to-video loss ^^^^→^^and a video-to-audio loss ^^^^→^^. In one example, the audio-to-video loss ^^^^→^^is expressed as follows: ^^→^ ^ / 345 (678∙ 6F9) ^^^ (^ , ^^) = - ∑- ", DE8→9(")D ∑^∈E8→9(") 12^ ∑:∈> 345 ( 68 97 ∙ 6: ⁄ ; )(4) where Gdefined as follows: G^→^(#) = {^ ∈ % | $^ ^" = $" } (5)Similar expressions can be ^ The cross-modallosses operate to “pull together” audio and video embeddings with the same label and “push apart” audio and video embeddings with different labels. In order to maintain modality- symmetric loss functions, we combine ^^^^→^^and ^^^^→^^with equal weights: ^^^^(^^, ^^) = 0.5 B ^^→^^^^ (^^, ^^) + ^^→^^^^ (^^, ^^)C (6)
[0038] As indicated in FIG.2, two different loss functions, ^^^^^and ^^^^^, operate on thetask-specific embeddings {ℎ^^^^ , ℎ^^^^ } and Iℎ^^^^ , ℎ^^^^ J, respectively:^^^^^ = ^^^^(^^ ^^^^ , ^^^^ ) (7a)^^ ^ ^^^^ = ^^^^^^^^^ , ^^^^ ^ (7b)As a result, {^^ , ^^^^^ } areinformation,I^^^^^ , ^^^^^ J are trained to primarily contain supervised information.As indicated in FIG.2, two different loss functions, ^6^^^and ^6^^^, operate on thecombined projections (embeddings) 240, 280 {^^, ^^}. In some examples, these loss functionsare as follows: ^6^^^ = ^^^^(^^, ^^) (8a)As a result, the system 200 isthe self-supervised informationfrom {^^ , ^^ } an ^ ^^^^ ^^^ d the supervised information from I^^^^ , ^^^^ J.Docket No.: D24032WO01
[0040] In some examples, the system 200 is trained using a multi-task learning approach designed to balance the multiple tasks / objectives mentioned above. Under this approach, a total loss function, ^KLKM^, to be optimized during the training process is constructed as a weighted sum of various individual loss functions described above. In one example, the total loss function can be expressed as follows: ^= N N 6 6 ^ ^KLKM^ 6 ^ ^^^^^^^ + N^^^^^^^ ^ + N^ ^N^^^^^^^ + N^^^^^^^ ^ (9)where the on the task-specific are used to control the balance between the self-supervised and supervised training objectives.
[0041] After the above-described training of the audio encoder 210 and the video encoder 250 is completed, the encoders 210, 250 can be used to generate video and audio (music) embeddings, denoted as ^^and ^^, for retrieval purposes. In some examples, a user can set and / or change the parameter(s) ^ used in the encoders 210, 250 to controllably adjust the emphasis placed on different types of information within the embedding vectors 240 (^^) and 280 (^^). For example, when ^ is set to 0, the embeddings ^^and ^^emphasize the self- supervised information learned by the system 200 during the training process, thereby diminishing the influence of the label-supervised content. On the other hand, when ^ is set to 1, the embeddings ^^and ^^emphasize the label-supervised content while diminishing the influence of the self-supervised information. By having the ability to choose any value of ^ from the range between 0 and 1, the user can advantageously fine-tune the balance between the self-supervised and label-supervised information in the retrieval process. This controllability confers a “customizable” property on the retrieval mechanism.
[0042] FIG.3 is a block diagram illustrating a process 300 of generating a music embeddings database 308 for video-to-music retrieval according to some examples. The process 300 employs the audio encoder 210 that has been trained as described above in reference to FIG.2. The generated music embeddings database 308 can be used for video-to-music retrieval, e.g., as described in more detail below in reference to FIG.4.
[0043] In some examples, the process 300 includes selecting the value of ^, which is then provided as a configuration parameter to the audio encoder 210. Recall that the value of ^ controls the relative weighting between self-supervised and label-supervised information in the embedding vectors, as previously described. Various options for selecting the value of ^ in theDocket No.: D24032WO01 process 300 include but are not limited to: (i) user selection; (ii) automatic or algorithmic selection; and (iii) a default value, e.g., the same value as that used during the training process.
[0044] In some examples, the music embeddings database 308 is generated as follows. Each music track 304 from a music library 302 is processed with the audio encoder 210 to generate a corresponding embedding (latent space vector) 306. The embedding 306 is analogous to the projection 240 (^^) described above in reference to FIG.2. A plurality of embedding 306 corresponding to a plurality of music tracks 304 from the music library 302 and generated in this manner is then sorted and organized in a suitable searchable format to create the music embeddings database 308.
[0045] FIG.4 is a block diagram illustrating a video-to-music retrieval process 400 configured to use the music embeddings database 308 according to some examples. The process 400 employs the video encoder 250 that has been trained as described above in reference to FIG. 2. In response to a video query 402, the process 400 generates a list 412 of K top-ranked music tracks of the music library 302. In various examples, the number K is a user-selected parameter that can be set to, for example, K = 1, 5, or 10. Using the list 412, the actual corresponding music tracks 304 can then be retrieved from the music library 302 (also see FIG.3).
[0046] Given the video track of the video query 402, the video encoder 250 operates to generate a corresponding video embedding 404. The video embedding 404 is analogous to the projection 280 (^^). A ranking module 410 operates to compare the video embedding 404 against the music embeddings 306 of the music embeddings database 308 using a suitable similarity metric and calculates similarity scores between the video embedding 404 and each of the music embeddings 306. In some examples, the used similarity metric is the cosine similarity metric. In other examples, other suitable similarity metrics can also be used. The music tracks 304 corresponding to the top K most similar embeddings 306 are then listed, in the ranked order, in the list 412 and retrieved from the music library 302.
[0047] FIG.5 is a block diagram illustrating a process 500 of generating a video embeddings database 508 for music-to-video retrieval according to some examples. The process 500 employs the video encoder 250 that has been trained as described above in reference to FIG.2. The generated video embeddings database 508 can be used for music-to-video retrieval, e.g., as described in more detail below in reference to FIG.6.
[0048] In some examples, the process 500 includes selecting the value of ^, which is then provided as a configuration parameter to the video encoder 250. As previously indicated, theDocket No.: D24032WO01 value of ^ controls the relative weighting between self-supervised and label-supervised information in the embedding vectors. Various options for selecting the value of ^ in the process 500 include but are not limited to: (i) user selection; (ii) automatic or algorithmic selection; and (iii) a default value, e.g., the same value as that used during the training process.
[0049] In some examples, the video embeddings database 508 is generated as follows. Each video track (clip) 504 from a video library 502 is processed with the video encoder 250 to generate a corresponding embedding (latent space vector) 506. The embedding 306 is analogous to the projection 280 (^^) described above in reference to FIG.2. A plurality of embedding 506 corresponding to a plurality of video tracks 504 from the video library 502 and generated in this manner is then sorted and organized in a suitable searchable format to create the video embeddings database 508.
[0050] FIG.6 is a block diagram illustrating a music-to-video retrieval process 600 configured to use the video embeddings database 508 according to some examples. The process 600 employs the audio encoder 210 that has been trained as described above in reference to FIG. 2. In response to a music query 602, the process 600 generates a list 612 of K top-ranked video tracks of the video library 502. In various examples, the number K is a user-selected parameter that can be set to, for example, K = 1, 5, or 10. Using the list 612, the actual corresponding video tracks 504 can then be retrieved from the video library 502 (also see FIG.5).
[0051] Given the music track of the audio query 602, the audio encoder 210 operates to generate a corresponding audio embedding 604. The audio embedding 604 is analogous to the projection 240 (^^). A ranking module 610 operates to compare the audio embedding 604 against the audio embeddings 506 of the video embeddings database 508 using a suitable similarity metric and calculates similarity scores between the audio embedding 604 and each of the audio embeddings 506. In some examples, the used similarity metric is the cosine similarity metric. In other examples, other suitable similarity metrics can also be used. The video tracks 604 corresponding to the top K most similar embeddings 506 are then listed, in the ranked order, in the list 612 and retrieved from the video library 502.
[0052] The value of ^ is an important factor in controlling the balance between the self- supervised and label-supervised information in the embeddings. As such, control over the value of ^ enables control over the cross-modal retrieval process. For example, when ^ is set to 0, the process 400 emphasizes self-supervised similarity, thereby retrieving music tracks 304 that most closely match the patterns learned independently by the model during training. Conversely,Docket No.: D24032WO01 when ^ is set to 1, the process 400 primarily focuses on the label-supervised similarity, thereby retrieving music tracks 304 that align closely with the labeled data. Similar observations apply to the process 600. By selecting a value of ^ between 0 and 1, the system allows the retrieval process to benefit from the strengths of both self-supervised and label-supervised learning paradigms. This flexibility beneficially provides the users with the ability to fine-tune the retrieval processes 400, 600 to meet their specific needs, adapting to different contexts where one type of information might be more valuable than the other. In some examples, this controllability provides the users with the power to customize the retrieval process, ensuring that the retrieval process aligns well with their specific goals and optimizes the outcome based on the selected balance between self-supervised and label-supervised learning.
[0053] FIG.7 is a flowchart illustrating a method 700 of media retrieval according to some examples. The method 700 is described below with continued reference to FIGS.2-7.
[0054] A block 702 of the method 700 includes setting a value of the weight parameter ^ (e.g., see the inputs labeled “^” in each of FIGS.3-6). In some examples, operations of the block 702 include receiving a first value of the weight parameter ^. In some examples, the received first value is provided via a user input.
[0055] A block 704 of the method 700 includes receiving an input media dataset (e.g., 402, FIG.4, or 602, FIG.6) of a first modality. The first modality is selected from the group consisting of an audio (e.g., music) modality and a video modality. In some examples, the input media dataset is an audio track. In some other examples, the input media dataset is a video track.
[0056] A block 706 of the method 700 includes obtaining, with a first-modality encoder, a first embedding (e.g., 404 or 604, FIGS.4, 6) representing the input media dataset in a latent space. In some examples, the first-modality encoder is the audio encoder 210. In some other examples, the first-modality encoder is the video encoder 250.
[0057] In some examples, the weight parameter ^ controls relative emphasis applied by the first-modality encoder (e.g., one of 210, 250) to first and second components of a first embedding. The first component corresponds to information learned by the first-modality encoder via self-supervised learning. The second component corresponds to information learned by the first-modality encoder via label-supervised learning. In some examples, the weight parameter ^ is limited to a numerical range [0, 1].Docket No.: D24032WO01
[0058] A block 708 of the method 700 includes selecting a subset of second embeddings (e.g., 412 or 612, FIGS.4, 6) based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space. Each of the second embeddings represents, in the latent space, a respective candidate media dataset (e.g., 306 or 506, FIGS.3, 6) of a different second modality. The second embeddings are obtained using a second-modality encoder. In some examples, the second-modality encoder is the audio encoder 210. In some other examples, the second-modality encoder is the video encoder 250. The first modality and the different second modality are selected from the group consisting of an audio modality and a video modality. In some examples, the similarity metric is a cosine metric.
[0059] In some examples, the weight parameter ^ further controls relative emphasis applied by the second-modality encoder to first and second components of a second embedding. The first component of the second embedding corresponds to information learned by the second- modality encoder via self-supervised learning. The second component of the second embedding corresponds to information learned by the second-modality encoder via label-supervised learning. In some examples, the first-modality encoder and the second-modality encoder are jointly trained using a second value of the weight parameter ^ that is different from the first value. In some examples, the first embedding is computed as a weighted sum of the first component and the second component. The respective weighting coefficients applied to the first and second components are determined based on the first value of the weight parameter ^.
[0060] In some examples, operations of the block 708 include receiving a size K of the subset of second embeddings, where K is a positive integer smaller than twenty. In some examples, K=1.
[0061] In some examples, the first-modality encoder and the second-modality encoder are jointly trained using a modality-symmetric contrastive loss function (e.g., see Eqs. (2)-(9)). In some examples, the modality-symmetric contrastive loss function is a weighted sum of a first loss function operating on task-specific embeddings and a second loss function operating on latent-space embeddings (e.g., see Eq. (9)). The first loss function is a weighted sum of a cross- modal self-supervised learning loss operating on the task-specific embeddings and a cross-modal label-supervised learning loss operating on the task-specific embeddings. The second loss function is a weighted sum of a cross-modal self-supervised learning loss operating on the latent- space embeddings and a cross-modal label-supervised learning loss operating on the latent-space embeddings. In some examples, values of at least a subset of weighting coefficients used in the weighting sums are selectable based on a training objective.Docket No.: D24032WO01
[0062] A block 710 of the method 700 includes identifying one or more output media datasets of the different second modality among the respective candidate media datasets based on the subset of second embeddings selected in the block 708. In some examples, each of the one or more output media datasets is a video track. In some other examples, each of the one or more output media datasets is an audio track.
[0063] A block 712 of the method 700 includes retrieving the one or more output media datasets identified in the block 710 from a media library. In some examples, the media library is the music library 302 (FIG.3). In some other examples, the media library is the video library 502 (FIG.5).
[0064] FIG.8 is a block diagram illustrating a computing device 800 one or more instances of which can be used to implement at least some of the above-described systems, methods, and processes according to some examples. The computing device 800 of FIG.8 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 800 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 802 and one or more storage devices 804). Additionally, in various embodiments, the computing device 800 may not include one or more of the components illustrated in FIG.8, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High- Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 800 may not include a display device 810, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 810 may be coupled.
[0065] The computing device 800 includes a processing device 802 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device 802 may include one or more digital signal processors (DSPs), application-specific integrated circuitsDocket No.: D24032WO01 (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.
[0066] The computing device 800 also includes a storage device 804 (e.g., one or more storage devices). In various embodiments, the storage device 804 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 804 may include memory that shares a die with the processing device 802. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 804 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 802), cause the computing device 800 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.
[0067] The computing device 800 further includes an interface device 806 (e.g., one or more interface devices 806). In various embodiments, the interface device 806 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 800 and other computing devices. For example, the interface device 806 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 800. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 806 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 806 for managing wireless communications may operate in accordance with a Global System for Mobile CommunicationDocket No.: D24032WO01 (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 806 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 806 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 806 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.
[0068] In some embodiments, the interface device 806 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 806 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 806 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 806 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 806 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 806 may be dedicated to wireless communications, and a second set of circuitry of the interface device 806 may be dedicated to wired communications.
[0069] The computing device 800 also includes battery / power circuitry 808. In various embodiments, the battery / power circuitry 808 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 800 to an energy source separate from the computing device 800 (e.g., to AC line power).
[0070] The computing device 800 also includes a display device 810 (e.g., one or multiple individual display devices). In various embodiments, the display device 810 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.Docket No.: D24032WO01
[0071] The computing device 800 also includes additional input / output (I / O) devices 812. In various embodiments, the I / O devices 812 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.
[0072] Depending on the specific embodiment, various components of the interface devices 806 and / or I / O devices 812 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 806 and / or I / O devices 812 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 802 and / or the storage device 804. In some additional examples, the interface devices 806 and / or I / O devices 812 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 802 and / or the storage device 804 into an analog form suitable for being transmitted through a communication channel.
[0073] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS.1-8, provided is an apparatus for media retrieval comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive an input media dataset of a first modality; obtain, with a first-modality encoder, a first embedding representing the input media dataset in a latent space; select a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space, each of the second embeddings representing, in the latent space, a respective candidate media dataset of a different second modality, wherein the second embeddings are obtained using a second- modality encoder; and identify one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings, wherein the first-modality encoder and the second-modality encoder are jointly trained using semi-supervised contrastive learning.
[0074] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS.1-8, provided is a method of media retrieval comprising: receiving an input media dataset of a first modality;Docket No.: D24032WO01 obtaining, with a first-modality encoder, a first embedding representing the input media dataset in a latent space; selecting a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space, each of the second embeddings representing, in the latent space, a respective candidate media dataset of a different second modality, wherein the second embeddings are obtained using a second- modality encoder; and identifying one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings, wherein the first-modality encoder and the second-modality encoder are jointly trained using semi-supervised contrastive learning.
[0075] In some embodiments of the above method, the first modality and the different second modality are selected from the group consisting of an audio modality and a video modality.
[0076] In some embodiments of any of the above methods, the input media dataset is an audio track; and wherein each of the one or more output media datasets is a video track.
[0077] In some embodiments of any of the above methods, the input media dataset is a video track; and wherein each of the one or more output media datasets is an audio track.
[0078] In some embodiments of any of the above methods, the method further comprises retrieving the identified one or more output media datasets from a media library.
[0079] In some embodiments of any of the above methods, the selecting comprises receiving a size K of the subset of second embeddings, where K is a positive integer smaller than twenty.
[0080] In some embodiments of any of the above methods, the method further comprises receiving a first value of a weight parameter, wherein the weight parameter controls relative emphasis applied by the first-modality encoder to first and second components of the first embedding, the first component corresponding to information learned by the first-modality encoder via self-supervised learning, the second component corresponding to information learned by the first-modality encoder via label-supervised learning, where the weight parameter is limited to a numerical range [0, 1].
[0081] In some embodiments of any of the above methods, the weight parameter further controls relative emphasis applied by the second-modality encoder to first and second components of a second embedding, the first component of the second embedding correspondingDocket No.: D24032WO01 to information learned by the second-modality encoder via self-supervised learning, the second component of the second embedding corresponding to information learned by the second- modality encoder via label-supervised learning.
[0082] In some embodiments of any of the above methods, the first-modality encoder and the second-modality encoder are jointly trained using a second value of the weight parameter that is different from the first value.
[0083] In some embodiments of any of the above methods, the received first value is provided via a user input.
[0084] In some embodiments of any of the above methods, the obtaining comprises computing the first embedding as a weighted sum of the first component and the second component; and wherein respective weighting coefficients applied to the first and second components are determined based on the first value.
[0085] In some embodiments of any of the above methods, the first-modality encoder and the second-modality encoder are jointly trained using a modality-symmetric contrastive loss function.
[0086] In some embodiments of any of the above methods, the modality-symmetric contrastive loss function is a weighted sum of a first loss function operating on task-specific embeddings and a second loss function operating on latent-space embeddings.
[0087] In some embodiments of any of the above methods, the first loss function is a weighted sum of a cross-modal self-supervised learning loss operating on the task-specific embeddings and a cross-modal label-supervised learning loss operating on the task-specific embeddings.
[0088] In some embodiments of any of the above methods, the second loss function is a weighted sum of a cross-modal self-supervised learning loss operating on the latent-space embeddings and a cross-modal label-supervised learning loss operating on the latent-space embeddings.
[0089] In some embodiments of any of the above methods, values of at least a subset of weighting coefficients used in the weighting sums are selectable based on a training objective.Docket No.: D24032WO01
[0090] In some embodiments of any of the above methods, the similarity metric is a cosine metric.
[0091] Some embodiments provide a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.
[0092] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0093] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0094] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0095] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted asDocket No.: D24032WO01 reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
[0096] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.
[0097] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0098] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.
[0099] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.
[0100] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.Docket No.: D24032WO01
[0101] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.
[0102] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”
[0103] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.
[0104] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”
[0105] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.
[0106] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable ofDocket No.: D24032WO01 communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.
[0107] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0108] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or aDocket No.: D24032WO01 similar integrated circuit in server, a cellular network device, or other computing or network device.
[0109] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0110] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
Docket No.: D24032WO01 CLAIMS What is claimed is:
1. A method of media retrieval, comprising: receiving an input media dataset of a first modality; obtaining, with a first-modality encoder, a first embedding representing the input media dataset in a latent space; selecting a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space, each of the second embeddings representing, in the latent space, a respective candidate media dataset of a different second modality, wherein the second embeddings are obtained using a second-modality encoder; and identifying one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings, wherein the first-modality encoder and the second-modality encoder are jointly trained using semi-supervised contrastive learning.
2. The method of claim 1, wherein the first modality and the different second modality are selected from the group consisting of an audio modality and a video modality.
3. The method of claim 2, wherein the input media dataset is an audio track; and wherein each of the one or more output media datasets is a video track.
4. The method of claim 2, wherein the input media dataset is a video track; and wherein each of the one or more output media datasets is an audio track.
5. The method of claim 1, further comprising retrieving the identified one or more output media datasets from a media library.
6. The method of claim 1, wherein the selecting comprises receiving a size K of the subset of second embeddings, where K is a positive integer smaller than twenty.Docket No.: D24032WO01 7. The method of claim 6, wherein K=1.
8. The method of claim 1, further comprising receiving a first value of a weight parameter, wherein the weight parameter controls relative emphasis applied by the first-modality encoder to first and second components of the first embedding, the first component corresponding to information learned by the first-modality encoder via self-supervised learning, the second component corresponding to information learned by the first-modality encoder via label-supervised learning, where the weight parameter is limited to a numerical range [0, 1].
9. The method of claim 8, wherein the weight parameter further controls relative emphasis applied by the second-modality encoder to first and second components of a second embedding, the first component of the second embedding corresponding to information learned by the second-modality encoder via self-supervised learning, the second component of the second embedding corresponding to information learned by the second-modality encoder via label- supervised learning.
10. The method of claim 9, wherein the first-modality encoder and the second-modality encoder are jointly trained using a second value of the weight parameter that is different from the first value.
11. The method of claim 8, wherein the received first value is provided via a user input.
12. The method of claim 8, wherein the obtaining comprises computing the first embedding as a weighted sum of the first component and the second component; and wherein respective weighting coefficients applied to the first and second components are determined based on the first value.
13. The method of claim 1, wherein the first-modality encoder and the second-modality encoder are jointly trained using a modality-symmetric contrastive loss function.
14. The method of claim 13, wherein the modality-symmetric contrastive loss function is a weighted sum of a first loss function operating on task-specific embeddings and a second loss function operating on latent-space embeddings.Docket No.: D24032WO01 15. The method of claim 14, wherein the first loss function is a weighted sum of a cross- modal self-supervised learning loss operating on the task-specific embeddings and a cross-modal label-supervised learning loss operating on the task-specific embeddings.
16. The method of claim 15, wherein the second loss function is a weighted sum of a cross- modal self-supervised learning loss operating on the latent-space embeddings and a cross-modal label-supervised learning loss operating on the latent-space embeddings.
17. The method of claim 16, wherein values of at least a subset of weighting coefficients used in the weighting sums are selectable based on a training objective.
18. The method of claim 1, wherein the similarity metric is a cosine metric.
19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of claim 1.
20. An apparatus for media retrieval, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive an input media dataset of a first modality; obtain, with a first-modality encoder, a first embedding representing the input media dataset in a latent space; select a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space, each of the second embeddings representing, in the latent space, a respective candidate media dataset of a different second modality, wherein the second embeddings are obtained using a second-modality encoder; and identify one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings, wherein the first-modality encoder and the second-modality encoder are jointly trained using semi-supervised contrastive learning.
Citation Information
Patent Citations
Media diversity recommendation method and device, computer equipment and storage medium
CN117312646A