Systems and methods for video representation learning using triplet training

The system addresses high computational costs in video representation learning by using triplet training to extract and process video, audio, and VAD features, generating embeddings and fingerprints for improved video annotation and retrieval.

JP2025527025APending Publication Date: 2025-08-15VIONLABS AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025511991
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-24
Filing Date
2023-08-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The high computational costs and challenges of unlabeled or inaccurate annotations in video representation learning for large amounts of video data hinder efficient video-based platforms and systems.

Method used

A system and method for video representation learning using triplet training, which extracts video, audio, and valence-arousal-dominance features, processes them through hierarchical and non-local attention networks, and generates embeddings and fingerprints for emotion, genre, and keyword predictions.

Benefits of technology

Enables efficient compact encoding of semantic information in videos, facilitating accurate video annotation, retrieval, and recommendation while reducing computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527025000001_ABST
    Figure 2025527025000001_ABST
Patent Text Reader

Abstract

A system and method for video representation learning using triplet training is provided. The system receives a video file and extracts features associated with the video file, such as video features, audio features, and valence-arousal-dominance (VAD) features. The system processes the video features, audio features, and VAD features using a hierarchical attention network to generate video embeddings, audio embeddings, and VAD embeddings, respectively. The system concatenates the video embeddings, audio embeddings, and VAD embeddings to create concatenated embeddings. The system processes the concatenated embeddings using a non-local attention network to generate a fingerprint associated with the video file. The system then processes the fingerprint to generate one or more of emotion predictions, genre predictions, and keyword predictions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 400,551, filed August 24, 2022, the entire disclosure of which is expressly incorporated herein by reference.

[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to the field of video representation learning. More particularly, this disclosure relates to systems and methods for video representation learning using triplet training. [Background technology]

[0003] With the rapid development in video production and the explosive growth of social media platforms, applications, and websites, video data has become a key component in the provision of online products and streaming services. However, the large amount of video data presents challenges for video-based platforms and systems that store and / or analyze and index large amounts of video content. In this regard, video representation learning can compactly encode semantic information in videos into a lower-dimensional space. The resulting embeddings are useful for video annotation, retrieval, and recommendation problems. However, performing machine learning of video representations remains challenging due to the high computational costs caused by the large amount of data and unlabeled or inaccurate annotations. Therefore, a system and method for video representation learning using triplet training that addresses the above and other needs is desirable. Summary of the Invention [Means for solving the problem]

[0004] The present disclosure relates to a system and method for video representation learning using triplet training. The system receives a video file (e.g., a portion or an entire film, a video clip, a preview video, or other suitable short or long video). The system extracts features associated with the video file. The features may include video features (also referred to as visual features), audio features, and valence-arousal-dominance (VAD) features. The system processes the video features, audio features, and VAD features using a hierarchical attention network to generate video embeddings, audio embeddings, and VAD embeddings, respectively. The system concatenates the video embeddings, audio embeddings, and VAD embeddings to create concatenated embeddings. The system processes the concatenated embeddings using a non-local attention network to generate a fingerprint associated with the video file. The system then processes the fingerprint to generate one or more of emotion predictions, genre predictions, and keyword predictions.

[0005] During the training process, the system generates multiple training samples. The system generates triplet training data associated with the multiple training samples. The triplet training data includes anchor data (e.g., vectors, points, etc.) that are the same as each of the multiple training samples, positive data that is similar to the anchor data, and negative data that is dissimilar to the anchor data. The system uses the triplet training data and a triplet loss (e.g., triplet neighborhood components analysis (NCA) loss) to train a fingerprint generator and / or a classifier. The fingerprint generator includes a hierarchical attention network and a non-local attention network. The triplet NCA loss can encourage the anchor-to-positive distance to be smaller than the anchor-to-negative distance, for example, by minimizing the anchor-to-positive distance while maximizing the anchor-to-negative distance.

[0006] The foregoing features of the present invention will become apparent from the following detailed description of the invention when read in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 illustrates one embodiment of a system of the present disclosure. [Figure 2] 1 is a flowchart illustrating the overall processing steps performed by the system of the present disclosure. [Figure 3] 1 is a flowchart illustrating one embodiment of the overall processing steps performed by the system of the present disclosure. [Figure 4A] 3 is a flowchart showing step 56 of FIG. 2 in more detail. [Figure 4B] 10 is a flowchart illustrating one embodiment of step 56 in more detail. [Figure 5] 3 is a flowchart showing step 60 of FIG. 2 in more detail. [Figure 6A]3 is a flowchart showing step 62 of FIG. 2 in more detail. [Figure 6B] 3 is a flowchart showing step 62 of FIG. 2 in more detail. [Figure 6C] 3 is a flowchart showing step 62 of FIG. 2 in more detail. [Figure 7] 1 is a flowchart illustrating the overall training steps of the present disclosure. [Figure 8A] 1 is a flowchart illustrating an exemplary process for extracting valence-arousal-dominance (VAD) features. [Figure 8B] 1 is a flowchart illustrating an example training process for extracting VAD features. [Figure 8C] FIG. 1 illustrates exemplary color-based VAD features. [Figure 8D] FIG. 1 illustrates exemplary sound-based VAD features. [Figure 9A] FIG. 1 illustrates examples of predicted emotions and video story descriptors from a video file. [Figure 9B] FIG. 1 illustrates examples of predicted emotions and video story descriptors from a video file. [Figure 10] 1 is a schematic diagram illustrating hardware and software components that may be utilized to implement the system of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008] The present disclosure relates to systems and methods for video representation learning using triplet training, as described in detail below in conjunction with FIGS.

[0009] Turning to the drawings, FIG. 1 illustrates one embodiment of a system 10 of the present disclosure. System 10 may be embodied as a central processing unit 12 (processor) in communication with a database 14. Processor 12 may include, but is not limited to, a computer system, a server, a personal computer, a cloud computing device, a smartphone, or any other suitable device programmed to execute the processes disclosed herein. Furthermore, system 10 may be embodied as a customized hardware component, such as a field programmable gate array ("FPGA"), an application specific integrated circuit ("ASIC"), an embedded system, or other customized hardware component, without departing from the spirit or scope of the present disclosure. It should be understood that FIG. 1 is but one possible configuration, and that system 10 of the present disclosure may be implemented using several different configurations.

[0010] Database 14 includes video files (e.g., portions of films or entire films, video clips, preview videos, or other suitable short or long videos) and video data associated with the video files, such as metadata associated with the video files, including, but not limited to, file formats, annotations, various information associated with the video files (e.g., personal information, access information for accessing the video files, subscription information, video length, etc.), volume of the video files, audio data associated with the video files, luminosity data associated with the video files (e.g., color, brightness, lighting, etc.), valence-arousal-dominance (VAD) models, etc. Database 14 also includes training data associated with neural networks for video representation learning (e.g., hierarchical attention networks, non-local attention networks, VAD models, and / or other involved networks or layers). The database 14 may further include one or more outputs from various components of the system 10 (e.g., output from the feature extractor 18a, the video feature module 20a, the audio feature module 20b, the VAD feature module 20c, the fingerprint generator 18b, the hierarchical attention network module 22a, the non-local attention network module 22b, the triplet training module 18c, the application module 18d, and / or other components of the system 10).

[0011] System 10 includes system code 16 (non-transitory computer-readable instructions) stored on a computer-readable medium and executable by hardware processor 12 or one or more computer systems. System code 16 can include various custom software modules that perform the steps / processes described herein, including, but not limited to, feature extractor 18a, video feature module 20a, audio feature module 20b, VAD feature module 20c, fingerprint generator 18b, hierarchical attention network module 22a, non-local attention network module 22b, triplet training module 18c, application module 18d, and / or other components of system 10. System code 16 can be programmed using any suitable programming language, including, but not limited to, C, C++, C#, Java, Python, or any other suitable language. Additionally, system code 16 may be distributed across multiple computer systems that communicate with each other via a communications network, and / or stored and executed on a cloud computing platform and remotely accessed by computer systems that communicate with the cloud platform. System code 16 may be in communication with database 14, which may be stored on the same computer system as system code 16 or on one or more other computer systems that communicate with system code 16.

[0012] 2 is a flowchart illustrating the overall processing steps 50 performed by the system 10 of the present disclosure. Starting at step 52, the system 10 receives a video file. For example, the system 10 may retrieve the video file from the database 14. The system 10 may access a video platform (e.g., a social media platform, a video website, a streaming service platform, etc.) to retrieve the video file. In another example, the system 10 may receive the video file from a user.

[0013] At step 54, system 10 extracts features associated with the video file. The features may include video features, audio features, and VAD features. For example, feature extractor 18a may extract features associated with the video file. Feature extractor 18a may process the video file to extract frame data and audio data from the video. Feature extractor 18a may utilize video feature module 20a, which includes an image feature extractor, to process the frame data and generate video features. Feature extractor 18a may utilize audio feature module 20b, which includes an audio feature extractor, to process the audio data and generate audio features.

[0014] The feature extractor 18a can utilize a VAD feature module 20c to process the audio data and frame data and generate VAD features. The VAD features include features related to "valence," ranging from dissatisfaction to happiness, which expresses pleasantness or unpleasantness toward something; "arousal," which is the level of effective activation ranging from sleepiness to excitement; and "dominance," which reflects the level of control over the emotional state, ranging from submission to dominance. For example, happiness has positive valence, and fear has negative valence. Anger is a high-arousal emotion, and sadness is a low-arousal emotion. Joy is a high-dominance emotion, and fear is a high-submission emotion. The VAD feature module 20c can process the audio data to determine an audio intensity level (e.g., high, medium, low) and the frame data to determine luminosity parameters (e.g., color, brightness, hue, saturation, light, etc.). The VAD features can include audio intensity level and luminosity parameters and / or other suitable features indicative of the VAD determined by the VAD feature module 20c. The VAD feature extraction process, the training process for VAD feature extraction, and examples of VAD features are described with respect to FIGS. 8A-8D.

[0015] At step 56, system 10 uses a hierarchical attention network to process the video features, audio features, and VAD features to generate video embeddings, audio embeddings, and VAD embeddings, respectively. Embeddings refer to low-dimensional data (e.g., low-dimensional vectors) transformed from high-dimensional data (e.g., high-dimensional vectors) so that the low-dimensional data and the high-dimensional data have similar semantic information. For example, fingerprint generator 18b may utilize hierarchical attention network module 22a to process the video features, audio features, and VAD features to generate video embeddings, audio embeddings, and VAD embeddings, respectively. Step 56 is further described in more detail with respect to Figures 4A and 4B.

[0016] In step 58, system 10 concatenates the video embedding, audio embedding, and VAD embedding. For example, system 10 may utilize fingerprint generator 18b to concatenate the video embedding, audio embedding, and VAD embedding to create a concatenated embedding.

[0017] In step 60, system 10 processes the concatenated embeddings using a non-local attention network to generate a fingerprint associated with the video file. The fingerprint refers to a unique feature vector associated with the video file. The fingerprint includes information associated with the audio data, frame data, and VAD data of the video file. The video file can be represented and / or identified by its corresponding fingerprint. Step 60 is further described in more detail with respect to FIG. 5.

[0018] At step 62, system 10 processes the fingerprint to generate emotion predictions, genre predictions, and keyword predictions. For example, system 10 may utilize application module 18d to apply the fingerprint to one or more classifiers (e.g., a stochastic gradient descent (SGD) classifier, a one-vs-all classifier such as a random forest classifier, and a multi-label classifier such as a probabilistic label tree) to predict an emotion associated with the video file (e.g., crime, emotional and inspirational, light-hearted and funny, etc.), a genre associated with the video file (e.g., action, comedy, drama, biopic, etc.), and / or keywords associated with the video file (also referred to as video story descriptors, e.g., thrilling, survival, underdog, etc.). Step 60 is further described in more detail with respect to FIG. 5.

[0019] FIG. 3 is a flowchart illustrating one embodiment 70 of the overall processing steps 50 performed by system 10 of the present disclosure. Starting at step 72, system 10 creates video features, audio features, and VAD features associated with the same video file. At step 74, system 10 inputs the video features, audio features, and VAD features into respective hierarchical attention networks (HANs). At step 76, each hierarchical attention network of system 10 outputs a video embedding, an audio embedding, and a VAD embedding. At step 78, system 10 concatenates the video embedding, the audio embedding, and the VAD embedding. At step 80, system 10 inputs the concatenated embeddings into a non-local attention network (NLA). At step 82, the non-local attention network (NLA) of system 10 outputs a fingerprint for the video file. At step 84, system 10 uses the fingerprint to predict the sentiment, genre, and keywords for the video file.

[0020] FIG. 4A is a flowchart illustrating step 56 of FIG. 2 in more detail. Starting at step 90, system 10 receives extracted features (e.g., video features, audio features, or VAD features described in FIGS. 2 and 3). At step 92, system 10 processes the extracted features using a recurrent neural network (RNN). RNNs are a class of artificial neural networks in which connections between nodes form directed or undirected graphs along a time sequence. RNNs can recognize sequential (or temporal) characteristics of data. At step 94, system 10 chunks the output data from the RNN. The chunking process refers to the process of taking individual data sets and grouping them into larger data sets. At step 96, system 10 applies a time-variant attention process to the chunked data. The time-variant attention process refers to a weighting process that adaptively assigns different weights to its input data. At step 98, system 10 processes the output data from the time-variant attention process using one or more additional RNNs. In step 100, system 10 applies an attention process to the output data from one or more additional RNNs. The attention process refers to a weighting process that assigns weights to its input data, emphasizing some portions of the input data while suppressing other portions of the input data. In step 102, system 10 applies the L of the output data from the attention process. 2 Calculate the norm. L 2 The norm (also called the Euclidean norm) calculates the distance of a vector coordinate from the origin of the vector space. In step 104, system 10 generates embeddings, including video embeddings, audio embeddings, and VAD embeddings (e.g., the video features, audio features, or VAD features described in Figures 2 and 3).

[0021] FIG. 4B is a flowchart illustrating one embodiment 120 of step 56 in more detail. Starting at step 122, system 10 receives extracted feature vectors (e.g., video features, audio features, or VAD features described in FIGS. 2 and 3). At step 124, system 10 processes each extracted feature vector using a recurrent neural network (RNN). At step 126, system 10 chunks the vectors output from the RNN. At step 128, system 10 applies an attention process (e.g., a time-distributed attention process) to weight the chunked vectors. At step 130, system 10 aggregates the weighted vectors to create a combined vector. At step 132, system 10 sequentially uses two additional RNNs to process each combined vector output from the attention process. At step 134, system 10 applies an additional attention process to weight the output vector from the additional RNN. In step 136, system 10 aggregates the weighted vectors output from the additional attention processes to create a final combined vector. In step 138, system 10 calculates the L of the final combined vector. 2 Calculate the norm. In step 140, the system 10 generates an embedding, for example, a video embedding, an audio embedding, or a VAD embedding (for example, the video embedding, audio embedding, or VAD embedding described in FIGS. 2 and 3).

[0022] 5 is a flowchart illustrating step 60 of FIG. 2 in more detail. Starting at step 150, system 10 receives a concatenated embedding (e.g., the concatenated embedding described in FIGS. 2 and 3). In step 152, system 10 applies an attention process to the output data from the attention process by weighting the concatenated embedding. In step 154, system 10 calculates the L of the weighted concatenated embedding. 2Calculate the norm. In step 156, the system 10 generates a fingerprint (eg, the fingerprint described in FIGS. 2 and 3).

[0023] 6A-6C are flowcharts illustrating step 62 of FIG. 2 in more detail. As shown in FIG. 6A, starting at step 160, system 10 inputs fingerprints associated with a video file (e.g., the fingerprints described in FIGS. 2 and 3) into a classifier. For example, system 10 may input the fingerprints into a one-versus-many classifier (e.g., an SGD classifier). In step 162, system 10 predicts one or more emotions associated with the video file. For example, system 10 may use an SGD classifier to place the video file into one or more particular emotion classifications (e.g., criminal, emotional, provocative, light-hearted, funny, etc.).

[0024] As shown in FIG. 6B, starting at step 164, system 10 inputs fingerprints associated with a video file (e.g., fingerprints described in FIGS. 2 and 3) into a classifier. For example, system 10 may input the fingerprints into a one-vs-all classifier (e.g., a random forest classifier). In step 166, system 10 predicts one or more genres associated with the video file. For example, system 10 may use a random forest classifier to place the video file into one or more particular genre classifications (e.g., action, comedy, drama, biography, etc.).

[0025] As shown in FIG. 6C, starting at step 168, system 10 inputs fingerprints associated with a video file (e.g., the fingerprints described in FIGS. 2 and 3) into a classifier. For example, system 10 may input the fingerprints into a multi-label classifier (e.g., a probabilistic label tree). In step 170, system 10 predicts one or more keywords associated with the video file. For example, system 10 may use a probabilistic label tree to label the video file with one or more video story descriptors (e.g., thrilling, survival, underdog, etc.).

[0026] FIG. 7 is a flowchart illustrating the overall training steps 200 of the present disclosure. Starting at step 202, the system 10 generates a plurality of training samples. The triplet training module 18c can construct the training samples using known video files, where the triplet training samples include, but are not limited to, known video features, known audio features, known VAD features, known fingerprints, emotion labels, audio labels, keyword labels, and other suitable known information. Emotion labels refer to labels (e.g., annotations, metadata, strings, text, etc.) that describe specific emotions. Genre labels refer to labels (e.g., annotations, metadata, strings, text, etc.) that describe specific genres. Keyword labels refer to labels (e.g., annotations, metadata, strings, text, etc.) that describe specific keywords that describe a video story. The labels can be generated manually by visualizing the formed fingerprints and clusters and / or collected from various internal or external sources.

[0027] In step 204, system 10 generates triplet training data associated with the plurality of training samples. For example, triplet training module 18c may include a triplet generator for generating anchor data (e.g., vectors, points) identical to the training samples, positive data similar to the anchor data, and negative data dissimilar to the anchor data.

[0028] In step 206, the system 10 uses the triplet training data and the triplet NCA loss to train a fingerprint generator and / or a classifier. The fingerprint generator includes a hierarchical attention network and a non-local attention network. For example, the system 10 can train the fingerprint generator 18b and one or more classifiers of the application module 18d individually / separately or end-to-end. The triplet NCA loss can encourage the anchor-positive distance to be smaller than the anchor-negative distance, for example, by minimizing the anchor-positive distance while maximizing the anchor-negative distance. The system 10 can train the fingerprint generator 18b and one or more classifiers of the application module 18d end-to-end so that the intermediate fingerprints are not universal but rather optimized for a specific application (e.g., emotion prediction, genre prediction, keyword prediction, etc.).

[0029] In step 208, system 10 deploys the trained fingerprint generator and / or the trained classifier for various applications. Examples are described with respect to FIGS.

[0030] FIG. 8A is a flowchart illustrating an example process 220 for extracting VAD features. Starting at step 222, system 10 determines video and audio features from a video file. For example, VAD feature module 20c may extract video and audio features from video feature module 20a and audio feature module 20b. At step 224, system 10 concatenates the video and audio features to create concatenated features. For example, VAD feature module 20c may concatenate the video and audio features to create concatenated features. At step 226, system 10 inputs the concatenated features into a VAD model. The VAD model may generate VAD features. The VAD model may be a neural network regression model having a long short-term memory (LSTM) network with dense layers. For example, VAD feature module 20c may include a VAD model, and / or database 14 may include a VAD model. The VAD feature module 20c may input the concatenated features to a VAD model. In step 228, the system 10 determines VAD features using the VAD model. For example, the VAD model may output VAD features.

[0031] FIG. 8B is a flowchart illustrating an example training process 240 for extracting VAD features. Starting at step 242, system 10 determines a training VAD dataset including VAD labels. For example, database 14 may include a training VAD dataset with VAD labels for one or more video files (e.g., 5-second clips). The VAD labels may be created manually. VAD feature module 20c and / or triplet training module 18c may retrieve the training VAD dataset from database 14. At step 244, system 10 extracts training video features and training audio features from the training VAD dataset. For example, VAD feature module 20c and / or triplet training module 18c may utilize video feature module 20a and audio feature module 20b to determine training video features and training audio features from the training VAD dataset. At step 246, system 10 concatenates the training video features and training audio features to create training concatenated features. For example, VAD feature module 20c and / or triplet training module 18c may concatenate training video features and training audio features to create training concatenated features.

[0032] At step 248, system 10 trains a VAD model based at least in part on the training concatenated features to generate a trained VAD model. For example, VAD feature module 20c and / or triplet training module 18c may optimize a loss function of the VAD model to generate VAD features indicative of VAD labels. At step 250, system 10 deploys the trained VAD model to generate VAD features for unlabeled video files. Examples of VAD features are described with respect to Figures 8C and 8D.

[0033] 8C and 8D show exemplary VAD features 260 and 280 based on color and sound, respectively. As shown in FIG. 8C, color information can be determined throughout the entire video from the beginning of the video to the end of the video. A corresponding stress level 264 can be determined based on the color information. For example, color information and stress level in area 1 can be generated from frame 266A, color information and stress level in area 2 can be generated from frame 266B, and color information and stress level in area 3 can be generated from frame 266C. The color information 262 and / or the corresponding stress level 264 can be used as VAD features 260. As shown in FIG. 8D, tone information 282 in the frequency domain can be determined from sound information in the video. The corresponding stress level 284 can be determined based on sounds / voices / audio from the video. The tone information 282 and / or the corresponding stress level 284 can be used as VAD features 280.

[0034] 9A and 9B show examples of predicted emotions 300 and video story descriptors 320 from a video file. As shown in Figure 9A, different scene emotions (e.g., joy 302, sadness 304, fear 306, and anger 308) are generated for a particular video clip of a movie. As shown in Figure 9B, different keywords (e.g., soldier, war, sacrifice, emotion, opposition) 322 are generated for a particular video clip of a movie.

[0035] FIG. 10 illustrates computer hardware and network components upon which system 400 may be implemented. System 400 may include multiple computational servers 402a-402n having at least one processor (e.g., one or more graphics processing units (GPUs), microprocessors, central processing units (CPUs), tensor processing units (TPUs), application-specific integrated circuits (ASICs), etc.) and memory for executing the computer instructions and methods described above (which may be embodied as system code 16). System 400 may also include multiple data storage servers 404a-404n for storing data. User devices 410 may include, but are not limited to, laptops, smartphones, and tablets for accessing and / or communicating with computational servers 402a-402n and / or data storage servers 404a-404n. System 400 may also include remote computing devices 406a-406n. The remote computing devices 406a-406n may provide various video files. The remote computing devices 406a-406n can include, but are not limited to, a laptop 406a, a computer 406b, and a mobile device 406n having an imaging device (e.g., a camera). The computational servers 402a-402n, the data storage servers 404a-404n, the remote computing devices 406a-406n, and the user device 410 can communicate via a communications network 408. Of course, the system 400 need not be implemented on multiple devices, and in fact, the system 400 can be implemented on a single device (e.g., a personal computer, a server, a mobile computer, a smartphone, etc.) without departing from the spirit or scope of the present disclosure.

[0036] While the systems and methods have thus been described in detail, it should be understood that the foregoing description is not intended to limit the spirit or scope thereof. It will be understood that the embodiments of the present disclosure described herein are merely exemplary, and that those skilled in the art may make any number of variations and modifications without departing from the spirit and scope of the present disclosure. All such variations and modifications, including those described above, are intended to be included within the scope of the present disclosure. What is desired to be protected by Letters Patent is set forth in the following claims.

Claims

1. In a system for video representation learning, a processor configured to receive a video file; system code executed by said processor; wherein the system code causes the processor to: extracting at least one video feature, at least one audio feature, and at least one valence-arousal-dominance (VAD) feature from the video file; processing the at least one video feature, the at least one audio feature, and the at least one VAD feature to generate video embeddings, audio embeddings, and VAD embeddings; concatenating the video embedding, the audio embedding, and the VAD embedding to create a concatenated embedding; processing the concatenated embedding to generate a fingerprint associated with the video file; processing the fingerprint to generate at least one of an emotion prediction, a genre prediction, or a keyword prediction for the video file; A system that allows the following to be performed.

2. 2. The system of claim 1, wherein the system code processes the at least one video feature, the at least one audio feature, and the at least one VAD feature using a hierarchical attention network to generate the video embedding, the audio embedding, and the VAD embedding.

3. The system of claim 1 , wherein the system code processes the concatenated embeddings using a non-local attention network to generate the fingerprint associated with the video file.

4. 2. The system of claim 1, wherein the system code processes the at least one video feature, the at least one audio feature, and the at least one VAD feature by processing the at least one video feature, the at least one audio feature, and the at least one VAD feature using a recurrent neural network (RNN) and chunked output data from the RNN.

5. The system of claim 4 , wherein the system code applies a time-distributed attention process to the chunked data and applies a time-distributed attention process to the chunked data.

6. 6. The system of claim 5, wherein the system code processes output data from the time-distributed attention process using one or more additional RNNs and applies an attention process to output data from the one or more additional RNNs.

7. The system code performs Logging of the output data from the attention process. 2 Calculate the norm, and 2 The system of claim 6 , wherein the norm is used to generate the embedding.

8. The system code applies an attention process to the concatenated embeddings and calculates the L of the output data from the attention process. 2 Calculate the norm, and 2 The system of claim 1 , wherein the connected embedding is processed by using a norm to generate the fingerprint.

9. 2. The system of claim 1, wherein the system code processes the fingerprint to generate the at least one of the emotion prediction, the genre prediction, or the keyword prediction by inputting the fingerprint into a classifier and predicting at least one of the emotion, genre, or keyword for the file.

10. 10. The system of claim 1, wherein the system code generates a plurality of training samples and triplet training data associated with the plurality of training samples, trains a fingerprint generator or a classifier using the triplet training data and a triplet loss, and deploys the trained fingerprint generator and / or the trained classifier.

11. 2. The system of claim 1 , wherein the system code determines video and audio features for the video file, concatenates the video and audio features to create concatenated features, inputs the concatenated features to a VAD model, and determines the at least one VAD feature using the VAD model.

12. 2. The system of claim 1, wherein the system code determines a training VAD dataset including VAD labels, extracts training video features and training audio features from the VAD dataset, concatenates the training video features and the training audio features to create training concatenated features, trains a VAD model based at least in part on the training concatenated features to generate a trained VAD model, and deploys the trained VAD model.

13. A method for video representation learning, extracting at least one video feature, at least one audio feature, and at least one valence-arousal-dominance (VAD) feature from the video file; processing the at least one video feature, the at least one audio feature, and the at least one VAD feature to generate video embeddings, audio embeddings, and VAD embeddings; concatenating the video embedding, the audio embedding, and the VAD embedding to create a concatenated embedding; processing the concatenated embedding to generate a fingerprint associated with the video file; processing the fingerprint to generate at least one of an emotion prediction, a genre prediction, or a keyword prediction for the video file; A method comprising:

14. 14. The method of claim 13, wherein the step of processing the at least one video feature, the at least one audio feature, and the at least one VAD feature further comprises using a hierarchical attention network to generate the video embedding, the audio embedding, and the VAD embedding.

15. The method of claim 13 , wherein the step of processing the concatenated embeddings further comprises using a non-local attention network to generate the fingerprint associated with the video file.

16. 15. The method of claim 14, wherein the step of processing the at least one video feature, the at least one audio feature, and the at least one VAD feature further comprises: processing the at least one video feature, the at least one audio feature, and the at least one VAD feature using a recurrent neural network (RNN) and chunked output data from the RNN.

17. 17. The method of claim 16, further comprising: applying a time-distributed attention process to the chunked data; and applying a time-distributed attention process to the chunked data.

18. 20. The method of claim 17, further comprising the steps of: processing output data from the time-distributed attention process using one or more additional RNNs; and applying an attention process to output data from the one or more additional RNNs.

19. L of output data from the attention process 2 Calculating a norm, and 2 and generating an embedding using the norm.

20. The step of processing the connected embeddings includes applying an attention process to the connected embeddings and L(i) of output data from the attention process. 2 Calculating a norm, and 2 and generating the fingerprint using a norm.

21. 14. The method of claim 13, wherein the step of processing the fingerprint to generate the at least one of the emotion prediction, the genre prediction, or the keyword prediction further comprises inputting the fingerprint into a classifier and predicting at least one of emotion, genre, or keywords for the file.

22. 14. The method of claim 13, further comprising generating a plurality of training samples and triplet training data associated with the plurality of training samples; training a fingerprint generator or a classifier using the triplet training data and triplet loss; and deploying the trained fingerprint generator and / or the trained classifier.

23. 14. The method of claim 13, further comprising determining video and audio features for the video file; concatenating the video and audio features to create concatenated features; inputting the concatenated features into a VAD model; and determining the at least one VAD feature using the VAD model.

24. 14. The method of claim 13, further comprising: determining a training VAD dataset including VAD labels; extracting training video features and training audio features from the VAD dataset; concatenating the training video features and the training audio features to create training concatenated features; training a VAD model based at least in part on the training concatenated features to generate a trained VAD model; and deploying the trained VAD model.