Captioning of videos using multiple cross-modality teachers

KR1020260139128APending Publication Date: 2026-09-21SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020267025026
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2025-01-21
Publication Date
2026-09-21

Smart Images

  • Figure P1020267025026_ABST
    Figure P1020267025026_ABST
Patent Text Reader

Abstract

Automatic captioning pipelines and methods for automatically annotating video data with subtitles obtainable using Automatic Speech Recognition (ASR). An automatic captioning pipeline with multimodal data inputs extends a dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large-scale video-language dataset, along with an automatic captioning approach that utilizes multimodal inputs such as text video descriptions, subtitles, and individual video frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application claims priority to U.S. application serial number 18 / 422,721 filed on January 25, 2024, the contents of which are incorporated herein in their entirety by reference.

[0002] The subject of the present invention relates to the automatic captioning of large-scale visual data, such as videos, using text descriptions. Background Technology

[0003] The quality of the data and annotations provides an upper limit on the quality of the downstream model. While large-scale text corpora and image-text pairs exist, high-quality video-text data remains difficult to collect. Brief explanation of the drawing

[0004] In the drawings, not necessarily to scale, and similar components may be described by similar reference numerals in different drawings. Some non-limiting examples are illustrated in the drawings of the attached drawings.

[0005] FIG. 1 includes examples of captioning using a large dataset according to the present disclosure and examples depicting an exemplary dataset;

[0006] FIG. 2 is a flowchart illustrating the steps of an automatic captioning method using a pipeline;

[0007] Figure 3 is a chart containing a list of eight teacher models showing excellent performance;

[0008] FIG. 4 is an example depicting a student captioning model including a vision branch and a text branch utilizing multimodal inputs including vision and text;

[0009] Figure 5 is an example illustrating captioning predictions with and without text inputs;

[0010] Figure 6 is an example depicting examples of generated video samples;

[0011] Figure 7 is a chart depicting a distribution plot of video lengths after video clip splitting;

[0012] FIG. 8 is a block diagram of a machine in which commands for performing one or more of the methodologies described in this specification can be executed. Specific details for implementing the invention

[0013] The present disclosure includes examples of automatic captioning pipelines and methods for automatically annotating video data with captions that can be obtained using Automatic Speech Recognition (ASR). An automatic captioning pipeline having inputs of multimodal data expands a dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset in conjunction with an automatic captioning approach that utilizes multimodal inputs such as textual video descriptions, captions, and individual video frames.

[0014] Multiple high-resolution videos can be curated from publicly available datasets, where they are segmented into semantically consistent video clips, and then multiple cross-modality teacher models are applied to obtain captions for each video clip. A retrieval model is fine-tuned for a relatively small subset of manually selected video clips to determine the optimal caption for each video clip, and then the retrieval model is applied to the entire dataset to select the optimal caption as annotation.

[0015] In the examples described herein, 70 million (70M) video pairs are generated with high-quality text captions and are generally referred to in this disclosure as the Panda-70M dataset. The value of the dataset is exemplified for three downstream tasks: video captioning, video and text retrieval, and text-driven video generation. Models trained on the data achieve substantially higher scores in most metrics across all tasks.

[0016] Additional purposes, advantages, and novel features of the examples will be presented in part in the following description, and may become apparent to a person skilled in the art upon review of the following description and the accompanying drawings, or may be learned through the production or operation of the examples. The purposes and advantages of the subject matter of the invention may be realized and achieved by the methodologies, means, and combinations specifically indicated in the appended claims.

[0017] In the detailed description below, numerous specific details are presented by example to provide a thorough understanding of the relevant teachings. However, it should be apparent to a person skilled in the art that these teachings can be practiced without such details. In other instances, to avoid unnecessarily obscuring aspects of these teachings, well-known methods, procedures, components, and circuits have been described at a relatively high level without detail.

[0018] As used herein, the term “coupled” refers to any logical, optical, physical, or electrical connection, link, or similarity in which signals or light generated or supplied by one system element are transmitted to another coupled element. Unless otherwise described, coupled elements or devices do not necessarily have to be directly connected to one another and may be separated by intermediate components, elements, or communication media capable of modifying, manipulating, or transmitting light or signals.

[0019] Now, refer in detail to the examples illustrated in the attached drawings and discussed below.

[0020] Large-scale multimodal learning generally requires very large amounts of data and high computational demands. Most breakthroughs are achieved by large-scale computing infrastructure, large-scale models, and large-scale data. Due to these three components, powerful text-image and image-text models exist. Scaling model size or computational demands is difficult and costly, but these challenges can generally be addressed within a finite amount of engineering time. Scaling data is relatively more difficult because it takes time for humans to analyze each sample.

[0021] The computing community has collected millions, even billions, of datasets representing image-text pairs. In comparison, video-text pairs are more difficult to acquire. First, annotating videos takes more time because the annotator must watch the entire video before labeling. Second, videos often consist of multiple scenes or events spliced ​​together and contain temporally changing content. Finally, meta-information itself, such as subtitles, video descriptions, or voiceovers, is often too broad and not properly aligned in time to accurately describe the video. For example, several 100 million datasets, such as HD-VILA-100M and HowTo100M, are annotated by ASR. However, these datasets, such as the HD-VILA-100M dataset demonstrated in Figure 1, often contain subtitles that fail to capture the main content and actions presented in the video. This limits the value of these datasets for multimodal training. Some publicly available datasets are summarized in Table 1. Some have low resolution, some use ASR for captioning, some contain data from limited domains, some are small, and some provide short captions. Table 1

[0022] The present disclosure includes examples of captioning pipelines that automatically annotate video data with captions generated by ASR to generate large datasets. Automatic captioning pipelines with inputs of multimodal data expand datasets of high-quality video-caption pairs. In one example, the Panda-70M dataset is generated. Some samples of the Panda-70M dataset are illustrated in 102 of FIG. 1. The Panda-70M dataset contains high-resolution videos from public domain sources with rich captions averaging 13.2 words per caption. Manually annotating 70M videos would be costly and time-consuming, but the examples of the present disclosure utilize automatic annotation. An important insight is that a video is typically provided with information from multiple modalities that can assist automatic captioning utilities. This includes the video's title, description, captions, individual static frames, and the video itself. If only one modality is used, the value of this data cannot be fully maximized. In contrast, the present disclosure utilizes different combinations of multimodal data as inputs to multiple cross-modality teachers. When different models are used to caption a subset of videos and the results are evaluated by showing them to humans, it is found that no single model can generate good captions for more than 35% of the videos. However, by collecting all captions from different models together, it is observed that 88.8% of the videos can be annotated with at least one good caption.

[0023] FIG. 2 is a flowchart (201) illustrating the steps of an exemplary method of an automatic captioning pipeline (200) for constructing a Panda-70M dataset by selecting 3.8M high-resolution videos (202) from the HD-VILA-100M dataset and processing them using the following three steps. First, a semantic awareness video segmentation algorithm (204) segments long videos (202) into semantically consistent clips (206) while maintaining a balance between semantic consistency and the duration of the video clips. A total of 70M semantically consistent video clips are obtained. Second, various cross-modality teacher models (208) are used to predict multiple candidate captions for the video clips, including image captioning models and image / video VQA (visual-question-answering) models that have additional text inputs (210), such as video descriptions and subtitles. Finally, a dataset of 100,000 videos is collected in which human annotators select the optimal caption for each video. This dataset is then used to fine-tune a detailed video-to-text (V2T) retrieval model (212) and to select the correct caption for annotating the entire dataset.

[0024] Running multiple teacher models (208) for all data can be relatively expensive and time-consuming. To achieve efficient video annotation on a large scale, in one example, the system uses a student captioning model trained to distill knowledge from teacher models (208). The student model adopts a 2-branch architecture that takes both visual inputs and text inputs to use multimodal information to provide advantages to the captioning process.

[0025] Extensive experiments demonstrate that pre-training with the Panda-70M dataset facilitates several downstream tasks, including video captioning, video and text retrieval, and text-driven video generation. Training the student model using the knowledge distillation method facilitates the development of a robust student model, which performs better than any single teacher model by more than 7.7% preference rate, as shown in Table 4 below, and is further enhanced by additional text inputs such as video descriptions and subtitles.

[0026] The desired video samples in the video-captioning dataset must possess two somewhat conflicting characteristics. On the one hand, the video must be semantically consistent so that the caption can accurately represent its semantic content without ambiguity. On the other hand, the video must not be too short or fragmentary to contain meaningful motion content that is beneficial to downstream tasks such as text-video generation. To achieve both goals, in some examples, the system uses a two-stage semantic awareness segmentation algorithm (204) to segment a long video (202) into semantically consistent clips (206). In the first stage, the long video (202) is segmented based on scene transition detection, as semantics often change when a new scene begins. In the second stage, adjacent clips are spliced ​​together if they were incorrectly separated by the first stage, ensuring that the videos do not become too short. To this end, embeddings of video frames are extracted using ImageBind available in Meta, Menlo Park, California, and adjacent clips of embeddings can be merged if the last frame of the previous clip and the first frame of the next clip are close. Additional procedures handle 1) long videos without scene transitions, 2) videos using fade-in or fade-out transitions that are not typically detected as scene transitions, and 3) the removal of duplicate clips to increase the diversity of the dataset.

[0027] In the first stage (Stage 1), splitting is performed based on scene transition detection. In one example, a long video (202) is split into video clips (206) using a tool (e.g., PySceneDetect) for detecting shot changes in the video that automatically splits the video into separate clips. A utility such as ContentDetector compares the content difference between adjacent frames with a set threshold / score, and if it exceeds this, a scene transition is triggered. cutscene_threshold is set to 25 and min_scene_len is set to 15 frames. A two-stage post-processing algorithm processes 1) long videos (206) that have transitions such as fade-in and fade-out effects that cannot be detected as scene transitions, or 2) unedited footage that does not contain any scene transitions.

[0028] In the first step, the maximum length of the video clip (206) is set to 5 seconds. If the video clip (206) is longer than the maximum length, the first 5 seconds are recursively cut into a new video clip (206) until the remaining video satisfies the condition. The video clip (206) must be semantically consistent. Accordingly, ImageBind features are extracted from the start and end frames, and if the two frames are perceptually different, the video clip (206) is removed. Specifically, given an n-frame video clip C, features f(CA) and f(CB) are extracted for the frames CA and CB, which are the 0.1 × n and 0.9 × n frames of the video clip (206). The video clips (206) It is maintained if it satisfies the condition. As such, video clips (206) that have transition effects or have dramatically different semantics at the beginning and end are excluded.

[0029] In the second stage (Stage 2), splicing is performed based on semantic similarity. To avoid fragmenting the long video (202) by Stage 1, adjacent video clips (206) are spliced ​​into one if they are semantically similar. Formally, given two adjacent video clips C1 and C2 in order, If these conditions are met, they are connected to a video clip (206).

[0030] Post-processing is performed in the following steps to stabilize the quality and variety of video clips (206). First, clips shorter than 2 seconds or containing only a little motion (i.e., Video clips (206) are excluded. For videos longer than 60 seconds, only the first 60 seconds are retained. Next, each video clip (206) is represented by the average ImageBind features extracted in Stage 1, and to increase the diversity of video samples, only video clips (206) that are semantically different from the preceding video clips (206) (i.e., Euclidean distance > 0.3) are retained. Finally, the first and last 10% of the video clips (206) are trimmed because the beginning and end parts generally contain unstable camera movements or transition effects. Through the segmentation algorithm, 3,790,459 long videos (202) are divided into 70,817,315 video clips (206) with an average video clip duration of 8.477 seconds. A distribution plot of video lengths is shown in FIG. 7.

[0031] To quantitatively verify the semantic consistency of a video clip (206), a maximum learning LPIPS that highlights the most important perceptual changes within the video clip (206) is used. Formally, given a video clip (206) of length n seconds, video frames are subsampled every second and keyframes are denoted as {f1, ..., fn}. The maximum learning LPIPS is formulated by the following Equation 1: (1) Here, LPIPS(·, ·) is the perceptual similarity of two images. As shown in Table 2, the segmentation maintains a longer video length than vanilla scene transition detection while achieving better semantic consistency than segmentation based on the alignment of subtitle sentences. Table 2

[0032] The videos in the HD-VILA-100M dataset contain rich multimodal information beneficial for video captioning. Therefore, in addition to the video itself, text information such as video titles, descriptions, and subtitles, as well as static images such as individual video frames, can be used for captioning. Based on these observations, various captioning models with inputs of different modalities can be utilized.

[0033] An automatic captioning pipeline (200) having multiple cross-modality teacher models (208) is used. In one example, a large pool containing 31 captioning teacher models (208) is used. Since inferring all teacher models (208) for 70 million video clips is costly, a shortlist of 8 teacher models (208) that show superior performance is constructed based on user research. The list is illustrated on the y-axis of FIG. 3. The teacher models (208) consist of 5 basic models with different pre-training weights and input information. The 5 basic models include Video LLaMA (Video VQA), VideoChat (Video VQA), VideoChat Text (a natural language model that first texts video content and summarizes it using a Large Language Model (LLM)), BLIP-2 (Image Captioning), and MiniGPT-4 (Image VQA), as illustrated in Table 3. To implement video captioning by teacher models (208) having various modalities, separate captioning processes tailored to each modality are formulated. For example, in the case of VQA models, in addition to visual data, prompts containing additional text information such as video descriptions and subtitles are input, and the teacher models (208) are asked to summarize all multimodal inputs into a single sentence. Table 3

[0034] The captioning process of each teacher model (208) lists the inference details of each basic model. Three suitable basic models include Video-LLaMA, VideoChat, and VideoChat Text. These models are described below in order.

[0035] Video-LLaMA: For all teacher models (208), the system uses the vision branch rather than the audio branch and uses Vicuna-7B as the LLM. Two formula weights are used, including pre-trained weights trained on 2.5M video-text pairs and LLaVA-CC3M, and fine-tuned weights further fine-tuned on command tuning data.

[0036] VideoChat: The system uses Vicuna-7B as the LLM and follows the official codebase for the rest of the settings.

[0037] VideoChat Text: The model is a Natural Language Processing (NLP)-based algorithm that first textifies video content into video tags by the Swin Transformer, into dense captions by GRiT, and into plain captions by the T5 language model. The original codebase uses ChatGPT-4 as a chatbot to implement VQA. LLaMA can be used for large-scale captioning.

[0038] Teacher models (208) using different modalities perform well on different types of videos. For example, video teacher models (208) can perform better on videos with complex dynamics due to additional modules for processing temporal information. On the other hand, for videos containing rare and uncommon objects, image teacher models (208) can accurately caption them because they are trained using datasets of large image-text pairs. Finally, for videos that are difficult to understand visually, VQA teacher models (208) have an advantage because they can utilize additional text inputs. This is supported by a numerical evaluation conducted in a user study in which participants were asked to select the optimal caption from eight candidate teacher models. The selection rates for each teacher model are plotted in FIG. 3. The results show that optimal captions are generated by various teacher models (208). On the other hand, the highest selection rate of the single teacher BLIP-2 with opt6.7b is only 17.85%, which suggests the limited captioning ability of the single teacher model (208) for various videos.

[0039] Given multiple candidate captions for a video, the caption that best aligns with the video content is used. Commonly available retrievals often fail to select the optimal result. One reason is that common models are trained using contrastive learning purposes. They are required to distinguish one sample from other completely unrelated samples. In contrast, in the present disclosure, all candidate captions are highly relevant to the video samples, and for optimal performance, it is necessary to have the teacher model (208) identify subtle differences within each caption.

[0040] The retrieval model (212) is tailored to a “fine” retrieval scenario, and a subset of 100,000 videos is collected, and human annotators select captions containing the most accurate and detailed information about the main content of the videos. This subset is split to obtain 98,000 training samples and 2,000 test samples. An Unmasked Teacher (UMT) is fine-tuned for this dataset, and hard negative mining is implemented for contrast loss, where 7 captions not selected by the annotators constitute the hard negative samples and are assigned larger training weights.

[0041] The retrieval performance of UMTs is quantitatively evaluated on the test set with and without fine-tuning, and the experiments indicate that fine-tuned UMTs can achieve an accuracy of 35.90% R@1, significantly outperforming pre-trained UMTs with 21.82% R@1. It is noted that another human consensus annotation achieves only 48% R@1, as retrieval can be subjective when there is more than one caption with equally superior performance. As shown in Figure 3, fine-tuned UMTs can select captions that are distributed similarly to human-selected captions.

[0042] Although the automatic captioning pipeline (200) can generate accurate captions, the high computational demands may hinder the ability to scale the dataset to a larger size. For example, 8 + 1 different models may be required to annotate a single video clip. According to the present disclosure, a student captioning model is trained on the Panda-70M dataset to distill knowledge from a number of teacher models (208).

[0043] As illustrated in FIG. 4, the student captioning model (400) includes a vision branch (402) and a text branch (404) that utilize multimodal inputs including vision and text. In the case of the vision branch (402), Video-LLaMA is used to extract LLM-compatible video embeddings. In the case of the text branch (404), text embeddings extracted by a text encoder (406) are directly input into the LLM (408). However, this can lead to two problems: first, text prompts containing video descriptions and subtitles may become too long, dominating the decisions of the LLM (408) and imposing a high computational burden; and second, information from descriptions and subtitles may be noisy and not necessarily aligned with the content of the video. To address these problems, exemplary systems include a text Q-former (410) that extracts fixed-length text embeddings and better connects the video embeddings and text embeddings. The text Q-former (410) has the same architecture as the query transformer of BLIP-2. During student captioning model training, gradient propagation from the text branch (404) to the vision branch (402) is blocked, and training is based on the visual encoder (412) and the Video Q-former & Liner (414) for the video input.

[0044] Some samples of the Panda-70M dataset are shown in 102 of Fig. 1. To quantitatively evaluate the effectiveness of the Panda-70M dataset, pre-training performance is tested for three downstream applications: video captioning, video and text retrieval, and video generation.

[0045] To evaluate the performance of the student captioning model (400) for video captioning, Video-LLaMA was used with the vision branch (402) as the base model. Two pre-trained weights were compared: formal weights jointly trained on 2.5M video-text pairs and 595K image-text pairs, and weights trained from scratch on the Panda-2M dataset. The Panda-2M dataset is a randomly sampled subset of the Panda-70M dataset and shares the same amount of training samples as the formal weights. The student captioning model (400) is also trained with both the video branch and the text branch (402 and 404), respectively, on the full Panda-70M dataset for better captioning performance. For all models, the same backbone is used, including Vicuna-7B as the LLM (408), ViT and the video Q-former as the visual encoder (412), and the linear projection layer of MiniGPT-4. For the Panda-2M pre-training, only video and caption data are used without other text information for fair comparison. For the training of the student captioning model (400), in addition to the video, metadata and subtitles were also randomly input into the student captioning model (400).

[0046] Zero-shot video captioning was tested against two benchmarks: MSR-VTT and MSVD. For MSR-VTT, the dataset included 10,000 videos with 20 manually annotated captions for each video, and results for 2,990 test splits were reported. For MSVD, the dataset consisted of 1,970 videos with a total of 80,000 descriptions, and figures for 670 test videos were reported. Note that training or validation videos from downstream datasets were not used. To quantitatively evaluate the quality of the output captions, standard protocols were followed, and BLEU-4, ROGUE-L, METEOR, and CIDEr are reported. All metrics were calculated using the pycocoevalcap package. BERTScore was also calculated to evaluate the contextual similarity between the actual data and the predicted captions for each token. The results are shown in Table 4. For a fair comparison, no additional text information was input into the student model (400) during inference on downstream datasets. FIG. 5 shows video samples from the test set of the Panda-70M dataset in 500 and several captions predicted from different teacher models (208) for qualitative comparison. Table 4

[0047] As shown in Table 4, Video-LLaMA with Panda-2M pre-trained weights achieves significantly superior performance compared to the official weights. Numerically, these pre-trained weights yield improvements of 18.5% and 19.6% for MSR-VTT and MSVD, respectively, in terms of B-4. As shown in Figure 5, the captions from the original Video-LLaMA contain irrelevant and general information such as date and location. In contrast, the predictions according to the present disclosure are better aligned with the video content.

[0048] To evaluate the performance of the student model (400), a user study was conducted in which participants were asked to select the optimal caption from 10 candidates for each video. The 10 captions were predicted by 8 teacher models (208) and 2 student models with different inputs (with and without text). The preference rate of each model was reported, and the R@1 accuracy of the fine-tuned UMT (i.e., all teachers) is shown in Table 5. Video samples were collected from a test set that was not shown during the training of the student model and UMT. As shown in Table 5, the student model (400) outperformed any single model and achieved performance comparable to all teacher models (208). Table 5

[0049] The student model (400) supports cross-modality inputs to utilize additional text information. This is supported by the numerical evaluation in Table 5, where the student model (400) with both video and text inputs performs better than the corresponding model with only video inputs by a 3% preference rate. Qualitatively, captioning predictions are illustrated in FIG. 5 with and without text inputs. While a prediction with pure video inputs may include partial content of the video such as "cactus," the corresponding model with both video and text inputs may more comprehensively include keywords such as "succulents" and "different species" from the video title, description, and subtitles.

[0050] UMT was used as the base model to evaluate performance on video and text retrieval. Standard protocols commonly use 3 million images and 2.5 million videos from CC3M as pre-training datasets. Therefore, for a fair comparison, a Panda-5M subset sharing the same number of training samples as the standard pre-training dataset was randomly sampled. For both datasets, the same backbone consisting of ViT-L / 16 and BERTlarge was used. Formula weights for the standard datasets were used, while the model for the Panda-5M dataset was trained from scratch.

[0051] Zero-shot retrieval was tested against three benchmarks: MSR-VTT, DiDeMo, and MSVD. For MSR-VTT, a standard protocol was followed to evaluate 1K test splits. For DiDeMo, 10K Flickr videos with a total of 40K dense captions were included. As in previous standards, paragraph-video retrieval was evaluated by concatenating all sentence descriptions of a single video into a single query. Results for 1K test sets were reported. For MSVD, results for 670 test videos were reported. Standard metrics were utilized, and R@1, R@5, and R@10 accuracy for both text-video and video-text retrieval are reported in Table 6. Table 6

[0052] Table 6 illustrates that pre-training on the Panda-5M dataset outperforms the formula weights on most datasets. In particular, pre-training yields a 7.0% improvement in terms of R@1 on MSR-VTT. Although pre-training shows lower performance for video-text retrieval on MSVD, Table 6 highlights that MSVD is an initially established dataset and some videos are low resolution (320 × 240px), causing a domain gap with the dataset.

[0053] To evaluate the effectiveness of text-to-video generation, AnimateDiff was used as the base model, and official release weights trained on 2.5 million text-to-video pairs were compared with weights trained on the Panda-2M dataset, a 2.5 million subset of Panda-70M. The official codebase was followed, and Stable Diffusion v1.5 (SD) was used as the base text-to-image (T2I) generator. During training, the T2I modules were fixed, and motion modeling modules were trained. For each training video, 16 frames were sampled with a stride of 4, then resized to a resolution of 256 × 256px and center-cropped.

[0054] To evaluate the models, evaluation protocols for zero-shot evaluation of UCF101 and MSR-VTT were followed. Specifically, 16-frame videos were generated at a resolution of 256 × 256px. For UCF101, a text prompt was generated for each class, and 10,000 videos were generated that shared the same class distribution as the original dataset. Video Distance (FVD) was calculated for the I3D embeddings. For MSR-VTT, video samples were generated for each of the 59,800 test prompts, and CLIP similarity (CLIPSim) was calculated. The figures are reported in Table 7. The generated video samples are shown in 600 of Fig. 6. To visualize the results, the official codebase was followed, and SD T2I was replaced with TUSUN, a personalized Dreambooth weight. Note that the test prompts and video samples from AnimateDiff with the official weights (top row of Fig. 6) were taken from the AnimateDiff project page. Table 7

[0055] Pre-training on the Panda-2M dataset consistently demonstrated superior performance across both metrics compared to the formula weights. As highlighted, pre-training yielded a lower FVD of 77.4 in UCF101, outperforming modern models pre-trained on datasets within 10 million in terms of FVD. Qualitatively, pre-training on the dataset leads to the generation of watermark-free videos with more meaningful motion and realistic appearance.

[0056] FIG. 8 is a schematic representation of a machine (800) on which instructions (810) (e.g., software, program, application, applet, app, or other executable code) may be executed to cause the machine (800) to perform any one or more of the methodologies discussed herein. For example, instructions (810) may cause the machine (800) to perform any one or more of the methods described herein, including an automatic captioning pipeline (200). Instructions (810) transform a general, unprogrammed machine (800) into a specific machine (800) programmed to perform the described and illustrated functions in the described manner. The machine (800) may operate as a standalone device or may be coupled to other machines (e.g., networked). In a networked deployment, the machine (800) may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine (800) may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing commands (810) that specify actions to be taken by the machine (800) sequentially or otherwise. Additionally, while only a single machine (800) is exemplified, the term “machine” should also be taken to include a collection of machines that execute commands (810) individually or jointly to perform any one or more of the methodologies discussed herein.In some examples, the machine (800) may also include both client and server systems, and specific operations of a specific method or algorithm are performed on the server side and specific operations of a specific method or algorithm are performed on the client side.

[0057] The machine (800) may include processors (804), memory (806), and input / output I / O components (802) that can be configured to communicate with each other via a bus (840). In the example, the processors (804) (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio-frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor (808) and a processor (812) that execute instructions (810). The term "processor" is intended to include multi-core processors that may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although FIG. 8 shows multiple processors (804), the machine (800) may include a single processor having a single core, a single processor having multiple cores (e.g., a multi-core processor), multiple processors having a single core, multiple processors having multiple cores, or any combination thereof.

[0058] Memory (806) includes main memory (814), static memory (816), and storage unit (818) accessible to processors (804) via a bus (840). The main memory (806), static memory (816), and storage unit (818) store instructions (810) for any one or more of the methodologies or functions described herein. The instructions (810) may also reside, in whole or in part, in the main memory (814), in the static memory (816), in a machine-readable medium (820) within the storage unit (818), inside at least one of the processors (804) (e.g., in the processor's cache memory), or in any suitable combination thereof during execution by the machine (800).

[0059] The I / O components (802) may include various components such as receiving input, providing output, generating output, transmitting information, exchanging information, and capturing measurements. The specific I / O components (802) included in a specific machine will vary depending on the type of machine. For example, portable machines such as mobile phones may include touch input devices or other such input mechanisms, whereas headless server machines are unlikely to include such touch input devices. It will be understood that the I / O components (802) may include many other components not shown in FIG. 8. In various examples, the I / O components (802) may include user output components (826) and user input components (828). User output components (826) may include visual components (e.g., displays such as a plasma display panel (PDP), light-emitting diode (LED) display, liquid crystal display (LCD), projector, or cathode ray tube (CRT)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input components (828) may include alphanumeric input components (e.g., keyboard, touch screen configured to receive alphanumeric input, optical keyboard, or other alphanumeric input components), point-based input components (e.g., mouse, touchpad, trackball, joystick, motion sensor, or other pointing mechanism), tactile input components (e.g., physical button, touch screen providing location and force of touches or touch gestures, or other tactile input components), audio input components (e.g., microphone), etc.

[0060] In additional examples, I / O components (802) may include biometric components (830), motion components (832), environment components (834), or location components (836) among various other components. For example, biometric components (830) include components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identify a person (e.g., voice identification, retinal identification, face identification, fingerprint identification, or electroencephalogram-based identification). Motion components (832) include acceleration sensor components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes).

[0061] Environmental components (834) include, for example, one or more cameras (with still image / photo and video capabilities), light sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting the concentration of hazardous gases for safety or for measuring pollutants in the atmosphere), or other components capable of providing indications, measurements, or signals corresponding to the surrounding physical environment.

[0062] The position components (836) include position sensor components (e.g., GPS receiver components), altitude sensor components (e.g., altimeters or barometers that detect atmospheric pressure from which altitude can be derived), direction sensor components (e.g., magnetometers), etc.

[0063] Communication can be implemented using various technologies. The I / O components (802) also further include communication components (838) operable to connect the machine (800) to a network (822) or devices (824) through respective couplings or connections. For example, the communication components (838) may include a network interface component or other suitable device for interfacing with the network (822). In additional examples, the communication components (838) may include wired communication components, wireless communication components, cellular communication components, near-field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components for providing communication through other modalities. The devices (824) may be any of other machines or various peripheral devices (e.g., peripheral devices coupled via USB).

[0064] Additionally, the communication components (838) may include components capable of detecting identifiers or detecting identifiers. For example, the communication components (838) may include Radio Frequency Identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., optical sensors for detecting multidimensional barcodes such as one-dimensional barcodes like Universal Product Code (UPC) barcodes, Quick Response (QR) codes, Aztec codes, Data Matrix, Data Glyph, Maxi Code, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or acoustic detection components (e.g., microphones for identifying tagged audio signals). Additionally, various information such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, or location by detecting NFC beacon signals that can indicate a specific location may be derived through the communication components (838).

[0065] Various memories (e.g., main memory (814), static memory (816), and memory of the processors (804)) and storage unit (818) may store one or more instruction sets and data structures (e.g., software) that implement any one or more of the methodologies or functions described herein. These instructions (e.g., instructions (810)) cause various operations to implement the disclosed examples, including the automatic captioning pipeline (200), when executed by the processors (804).

[0066] Commands (810) may be transmitted or received through a network (822) using a transmission medium, through a network interface device (e.g., a network interface component included in the communication components (838)) and using any one of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, commands (810) may be transmitted to or received by devices (824) using a transmission medium through coupling (e.g., peer-to-peer coupling).

[0067] Terms and expressions used in this specification shall be understood to have the ordinary meanings given to such terms and expressions with respect to each area of ​​inquiry and study, except where specific meanings are otherwise provided in this specification. Relational terms such as first, second, etc., may be used solely to distinguish one entity or action from another without necessarily requiring or implying such actual relationship or order between such entities or actions. "comprises," "comprising," "includes," "including," or any other variation thereof are intended to cover non-exclusive inclusion, and accordingly, a process, method, article, or device comprising a list of elements or steps may not only include such elements or steps but may also include other elements or steps not explicitly listed or inherent in such process, method, article, or device. An element preceded by "a" or "an" does not exclude, without additional restriction, the presence of additional identical elements in a process, method, article, or device comprising that element.

[0068] Unless otherwise specified, all measurements, values, grades, positions, sizes, dimensions, and other specifications set forth herein, including the following claims, are not exact but approximations. Such quantities are intended to have a reasonable range consistent with the functions to which they relate and customary practices in the art to which they belong. For example, unless otherwise explicitly stated, parameter values ​​or similar values ​​may vary by ± 10% from the specified amount.

[0069] Furthermore, in the foregoing detailed description, various features are grouped together in various examples for the purpose of simplifying the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are explicitly stated in each claim. Rather, as reflected in the following claims, the subject matter to be protected is less than all the features of any single disclosed example. Accordingly, the following claims are incorporated into the detailed description as each claim constitutes an independently claimed subject matter.

[0070] Although the foregoing has described the best mode and other examples, it is understood that various modifications may be made, the subject matter disclosed herein may be embodied in various forms and examples, and may be applied to numerous applications comprising only a part described herein. It is intended to claim any modifications and variations that fall within the true scope of the concepts by the following claims.

Claims

Claim 1 A pipeline configured to automatically annotate video data with subtitles, comprising: curating multiple high-resolution videos from a publicly available dataset; splitting the videos into semantically consistent video clips; and applying multiple cross-modality teacher models to obtain captions for each of the video clips. Claim 2 In claim 1, the pipeline is configured to use automatic speech recognition (ASR) to automatically annotate the video data with subtitles. Claim 3 In claim 1, the above teacher models have different pre-training weights and are configured to receive different input information, a pipeline. Claim 4 In claim 1, the pipeline is configured to fine-tune a subset of video clips using a retrieval model in which the exact caption of each of the video clips can be manually selected. Claim 5 In paragraph 4, the above subset is configured to be divided to obtain training samples and test samples, a pipeline. Claim 6 In paragraph 4, the pipeline is configured to use the retrieval model for the entire dataset to select an accurate caption as an annotation after fine-tuning the subset. Claim 7 In claim 1, the pipeline is configured to use inputs of multimodal data for the teacher models and scale up the dataset into high-quality video-caption pairs. Claim 8 In paragraph 7, the multimodal data comprises a textual video description, subtitles, and individual video frames, in a pipeline. Claim 9 In claim 8, the pipeline is configured to enable a student captioning model to learn about the dataset in order to distill knowledge from the teacher models. Claim 10 In claim 9, the student captioning model comprises a vision branch and a text branch configured to utilize multimodal inputs, and during the training of the student captioning model, gradient propagation from the text branch to the vision branch is blocked, in a pipeline. Claim 11 A pipeline according to claim 10, wherein the vision branch is configured to be used to extract a large language model (LLM) compatible video embedding, and the text branch comprises a text Q-former and a text encoder configured to extract a fixed-length text embedding to link the video embedding and the text embedding. Claim 12 A method for using a pipeline to automatically annotate video data with subtitles, comprising the steps of: selecting a plurality of high-resolution videos from a publicly available dataset; splitting the videos into semantically consistent video clips; and applying a plurality of cross-modality teacher models to obtain captions for each of the video clips. Claim 13 In paragraph 12, the pipeline uses automatic speech recognition (ASR) to automatically annotate the video data with subtitles. Claim 14 In paragraph 12, the above teacher models have different pre-training weights and receive different input information, a method. Claim 15 In paragraph 12, the pipeline fine-tunes a subset of video clips using a retrieval model in which the exact caption for each of the video clips can be manually selected. Claim 16 In paragraph 15, the above subset is divided to obtain training samples and test samples, a method. Claim 17 In paragraph 15, a method in which, after fine-tuning the subset, the pipeline uses the retrieval model for the entire dataset to select an accurate caption as an annotation. Claim 18 In claim 12, the pipeline uses inputs of multimodal data for the teacher models and expands the dataset into high-quality video-caption pairs. Claim 19 In paragraph 18, the method wherein the multimodal data comprises text video descriptions, subtitles, and individual video frames. Claim 20 A non-transient computer-readable medium storing program code, wherein, when the program code is executed, the non-transient computer-readable medium operates to cause a pipeline to perform the steps of: selecting a plurality of high-resolution videos from a publicly available dataset; splitting the videos into semantically consistent video clips; and applying a plurality of cross-modality teacher models to obtain captions for each of the video clips.