SISTEMAS E MÉTODOS PARA IDENTIFICAÇÃO AUTOMÁTICA DE CLIPES DE VÍDEO DIGITAIS QUE RESPONDEM A CONSULTAS DE BUSCAS ABSTRATAS

BR112025019964A2Pending Publication Date: 2026-08-04NETFLIX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025019964
Authority / Receiving Office
BR · BR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-20
Filing Date
2024-03-20
Publication Date
2026-08-04

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosed computer-implemented methods and systems include implementations that automatically generate and train a video clip classifier model to identify video clips that respond to a specific search query for a desired depiction that can include abstract, context-dependent, and / or subjective terms. For example, the methods and systems described herein generate and update a digital content understanding graphical user interface to facilitate the process of generating a corpus of training digital video clips, training a video clip classifier model with the training digital video clips, and applying the video clip classifier model to new digital video clips. Various other methods, systems, and computer-readable media are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

1 / 45 Systems and methods for automatic identification of digital video clips that respond to abstract cross-reference search queries.

[0001] This application claims priority over U.S. Application No. 18 / 186,467, filed March 20, 2023, the disclosure of which is incorporated in its entirety by this reference. BACKGROUND

[0002] Digital media is increasingly consumed in many different forms. For example, users enjoy watching TV episodes and movies, as well as trailers, previews, and clips of those TV episodes and movies. For example, a movie trailer typically includes scenes from the movie that are put together in a way that sparks the interest of a potential viewer. Similarly, a preview of a season of TV episodes may include scenes from the season's episodes that foreshadow plot points and cliffhangers.

[0003] However, generating trailers and previews can give rise to several technological problems. For example, a movie trailer might be generated as a result of a process involving a user manually searching through a film's scenes for video clips that include a certain type of scene, a certain object, a certain character, a certain emotion, and so on. In some cases, the user might utilize a search tool to help sort through the thousands of scenes that movies and TV shows typically include. Despite this, existing search tools generally search video clips by attempting to associate the clip images with a text-based search query. This approach, however, is often unable to handle nuanced search queries for anything beyond specific objects or people included in a given scene.

[0004] Therefore, these existing search tools are often inaccurate. For example, search tools Petition 870250084254, dated 09 / 18 / 2025, page 19 / 91 2 / 45 of existing search engines are often limited in terms of search capabilities. To illustrate, a search tool might be able to match frames from a digital video (e.g., a movie) to a received search query for a specific term, such as a search query for a particular object or character. As search queries become more detailed, subjective, and context-dependent, standard search tools may fail to return accurate results. Additional resources must be spent manually analyzing these inaccurate results to find digital video clips that correctly answer the search query.

[0005] Furthermore, standard search tools for finding specific clips in a digital video are often inflexible. For example, as mentioned above, standard search tools are usually restricted to simple image-based searches and / or basic keyword searches. Therefore, these tools lack the flexibility needed to perform searches based on more abstract concepts, such as searches for specifically portrayed emotions, shot types, and the overall feeling of the scene.

[0006] Furthermore, existing search methodologies are generally inefficient. As discussed above, some search methods are entirely manual and require users to manually extract clips from digital videos. Other methodologies may include search tools that can identify digital video clips that answer certain types of search queries, but these tools utilize an excessive number of processor cycles and memory resources to perform searches limited to concrete search terms. In some cases, search methodologies may include machine learning components, but these components are usually created and trained manually, a process that requires large amounts of time and computing resources. Petition 870250084254, dated 09 / 18 / 2025, page 20 / 91 3 / 45 SUMMARY

[0007] As will be described in more detail below, this disclosure describes embodiments that automatically identify digital video clips that respond to abstract search queries for use in digital video assets, such as trailers and previews. In one example, a computer-implemented method for automatically predicting classification categories for digital video clips that indicate whether the digital video clips respond to an abstract search query may include generating, by applying a video clip classifier model to a training corpus of digital video clips that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the training corpus of digital video clips,To retrain the video clip classifier model based on user confirmations, detected through the digital content comprehension graphical user interface, regarding the classification score accuracies generated by the video clip classifier model that correspond to the training digital video clips; to analyze a digital video into digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface; and to generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

[0008] In addition, in some examples, the method may additionally include generating the corpus of digital video clips from Petition 870250084254, dated 09 / 18 / 2025, page 21 / 91 4 / 45 training, iteratively receiving search input related to at least one of an object, action, shot type, editing technique, character, or story theme, and identifying, within a repository of digital training video clips, a plurality of digital training video clips that respond to the received search input.

[0009] In some instances, the classification category prediction displays in the digital content comprehension graphical user interface include a playback window loaded within a training digital video clip corresponding to classification category prediction displays. The classification category prediction displays may additionally include a title of a digital video from which the displayed training digital video clip came, and an option to positively or negatively confirm the same training digital video clip.Generating classification category prediction displays within the digital content comprehension graphical user interface may additionally include classifying the classification category prediction displays into high confidence levels and low confidence levels, and updating the classification category prediction displays within the digital content comprehension graphical user interface according to the high confidence levels and the low confidence levels.

[0010] In addition, in some examples, the method may also include detecting user confirmations regarding the accuracy of the classification scores generated by the video clip classifier model, detecting at least one of (1) a first user input corresponding to a positive confirmation of a first video clip included in the digital training video clips, or (2) a second user input corresponding to a negative confirmation of Petition 870250084254, dated 09 / 18 / 2025, page 22 / 91 5 / 45 first video clip included in the training digital video clips. The method may also include detecting the digital video selection through the digital content comprehension graphical user interface, detecting a selection of at least one short-form digital video, one long-form digital video, or a season of short-form digital videos. Furthermore, analyzing digital video in digital video clips may include analyzing digital video in portions of continuous digital video footage between two cuts.

[0011] In some examples, generating suggested digital video clip displays that replace classification category prediction displays within the digital content understanding graphical user interface may include generating input vectors based on the digital video clips, applying the retrained video clip classifier model to the generated input vectors, receiving, from the retrained video clip classifier model, classification scores for the digital video clips that match the received search input, generating, for the digital video clips, suggested video clip displays, and replacing classification category prediction displays with the suggested digital video clip displays for the digital video clips within the digital content understanding graphical user interface according to the classification scores for the digital video clips.

[0012] Some examples described here include a system with at least one physical processor and physical memory, including computer executable instructions that, when executed by at least one physical processor, cause at least one physical process to perform various actions. In at least one example, the computer executable instructions, when executed by at least one physical processor, cause at least one Petition 870250084254, dated 09 / 18 / 2025, page 23 / 91 6 / 45 physical processor performs actions including generating, through the application of a video clip classifier model to a corpus of training digital video clips that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the training corpus of digital video clips, retraining the video clip classifier model based on user confirmations, detected through the digital content comprehension graphical user interface, regarding the accuracies of classification scores generated by the video clip classifier model that correspond to the training digital video clips, analyzing a digital video in digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface,and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content understanding graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

[0013] In some examples, the method described above is encoded as computer-readable instructions in a computer-readable medium. In one example, the computer-readable instructions, when executed by at least one processor of a computing device, cause the computing device to generate, through the application of a video clip classifier model to a corpus of digital training video clips that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the corpus. Petition 870250084254, dated 09 / 18 / 2025, page 24 / 91 7 / 45 of training digital video clips, retrain the video clip classifier model based on user confirmations detected through the digital content understanding graphical user interface, regarding the classification score accuracies generated by the video clip classifier model that match the training digital video clips, analyze a digital video in digital video clips in response to the detection of a digital video selection through the digital content understanding graphical user interface, and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content understanding graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

[0014] In one or more examples, features of any of the embodiments described herein are used in combination with each other in accordance with the general principles described herein. These and other embodiments, features and advantages will be better understood after reading the following detailed description together with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The attached drawings illustrate a series of exemplary embodiments and form part of the descriptive report. Together with the description that follows, these drawings demonstrate and explain various principles of the present disclosure.

[0016] FIG. 1 is a block diagram of an exemplary environment for implementing a digital content comprehension system according to one or more implementations.

[0017] FIG. 2 is a flow diagram of an exemplary computer-implemented method for automatically generating classification category predictions indicating digital video clips that respond to specific search queries. Petition 870250084254, dated 09 / 18 / 2025, page 25 / 91 8 / 45 and abstract according to one or more implementations.

[0018] FIGS. 3A-3N illustrate a graphical user interface for understanding digital content during the process of generating and using a video clip classifier model to identify specific digital video clips according to one or more implementations.

[0019] FIG. 4 is a detailed diagram of the digital content comprehension system according to one or more implementations.

[0020] Throughout the drawings, identical reference characters and descriptions indicate similar, but not necessarily identical, elements. Although the exemplary embodiments described in this document are susceptible to various modifications and alternative forms, specific embodiments have been shown as examples in the drawings and will be described in detail in this document. However, the exemplary embodiments described in this document are not intended to be limited to the particular forms disclosed. Instead, this disclosure covers all modifications, equivalents, and alternatives that fall within the scope of the appended claims. DETAILED DESCRIPTION OF EXEMPLARY MODALITIES

[0021] As mentioned above, rapidly generating digital media assets, such as trailers and previews, is generally desirable. For example, content creators often need to be able to quickly identify video clips from a film that answer specific search queries in order to efficiently build a trailer for the film that conveys the desired story, emotion, tone, etc. Existing methods for querying digital video clips generally involve search tools that lack the ability to handle detailed or abstract search queries. In some cases, a search tool may incorporate components of Petition 870250084254, dated 09 / 18 / 2025, page 26 / 91 9 / 45 machine learning. However, these components are usually built individually and trained in slow, inefficient, and computationally expensive processes.

[0022] To remedy these problems, the present disclosure describes implementations that can automatically generate and train a video clip classifier model to identify video clips that respond to a specific search query that may include abstract, context-dependent, and / or subjective terms. For example, the implementations described here can generate a digital content comprehension graphical user interface that guides the training data generation process, build a video clip classifier model, train the video clip classifier model, and apply the video clip classifier model to new video clips.The implementations described here can identify digital training video clips that respond both positively and negatively to a received search query and can generate classification category predictions for each of the identified digital training video clips. The implementations described here can additionally receive confirmations regarding the accuracy of these predictions through the digital content comprehension graphical user interface. The implementations described here can further train the video clip classifier model based on the confirmations received through the digital content comprehension graphical user interface.Finally, the implementations described here can additionally apply the trained video clip classifier model to newly analyzed video clips from a movie or TV episode to determine which video clips correspond to the term, notion, or moment for which the video clip classifier was trained.

[0023] In more detail, the systems and methods Petition 870250084254, dated 09 / 18 / 2025, page 27 / 91 The disclosed 10 / 45 systems and methods offer an efficient methodology for generating a corpus of digital training video clips to train a video clip classifier model. For example, the disclosed systems and methods allow a user to search for digital training video clips that answer search queries associated with a representation of a specific moment. To illustrate, if the specific moment is “thoughtful clips,” the disclosed systems and methods can allow the user to search for digital training video clips that answer search queries that positively inform about that specific moment, such as “silence,” “sitting,” “slow walking,” “soft music,” and “close-up face.” The disclosed systems and methods can additionally allow the user to search for digital training video clips that answer search queries that negatively inform about that specific moment, such as “loud,” “action,” “explosions,” and “group shots.”By utilizing all these digital training video clips, the disclosed systems and methods enable the creation of a video clip classifier model that is precisely trained to a specific definition of a particular moment. By additionally allowing for the rapid labeling of low-confidence classification predictions generated by the video clip classifier model during training, the disclosed systems and methods efficiently enable further improvement of the video clip classifier model.

[0024] Once trained, the disclosed systems and methods can apply the video clip classifier model to additional digital video clips. For example, the disclosed systems and methods can apply the video clip classifier model to the user-specified digital video (e.g., a TV episode, a season of TV episodes) to generate classification predictions for video clips. Petition 870250084254, dated 09 / 18 / 2025, page 28 / 91 11 / 45 from the digital video indicated by the user. Due to the way the disclosed systems and methods generate the training corpus for the video clip classifier model, the predictions generated by the video clip classifier model are precisely adapted to how the user defined the specific moment in which they are interested.

[0025] Features of any of the implementations described herein may be used in combination with one another, in accordance with the general principles described herein. These and other implementations, features and advantages will be better understood after reading the following detailed description together with the attached drawings and claims.

[0026] The following will provide, with reference to FIGS. 1-4, detailed descriptions of a digital content comprehension system that can quickly and efficiently generate trained video clip classifier models that identify video clips that answer specific search queries. For example, an exemplary network environment is illustrated in FIG. 1 to show the digital content comprehension system operating in connection with various devices while generating and training video clip classifier models. FIG. 2 illustrates the steps performed by the digital content comprehension system during this process. FIGS. 3A-3N illustrate a digital content comprehension graphical user interface generated by the digital content comprehension system as it guides the process of generating, training, and applying a video clip classifier model. Finally, FIG.4 provides additional details regarding the features and functionalities of the digital content comprehension system.

[0027] As we have just mentioned, FIG. 1 illustrates an exemplary network environment that implements 100 aspects of this disclosure. For example, the network environment 100 may include server(s) 104, a client computing device 106 and a Petition 870250084254, dated 09 / 18 / 2025, page 29 / 91 12 / 45 network 112. As shown further, the server(s) 104 and the client computing device 106 may include memory 114, additional items 116 and a physical processor 118.

[0028] In at least one implementation, a digital content comprehension system 102 may be implemented in the memory 114 of the server(s) 104. In some implementations, the client computing device 106 may also include a web browser 108 installed in memory 114. As shown in FIG. 1, the client computing device 106 and the server(s) 104 may communicate via the network 112 to transmit and receive digital content data.

[0029] In one or more implementations, the client computing device 106 may include any type of computing device. For example, the client computing device 106 may include a desktop computer, a laptop, a tablet, a smartphone, a smart wearable device, an augmented reality device, and / or a virtual reality device. In at least one implementation, the installed web browser 108 may access websites, download content, render web page displays, and so forth.

[0030] As shown in FIG. 1, the network environment 100 may include the digital content comprehension system 102. In one or more implementations, the digital content comprehension system 102 may generate and provide a digital content comprehension graphical user interface for the client computing device 106. In one or more implementations, the digital content comprehension system 102 may generate and train video clip classifier models in response to different types of interactions detected through the digital content comprehension graphical user interface. Finally, the digital content comprehension system 102 may automatically identify video clips from a digital video (e.g., a movie or TV episode) that depict an object, Petition 870250084254, dated 09 / 18 / 2025, page 30 / 91 13 / 45 person, moment, scene type, etc. specified.

[0031] In at least one implementation, the digital content comprehension system 102 may utilize a digital content repository 110 stored in additional items 116 on server(s) 104. For example, the digital content repository 110 may store and maintain digital training video clips. The digital content repository 110 may store and maintain digital videos, such as digital movies and TV episodes. The digital content repository 110 may maintain digital training video clips, digital videos, and other digital content (e.g., digital audio files, digital text such as film scripts, digital photographs) in any of several organizational schemes, such as, but not limited to, alphabetical order, by runtime, by genre, by type, etc.

[0032] As mentioned above, the client computing device 106 and the server(s) 104 can be communicatively coupled via the network 112. The network 112 can represent any type or form of communication network, such as the Internet, and can include one or more physical connections, such as a LAN, and / or wireless connections, such as a WAN.

[0033] Although FIG. 1 illustrates network environment components 100 in one arrangement, other arrangements are possible. For example, in one implementation, the digital content comprehension system 102 may operate as a native application that can be installed on the client computing device 106. In another implementation, the digital content comprehension system 102 may operate on multiple servers. Furthermore, in some implementations, the digital content comprehension system 102 may operate within a larger digital content system that transmits digital content to client transmission devices.

[0034] In one or more implementations, the methods and steps Petition 870250084254, dated 09 / 18 / 2025, page 31 / 91 14 / 45 performed by the digital content comprehension system 102 refer to multiple terms. For example, the term “digital video” can refer to a digital media item. In one or more implementations, a digital video includes audio and visual data, such as image frames synchronized with an audio soundtrack. As used here, the term “digital video clip” can refer to a portion of a digital video. For example, a digital video clip can include image frames and synchronized audio for footage that occurs between cuts or transitions within the video. In one or more implementations, a “short-form digital video” can refer to an episodic digital video, such as an episode of a television program. It follows that a “short-form digital video season” can refer to a collection of episodic digital videos.For example, a season of episodic digital videos can include any number of short-form digital videos (e.g., 10 episodes - 22 episodes). Additionally, as used here, a "long-form digital video" can refer to a non-episodic digital video, such as a film.

[0035] As used herein, a “search query” can refer to a word, phrase, image, or sound that correlates with one or more repository entries. For example, a search query might include a title or identifier of a digital video stored in repository 110. Additionally, as used herein, a “desired representation” can refer to a specific type of search query against which a video clip classifier model can be trained. For example, a desired representation might include an object, character, actor, filming technique, sentiment, action, etc., that a video clip classifier model can be trained to identify within a video clip. In the examples described herein, a desired representation might be correlated with the name or title of a clip classifier. Petition 870250084254, dated 09 / 18 / 2025, page 32 / 91 15 / 45 video.

[0036] As used herein, the term “video clip classifier model” can refer to a computational model that can be trained to generate predictions. For example, as described in connection with the examples herein, a video clip classifier model might be a binary classification machine learning model that can be trained to generate predictions indicating whether a video clip shows a specific desired representation. In at least one implementation, the video clip classifier model might generate such a prediction in the form of a classification score (e.g., between zero and one) that indicates a level of confidence about whether a video clip shows a specific desired representation. Thus, a video clip classifier model might indicate a high level of confidence that a video clip includes a desired representation by generating a classification score close to one (e.g., 0.90).On the other hand, a video clip classifier model may indicate a low level of confidence that a video clip includes a desired representation, generating a classification score close to zero (e.g., 0.1).

[0037] As used herein, a “corpus of digital video clips for training” can refer to a collection of digital video clips that are used to train a video clip classifier model. For example, digital video clips for training may include video clips that positively match a search query or desired representation (i.e., video clips that include the desired representation). Digital video clips for training may also include video clips that negatively match the search query or desired representation (i.e., video clips that do not include the desired representation). When training the classifier model Petition 870250084254, dated 09 / 18 / 2025, page 33 / 91 16 / 45 of video clips with these video clips, the video clip classifier model can learn to determine whether or not a video clip includes a desired representation.

[0038] As used herein, “user confirmations” may refer to user input associated with training digital video clips and / or classification category predictions. For example, the 102 digital content comprehension system may generate a digital content comprehension graphical user interface that includes selectable confirmation options. Using these options, the 102 digital content comprehension system may detect user selections that positively confirm a training digital video clip, indicating that a training digital video clip should be included as a positive training example for a video clip classifier model. The 102 digital content comprehension system may also detect, through these options, a positive confirmation of a classification category prediction, indicating that the classification category prediction correctly includes a desired representation.Furthermore, the 102 digital content comprehension system can detect user selections that negatively confirm a training digital video clip, indicating that the training digital video clips should be included as a negative training example for the video clip classifier model. The 102 digital content comprehension system can also detect a negative confirmation of a classification category prediction, indicating that the classification category prediction incorrectly fails to include the desired representation.

[0039] As used herein, the term “classification category prediction” may refer to a digital training video clip and its corresponding classification score. The Petition 870250084254, dated 09 / 18 / 2025, page 34 / 91 The 17 / 45 digital content comprehension system 102 can generate the digital content comprehension graphical user interface, including a classification category prediction display that includes various relevant information for a training digital video clip. For example, the classification category prediction display might include the training digital video clip loaded in a playback control, a title of the digital video from which the training digital video clip came, the classification score of the training digital video clip, options to positively or negatively confirm the classification score for the classification category prediction, and other information.

[0040] Similarly, as used herein, the term “suggested digital video clip” may refer to a digital clip that is not part of the training corpus but is determined by a video clip classifier model trained to include the desired representation for which the video clip classifier model was trained. For example, the 102 digital content comprehension system may generate suggested digital video clip displays that include information similar to that included in the classification category prediction displays.In at least one implementation, the 102 digital content comprehension system can generate digital video clip display suggestions without options to positively or negatively confirm classification scores, since the 102 digital content comprehension system typically provides digital video clip display suggestions after training the associated video clip classifier model.

[0041] As mentioned above, FIG. 2 is a flow diagram of an exemplary computer-implemented method for automatically generating classification category predictions indicating digital video clips that respond to queries of Petition 870250084254, dated 09 / 18 / 2025, page 35 / 91 18 / 45 specific and potentially abstract searches. The steps shown in FIG. 2 can be performed by any suitable computer executable code and / or computing system, including the system(s) illustrated in FIG. 4. In one example, each of the steps shown in FIG. 2 can represent an algorithm whose structure includes and / or is represented by multiple substeps, examples of which will be provided in more detail below.

[0042] As illustrated in Figure 2, in step 202, the digital content comprehension system 102 can generate, by applying a video clip classifier model to a training corpus of digital video clips that respond to a received search input, classification category predictions for the training corpus of digital video clips. For example, the digital content comprehension system 102 can generate the training corpus of digital video clips that includes digital video clips that respond positively to a search query (e.g., a desired video clip representation) and digital video clips that respond negatively to a search query. The digital content comprehension system 102 can additionally receive user confirmations regarding the accuracy of the training digital video clips that match the search query.The 102 digital content comprehension system can additionally apply a video clip classifier model to the corpus of training digital video clips to generate classification category predictions for each of the training digital video clips, where the classification category prediction for a training digital video clip indicates a probability that the training digital video clip portrays the desired representation (e.g., a person, place, object, feeling, filming technique). Petition 870250084254, dated 09 / 18 / 2025, page 36 / 91 19 / 45

[0043] Furthermore, in step 204, the digital content comprehension system 102 can retrain the video clip classifier model based on user confirmations, detected through a digital content comprehension graphical user interface, regarding the accuracy of the classification scores generated by the video clip classifier model that correspond to the training digital video clips. For example, to further train the video clip classifier model to accurately generate classification category predictions regarding the desired representation, the digital content comprehension system 102 can generate a display including the classification category predictions.The 102 digital content comprehension system can generate the display so that each classification category prediction includes an indication of its associated digital video clip, as well as an option for a user to indicate whether the classification score associated with the classification category prediction is accurate. In response to detecting user confirmations regarding the accuracy of a threshold number of classification category predictions, the 102 digital content comprehension system can retrain the video clip classifier model based on the confirmations.

[0044] Furthermore, in step 206, the digital content comprehension system 102 can analyze a digital video into digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface. For example, in response to retraining the video clip classifier model, the digital content comprehension system 102 can apply the video clip classifier model to video clips that are not part of the digital video clip training corpus. Thus, the digital content comprehension system Petition 870250084254, dated 09 / 18 / 2025, page 37 / 91 20 / 45 102 can detect a user's selection of a digital video (e.g., a movie, a TV episode, a season of TV episodes), and then analyze the selected digital video into digital video clips. In at least one implementation, the 102 digital content comprehension system analyzes a digital video clip to include the footage between two cuts or film transitions.

[0045] Furthermore, in step 208, the digital content comprehension system 102 can generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips. For example, the digital content comprehension system 102 can apply the retrained video clip classifier model to the analyzed digital video clips of the selected digital video to generate suggested digital video clip displays with classification scores that indicate whether or not the associated digital video clips respond to the received search query by portraying a desired representation.Thus, the combination of automated and user-guided steps in the process of generating digital video clip suggestions can result in highly accurate digital video clip suggestions, even when the related search query is abstract, subjective, and / or context-dependent.

[0046] As discussed above, the digital content comprehension system 102 generates and provides a digital content comprehension graphical user interface to the client computing device 106 to guide the process of generating and using a video clip classifier model to identify specific digital video clips. FIGS. 3A-3N illustrate the digital content comprehension graphical user interface. Petition 870250084254, dated 09 / 18 / 2025, page 38 / 91 21 / 45 digital content generated and updated by the digital content comprehension system 102 during this process. For example, as shown in FIG. 3A, a graphical user interface for digital content comprehension 304 is displayed via the web browser 108 on the client computing device 106.

[0047] In one or more implementations, the digital content comprehension system 102 can generate the digital content comprehension graphical user interface 304, including various options associated with video clip classifier models. For example, the digital content comprehension system 102 can generate the digital content comprehension graphical user interface 304, including a list 308 of existing video clip classifier models. In response to a detected selection of any of the existing video clip classifier models in the list 308, the digital content comprehension system 102 can make the selected video clip classifier model available for further training and / or application to analyzed video clips from a digital video (e.g., a movie or TV episode).To illustrate, in response to a detected selection of the closeup video clip classifier model in list 308, the digital content comprehension system 102 can make this model available for application to digital video clips analyzed from a digital video. The close-up video clip classifier model can then generate classification category predictions for each of the digital video clips, where the classification category predictions indicate whether each of the digital video clips depicts a close-up shot of people, objects, scenes, etc.

[0048] In addition to providing access to existing video clip classifier models, the digital content comprehension system 102 can additionally generate the digital content comprehension graphical user interface 304, Petition 870250084254, dated 09 / 18 / 2025, page 39 / 91 22 / 45 including options for generating a new video clip classifier model. For example, in response to the user entering a new video clip classifier model title (e.g., “Happy Takes”) in text input box 306 and a detected selection of the “Create New Model” button 309, the digital content comprehension system 102 can initiate the process of generating a new video clip classifier model. In one or more implementations, the text entered in text input box 306 can indicate a desired search query or representation that will be the focus of the new video clip classifier model.

[0049] Furthermore, as shown in FIG. 3A, the digital content comprehension system 102 can additionally generate the digital content comprehension graphical user interface 304, including a model search text box 307. In one or more implementations, the digital content comprehension system 102 can search the digital content repository 110 for one or more video clip classifier models that respond to a search query received through the model search text box 307. In at least one implementation, the digital content comprehension system 102 can search for models with names similar to or related to the search query received through the model search text box 307. In additional implementations, the digital content comprehension system 102 can search for models that have been previously applied to TV episodes and / or films indicated through the model search text box 307.Furthermore, the digital content comprehension system 102 can search for models that are associated with a desired representation of a specific moment indicated by the model search text box 307. In this way, the digital content comprehension system 102 can provide quick access to previously trained and used video clip classifier models that can... Petition 870250084254, dated 09 / 18 / 2025, pp. 40 / 91 23 / 45 have been configured by other users.

[0050] As shown in FIG. 3B, the digital content comprehension system 102 can update the digital content comprehension graphical user interface 304 after the start of this process to include the title of the video clip classifier model being generated (e.g., “Happy Takes”) and a display including tabs 312a (e.g., “Choose Candidates”), 310b (e.g., “Confirm Choices”), 310c (e.g., “Build and Improve”), and 310d (e.g., “Use / Publish”). In one or more implementations, the digital content comprehension system 102 can guide the user of the client computing device 106 through the process of generating the new video clip classifier model based on the functionality present in each of the tabs 310a-310d.

[0051] In response to a detected selection from tab 310a (e.g., “Choose Candidates”), the digital content comprehension system 102 can update the digital content comprehension graphical user interface 304 to include a search query input field 312 and a “Search” button 314. As shown in FIG. 3C, the digital content comprehension system 102 can detect a search entry within the search query input field 312 (e.g., “laughing”) and a selection of the “Search” button 314. In response to this, the digital content comprehension system 102 can identify one or more digital training video clips, as shown in FIG. 3D.For example, the display of digital training video clip 316a might include a digital training video clip 318a loaded in the video playback window, a title ID 320a, a digital video clip title 322a, a start timestamp 324a, a duration 326a, a rating score 328a, and a label option 330a.

[0052] In more detail, the digital content comprehension system 102 can identify the digital video clip. Petition 870250084254, dated 09 / 18 / 2025, pp. 41 / 91 24 / 45 training 318a by performing a search in the digital content repository 110. For example, the digital content repository 110 may store training digital video clips that include a series of digital video frames and metadata, including the title ID and title of the digital video from which the clip came, the timestamp where the clip begins within the digital video, and the duration of the clip within the digital video. In this way, the digital content comprehension system 102 can identify the training digital video clip 318a by performing a visual search of the frames of the training digital video clip stored in the digital content repository 110 for those that answer the search query in the search query input field 312.For example, the digital content comprehension system 102 can use computer vision techniques to analyze frames from training digital video clips in the digital content repository 110 for those that depict the object, character, or topic of the search query. In some implementations, the digital content comprehension system 102 can additionally identify the training digital video clip 318a by searching the metadata associated with training digital video clips in the digital content repository 110 for terms and other data that match the search query.

[0053] In one or more implementations, the digital content comprehension system 102 can generate the training digital video clip display 316a shown in FIG. 3D so that the user of the client computing device 106 can view the training digital video clip 318a via a video playback control within the training digital video clip display 316a and determine whether the training digital video clip 318a should be used to train the new video clip classifier model. For example, the digital content comprehension system 102 can Petition 870250084254, dated 09 / 18 / 2025, page 42 / 91 25 / 45 Add the training digital video clip 318a to a corpus of training digital video clips to train the new video clip classifier model in response to the detection of a positive radio button selection 331b in the label option 330a. Conversely, the digital content comprehension system 102 can keep the training digital video clip 318a out of the corpus of training digital video clips in response to the absence of a selection made in the label option 330a.

[0054] In at least one implementation, the digital content comprehension system 102 can identify training digital video clips that respond positively to the desired representation indicated by the title of the new video clip classifier model (e.g., a received search query). For example, as shown in FIG. 3D, the digital content comprehension system 102 identified training digital video clip 318a depicting at least one smiling person. In a further implementation, the digital content comprehension system 102 can additionally identify training digital video clips that respond negatively to the desired representation indicated by the title of the new video clip classifier model. For example, as shown in FIG.3E, the digital content comprehension system 102 can generate the display of training digital video clip 316b and add training digital video clip 318b to the same corpus of training digital video clips, even though training digital video clip 318b responds negatively to the video clip classifier title “Happy Takes”—in other words, training digital video clip 318b responds to the search query “sad,” which is an antonym of “laughing” and / or “happy.” The digital content comprehension system 102 can determine, for example, that training digital video clip 318b is a… Petition 870250084254, dated 09 / 18 / 2025, page 43 / 91 26 / 45 negative training digital video clip that matches the video clip classifier model “Happy Takes” in response to a detected selection of the negative radio button 331a in the label option 330b.

[0055] In one or more implementations, the digital content comprehension system 102 can build the training corpus of digital video clips in multiple iterations. For example, a user of the client computing device 106 can add several terms to the search query input field 312 in multiple iterations. Each of the terms entered by the user can be associated — positively or negatively — with the desired representation indicated by the title of the new video clip classifier model “Happy Takes”. To illustrate, the digital content comprehension system 102 can search for training digital video clips that respond to positively associated terms, such as “smiling”, “laughing”, “sunny”, “singing”, and “hugging”. The digital content comprehension system 102 can additionally search for training digital video clips that respond to negatively associated terms, such as “sad”, “angry”, “gloomy”, and “fighting”.In some implementations, these terms are entered by the user of the client computing device 106. In additional implementations, the digital content comprehension system 102 may identify, suggest and / or enter the same terms.

[0056] Once the digital content comprehension system 102 has built the corpus of training digital video clips—both positive and negative training digital video clips—the digital content comprehension system 102 can provide the user of the client computing device 106 with an opportunity to adjust the corpus of training digital video clips. For example, as shown in Petition 870250084254, dated 09 / 18 / 2025, page 44 / 91 27 / 45 FIG. 3F, in response to a detected selection from tab 310b, the digital content comprehension system 102 can update the digital content comprehension graphical user interface 304 to include the corpus 333 of displays of training digital video clips 316a, 316c (for example, as training digital video clip 318a and training digital video clip 318c). From this display, the digital content comprehension system 102 allows the user of the client computing device 106 to verify the positive and negative labels indicated by the label options 330a, 330b associated with each of the training digital video clips 318a, 318c, respectively.

[0057] With the corpus of digital video clips for training generated and verified, the digital content comprehension system 102 can build the new video clip classifier model (e.g., the “Happy Takes” video clip classifier model). For example, as shown in FIG. 3G and in response to a detected selection from tab 310c, the digital content comprehension system 102 can generate and provide the “Build” button 336 within the digital content comprehension graphical user interface 304. In response to a detected selection of the “Build” button 336, the digital content comprehension system 102 can generate and begin training the new video clip classifier model.

[0058] For example, as shown in FIG. 3H, the digital content comprehension system 102 can build the new video clip classifier model and then apply the video clip classifier model to the training corpus of digital video clips. As a result, the video clip classifier model generates classification category predictions for each of the training digital video clips within various confidence levels. To illustrate, Petition 870250084254, dated 09 / 18 / 2025, page 45 / 91 Figure 3H shows the classification category prediction display area 342, including classification category prediction displays—including a classification category prediction display 344a—classified into positive confidence levels and negative confidence levels. For example, the various confidence levels include a high positive confidence level 338a and a low positive confidence level 338b, and a high negative confidence level 340a and a low negative confidence level 340b.

[0059] In one or more implementations, the digital content comprehension system 102 classifies classification category predictions under the high positive confidence level 338a in response to the video clip classifier model, generating prediction scores for those classification category predictions that are greater than a threshold value. For example, as shown in FIG. 3I, the digital content comprehension system 102 classifies classification category prediction displays 344b, 344a, and 344c under the high positive confidence level 338a in response to the classification scores 328b, 328a, and 328c associated with these classification category predictions being greater than a predetermined threshold.In at least one implementation, classification scores 328b, 328a, and 328c indicate the probability that the classification category prediction associated with the training digital video clip will include the desired representation indicated by the video clip classifier model title (e.g., “Happy Takes”). In other words, the probability that the training digital video clip will portray the concept, both positively and negatively, indicated by the corpus of training digital video clips. In at least one implementation, as shown in FIG. 3I, the 102 digital content comprehension system can display classification category prediction displays 344a-344c within the area. Petition 870250084254, dated 09 / 18 / 2025, page 46 / 91 29 / 45 of category 342 ranking prediction display in ranked order based on their ranking scores 328a-328c.

[0060] As mentioned above, the 102 digital content comprehension system also classifies negative classification category predictions into confidence levels (i.e., confidence levels about whether the digital video clips in the training corpus do not include a desired representation associated with the video clip classifier model title). For example, as shown in FIG. 3J and in response to a detected selection of the high negative confidence level 340a, the 102 digital content comprehension system can update the classification category prediction display area 342 with the classification category prediction displays 346a, 346b, and 346c.As indicated by classification scores 328a-328c, the predictions of the classification category under the high negative confidence level 340a are associated with video clips in which the video clip classifier model has too much confidence and do not portray the concept or theme indicated by the video clip classifier model's title (e.g., "Happy Takes").

[0061] Furthermore, as shown in FIG. 3K and in response to a detected selection of the low positive confidence level 338b, the digital content comprehension system 102 can update the classification category prediction display area 342 with classification category predictions with intermediate or borderline classification scores 328a, 328b. For example, the digital video clips associated with the classification category prediction displays 348a and 348b may not clearly or explicitly represent the concept or theme indicated by the video clip classifier model title. To illustrate, a digital video clip may depict Petition 870250084254, dated 09 / 18 / 2025, page 47 / 91 30 / 45 a smiling character, but in a scene with a negative context. In this way, the video clip classifier model can generate prediction scores for these digital video clips that are neither strongly positive nor strongly negative. For such classification category 348a348b prediction displays, the digital content comprehension system 102 can receive additional accuracy confirmations through the label option 330a, 330b. For example, the client computing device user 106 can indicate that the classification category 348a prediction display is positively associated with the concept “Happy Takes” by selecting the positive radio button 331b.

[0062] Similarly, as shown in FIG. 3L and in response to a detected selection of the low negative confidence level 340b, the digital content comprehension system 102 can update the classification category prediction display area 342 with a classification category prediction display 350a with an intermediate classification score in relation to negative concepts related to the concept or theme indicated by the video clip classifier model title. Just as in the low positive confidence level 338b, the classification category prediction display 350a under the low negative confidence level 340b can be associated with a digital video clip that may or may not depict a negative concept that corresponds to Happy Shots. The digital content comprehension system 102 can explicitly label the classification category prediction display 350a in response to detected selections via the label option 330a.

[0063] In one or more implementations, the digital content comprehension system 102 can retrain the video clip classifier model based on user confirmations detected through label options 331a, 331b under Petition 870250084254, dated 09 / 18 / 2025, page 48 / 91 31 / 45 the high positive confidence level 338a, the low positive confidence level 338b, the high negative confidence level 340a, and the low negative confidence level 340b. For example, the digital content comprehension system 102 can rename digital video clips within the training corpus of digital video clips to reflect user confirmations. The digital content comprehension system 102 can additionally train the video clip classifier model with the updated corpus.

[0064] Furthermore, the 102 digital content comprehension system can reapply the video clip classifier model through additional training cycles. For example, the 102 digital content comprehension system can reapply the video clip classifier model to the training corpus of digital video clips, even after all training digital video clips have been labeled. In each additional training cycle, the user can re-label the divergent predictions generated by the video clip classifier model. In some implementations, divergent predictions generated by the video clip classifier model may signal the need for additional training digital video clips to be added to the training corpus so that the video clip classifier model can better learn a specific concept.

[0065] With the trained video clip classifier model, the 102 digital content comprehension system can apply the video clip classifier model to digital video clips that are not part of the digital video clip training corpus. For example, as shown in FIG. 3M and in response to a detected selection from tab 310d, the 102 digital content comprehension system can configure the application of the video clip classifier model. For example, the Petition 870250084254, dated 09 / 18 / 2025, page 49 / 91 The 32 / 45 digital content comprehension system 102 can update the 304 digital content comprehension user interface to include options 352a, 352b, 352c, and 352d. For example, in response to a detected selection of option 352b, 352c, or 352d, the 102 digital content comprehension system can publish the video clip classifier model to multiple outputs.

[0066] In response to a detected selection of option 352a, the digital content comprehension system 102 can provide title input 354, take number input 356, and sorting input 358. For example, the digital content comprehension system 102 can identify a short-form digital video (e.g., a TV episode), a long-form digital video (e.g., a movie), or a season of short-form digital videos, according to the user input detected via title input 354. In one or more implementations, the digital content comprehension system 102 can identify the digital video indicated by title input 354 based on a numeric identifier, a title, a genre, and / or a keyword.

[0067] After identifying the digital video indicated by title entry 354, the digital content comprehension system 102 can analyze the digital video into digital video clips. For example, since the video clip classifier model can be trained to operate in connection with video clips rather than complete digital videos, the digital content comprehension system 102 can analyze the digital video into clips by identifying portions of continuous digital video footage between two cuts within the complete digital video. To illustrate, the digital content comprehension system 102 can identify transitions between two shots or scenes as cuts and can analyze the digital video based on the identified transitions. In this way, the analyzed video clips can have different durations and include different representations. Petition 870250084254, dated 09 / 18 / 2025, pp. 50 / 91 33 / 45

[0068] The 102 digital content comprehension system can additionally apply the video clip classifier model to the analyzed digital video clips to generate video clip classification scores. As discussed above, the classification score of each digital video clip can predict the probability that this clip includes the desired representation (e.g., the desired object, the desired subject, the desired character, the desired emotion, etc.) that the video clip classifier model was trained to identify. In the example shown in FIGS. 3A-3N, the video clip classifier model can generate scores for digital video clips, predicting whether each one represents a “Happy Take.”

[0069] In one or more implementations, the digital content comprehension system 102 can present the results of the video clip classifier model according to the 356 shot input and the 358 sorting input. For example, the digital content comprehension system 102 can generate a display of the analyzed digital video clips from the digital video (e.g., “81031991 The Witcher: “Season 2”) that includes the 10 highest-scoring digital video clips in sorting order (e.g., from highest score to lowest score). In at least one implementation, the digital content comprehension system 102 can analyze the digital video, apply the video clip classifier model, and generate the results display in response to a detected selection of the “Get 360 Shots” button.

[0070] To illustrate, as shown in FIG. 3N, the digital content comprehension system 102 can generate the display of results 361, including the suggested digital video clip displays 362a and 362b. As shown, the digital content comprehension system 102 can classify Petition 870250084254, dated 09 / 18 / 2025, pp. 51 / 91 34 / 45 as suggested digital video clip displays 362a, 362b according to their ranking scores 328a, 328b (e.g., as dictated by sorting entry 358).

[0071] As mentioned above and as shown in FIG. 4, the digital content comprehension system 102 performs several functions related to the automatic identification of digital video clips for inclusion in media assets, such as trailers and previews. FIG. 4 is a block diagram 400 of the digital content comprehension system 102 operating in the memory 114 of the server(s) 104 while performing these functions. Thus, FIG. 4 provides additional details regarding these functions. For example, as shown in FIG. 4, the digital content comprehension system 102 may include a digital video analysis manager 402, a video clip classifier model manager 404, and a graphical user interface manager 406.

[0072] In certain implementations, the digital content comprehension system 102 may represent one or more software applications, modules, or programs that, when executed by a computing device, may cause the computing device to perform one or more tasks. For example, and as will be described in greater detail below, one or more of the digital video analysis managers 402, the video clip classifier model manager 404, or the graphical user interface manager 406 may represent software stored and configured to run on one or more computing devices, such as the server(s) 104. One or more of the digital video analysis manager 402, the video clip classifier model manager 404, and the graphical user interface manager 406 of the digital content comprehension system 102 shown in FIG.4 can also represent all or portions of one or more special-purpose computers designed to perform one or more tasks. Petition 870250084254, dated 09 / 18 / 2025, pp. 52 / 91 35 / 45

[0073] As mentioned above and as shown in FIG. 4, the digital content comprehension system 102 may include the digital video analysis manager 402. In one or more implementations, the digital video analysis manager 402 handles tasks associated with separating a digital video into clips. For example, the digital video analysis manager 402 may identify scene cuts or transitions within the digital video. In at least one implementation, the digital video analysis manager 402 may additionally generate digital video clips, including the footage between the identified cuts or transitions. In some implementations, the digital video analysis manager 402 adds metadata to the digital video clips, which may include the title of the digital video from which the clips were analyzed, timestamps where the clips begin in their associated digital videos, a duration of the digital video clips, and so on.

[0074] As mentioned above and as shown in FIG. 4, the digital content comprehension system 102 may include the video clip classifier model manager 404. In one or more implementations, the video clip classifier model manager 404 handles tasks associated with the generation, training, and application of a video clip classifier model. For example, the video clip classifier model manager 404 may generate a video clip classifier model including a binary classifier machine learning model. The video clip classifier model manager 404 may further train the video clip classifier model based on a training corpus of digital video clips to make binary predictions.After the video clip classifier model is trained, the 404 video clip classifier model manager can apply the video clip classifier model to video clips. Petition 870250084254, dated 09 / 18 / 2025, pp. 53 / 91 36 / 45 digital signatures that are not part of the training corpus.

[0075] Although the examples and implementations discussed here include video clip classifier models, in other implementations, the 102 digital content comprehension system can generate, train, and apply other types of classifier models. For example, the 102 digital content comprehension system can generate and train audio clip classifier models and / or script text classifier models. Similarly, while the implementations discussed here work in connection with digital video clips, other implementations may work in connection with short-form digital videos and / or other longer digital video segments. Furthermore, in other implementations, the 404 video clip classifier model manager can generate a video clip classifier model including a machine learning model that is different from and / or more sophisticated than a binary classifier machine learning model.

[0076] Furthermore, the examples discussed here focus on identifying video clips for generating video assets such as previews and trailers. In further implementations, the 404 video clip classifier model manager generates and trains video clip classifier models to identify clips within a digital video that include undesirable content (e.g., profanity, nudity, violence). Based on these clip identifications, other systems can assign ratings to digital videos, issue parental advisories associated with digital videos, filter digital videos, etc.

[0077] In addition, in one or more implementations, the 404 video clip classifier model manager can train and retrain a video clip classifier model in a non-linear fashion and in multiple iterations. In other Petition 870250084254, dated 09 / 18 / 2025, pp. 54 / 91 37 / 45 words, the 404 video clip classifier model manager may not generate and train the video clip classifier model in a specific sequence with respect to the creation of the training corpus and the application of the video clip classifier model to untrained digital video clips. To illustrate, the 404 video clip classifier model manager allows training and retraining at any point in the process represented by FIGS. 3A-3N. For example, the 404 video clip classifier model manager can retrain a video clip classifier even after it has been published to one or more outputs and / or applied to new untrained digital video clips to generate classification predictions.In this way, the 404 video clip classifier model manager encourages continuous retraining of a video clip classifier model to improve the efficiency and accuracy of that model.

[0078] In one or more implementations, the 404 video clip classifier model manager may additionally handle tasks associated with generating a corpus of training digital video clips. For example, the 404 video clip classifier model manager may fetch repository 110 based on search queries, receive user acknowledgments associated with training digital video clips, and generate a corpus of training digital video clips based on the user acknowledgments. In at least one implementation, the 404 video clip classifier model manager may allow modifications to a corpus of training digital video clips at any point during the process illustrated in FIGS. 3A-3N.

[0079] As mentioned above and as shown in FIG. 4, the digital content comprehension system 102 may include the graphical user interface manager 406. In one or more implementations, the graphical interface manager Petition 870250084254, dated 09 / 18 / 2025, pp. 55 / 91 User interface 406 generates and updates the digital content comprehension graphical user interface 304. For example, the user interface manager 406 may include or be associated with a web server that generates web browser instructions that cause the web browser 108 to render the digital content comprehension graphical user interface 304 on the client computing device 106. The user interface manager 406 may further update the digital content comprehension graphical user interface 304 based on outputs from the digital video analysis manager 402 or the video clip classifier model manager 404. The user interface manager 406 may further update the digital content comprehension graphical user interface 304 based on user interactions detected in connection with the web browser 108 on the client computing device 106.

[0080] As shown in FIGS. 1 and 4, the digital content comprehension system 102 and the client computing device 106 may include one or more physical processors, such as the physical processor 118. The physical processor 118 may generally represent any type or form of processing unit implemented in hardware capable of interpreting and / or executing computer-readable instructions. In one implementation, the physical processor 118 may access and / or modify one or more of the components of the digital content comprehension system 102. Examples of physical processors include, without limitation, microprocessors, microcontrollers, Central Processing Units (CPUs), Field Programmable Gate Arrays (FPGAs) implementing softcore processors, Application-Specific Integrated Circuits (ASICs), parts of one or more thereof, variations or combinations thereof, and / or any other suitable physical processor. Petition 870250084254, dated 09 / 18 / 2025, pp. 56 / 91 39 / 45

[0081] In addition, the server(s) 104 and the client computing device 106 may include memory 114. In one or more implementations, memory 114 generally represents any type or form of volatile or non-volatile storage device or medium capable of storing computer-readable data and / or instructions. In one example, memory 114 may store, load, and / or maintain one or more components of the digital content comprehension system 102. Examples of memory 114 may include, without limitation, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drives (HDDs), solid-state drives (SSDs), optical disk drives, caches, variations or combinations thereof, and / or any other suitable storage memory.

[0082] Furthermore, as shown in FIG. 4, the server(s) 104 and the client computing device 106 may include additional items 116. On the server(s) 104, the additional items 116 may include the digital content repository 110. As mentioned above, the digital content repository 110 may include digital training video clips, corpora of digital training video clips, and other digital videos. In some implementations, the digital content repository 110 may also include other types of digital content, such as dialogue audio, soundtrack, digital script text, and still images.

[0083] In summary, the 102 digital content comprehension system allows for the accurate and efficient generation of assets based on video clips, such as trailers and previews. For example, the 102 digital content comprehension system generates robust corpora from training digital video clips associated with desired representations that may include abstract and / or subjective ideas. As discussed above, the system Petition 870250084254, dated 09 / 18 / 2025, pp. 57 / 91 The 40 / 45 digital content comprehension 102 system creates greater efficiency in training and using video clip classifier models by labeling high-confidence training data—which includes video clips that respond positively to the desired representation and negatively to the desired representation—and low-confidence training data. The 40 / 45 digital content comprehension 102 system further trains video clip classifier models using these generated digital training video clips. Finally, the 40 / 45 digital content comprehension 102 system can apply trained video clip classifier models to new digital video clips and publish the trained video clip classifier models for use through additional outputs.In one or more implementations, as discussed herein, the 102 digital content comprehension system facilitates the process of generating, training, and applying video clip classifier models by generating and updating a digital content comprehension graphical user interface. Examples of Modalities

[0084] Example 1: A computer-implemented method for generating classification category predictions for digital video clips as to whether the digital video clips represent a specified object, subject, scene type, emotion, and so on. For example, the method might include generating, by applying a video clip classifier model to a corpus of training digital video clips that respond to a received search input and within a digital content comprehension graphical user interface, classification category prediction displays for the training corpus of digital video clips, retraining the video clip classifier model based on user confirmations, detected through the digital content comprehension graphical user interface, regarding the accuracies of Petition 870250084254, dated 09 / 18 / 2025, pp. 58 / 91 41 / 45 classification scores generated by the video clip classifier model that correspond to the training digital video clips, analyze a digital video in digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface, and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

[0085] Example 2: The computer-implemented method of Example 1, including additionally generating the corpus of digital training video clips, iteratively receiving search input related to at least one of an object, an action, a shot type, an editing technique, a character, or a story theme, and receiving search input related to at least one of an object, an action, a shot type, an editing technique, a character, or a story theme; and identifying, within a repository of digital training video clips, a plurality of digital training video clips that respond to the received search input.

[0086] Example 3: The computer-implemented method of either Examples 1 and 2, wherein the classification category prediction displays within the digital content comprehension graphical user interface comprise a playback window loaded with a training digital video clip corresponding to the classification category prediction, a title of a digital video from which the training digital video clip corresponding to the classification category prediction came, and an option to positively or negatively confirm the digital video clip of Petition 870250084254, dated 09 / 18 / 2025, pp. 59 / 91 42 / 45 training corresponding to the classification category prediction.

[0087] Example 4: The computer-implemented method of any of Examples 1 through 3, wherein generating classification category prediction displays within the digital content understanding graphical user interface additionally includes classifying the classification category prediction displays into high confidence levels and low confidence levels and updating the classification category prediction displays within the digital content understanding graphical user interface according to the high confidence levels and the low confidence levels.

[0088] Example 5: The computer-implemented method of any of Examples 1 through 4, including additionally detecting user confirmations as to the accuracy of classification scores generated by the video clip classifier model, detecting at least one of a first user input corresponding to a positive confirmation of a first video clip included in the digital training video clips, or a second user input corresponding to a negative confirmation of the first video clip included in the digital training video clips.

[0089] Example 6: The computer-implemented method of any of Examples 1 to 5, including additionally detecting the digital video selection by means of the digital content comprehension graphical user interface, detecting a selection of at least one short-form digital video, one long-form digital video, or a season of short-form digital videos.

[0090] Example 7: The computer-implemented method of any of Examples 1 through 6, wherein analyzing digital video into digital video clips includes analyzing digital video into portions of continuous digital video footage between two Petition 870250084254, dated 09 / 18 / 2025, pp. 60 / 91 43 / 45 cuts.

[0091] Example 8: The computer-implemented method of any of Examples 1 through 7, in which generating the suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface includes generating input vectors based on the digital video clips, applying the retrained video clip classifier model to the generated input vectors, receiving, from the retrained video clip classifier model, classification scores for the digital video clips that match the received search input, generating, for the digital video clips, suggested digital video clip displays,and replace category rating prediction displays with suggested digital video clip displays within the digital content comprehension user interface according to the rating scores for the digital video clips.

[0092] In some examples, a system may include at least one processor and physical memory including computer executable instructions that, when executed by at least one processor, cause at least one processor to perform various acts. For example, the computer executable instructions may cause at least one processor to perform acts including generating, by applying a video clip classifier model to a corpus of digital training video clips that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the corpus of digital training video clips, retraining the video clip classifier model based on user confirmations, detected through the graphical user interface of Petition 870250084254, dated 09 / 18 / 2025, pp. 61 / 91 44 / 45 Digital content comprehension, regarding the accuracy of classification scores generated by the video clip classifier model that correspond to the training digital video clips, analyze a digital video into digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface, and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

[0093] In addition, in some examples, a non-transient computer-readable medium may include one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to perform various acts.For example, one or more computer-executable instructions can cause the computing device to generate, by applying a video clip classifier model to a corpus of training digital video clips that respond to a received search input associated with a desired representation and within a digital content understanding graphical user interface, classification category prediction displays for the training corpus of digital video clips, retrain the video clip classifier model based on user confirmations, detected through the digital content understanding graphical user interface, regarding the accuracies of classification scores generated by the video clip classifier model that correspond to the training digital video clips, analyze a digital video in digital video clips in response to the detection of a digital video selection through the. Petition 870250084254, dated 09 / 18 / 2025, pp. 62 / 91 45 / 45 digital content comprehension graphical user interface, and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

[0094] Unless otherwise indicated, the terms connected and coupled to (and their derivatives), as used in the descriptive report and claims, should be interpreted as allowing both direct and indirect connections (i.e., through other elements or components). Furthermore, the terms a or an, as used in the descriptive report and claims, should be interpreted as meaning at least one of. Finally, for ease of use, the terms including and having (and their derivatives), as used in the descriptive report and claims, are interchangeable and have the same meaning as the word comprising. Petition 870250084254, dated 09 / 18 / 2025, pp. 63 / 91

Claims

1 / 10 CLAIMS 1. A computer-implemented method, characterized in that it comprises: generating, by applying a video clip classifier model to a corpus of training digital video clips that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the training corpus of digital video clips; retraining the video clip classifier model based on user confirmations, detected through the digital content comprehension graphical user interface, regarding the accuracy of the classification scores generated by the video clip classifier model that correspond to the training digital video clips.Analyze a digital video into digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface; and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

2. A computer-implemented method according to claim 1, characterized in that it further comprises generating the corpus of digital training video clips by iteratively: receiving search input relating to at least one of an object, action, scene type, editing technique, character, or story theme; and identifying, within a repository of digital training video clips, a plurality of digital training video clips that respond to the received search input.

3. A computer-implemented method according to claim 1, characterized in that the classification category prediction displays within the digital content comprehension graphical user interface comprise a playback window loaded with a training digital video clip corresponding to the classification category prediction display, a title of a digital video from which the training digital video clip corresponding to the classification category prediction display came, and an option to positively or negatively confirm the training digital video clip corresponding to the classification category prediction display.

4. A computer-implemented method according to claim 3, characterized in that generating the classification category prediction displays within the digital content comprehension graphical user interface further comprises: classifying the classification category prediction displays into high confidence levels and low confidence levels; and updating the classification category prediction displays within the digital content comprehension graphical user interface according to the high confidence levels and low confidence levels.

5. A computer-implemented method according to claim 3, characterized in that it further comprises detecting user confirmations regarding the accuracy of the classification scores generated by the video clip classifier model, detecting at least one of: a first user input corresponding to a positive confirmation of a first video clip included in the digital training video clips; or a second user input corresponding to a negative confirmation of the first video clip included in the digital training video clips.

6. A computer-implemented method according to claim 1, characterized in that it further comprises detecting the selection of digital video by means of a digital content comprehension graphical user interface, detecting a selection of at least one short-form digital video, one long-form digital video, or a series of short-form digital videos.

7. A computer-implemented method according to claim 1, characterized in that analyzing digital video in digital video clips comprises analyzing digital video in portions of continuous digital video footage between two cuts.

8. A computer-implemented method according to claim 1, characterized in that generating suggested digital video clip displays that replace classification category prediction displays within the digital content comprehension graphical user interface comprises: generating input vectors based on digital video clips; applying the retrained video clip classifier model to the generated input vectors; receiving, from the retrained video clip classifier model, classification scores for the digital video clips that match the received search input; generating, for the digital video clips, clip displays. Petition 870250084254, dated 09 / 18 / 2025, p.66 / 91 4 / 10 suggested video; and replace the classification category prediction displays with the suggested digital video clip displays for the digital video clips within the digital content comprehension graphical user interface according to the classification scores for the digital video clips.

9. A system characterized in that it comprises: at least one physical processor; and physical memory comprising computer executable instructions which, when executed by at least one physical processor, cause the at least one physical processor to perform acts comprising: generating, by applying a video clip classifier model to a corpus of digital training video clips that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the corpus of digital training video clips;Retrain the video clip classifier model based on user confirmations, detected through the digital content comprehension graphical user interface, regarding the accuracy of the classification scores generated by the video clip classifier model that correspond to the training digital video clips; analyze a digital video in digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface; and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

10. A system according to claim 9, characterized in that additional computer-executable instructions, when executed by at least one physical processor, cause that physical processor to generate the corpus of digital training video clips by iteratively: receiving search input related to at least one of an object, action, scene type, editing technique, character, or story theme; and identifying, within a repository of digital training video clips, a plurality of digital training video clips that respond to the received search input.

11. System, according to claim 9, characterized in that the classification category prediction displays within the digital content comprehension graphical user interface comprise a playback window loaded with a training digital video clip corresponding to the classification category prediction, a title of a digital video from which the training digital video clip corresponding to the classification category prediction came, and an option to positively or negatively confirm the training digital video clip corresponding to the classification category prediction.

12. System according to claim 11, characterized in that additional computer-executable instructions, when executed by at least one physical processor, cause that physical processor to generate the classification category prediction displays within the digital content comprehension graphical user interface by: Petition 870250084254, dated 09 / 18 / 2025, p. 68 / 91 6 / 10 classifying the classification category prediction displays into positive confidence levels and negative confidence levels; and updating the classification category prediction displays within the digital content comprehension graphical user interface according to the positive confidence levels and negative confidence levels.

13. System according to claim 11, characterized in that additional computer-executable instructions, when executed by at least one physical processor, cause that physical processor to detect user confirmations regarding the accuracy of the classification scores generated by the classified video clip model, detecting at least one of: a first user input corresponding to a positive confirmation of a first video clip included in the digital training video clips; or a second user input corresponding to a negative confirmation of the first video clip included in the digital training video clips.

14. System according to claim 9, characterized in that additional computer-executable instructions, when executed by at least one physical processor, cause that physical processor to detect the selection of digital video through the digital content comprehension graphical user interface, detecting a selection of at least one short-form digital video, one long-form digital video, or a season of short-form digital videos.

15. System, according to claim 9, characterized in that additional computer-executable instructions, when executed by at least one physical processor, cause the at least one physical processor to analyze the digital video into digital video clips, analyzing the digital video in portions of continuous digital video footage between two cuts.

16. A system according to claim 9, characterized in that additional computer-executable instructions, when executed by at least one physical processor, cause that physical processor to generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface by: generating input vectors based on the digital video clips; applying the retrained video clip classifier model to the generated input vectors; receiving, from the retrained video clip classifier model, classification scores for the digital video clips that match the received search input; generating, for the digital video clips, suggested digital video clip displays;and replace the classification category prediction displays with suggested digital video clip displays within the digital content comprehension graphical user interface according to the classification scores for the digital video clips.

17. Non-transient computer-readable medium, characterized in that it comprises one or more computer-executable instructions which, when executed by at least one processor of a computing device, cause the computing device to: generate, applying a video clip classifier model to a corpus of digital training video clips Petition 870250084254, dated 09 / 18 / 2025, p.70 / 91 8 / 10 that respond to a received search input associated with a desired representation and within a digital content comprehension graphical user interface, classification category prediction displays for the training corpus of digital video clips; retrain the video clip classifier model based on user confirmations, detected through the digital content comprehension graphical user interface, regarding the accuracy of the classification scores generated by the video clip classifier model that correspond to the training digital video clips.Analyze a digital video into digital video clips in response to the detection of a digital video selection through the digital content comprehension graphical user interface; and generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface based on the application of the retrained video clip classifier model to the digital video clips.

18. Non-transient computer-readable medium according to claim 17, characterized in that it comprises one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to generate the corpus of digital training video clips by iteratively: receiving search input relating to at least one of an object, an action, a scene type, an editing technique, a character, or a story theme; and identifying, within a repository of digital training video clips, a plurality of digital training video clips that respond positively to the received search input.

19. Non-transient computer-readable medium according to claim 17, characterized in that the classification category prediction displays within the digital content comprehension graphical user interface comprise a playback window loaded with a training digital video clip corresponding to the classification category prediction display, a title of a digital video from which the training digital video clip corresponding to the classification category prediction display came, and an option to positively or negatively confirm the training digital video clip corresponding to the classification category prediction display.

20. Non-transient computer-readable medium according to claim 17, characterized in that it comprises one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to generate suggested digital video clip displays that replace the classification category prediction displays within the digital content comprehension graphical user interface by: generating input vectors based on the digital video clips; applying the retrained video clip classifier model to the generated input vectors; receiving, from the retrained video clip classifier model, classification scores for the digital video clips that match the received search input; generating, for the digital video clips, suggested digital video clip displays;and replace the category prediction displays of Petition 870250084254, dated 09 / 18 / 2025, p. 72 / 91 10 / 10 classification with the suggested digital video clip displays for the digital video clips within the digital content comprehension graphical user interface according to the classification scores for the digital video clips. Petition 870250084254, dated 09 / 18 / 2025, p. 73 / 91;