Interact with semantic video clips using interactive tiles.

By using machine learning to detect video features and build graph models, interactive tiles and editor interfaces are generated, solving the problems of traditional video editing tools being tedious and inefficient, and achieving a more intuitive and flexible video editing experience.

CN115407912BActive Publication Date: 2026-03-13ADOBE INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Traditional video editing tools rely on the selection of specific video frames or time ranges, resulting in a tedious and inefficient interaction method that exceeds the skill level of many users and lacks flexibility and efficiency.

Method used

By using machine learning models to detect features in videos, it generates default segments, search segments, capture alignment point segments, and thumbnail segments, builds a graph model to calculate the shortest path, and provides an interactive tile and editor interface that allows users to select, trim, and replay video clips on a semantic basis.

Benefits of technology

It provides a more intuitive and flexible way to interact with videos, allowing users to quickly select, trim, and export semantically meaningful video clips, improving editing efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115407912B_ABST
    Figure CN115407912B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention relate to interacting with semantic video segments via interactive tiles. The embodiments relate to interactive tiles representing video segments that are segments of a video. In some embodiments, each interactive tile represents a different video segment from a specific video segment (e.g., a default video segment). Each interactive tile includes a thumbnail (e.g., the first frame of the video segment represented by the tile), a script from the beginning of the video segment, a visualization of detected faces in the video segment, and one or more faceted timelines that visualize the categories of detected features (e.g., visualizations of detected visual scenes, audio classifications, and visual artifacts). In some embodiments, interacting with a specific interactive tile can navigate to the corresponding portion of the video, add the corresponding video segment to the selection, and / or scan through the tile thumbnail.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is a continuation-in-part of U.S. Patent Application No. 17 / 017,344, filed September 10, 2020, entitled “Segmentation and Hierarchical Clustering of Video,” the entire contents of which are incorporated herein by reference. Background Technology

[0003] In recent years, video usage has surged, finding diverse applications across virtually every industry, from film and television to advertising and social media. Businesses and individuals frequently create and share video content in a wide range of contexts, including, to name just a few, presentations, tutorials, commentaries, news and sports clips, blogs, product reviews, recommendations, comedy, dance, music, movies, and video games. Videos can be captured using cameras, generated using animation or rendering tools, edited using various types of video editing software, and shared through a variety of channels. In fact, recent advancements in digital cameras, smartphones, social media, and other technologies have provided numerous new methods that make it easier for even beginners to capture and share videos. With these new methods of capturing and sharing video, the demand for video editing features is also increasing.

[0004] Typically, video editing involves selecting video frames and performing some type of action on those frames or associated audio. Some common operations include importing, trimming, cropping, rearranging, applying transitions and effects, adjusting colors, adding titles and graphics, exporting, and so on. Video editing software like Pro and Adobe Premiere Element typically includes a graphical user interface (GUI) that presents a video timeline representing the video frames and allows users to select specific frames and perform actions on them. However, traditional video editing can be tedious, challenging, and even beyond the skill level of many users. Summary of the Invention

[0005] Embodiments of this invention relate to video segmentation and various interactive methods for video browsing, editing, and playback. In one example embodiment, video is ingested by detecting various features using one or more machine learning models, and one or more video segments are generated based on the detected features. The detected features serve as the basis for one or more video segments, such as default segments, search segments based on user queries, snap-alignment segmentation that identifies selected snap points, and / or thumbnail segmentation that identifies which parts of the video timeline are illustrated using different thumbnails. Finder and editor interfaces display the different segments, providing users with the ability to browse, select, edit, and / or play back semantically meaningful video clips.

[0006] In some embodiments, video segmentation is computed by determining candidate boundaries based on feature boundaries detected from one or more feature tracks; modeling different segmentation options by constructing a graph having nodes representing candidate boundaries, edges representing candidate segments, and edge weights representing cut costs; and computed by solving a shortest path problem to find paths through edges (segments) that minimize the sum of edge weights (cut costs) along that path. In some embodiments, one or more aspects of the segmentation routine depend on the type of segmentation (e.g., default, search, snap alignment, thumbnail). As a non-limiting example, candidate cut points are determined differently, different edges are used for different types of segmentation, and / or different cut costs are used for edge weights. In one example embodiment, default segmentation is computed based on a desired set of detected features, such as detected sentences, faces, and visual scenes.

[0007] In some embodiments, the finder interface displays video segments (such as a default segment) with interactive tiles representing video clips within the video segment and features detected in each video clip. In some embodiments, each interactive tile represents a different video clip from a specific video segment (e.g., the default video segment) and includes a thumbnail (e.g., the first frame of the video clip represented by the tile), a transcript from the beginning of the video clip, a visualization of the video clip, a visualization of faces detected in the video clip, and / or one or more faceted timelines visualizing the categories of detected features (e.g., detected visual scenes, audio classifications, visualizations of visual artifacts). In one embodiment, different ways of interacting with a specific interactive tile are used to navigate to the corresponding portion of the video, first select to add the corresponding video clip, and / or scrub through the tile thumbnails.

[0008] In some embodiments, search segments are computed based on a query. Initially, a first segment, such as a default segment, is displayed (e.g., as interactive tiles in the finder interface, as a video timeline in the editor interface), and the default segment is re-segmented in response to a user query. A query can take the form of keywords and one or more selected facets from detected feature categories. Keywords are searched to find detected script words, detected object or action tags, or detected audio event tags that match the keywords. Selected facets are searched to find detected instances of the selected facets. Each video segment matching the query is re-segmented by solving a shortest path problem of a graph modeled over the different segmentation options. The finder interface updates the interactive tiles to represent the search segments. Thus, the search is used to decompose the interactive tiles to represent smaller units of the query-based video.

[0009] In some embodiments, the finder interface is used to browse videos and add video clips to the selection. After adding the desired set of video clips to the selection, the user switches to the editor interface to perform one or more enhancements or other editing operations on the selected video clips. In some embodiments, the editor interface initializes the video timeline using representations of the video clips selected from the finder interface, such as a composite video timeline representing a synthesized video formed from the selected video clips, wherein the boundaries of the video clips are illustrated as the bottom layer of the composite timeline. In some embodiments, the composite video timeline includes visualizations of detected features and corresponding feature ranges to aid in selecting, trimming, and editing video clips.

[0010] Some embodiments involve capture alignment point segments that define the location of a selectable capture alignment point for selecting video segments through interaction with the video timeline. Candidate capture alignment points are determined based on the boundaries of a feature range in the video, indicating when instances of detected features are present in the video. In some embodiments, the interval between candidate capture alignment points is penalized because it is separated by a minimum duration corresponding to the minimum pixel interval between consecutive capture alignment points on the video timeline. The capture alignment point segments are computed by solving a shortest path problem for a graph modeled with different capture alignment point locations and intervals. When a user clicks or taps and drags on the video timeline, the selected capture alignment is moved to a capture alignment point defined by the capture alignment point segments. In some embodiments, the capture alignment point is displayed during the drag operation and disappears when the drag operation is released.

[0011] Some embodiments involve thumbnail segments that define the locations of displayed thumbnails on a video timeline. Candidate thumbnail locations are determined based on boundaries of video feature ranges that indicate when instances of detected features are present in the video. In some embodiments, candidate thumbnail intervals are penalized because they are separated by a minimum duration corresponding to the minimum pixel interval (e.g., the width of the thumbnail) between consecutive thumbnail locations on the video timeline. Thumbnail segments are computed by solving a shortest path problem for a graph modeled by different thumbnail locations and intervals. Thus, the video timeline displays thumbnails at locations defined by the thumbnail segments, where each thumbnail depicts a portion of the video associated with that location.

[0012] In some embodiments, the editor interface provides any number of editing functions for the selected video clip. Depending on the implementation, available editing functions include stylistic improvements to the content (e.g., wind noise reduction), improvements to the duration of the effects of omitted content (e.g., "hiding" a shot area, removing profanity, creating a time delay, shortening to n seconds), and / or contextual functions depending on the selected content (e.g., removing words from content with a corresponding script or adding a beep). Typically, the editor interface provides any suitable editing functionality, including rearranging, cropping, applying transitions or effects, adjusting colors, adding titles or graphics, etc. The resulting composite video can be played back, saved, exported, or otherwise manipulated.

[0013] Thus, this disclosure provides an intuitive video interaction technology that allows users to easily select, trim, replay, and export semantically meaningful video clips at the desired granularity, thereby providing creators and consumers with a more intuitive structure for interacting with videos.

[0014] This synopsis is provided to introduce, in a simplified form, a series of concepts that will be further described in the detailed description below. This synopsis is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. Attached Figure Description

[0015] The present invention will now be described in detail with reference to the accompanying drawings, wherein:

[0016] Figure 1A-Figure 1B This is a block diagram of an example computing system for video editing or playback according to an embodiment of the present invention;

[0017] Figures 2A-2F This is an illustration of an example technique for calculating candidate boundaries based on detected feature boundaries according to an embodiment of the present invention;

[0018] Figures 3A-3CThis is an illustration of an example technique for constructing a graph with paths representing different candidate segments according to an embodiment of the present invention;

[0019] Figures 4A-4C This is an illustration of an example technique for calculating boundary shearing costs based on a visual scene, according to an embodiment of the present invention.

[0020] Figure 5 This is an illustration of an example technique for calculating boundary shearing costs based on script features of segmented clustering, according to an embodiment of the present invention.

[0021] Figure 6 This is an illustration of an example technique for calculating boundary shearing costs based on detected faces, according to an embodiment of the present invention.

[0022] Figures 7A-7D This is an illustration of an example technique for calculating boundary shearing costs based on queries, according to an embodiment of the present invention;

[0023] Figure 8 This is an illustration of an example technique for calculating interval shearing cost based on the incoherence of detected features, according to an embodiment of the present invention.

[0024] Figures 9A-9B This is an illustration of an example technique for calculating interval shearing cost based on a query, according to an embodiment of the present invention;

[0025] Figure 10 This is an illustration of an example finder interface for browsing default segments and / or searching segments according to an embodiment of the present invention;

[0026] Figures 11A-11B This is an illustration of an example interactive tile according to an embodiment of the present invention;

[0027] Figures 12A-12D This is an illustration of an example faceted search menu according to an embodiment of the present invention;

[0028] Figure 13 This is an illustration of an example search in the finder interface according to an embodiment of the present invention;

[0029] Figure 14 This is an illustration of an example editor interface for video editing according to an embodiment of the present invention;

[0030] Figures 15A-15B This is an illustration of an example selection box selected using snap alignment according to an embodiment of the present invention;

[0031] Figures 16A-16B This is a flowchart illustrating a method for generating video segments using a graph model according to an embodiment of the present invention;

[0032] Figure 17 This is a flowchart illustrating a method for dividing a video into search segments according to an embodiment of the present invention;

[0033] Figure 18 This is a flowchart illustrating a method for navigating video using interactive tiles according to an embodiment of the present invention;

[0034] Figure 19 This is a flowchart illustrating a method for adding to the selection of video segments according to an embodiment of the present invention;

[0035] Figure 20 This is a flowchart illustrating a method for capturing and aligning selection boundaries for selecting video segments according to an embodiment of the present invention;

[0036] Figure 21 This is a flowchart illustrating a method for presenting a video timeline with thumbnails at locations defined by thumbnail segments according to an embodiment of the present invention;

[0037] Figure 22 This is a flowchart illustrating another method for presenting a video timeline with thumbnails at locations defined by thumbnail segments, according to an embodiment of the present invention; and

[0038] Figure 23 This is a block diagram of an example computing environment suitable for implementing embodiments of the present invention. Detailed Implementation

[0039] Overview

[0040] Video files, clips, or projects can typically be broken down into visual and audio elements. For example, video can be encoded or otherwise identified as a video track consisting of a series of still images (e.g., video frames) and an accompanying audio track consisting of one or more audio signals. Conventionally, video editing tools provide an interface that allows users to perform time-based editing on selected video frames. That is, conventional video editing typically involves representing the video as a series of fixed units with equal durations (e.g., multiple video frames) and presenting a video timeline that allows users to select and interact with specific video frames. However, this interaction method, which relies on the selection of specific video frames or corresponding time ranges, is inherently slow and granular, leading to editing workflows that are often perceived as tedious, challenging, or even beyond the skill level of many users. In other words, time-based video editing that requires selecting specific video frames or time ranges provides a fixed-granularity interaction method, resulting in an inflexible and inefficient interface. Thus, there is a need for an improved interface and improved interaction method for video editing tools.

[0041] Therefore, embodiments of the present invention relate to video segmentation and various interactive methods for video browsing, editing, and playback. In one example embodiment, video is ingested by detecting various features using one or more machine learning models, and one or more video segments are generated based on the detected features. More specifically, one or more machine learning models are used to process the video to detect features (e.g., scripts, linguistic features, speakers, faces, audio classification, visually similar scenes, visual artifacts, video objects or actions, audio events, software log events) and corresponding feature ranges where the detected features exist. The detected features serve as the basis for one or more video segments, such as default segments, search segments based on user queries, capture alignment point segments that identify selected capture alignment points, and / or thumbnail segments that identify which parts of the video timeline are illustrated with different thumbnails. Finder and editor interfaces display the different segments, providing users with the ability to browse, select, edit, and / or play back semantically meaningful video clips. Thus, this technology provides new methods for creating, editing, and consuming videos, offering creators and consumers a more intuitive structure for interacting with videos.

[0042] In one example implementation, candidate cut points are identified to identify video segments based on detected feature boundaries; different segmentation options are modeled by constructing a graph with nodes representing candidate cut points, edges representing candidate segments, and edge weights representing cut costs; and video segments are computed by solving a shortest path problem to find paths through edges (segments) that minimize the sum of edge weights (cut costs) along that path. In some embodiments, depending on the use case, the segmentation routine accepts different input parameters, such as specified feature tracks (e.g., predetermined, user-selected, etc.), user queries, the target minimum or maximum length of the video segments (which in some cases depends on the scaling level), the range of the video to be segmented, etc. Additionally or alternatively, in some cases, one or more aspects of the segmentation routine depend on the type of segmentation. As a non-limiting example, candidate cut points are determined differently, different edges are used for different types of segmentation, and / or different cut costs are used for edge weights. In some embodiments, the output is a representation of the complete set of disjoint (i.e., non-overlapping) video segments (i.e., covering the entire and / or specified range of the video).

[0043] In one example embodiment, the default segmentation is calculated based on a desired set of detected features, such as detected sentences, faces, and visual scenes. The finder interface displays the default segmentation with interactive tiles representing video clips within the default segmentation and the features detected in each video clip. In some embodiments, the finder interface displays the categories of features used by the default segmentation, accepts a selection of a desired set of feature categories, recalculates the default segmentation based on the selected feature categories, and updates the interactive tiles to represent the updated default segmentation. In an example implementation with search, the finder interface accepts queries (e.g., keywords and / or faceting), generates or triggers search segments to re-segment the default segmentation based on the query, and updates the interactive tiles to represent the search segments. Thus, search is used to break down the interactive tiles into smaller units representing the video based on the query. Additionally or alternatively, sliders or other interactive elements display input parameters (e.g., target minimum and maximum length of the video clips) for allowing the user to interactively control the size of the video clips represented by the tiles.

[0044] In one example embodiment, each interactive tile represents a video segment from a specific segment (e.g., default or search) and features detected in the video segment. For example, each tile displays a thumbnail (e.g., the first frame of the video segment), a script from the beginning of the video segment, a visualization of detected faces in the video segment, and one or more faceted timelines visualizing the categories of the detected features (e.g., visualizations of detected visual scenes, audio classifications, and visual artifacts). Clicking on the visualization of detected features in the interactive tile navigates to the video section containing the detected features.

[0045] In some embodiments, the finder interface includes a selected clip panel where a user can add video clips to the selection. Depending on the implementation, video clips are added to the selection in various ways, such as by dragging tiles into the selected clip panel, activating buttons or other interactive elements in interactive tiles, visual interaction with features detected in interactive tiles (e.g., right-clicking on a visualization to activate a context menu and adding the corresponding video clip to the selection from the context menu), interaction with scripts (e.g., highlighting, right-clicking to activate a context menu, and adding from the context menu to the selection), and / or other methods. In some embodiments, the finder interface includes one or more buttons or other interactive elements that switch to the editor interface to initialize the video timeline using a representation of the selected video clip in the selected clip panel.

[0046] In one example embodiment, the editor interface provides a composite video timeline representing a composite video formed from video segments selected in the finder interface, where the boundaries of the video segments are illustrated as the bottom layer of the composite timeline. In some embodiments, dragging along the composite timeline snaps the selected boundaries to snap alignment points defined by snap alignment point segmentation and / or the current zoom level. In an example implementation with search, the editor interface accepts queries (e.g., keywords and / or faceting), generates search segments that segment the video segments in the composite video based on the search, and presents a visualization of the search segments (e.g., by illustrating the boundaries of the video segments as the bottom layer of the composite timeline).

[0047] In some embodiments, the synthesized video timeline includes visualizations of detected features and corresponding feature ranges to aid in trimming. For example, when dragging on the synthesized timeline, visualizations of detected features can help inform which parts of the video to select (e.g., video segments containing visualized audio classifications). Additionally or alternatively, when dragging on the synthesized timeline, snap-align points—defined by snap-align point segments that represent certain feature boundaries—are visualized on the timeline (e.g., as vertical lines on the synthesized timeline), illustrating which parts of the video would be good snap-align points. In yet another example, clicking on a visualization of detected features (e.g., bars representing video segments with detected artifacts) prompts selection of segments of the synthesized video containing that detected feature. Thus, the editor interface provides an intuitive interface for selecting semantically meaningful video segments for editing.

[0048] In some embodiments, the composite timeline represents each video segment in the composite video, where one or more thumbnails illustrate representative video frames. In one example implementation, each video segment includes one or more thumbnails at timeline locations identified by thumbnail segmentation and / or the current zoom level, with longer video segments more likely to include multiple thumbnails. For example, as the user zooms in, more thumbnails appear at semantically meaningful locations, while any already visible thumbnails remain in place. Thus, thumbnails serve as landmarks to aid in navigating the video and selecting video segments.

[0049] In one example implementation, the editor interface provides any number of editing functions for the selected video clip. Depending on the implementation, available editing functions include stylistic improvements to transform content (e.g., wind noise reduction), improvements to the duration of the effects of omitted content (e.g., "hiding" shot areas, removing profanity, creating time delays, shortening to n seconds), and / or contextual functions depending on the selected content (e.g., removing words from content with corresponding scripts or adding beeps). Typically, the editor interface provides any suitable editing functionality, including rearranging, cropping, applying transitions or effects, adjusting colors, adding titles or graphics, etc. The resulting composite video can be played back, saved, exported, or otherwise manipulated.

[0050] Thus, this disclosure provides an intuitive video interaction technology that allows users to easily select, trim, replay, and export semantically meaningful video segments at a desired granularity. Instead of simply providing a video timeline segmented into fixed units (e.g., a frame, a second) of equal duration in a manner separate from semantic meaning, it reveals video segments with unequal durations and boundaries at semantically meaningful locations, not arbitrary ones, based on detected features. Therefore, this video interaction technology offers a more flexible and efficient interaction method, allowing users to quickly identify, select, and manipulate meaningful video blocks of potential interest. In this way, editors can now work faster, and consumers can now jump to sections of interest without watching the video.

[0051] Example video editing environment

[0052] Now for reference Figure 1A A block diagram of an example environment 100 suitable for implementing embodiments of the present invention is shown. Generally, environment 100 is suitable for video editing or playback and, among other things, facilitates video segmentation and interaction with the resulting video segments. Environment 100 includes a client device 102 and a server 150. In various embodiments, client device 102 and / or server 150 are any kind of computing device capable of facilitating video editing or playback, such as those referenced below. Figure 23 The computing device 2300 is described. Examples of computing devices include personal computers (PCs), laptops, mobile or mobile devices, smartphones, tablets, smartwatches, wearable computers, personal digital assistants (PDAs), music players or MP3 players, global positioning systems (GPS) or devices, video players, handheld communication devices, gaming devices or systems, entertainment systems, in-vehicle computer systems, embedded system controllers, cameras, remote controls, barcode scanners, computerized measuring devices, electronics, consumer electronic devices, workstations, some combination thereof, or any other suitable computing device.

[0053] In various implementations, components of environment 100 include computer storage media for storing information, including data, data structures, computer instructions (e.g., software program instructions, routines, or services), and / or models (e.g., machine learning models) used in some embodiments of the technology described herein. For example, in some implementations, client device 102, server 150, and / or storage device 190 include one or more data repositories (or computer data storage devices). Furthermore, although client device 102, server 150, and storage device 190... Figure 1A Each is depicted as a single component, but in some embodiments, the client device 102, server 150, and / or storage device 190 are implemented using any number of data repositories, and / or using cloud storage.

[0054] Components of environment 100 communicate with each other via network 103. In some embodiments, network 103 includes one or more local area networks (LANs), wide area networks (WANs), and / or other networks. Such networking environments are common in offices, enterprise-wide computer networks, intranets, and the Internet.

[0055] exist Figure 1A and Figure 1B In the example illustrated, client device 102 includes a video interaction engine 108, and server 150 includes a video segmentation tool 155. In various embodiments, the video interaction engine 108, video segmentation tool 155, and / or... Figure 1A and Figure 1B Any element illustrated herein is incorporated into or integrated into (e.g., corresponding applications on client device 102 and server 150, respectively) or into (multiple) attachments or plugins of (multiple) applications. In some embodiments, (multiple) applications are any applications capable of facilitating video editing or playback, such as standalone applications, mobile applications, web applications, etc. In some implementations, (multiple) applications include web applications, for example, those accessible via a web browser, at least partially hosted on a server side, etc. Additionally or alternatively, (multiple) applications include dedicated applications. In some cases, applications are integrated into the operating system (e.g., as a service). Example video editing applications include Adobe Premiere Pro and Adobe Premiere Elements.

[0056] In various embodiments, the functionality described herein is distributed across any number of devices. In some embodiments, video editing application 105 is hosted at least partially on a server side, enabling video interaction engine 108 and video segmentation tool 155 to coordinate (e.g., via network 103) to perform the functionality described herein. In another example, video interaction engine 108 and video segmentation tool 155 (or portions thereof) are integrated into a general-purpose application executable on a single device. While some embodiments have been described with respect to applications(s), in some embodiments any functionality described herein is additionally or alternatively integrated into an operating system (e.g., as a service), a server (e.g., a remote server), a distributed computing environment (e.g., as a cloud service), etc. These are merely examples, and any suitable distribution of functionality among these or other devices can be implemented within the scope of this disclosure.

[0057] First, through Figure 1A and Figure 1B The configuration illustrated in the diagram begins with a high-level overview of the example workflow. Client device 102 is a desktop, laptop, or mobile device (such as a tablet or smartphone), and video editing application 105 provides one or more user interfaces. In some embodiments, a user accesses video through video editing application 105 and / or otherwise uses video editing application 105 to identify the location where the video is stored (whether it is a local location on client device 102 or a remote location such as storage device 190, etc.). Additionally or alternatively, the user uses the video recording capabilities of client device 102 (or some other device) and / or some application (e.g., Adobe BEHANCE) running at least partially on that device to record video. In some cases, video editing application 105 uploads video (e.g., to an accessible storage device 190 for video file 192) or otherwise transmits the location of the video to server 150, and video segmentation tool 155 receives or accesses the video and performs one or more ingestion functions on the video.

[0058] In some embodiments, video segmentation tool 155 extracts various features from the video (e.g., scripts, linguistic features, speakers, faces, audio classification, visually similar scenes, visual artifacts, video objects or actions, audio events, software log events), and generates and stores representations of the detected features, corresponding feature ranges for the detected features, and / or corresponding confidence levels (e.g., detected feature 194). In one example implementation, based on the detected features, video segmentation tool 155 generates and stores representations of one or more segments of the video (e.g., multiple video segments 196), such as default segments, search segments based on user queries, capture alignment point segments, and / or thumbnail segments. In some cases, one or more segments are generated at multiple granularity levels (e.g., corresponding to different zoom levels). In some embodiments, some segments are generated at ingestion time. In some cases, some or all segments are generated at some other time (e.g., on demand). Thus, the video segmentation tool 155 and / or the video editing application 105 access the video (e.g., one of the video files 192) and generate and store one or more segments of the video (e.g., multiple video segments 196), the semantically meaningful video segments (e.g., video files 192) of the multiple video segments, and / or a representation of the video in any suitable storage location (such as storage device 190, client device 102, server 150, a combination thereof, and / or other locations).

[0059] In one example embodiment, video editing application 105 (e.g., video interaction engine 108) provides one or more user interfaces with one or more interactive elements that allow users to interact with the ingested video, and more specifically with one or more video segments 196 (e.g., semantically meaningful video clips, capture alignment points, thumbnail positions) and / or detected features 194. Figure 1B The illustration shows an example implementation of a video interaction engine 108, which includes a video browsing tool 110 and a video editing tool 130.

[0060] In one example implementation, the video browsing tool 110 provides a finder interface that displays a default segmentation with interactive tiles 112 representing video segments within the default segmentation and features 194 detected in each video segment. In an example implementation with search functionality, the video browsing tool 110 includes a search resegmentation tool 118 that accepts queries (e.g., keywords and / or facets), generates, or otherwise triggers the resegmentation of the default segmentation based on the query (e.g., by...). Figure 1AThe search segmentation component 170 is generated, and the interactive tiles 112 are updated to represent the search segments. The video browsing tool 110 provides a selected clip panel 114 where the user can add video clips to the selection.

[0061] The video editing tool 130 provides an editor interface with a composite editing timeline tool 132, which presents a composite video timeline representing the composite video formed from video segments selected in the finder interface. In some embodiments, a selection box selection and snap alignment tool 136 detects dragging operations along the composite timeline and snaps the selection boundaries to snap alignment points defined by snap alignment point segments and / or the current zoom level. In some cases, the composite editing timeline tool 132 includes a feature visualization tool 134, which presents a visualization of detected features 194 and corresponding feature ranges to aid in trimming. Additionally or alternatively, a thumbnail preview tool 138 represents one or more thumbnails showing representative video frames at positions on the composite timeline identified by thumbnail segments and / or the current zoom level. Thus, the video editing tool 130 enables the user to navigate the video and select semantically meaningful video segments.

[0062] Depending on the implementation, the video editing tool 130 and / or the video interaction engine 108 may perform any number and type of operations on the selected video segment. As a non-limiting example, the selected video segment may be played back, deleted, edited in some other way (e.g., by rearranging, cropping, applying transitions or effects, adjusting colors, adding titles or graphics), exported, and / or otherwise manipulated. Therefore, in various embodiments, the video interaction engine 108 provides interface functionality that allows users to select, navigate, play, and / or edit videos based on interactions with semantically meaningful video segments and their detected features.

[0063] Example video segmentation technology

[0064] return Figure 1A In some embodiments, the video segmentation tool 155 generates representations of one or more segments of the video. Typically, one or more segments are generated at any appropriate time, such as when ingesting or initially processing the video (e.g., default segmentation), when a query is received (e.g., searching for segments), when the video timeline is displayed, when the user interface is activated, and / or at some other time.

[0065] exist Figure 1AIn the example illustrated, video ingestion tool 160 ingests video (e.g., a video file, a portion of a video file, a video represented by a project file, or otherwise identified). In some embodiments, ingesting video includes extracting one or more features from the video and / or generating one or more segments of the video, which identify corresponding semantically meaningful video fragments and / or fragment boundaries. Figure 1A In the implementation illustrated, the video ingestion tool 160 includes multiple feature extraction components 162 and a default segmentation component 164. In this implementation, the multiple feature extraction components 162 detect one or more features from the video, and the default segmentation component 164 triggers the video segmentation component 180 to generate default video segments during video ingestion.

[0066] At a high level, video ingestion tool 160 (e.g., multiple feature extraction components 162) uses one or more machine learning models, natural language processing, digital signal processing, and / or other techniques to detect, extract, or otherwise determine various features from the video (e.g., scripts, linguistic features, speakers, faces, audio classification, visually similar scenes, visual artifacts, video objects or actions, audio events, software log events). In some embodiments, multiple feature extraction components 162 include one or more machine learning models for each of a plurality of feature categories to be detected. Thus, video ingestion tool 160 and / or the corresponding multiple feature extraction components 162 extract, generate, and / or store representations of detected features (e.g., facets), the corresponding feature ranges where the detected features exist, and / or the corresponding confidence levels for each category.

[0067] In some embodiments, one or more feature categories (e.g., speaker, face, audio classification, visually similar scenes, etc.) have their own feature orbitals representing instances of features detected within the feature category (e.g., facets such as unique faces or speaker). As a non-limiting example, for each feature category, the representation of a detected feature (e.g., detected feature 194) includes a list, array, or other representation of each instance of a facet detected within that feature category (e.g., detected faces) (e.g., each unique face). In one example implementation, each instance of a detected facet is represented using the feature range in which the instance was detected (e.g., start and stop timestamps for each instance), a unique value identifying the facet to which the instance belongs (e.g., a unique value for each unique face, speaker, visual scene, etc.), a corresponding confidence level quantifying the prediction confidence or probability, and / or a representation of other characteristics.

[0068] In some embodiments, feature extraction components 162 extract script and / or linguistic features from an audio track associated with the video. In one example implementation, any known speech-to-text algorithm is applied to the audio track to generate a script of speech, detect speech segments (e.g., corresponding to words, sentences, utterances of continuous speech separated by audio gaps, etc.), and detect non-speech segments (e.g., pauses, silences, or non-speech audio). In some embodiments, speech activity detection (e.g., applied to the audio track, to detected non-speech segments) is applied to detect and / or classify audio track segments with non-word human sounds (e.g., laughter, audible breathing, etc.). In some cases, the script and / or detected script segments are associated with the timeline of the video, and the script segments are associated with corresponding time ranges. In some embodiments, any known topic segmentation technique (semantic analysis, natural language processing, applied language models) is used to segment or otherwise identify video segments that may contain similar topics, and the detected speech segments are associated with a score indicating how likely the speech segment is to end a topic segment.

[0069] In some embodiments, the feature extraction component 162 includes one or more machine learning models that detect unique speakers from audio tracks associated with a video. In one example implementation, any known speech recognition, speaker identification, or speaker segmentation clustering techniques are applied to detect unique voiceprints (e.g., within a single video, or across a collection of videos) and to segment or otherwise identify portions of the audio tracks based on speaker identity. Example techniques used in speech recognition, speaker identification, or speaker segmentation clustering employ frequency estimation, pattern matching, vector quantization, decision trees, hidden Markov models, Gaussian mixture models, neural networks, and / or other techniques. In some embodiments, the feature extraction components(s) 162 apply speaker segmentation clustering techniques, such as those described in ActiveSpeakers in Context by Juan Leon Alcazar, Fabian Caba, Long Mai, Federico Perazzi, Joon-Young Lee, Pablo Arbelaez, and Bernard Ghanem (Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE / CVF Conference Proceedings on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12465-12474). As an adjunct or alternative to using audio signatures to detect speakers, in some embodiments, one or more machine learning models are used to determine which detected face is speaking by detecting mouth movements on detected faces. In one example implementation, each instance of a speaker detected in a video is associated with a corresponding temporal range of the video in which the speaker was detected, a corresponding confidence level quantifying the prediction confidence or probability, and / or a thumbnail of the detected face of the speaker. Additionally or alternatively, detected speech segments (e.g., words, phrases, sentences) and / or other script features are associated with the corresponding representations of the detected speakers to generate a script for segmentation and clustering.

[0070] In some embodiments, the feature extraction component 162 includes one or more machine learning models for detecting unique faces from video frames of a video. In one example implementation, any known face detection technique (e.g., RetinaFace) is applied to detect unique faces in each video frame and / or across time. For example, each video frame is processed by (e.g., using one or more neural networks) segmenting each face from the background, aligning each face, detecting the location of facial landmarks (e.g., eyes, nose, mouth), and generating (e.g., vector) representations of the detected facial landmarks. In some embodiments, detected faces from different frames (e.g., within a single video, or across a set of videos) and having similar representations (e.g., separated by a certain threshold distance, clustered based on one or more clustering algorithms) are determined to belong to the same identity. In one example implementation, each instance of a detected face is associated with a corresponding time range across the video frames containing the detected face and / or a corresponding confidence level that quantifies the predicted confidence or probability.

[0071] In some embodiments, the feature extraction component 162 includes one or more machine learning models for extracting audio classifications from audio tracks associated with a video. Any known sound recognition technique is applied to detect any number of audio classifications (e.g., music, speech, others). In one example implementation, each frame of audio data from an audio track is encoded into a vector representation (e.g., using linear predictive encoding and decoding) and classified by one or more neural networks. In some embodiments, the audio timeline is categorized into the detected audio classifications (e.g., music, speech, or others). In one example implementation, consecutive audio frames with the same classification are grouped together and associated with a corresponding (e.g., average) confidence level across a corresponding time range and / or a quantized prediction confidence or probability.

[0072] In some embodiments, the feature extraction component 162 includes one or more machine learning models that detect visually similar scenes from video frames of a video. In one example implementation, each video frame is processed (e.g., via one or more neural networks) to extract a corresponding (e.g., vector) representation of the visual features in the video frame, and the representations of different video frames are clustered temporally (e.g., fixed or variable) into multiple visual scenes using any suitable clustering algorithm (e.g., k-means clustering). In some embodiments, each visual scene is associated with a corresponding temporal range spanning video frames within the visual scene. In some cases, each scene transition is assigned a transition confidence, for example, by computing a distance metric comparing the representations of visual features in video frames surrounding the transition.

[0073] In some embodiments, the feature extraction component 162 includes one or more machine learning models for detecting visual artifacts in video frames of a video. Any known visual detection technique is applied to identify one or more categories of visual artifacts from the video frames. In one example implementation, one or more neural network classifiers detect corresponding categories of visual artifacts, such as unstable camera motion across video frames, camera occlusion in a given video frame, blurring in a given video frame, compression artifacts in a given video frame, lack of motion (e.g., empty video frames, no visual change across video frames), etc. In some embodiments, each instance of a detected visual artifact is associated with a corresponding temporal range across the video frames in which the visual artifact was detected and / or a corresponding confidence level that quantifies the prediction confidence or probability.

[0074] In some embodiments, the feature extraction component 162 includes one or more machine learning models that detect objects or actions from video frames of a video. Any known object or action recognition technique is applied to visually extract one or more categories of objects or actions from one or more video frames. In one example implementation, one or more neural network classifiers detect the presence of any number of object categories (e.g., hundreds, thousands, etc.) in each video frame. Additionally or alternatively, one or more neural network classifiers detect the presence of any number of action categories in the video frame sequence (e.g., low-level movements such as standing, sitting, walking, and talking; high-level events such as eating, playing, and dancing; and / or others). In some embodiments, each instance of a detected object or action category is associated with a corresponding time range across the video frames in which the object or action was detected, a corresponding confidence level quantifying the prediction confidence or probability, and / or one or more searchable keywords (e.g., tags) representing the category.

[0075] In some embodiments, the feature extraction component 162 includes one or more machine learning models for detecting audio events from an audio track associated with the video. Any known sound recognition technique can be applied to detect any number of audio event classifications (e.g., alarms, laughter, ringing, applause, coughing, buzzing, horns, barking, gunshots, sirens, etc.). In one example implementation, each frame of audio data from the audio track is encoded into a vector representation (e.g., using linear predictive codec) and classified by one or more neural networks. In one example implementation, consecutive audio frames with the same classification are grouped together and associated with a corresponding time range across the audio frames, a corresponding confidence level quantifying the prediction confidence or probability, and / or one or more searchable keywords (e.g., tags) representing the classification.

[0076] In some embodiments, feature extraction components 162 extract log events represented in one or more time logs (such as software usage logs) associated with the video. Various implementations involve different types of time logs and / or log events. For example, in one implementation involving a screen capture or screenshot video of a tutorial for creative software such as Adobe Photoshop or Adobe Fresho, a software usage log generated by the creative software at the time of screen capture or screenshotting is read to identify the time of detected log events, such as tool events (e.g., use of a specific software tool, such as selecting a brush, creating a layer, etc.). In an example game implementation, the software usage log is read to identify the time of detected software log events such as leveling up or defeating an enemy. While the foregoing examples involve time logs with log events derived from video frames, this is not necessarily the case. For example, in an implementation of live chat or a chat stream associated with live streaming video, the corresponding user chat log or session is read to identify the time of events such as chat messages on a specific topic. In a sample video streaming implementation (whether live streaming or watching archived video), usage logs representing how (multiple) users) watch the video are read to identify the timing of detected interaction events (such as navigation events, e.g., play, pause, skip)). Typically, any type of time log and / or metadata can be read to identify log events and their corresponding times. In one sample implementation, each instance of the extracted log event is associated with a corresponding time range spanning the video segment where the log event occurred and / or one or more searchable keywords (e.g., tags) representing the log event (e.g., for tool events, software tool names, or actions).

[0077] exist Figure 1A In the implementation illustrated, video segmentation component 180 executes a segmentation routine that accepts various input parameters, such as identifiers of specified feature tracks (e.g., predetermined, user-selected, and / or otherwise), user queries (e.g., keywords and / or selected facets), the target minimum or maximum length of the video segment (which in some cases depends on the zoom level), the range of the video to be segmented, etc. Video segmentation component 180 uses the input parameters to construct a graph with nodes representing candidate cut points, edges representing candidate segments, and edge weights representing cut costs. Video segmentation component 180 computes video segments by solving a shortest path problem to find the path (segment) that minimizes the sum of edge weights (cut costs).

[0078] In some embodiments, one or more aspects of the segmentation routine depend on the type of segmentation (e.g., default, search, capture alignment, thumbnail). For example, and as described in more detail below, depending on the type of segmentation, candidate cut points may be determined differently, different edges may be used, different cut costs may be used for edge weights, different target minimum or maximum video segment lengths may be used, scaling levels may affect (or not affect) target minimum or maximum video segment lengths, user queries may affect (or not affect) segmentation, and / or other relevances are possible. Therefore, in some embodiments, the default segmentation component 164, the search segmentation component 170, the capture alignment segmentation component 172, and / or the thumbnail segmentation component 174 trigger the video segmentation component 180 to generate corresponding video segments using the corresponding candidate cut points, edges, edge weights, target minimum or maximum video segment lengths, query applicability, and / or other aspects. Additionally or alternatively, separate segmentation routines are executed. In the example implementation, the default segmentation is determined at ingestion time. In another example implementation, search, snap alignment points, and / or thumbnail segments are determined on demand (e.g., when a query is received, when the editor interface is loaded, or when zooming into the video timeline). However, these are just examples, and other implementations can determine any type of segmentation at any appropriate time.

[0079] In one example embodiment, the video segmentation component 180 outputs a representation of the complete set of non-overlapping (i.e., non-intersecting) video segments (i.e., covering the entire and / or specified range of the video), and / or timestamps or some other representation of the video segment boundaries (e.g., cut points). In some implementations of determining search segments, additionally or alternatively, the output includes a representation of whether each video segment in the search segment is a match for a specific query (e.g., whether the video segment is “on” or “off” with respect to the query), a representation of which(s) features and / or values ​​of the matching segments match the query, and / or a representation of at which(s) the query was matched within each matching segment.

[0080] The following embodiments relate to the implementation of a video segmentation component 180, which has one or more common aspects across different types of segmentation. For illustrative purposes, example operations of this implementation of the video segmentation component 180 are described with respect to default segmentation and search segmentation. In this implementation, the video segmentation component 180 includes a candidate boundary selection component 182, a graph construction component 184, a shear cost calculation component 186, and a path optimization component 188. At a high level, the candidate boundary selection component 182 identifies candidate boundaries from the boundaries of feature ranges in a specified feature track. The graph construction component 184 constructs a graph with paths representing different segmentation options, the shear cost calculation component 186 calculates edge weights for edges between nodes in the graph, and the path optimization component 188 solves the shortest path problem along the graph to compute the optimal segmentation.

[0081] Candidate boundary selection component 182 identifies candidate boundaries from the boundaries (feature boundaries) of feature ranges within a specified range of the video and on a specified feature track. In an example default segmentation, the feature track includes detected sentences (e.g., from a script), detected faces, and detected visual scenes, and the specified range of the video is the entire video. In an example search segmentation, the feature track includes the same feature track as the default segmentation, and the specified range of the video is each video segment in the default segmentation (i.e., the search segmentation operates on each video segment in the default segmentation, thus re-segmenting the default segmentation). In another example search segmentation, the feature track includes the same or different feature track as the default segmentation, and the specified range of the video is the entire video (i.e., the search segmentation creates new segments independent of the default segmentation).

[0082] In one example implementation, the candidate boundary selection component 182 identifies instances of detected features that overlap with a specified range of the video from a specified feature track, adjusts the feature boundaries to capture neighboring feature boundaries aligned with the preferred feature, and identifies candidate boundaries from the remaining feature boundaries. Figures 2A-2F This is an illustration of an example technique for calculating candidate boundaries from detected feature boundaries. Figures 2A-2F The illustration shows a portion of the example-specified feature tracks, sentence track 210 and face track 220 (showing two facets, face 1 and face 2).

[0083] Figure 2A This is an illustration of an example technique for identifying instances of detected features that overlap with a specified range of the video to be segmented from a specified feature track. In this example, all detected features from sentence track 210 and face track 220 that overlap with a specified range between start and end times are identified (overlapping features are shaded). In some embodiments where feature tracks are represented by a list of detected instances and corresponding ranges, identifying overlapping features includes iterating through the listed ranges and generating a representation of those ranges that at least partially overlap with a specified range of the video to be segmented.

[0084] Figures 2B-2CThis is an illustration of an example technique for adjusting feature boundaries to capture neighboring feature boundaries aligned to a priority feature. In some embodiments, feature tracks are prioritized (e.g., in a priority list) based on which feature categories are considered most important or determined to have the highest quality data. In the example implementation, the priority list of feature tracks includes script features (e.g., sentences), then visual scenes, and then faces. The candidate boundary selection component 182 iterates over the specified feature tracks in priority order, starting with the most important feature track. For each priority feature boundary of overlapping features from the priority feature tracks, the candidate boundary selection component 182 merges those feature boundaries from other feature tracks that are within a threshold merging distance (e.g., 500 ms) of the priority feature boundary into the priority feature boundary. Additionally or alternatively, the candidate boundary selection component 182 truncates and / or discards feature boundaries that fall outside a specified range of the video to be segmented. Figure 2B In the context of the sentence track 210, feature boundaries 224 and 226 from the face track 220 are within the threshold merging distance of the priority feature boundary 222, therefore feature boundaries 224 and 226 are captured and aligned to the priority feature boundary 222, as shown below. Figure 2C As shown in the diagram.

[0085] In some embodiments, the feature boundaries of features detected within a feature track are merged into neighboring feature boundaries outside the detected features that are within a threshold merging distance, for example, to create a span that includes two detected features. Additionally or alternatively, the feature boundaries of features detected within a feature track are merged into neighboring feature boundaries predicted at a higher confidence level and located within a threshold merging distance. For example, in Figure 2B In the context of face orbit 220, feature boundary 232 (for face 2) is within a threshold merging distance of feature boundary 234 (for face 1) from face orbit 220; therefore, feature boundary 232 is snapped and aligned to feature boundary 234, as shown below. Figure 2C As shown in the diagram.

[0086] Figure 2D This is an illustration of an example technique for identifying candidate boundaries for default segmentation from remaining feature boundaries. In some embodiments, the candidate boundary selection component 182 evaluates the remaining feature boundaries after adjustment, determines whether the distance between two consecutive feature boundaries is greater than a threshold gap duration (e.g., 500 milliseconds), and if so, inserts one or more (e.g., equally spaced) boundaries into the gap such that the distance between two consecutive feature boundaries is less than the threshold gap duration. In some embodiments, to improve performance, the candidate boundary selection component 182 divides a specified range of the video into a number of equal-length intervals (…). Figure 2DThe three remaining feature boundaries are selected, and for each interval, the top N boundaries in that interval are selected based on a certain scoring function. In an example implementation, a shear cost is calculated for each remaining feature boundary (as described in more detail below with respect to the shear cost calculation component 186), and the candidate boundary selection component 182 selects the top N boundaries with the lowest boundary shear cost.

[0087] Figures 2E-2F This is an illustration of an example technique for identifying candidate boundaries for search segments from remaining feature boundaries. In an example embodiment with search, the user enters a query including one or more keywords and / or selected facets (described in more detail below). Regardless of whether exact matching and / or fuzzy matching are used, the candidate boundary selection component 182 searches for detected features with associated text or values ​​(e.g., script, object or action tags, audio event tags, log event tags, etc.) that match the keywords(s). In the example facet search, the user selects a specific face, speaker, audio category (e.g., music, speech, others), visual scene, visual artifacts, and / or other facets, and the candidate boundary selection component 182 searches for detected instances of the selected facet(s). Figure 2F The illustration shows an example query (“hello”) with matching features 212 and 214 from sentence track 210. In one example implementation, candidate boundary selection component 182 selects candidate boundaries as those remaining feature boundaries that are within a certain threshold distance (e.g., 5 seconds) of the boundaries of the matching features (query matching boundaries).

[0088] Return to Figure 1A The graph building component 184 constructs a graph with paths representing different candidate segments. Figures 3A-3C This is an illustration of an example technique for constructing a graph with paths representing different candidate segments. In one example implementation, candidate boundaries are represented by and to indicate candidate segments (e.g., Figure 3B The edges of video clip 365a in the video (e.g., Figure 3B The nodes connected by edge 365a in the equation (e.g., Figure 3A Node 350 in the graph is used to represent this. Different candidate segments are represented by different paths in the graph from the start time to the end time. For example, in Figure 3C In the graph, edges 362a, 365a, and 368a form paths representing segments 362b, 365b, and 368b. In some embodiments, to force the generation of non-overlapping segments (e.g., paths that move forward in time), the graph is constructed as a directed graph having edges that connect only nodes that move forward in time, or having infinite edge weights that move backward in time.

[0089] In some embodiments, Figure 1AThe shear cost calculation component 186 calculates edge weights for edges in the graph. In one example embodiment, the penalty or shear cost of a candidate segment is quantified (e.g., over some normalized range, such as 0 to 1 or -1 to +1) for the edge weights representing edges representing candidate segments. As a non-limiting example, the shear cost calculation component 186 determines the edge weights between two nodes as the (e.g., normalized) sum of the boundary shear cost and the interval shear cost for the candidate segment. In some embodiments, the boundary shear cost penalizes the candidate boundary for being at a “bad” shear point (e.g., within a feature detected from another feature track), and the interval shear cost penalizes the candidate segment based on negative characteristics of the candidate segment’s span (e.g., having a length beyond the target minimum or maximum length, incoherence of features overlapping with other feature tracks, overlap of both query “on” and “off” features). In some embodiments, the boundary shear cost for a candidate segment is the boundary shear cost for the leading boundary of the candidate segment (e.g., when candidate segmentation is completed, candidate segments in the candidate segment are therefore adjacent, and the trailing boundary is therefore included in the next segment).

[0090] In a sample default segmentation, the shearing cost calculation component 186 penalizes shearing within the detected feature range (e.g., sentence, facial appearance, visual scene), determines the boundary shearing cost differently relative to different feature tracks, and / or calculates the overall boundary shearing cost as a combination (e.g., sum, weighted sum) of the individual contributions from each feature track. As a non-limiting example where the specified features include script features, visual scene, and face, the boundary shearing cost for the boundary is calculated as follows:

[0091] Boundary clipping cost = (3.0 * script boundary clipping cost + 2.0 * visual scene boundary clipping cost + face boundary clipping cost) / 6.0. (Equation 1)

[0092] Figures 4A-4C This is an illustration of an example technique for calculating boundary shearing costs based on a visual scene. Figures 4A-4C An example visual scene track 405 is illustrated, where different visual scenes are represented by different patterns. In some embodiments, a good candidate boundary is (1) close to a visual scene boundary, and / or (2) that the visual scene boundary has a high transition confidence. In an example implementation considering candidate boundaries with respect to visual scene track 405, shear cost calculation component 186 identifies the nearest visual scene boundary from visual scene track 405. Figure 4A ), and calculate the distance to the nearest boundary relative to the duration of the visual scene 410. Figure 4B ).exist Figure 4BIn the example illustrated, the distance to the nearest boundary (1000 ms) relative to the duration (5000 ms) of the visual scene 410 is 1000 / 5000 = 0.2. In some embodiments, the shear cost calculation component 186 retrieves the transition confidence of the nearest boundary (in Figure 4B In the example, it is 0.9), and the boundary shear cost is calculated, for example, based on the transition confidence of the nearest boundary, as follows:

[0093] Visual scene boundary clipping cost = -transformation confidence * (1.0 - 2 * relative distance to the boundary) (Equation 2)

[0094] exist Figure 4B In the example shown in the figure, the visual scene boundary clipping cost = 0.9*(1.0-2*0.2) = -0.54.

[0095] Figure 4C The illustration depicts a critical situation where scene 420 has candidate boundaries, with the nearest visual scene boundary located at the endpoint of a specified range of the video. In some embodiments, if the transition confidence of the nearest visual scene boundary is undefined or empty, the clipping cost calculation component 186 calculates the boundary clipping cost, for example, based on the distance to the nearest boundary relative to the duration containing the visual scene, as follows:

[0096] Visual scene boundary clipping cost = -(1.0 - 2.0 * (relative distance to the nearest visual scene boundary)) (Equation 3)

[0097] When the relative distance to the nearest visual scene boundary is zero (the boundary is at the endpoint of the specified range), Equation 3 solves to -1.0. When the relative distance to the nearest visual scene boundary is 0.5 (the candidate boundary is in the middle of the candidate fragment), Equation 3 solves to +1.0.

[0098] Figure 5 This is an illustration of an example technique for calculating boundary shearing costs based on script features of segmentation clustering. Figure 5 The illustration shows example script tracks 505 with different script features for segmentation clustering based on detected speakers (different speakers are represented by different patterns). In one example implementation, the cut-off cost calculation component 186 penalizes candidate boundaries located within speech segments (e.g., word segments, phrase segments), candidate boundaries located during speaker appearance but between sentences, and / or candidate boundaries for unlikely ending topic segments at the end of speech segments. In one embodiment, the cut-off cost calculation component 186 determines whether scripts and / or scripts for segmentation clustering are available. If script features are unavailable, there is no contribution to the cut-off cost based on script features. Figure 5In the example implementation illustrated, where the cutoff cost is normalized between -1 (good) and +1 (bad), candidate boundaries 510 located in the middle of a word are scored with a large cutoff cost (e.g., infinite), candidate boundaries 520 located in the middle of a phrase but between words are scored with a medium cutoff cost (e.g., 0), candidate boundaries 530 located in the middle of a speaker but between sentences are scored with a cutoff cost equal to or proportional to the predicted probability of the topic segment ending in the previous speech segment (e.g., -0.2), and / or candidate boundaries 540 located at speaker changes are scored with a low cutoff cost (e.g., -1.0). Additionally or alternatively, candidate boundaries at the end of a speech segment (e.g., candidate boundaries 530, 540) are scored with a score equal to or proportional to the predicted probability of the topic segment ending in the previous speech segment (e.g., ...). Figure 5 The score is based on the "-segment end score" of the preceding sentence.

[0099] Figure 6 This is an illustration of an example technique for calculating boundary shearing costs based on detected faces. Figure 6 The illustration shows an example facial orbital 220 and the detected feature range for each detected facial appearance. In one example implementation, a shear cost calculation component 186 penalizes candidate boundaries when they fall within the feature range of a detected facial appearance. As a non-limiting example, candidate boundaries falling within the detected facial appearances (e.g., boundaries 610 and 620) are scored with a high shear cost (e.g., +1.0), and candidate boundaries not falling within the detected facial appearances (e.g., boundaries 610 and 620) (e.g., boundary 640) are scored with a low shear cost (e.g., -1.0).

[0100] In a sample search segment, the "good" boundary will transition between features that match the query (e.g., the query "on" feature) and features that do not match the query (e.g., the query "off" feature). Thus, in some embodiments, the shear cost calculation component 186 penalizes candidate boundaries located within matching features (e.g., the query "on" feature) and / or candidate boundaries located away from the query on / off transition. Figures 7A-7D This is an illustration of an example technique for calculating boundary shearing costs based on queries. Figures 7A-7D The illustration shows example matching (query "on") features 710 and 720. In one example implementation, if the candidate boundary is within the matching (query "on") feature, the shear cost calculation component 186 is scored with a high shear cost (e.g., +1.0). Figure 7A In the process, candidate boundary 790 is within matching feature 710, so shear cost calculation component 186 scores candidate boundary 790 as having high shear cost (e.g., +1.0).

[0101] In some embodiments, if the candidate boundary is not located inside a matching feature, the shear cost calculation component 186 considers the two nearest and / or adjacent matching features and their lengths. For example, in Figure 7B In this context, candidate boundary 794 lies outside matching features 710 and 720. In one example implementation, the shear cost calculation component 186 considers windows on each side of candidate boundary 794 (e.g., windows whose size corresponds to half the average length of adjacent query "open" features). Figure 7B In the example illustrated, the shear cost calculation component 186 determines the size of windows 730a and 730b to be 0.5 * (5000 ms + 3000 ms) / 2 = 2000 ms, a 2000 ms window on either side of the candidate boundary 794. The cost calculation component 186 calculates the query "on" time in windows 730a and 730b. Figure 7C In the above, window 730a overlaps with the matching (query "on") feature 710 for 1800 milliseconds (query "on" time = 1800 milliseconds), and window 730b does not overlap with the matching (query "on") feature 720 (query "on" time = 0 milliseconds). Finally, for example, cost calculation component 186 calculates the shearing cost based on the amount of time windows 730a and 730b overlap with the matching (query "on") feature as follows:

[0102]

[0103] exist Figure 7C In the example shown in the figure, Equation 4 is solved as 1.0 - 2 * 1800 / 2000 = -0.8. Figure 7D The illustration shows another example where candidate boundary 796 is not within the query "on" feature. In this example, candidate boundary 796 is centered on the query "off" feature, windows 740a and 740b overlap with the matching (query "on") feature for 500 milliseconds each, and Equation 4 is solved as 1.0 - 2*0 / 2000 = 1.0.

[0104] Turning now to the example interval shearing cost, in some embodiments, the shearing cost calculation component 186 assigns an interval shearing cost that penalizes candidate segments for having a length exceeding the target minimum or maximum length, incoherence of overlapping features with other feature tracks, both overlapping query "on" and "off" features, and / or other characteristics. In one example implementation, the shearing cost calculation component 186 (e.g., based on target length, incoherence, partial query matching, etc.) calculates the overall interval shearing cost as a combination (e.g., a sum, a weighted sum) of the individual contributions from individual items.

[0105] In some embodiments, the cut cost calculation component 186 calculates interval cut costs for candidate segments with lengths exceeding the target minimum or maximum length. In some embodiments, the target minimum or maximum length is fixed for a particular type of segment (e.g., a target segment length ranging from 15 seconds to 5 times the video duration), proportional to or otherwise dependent on the input scaling level (e.g., the scaling level for the composition timeline in the editor interface, discussed in more detail below), displayed by an interactive control element (e.g., a slider or button allowing the user to set or adjust the target minimum or maximum length), mapped to discrete or continuous values, and so on. In one example implementation, the search segment has fixed target minimum and maximum segment lengths unaffected by the scaling level. In another example implementation, the capture alignment point segment has target minimum and maximum segment lengths mapped to the scaling level, and the target minimum and maximum segment lengths decrease as the user zooms in (e.g., to the composition timeline in the editor interface), resulting in more capture alignment points for smaller video segments.

[0106] Depending on the specified and / or determined target minimum and maximum segment lengths, the shear cost calculation component 186 penalizes candidate segments outside the target length. In some embodiments, the shear cost calculation component 186 uses hard constraints to assign infinite shear costs to candidate segments outside the target length. In some embodiments, the shear cost calculation component 186 uses soft constraints to assign large shear costs to candidate segments outside the target length (e.g., 10, 100, 1000, etc.).

[0107] In some embodiments, the cut cost calculation component 186 calculates the interval cut cost for candidate segments based on the incoherence of overlapping features from other feature tracks. In some cases, a “good” video segment contains coherent or similar content about each feature track. Thus, in one example implementation, the cut cost calculation component 186 penalizes candidate segments that lack coherence in overlapping regions of another feature track (e.g., feature changes detected in the overlapping region). In some embodiments, the cut cost calculation component 186 calculates the interval cut cost based on incoherence for each feature track, calculates the interval incoherence cut cost differently for different feature tracks, and / or calculates the overall interval incoherence cut cost by combining contributions from each feature track (e.g., summation, weighted summation, etc.).

[0108] Figure 8This is an illustration of an example technique for calculating interval shearing cost based on detected feature inconsistencies. In one example implementation, shearing cost calculation component 186 calculates the interval inconsistency shearing cost for a candidate segment based on the number of feature transitions detected in a region of a feature track that overlaps with the candidate segment. For example, in Figure 8 In the candidate segment 810, two transitions (feature boundaries 855 and 865) overlap with facial track 220, so candidate segment 810 will contain sub-segment 850 (with face 1 and face 2), sub-segment 860 (without face), and sub-segment 870 (with face 2). In contrast, candidate segment 820 overlaps with zero transitions in facial track 805, so candidate segment 820 will only contain sub-segment 880 (with face 1 and face 2). In an example implementation, the shear cost calculation component 186 counts the number of feature transitions detected in the region of the feature track that overlaps with each candidate segment, normalizes the number of transitions (e.g., by the maximum number of transitions with respect to that feature track), and assigns the normalized number of transitions as the interval discontinuity shear cost for the candidate segment.

[0109] In some embodiments, some transitions (e.g., feature boundaries) are associated with a measure of the strength of the transition (e.g., a segment end score quantifying the probability that a previous speech segment ends the topic segment, a confidence level that the speech segment was spoken by a new speaker, or a measure of the similarity of frames in a visual scene). Thus, in some cases, the clipping cost calculation component 186 penalizes candidate segments based on a count of overlapping feature transitions weighted by the strength of the transition. This can be used to reduce incoherence clipping costs based on the incoherence of detected features, such as in cases where sentences change but the topics are similar, or where visual scenes change but appear similar.

[0110] In some embodiments, the cut cost calculation component 186 calculates the interval cut cost for candidate segments based on the query. In some cases, a "good" video segment is either completely on or completely off with respect to the query (e.g., to encourage clean search results). For example, in some implementations, if a user queries "elephant," ideally the returned segments would always contain an elephant or not contain one at all. Thus, in one example implementation, the cut cost calculation component 186 only penalizes candidate segments that are partially on and partially off (segments that sometimes contain an elephant). Figures 9A-9B This is an illustration of an example technique for calculating interval shearing costs based on queries. Figure 9A In the text, candidate fragments 910 and 920 are partially on and partially off, respectively, and... Figure 9BIn this example, candidate segments 930 and 940 are either fully on or fully off. In one example implementation, the shear cost calculation component 186 counts and normalizes the number of transitions between on and off segments, as described above. In another example, the shear cost calculation component 186 calculates the percentage of overlapping query on time in candidate segments relative to the total time, the percentage of overlapping query off time in candidate segments relative to the total time, and assigns a lower shear cost the closer the percentage is to zero or one.

[0111] In conclusion and return Figure 1A The shearing cost calculation component 186 calculates the shearing cost for different candidate segments based on boundary and / or interval contributions, and assigns the shearing cost for a candidate segment as the edge weight representing the edge of that candidate segment. Then, for example using dynamic programming, the path optimization component 188 solves the shortest path problem along the graph to compute the optimal segmentation. In an example implementation, candidate paths are selected to produce complete and non-overlapping segments. For each candidate path, the path optimization component 188 sums the edge weights and selects the path with the smallest sum as the optimal path, which represents the optimal segmentation.

[0112] The foregoing discussion relates to an example implementation of video segmentation component 180, which is triggered by default segmentation component 164 to compute example default segments and by search segmentation component 170 to compute example search segments. Another example video segment is a capture alignment point segment that identifies the location of a selected capture alignment point for a video. As explained in more detail below, an example use of capture alignment point segments is in a user interface having a video timeline representing the video (e.g., a composite timeline representing a selected video segment in an editor interface), where the capture alignment point identified by the capture alignment point segment is illustrated on the timeline and / or used to capture the selection of the video segment for alignment as the user drags along the corresponding portion of the timeline or script. In various embodiments, the capture alignment point segment is computed at any suitable time (e.g., when the video timeline is displayed, the editor interface is activated, a video segment to be represented by the composite timeline is identified, and / or at some other time). Figure 1A and Figure 1B In the example embodiment illustrated, the video interaction engine 108 of the video editing application 105 (e.g., video editing tool 130) communicates with the capture alignment point segmentation component 172 to trigger the video segmentation component 180 to calculate the capture alignment point segmentation.

[0113] In one example implementation of capture alignment point segmentation, video segmentation component 180 executes a segmentation routine that accepts various input parameters, such as specified feature tracks (e.g., predetermined, user-selected, etc.), target minimum or maximum length of the video segment (which in some cases depends on the zoom level, an interactive control element presented to the user, etc.), the range of the video to be segmented (e.g., each video segment designated for editing, represented by the composite timeline, etc.), etc. In some embodiments, video segmentation component 180 computes individual capture alignment point segments for each video segment represented by the composite timeline or otherwise designated for editing.

[0114] In one example implementation of capture alignment point segmentation, video segmentation component 180 uses any of the techniques described herein to perform segmentation routines. For example, candidate boundary selection component 182 of video segmentation component 180 identifies candidate boundaries as candidate capture alignment points from the boundaries of feature ranges in a specified feature track. In one example embodiment, if no detected feature or feature range is available, candidate boundary selection component 182 returns candidate capture alignment points with regular spacing. If a detected feature and feature range are available, candidate boundary selection component 182 considers whether a script feature is available. If a script feature is unavailable, candidate boundary selection component 182 calculates candidate capture alignment points with regular spacing (e.g., approximately 500 milliseconds apart) and then captures and aligns those points to the nearest feature boundary (e.g., 250 milliseconds) within a specified feature track that is within a capture alignment threshold.

[0115] In example embodiments where script features are available, candidate boundary selection component 182 iteratively adds script feature boundaries (e.g., word boundaries) as candidate capture alignment points through feature boundaries for script features (e.g., words). Additionally or alternatively, when the gap between consecutive script feature boundaries (e.g., representing word duration and / or the gap between words) is greater than a certain threshold (e.g., 500 milliseconds), candidate boundary selection component 182 adds regularly spaced candidate capture alignment points (e.g., spaced approximately 500 milliseconds) into the gap. In some embodiments, candidate boundary selection component 182 captures and aligns the added points to the nearest non-script feature boundary in a specified feature track that is within a capture alignment threshold (e.g., 250 milliseconds). These are just a few ways to designate candidate boundaries as candidate capture alignment points, and additionally or alternatively, any other techniques for identifying candidate capture alignment points, including other techniques described herein for identifying candidate boundaries, may be implemented.

[0116] In some embodiments, the graph construction component 184 of the video segmentation component 180 constructs a graph having nodes representing candidate capture alignment points, edges representing candidate intervals (e.g., candidate segments) between capture alignment points, and edge weights calculated by the cut cost calculation component 186 of the video segmentation component 180. In one example implementation, the cut cost calculation component 186 assigns cut costs to candidate segments that encourage capture alignment at “good” points and / or prevent capture alignment at “bad” points. As a non-limiting example, the cut cost calculation component 186 determines the edge weight between two nodes as the sum of the boundary cut cost (e.g., as described above) and the interval cut cost (e.g., normalized) for the candidate segment. Regarding interval cut costs, in some cases, capture alignment points that are too close together may not be helpful. Thus, in one example embodiment, the target minimum length between capture alignment points (e.g., represented by the target minimum video segment length) is determined based on a minimum pixel interval that, in some cases, depends on the zoom level of the viewed video timeline. For example, a specified minimum pixel interval (e.g., 15 pixels) is mapped to a corresponding duration on the timeline (e.g., based on the activity zoom level), and this duration is used as the target minimum interval between capture alignment points. In some cases, the target minimum interval is used as a hard constraint (e.g., candidate segments shorter than the minimum interval are assigned an infinite interval cut cost), a soft constraint (e.g., candidate segments longer than the minimum interval are assigned a large interval cut cost, such as 10, 100, 1000, etc.), or others.

[0117] Thus, the cutting cost calculation component 186 of the video segmentation component 180 calculates the edge weights of the edges between nodes in the graph, and the path optimization component 188 of the video segmentation component 180 solves the shortest path problem along the graph to calculate the optimal segmentation, and the segment boundary obtained in the optimal segmentation represents the optimal capture alignment point based on the cutting cost.

[0118] Another example of video segmentation is thumbnail segmentation, which identifies positions on the video timeline for visualization using different thumbnails. Figure 1A In the example embodiment illustrated, thumbnail segmentation component 174 triggers video segmentation component 180 to calculate thumbnail segments. As explained in more detail below, an example use of thumbnail segmentation is in a user interface having a video timeline representing video (e.g., a composite timeline representing a selected video segment in an editor interface), where the thumbnail position identified by the thumbnail segment is illustrated using the corresponding video frame (thumbnail) on the timeline. In various embodiments, thumbnail segments are calculated at any suitable time (e.g., when the video timeline is displayed, the editor interface is activated, a video segment to be represented by the composite timeline is identified, and / or at some other time). Figure 1A and Figure 1B In the example embodiment illustrated, the video interaction engine 108 of the video editing application 105 (e.g., video editing tool 130) communicates with the thumbnail segmentation component 174 to trigger the video segmentation component 180 to calculate thumbnail segments.

[0119] In one example implementation of thumbnail segmentation, video segmentation component 180 uses any of the techniques described herein to perform segmentation routines. In some embodiments, video segmentation component 180 performs segmentation routines similar to the example implementation of capture alignment point segmentation described above, with the following additional or alternative aspects. For example, candidate boundary selection component 182 of video segmentation component 180 identifies candidate boundaries as candidate thumbnail locations from the boundaries of feature ranges in a specified feature track, and graph construction component 184 of video segmentation component 180 constructs a graph having nodes representing candidate thumbnail locations, edges representing candidate intervals (e.g., candidate segments) between thumbnail locations, and edge weights calculated by cut cost calculation component 186 of video segmentation component 180.

[0120] In one example implementation, the cut cost calculation component 186 assigns cut costs to candidate segments, which encourages thumbnails to be placed in “good” positions. As a non-limiting example, the cut cost calculation component 186 determines the edge weights between two nodes as a (e.g., a normalized) sum of the boundary cut costs for the candidate segments (e.g., penalizing candidate thumbnail positions that fall within a detected feature range or within a video portion with detected high movement, etc.) and the interval cut costs (e.g., penalizing candidate thumbnail positions where the visual difference between two consecutive thumbnails is small, penalizing thumbnail intervals corresponding to the minimum pixel interval of the thumbnails, based on scaling level, etc.).

[0121] Regarding boundary clipping costs, in some cases, to prevent thumbnails from being displayed at “bad” clipping points (e.g., within visual features detected from another feature track), the clipping cost calculation component 186 assigns low boundary clipping costs to candidate thumbnail locations based on proximity to visual feature boundaries (e.g., faces, scenes), assigns high boundary clipping costs to candidate thumbnail locations falling within the detected feature range, and / or assigns high boundary clipping costs to candidate thumbnail locations falling within video portions with high movement detected (e.g., detected using one or more machine learning models using feature extraction components 162).

[0122] Regarding interval clipping cost, in some cases, to encourage the display of thumbnails with different visual content, the clipping cost calculation component 186 determines the interval clipping cost for a candidate thumbnail position based on the visual similarity and / or difference between the visual content of two consecutive candidate thumbnails corresponding to the start and end boundaries of the candidate interval / segment. In an example involving facial or visual scene transitions, the clipping cost calculation component 186 calculates a measure of the similarity or difference between candidate thumbnails / video frames at thumbnail positions corresponding to the start and end boundaries of the candidate interval / segment, and penalizes thumbnail positions where consecutive thumbnails are within a threshold similarity. Additionally or alternatively, in some cases, the spacing between thumbnails cannot be closer than the width of the thumbnails. Thus, in one example embodiment, the target minimum thumbnail interval (e.g., represented by the target minimum video segment length) is determined based on the minimum pixel interval (e.g., the desired thumbnail width), which in some cases depends on the zoom level of the viewed video timeline. For example, a specified minimum thumbnail interval is mapped to a corresponding duration on the timeline (e.g., based on the activity zoom level), and this duration is used as the target minimum interval between thumbnails. In some cases, the target minimum interval is used as a hard constraint (e.g., candidate intervals / segments shorter than the minimum interval are assigned an infinite interval cut cost), a soft constraint (e.g., candidate intervals / segments longer than the minimum interval are assigned a large interval cut cost, such as 10, 100, 1000, etc.), or others.

[0123] Thus, the cutting cost calculation component 186 of the video segmentation component 180 calculates the edge weights between nodes in the graph, and the path optimization component 188 of the video segmentation component 180 solves the shortest path problem along the graph to calculate the optimal segmentation, and the segmentation boundary obtained in the optimal segmentation represents the optimal thumbnail position based on the cutting cost.

[0124] Additionally or alternatively, in some embodiments, the video segmentation component 180 calculates multiple levels of capture alignment point segments corresponding to different target video segment lengths (e.g., corresponding to different zoom levels, different input levels set by interactive control elements presented to the user, etc.). In some embodiments, lower-level capture alignment point segments include capture alignment points from higher-level segments plus additional capture alignment points (e.g., the input to a lower-level capture alignment point segment is a higher-level capture alignment segment, the capture alignment point segments operate on each video segment starting from the next lower level, etc.). These are just a few examples, and other implementations are contemplated within the scope of this disclosure.

[0125] In some embodiments, the video segmentation component 180 calculates multiple levels of segments corresponding to different zoom levels for a specific type of segment (e.g., capture alignment point segmentation, thumbnail segmentation). For example, when a user zooms in on the video timeline, in some cases, existing capture alignment points or thumbnails from higher-level segments are included in lower-level segments. Similarly, when a user zooms out on the video timeline, capture alignment points or thumbnails from higher-level segments are a subset of capture alignment points or thumbnails from lower-level segments. In some embodiments, the graph construction component 184 of the video segmentation component 180 constructs a graph to implement such a hierarchy, and different target video segment lengths are determined for different levels of the hierarchy (e.g., corresponding to different zoom levels, different input levels set by interactive control elements presented to the user, etc.). Thus, in some embodiments, one or more video segments are inherently hierarchical.

[0126] In some embodiments, video segmentation component 180 (or some other component) uses one or more data structures to generate representations of the computed video segment(s) 196. In one example implementation, the video segments of the video segment(s) 196 are identified by various values ​​that represent or reference timeline positions (e.g., boundary positions, IDs, etc.), segment durations, capture alignment points, or intervals between thumbnails, and / or other representations. In one example implementation involving hierarchical segmentation, a two-dimensional array is used to represent hierarchical segmentation, where the dimensions of the array correspond to different levels of segmentation, and the value stored in each dimension of the array represents a video segment in the corresponding hierarchical level.

[0127] In some cases, a single copy of the video and a representation of the boundary positions for one or more segments are preserved. Additionally or alternatively, in example embodiments involving a specific type of video segmentation of a video file (e.g., a default video segment), for efficiency purposes, the video file is broken into segments at the boundary positions of video clips from (e.g., the default) video segment and / or at feature boundaries from a feature track (e.g., visual scene boundaries). For example, as a motivation, a user might start or stop playback at the boundary of a video clip from the default video segment. Conventional techniques for generating segments with uniform spacing might require starting or stopping the video in the middle of a segment, which in turn leads to codec and / or playback inefficiencies. Similarly, uniformly spaced segments might require re-encoding and therefore be more expensive to export. Thus, in many cases, using the boundaries from one or more video segments (e.g., the default segment) and / or feature boundaries from a feature track (e.g., visual scene boundaries) as keyframes to start a new segment will make playback, stitching, and / or export operations more computationally efficient.

[0128] Interact with video segments

[0129] For example, previous chapters described sample techniques for video segmentation, preparing for video editing or other video interactions. By identifying semantically meaningful locations within the video, video segmentation tool 155 generates a structured representation of the video, which, for example, is achieved through... Figure 1A and Figure 1B The video interaction engine 108 in the video editing application 105 provides an efficient and intuitive structure for interacting with videos.

[0130] The video interaction engine 108 provides interface functionality that allows users to select, navigate, play, and / or edit videos by interacting with one or more segments of the video and / or detected features of the video. Figure 1B In the example implementation, the video interaction engine 108 includes a video browsing tool 110 providing a finder interface and a video editing tool 130 providing an editor interface. The video browsing tool 110 (e.g., a finder interface) and / or the video editing tool 130 (e.g., an editor interface) present one or more interactive elements that provide various interactive methods for selecting, navigating, playing, and / or editing videos based on one or more video segments 196. Figure 1B In this context, the video browsing tool 110 (e.g., a finder interface) includes various tools such as interactive tiles 112, a selected clip panel 114, a default resegmentation tool 116, a search resegmentation tool 118, a scripting tool 120, a segmented timeline tool 122, and a video playback tool 124. Figure 1B In this context, the video editing tool 130 (e.g., an editor interface) includes various tools such as a composite clip timeline tool 132, a search and re-segmentation tool 142, and a video playback tool 144. In various embodiments, these tools are implemented using code that corresponds to the presentation of interactive elements and the detection and interpretation of inputs interacting with interactive elements(s).

[0131] Regarding the video browsing tool 110 (e.g., a finder interface), the interactive tiles 112 represent video segments in the default segmentation and features detected in each video segment (e.g., Figure 1AThe user can select a video segment represented by interactive tile 112 and / or the detected feature 194 in interactive tile 112 to jump to the corresponding part of the video and / or add the corresponding video segment to the selected clip panel 114. The default resegmentation tool 116 recalculates the default segmentation based on the selected feature category (e.g., feature track) and updates the interactive tile 112 to represent the updated default segmentation. The search resegmentation tool 118 triggers a search segmentation that resegments the default segmentation based on a query. The script tool 120 presents a portion of a script corresponding to the video portion displayed in the video playback tool 124. In some embodiments, the user can select a portion of the script and add the corresponding video segment to the selected clip panel 114. The segmented timeline tool 122 provides a video timeline of the video segmented based on the active segmentation, and the video playback tool 124 plays back the selected portion of the video.

[0132] Regarding the video editing tool 130 (e.g., the editor interface), the composite clip timeline tool 132 presents a composite video timeline representing the composite video formed from video segments selected in the finder interface. In this example, the composite clip timeline tool 132 includes a feature visualization tool 134 representing features detected on the timeline, a selection box selection and snap alignment tool 136 representing capture alignment points on the timeline and / or aligning selected captures to those points, a thumbnail preview tool 138 representing thumbnails on the timeline, and a zoom / scroll bar tool 140 controlling the zoom level and positioning of the timeline. The search resegmentation tool 142 triggers a search segmentation that resegments video segments in the composite video based on a query. The video playback tool 144 plays back selected portions of the video. The editor panel 146 provides any number of editing functions for the selected(multiple) video segments(s), such as style improvements for changing content, improvements for the duration of the effects of omitted content, and / or contextual functions depending on the selected content. The functionality of video browsing tool 110, video editing tool 130, and other sample video interactive tools is described below. Figure 10 Figure 15 describes this in more detail.

[0133] Now go to Figure 10 , Figure 10 This is an illustration of a sample finder interface 1000 used for browsing default and / or search segments. Figure 10 In the example illustrated, the finder interface 1000 includes a video timeline 1005 (e.g., by...). Figure 1B The segmented timeline tool 122 controls the video frames 1010 (e.g., by...). Figure 1B The video playback tool 124 is controlled), and the interactive tiles 1020 (e.g., by...) Figure 1B Interactive tile 112 control), search bar 1060 (e.g., controlled by...) Figure 1B The search re-segmentation tool 118 is controlled by the script 1080 (e.g., by...). Figure 1B The script tool 120 controls the selection of the clip panel 1090 (e.g., controlled by the script tool 120). Figure 1B The selected clipboard 114 controls this.

[0134] In example use cases, such as using File Explorer to identify the location of a video (not depicted), the user loads the video for editing. In some cases, after receiving a command to load a video, the video is ingested to generate one or more segments (e.g., via...). Figure 1A The video ingestion tool 160 and / or video segmentation component 180 are used, and a default segment is loaded. In one example implementation, when the video is loaded and / or the user opens the finder interface 1000, the finder interface 1000 presents a video timeline 1005 representing active video segments such as the default segment (e.g., by displaying segment boundaries as the underlying layer), and / or the finder interface 1000 uses interactive tiles 1020 (e.g., interactive tile 112 of FIG. 1) representing video segments in active video segments (such as the default segment) to present a visual overview of the video. In one example implementation, the default segment is calculated based on detected sentences, faces, and visual scenes, and the interactive tiles 1020 are arranged in a grid consisting of rows and columns.

[0135] In some embodiments, the finder interface 1000 includes one or more interactive elements (e.g., by...). Figure 1B The default resegmentation tool 116 (controlled by the default segmentation tool) displays one or more input parameters for the default segmentation, allowing the user to modify the visual overview (e.g., by specifying one or more feature tracks for the default segmentation). Example visual overviews include those focused on people (e.g., based on detected faces, detected speakers, and / or detected script features), those focused on the visual scene (e.g., based on detected visual scene and / or detected script features), and those focused on sound (e.g., based on detected audio classification and / or detected script features). Based on one or more specified feature tracks and / or overviews, the default segmentation is recalculated, and the interactive tiles 1020 are updated to represent the updated default segmentation. Thus, resegmenting the default segmentation allows the user to quickly visualize the video in different ways.

[0136] In the finder interface 1000, a user can scan the video timeline 1005 (which updates video frames 1010), scan scripts 1080, or browse interactive tiles 1020. Each interactive tile 1020 (e.g., interactive tile 1030) includes a thumbnail (e.g., a thumbnail 1032 of the first video frame of the video segment represented by interactive tile 1030) and representations of one or more detected features and / or corresponding feature ranges, such as a script from the beginning of the video segment (e.g., script 1034), a detected face from the video segment (e.g., face 1036), and one or more faceted timelines in its own faceted timeline (e.g., visual scene timeline 1038, faceted audio timeline 1040). In some embodiments, the faceted timeline represents detected facets in a particular category of detected features (e.g., visual scene, audio classification) and their corresponding positions within the video segment. Each interactive tile 1020 in the interactive tile 1020 allows the user to navigate the video by clicking on a facet in the faceted timeline, which jumps the video frame 1010 to the corresponding part of the video. In some embodiments, the user can customize the visualization features in the interactive tile 1020 by turning the visualization of features for specific categories on / off (e.g., controlling the visualization of people, sounds, visual scenes, and visual artifacts respectively by clicking buttons 1062, 1064, 1066, or 1068).

[0137] Figures 11A-11BThese are illustrations of example interactive tiles 1110 and 1150. Interactive tile 1110 includes a thumbnail 1115 of the first video frame of the video segment represented by interactive tile 1110, a script 1120 (represented by virtual text) from the beginning of the video segment, a visual scene timeline 1125, a faceted audio timeline 1140, clip duration, and an add button 1148 for adding the video segment represented by interactive tile 1110 to the selection. In this example, the visual scene timeline 1125 is faceted based on instances of detected visual scenes appearing in the video segment represented by interactive tile 1110. More specifically, segment 1130 represents one visual scene, and segment 1135 represents another visual scene. Furthermore, the faceted audio timeline 1140 is faceted based on instances of detected audio categories (e.g., music, speech, others) appearing in the video segment represented by interactive tile 1110. More specifically, segment 1144 represents one audio category (e.g., speech), while segment 1142 represents another audio category (e.g., music). In this example, a user can click on one of the facets from the faceted timeline (e.g., segments 1130, 1135, 1142, 1144) to jump to that segment of the video. In another example, interactive tile 1150 includes a representation of a detected face 1160 in a video segment represented by interactive tile 1160. Visualization of detected features such as these helps users navigate the video without replaying it.

[0138] In some embodiments, hovering over a portion of an interactive tile (such as a faceted timeline, a thumbnail, and / or anywhere within the interactive tile) updates the thumbnail in the interactive tile, or presents a pop-up with a thumbnail of the corresponding portion of the video (e.g., Figure 10 (Pop-up thumbnails 1055). For example, a horizontal input positioning (e.g., the x-position of mouse or click input) relative to the total width of the interactive tile (e.g., the faceted timeline, thumbnails, or interactive tiles hovered over) is mapped to a percentage offset in the video segment represented by the interactive tile, and the corresponding thumbnail is found and displayed. In some embodiments, uniformly sampled video frames may be used as thumbnails. In other embodiments, available thumbnails are identified by thumbnail segments (e.g., the most recent available thumbnail or available thumbnails within a threshold distance of the horizontal input positioning are returned). Thus, a user can scan through a set of thumbnails by hovering over one or more portions of the interactive tile.

[0139] exist Figure 10In the embodiment illustrated, the finder interface 1000 includes a search bar 1060. In this example, the search bar 1060 accepts one or more keywords (e.g., entered in the search field 1070) and / or (e.g., entered via corresponding menus accessible via buttons 1062, 1064, 1066, and 1068) one or more selected facets, triggering a search segmentation that re-segments the default segments based on the query. In one example implementation, the user types one or more keywords in the search field 1070 and / or selects one or more facets via menus or other interactive elements representing feature categories (feature tracks) and / or corresponding facets (detected features). Figure 10 Interacting with buttons 1062, 1064, 1066, and 1068 (e.g., by hovering over the menu, left-clicking a button corner, or right-clicking) activates the corresponding menu, which displays the detected person (button 1062), the detected sound (button 1064), the detected visual scene (button 1066), and the detected visual artifact (button 1068). Figures 12A-12D This is an illustration of an example faceted search menu. Figure 12A An example menu with detected faces is shown (e.g., by...). Figure 10 Button 1062 is activated. Figure 12B An example menu with detected sounds is shown (e.g., by...). Figure 10 Button 1064 is activated. Figure 12C An example menu with a detected visual scene is shown (e.g., by...). Figure 10 Button 1066 is activated. Figure 12D An example menu with detected visual artifacts is shown (e.g., by...). Figure 10 (Button 1068 is activated). This allows users to navigate... Figures 12A-12D The faceted search menu shown in the figure allows you to select one or more facets (e.g., specific faces, voice categories, visual scenes, and / or visual artifacts), enter one or more keywords into the search field 1070, and run the search (e.g., by clicking on a facet, the existing faceted search menu, clicking the search button, etc.).

[0140] In some embodiments, a typed keyword search triggers a search for detected features that have associated text or values ​​matching the keyword (e.g., script, object or action tag, audio event tag, log event tag, etc.), and / or a selected facet triggers a search for detected instances of the selected facet(s). In one example implementation, search bar 1060 triggers... Figure 1A The search segmentation component 170 and / or the video segmentation component 180 are used to calculate search segments, which are calculated by the search segmentation component 170 and / or the video segmentation component 180. Figure 10The default segmentation represented by interactive tile 1020 is re-segmented, thereby updating interactive tile 1020 to represent video segments of the resulting search segment. In this example, the search is used to break down interactive tiles that match the query to represent smaller units of video based on the query. In other words, tiles that match the query are broken down into smaller video segments, while tiles that do not match remain unchanged. In some embodiments, tiles that match the query and are broken down into smaller video segments are animated, for example, as illustrated in the diagram of the broken-down tiles.

[0141] In some embodiments, the finder interface 1000 emphasizes interactive tiles representing matching video clips (query open clips). For example, Figure 11B Interactive tiles 1150 are illustrated with fragments 1170 indicating that the tile is a match for the query. Other examples of emphasis include outlining, adding fill (e.g., transparent fill), etc. In some embodiments, matching tiles additionally or alternatively indicate why the tile is a match, such as by presenting or emphasizing representations of features(s) that match the query (e.g., matching faces, visual scenes, sound classifications, objects, keywords, etc.). For example, Figure 11B Interactive tile 1150 highlights one of the faces in face 1160, indicating that the face matches the query. In another example, matches with a script are used to highlight, underline, or otherwise emphasize matching words in the tile. In this way, users can easily see which interactive tiles match the query and why.

[0142] In some cases, the size of the video clip a user wants when searching for content can vary depending on the task. For example, if a user wants to find a clip of a child laughing, they might only want a few seconds of search results, but if they want to find a clip of a rocket launch, they might want longer search results. Thus, in some embodiments, the finder interface 1000 provides a slider or other interactive element (not shown) that displays input parameters for segmentation (e.g., the target minimum and maximum length of the video clip), allowing the user to interactively control the size of the video clips generated by segmentation and represented by interactive tiles 1020. In some embodiments, one or more interactive tiles (e.g., each tile) provide their own slider or other interactive element (e.g., a handle) that displays input parameters, allowing the user to interactively control the size of the video clip(s) represented by a particular tile. Therefore, various embodiments provide one or more interactive elements that allow the user to break down a tile into smaller parts locally (per tile) and / or globally (all tiles).

[0143] Script 1080 presents a script for the video and highlights the active portion 1085 of the script. In some embodiments, script 1080 provides a segmentation clustering script that represents a portion of the active portion of script 1085 of the detected speaker.

[0144] The selected clip panel 1090 represents the video clip added by the user to the selection. In one example implementation, the user can add a video clip to the selection by: dragging an interactive tile into the selected clip panel 1090, or clicking the + button in the interactive tile (e.g., ...). Figure 11A The interactive tile 1110 contains buttons 1148, and the interactive tiles contain visual interactions with detected features or facets (e.g., by highlighting a portion of script 1080, right-clicking to activate a context menu, and adding selections from the context menu, right-clicking, etc.). Figure 11A The visualization of one of the segments 1130, 1135, 1142, 1144, or facet 1160 can be used to activate the context menu and add the corresponding subset of the video segments from the context menu to the selection, and / or other methods. Figure 10 In the middle, the selected clip panel 1090 displays a list of thumbnails or selected video clips.

[0145] Once a set of video clips is selected, the user can switch to the editor interface to perform one or more editing functions. Figure 10 In the example shown, the finder interface 1000 provides one or more navigation elements (e.g., finder button 1095, editor button 1097, edit button in the selected clip panel 1090, etc.) for navigating between the finder and editor interfaces.

[0146] Figure 13 This is an illustration of an example search in the finder interface. In this example, the user enters the query "giggle" into search field 1310, which triggers a search segment that highlights matching tiles (e.g., matching tile 1320) and matching portions of scripts (e.g., script area 1325). In this example, the user adds matching video clips represented by tiles 1332, 1334, and 1336 to the selected clip panel 1340, which is represented by thumbnails 1342, 1344, and 1346. Figure 13 In this process, the user adds another matching video clip by clicking the add button 1330 in the corresponding interactive tile. Once the user has finished adding the video clip to the selection in the selected clip panel 1340, the user clicks the button 1350 (Edit Your Clip) to switch to the editor interface, such as... Figure 14 The interface shown in the figure.

[0147] Figure 14 This is an illustration of a sample editor interface 1400 used for video editing. Figure 14 In the example shown, the editor interface 1400 includes a video timeline 1405 (e.g., by...). Figure 1B The composite editing timeline tool 132 controls the search bar 1450 (e.g., by...). Figure 1B The search re-segmentation tool 142 (controlled by) and the editor panel 1460 (e.g., by) Figure 1B The editor panel 146 controls the video timeline 1405, which includes thumbnails 1410 (e.g., by...). Figure 1B The thumbnail preview tool 138 controls the faceted audio timeline 1420 and the faceted artifact timeline 1430 (e.g., controlled by the thumbnail preview tool 138). Figure 1B The feature visualization tool 134 controls the selection box 1440 (e.g., by...). Figure 1B The selection box and snap alignment tools (136 control) are used for selection.

[0148] In one example implementation, the editor interface 1400 presents a video timeline 1405 representing active video segments (e.g., by displaying segment boundaries as the underlying layer). In one example use case, such as using a file explorer to identify the location of the video (not depicted), the user loads the video for editing, and the video is ingested to generate one or more segments (e.g., via...). Figure 1A The video ingestion tool 160 and / or video segmentation component 180), and / or editor interface 1400 initialize the video timeline 1405 using segments of the entire video (e.g., default segments). In another example use case, a subset of video clips in the video has been previously selected or otherwise designated for editing (e.g., when switching from a finder interface where one or more video clips have been added to a selected clip panel, or when there are already editing items at load). When editor interface 1400 is opened and / or an existing editing item is loaded, editor interface 1400 initializes the video timeline 1405 using the designated video clips. In some cases, video timeline 1405 represents a composite video formed by those video clips designated for editing and / or re-segmentation (e.g., searching for segments) of those video clips. At a higher level, users can select, edit, move, delete, or otherwise manipulate video clips from video timeline 1405.

[0149] In one example implementation, the finder and editor interfaces are linked by one or more navigation elements (e.g., Figure 10 Finder button 1095 and editor button 1097, Figure 14The system includes a finder button 1495 and an editor button 1497, which switch between the finder and editor interfaces. In this implementation, the user can use the finder interface to browse the video and add video clips to the selection, switch to the editor interface, and perform one or more enhancements or other editing operations. In some embodiments, when the user specifies a video clip for editing from the finder and switches to the editor interface, the editor interface creates a representation of the composite video arranged in chronological order, in the order in which the selected video clips were added to the selection, in some specified order (e.g., the order in which the selected video clips are arranged in the selected clip panel), and / or in other ways. In some embodiments, the editor interface 1400 (e.g., the video timeline 1405) represents the boundaries of the specified video clips and / or represents features detected in the specified video clips. In an example implementation of the editor interface 1400, the user can only browse the content that has been added to the composite video; the user can trim content that is no longer desired, but in order to add new content to the composite video, the user returns to the finder interface.

[0150] In some embodiments of the editor interface 1400, a user can scan through the video timeline 1405 and jump to different parts of the composite video by clicking on the timeline. Additionally or alternatively, a user can jump to different parts of the composite video by scanning the script 1445 and clicking on specific sections (e.g., words). In some embodiments, the script is presented side-by-side with the video, at the top of the video (e.g., as shown in the image). Figure 14 (in Chinese) and / or otherwise presented.

[0151] In some embodiments, to help identify specific portions of the synthesized video, the video timeline 1405 represents one or more detected features and / or their location within the synthesized video (e.g., corresponding feature ranges). In some embodiments, the video timeline 1405 utilizes corresponding faceted timelines to represent each category of detected features, where the faceted timeline represents the detected facets (e.g., faces, audio classifications, visual scenes, visual artifacts, objects, or actions, etc.) and their corresponding locations within the video segment. In some embodiments, a user can customize the visualized features on the video timeline 1405 by turning the visualization of features for specific categories on / off (e.g., controlling the visualization of people, sounds, visual scenes, and visual artifacts by clicking buttons 1462, 1464, 1466, and 1468, respectively). Figure 14In the embodiment illustrated, the user has activated visualizations for sound (e.g., via button 1464) and visual artifacts (e.g., via button 1468), so the video timeline 1405 includes a faceted audio timeline 1420 and a faceted artifact timeline 1430. In some embodiments, clicking on a specific facet from the faceted timeline is used to select the corresponding portion of the video (e.g., as illustrated in selection box 1440 for a video portion with a selected visual artifact). For example, in this way, the user can easily select and delete the video portion with detected visual artifacts. Typically, snap alignment to semantically meaningful snap alignment points helps the user trim quickly.

[0152] In some embodiments, a portion of the composite video represented by video timeline 1405 is selectable through interaction with video timeline 1405 and / or script 1445. Typically, selections are emphasized in any suitable manner, such as outlining (e.g., using dashed lines), adding fill to the selected area (e.g., transparent fill), and / or other methods. In one example implementation, a selection (e.g., a selection box selection, such as selection box selection 1440) is created by clicking or tapping and dragging on a video segment represented in video timeline 1405 or on script 1445. In some embodiments, a selection made on one element (video timeline 1405 or script 1445) additionally emphasizes (e.g., highlights) a corresponding portion of another element (not shown). In some cases, a selection can be edited after it has been drawn by clicking and dragging the start and / or end points of the selection. In one example implementation, a selection drag operation (e.g., along video timeline 1405, script 1445) aligns the selection boundary snap to a snap alignment point defined by (e.g., calculated as described above) snap alignment point segments and / or the current zoom level. In some embodiments, video timeline 1405 presents a visualization of the snap alignment point defined by the snap alignment point segments and / or the current zoom level. In some cases, the snap alignment point is displayed only during the drag operation (e.g., on video timeline 1405) so that the snap alignment point displayed on video timeline 1405 disappears when the drag operation is released.

[0153] Figures 15A-15B This is an example illustration selected using the snap-aligned selection box. Figure 15A In the example, the user has created a selection box 1520 by clicking and dragging the cursor 1510 to the left. Selection box 1520 snaps to an alignment point indicated by a vertical bar (e.g., snap alignment point 1530). In this example, the user has previously added video clips that match the query "giggle," so some video clips are matched based on the giggle detected in the audio track. Figures 15A-15BIn the middle, the faceted audio timeline 1540 represents different detected audio categories in different ways and with the volume level of the audio track (e.g., by using one color to represent detected music, a second color to represent detected speech, and a third color to represent other sounds). Figure 15B In the segmented audio timeline 1540, the detected giggling sounds 1550, 1560, and 1570 are represented as other sounds. However, the detected giggling sound 1560 originates from a video clip with a detected visual artifact 1580. To remove this clip, the user moves the selection box 1520 to... Figure 15B The updated location shown in the image ensures that video segments with other detected giggling sounds (1550 and 1570) are ignored. This allows the user to delete... Figure 15B Select the video segment within 1520 using the selection box in the composite video to remove the video segment with detected visual artifacts 1580 from the composite video.

[0154] return Figure 14 In some embodiments, the video timeline 1405 displays one or more thumbnails (e.g., thumbnail 1410). In one example implementation, when a user first opens the editor interface 1400, the editor interface 1400 (e.g., video timeline 1405) uses a thumbnail at the beginning of each video segment to represent each video segment in the video. Additionally or alternatively, the editor interface 1400 (e.g., video timeline 1405) represents thumbnails at video locations identified by thumbnail segments (e.g., calculated as described above). In one example implementation, each video segment has at least one thumbnail, and longer video segments are more likely to include multiple thumbnails. In some embodiments, when a user zooms in on the video timeline 1405, more thumbnails appear (e.g., based on thumbnail segments calculated at that zoom level), and / or thumbnails visible at higher zoom levels remain in place. Thus, the thumbnails on the video timeline 1405 serve as landmarks to aid in navigating the video and selecting video segments.

[0155] In some embodiments, the video timeline 1405 includes zoom / scroll bar tools (e.g., by...). Figure 1BThe zoom / scroll tool 1400 controls the zoom level and allows the user to change the zoom level and scroll to different positions on the video timeline 1405. In some cases, the characteristics of the capture alignment points, thumbnails, and / or visualizations depend on the zoom level. In one example implementation, as the user zooms in on the video timeline, different capture alignment point segments and / or thumbnail segments are calculated or looked up based on the zoom level (e.g., different zoom levels are mapped to different value parameters for one or more segments, such as the target video segment length, the target minimum interval between capture alignment points, and the target minimum interval between thumbnails). Additionally or alternatively, the editor interface 1400 provides one or more interactive elements that display one or more segmentation parameters for the capture alignment point segments and / or thumbnail segments, thereby enabling the user to control the granularity of the capture alignment point segments and / or thumbnail segments. In some embodiments, zooming in and out on the video timeline 1405 expands and reduces the detail of the visualization features on the video timeline 1405. In one example implementation, the high zoom level displays the audio faceted timeline representing the MSO (Music / Speech / Other) audio category, and the zoom operation adds representations of detected audio events to the video timeline 1405. Additionally or alternatively, higher zoom levels integrate feature ranges or other feature visualizations that are expanded to display more detail at lower zoom levels.

[0156] In some embodiments, the editor interface 1400 accepts queries (e.g., keywords and / or facets), triggers temporary search segments that segment video segments in the synthesized video based on the query, and presents a visualization of the search segments (e.g., by illustrating the boundaries of their video clips as the underlying layer of the video timeline 1405). Figure 14 In the embodiment illustrated, the editor interface 1400 includes a search bar 1450. In this example, the search bar 1450 accepts queries in the form of one or more keywords (e.g., entered in the search field 1470) and / or one or more selected facets (e.g., entered via corresponding menus accessible via buttons 1462, 1464, 1466, and 1468), and triggers a search segmentation process that re-segments video segments in the synthesized video based on the query. In one example implementation, the user types one or more keywords in the search field 1470 and / or selects one or more facets via menus or other interactive elements representing feature categories (feature tracks) and / or corresponding facets (detected features). Figure 14Interacting with buttons 1462, 1464, 1466, and 1468 (e.g., by hovering over the menu, left-clicking a button corner, or right-clicking) activates the corresponding menu, which displays the detected person (button 1462), detected sound (button 1464), detected visual scene (button 1466), and detected visual artifact (button 1468). Thus, the user can navigate the faceted search menu, select one or more facets (e.g., a specific face, sound category, visual scene, and / or visual artifact), enter one or more keywords into search field 1470, and run a search (e.g., by clicking on a facet, an existing faceted search menu, clicking the search button, etc.).

[0157] In some embodiments, when a user queries by aspect or keyword, search bar 1450 triggers a temporary search segment and highlights a matching video segment in video timeline 1405. In this example, the search segment is considered temporary because it does not perform any destructive operations on video clips in the composite video. If the user performs another query by adding or removing keywords or facets, search bar 1450 triggers a new temporary search segment. If the user deletes or removes the query, the search segment disappears, and video timeline 1405 re-invokes the representation of the composite video as it was before the search. In some embodiments, keyword and facet queries persist until the user deletes or removes them, clearing any search result highlighting. In one example implementation, the search state does not persist as the user switches back and forth between the finder and editor interfaces.

[0158] In some embodiments, the search segments in the editor interface 1400 take into account any existing video segments in the composite video without altering any of their boundaries (e.g., the search segments operate independently on each video segment in the composite video). If a user performs an action on a search result (e.g., deleting a matching video segment), a new edit boundary is created to reflect that action. In other words, if a user searches for "tree" and deletes a video segment depicting a tree in one part of the composite video but not in another, the part where the user performed the action (in this case, deletion) will have a new segment boundary, but the other parts will not. In some embodiments, the search results are affected by the zoom level (e.g., shown in more precision or zoom level detail) and / or the corresponding area on the video timeline 1405 is illustrated to show where the query is active in the composite video.

[0159] Therefore, various embodiments of the video timeline 1405 present a high-level overview of the visual and audio content contained within the synthesized video, depending on the feature category switched to, the search criterion zoom level, and the screen size. Thus, in various embodiments, users can simultaneously view detected features, capture alignment points, video thumbnails, and / or search results to help them select good cut points.

[0160] In one example implementation, after selecting one or more video segments, one or more editing functions provided by the editor interface 1400 are used to edit the selected video segments. For example, the editor panel 1460 provides any number of editing functions for the selected video segments. Depending on the implementation, available editing functions include style improvements to transform content (e.g., wind noise reduction), improvements to the duration of the effects of omitted content (e.g., "hiding" a shot area, removing profanity, creating a time delay, shortening to n seconds), and / or contextual functions depending on the selected content (e.g., removing words from content with a corresponding script or emitting a beep). In some embodiments, improvements to video attributes are declarative and non-destructive. For example, if a selection is made and overlaps portions of a composite video with previously applied attributes, any newly applied attribute will overwrite conflicting attributes with a new value. In various embodiments, the editor panel 1460 provides any suitable editing functionality, including rearranging, cropping, applying transitions or effects (e.g., changing speed, volume), adjusting colors, adding titles or graphics, etc.

[0161] In this way, the resulting composite video can be played back, saved, exported, or other operations can be performed. In one example, video segments from the composite video can be played back (e.g., after clicking the play button), skipping video segments not in the composite video. In another example, video segments from the composite video can be exported. Depending on the implementation, any known tools or techniques can be used to perform any type of operation on the video segments in the composite video.

[0162] The aforementioned video segmentation and interaction techniques are merely examples. Other variations, combinations, and sub-combinations are contemplated within the scope of this disclosure.

[0163] Example Flowchart

[0164] Now refer to Figure 16- Figure 22The document provides flowcharts illustrating various methods. Each block of methods 1600 to 2200, and any other methods described herein, includes a computational process performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. To name just a few, these methods can be provided by a standalone application, service, or managed service (independently or in combination with another managed service), or as a plug-in to another product.

[0165] First turn Figures 16A-16B These figures illustrate a method 1600 for generating video segments using a graph model, according to embodiments described herein. Initially, at box 1610, one or more feature tracks are accessed. The one or more feature tracks represent instances of features detected in the video and the corresponding time ranges when the instances are present in the video. To generate optimal video segments, a graph model is used to evaluate candidate segments. More specifically, at box 1620, a graph representation is generated using nodes and edges, such that nodes represent candidate boundaries selected from the boundaries of the time ranges, and edges represent candidate video segments and have edge weights representing the cutting cost for the candidate video segments. Boxes 1640 and 1650 illustrate possible ways of performing at least a portion of box 1620.

[0166] At box 1640, candidate boundaries are selected from a subset of the boundaries of the time range when the detected features are present in the video. Figure 16B Boxes 1642-1646 illustrate possible ways of performing at least a portion of box 1640. At box 1642, instances of detected features overlapping a specified range of the video to be segmented are identified from a specified feature track. These identified features can be considered “overlapping features” because they exist in the video during a time range that overlaps with a specified range of the video to be segmented (e.g., the entire video, a specific video segment). At box 1644, the boundaries of the time range where the identified instances (overlapping features) exist are adjusted to capture the nearest boundary of the time range where the preferred feature exists, forming a subset of the time range boundaries. In one example implementation, the priority list of features includes script features (e.g., sentences), then visual scenes, and then faces. At box 1646, candidate boundaries are identified from the subset.

[0167] At box 1650, calculate the shearing cost for the edge weights in the graph. Figure 16BBoxes 1652-1656 illustrate possible ways of performing at least a portion of box 1650. At box 1652, a boundary clipping cost is determined for a candidate fragment. In one example embodiment, the boundary clipping cost penalizes the leading boundary of a candidate fragment based on the negative characteristics of the leading boundary. In some embodiments, the boundary clipping cost penalizes the leading boundary of a candidate fragment because it is within a detected feature (e.g., a detected word) from another feature track, or in another scene, based on the distance from the boundary of the time range to the detected visual scene boundary. In some embodiments, Equation 1 is used to calculate the boundary clipping cost. At box 1654, an interval clipping cost is determined for a candidate fragment. In one example embodiment, the interval clipping cost penalizes the candidate fragment based on the negative characteristics of the candidate fragment's span. In some embodiments, the interval clipping cost penalizes candidate fragments whose length is outside the target minimum or maximum length, candidate fragments with incoherence to overlapping features from other feature tracks, and / or candidate fragments that overlap with both query on and off features. At box 1656, the edge weight between two nodes is determined as the sum of the boundary clipping cost and the interval clipping cost for the candidate fragment.

[0168] return Figure 16A At box 1660, the shortest path through the nodes and edges is calculated. At box 1670, the representation of the video segment corresponding to the shortest path is rendered.

[0169] Now switch to 17. Figure 17 The illustration depicts a method 1700 for segmenting a video into search segments according to an embodiment of the present invention. Initially, at box 1710, the default segments of the video are rendered. At box 1720, a query is received. At box 1730, detected features of the video are searched for matching features that match the query. At box 1740, the default segments are re-segmented into search segments based on the query. Box 1745 illustrates a possible manner of performing at least a portion of box 1740. At box 1745, each video segment in the default segments, including at least one matching feature, is segmented based on the query. At box 1750, the rendering is updated to represent the search segments.

[0170] Now go to Figure 18 , Figure 18The illustration depicts a method 1800 for navigating a video using interactive tiles according to an embodiment of the present invention. Initially at box 1810, interactive tiles are rendered, representing (i) segments of video and (ii) instances of detected features of the video segments. At box 1820, a click or tap is detected in one of the interactive tiles. More specifically, the click or tap is on a visualization of one of a plurality of instances of one of the detected features detected from one of the video segments represented by the interactive tiles. At box 1830, the video is navigated to the portion of the video where the instance of the detected feature is present.

[0171] Now go to Figure 19 , Figure 19 A method 1900 for adding to a selection of video segments according to an embodiment of the present invention is illustrated. Initially at box 1910, an interactive tile is rendered, which represents (i) a segment of video and (ii) an instance of a detected feature of the video segment. At box 1920, interaction with one of the interactive tiles representing one of the video segments is detected. Boxes 1930-1950 illustrate possible ways of performing at least a portion of box 1920. At box 1930, a click or tap (e.g., a right-click or tap and hold) is detected in the interactive tile. This click or tap is on a visualization of one instance of one of a plurality of instances of one of the detected features detected from one of the video segments represented by the interactive tile. At box 1940, a context menu is activated in response to the click or tap. At box 1950, input of an option to select from the context menu is detected. This option is used to add a portion of the video segment corresponding to that instance of the detected feature to the selection. At box 1960, in response to the detection of an interaction, at least a portion of the video clip is added to the selection of the video clip.

[0172] Now go to Figure 20 , Figure 20 The illustration depicts a method 2000 for capturing and aligning selection boundaries of a selected video segment according to an embodiment of the present invention. Initially, at box 2010, a first segment of the video timeline is rendered. At box 2020, a representation of a second segment of the first segment is generated using one or more feature tracks, which represent instances of features detected in the video and feature ranges indicating when those instances are present in the video. At box 2030, in response to a dragging operation along the video timeline, the selection boundaries of a selected portion of the video are captured and aligned to a capture alignment point defined by the second segment.

[0173] Now go to Figure 21 , Figure 21The illustration depicts a method 2100 for presenting a video timeline with thumbnails at locations defined by thumbnail segments according to an embodiment of the invention. Initially, at box 2110, the presentation of a first segment of the video timeline is initiated. At box 2120, the generation of a representation of the video's thumbnail segments is triggered. The generation of this representation uses one or more feature tracks that represent instances of features detected in the video and feature ranges indicating when those instances are present in the video. At box 2130, the video timeline is updated to represent one or more thumbnails of the video at locations on the video timeline defined by the thumbnail segments.

[0174] Now go to Figure 22 , Figure 22 The illustration depicts a method 2200 for presenting a video timeline with thumbnails at locations defined by thumbnail segments. Initially, at box 2210, the video timeline is presented. At box 2220, a representation of the video's thumbnail segments is accessed. These thumbnail segments define thumbnail positions on the video timeline at the boundaries of a range of features at which instances of detected features are present in the video. At box 2230, the presentation is updated to include a thumbnail at one of the thumbnail positions defined by the thumbnail segments, and this thumbnail depicts the portion of the video associated with that thumbnail position.

[0175] Example operating environment

[0176] Having described an overview of embodiments of the invention, we now describe example operating environments in which embodiments of the invention may be implemented, in order to provide a general context for the various aspects of the invention. Specific references will now be made. Figure 23 An example operating environment for implementing embodiments of the present invention is shown and is generally designated as computing device 2300. Computing device 2300 is merely one example of a suitable computing environment and is not intended to imply any limitation on the scope of the purpose or functionality of the invention. Computing device 2300 should also not be construed as having any dependencies or requirements associated with any of the components or combinations illustrated.

[0177] This invention can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a cellular phone, personal data assistant, or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This invention can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This invention can also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.

[0178] refer to Figure 23 The computing device 2300 includes a bus 2310 that is directly or indirectly coupled to the following devices: a memory 2312, one or more processors 2314, one or more presentation components 2316, an input / output (I / O) port 2318, an input / output component 2320, and an exemplary power supply 2322. The bus 2310 may represent one or more buses (such as an address bus, a data bus, or a combination thereof). Although for clarity... Figure 23 The various boxes are represented by lines, but the actual definition of the various components is not so clear, and metaphorically, the lines would be more accurately described as gray and blurred. For example, one can consider the presentation components of a display device as I / O components. Furthermore, a processor has memory. The inventors recognize this as the nature of the art and reiterate... Figure 23 The diagrams are merely illustrative examples of computing devices that can be used in conjunction with one or more embodiments of the present invention. There is no distinction between categories such as “workstation,” “server,” “laptop,” “handheld device,” etc., because all these categories are... Figure 23 Within the scope of and referred to as "computing device".

[0179] Computing device 2300 typically includes a variety of computer-readable media. Computer-readable media can be any available medium accessible by computing device 2300 and includes volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 2300. Computer storage media itself does not include signals. Communication media typically embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and includes any information delivery medium. The term "modulated data signal" means a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, communication media includes wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0180] Memory 2312 includes computer storage media in the form of volatile and / or non-volatile memory. Memory can be removable, non-removable, or a combination of both. Example hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Computing device 2300 includes one or more processors that read data from various entities such as memory 2312 or I / O components 2320. Multiple presentation components 2316 present data indications to a user or other device. Example presentation components include display devices, speakers, printing components, vibration components, etc.

[0181] I / O port 2318 allows computing device 2300 to be logically coupled to other devices including I / O component 2320, some of which may be built-in. Illustrative components include microphones, joysticks, game controllers, satellite antennas, scanners, printers, wireless devices, etc. I / O component 2320 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some instances, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometrics, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (described in more detail below) associated with the display of computing device 2300. Computing device 2300 may be equipped with depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. In addition, computing device 2300 may be equipped with accelerometers or gyroscopes capable of detecting motion. The output of the accelerometer or gyroscope can be provided to the display of the computing device 2300 to render immersive augmented reality or virtual reality.

[0182] The embodiments described herein support video editing or playback. The components described herein refer to integrated components of a video editing system. Integrated components refer to the hardware architecture and software framework that support the functionality of using a video editing system. Hardware architecture refers to physical components and their interrelationships, while software framework refers to the software that provides functionality that can be implemented using the hardware embodied on the device.

[0183] End-to-end software-based video editing systems can operate within video editing system components to manipulate computer hardware to provide video editing system functionality. At a low level, the hardware processor executes instructions selected from a set of machine language (also known as machine code or native) instructions for a given processor. The processor recognizes native instructions and performs corresponding low-level functions related to, for example, logic, control, and memory operations. Low-level software written in machine code can provide more complex functionality for higher-level software. As used herein, computer-executable instructions include any software, including low-level software written in machine code, higher-level software such as application software, and any combination thereof. In this respect, video editing system components can manage resources and provide services for the functionality of the video editing system. Embodiments of the invention are contemplated in any other variations and combinations thereof.

[0184] While some implementations of neural networks are described, typical implementations can be implemented using any type(s) of machine learning models, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutions, recurrent loops, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machines, etc.) and / or other types of machine learning models.

[0185] Various components in this disclosure have been identified, and it should be understood that any number of components and arrangements can be used to achieve the desired functionality within the scope of this disclosure. For example, for clarity of concept, components in the embodiments depicted in the figures are shown by lines. Other arrangements of these and other components can also be implemented. For example, while some components are depicted as single components, many elements described herein can be implemented as discrete or distributed components or combined with other components, and implemented in any suitable combination and location. Some elements may be omitted entirely. Furthermore, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, various functions can be performed by a processor executing instructions stored in memory. Thus, other arrangements and elements (e.g., machines, interfaces, functions, sequences, and groups of functions, etc.) may be used in addition to or in place of those shown.

[0186] The subject matter of this invention has been specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of this patent. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways in combination with other current or future techniques to include different steps or combinations of steps similar to those described in this document. Furthermore, although the terms “step” and / or “box” may be used herein to mean different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the individual steps is explicitly described.

[0187] The invention has been described in conjunction with specific embodiments, which are intended to be illustrative in all respects and not restrictive. Alternative embodiments will become apparent to those skilled in the art without departing from the scope of the invention.

[0188] As can be seen from the foregoing, the present invention is well-suited to achieving all the aforementioned objects and objectives, as well as other obvious and inherent advantages for the system and method. It should be understood that certain features and sub-combinations are useful and can be used without reference to other features and sub-combinations. This is contemplated by the claims and is within the scope of the claims.

Claims

1. One or more computer storage media storing computer-usable instructions, which, when used by one or more computing devices, cause the one or more computing devices to perform operations, said operations including: This results in the presentation of: (i) multiple interactive tiles representing multiple video segments and (ii) interactive elements that accept user input to toggle a visualization within each of the multiple interactive tiles, wherein the toggled visualization represents multiple detected features of a specified category that are detected from the corresponding video segment. Detecting a click or tap on an instance of one of the detected features in one of the multiple interactive tiles representing one of the multiple video segments; as well as Navigate to a portion of the video where the instance of the detected feature exists.

2. The computer storage medium of claim 1, wherein the designated category includes detected people, and the interactive element toggles visualization of each detected face or detected speaker within each interactive tile of a video clip, the detected face or the detected speaker being detected from the video clip.

3. The computer storage medium of claim 1, wherein the designated category includes detected sounds, and the interactive element toggles to represent visualization of each detected sound within each interactive tile of the video segment, the detected sounds being detected from the video segment.

4. The computer storage medium of claim 1, wherein the designated category includes detected visual scenes, and the interactive element toggles to represent visualization of each detected visual scene within each interactive tile of the video clip, the detected visual scene being detected from the video clip.

5. The computer storage medium of claim 1, wherein the designated category includes detected visual artifacts, and the interactive element toggles to represent visualization of each detected visual artifact within each interactive tile of a video segment, the detected visual artifacts being detected from the video segment.

6. The computer storage medium of claim 1, wherein the presentation further comprises a second interactive element that accepts user input controlling a global target duration, the global target duration being for all video segments of the plurality of video segments represented by the plurality of interactive tiles.

7. The computer storage medium of claim 1, wherein each of the plurality of interactive tiles includes a corresponding interactive element that controls the duration of the corresponding video segment represented by the interactive tile.

8. A computerized method, comprising: This results in the presentation of: (i) multiple interactive tiles representing multiple video segments and (ii) interactive elements that accept user input to toggle a visualization within each of the multiple interactive tiles, wherein the toggled visualization represents multiple detected features of a specified category that are detected from the corresponding video segment. Detecting interaction with one of the plurality of interactive tiles, the interactive tile representing one of the plurality of video segments; as well as In response to the detection of the interaction, at least a portion of the video clip is added to the selection of multiple video clips.

9. The computerized method of claim 8, wherein detecting the interaction with the interactive tile comprises: Detect drag operations that involve dragging the interactive tile to a panel representing the selection of the plurality of video clips, or detect the activation of the interactive element in the interactive tile.

10. The computerized method of claim 8, wherein detecting the interaction with the interactive tile comprises: Detecting a click or tap on an instance of one of the plurality of detected features in the interactive tiles representing the video segment; The context menu is activated in response to the click or tap. as well as The input of an option selected from the context menu is detected, the option being used to add a portion of the video segment corresponding to the instance of the detected feature to the selection.

11. The computerized method of claim 8, wherein the interactive tile includes a thumbnail representing a frame of the video segment, a portion of a script of the video segment, and a visualization of a plurality of detected faces in the video segment.

12. The computerized method of claim 8, wherein the interactive tile is configured to scan different thumbnails of frames of the video segment in response to an input that hovers over one or more portions of the interactive tile.

13. A computer system, comprising: One or more hardware processors and a memory configured to provide computer program instructions to the one or more hardware processors; A video interaction engine, configured to use the one or more hardware processors to perform operations, including: This results in the presentation of: (i) multiple interactive tiles representing multiple video segments of a selected video file or video editing project and (ii) interactive elements that accept user input to toggle a visualization within each of the multiple interactive tiles, wherein the toggled visualization represents multiple detected features of a specified category that are detected from the corresponding video segment; Detecting interaction with one of the plurality of interactive tiles, the interactive tile representing one of the plurality of video segments; and In response to the detection of the interaction, an operation involving the video segment is performed.

14. The computer system of claim 13, wherein the interactive element switches the plurality of interactive tiles to visualize detected faces or speakers, detected sounds, detected visual scenes, or detected visual artifacts detected from the corresponding video segment.

15. The computer system of claim 13, wherein the interactive tile is configured to scan different thumbnails of frames of the video segment in response to input hovering over one or more portions of the interactive tile.

16. The computer system of claim 13, wherein detecting the interaction with the interactive tile comprises: Detecting a click or tap on an instance of one of the plurality of detected features in the interactive tiles representing the video segment, wherein the operation involving the video segment includes navigating to a portion of the video segment in which the instance of the detected feature exists.

17. The computer system of claim 13, wherein detecting the interaction with the interactive tile comprises: The operation involves detecting a drag operation that drags the interactive tile to a panel representing a selection of multiple video segments, or detecting the activation of a corresponding interactive element in the interactive tile, wherein the operation involving the video segments includes adding at least a portion of the video segment to the selection of the multiple video segments.

18. The computer system of claim 13, wherein the interactive tile includes a thumbnail representing a frame of the video segment, a portion of a script of the video segment, and a visualization of a plurality of detected faces in the video segment.

Citation Information

Patent Citations

  • Segmentation and hierarchical clustering of video

    US11450112B2

  • Apparatus and software system for and method of performing a visual-relevance-rank subsequent search

    US20100070483A1

  • Video editing apparatus and method for guiding video feature information

    US20130236162A1

  • Home movie maker

    US6400378B1