Systems and method for automatically generating modified video content
The system addresses inefficiencies in conventional video editing by using machine learning to classify shot types and automatically generate modified video content, ensuring regions of interest are maintained, thus enhancing efficiency and optimizing video content for specific display formats.
Patent Information
- Application Number
- PCT/US2024/017356
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-09-04
AI Technical Summary
Conventional video editing tools fail to dynamically generate modified video content that optimally retains regions of interest throughout the duration of a video, especially when converting from landscape to portrait orientation, leading to inefficient manual adjustments and resource consumption.
A system utilizing machine learning models to classify shot types and automatically generate modified video content by cropping, recomposing, or preserving aspect ratios based on the type of frame sequences, ensuring regions of interest are maintained.
Automated video content generation improves efficiency and reduces resource consumption by dynamically adapting to changing regions of interest, providing optimized short-form video content for specific display formats.
Smart Images

Figure US2024017356_04092025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHOD FOR AUTOMATICALLY GENERATING MODIFIEDVIDEO CONTENTTECHNICAL FIELD
[0001] Aspects and implementations of the present disclosure relate to automatically generating modified video content.BACKGROUND
[0002] A platform (e.g., a content sharing platform) can allow users to upload, view, and share digital content such as media items. Media items can include audio clips, movie clips, music video, images and other multimedia content. For example, a user can generate a video (e.g., using a client device) and can provide the video to the platform (e.g., via the client device) to be accessible by other users of the platform. User can use computing devices (such as smart phones, cellular phones, laptop computers, desktop computers, tablet computers, televisions, and the like) to use, play, and / or otherwise consume the video and other media items (e.g., watch digital videos, and / or listen to digital music).SUMMARY
[0003] The below summary is a simplified summary of the disclosure in order to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is intended neither to identify key or critical elements of the disclosure, nor to delineate any scope of the particular implementations of the disclosure or any scope of the claims. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] In some implementations, a system and method are disclosed for automatically generating modified video content. In an implementation, a method includes receiving a request to generate modified video content from a video item. The method further includes determining a type of a sequence of frames of the video item, the type pertaining to one or more regions of interest in a plurality of sequences of frames of the video item. In some implementations, the sequence of frames includes a shot of the video item. The method further includes generating the modified video content from the video item with a display configuration based on the type.
[0005] In some implementations, to determine the type of the sequence of frames of the video item, the method includes providing information associated with the first shot as input into one or more artificial intelligence (Al) models (sometimes referred to as machineleaming models). The one or more Al models are trained to predict, based on the information associated with the sequence of frames, a probability score for each type of a plurality of sequence of frame types. In some implementations, the information associated with the sequence of frames comprises at least one of bounding boxes indicative of facial features, bounding boxes indicative of persons, bounding boxes indicative of text, or one or more embeddings. The method further includes obtaining a plurality of outputs from the one or more Al models. The plurality of outputs comprises a plurality of probability scores each indicating a likelihood that the sequence of frames belongs to a respective type of the plurality of sequence of frames types. The method further includes determining the type based on the plurality of probability scores.
[0006] In some implementations, the type of the sequence of frames is one of a plurality of sequence of frames types. Each of the plurality of sequence of frames types is associated with a distinct set of region of interest characteristics. The plurality of sequence of frames types comprises a first type indicating that the sequence of frames includes one region of interest with a size below a threshold, a second type indicating that the sequence of frames includes one region of interest with a size above the threshold, and a third type indicating that the sequence of frames comprises a plurality of regions of interest.
[0007] In some implementations, the method includes determining that the type of the sequence of frames is the first type. The method further includes cropping the sequence of frames to the one region of interest to generate the display configuration of the modified video content.
[0008] In some implementations, the method includes determining that the type of the sequence of frames is the second type. The method further includes maintaining an original aspect ratio of the sequence of frames in the display configuration of the modified video content.
[0009] In some implementations, the method includes determining that the type of the sequence of frames is the third type. The method further includes including at least a subset of the plurality of regions of interest within the display configuration of the modified video content.
[0010] In some implementations, to include at least the subset of the plurality of regions of interest within the display configuration of the modified video content, the method includes identifying a secondary region of interest of the plurality of regions of interest. The method further includes removing content associated with the secondary region of interest from each frame of the sequence of frames. The method further includes identifying aprimary region of interest of the sequence of frames with the removed content. The method further includes including the primary region of interest and secondary region of interest within the configuration of the modified video content.
[0011] In some implementations, to identify the secondary region of interest, the method includes determining temporal features for the each of the plurality of regions of interest. The temporal features comprise at least a size, a location, or an overlap of the plurality of regions of interest. The method further includes determining spatial features for each of the plurality of regions of interest. The spatial features indicate an amount of movement of the plurality of regions of interest between frames of the sequence of frames. The method further includes determining the secondary region of interest based on the temporal features and the spatial features. The secondary region of interest is a spatially and temporally static region of interest of a lesser size than other regions of interest of the plurality of regions of interest.
[0012] In some implementations, a system including a memory device and a processing devcie coupled to the memory device can perform operations described herein. For example, the processing devcie of the system can perform any of the above-described methods. In some implementations, a non-transitory computer-readable storage medium including instructions for a server that, when executed by a processing device, cause the processing device to perform any of the above-described methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Aspects and implementations of the present disclosure will be understood more fully from the detailed description given below and from the accompanying drawings of various aspects and implementations of the disclosure, which, however, should not be taken to limit the disclosure to the specific aspects or implementations, but are for explanation and understanding only.
[0014] FIG. 1 illustrates an example system architecture, in accordance with aspects and implementations of the present disclosure.
[0015] FIG. 2 illustrates an example user interface (UI) of an editor of a video item, in accordance with aspects and implementations of the present disclosure.
[0016] FIG. 3 is a block diagram illustrating automatically generating modified video content from a video item, in accordance with implementations of the present disclosure.
[0017] FIG. 4 illustrates an example shot classification vector for determining a shot type, in accordance with implementations of the present disclosure.
[0018] FIG. 5 illustrates stages of an operation pipeline for determining secondary regions of interest of a video shot, in accordance with aspects and implementations of the present disclosure.
[0019] FIG. 6 illustrates a flow diagram of an example method of automatically generating modified video content from a video item, in accordance with aspects and implementations of the present disclosure.
[0020] FIG. 7 illustrates an example training engine for training and deployment of a deep neural network, in accordance with aspects and implementations of the present disclosure.
[0021] FIG. 8 is a block diagram illustrating an exemplary computer system, in accordance with aspects and implementations of the present disclosure.DETAILED DESCRIPTION OF THE DRAWINGS
[0022] Aspects of the present disclosure generally relate to automatically generating modified video content from a video item that is adapted for display in a particular format such as a short-form format. A platform (e.g., a content sharing platform, etc.) can enable a user to access a media item (e.g., a video item, an audio item, etc.) provided by another user of the content sharing platform (e.g., via a client device connected to the content sharing platform). For example, a client device associated with a first user, such as a content creator, of the content sharing platform can transmit the media item to the content sharing platform via a network. A client device associated with a second user of the content sharing platform can transmit a request to access the media item and the content sharing platform can provide the client device associated with the second user with access to the media item (e.g., by transmitting the media item to the client device associated with the second user, etc.) via the network.
[0023] Content sharing platforms typically allow users to upload media items (e.g., videos). When the video is uploaded to a content sharing platform, it can be uploaded in various orientations such as landscape, portrait, square, etc. with corresponding aspect ratios. Content sharing platforms can support media items with different aspect ratios, allowing users to upload media items in a suitable orientation. For example, users can upload videos to a content sharing platform in a landscape orientation as it matches an aspect ratio of most computer screens, televisions, and mobile devices when held horizontally. Landscape videos have a wider width than height and typically have a 16:9 aspect ratio. The aspect ratio refers to the proportional relationship between the width and the heights of a video frame. Forexample, a 16:9 aspect ratio indicates that the width of the video frame is 16 units, and the height is 9 units.
[0024] In many instances, users of the content sharing platform can repurpose content uploaded to the content sharing platform in a landscape orientation (16:9 aspect ratio) to a portrait orientation (e.g., aspect ratio 9: 16) to upload to another content sharing platform (e.g., a short-form content sharing platform) or to reupload to the same content sharing platform in the modified orientation. A short-form content sharing platform can refer to a platform that focuses on and / or supports brief, concise, and quickly consumable content, often with a limited duration (e.g., less than 60 seconds in length). Short-form content sharing platforms usually display videos in a portrait format (e.g., aspect ratio 9: 16) and are often designed for content viewing via a mobile device (e.g., a smartphone), where users typically hold their phone vertically while creating and consuming content. Short-form content sharing platforms are becoming increasingly popular as they capture user preferences for quick and easily-consumable digital content. In some instances, users may want to recompose multiple regions of interest of a long-form video into a short-form version of the video. For example, a long-form gaming video may include one region of interest depicting exciting and visually appealing moments of gameplay and another region depicting reactions and commentary from the content creator. Content creators may want to repurpose such long-form gaming videos into a short-form, portrait format including both regions of interest for transmission and upload to a short-form video conference platform.
[0025] Conventional video editing tools can allow users to modify various characteristics of a video item such as video orientation, an aspect ratio, a composition of visual elements, a frame location, and the like to generate modified video content. For example, some video editing tools can allow users to modify an aspect of a video item from a first aspect ratio (e.g., 16:9 aspect ratio) for viewing in a landscape orientation to a second aspect ratio (e.g., 9: 16 aspect ratio) for viewing in a portrait orientation. However, many of these tools utilize a manual process to allow users to modify video content. For example, such tools can present a video item to a user of the tool within a complex user interface (UI). The user can draw one or more adjustable boxes over a portion of the video item and manually interact (e.g., using an input device such as a touch screen, a cursor control device such as a mouse, or the like) with the box to select a desired region of the video item with a desired second aspect ratio. The tool can generate a modified video item according to the portion of the video item occupied by the adjustable box. Users can interact with the adjustable box to ensure that regions of interest of the video item are maintained in the modified video item. Regions ofinterest can refer to specific regions within the video item that stand out from their surroundings and are likely to capture a viewer’s attention or piques a viewer’s interest. However, the number and location of regions of interest of the video item can constantly change throughout the duration of the video. Accordingly, users can typically parse the video on a per-frame basis and modify the adjustable boxes for each frame of the video item to ensure changing regions of interest are retained in the modified version of the video item. This can burden users with additional tasks and require additional computing resources to support these tasks.
[0026] Some video editing tools can automatically generate modified video content from a video item and produce short-form versions of the video item without using user-adjustable boxes. However, such tools fail to dynamically generate modified video content based on changes to the number and location of regions of interest throughout the duration of the video. For example, a first portion of a video item may include a primary region of interest depicting gameplay footage including multiple pertinent visual elements (e.g., a game interface, characters, environments, and visual elements directly related to gameplay), and a second portion of the video item may include the same primary region of interest but with an added secondary region of interest displaying a content creator’s reaction to the gameplay footage reflected in the primary region of interest. Conventional video editing tools typically fail to distinguish between video portions with different numbers of regions of interest, when generating modified video content. For example, a user can request the video item to be modified from a landscape orientation (e.g., 16:9 aspect ratio) to a portrait orientation (e.g., 9: 16 aspect ratio) for optimal viewing on a short-form content sharing platform. In response to the request, a conventional video editing tool can automatically crop the primary region of interest of both the first portion and the second portion captured in landscape orientation (e.g., 16:9 aspect ratio) and produce a portrait orientation (e.g., 9: 16 aspect ratio), short-form version of the video. However, such a cropping of the video fails to include the secondary region of interest in the second portion that is separate from the primary region of interest. Thus, the user of the tool must manually adjust the modified video in portrait orientation to ensure the secondary region of interest is retained within frames of the short-form version video corresponding to the second portion of the original video. Modifying video content using existing tools can therefore be challenging and can provide modified video content that is not optimized for a particular display format. This can cause frustration for the user and further burden the user with additional tasks that unnecessarily consume computingresources, thereby decreasing overall efficiency and increasing overall latency of the platform.
[0027] Aspects and implementations of the present disclosure address the above and other deficiencies by automatically generating modified video content from a media item, such as a video item, while ensuring that regions of interest with characteristics changing throughout the original media item are reflected in the modified version of the media item according to those characteristics. In some implementations, the modified video content can be automatically generated from the media item based on different types of frame sequences in the media item. For example, a video item can consist of multiple sequences of frames, where one sequence of frames can be captured without interruption in a single shot, for example when a user is recording an event from a first angle, and a next sequence of frames can be captured without interruption in a next single shot, for example when a user is recording an event from a second angle. Different sequences of frames (e.g., different shots) can be of different types depending on different region-of-interest characteristics. That is, a type of a sequence of frames (e.g., a shot type) can refer to a type defined by a distinct set of region-of-interest characteristics in the respective sequence of frames (e.g., a respective shot). The distinct set of region-of-interest characteristics may include, for example, a distinct number of regions of interest in the shot, a distinct size of a region(s) of interest in the shot, and / or the like. For example, a first shot of a video item may be of a first shot type with a primary region of interest depicting gameplay footage including multiple pertinent visual elements (e.g., a game interface, characters, environments, and visual elements directly related to gameplay), and a second shot of the video item may be of a second shot type with the same primary region of interest as the first shot but with an added secondary region of interest displaying a content creator’s reaction to the gameplay footage displayed in the primary region of interest.
[0028] A content sharing platform can allow users to upload, consume, share, search for, comment on, and otherwise engage with media items. A platform, such as content sharing platform, can allow users to request that a modified video content from video items and be uploaded to the platform. In some embodiments, the platform can provide a user interface (UI) for presentation on a client device to allow a user to request that modified video content be automatically generated from a video item for uploading and sharing on the platform. For example, in response to a user interaction with a UI element of the UI, the platform can automatically modify an aspect ratio of the video item from a 16:9 aspect ratio (e.g., for landscape orientation viewing) to a 9: 16 aspect ratio (e.g., for portrait orientation viewing)while ensuring shots of different types are handled differently to create a visually pleasing modified video content. The platform can present the generated modified video item within the UI provided for display on the client device for the user to upload, share, view, etc. In some embodiments, the modified video item can be a short-form version of the video item. A short-form version of a video item can refer to a brief, concise, and quickly consumable video clip of limited duration (e.g., less than 60 seconds in length). Such a video clip can be intended for content viewing via, for example, a mobile device (e.g., a smartphone) that is frequently held vertically by users, necessitating its conversion into a portrait format (e.g., aspect ratio 9: 16) for better use of screen space and improved viewing experience of users.
[0029] The platform can classify shots as one of multiple shot types during modified video content generation. For example, shot types can include shot types suitable for cropping, shot types suitable for aspect ratio preservation, and shot types suitable to be recomposed. A shot type suitable for cropping may be a type of a shot that only contains a single, central region of interest that does not occupy a significant portion of the entire frame of the shot (e.g., has a size below a threshold). A shot type suitable for aspect ratio preservation can be a type of a shot that contains a single region of interest that occupies a significant portion of the horizontal width of the shot (e.g., has a horizontal width above a threshold) which would be lost if the aspect ratio of the shot were changed.
[0030] A type of a shot suitable to be recomposed can be a type of a shot that contains two or more distinct regions of interest that can be individually cropped and recomposed to preserve important visual content of these distinct regions of interest. For example, the shot of such type may be a shot of frames of a gaming video stream with a primary region of interest and a secondary region of interest. The primary region of interest may include one or more pertinent visual elements associated with gameplay (e.g., a game stream showcasing the in-game environment, characters, actions, etc.) and may allow viewers to follow the progress of the game. The secondary region of interest, for example, can include a visual representation of a video stream of the player or of another individual viewing and reacting to content depicted in the gameplay primary region of interest. For example, the secondary region of interest can capture the individual’s face, reactions, and a portion of the individual’s surrounding environment.
[0031] In some embodiments, the platform can automatically generate modified video content from a video item with a display configuration based on one or more shot types. For example, in response to a determination that a first shot of a video item is of a shot type suitable to be recomposed, the platform can individually crop both the primary region ofinterest and the secondary region of interest and reassemble (e.g., recompose) the shot to include both the primary region of interest and secondary region of interest in a display configuration of the modified video content. In response to a determination that a second shot of the video item is of a shot type suitable for cropping, the platform can automatically crop the single region of interest in the display configuration of the modified video content. In response to a determination that a third shot of the video item is of a shot type suitable for aspect ratio preservation, the platform can ensure the display configuration of the modified video content maintains the original horizontal width of the video, which would be lost if the aspect ratio of the video were changed.
[0032] In some embodiments, the platform can utilize one or more machine learning models to determine shot types. The machine learning models can be trained to predict, based on a given shot, a shot type that would determine a layout / display configuration of the shot when generating modified video content. In some embodiments, the one or more machine learning models can receive input signals relevant to a shot of the video item as input and provide probability scores for each shot type representing how the shot should be modified. The input signals can include per-frame detection signals (e.g., bounding boxes indicating facial features, bounding boxes indicating persons, bounding boxes indicating text, etc.), frame-level embedding, video-level embeddings, and / or video / shot metadata.
[0033] Automatically generating modified video content from video items based on shot types improves generation of modified video content adapted to a particular device and / or display configuration and improves an overall user experience with the content sharing platform as users can easily upload modified video content, such as short-form versions of long-form videos, which are especially useful for small-screen devices (e.g., mobile phones). In addition, aspects and implementations of the present disclosure result in more efficient use of processing resources by automatically generating modified video content from a video item upon a user request rather than the users repeatedly modifying video items manually to capture regions of interest on a per-frame basis, thereby avoiding unnecessary consumption of resources to support iterative manual modifications to obtain suitable modified video content from a video item.
[0034] Although the description herein often refers to video items as an example type of media item, it is appreciated that aspects and implementations of the present disclosure can apply to other types of media items such as images, audio, and other multi-media without deviating from the scope of the present disclosure. Additionally, although the description herein often refers to classifying shot types, it is appreciated that aspects and implementationsof the present disclosure can generally apply to any sequence of frames of a video regardless of whether the sequence of frames corresponds to a shot of the video item.
[0035] FIG. 1 illustrates an example system architecture 100, in accordance with implementations of the present disclosure. The system architecture 100 (also referred to as “system” herein) includes client devices 102A through 102N (referred to generally as “client device(s) 102” herein), a data store 110, a platform 120, and / or a server machine 150 each connected to a network 108. In implementations, network 108 can include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., Ethernet network), a wireless network (e.g., an 802.11 network or a Wi-Fi network), a cellular network (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or a combination thereof.
[0036] In some embodiments, platform 120 can be a content sharing platform that allows users to consume, upload, share, search for, approve of (“like”), dislike, and / or comment on media items 112. Platform 120 can include a website (e.g., a webpage) or application back- end software used to provide a user with access to media items 112 (e.g., via client devices 102). A media item 112 can be consumed via the Internet or via a mobile device application, such as a content viewer 103 A through 103N (referred to generally as “content viewer(s) 103” herein) of client device 102. In some embodiments, a media item 112 can correspond to a media file (e.g., a video file, and audio file, etc.). In other or similar embodiments, a media item 112 can correspond to a portion of a media file (e.g., a portion or a chunk of a video file, an audio file, etc.). As discussed previously, a media item 112 can be requested for presentation to users of the platform by a user of the platform 120. As used herein, “media,” media item,” “online media item,” “digital media,” “digital media item,” “content,” and “content item” can include an electronic file that can be executed or loaded using software, firmware or hardware configured to present the digital media item to an entity. In one implementation, the platform 120 can store the media items 112 using the data store 110. In another implementation, the platform 120 can store media item 112 or fingerprints as electronic files in one or more formats using data store 110. Platform 120 can provide media item 112 to a user associated with a client device (e.g., client device 102A) by allowing access to media item 112 (e.g., via a content sharing platform application), transmitting the media item 112 to the client device 102A, and / or presenting or permitting presentation of the media item 112 at a content viewer 103 A of client device 102 A.
[0037] In some embodiments, media item 121 can be a video item. A video item refers to a set of sequential video frames (e.g., image frames) representing a scene in motion. Forexample, a series of sequential video frames can be captured continuously or later reconstructed to produce animation. Video items can be provided in various formats including, but not limited to, analog, digital, two-dimensional and three-dimensional video. Further, video items can include movies, video clips, video streams, or any set of images (e.g., animated images, non-animated images, etc.) to be displayed in sequence. In some embodiments, a video item can be stored (e.g., at data store 110) as a video file that includes a video component and an audio component. The video component can include video data that corresponds to one or more sequential video frames of the video item. The audio component can include audio data that corresponds to the video data.
[0038] In some implementations, data store 110 is a persistent storage that is capable of storing data as well as data structures to tag, organize, and index the data. Data can include audio data and / or video data, in accordance with embodiments described herein. Data store 110 can be hosted by one or more storage devices, such as main memory, magnetic or optical storage based disks, tapes or hard drives, NAS, SAN, and so forth. In some implementations, data store 110 can be a network-attached file server, while in other embodiments data store 110 can be some other type of persistent storage such as an object-oriented database, a relational database, and so forth, that can be hosted by platform 120 or one or more different machines (e.g., server machines 130-160) coupled to the platform 120 via network 108. Data store 110 can include a media cache that stores copies of media items that are received from the platform 120. In one example, media item 112 can be a file that is downloaded from platform 120 and can be stored locally in media cache. In another example, media item 112 can be streamed from platform 120 and can be stored as an ephemeral copy in memory of one or more of server machine 130-160.
[0039] The client devices 102 can each include computing devices such as personal computers (PCs), laptops, mobile phones, smart phones, tablet computers, netbook computers, network-connected televisions, etc. In some implementations, client devices 102 can also be referred to as “user devices.” Each client device 102 can include a content viewer 103. The content viewer 103 can include a web browser and / or the client application to present, on a client device 102, a user interface (UI) 124 A through 124N (referred to generally as “UI(s) 124” herein) for users to view or upload content, such as images, video items, web pages, documents, etc. For example, the content viewer 103 can be a web browser that can access, retrieve, present, and / or navigate content (e.g., web pages such as Hyper Text Markup Language (HTML) pages, digital media items, etc.) served by a web server. The content viewer 103 can render, display, and / or present the content to a user. The contentviewer 103 can also include an embedded media player (e.g., a Flash® player or an HTML5 player) that is embedded in a web page (e.g., a web page that can provide information about a product sold by an online merchant). In another example, the content viewer 103A-N can be a standalone application (e.g., a mobile application or app) that allows users to view digital media items (e.g., digital video items, digital images, electronic books, etc.). According to aspects of the disclosure, the content viewer 103 can be a content platform application for users to record, edit, and / or upload content for sharing on platform 120. As such, the content viewers 103 and / or the UIs 124 associated with the content viewers 103 can be provided to client devices 102 by platform 120. In one example, the content viewers 103 can be embedded media players that are embedded in web pages provided by the platform 120.
[0040] Platform 120 can include multiple channels (e.g., channels A through Z). A channel can include one or more media items 121 available from a common source or media items 121 having a common topic, theme, or substance. Media item 121 can be digital content chosen by a user, digital content made available by a user, digital content uploaded by a user, digital content chosen by a content provider, digital content chosen by a broadcaster, etc. For example, a channel X can include videos Y and Z. A channel can be associated with an owner, who is a user that can perform actions on the channel. Different activities can be associated with the channel based on the owner’s actions, such as the owner making digital content available on the channel, the owner selecting (e.g., liking) digital content associated with another channel, the owner commenting on digital content associated with another channel, etc. The activities associated with the channel can be collected into an activity feed for the channel. Users, other than the owner of the channel, can subscribe to one or more channels in which they are interested. The concept of “subscribing” can also be referred to as “liking,” “following,” “friending,” and so on.
[0041] In some embodiments, system 100 can include one or more third party platforms (not shown). In some embodiments, a third party platform can provide other services associated with media items 121. For example, a third party platform can include an advertisement platform that can provide video and / or audio advertisements. In another example, a third party platform can be a video streaming service provider that provides a media streaming service via a communication application for users to play videos, TV shows, video clips, audio, audio clips, and movies, on client devices 102 via the third party platform.
[0042] In some embodiments, a client device 102 can transmit a request to platform 120 for access to a media item 121. Platform 120 can identify the media item 121 of the request (e.g., at data store 110, etc.) and can provide access to the media item 121 via the UIs 124 ofthe content viewers 103 provided by platform 120. In some embodiments, the requested media item 121 can have been generated by another client device 102 connected to platform 120. For example, client device 102A can generate a video item (e.g., via an audiovisual component, such as a camera, of client device 102 A) and provide the generated video item to platform 120 (e.g., via network 108) to be accessible by other users of the platform. In other or similar embodiments, the requested media item 121 can have been generated using another device (e.g., that is separate or distinct from client device 102A) and transmitted to client device 102A (e.g., via a network, via a bus, etc.). Client device 102A can provide the video item to platform 120 (e.g., via network 108) to be accessible by other users of the platform, as described above. Another client device, such as client device 102B, can transmit the request to platform 120 (e.g., via network 108) to access the video item provided by client device 102A, in accordance with the previously provided examples.
[0043] In some embodiments, platform 120 can include a user interface (UI) manager 161. The UI manager 161 can generate and provide a UI for display at a client device 102. The UI manager can provide a UI to allow users of the platform 120 to edit media items 112 using one or more UI elements included within the UI. For example, the UI manager 161 can provide a UI for display at a client device 102 that includes a UI element to request that modified video content be automatically generated from a media item 121. Responsive to receiving an indication of a user interaction with the UI element of the displayed UI, UI manager 161 can transmit a command causing modified media content generator 151 to automatically generate modified video content from the media item 121, as described in greater detail below. Server machine 160 can additionally or alternatively include UI manager 161.
[0044] Training data generator 131 (e.g., residing at server machine 130) can generate training data to be used to train artificial intelligence (Al) models 160A-N. Models 160A-N can include Al models used or otherwise accessible to modified media content generator 151. In some embodiments, training data generator 131 can generate the training data based on video frames of training media items and / or training images (e.g., stored at data store 110 or another data store connected to system 100 via network 108) and / or data associated with one or more client devices that accessed the training media items.
[0045] Server machine 140 can include a training engine 141. Training engine 141 can train Al models 160A-N using the training data from training data generator 131. In some embodiments, the Al models 160A-N can refer to model artifacts created by the training engine 141 using the training data that includes training inputs and corresponding target outputs(correct answers for respective training inputs). The training engine 141 can find patterns in the training data that map the training input to the target output (the answer to be predicted), and provide the machine learning models 160A-N that captures these patterns. The machine learning models 160A-N can be composed of, e.g., a single level of linear or non-linear operations (e.g., a Convolutional Neural Network (CNN) or other deep network, e.g., a machine learning model that is composed of multiple levels of non-linear operations). An example of a deep network is a neural network with one or more hidden layers, and such a machine learning model can be trained by, for example, adjusting weights of a neural network in accordance with a backpropagation learning algorithm or the like. In other or similar embodiments, the Al models 160A-N can refer to model artifacts that are created by training engine 141 using training data that includes training inputs. Training engine 141 can find patterns in the training data, identify clusters of data that correspond to the identified patterns, and provide the Al models 160A-N that captures these patterns. Al models 160A-N can use one or more of clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, k-nearest neighbor algorithm (k-NN), linear regression, random forest, neural network (e.g., artificial neural network), a boosted decision forest, etc. In some embodiments, the Al models 160A-N include one or more generative Al models. A generative Al model can deviate from a machine learning model based on the generative Al model’s ability to generate new, original data, rather than making predictions based on existing data patterns. A generative Al model can include a generative adversarial network (GAN), a variational autoencoder (VAE), a large language model (LLM), or a diffusion model. In some instances, a generative Al model can employ a different approach to training or learning the underlying probability distribution of training data, compared to some machine learning models. For instance, a GAN can include a generator network and a discriminator network. The generator network attempts to produce synthetic data samples that are indistinguishable from real data, while the discriminator network seeks to correctly classify between real and fake samples. Through this iterative adversarial process, the generator network can gradually improve its ability to generate increasingly realistic and diverse data.
[0046] In some embodiments, one or more Al models 160A-N can be trained to predict, based on a given image or frame, such as a frame of a media item 121, bounding boxes for the given frame that indicate regions of interest of the given frame. In some embodiments, one or more of the Al models 160A-N can be trained to predict bounding boxes of a certain type. For example, Al model 160A can be an object detection model to predict bounding boxes that indicate objects. Al model 160B can be a face detection model that predictsbounding boxes that indicate facial features. Al model 160C can be a boundary detection model to predict shot boundaries. In some embodiments, one or more of the Al models 160A- N can be trained to predict a shot type. For example, Al model 160D can be a shot classification to predict shot types. In some embodiments, one or more of the Al models 160A-N can be trained to predict bounding boxes associated with regions of interest. For example, Al model 160E can be a primary region of interest detector to predict bounding boxes that indicate primary regions of interest. Al model 160F can be a secondary region of interest detector to predict bounding boxes associated with a secondary region of interest. The models may be trained to predict the various outputs based on labelled training data. For example, a model may be trained to predict bounding boxes of a certain type based on training example frames with labelled bounding boxes of the certain type.
[0047] Server machine 150 can include a modified media content generator 151. Modified media content generator 151 can dynamically (e.g., immediately or not more than a threshold number of seconds after a modification request) generate modified media content based on one or more regions of interest of a media item 121. Modified media content generator 151 can provide one or more shots of media item 112 as input to one or more trained Al models 106A-N. The modified media content generator 151 can generate modified video content, such as short-form versions of the media item 121 based on the outputs of the Al models indicating shot types. Modified media content generator 151 can crop frames of the one or more shots of media item 121 to ensure regions of interest are present within the cropped frames based on the shot type, as described in detail below with respect to FIG. 3-5.
[0048] It should be noted that although FIG. 1 illustrates modified media content generator 151 and UI manager 161 as part of platform 120, in additional or alternative embodiments, modified media content generator 151 and UI manager 161 can reside on one or more server machines that are remote from platform 120 (e.g., server machine 150, server machine 160). It should be noted that in some other implementations, the functions of server machines 130, 140, 150, 160 and / or platform 120 can be provided by a fewer number of machines. For example, in some implementations, components and / or modules of any of server machines 130, 140, 150, 160 can be integrated into a single machine, while in other implementations components and / or modules of any of server machines 130, 140, 150, 160 can be integrated into multiple machines. In addition, in some implementations, components and / or modules of any of server machines 130, 140, 150, 160 can be integrated into platform 120.
[0049] In general, functions described in implementations as being performed by platform 120 and / or any of server machines 130, 140, 150, 160 can also be performed on the client devices 102 in other implementations. In addition, the functionality attributed to a particular component can be performed by different or multiple components operating together. Platform 120 can also be accessed as a service provided to other systems or devices through appropriate application programming interfaces, and thus is not limited to use in websites.
[0050] Although implementations of the disclosure are discussed in terms of platform 120 and users of platform 120 accessing a video item, implementations can also be generally applied to media items generally. Further, implementations of the disclosure are not limited to content sharing platforms that allow user to generate, share, view, and otherwise consume media items such as video items.
[0051] In implementations of the disclosure, a “user” can be represented as a single individual. However, other implementations of the disclosure encompass a “user” being an entity controlled by a set of users and / or an automated source. For example, a set of individual users federated as a community in a social network can be considered a “user.” In another example, an automated consumer can be an automated ingestion pipeline of platform 120.
[0052] Further to the descriptions above, a user can be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein can enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content or communications from a server. In addition, certain data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity can be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location can be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user can have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0053] FIG. 2 illustrates an example user interface (UI) 200 of an editor of a video item, in accordance with some embodiments of the present disclosure. Users of a content sharing platform (e.g., platform 120 of FIG. 1) can interact with the UI 200 to request generation of modified video content from media items (e.g., media items 121 of FIG. 1). The UI 200 canbe generated by one or more processing devices of server machine 130, server machine 140, server machine 150, and / or server machine 160. In some embodiments, the UI 200 can be generated by user interface (UI) manager 161 of FIG. 1, for presentation at a client device (e.g., client devices 102A-102N.) The UI manager 161 can provide the UI 200 to enable users of the content sharing platform to generate a modified video content from a media item based on a user interaction with the UI 200. In some embodiments, the UI 200 can be generated by a client device (e.g., content viewer application 103 on client device 102).
[0054] UI 200 can include multiple regions, including a region 210 to display one or more media items (e.g., video items) corresponding to video data captured and / or streamed by client devices, such as client devices 102A-102N of FIG. 1. The UI 200 can further include a region 220 for providing selectable layouts and options to adjust a composition of the video item. The region 220 includes a UI element 222, a UI element 224, a UI element 226, a UI element 228, and a UI element 230. Responsive to a user interaction with one of the UI elements displayed with the region 220, one or more components (e.g., modified media content generator 151 of FIG. 1) of the platform can generate modified video content from the video item. A UI manager, such as UI manager 161 of FIG. 1, can provide the modified video content for display at region 210 of the UI 200.
[0055] In an illustrative example, in response to a user interaction with UI element 222 of the region 220, a processing device (e.g., modified media content generator 151 of FIG. 1) can automatically generate modified video content from the video item. To automatically generate modified video content from the video item, the processing device can analyze one or more shots of the video item to determine shot types and modify the one or more shots based on one or more shot types. The processing device can automatically generate modified video content of the video item while maintaining focus on one or more regions of interest indicated by the shot type. In some embodiments, the processing device can modify an aspect ratio of the video item to an aspect ratio suitable for a short-form video. For example, the processing device can adjust the aspect ratio of the media item from a 16:9 aspect ratio (landscape) to a 9: 16 aspect ratio (portrait). In some embodiments, to determine shot type, the processing device can provide one or more shots of the video item as input to one or more Al models. The processing device can analyze outputs of the one or more Al models, to determine shot types for shots of the video item and generate one or more layouts of modified video content based on predicted shot types, as described in detail below with respect to FIG. 3. The UI 200 can be updated to display the modified video content at the region 210.
[0056] In some embodiments, the region 210 can be split to display multiple regions of the video item. For example, the region 210 can be split to display a first portion of the video item at an upper portion of the region 210 and a second portion of the video item at a lower portion of the region 210. In an illustrative example, a first area of a video item can be dedicated to display a view of a content creator captured via a webcam, a camera, or the like. A second area of the video item can be dedicated to display other visual content shared by the content creator such as a gameplay footage, content produced by another content creator, or the like. The processing device can identify a first region of interest associated with the first area of the video item and second region of interest associated with the second area of the video item. The processing device can modify a composition of the first area and the second area according to the first region of interest and the second region of interest, and obtain a first modified area of the video item and a second modified area of the video item. The first modified area of the video can be provided for display at the lower portion of the region 210 and the second modified area of the video can be provided for display at the upper portion of the region 210.
[0057] In some embodiments, UI element 224, UI element 226, UI element 228, and UI element 230 can be interactive UI elements that, in response to a user interaction, provide one or more pre-defined layouts of region 210 for a user to manually adjust and display a modified video item. For example, in response to a user interaction with UI element 224, the processing device can modify an aspect ratio of the video item from a first aspect ratio (e.g., 16:9) to a second aspect ratio (e.g., 9: 16). To fit the narrower second aspect ratio, the processing device can remove content from the sides of the video, thereby reducing the width of the video to fit within the second aspect ratio. In some embodiments, the processing device can provide a window over an area of the video where content within the window is retained and content outside of the window is removed. In some embodiments, the window is a static window. For example, the processing device can provide a static window that remains at a center portion of the video item for each frame of the video. In some embodiments, the window can be presented within the region 210 of the video UI 200 and a user can manually modify (e.g., using an input device such as a touch screen, a cursor control device such as a mouse, or the like) the window to capture a desired region of the video item for displaying within region 210 on a per-frame basis in order to ensure regions of interest are maintained in the modified composition of the video item.
[0058] In another example, in response to a user interaction with UI element 226, the processing device can cause an upper region of the region 210 to display a first window of afirst area of the video item and a lower region of the region 210 to display a second window of the second area of the video item. In some embodiments, a user can manually adjust (e.g., using an input device such as a touch screen, a cursor control device such as a mouse, or the like) the first window and the second window to capture a desired first area and second area of the video item. The UI element 228 and the UI element 230 can provide similar functionality using different predefined layouts. For example, in response to a user interaction with UI element 226, the processing device can cause an upper region of the region 210 to display the first window of the first area of the video item and lower region of the region 210 to display the second window of the second area of the video item. In response to a user interaction with UI element 230, the processing device can cause an upper region of the region 210 to display a first window of a first area of the video item and lower region of the region 210 to display a second window of the second area of the video item, where the upper region is larger than the lower region.
[0059] FIG. 3 is a block diagram 300 illustrating automatically generating modified video content from a video item, in accordance with implementations of the present disclosure. A shot boundary detector 302 can receive a video item 301, such as media item 121 of FIG. 1, as input. In some embodiments, the shot boundary detector 302 can use an Al model trained to predict shot boundaries. Shot boundary detection is a computer vision technique to identify transitions or boundaries between consecutive shots in a video sequence such as video item 301. A shot in the video item 301 refers to a continuous sequence of frames captured by a single camera without interruption. By accurately detecting shot boundaries, the video item 301 can be segmented into individual shots, allowing for further analysis of the individual shots. The shot boundary detector 302 can use the Al model trained on historical data such as videos with labeled shot boundaries. Visual and / or audio features can be extracted from each video frames of the historical videos that represent the content of each of the frames. The features (e.g., video and / or audio features) and corresponding labeled shot boundaries can be used to train the shot boundary detector 302. The trained Al model used by the shot boundary detector 302 can process each frame of the video item 301 and predict shots of the video item 301 based on learned patterns from the training data. The shot boundary detector 302 can transmit one or more shot boundaries 304 A through 304N (referred to generally as “shots 304” herein) of the video item 301 as input to a shot classifier 306. In some embodiments, the input of the Al model used by the shot boundary detector 302 can be a series of time-stamped video frames of the video item 301, and the output of the Almodel used by the shot boundary detector 302 can be a list of timestamps identifying shot boundaries 304 detected within the video item 301.
[0060] It is appreciated that various other shot boundary detection computer vision techniques and signal processing algorithms can be employed to determine shot boundaries without deviating from the scope of the present disclosure. For example, a pixel-based approach can calculate pixel metrics such as pixel intensity or pixel color to identify shot boundaries based on threshold differences of pixel metrics. In another example, a histogrambased analysis can compare a color distribution between consecutive frames to determine shot boundaries. In yet another example, a motion-based approach can analyze motion vectors between frames to identify shot boundaries. In still another example, temporal analysis can use various metrics such as mean squared error (MSE) or structural similarity index (SSIM) to quantify a similarity level or a difference level between consecutive frames indicative of a shot boundaries. In some embodiments, a CNN-based model may provide shot boundaries 304. It can be noted that one or more of the above-reference approaches can be used individually or in combination to identify shot boundaries of the video item 301.
[0061] The shot classifier 306 can receive shots 304 as input. The shot classifier 306 can identify video visual signals and metadata associated with a single shot 304. The shot classifier 306 can classify the shot as being associated with one of multiple shot types. In some embodiments, shot classifications can include a shot 304 classified as the type suitable for cropping, a shot 304 classified as the type suitable for aspect ratio preservation, and a shot 304 classified as the type suitable to be recomposed.
[0062] A shot type suitable for cropping may be determined for a shot 304 that only contains a single, central region of interest that does not occupy the entirety of the frame.
[0063] A shot type suitable for aspect ratio preservation may be determined for a shot 304 that contains important visual information that occupies the horizontal width of the shot 304 which would be lost if the aspect ratio of the shot were changed. For example, the shot 304 may include a single region of interest that occupies the entire horizontal width of the shot 304.
[0064] A shot type suitable to be recomposed can be determined for a shot 304 that contains two or more distinct regions of interest that can be individually cropped and recomposed to preserve the important visual content contained within the two or more distinct regions of interest. For example, the shot 304 may be a shot of a gaming video stream with a gameplay region of interest and a video stream region of interest. The gameplay region of interest may capture a rendering of an avatar of a video game controlled by a player (e.g., acontent creator). The gameplay region of interest can include, for example, a game stream showcasing the in-game environment, characters, and actions. The gameplay region of interest may generally include a primary region of interest in which viewers can follow the progress of the game. The gameplay region of interest can be captured via software, such as screen capture software, or hardware, such as a capture card.
[0065] The video stream region of interest can include a visual representation of a video stream, for example, of the player or of another individual viewing and reacting to content depicted in the gameplay region of interest. The video stream region of interest can capture the individual’s face, reactions, and a portion of the surrounding environment. The stream region of interest can be captured via a dedicated capture device such as dedicated camera or a webcam. It is appreciated that gameplay region of interest and video stream region of interest are used herein by way of example, and not by way of limitation, noting that a shot 304 classified as associated with the type suitable to be recomposed can be any shot that includes two distinct regions of interest.
[0066] The shot classifier 306 may leverage multiple signals to determine a classification of a shot 304, such as the shot 304 itself, bounding boxes (e.g., bounding boxes indicative of persons, facial features, text, etc. and associate confidence scores), embedding, and / or video metadata.
[0067] In some embodiments, a face detection model (not illustrated) can be coupled to receive the shot 304 as input and provide one or more bounding boxes associated with facial features to the shot classifier 306. The face detection model can be a Al model trained to predict (e.g., identify and locate) faces within a frame of a video item or an image. The face detection model 316 can be trained with a historical dataset of images and / or video frames. The historical dataset can be labeled with bonding boxes around facial features contained with the images / video frames of the historical dataset. After training and deployment, the face detection model can process unlabeled data such as frames of the shot 304 and predict presence and location of facial features within frames of the shot 304. The face detection model can identify one or more bounding boxes indicative of a location of one or more facial features within a respective frame of the shot 304 and provide them to the shot classifier 306.
[0068] In some embodiments, a person detection model (not illustrated) can be coupled to receive the shot 304 as input and provide one or more bounding boxes associated with persons to the shot classifier 306. A text detection model (not illustrated) can be coupled to receive the shot 304 as input and provide one or more bounding boxes associated with text to the shot classifier 306. The person detection model and the text detection model can betrained with historical dataset of images and / or video frames using similar methodologies described above with respect to the face detection model. The face detection model, the person detection model, and the text detection model can leverage various deep learning techniques, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and the like to learn complex patterns and features for face detection, person detection, and text detection, respectively.
[0069] In some embodiments, the shot classifier 306 may analyze bounding boxes provided by the above-described face, person, and text detectors to determine regions of interest. In some embodiments, multiple bounding boxes may be integrated into a single region of interest to provide a holistic view of pertinent visual elements within the shot 304 and information (e.g., width, height, location, etc.) associated with the regions of interest.
[0070] In some embodiments, the shot classifier 306 may utilize one or more embeddings associated with the shot 304 to classify the shot 304. Embedding may include frame-level embeddings and video-embeddings. Frame-level embeddings can refer to the representation of individual frames of the shot 304 as vectors in a high-dimensional space. For example, each frame of the shot 304 can be provided as input to a neural network, such as a CNN, trained to extract features and transform frames of the shot 304 into a numerical embedding. The numerical frame-level embedding can capture spatial information (e.g., colors, textures, patterns, etc.) present in each frame of the shot 304 and temporal information to capture motion and dynamic within the shot 304. Video-level embeddings can refer to the representation of the video item associated with the shot 304 as a single vector. Video-level embeddings can be generated by aggregating frame-level embeddings of the video using various techniques such as mean pooling, max pooling, RNNs, and the like. Video-level embeddings can provide a concise, mathematical representation of information present in a video.
[0071] The shot classifier 306 can utilize embeddings (frame-level and / or video-level embeddings) to identify and localize specific objects or entities within frames of the shot 304. By analyzing frame-level embeddings, the shot classifier can determine regions of interest across frames of the shot 304 and information (e.g., location, size, etc.) associated with the regions of interest. Frame-level embeddings can be useful for identifying subtle details and / or patterns that contribute to defining regions of interest of the shot 304.
[0072] In some embodiments, the classifier 306 can provide each of the shot 304, bounding boxes (e.g., indicative of faces, persons, or text), embeddings (e.g., frame-level, video-level, etc.), and / or metadata as input into a shot classification model trained to predict aprobability score for each of the multiple shot types (e.g., a shot type suitable for cropping, a shot type suitable for aspect ratio preservation, a shot type suitable to be recomposed, etc.).
[0073] In some embodiments, an untrained shot classification model can be trained using supervised learning. In at least one embodiment, the untrained shot classification model is trained in a supervised manner and processes input data (e.g., frame-level embeddings, shotlevel embeddings, video shots, bounding boxes, metadata, etc.) based on a training dataset and compares resulting outputs against a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through the untrained shot classification model. In some embodiments, a training framework can adjust weights that control the untrained shot classification model. In at least one embodiment, the training framework includes tools to monitor how well the untrained shot classification model is converging towards a as trained model, suitable to generating correct probability scores for one or more shot classifications. In some embodiments, the training framework trains the untrained shot classification model repeatedly while adjusting weights to refine an output of the untrained shot classification model using a loss function and adjustment algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework trains the untrained shot classification model until a desired accuracy is received. In at least one embodiment, the trained shot classification model may then be deployed to implement any number of Al operations.
[0074] In at least one embodiment, shot classifier 306 may provide bounding boxes (e.g., indicative of faces, persons, or text) associated with the shot 304, embeddings associated with the shot 204, metadata associated with the shot 304, or the shot 304 itself as input to the trained shot classification model. Shot classifier 306 may obtain one or more outputs of the trained shot classification model that may include probability scores for each shot classification (shot type). The probability scores may indicate a predicted probability that the shot 304 is associated with a corresponding shot classification. For example, the classifier 306 can obtain a probability score that indicates whether the shot 304 is of the type that is suitable for cropping, a probability score that indicates whether the shot 304 is of the type that is suitable for aspect ratio preservation, and a probability that indicates whether the shot 304 is of the type that is suitable to be recomposed.
[0075] Probability scores may indicate a likelihood that the shot 304 belongs to the corresponding classification. The classification model may be trained using shots labeled with a corresponding shot type. The labeled dataset may be used for training and evaluating the classification model. A machine learning algorithm, such as nearest neighbors, decision tree, random forest, and / or the like can be used for shot type classification. The classificationmodel may provide, as output, an N-bit vector representing the probability score for each of the shot classifications. For example, the classification model may output a 3 -bit vector representing probability scores for each of the shot type suitable for cropping, the shot type suitable for aspect ratio preservation, and the shot type suitable to be recomposed, as described below with respect to FIG. 4.
[0076] FIG. 4 illustrates an example shot classification vector 400, in accordance with implementations of the present disclosure. Specifically, FIG. 4 illustrates a shot classification vector 400 with an entry 402 corresponding to a probability score (0.0) that the shot is of the type suitable for cropping, an entry 404 corresponding to a probability score (0.1) that the shot is of the type suitable for aspect ratio preservation, and an entry 406 corresponding to a probability score (0.9) that the shot is of the type suitable to be recomposed. It is appreciated that the 3 -bit shot classification vector 400 is illustrated herein by way of example, and not by way of limitation, not that the techniques described herein can generally be applied to a N-bit shot classification vector, where N is the number of shot classifications.
[0077] Returning to FIG. 3, the shot classifier 306 can determine a shot classification based on the 3 -bit probability vector received as output from the trained shot classification model. In some embodiments, the shot classifier 306 can determine the shot classification (shot type) based on greatest predicted probability score. Using the shot classification vector 400 as an example, the shot classifier may determine the corresponding shot is of the type suitable to be recomposed as the entry 406 corresponding to that probability score (0.9) is greater than the probabilities score of the other entries of the shot classification vector. In some embodiments, the shot classifier 306 can determine the shot classification using a thresholding technique. For example, if the probability that the shot 304 is of the type suitable to be recomposed or the probability that the shot 304 is of the type suitable for cropping is greater than a threshold probability score of 0.9, the shot classifier 306 can determine the shot classification according to the greater of the two probability scores. The threshold probability score may be provided by a user (e.g., a developer). Otherwise, the shot classifier 306 can classify the shot 304 as being of the type suitable for aspect ratio preservation, and cause the shot 304 to maintain its original aspect ratio and frame layout.
[0078] Responsive to determining that the shot 304 is of the type suitable for cropping, the shot classifier 306 may provide the shot as input to a primary region of interest detector 308. In some embodiments, the primary region of interest detector 308 may use an Al model trained to predict, based on a given input shot (e.g., shot 304), multiple bounding boxes indicating a primary region of interest that can include one or more pertinent visual elements.Pertinent visual elements, for example, can include persons, person-like characters (e.g., player characters in a video game, non-player characters in a video game, etc.), text (e.g., captions, titles, etc.), objects (e.g., a vehicle in racing game, a projectile in a video game, etc.), facial features, and / or other regions of a shot that may be visually significant or important to a viewer of the shot. In some embodiments, the primary region of interest detector 308 can perform a union of the one or more predicted bounding boxes to determine a primary region of interest that includes the one or more detected bounding boxes. In some embodiments, the primary bounding box may be the same throughout the entirety of the shot 304. In some embodiments, the primary bounding box may be dynamically (e.g., for each frame of the 304) updated as bounding boxes indicating changing location of pertinent visual elements throughout the shot 304. The determined primary region of interest associated with the shots 304 classified as being of the type suitable for cropping can be provided to the layout generator. Accordingly, the layout generator 312 can receive a primary region of interest as input for shots 304 classified as being of the type suitable for cropping. Primary region of interest detector may, for example, be trained to detect primary regions of interest based on labelled training examples as described above.
[0079] Responsive to determining that the shot 304 is being of the type suitable to be recomposed, the shot classifier 306 can provide the shot as input to the secondary region of interest detector 310. In some embodiments, the secondary region of interest detector 310 may be an Al model trained to predict, based on a given input shot (e.g., shot 304), one or more bounding boxes indicative of regions of the shot that are secondary to a primary region of interest of the shot. A secondary region of interest may be defined as a spatially and temporally static region within the shot 304 that is supplementary to the primary region of interest. In some embodiments, a region that is supplementary to a primary region of interest is a region that occupies a lesser area than the primary region of interest. A region of interest is spatially and temporally static if the area of the region does not change over time and / or there are limited movements within that region. A region of interest is temporally static if the characteristics of the region do not change over a sequence of frames of the shot 304. For example, a secondary region of interest may include a region corresponding to a video stream from a content creator’s capture device imposed on top of a gameplay footage, a region depicting a news anchor in a newsroom, a region depicting a teacher in a classroom video, and the like. In some embodiments, the secondary region of interest detector 310 can be implemented according to an operation pipeline described below with respect to FIG. 5.
[0080] FIG. 5 illustrates an operation pipeline 500 for determining secondary regions of interest of a video shot, in accordance with aspects and implementations of the present disclosure. The operation pipeline 500 is a multi-stage pipeline including a face / person detection stage 502, a tracking stage 504, a filtering and global clustering stage 506, and a candidate scoring and ranking stage 508.
[0081] At the face / person detection stage 502, processing logic may receive one or more shots 512 as input, where the shots 512 include frames 514. The frames 514 of the shots 512 can be provided as input to a trained face and / or person Al model. The face detection model can be trained to predict bounding boxes corresponding to facial features in an image or video frame. The person detection model may be trained to predict bounding boxes corresponding to human features. Various deep learning architecture, such as Convolutional Neural Networks (CNNs) may be employed for face and person detection. In some embodiments, the face and person Al models can be trained using labeled historical datasets including images with annotated bounding boxes. The face and / or person Al can be trained according to various techniques described below with respect to FIG. 7. As illustrated at stage 502, one or more bounding boxes may be predicted for each of the frames 514.
[0082] At the tracking stage 504, processing logic may identify bounding boxes across the frames 514 that belong to the same person or object. In some embodiments, to identify and group bounding boxes, the processing logic may measure spatial overlap of bounding boxes between frames 514. For example, the processing logic may calculate an Intersection over Union (loU) of bounding boxes between frames 514 where the calculated loU value provides a measure of spatial overlap of bounding boxes between frames. In some embodiments, a threshold loU value may be set (e.g., by a developer, a user, etc.) to determine whether bounding boxes of adjacent frames are associated with the same person. If the loU value exceeds the set threshold, the bounding boxes are considered associated with the same person. In some embodiments, using the above-described loU thresholding technique, bounding boxes associated with the same person may be linked across multiple frames to form a bounding box track. The processing logic may assign a unique identifier for each bounding box track to allow the operation pipeline 500 to track associated bounding boxes across frames 514 and shots 512 of a video. In some embodiments, at the tracking stage 504, the processing logic may merge tracks of bounding boxes that have a similar width, height, and / or center into a single track. In the example illustrated with respect to FIG. 5, the processing logic may identify six different tracks in which the first track includesbounding boxes associated with a first person, the second track includes bounding boxes associated with a second person, and so forth.
[0083] At the filter and global cluster stage 506, the processing logic may filter bounding boxes within tracks and combine tracks from different shots 512 to obtain combined tracks (referred to as “clusters” herein). To filter bounding boxes within tracks, the processing logic may calculate the standard deviation of the width (oWidth), height (oHeight), a center x- coordinate (crCenterX), and a center y-coordinate (oCenterY) of bounding boxes within each track. The processing logic may filter (e.g., remove) bounding boxes with calculated standard deviations above a filtering threshold (e.g., as set by a developer, a user, etc.). To cluster tracks from different shots 512, the processing logic may group tracks from different shots that share a similar width, height, and / or center coordinate to form global clusters of tracks as candidates for selecting tracks that are most likely to be a secondary region of interest. For example, the processing logic may perform the above-described filtering and clustering operations to produce a cluster 516 and a cluster 518 as track candidates.
[0084] At the candidate scoring and ranking stage 508 of the operation pipeline, processing logic may extract features from cluster 516 and cluster 518 to determine candidate scores that indicate a likelihood that the clusters 516 and 518 are secondary regions of interest. In some embodiments, the processing logic may extract and combine features according to the following candidate scoring function (1):(1) Scores = X ■ W + b where, X are the combined (e.g., concatenated) features associated with bounding boxes of a given cluster, W are weights, and b is a bias. In some embodiments, tracks with bounding boxes that remain spatially and temporally static across frames and are of a lesser area (e.g., as compared to a primary region of interest) receive higher candidate scores.
[0085] The features X can include spatial features and / or temporal features for candidate track scoring. Spatial features may include a size, location, and overlap of bounding boxes included within the clusters. Size may include a width and height of bounding boxes. In some embodiments, width may be defined as follows: Pwidth = Ws / W , where Wsis the width of bounding boxes within the cluster and W is the width of the frame. In some embodiments, height may be defined as Pheight = Hs / H , where Hsis the height of bounding boxes with the cluster and H is the height of the frame. In some embodiments, the Pwidth and Pheight can be inversely correlated with candidate scores. For example, as Pwidth and Pheightincrease, the candidate score may decrease; and as Pwidth and Pheight decrease, the candidate score may increase.
[0086] Location may include coordinates of a center of the bounding boxes included within the cluster. In some embodiments, the location may include a Pxcomponent and a Pycomponent. The Pxcomponent may be calculated as follows: Px= xcenter / W , where xcenter is the x-coordinate of the center of the bounding boxes in the cluster and W is the width of the frame. The Pycomponent may calculated as follows: Py= ycenter / W , where yCenteris the y-coordinate of the center of the bounding boxes in the cluster and W is the width of the frame. In some embodiments, coordinates closer to the center of a respective frame may generally decrease a candidate score while coordinates closer to edges of the respective frame may generally increase a candidate score.
[0087] Overlap of bounding boxes may include two components: P^veriap and P20veriaP. The PpOveriap component may be calculated according to the follow equation: P1OVeriap=AW / Wp , where AW is the amount of width of overlap between the bounding boxes within a candidate cluster and a primary region of interest of the respective frame, and Wpis the width of the primary region of interest. The PsOveriaPcomponent may be calculated according to the follow equation: P2overiap=AW / WS, where AW is the amount of width of overlap between the bounding boxes within the candidate cluster and the primary region of interest of the respective frame, and Wsis the width of the secondary region of interest. In some embodiments, P ^veriap and P20veriaPmay be inversely correlated with candidate scores. For example, as overlap with the primary region of interest increases, the candidate score may decrease; and as overlap with the primary region of interest decreases, the candidate score may increase.
[0088] In some embodiments, temporal features may include the following components: Pshots and Pduration. Pshots may be calculated using the following equation: Pshots = Shotsp / ShotsT, where Shotspis the number of shots present in the candidate cluster, and ShotSTis the total number of shots in the video. In some embodiments, Pshots may be positively corelated with candidate scores. For example, a cluster with a higher number of shots present may obtain a higher score than a cluster with a lower number of shots presents. Pduration may be calculated using the following equation: Pduration = Secondsp / SecondsT, where Secondspis the number of seconds a track in the cluster is present within the video and Secondsps the total duration of the video. In some embodiments, Pduration may be positively correlated with a candidate score. For example, acluster with a greater number of seconds present within the video may obtain a higher candidate score than a cluster with a lesser number of second present within the video.
[0089] In some embodiments, the weights (W) and bias (b) of the candidate scoring function can be determined during a training process to determine weights and biases that allow the candidate scoring function to make accurate predictions on new data, such as shots 512. The training process may involve a loss function to measure the difference between predictions and true labeled historical data. In some embodiments, the labeled historical data may include historical shots from a video with person tracks labeled as secondary regions of interest. In some embodiments, the labeled historical data may include historical shots from a video with labeled with binary information based on whether a secondary region of interest is present.
[0090] At the candidate scoring stage 508, processing logic may determine whether the candidate clusters are secondary regions of interest based on scores calculated according to the candidate scoring function (1). In some embodiments, processing logic may determine one or more secondary regions of interest using a threshold score. For example, processing logic may determine that cluster 516 and cluster 518 are secondary regions of interest responsive to a determination that the candidate scores associated with cluster 516 and cluster 518 exceed a threshold candidate score. In some embodiments, processing logic may determine a single secondary region of interest based on a maximum score. For example, processing logic may determine that cluster 516 is a secondary region of interest responsive to a determination that the score associated with cluster 516 is the greater than the score associated with cluster 518.
[0091] Returning to FIG. 3, responsive to determining one or more secondary regions of interest, the one or more secondary regions of interest may be cropped (e.g., removed) from the shots 304 and provided as input to the layout generator 312. In some embodiments, the secondary region of interest detector 310 can provide the shots 304 with the one or more secondary regions of interest removed as input to the primary region of interest detector 308. The primary region of interest detector 308 may perform a union using the remaining bounding boxes in the shots 304 to determine a primary region of interest associated with the shots 304 classified as being of the type suitable to be recomposed, and the primary region of interest can be provided to the layout generator 312. Accordingly, the layout generator 312 can receive one or more secondary regions of interest and a primary region of interest as input for shots 304 classified as being of the type suitable to be recomposed.
[0092] If the shot 304 is classified as being of the type suitable for aspect ratio preservation, the shot classifier 306 may directly provide the shot type and the shot 304 as input to the layout generator 312. The layout generator may produce one or more layouts / display configurations for modified versions of the video item 301 using the shots 304, shot type, and regions of interest. In some embodiments, the layout generator 312 may receive a shot type, one or more regions of interest (e.g., one or more secondary regions of interest and / or a primary region of interest), video item 301, a start time, an end time, and audio corresponding to the video 301 as input. In some embodiments, the layout generator 312 may generate the one or more display configurations (referred to generally as “layouts” herein) for modified video content of the video item 301 based on a shot type.
[0093] In some embodiments, the layout generator 312 may receive an indication from shot classifier 306 that the shot 304 is of the type suitable for aspect ratio preservation. Accordingly, the layout generator may generate one or more layouts / display configuration that preserve the entire horizontal width of the frame. For example, the video item 301 may have a 16:9 aspect ratio (for landscape viewing). The layout generator 312 may generate a layout suitable for portrait viewing while preserving the 16:9 aspect ratio by placing letterbox bars (e.g., black bars, blurred bars, etc.) at the top and / or bottom of the frame to preserve the original 16:9 aspect ratio.
[0094] In some embodiments, the layout generator 312 may receive an indication that shot 304 is of the type suitable for cropping. Because the shot 304 has been classified as suitable for cropping, the layout generator 312 may also receive a primary region of interest from the primary region of interest detector 308. The layout generator may generate one or more layouts / display configurations that include content from the shots 304 cropped to the primary region of interest.
[0095] In some embodiments, the layout generator 312 may receive an indication from the shot classifier 306 that the shot 304 is suitable to be recomposed. Because the shot 304 has been classified as suitable to be recomposed, the layout generator 312 may also receive a primary region of interest corresponding to the shot 304 from the primary region of interest detector 308, and one or more secondary regions of interest corresponding to the shot 304 from the secondary region of interest detector 310. The layout generator may reassemble (e.g., recompose) the cropped primary region of interest and the one or more cropped secondary regions of interest to preserve visually pertinent content that may be occurring within the regions of interest. In some embodiments, the layout generator 312 may combine the regions of interest to generate layout / display configuration for modified video content.
[0096] In an illustrative example, the layout generator 312 may composite the one or more secondary regions of interest on top of the primary region of interest to generate a first layout. In some embodiments, the layout generator may generate more than one layout. For example, the layout generator 312 may generate a second layout by including the primary region of interest at a first region of the second layout and including the one or more secondary regions of interest at a second, smaller region of the second layout. In some embodiments, the layout generator 312 may generate layouts based on one or more predefined layout designs, such as including composing the secondary regions of interest and the primary region of interest into separate regions (e.g., an upper region and a lower region) with equal areas, composing the secondary regions of interest and the primary region of interest into separate regions with unequal areas, compositing the secondary region of interest on top of the primary region of interest, etc.
[0097] The layout generator 312 may provide one or more generated layouts of modified video shots 304 as input to the layout picker 314 with an identifier associated with the video item 301 and a start and end time of the included shots 304. In some embodiments, the layout picker 314 may be an Al model trained to predict a quality score for the one or more provided layouts. In some embodiments, the layout generator and / or the layout picker 314 may be trained and evaluated according to one or more quality metrics that indicate a quality level of the generated layouts. The quality metrics may include, but are not limited to, whether important visual elements (e.g., text, important persons, important objects, etc.) are included in the generated layout; whether some visual elements distorted (e.g., too zoomed in, too zoomed out, etc.); whether a visual element is duplicated; whether the layout is not suitable for the content shown in the shots; and the like. In some embodiments, human evaluators may use such quality metrics to evaluate layouts produced by the layout generator 312 and picked by the layout picker 314 to refine the output of layout generator 312 and layout picker 314. In some embodiments, the quality metrics may be used to train Al models associated with primary region of interest detector 308 and / or primary region of interest detector 310.
[0098] In some embodiments, the layout picker 314 can select a layout generated by the layout generator 312 with the greatest quality score. In some embodiments, a modified video content from shots 304 may be generated using the selected layout and provided to the client device 102. In response to generating the modified video content from the shots 304, user interface (UI) can be provided, for display on the client device 102, to present the modified video content from the shots 304.
[0099] FIG. 6 illustrates a flow diagram of an example method of automatically generating modified video content from a video item, in accordance with aspects and implementations of the present disclosure. Method 600 may be performed by processing logic that can include hardware (circuitry, dedicated logic, etc.), software (e.g., instructions run on a processing device), or a combination thereof. In one implementation, some or all the operations of method 600 can be performed by one or more components of system 100 of FIG. 1. In at least one embodiment, some or all of the operations of method 600 can be performed by a server device or a client device.
[0100] For simplicity of explanation, the methods of this disclosure are depicted and described as a series of acts. However, acts in accordance with this disclosure can occur in various orders and / or concurrently, and with other acts not presented and described herein. Furthermore, not all illustrated acts can be required to implement the methods in accordance with the disclosed subject matter. In addition, those skilled in the art will understand and appreciate that the methods could alternatively be represented as a series of interrelated states via a state diagram or events. Additionally, it should be appreciated that the methods disclosed in this specification are capable of being stored on an article of manufacture to facilitate transporting and transferring such methods to computing devices. The term “article of manufacture,” as used herein, is intended to encompass a computer program accessible from any computer-readable device or storage media.
[0101] At operation 602, the processing logic can receive a request to generate modified video content from a video item, such as media item 121 of FIG. 1.
[0102] At operation 604, the processing logic can determine a type of a sequence of frames of the video item, the type pertaining to one or more regions of interest in a plurality of sequences of frames. In some embodiments, the sequence of frames is a shot of the video item.
[0103] In some embodiments, to determine the type of the sequence of frames of the video item, the processing logic can provide information associated with the sequence of frames as input into one or more Al models. In some embodiments, the information associated with the sequence of frames includes at least one of bounding boxes indicative of facial features, bounding boxes indicative of persons, bounding boxes indicative of text, or one or more embeddings. The one or more Al models are trained to predict, based on the information associated with the sequence of frames, a probability score for each type of multiple types. The processing logic can obtain outputs from the one or more Al models. The outputs include multiple probability scores each indicating a likelihood that the sequence offrames belongs to a respective type of the multiple types. The processing logic can determine the type based on the plurality of probability scores.
[0104] At operation 606, the processing logic can generate the modified video content from the video item with a display configuration based on the type. In some embodiments, the processing logic can determine that the type of the sequence is a first type (also referred to as a “shot type suitable for cropping” herein). The first type indicates that the sequence of frames includes one region of interest with a size below a threshold. The processing logic can crop the sequence of frames to the one region of interest to generate the display configuration of the modified video content.
[0105] In some embodiments, the processing logic can determine that the type of the sequence of frames is a second type (also referred to as “shot type suitable for aspect ratio preservation” herein). The second type indicates that the sequence of frames includes one region of interest with a size above the threshold (e.g., occupies a predetermined portion of a horizontal width of the sequence of frames). The processing logic can maintain an original aspect ratio of the sequence of frames in the display configuration of the modified video content.
[0106] In some embodiments, the processing logic can determine that the type of the sequence of frames isa third type (also referred to as a “shot type suitable to be recomposed” herein). The third type indicates that the sequence of frames includes multiple regions of interest. The processing logic can include at least a subset of the multiple of regions of interest within the display configuration of the modified video content.
[0107] In some embodiments, to include at least the subset of the multiple regions of interest within the display configuration of the modified video content, the processing logic can identify a secondary region of interest of the multiple regions of interest. The processing logic can remove content associated with the secondary region of interest from each frame of the sequence of frames. The processing logic can identify a primary region of interest of the sequence of frames with the removed content and include the primary region of interest and secondary region of interest within the configuration of the modified video content.
[0108] In some embodiments, to identify the secondary region of interest, the processing logic can determine temporal features for the each of the multiple regions of interest. The temporal features include at least a size, a location, or an overlap of the multiple regions of interest. The processing logic can further determine spatial features for each of the multiple of regions of interest. The spatial features indicate an amount of movement of the multiple regions of interest between frames of the sequence of frames. The processing logic canfurther determine the secondary region of interest based on the temporal features and the spatial features. The secondary region of interest is a spatially and temporally static region of interest of a lesser size than other regions of interest of the plurality of regions of interest.
[0109] FIG. 7 illustrates an example training engine 141 for training and deployment of a deep neural network, in accordance with aspects and implementations of the present disclosure. In some embodiments, untrained neural network 706 is trained using a training dataset 702. In some embodiments, training data generator 131 can be configured to generate training dataset 702. In some embodiments, training framework 704 is a PyTorch framework, whereas in other embodiments, training framework 704 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deepleaming4j, or other training framework. In at least one embodiment, training framework 704 trains an untrained neural network 706 and enables it to be trained using processing resources described herein to generate a trained neural network 708. The training framework 704 can be used to generate a trained neural network models associated with shot boundary detector 302, shot classifier 306, primary region of interest detector 308, secondary region of interest detector 310, layout generator 312, and layout picker 314. In some embodiments, weights can be chosen randomly or by pre-training using a deep belief network. In some embodiments, training can be performed in either a supervised, partially supervised, or unsupervised manner.
[0110] In some embodiments, untrained neural network 706 is trained using supervised learning, wherein training dataset 702 includes an input paired with a desired output for an input, or where training dataset 702 includes input having a known output and an output of neural network 706 is manually graded. For example, shot classifier 306 can be trained using supervised learning, where training dataset 702 includes an input of historical shots paired with a desired output that indicates a shot type of the respective shot. The dataset can include diverse shots capturing variations in location, orientation, and presence of various features. In at least one embodiment, untrained neural network 706 is trained in a supervised manner and processes inputs from training dataset 702 and compares resulting outputs against a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through untrained neural network 706. In at least one embodiment, training framework 704 adjusts weights that control untrained neural network 706. In at least one embodiment, training framework 704 includes tools to monitor how well untrained neural network 706 is converging towards a model, such as trained neural network 708, suitable to generating correct answers, such as in result 714, based on input data such as a new dataset 712. In at least one embodiment, training framework 704 trains untrained neural network 706repeatedly while adjust weights to refine an output of untrained neural network 706 using a loss function and adjustment algorithm, such as stochastic gradient descent. In at least one embodiment, training framework 704 trains untrained neural network 706 until untrained neural network 706 achieves a desired accuracy. In at least one embodiment, trained neural network 708 can then be deployed to implement any number of machine learning operations.
[0111] In at least one embodiment, untrained neural network 706 is trained using unsupervised learning, wherein untrained neural network 706 attempts to train itself using unlabeled data. In at least one embodiment, unsupervised learning training dataset 702 will include input data without any associated output data or “ground truth” data. In at least one embodiment, untrained neural network 706 can learn groupings within training dataset 702 and can determine how individual inputs are related to training dataset 702. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in trained neural network 708 capable of performing operations useful in reducing dimensionality of new dataset 712.
[0112] In at least one embodiment, semi-supervised learning can be used, which is a technique in which training dataset 702 includes a mix of labeled and unlabeled data. In at least one embodiment, training framework 704 can be used to perform incremental learning, such as through transferred learning techniques. In at least one embodiment, incremental learning enables trained neural network 708 to adapt to new dataset 712 without forgetting knowledge instilled within trained neural network 708 during initial training.
[0113] FIG. 8 is a block diagram illustrating an exemplary computer system 800, in accordance with implementations of the present disclosure. The computer system 800 can correspond to platform 120 and / or client devices 102A-N, described with respect to FIG. 1. Computer system 800 can operate in the capacity of a server or an endpoint machine in endpoint-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine can be a television, a personal computer (PC), a tablet PC, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0114] The example computer system 800 includes a processing device (processor) 802, a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random accessmemory (DRAM) such as synchronous DRAM (SDRAM), double data rate (DDR SDRAM), or DRAM (RDRAM), etc.), a static memory 806 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 818, which communicate with each other via a bus 840.
[0115] Processor (processing device) 802 represents one or more general-purpose processing devices such as a microprocessor, central processing unit, or the like. More particularly, the processor 802 can be a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or processors implementing a combination of instruction sets. The processor 802 can also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. The processor 802 is configured to execute instructions 805 for performing the operations discussed herein.
[0116] The computer system 800 can further include a network interface device 808. The computer system 800 also can include a video display unit 810 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an input device 812 (e.g., a keyboard, and alphanumeric keyboard, a motion sensing input device, touch screen), a cursor control device 814 (e.g., a mouse), and a signal generation device 820 (e.g., a speaker).
[0117] The data storage device 818 can include a non-transitory machine-readable storage medium 824 (also non-transitory computer-readable storage medium) on which is stored one or more sets of instructions 805 embodying any one or more of the methodologies or functions described herein. The instructions can also reside, completely or at least partially, within the main memory 804 and / or within the processor 802 during execution thereof by the computer system 800, the main memory 804 and the processor 802 also constituting machine-readable storage media. The instructions can further be transmitted or received over a network 830 via the network interface device 808.
[0118] In one implementation, the instructions 805 include instructions for automatically generating modified video content from a video item. While the computer-readable storage medium 824 (machine-readable storage medium) is shown in an exemplary implementation to be a single medium, the terms “computer-readable storage medium” and “machine- readable storage medium” should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store the one or more sets of instructions. The terms “computer-readable storage medium” and“machine-readable storage medium” shall also be taken to include any medium that is capable of storing, encoding or carrying a set of instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present disclosure. The terms “computer-readable storage medium” and “machine-readable storage medium” shall accordingly be taken to include, but not be limited to, solid-state memories, optical media, and magnetic media.
[0119] Reference throughout this specification to “one implementation,” “one embodiment,” “an implementation,” or “an embodiment,” means that a particular feature, structure, or characteristic described in connection with the implementation and / or embodiment is included in at least one implementation and / or embodiment. Thus, the appearances of the phrase “in one implementation,” or “in an implementation,” in various places throughout this specification can, but are not necessarily, referring to the same implementation, depending on the circumstances. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more implementations.
[0120] To the extent that the terms “includes,” “including,” “has,” “contains,” variants thereof, and other similar words are used in either the detailed description or the claims, these terms are intended to be inclusive in a manner similar to the term “comprising” as an open transition word without precluding any additional or other elements.
[0121] As used in this application, the terms “component,” “module,” “system,” or the like are generally intended to refer to a computer-related entity, either hardware (e.g., a circuit), software, a combination of hardware and software, or an entity related to an operational machine with one or more specific functionalities. For example, a component can be, but is not limited to being, a process running on a processor (e.g., digital signal processor), a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a controller and the controller can be a component. One or more components can reside within a process and / or thread of execution and a component can be localized on one computer and / or distributed between two or more computers. Further, a “device” can come in the form of specially designed hardware; generalized hardware made specialized by the execution of software thereon that enables hardware to perform specific functions (e.g., generating interest points and / or descriptors); software on a computer readable medium; or a combination thereof.
[0122] The aforementioned systems, circuits, modules, and so on have been described with respect to interact between several components and / or blocks. It can be appreciated thatsuch systems, circuits, components, blocks, and so forth can include those components or specified sub-components, some of the specified components or sub-components, and / or additional components, and according to various permutations and combinations of the foregoing. Sub-components can also be implemented as components communicatively coupled to other components rather than included within parent components (hierarchical). Additionally, it should be noted that one or more components can be combined into a single component providing aggregate functionality or divided into several separate subcomponents, and any one or more middle layers, such as a management layer, can be provided to communicatively couple to such sub-components in order to provide integrated functionality. Any components described herein can also interact with one or more other components not specifically described herein but known by those of skill in the art.
[0123] Moreover, the words “example” or “exemplary” are used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the words “example” or “exemplary” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X employs A or B” is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then “X employs A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form.
[0124] Finally, implementations described herein include collection of data describing a user and / or activities of a user. In one implementation, such data is only collected upon the user providing consent to the collection of this data. In some implementations, a user is prompted to explicitly allow data collection. Further, the user can opt-in or opt-out of participating in such data collection activities. In one implementation, the collect data is anonymized prior to performing any analysis to obtain any statistical patterns so that the identity of the user cannot be determined from the collected data.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: receiving a request to generate modified video content from a video item; determining a type of a sequence of frames of the video item, the type pertaining to one or more regions of interest in a plurality of sequences of frames of the video item; and generating the modified video content from the video item with a display configuration based on the type.
2. The method of claim 1, wherein the sequence of frames comprises a shot of the video item.
3. The method of claim 1, wherein determining the type of the sequence of frames of the video item comprises: providing information associated with the sequence of frames as input into one or more artificial intelligence (Al) models, wherein the one or more Al models are trained to predict, based on the information associated with the sequence of frames, a probability score for each type of a plurality of sequence of frames types; obtaining a plurality of outputs from the one or more Al models, wherein the plurality of outputs comprises a plurality of probability scores each indicating a likelihood that the sequence of frames belongs to a respective type of the plurality of sequence of frames types; and determining the type based on the plurality of probability scores.
4. The method of claim 3, wherein the information associated with the sequence of frames comprises at least one of bounding boxes indicative of facial features, bounding boxes indicative of persons, bounding boxes indicative of text, or one or more embeddings.
5. The method of claim 1, wherein the type of the sequence of frames is one of a plurality of sequence of frames types, wherein each of the plurality of sequence of frames types is associated with a distinct set of region of interest characteristics, and wherein the plurality of sequence of frames types comprises a first type indicating that the sequence of frames includes one region of interest with a size below a threshold, a second type indicatingthat the sequence of frames includes one region of interest with a size above the threshold, and a third type indicating that the sequence of frames comprises a plurality of regions of interest.
6. The method of claim 5, further comprising: determining that the type of the sequence of frames is the first type; and cropping the sequence of frames to the one region of interest to generate the display configuration of the modified video content.
7. The method of claim 5, further comprising: determining that the type of the sequence of frames is the second type; and maintaining an original aspect ratio of the sequence of frames in the display configuration of the modified video content.
8. The method of claim 5, further comprising: determining that the type of the sequence of frames is the third type; and including at least a subset of the plurality of regions of interest within the display configuration of the modified video content.
9. The method of claim 8, wherein including at least the subset of the plurality of regions of interest within the display configuration of the modified video content comprises: identifying a secondary region of interest of the plurality of regions of interest; removing content associated with the secondary region of interest from each frame of the sequence of frames; identifying a primary region of interest of the sequence of frames with the removed content; and including the primary region of interest and secondary region of interest within the configuration of the modified video content.
10. The method of claim 9, wherein identifying the secondary region of interest comprises: determining temporal features for the each of the plurality of regions of interest, wherein the temporal features comprise at least a size, a location, or an overlap of the plurality of regions of interest;determining spatial features for each of the plurality of regions of interest, wherein the spatial features indicate an amount of movement of the plurality of regions of interest between frames of the sequence of frames; and determining the secondary region of interest based on the temporal features and the spatial features, wherein the secondary region of interest is a spatially and temporally static region of interest of a lesser size than other regions of interest of the plurality of regions of interest.
11. A system comprising: a memory device; and a processing device coupled to the memory device, the processing device to perform operation comprising: receiving a request to generate modified video content from a video item; determining a type of a sequence of frames of the video item, the type pertaining to one or more regions of interest in a plurality of sequences of frames of the video item; and generating the modified video content from the video item with a display configuration based on the type.
12. The system of claim 11, wherein the sequence of frames comprises a shot of the video item.
13. The system of claim 11, wherein determining the type of the sequence of frames of the video item comprises: providing information associated with the sequence of frames as input into one or more artificial intelligence (Al) models, wherein the one or more Al models are trained to predict, based on the information associated with the sequence of frames, a probability score for each type of a plurality of sequence of frames types; obtaining a plurality of outputs from the one or more Al models, wherein the plurality of outputs comprises a plurality of probability scores each indicating a likelihood that the sequence of frames belongs to a respective type of the plurality of sequence of frame types; and determining the type based on the plurality of probability scores.
14. The system of claim 13, wherein the information associated with the sequence of frames comprises at least one of bounding boxes indicative of facial features, bounding boxes indicative of persons, bounding boxes indicative of text, or one or more embeddings.
15. The system of claim 11, wherein the type of the sequence of frames is one of a plurality of sequence of frames types, wherein each of the plurality of sequence of frames types is associated with a distinct set of region of interest characteristics, and wherein the plurality of sequence of frames types comprises a first type indicating that the sequence of frames includes one region of interest with a size below a threshold, a second type indicating that the sequence of frames includes one region of interest with a size above the threshold, and a third type indicating that the sequence of frames comprises a plurality of regions of interest.
16. The system of claim 15, further comprising: determining that the type of the sequence of frames is the first type; and cropping the sequence of frames to the one region of interest to generate the display configuration of the modified video content.
17. The system of claim 15, further comprising: determining that the type of the sequence of frames is the second type; and maintaining an original aspect ratio of the sequence of frames in the display configuration of the modified video content.
18. The system of claim 15, further comprising:Determining that the type of the sequence of frames is the third type; and including at least a subset of the plurality of regions of interest within the display configuration of the modified video content.
19. The system of claim 18, wherein including at least the subset of the plurality of regions of interest within the display configuration of the modified video content comprises: identifying a secondary region of interest of the plurality of regions of interest; removing content associated with the secondary region of interest from each frame of the sequence of frames;identifying a primary region of interest of the sequence of frames with the removed content; and including the primary region of interest and secondary region of interest within the configuration of the modified video content.
20. A non-transitory computer-readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform operations comprising: receiving a request to generate modified video content from a video item; determining a type of a sequence of frames of the video item, the type pertaining to one or more regions of interest in a plurality of sequences of frames of the video item; and generating the modified video content from the video item with a display configuration based on the type.
Citation Information
Patent Citations
Automated video cropping
US10834465B1
Intelligent reframing
US11595614B1
Cited By
Generation of candidate video elements
US12700145B2
Generation of candidate video elements
US20250292442A1
Smartphone photography adapted for screen orientations and social media
US20250336057A1