Personalized Automatic Video Cropping
By using machine learning models to generate facial signals and adjust crop scores, the problem of device display aspect ratio and orientation mismatch with video is solved, automatic video cropping is achieved, and video playback experience is improved.
Patent Information
- Application Number
- CN202080065525.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-13
- Filing Date
- 2020-12-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-12-08
AI Technical Summary
When the display aspect ratio and/or orientation on the device does not match the aspect ratio of the video or image, it is often necessary to add frames up and down to accommodate the display, resulting in a poor viewing experience.
Automatic video cropping is achieved by using a trained machine learning model, face signals are generated and cropped values are adjusted to determine the position of the clip area of the video. Based on motion costs and facial signals, the model generates a minimum cost path and outputs cropped video with appropriate aspect ratios or orientations.
It realizes that when the device display aspect ratio and orientation do not match the video, the video cropping area is automatically adjusted, which improves the video playback experience and ensures that the video can be displayed appropriately on different devices.
Smart Images

Figure CN114402355B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 948,179, filed on December 13, 2019, entitled "Personalized Automatic Video Cropping", the entire content of which is incorporated herein by reference. Background Art
[0003] When viewing videos (and images) on a device, the display aspect ratio and / or orientation of the device may often not match the aspect ratio of the media. As a result, the media is typically letterboxed for display (e.g., having large black borders on the sides and a reduced video size or still image size between the borders). In some cases, a viewer software application may crop the original media to avoid letterboxing.
[0004] The background art provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent described in this background art section, the work of the currently named inventors and aspects that may not constitute prior art at the time of filing are neither expressly nor implicitly considered prior art to the present disclosure. Summary of the Invention
[0005] Some implementations may include a method. The method may include: obtaining an input video including a plurality of frames, and determining a per-frame crop score for each of one or more candidate crop regions in each frame of the input video. The method may also include using a trained machine learning model to generate a face signal for each of one or more candidate crop regions within each frame of the input video, and adjusting each per-frame crop score based on the face signal of the one or more candidate crop regions. In some implementations, the face signal may indicate whether at least one significant face is detected in the candidate crop region.
[0006] The method may further include determining a minimum cost path representing the crop region location for the input video based on a motion cost and the adjusted per-frame crop scores for the one or more candidate crop regions, and generating crop key frames corresponding to the crop region locations along the minimum cost path, where the crop key frames include a start frame, an end frame, and a crop region location. The method may also include outputting a modified video having one or more of an output aspect ratio or an output orientation different from a corresponding input aspect ratio or an input orientation of the input video, where the aspect ratio or input orientation is a parameter used during the capture of the input video.
[0007] In some implementations, adjusting each per-frame cropping score includes one of the following: increasing the per-frame cropping score by a first value if it is determined that a face is present in a candidate cropping region corresponding to the per-frame cropping score; or increasing the per-frame cropping score by a second value if it is determined that at least one significant face is present in the candidate cropping region corresponding to the per-frame cropping score, where the second value is greater than the first value.
[0008] The method may also include determining a quality score for a cropping key frame and performing automatic video cropping of the input video based on the quality score. The method may also include determining a confidence score for the cropping key frame and performing automatic video cropping of the input video based on the confidence score.
[0009] In some implementations, determining the per-frame cropping score includes determining one or more of the following for each candidate cropping region: an aesthetics score, a face analysis score, or the presence of an active speaker. In some implementations, generating a cropping key frame includes interpolating between two key frames. In some implementations, the interpolation includes applying a Bézier spline.
[0010] In some implementations, generating a face signal includes accessing one or more personalized parameters. In some implementations, the one or more personalized parameters include face recognition information for one or more significant faces. In some implementations, outputting the modified video includes displaying the modified video on a display.
[0011] The method may also include, before obtaining the input video, receiving a video playback command at the device and, in response to receiving the video playback command, detecting the device orientation and the display aspect ratio for the device. The method may also include determining a cropping region based on the device orientation and the display aspect ratio for the device.
[0012] Some implementations may include a non-transitory computer-readable medium having software instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations. The operations may include obtaining an input video including a plurality of frames, determining a per-frame cropping score for one or more candidate cropping regions in each frame of the input video, and using a trained machine learning model to generate a face signal for one or more candidate cropping regions within each frame of the input video. In some implementations, the face signal may indicate whether at least one significant face is detected in the candidate cropping region. The operations may also include adjusting each per-frame cropping score based on the face signal for one or more candidate cropping regions and determining a minimum cost path representing the location of the cropping region for the input video based on a motion cost and the adjusted per-frame cropping scores for one or more candidate cropping regions.
[0013] The operation may also include generating crop keyframes corresponding to the positions of the crop regions along the minimum cost path, where the crop keyframes include a start frame, an end frame, and the crop region positions, and outputting a modified video having one or more of an output aspect ratio or an output orientation different from a corresponding input aspect ratio or input orientation of the input video, where the aspect ratio or input orientation is a parameter used during the capture of the input video.
[0014] In some implementations, adjusting each per-frame crop score includes one of the following: increasing the per-frame crop score by a first value if it is determined that a face is present in the candidate crop region corresponding to the per-frame crop score; or increasing the per-frame crop score by a second value if it is determined that at least one significant face is present in the candidate crop region corresponding to the per-frame crop score, where the second value is greater than the first value.
[0015] The operation may also include determining a quality score for the crop keyframes and performing automatic video cropping of the input video based on the quality score. The operation may also include determining a confidence score for the crop keyframes and performing automatic video cropping of the input video based on the confidence score.
[0016] In some implementations, determining the per-frame crop score includes determining one or more of the following for each candidate crop region: an aesthetics score, a face analysis score, or the presence of an active speaker. In some implementations, generating the crop keyframes includes interpolating between two keyframes. In some implementations, the interpolation includes applying a Bézier spline. In some implementations, generating the face signal includes accessing one or more personalized parameters.
[0017] Some implementations may include a system including one or more processors coupled to a non-transitory computer-readable medium storing software instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations may include obtaining an input video including a plurality of frames, determining a per-frame crop score for one or more candidate crop regions in each frame of the input video, and using a trained machine learning model to generate a face signal for one or more candidate crop regions within each frame of the input video. In some implementations, the face signal may indicate whether at least one significant face is detected in the candidate crop region. The operation may also include adjusting each per-frame crop score based on the face signals for the one or more candidate crop regions and determining a minimum cost path representing the position of the crop region for the input video based on the motion cost and the adjusted per-frame crop scores for the one or more candidate crop regions.
[0018] The operation can also include generating crop keyframes corresponding to the positions of crop regions along a minimum cost path, where the crop keyframes include a start frame, an end frame, and a crop region position, and outputting a modified video having one or more of an output aspect ratio or an output orientation different from a corresponding input aspect ratio or an input orientation of the input video, where the aspect ratio or the input orientation is a parameter used during the capture of the input video. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a block diagram of an example network environment that can be used for one or more implementations described herein.
[0020] Figure 2A and 2B are diagrams of a landscape video format and a letterboxed format.
[0021] Figure 3A and 3B illustrate a crop rectangle on a horizontal video and a cropped video displayed in a portrait orientation according to some implementations.
[0022] Figure 4 is a flowchart showing a process for automatically cropping a video according to some implementations.
[0023] Figure 5 is a flowchart of an example method for automatically cropping a video according to some implementations.
[0024] Figure 6 is a block diagram of an example device that can be used for one or more implementations described herein.
[0025] Figure 7 is a diagram showing an example automatic video crop path according to some implementations.
[0026] Figures 8A - 8C is a diagram showing an example video in which a crop region is moved to a different position according to some implementations. DETAILED DESCRIPTION
[0027] Some implementations described herein relate to methods, systems, and computer-readable media for automatically cropping videos. The described implementations can use a trained machine learning model to automatically crop videos using personalized parameters. The training data for the model can include personalized information for a user accessed with the user's permission. The personalized information can include facial recognition information for faces stored in local storage (e.g., on a device).
[0028] Some implementations described herein relate to methods, systems, and computer-readable media for automatically performing personalized video cropping. Different video platforms and delivery devices can have different aspect ratios, including 4:3 (landscape), 9:16 (portrait), and 1:1 (square), where the first number refers to the width of the video and the second number refers to the height of the video.
[0029] The techniques described can automatically perform personalized video cropping (e.g., cropping to display a landscape-oriented video in a portrait or square format) during video playback without interrupting the user based on personalized parameters such as faces that are important to the user (e.g., recognized, familiar, or known to the user based on images / videos in the user's image library). The user can have an image and video library (e.g., images stored and managed by an image management software application) that includes multiple images and / or videos captured by the user or otherwise added to the library (e.g., images shared with the user by other users). The library can be local to the user device (e.g., smartphone) and / or on a server (e.g., cloud-based image / video hosting service). For example, the user can use one or more devices such as a smartphone, digital camera, wearable device, etc. to capture individual images and / or videos of people and store such images in their library. In some implementations, important faces can also include faces that are not in the user's library but are identified based on other image sources - such as the user's social graph, social media accounts, email accounts, the user's electronic address book - accessed with the user's permission.
[0030] If the user refuses to grant permission to access the library or any other source, then those sources are not accessed and important face determination is not performed. Additionally, the user can exclude one or more faces from being recognized and / or included as important faces. Further, the user can be provided with options (e.g., a user interface) to manually indicate important faces. As used herein, the term face can refer to a human face and / or any other face that can be detected using face detection techniques (e.g., the face of a pet or other animal).
[0031] Some image management applications may include features that are enabled with user permission to detect the faces of people and / or pets in images or videos. If the user permits, such an image management application can detect the faces in the images / videos in the user's library and determine the frequency of occurrence of each face. For example, the faces of individuals such as a spouse, siblings, parents, close friends, etc. may occur at a high frequency in the images / videos in the user's library, while other people such as bystanders (e.g., in a public place) may not occur as frequently. In some implementations, faces that occur at a high frequency can be identified as important faces. In some implementations, for example, when the library enables the user to tag or label faces, the user can indicate the names (or other information) of the faces that appear in their library. In these implementations, the faces for which the user has provided a name or other information can be identified as important faces. Other faces detected in the image library can be identified as unimportant.
[0032] According to various implementations, video cropping and determination of important faces are performed locally on the client device and do not require a network connection. The described techniques enable an improved video playback experience when viewing a video in an aspect ratio different from that of the video or in a device orientation (e.g., portrait vs. landscape) that is not suitable for the video to be captured or stored therein. The described techniques can be implemented on any device (e.g., a mobile device) that plays back a video.
[0033] In some implementations, automatic video cropping can include three stages: per-frame cropping score, temporal coherence, and motion smoothing. The per-frame cropping score can be image-based, which can be noisy and include various different scores. In some implementations, a heuristic combination can be used to generate a single per-frame score for candidate cropping regions. The second stage can include temporal coherence, which can include operations to smooth an optimal path through spatial and temporal fitting. Temporal coherence can include a representation of scene motion. The third stage can include motion smoothing and heuristic combination. In this stage, local optimization can be performed for aspects of the video that may not be globally tractable. Specific heuristics and rules can be applied to handle specific cases.
[0034] Figure 1 A block diagram of an example network environment 100 that can be used in some implementations described herein is shown. In some implementations, network environment 100 includes one or more server systems, e.g., Figure 1 server system 102 in the example of. Server system 102 can communicate with, for example, network 130. Server system 102 can include server device 104 and database 106 or other storage devices. In some implementations, server device 104 can provide video application 158.
[0035] The network environment 100 may also include one or more client devices, such as client devices 120, 122, 124, and 126, which may communicate with each other and / or with the server system 102 and / or the second server system 140 via the network 130. The network 130 may be any type of communication network, including one or more of the Internet, a local area network (LAN), a wireless network, a switch or hub connection, etc.
[0036] For ease of illustration, Figure 1 One box is shown for the server system 102, the server device 104, and the database 106, and four boxes are shown for the client devices 120, 122, 124, and 126. The server boxes 102, 104, and 106 may represent multiple systems, server devices, and network databases, and these block boxes may be provided in a configuration different from the one shown. For example, the server system 102 may represent multiple server systems that may communicate with other server systems via the network 130. In some implementations, the server system 102 may include, for example, a cloud-hosted server. In some examples, the database 106 and / or other storage devices may be provided in a server system box separate from the server device 104 and may communicate with the server device 104 and other server systems via the network 130.
[0037] Moreover, any number of client devices may exist. Each client device may be any type of electronic device capable of communication, such as, for example, a desktop computer, a laptop computer, a portable or mobile device, a cellular phone, a smart phone, a tablet computer, a television, a TV set-top box or entertainment device, a wearable device (such as display glasses or goggles, a watch, headphones, an armband, jewelry, etc.), a personal digital assistant (PDA), etc. In some implementations, the network environment 100 may not have all of the components shown and / or may have other elements including other types of elements in place of or in addition to those described herein.
[0038] In various implementations, end users U1, U2, U3, and U4 may use the respective client devices 120, 122, 124, and 126 to communicate with the server system 102 and / or with each other. In some examples, users U1, U2, U3, and U4 may interact with each other via applications running on the respective client devices and / or the server system 102 and / or via network services (such as social network services or other types of network services) implemented on the server system 102. For example, the respective client devices 120, 122, 124, and 126 may convey data to and from one or more server systems, such as the server system 102.
[0039] In some implementations, the server system 102 may provide appropriate data to the client devices such that each client device can receive the communicated or shared content uploaded to the server system 102. In some examples, users U1 - U4 may interact via audio or video conferencing, audio, video, or text chat, or other communication modes or applications.
[0040] The network services implemented by the server system 102 may include systems that allow users to perform various communications, form links and associations, upload and publish shared content such as images, text, video, audio, and other types of content, and / or perform other functions. For example, the client devices may display the received data, such as content posts that are sent or streamed to the client devices and originate from different client devices (or directly from different client devices) or from the server system and / or network services via the server and / or network services.
[0041] In some implementations, any one of the client devices 120, 122, 124, and / or 126 may provide one or more applications. For example, as Figure 1 shown, the client device 120 may provide an automatic video cropping application 152. The client devices 122 - 126 may also provide similar applications. The automatic video cropping application 152 may be implemented using the hardware and / or software of the client device 120. In different implementations, the automatic video cropping application 152 may be, for example, an independent client application executed on any one of the client devices 120 - 124. The automatic video cropping application 152 may provide various functions related to videos, such as automatically cropping a video to change from one aspect ratio to another, etc.
[0042] The user interfaces on the client devices 120, 122, 124, and / or 126 may be implemented to display user content and other content, including images, videos, data, and other content, as well as communications, settings, notifications, and other data. Such user interfaces may be displayed using software on the client devices, software on the server devices, and / or a combination of client software and server software executed on the server device 104. The user interface may be displayed by a display device of the client device - for example, a touch screen or other display screen, a projector, etc. In some implementations, the server may simply enable the user to stream / download videos over the network and, with the user's permission, enable the upload / storage of videos sent by the user.
[0043] Other implementations of the features described herein can use any type of system and / or service. For example, instead of or in addition to a social networking service, other network services (e.g., connected to the Internet) can be used. Any type of electronic device can utilize the features described herein. Some implementations can provide one or more of the features described herein on one or more client or server devices that are disconnected from or intermittently connected to a computer network.
[0044] Figure 2A is a diagram of a video in landscape format according to some implementations. Figure 2B is when the video is viewed on the device in portrait orientation with the video framed left and right Figure 2A of the video shown in. Figure 2A Shows user device 202 displaying a video in landscape mode 204. Figure 2B Shows that the device is now in portrait mode and the video 204 has been framed left and right 206 for display in portrait orientation. As Figure 2B shown, when the video is framed left and right for display in portrait orientation, the portion of the device's display screen occupied by the video is substantially smaller.
[0045] Figure 3A is a diagram showing a cropping rectangle on a horizontal video according to some implementations. Figure 3B is a diagram showing a cropped video displayed in portrait orientation according to some implementations. Figure 3A Shows a single frame of a video in landscape orientation with a cropping area 302 shown as a dashed line. The cropping area 302 is located at x position 304. As described below, the automatic cropping process generates the x position for the cropping area (landscape to portrait cropping) across the temporal domain of the video. More generally, the process generates a cropping window (x, y, width, height) relative to the source video.
[0046] Figure 4 is a flowchart showing an example method 400 of automatically cropping a video personalized for a user according to some implementations. In some implementations, method 400 can be implemented, for example, on a server system 102 as shown in Figure 1 shown. In some implementations, some or all of method 400 can be implemented on a device as shown in Figure 1Implemented on one or more of the client devices 120, 122, 124, or 126 shown, on one or more server devices, and / or on both server and client devices. In the example described, the implementation system includes one or more digital processors or processing circuits ("processors") and one or more storage devices (e.g., database 106 or other storage devices). In some implementations, different components of one or more servers and / or clients may perform different blocks or other portions of method 400. In some examples, a first device is described as performing a block of method 400. Some implementations may have one or more blocks of method 400 performed by one or more other devices (e.g., other client devices or server devices) that may send results or data to the first device.
[0047] In some implementations, method 400 or portions of the method may be initiated automatically by the system. In some implementations, the implementation system is the first device. For example, the method (or portions thereof) may be performed based on one or more specific events or conditions - e.g., playback of a video on a client device, preparation of a video for upload from a client device, and / or one or more other conditions specified in settings that may be read by the method.
[0048] Method 400 may begin at block 402. At block 402, it is checked whether user consent (e.g., user permission) has been obtained to use user data in the implementation of method 200. For example, user data may include important faces and additional user criteria, user images in an image collection (e.g., images captured by the user, uploaded by the user, or otherwise associated with the user), information about the user's social network and / or contacts, user characteristics (identity, name, age, gender, occupation, etc.), social and other types of actions and activities, calendars and appointments, content, ratings, and opinions created or submitted by the user, the user's geographical location, historical user data, etc. Block 402 may be performed as part of an automatic video cropping framework layer such that blocks 404 and subsequent blocks are only invoked when user consent for performing the automatic cropping application has been obtained at the framework layer. If user consent has been obtained from the relevant users from whom user data may be used in method 400, then at block 404, the blocks of the method herein that can be implemented with the possible uses of the user data as described for those blocks are determined, and the method proceeds to block 406. If user consent has not been obtained, then at block 406, it is determined to implement the blocks without using user data, and the method proceeds to block 406. In some implementations, if user consent has not been obtained, the remainder of method 400 is not performed, and / or specific blocks that require user data are not performed. For example, if the user does not provide permission, blocks 412-414 are skipped. Also, the identification of important faces may be based on locally stored data and may be performed locally on the user device. The user may specify particular important faces to identify or not identify, remove the specification, or stop using automatic cropping based on important faces at any time.
[0049] At block 408, an input video is obtained. For example, a video stored in a memory on the user device is accessed. The video may include multiple frames. The input video has an orientation (vertical / horizontal) and an aspect ratio, such as 4:3, 16:9, 18:9, etc. For example, the aspect ratio may be selected at the time of video capture, e.g., based on the camera parameters of the device capturing the video. Block 408 may then be followed by block 410.
[0050] At block 410, a per-frame crop score is determined for one or more candidate crop regions for each frame of an input video. The candidate crop regions can be regions that match the viewing orientation of the device on which the video is being viewed and can have the same aspect ratio as the device, such that the video cropped to that region substantially fills the entire screen (or window if the video is being played in a windowed user interface). For example, if an input video that is landscape (horizontal dimension greater than vertical dimension) and has an aspect ratio of 4000×3000 pixels is to be displayed on a 2000×2000 pixel square display, each candidate crop region can be 3000×3000 pixels, thus matching the 3000 pixel dimension. The selected crop region of 3000×3000 pixels can be scaled to fit the square display, e.g., down to 2000×2000 pixels. Selecting a higher resolution crop region and then scaling can preserve most of the original content. Alternatively, a candidate crop region that matches the 2000×2000 pixel display can be selected.
[0051] The crop score can include one or more individual scores. When more than one score is used, there can be heuristics for determining how to combine the individual scores into a single score. The individual scores can include an aesthetic score (e.g., from 508), a face / person analysis-based score (e.g., 506), and / or an active speaker analysis (e.g., 504). Additionally, some implementations can include one or more additional scores based on object detection, pet or animal detection, or optical character recognition (OCR). For example, a crop region that includes a prominent object identified using object detection techniques can be assigned a higher score than a region in which no prominent object is detected or only a partial object is detected. For example, if the video depicts a natural scene, a crop region that has prominent objects such as trees, mountains, or other objects can be assigned a higher score than, for example, a crop region that includes only the sky and has no prominent objects.
[0052] In another example, a crop region that depicts a pet (e.g., a dog, cat, or other pet animal that is tagged in the user's personal image / video library and accessed with the user's permission) or other animal can be assigned a higher score than a region that excludes the pet or animal or only partially depicts the pet or animal. In yet another example, a region that includes text recognized using OCR can be assigned a higher score. For example, if the video includes a storefront with a sign that includes text, a crop region that includes the sign can be assigned a higher score than a crop region that excludes or only partially depicts the sign. The per-frame crop score can include the score for a candidate crop region (e.g., a crop rectangle at a given x position in the video frame). Block 410 can then be followed by block 412.
[0053] At 412, a facial signal is generated and a personalization score is determined. In some implementations, the facial signal can indicate whether at least one significant face is detected in a candidate crop region. In some implementations, the personalization score can include a score determined based on using face detection techniques to detect faces in a frame and determining whether at least one of the faces in the frame matches a significant face (e.g., as determined based on user-permitted data such as previous videos or photos of people in a user library, or as determined by other user-permitted signals such as a user's social graph connections or communication history in viewing emails, phone calls, chats, video calls, etc.). The personalization score can be determined based on a signal from a machine learning model that represents the extent of one or more significant faces in the candidate crop region, which extent can be determined by the machine learning model. In addition to determining one or more significant faces in the candidate crop region, the personalization score module can also determine the location of one or more significant faces within the candidate crop region, e.g., whether the face is located at the center of the candidate crop region, whether it is near the edge of the candidate crop region, etc. Block 412 can then be followed by block 414.
[0054] At 414, the per-frame crop score is adjusted based on the facial signal. For example, if a face is detected in the candidate crop region, the score for that region can be increased by a first factor. If the candidate crop region is detected to include a significant face, the score for that region can be increased by a second factor that is greater than the first factor.
[0055] In some implementations, the intersection of the crop region with a bounding box that includes the face can be determined. For example, face detection techniques can be used to determine the bounding box. The crop score can be adjusted based on the intersection. For example, in some implementations, a full intersection (where the entire face is within the crop region) can receive a full score boost, while a partial face (where a portion of the face is missing from the crop region) can receive a lower score boost, e.g., the score boost can be weighted by the ratio of the area of the intersection to the area of the face bounding box. Block 408 can then be followed by block 416.
[0056] At 416, a motion cost is determined. In some implementations, the motion cost can be a cost associated with selecting a candidate cropping region at a particular time (e.g., a particular timestamp in a video) given a potential cropping path at one or more previous times (e.g., an earlier timestamp in the video) and the motion present in the video. In some implementations, motion cost determination can include, for example, using optical flow or other techniques to analyze frame-to-frame motion of the cropping region. The results can be clustered into a small number of motion clusters (e.g., clusters that include the motion of cropping regions at a set of positions that are close to each other). In some implementations, sparse optical flow can perform better in non-textured regions. The motion can be reduced to a small number of clusters (e.g., clusters that move around regions that are relatively close to each other). The clustering may not provide temporal coherence.
[0057] In an example implementation, the motion cost can be calculated based on a comparison of the motion of the cropping region relative to a previous time with the motion of the best-matching motion cluster. For example, the best-matching motion cluster can be a cluster that has a spatial centroid within the cropping region and a motion vector that is most similar to the motion of the cropping region. The motion cost can be a function of the absolute difference in velocity between the best-matching motion cluster and the motion of the candidate cropping region. A cost value can be assigned to the moving cropping region and used to determine the motion cost. For example, Figures 8A - 8C A cropping region 804 in video 802 is shown, where the cropping region 802 moves to different positions (e.g., 808 and 812) during playback of the video based on the techniques described herein. Block 416 can then be followed by block 418.
[0058] At 418, a minimum-cost path is determined based on a per-frame cropping score and the motion cost determined at 416. In some implementations, the minimum-cost path can include obtaining candidate cropping regions based on the cropping score and performing a minimum-cost path-finding operation, which can include a cost for moving the cropping region (e.g., the cost for moving the cropping region can be based on the distance that the cropping region is continuously moved from frame to frame or the distance between a frame and a subsequent frame). The minimum-cost path is found by solving for the minimum-cost path. Figure 7 An example graph 700 representing the minimum-cost path and other factors plotted as a curve is shown, where the y-axis 702 is the x position of the cropping region within the video and the x-axis is time. Some implementations can include outlier removal where outliers are removed from the cropping region positions within the input video to smooth the path of the cropping region and remove discontinuities. This may be a byproduct of the minimum-cost path. Block 418 can then be followed by block 420.
[0059] At 420, crop keyframes are generated. The crop keyframes can include a start frame and an end frame, as well as the x position of the crop region within the input video. For example, crop path generation can be performed at 5 frames per second (fps) for a video with a frame rate of 30 fps. In this example, the keyframes can be generated at 5 fps, and interpolation techniques such as Bezier splines can be used to generate smooth interpolation at the full frame rate of 30 fps of the video. For example, there can be three keyframe portions of the video as shown in Figures 8A - 8C , where each crop keyframe includes a crop region at a different x position. Block 420 can then be followed by block 422.
[0060] At 422, a cropped video is output based on the input video and the crop keyframes. For example, the cropped video can be displayed on a display of a user device. In another example, the cropped video can be uploaded to a video sharing site, etc. The cropped video can have an aspect ratio or orientation different from that of the input video.
[0061] In some implementations, instead of or in addition to including the cropped video, the output can include the crop keyframes and paths. For example, the crop keyframes and paths can be stored associated with the video, e.g., as video metadata. When a viewer application initiates playback of the video, the crop keyframes and paths that match the aspect ratio or orientation of the viewer application (which can be based on the device on which the video is viewed) can be determined and provided to the viewer application. The viewer application can utilize the crop keyframes and paths to crop the video during playback. This implementation eliminates the need to generate a separate video asset (which matches the viewer application) when viewing the video in a viewer application that discerns the crop keyframes and paths, and can utilize this information to crop the video during playback.
[0062] The various blocks of method 400 can be combined, split into multiple blocks, or executed in parallel. For example, blocks 406 and 408 can be combined. In some implementations, these blocks can be executed in a different order. For example, blocks 404 - 408 and blocks 412 - 414 can be executed in parallel.
[0063] Method 400 or portions thereof can be repeated any number of times using additional inputs (e.g., additional videos). Method 400 can be implemented with the permission of a particular user. For example, a video playback user interface can be provided that enables the user to specify whether to enable automatic personalized video cropping. The user can be provided with information that performing automatic personalized video cropping during playback can utilize facial recognition using personalized parameters (e.g., stored on the user device), and an option to completely disable automatic personalized video cropping can be provided to the user.
[0064] Method 400 can be performed entirely on a client device that is playing back or uploading a video, including face detection and significant face determination under specific user permissions. Additionally, the face can be a person or other (e.g., animal or pet). Further, the technical benefit of performing automatic personalized video cropping on the device is that the described method does not require the client device to have an active Internet connection, thus allowing automatic video cropping even when the device is not connected to the Internet. Further, since the method is performed locally, it does not consume network resources. Even further, no user data is sent to a server or other third-party device. Thus, the described technology can solve the problem of video playback in an aspect ratio or orientation different from the aspect ratio or orientation in which the video was captured, where the benefits of personalized parameters are utilized in a manner that does not require sharing user data.
[0065] In some implementations, during playback of a video, a change in the orientation or aspect ratio of the device can be detected (e.g., when the user rotates the device 90 degrees during playback, or unfolds a foldable device to double the aspect ratio), and in response, the cropping can be adjusted (e.g., the cropping area can be adjusted to fit the desired output orientation and / or aspect ratio).
[0066] The described technology can advantageously generate a cropped video that is personalized for a user (e.g., the user who is viewing the cropped video). For example, consider a video in landscape orientation (where the width is greater than the height), with two people depicted on either side of the video, e.g., the first person appears near the left edge of the image and the second person is depicted near the right edge of the image. When such a video is viewed in portrait orientation on a smartphone or other device where the screen has a width less than the height, conventional non-personalized cropping can result in a cropping area that depicts one of the two people, which is selected independently of the viewing user. In contrast, personalized cropping as described herein can automatically select a specific person (e.g., with a significant face) as the focus of the video and select a cropping area that depicts that person. For example, it can be understood that different viewers may find different people depicted in the video to be significant, and thus, the cropping areas for different viewers may be different. More generally, when a video depicts multiple objects, according to the technology described herein, the cropping areas for different viewers can be personalized such that the subject of interest (e.g., significant face, pet, etc.) is retained in the cropped video.
[0067] Figure 5FIG. is a diagram of an example module 500 for automatically cropping a video according to some implementations. A video is obtained at 502. Initial per-frame scoring can be done using an active speaker analysis module 504, a person / facial analysis module 506, or an aesthetics scoring module 508. A personalized score combination module 512 generates a score including a personalized value based on personalized parameters. The combined personalized score can include individual scores from 504 - 508, which are combined with a score based on a personal criterion 510 that includes recognition of important faces that may be captured in the video and / or other criteria. The personalized score can include values from a machine learning model trained to discern important faces of a user in an image, and the model can provide an indication of important faces within a candidate cropping region. The indication of important faces can include the location of each important face identified within the cropping region. An important face can include a face of which the user has previously taken a video or photo using the device; a face that is at least one of the following: occurs at least a threshold number of times; appears in at least a threshold percentage of the user's library, e.g., in at least 5% of the images and videos in the library); appears at least a threshold frequency (e.g., for most years in the library of images / videos, at least once per year); etc.
[0068] In parallel with the personalized score combination 512, motion analysis 514 and / or motion and acceleration cost calculation 516 can be performed. The outputs of the motion cost calculation 516 and the personalized score combination 512 can be used by a minimum cost path finding 518. The output of the minimum cost path finding 518 can be further processed using local optimization and heuristics 520 and then used to crop key frames 522. The output of the minimum cost path finding 518 can also be used to calculate a quality or confidence score 524.
[0069] The quality or confidence score can be used to determine whether to automatically crop the video or not. For example, some videos cannot be well cropped into portraits. A quality or confidence metric indicating that the video cannot be well cropped can indicate to the system not to attempt video cropping and can instead fall back to displaying the video in an up-and-down framed format. In another example, there may be more than one important face in the video, and the faces may be arranged such that cropping would crop one or more important faces from the video. This can be another case where an automatic video cropping operation is not performed.
[0070] Some implementations can include additional input signals for determining where to locate the crop region. The additional input signals can include one or more of the following: video saliency (e.g., not just aesthetics), face quality, object of interest (e.g., person, animal, etc.), or personalization signals (e.g., important face). In some implementations, the system can attempt to programmatically determine who is important in the video using one or more of detecting a camera following a person, who is looking at the camera, the duration of the camera, etc.
[0071] Some implementations can include routines for smoothing camera acceleration. Some implementations can include keyframe interpolation using Bezier splines. In some implementations, the system can control the change of camera speed.
[0072] Figure 6 is a block diagram of an example device 600 that can be used to implement one or more features described herein. In one example, the device 600 can be used to implement a client device, such as Figure 1 any of the client devices 115 shown. Alternatively, the device 600 can implement a server device, such as server 104. In some implementations, the device 600 can be used to implement a client device, a server device, or both a client and a server device. The device 600 can be any suitable computer system, server, or other electronic or hardware device as described above.
[0073] One or more methods described herein can run in a stand-alone program, which can be executed on any type of computing device, as part of another program executed on any type of computing device, or as part of a mobile application (“app”) or a mobile app executed on a mobile computing device (e.g., cellular phone, smart phone, tablet computer, wearable device (wristwatch, armband, jewelry, headgear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted display, etc.), laptop computer, etc.).
[0074] In some implementations, device 600 includes a processor 602, a memory 604, and an input / output (I / O) interface 606. The processor 602 can be one or more processors and / or processing circuits to execute program code and control the basic operations of device 600. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. The processor can include a system having a general-purpose central processing unit (CPU) including one or more cores (e.g., single-core, dual-core, or multi-core configuration), a multi-processing unit (e.g., multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a dedicated circuit for implementing functions, a dedicated processor for implementing neural network model-based processing, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some implementations, the processor 602 can include one or more coprocessors for implementing neural network processing. In some implementations, the processor 602 can be a processor that processes data to produce a probabilistic output. For example, the output produced by the processor 602 may be inaccurate or may be accurate within the range of the expected output. The processing is not limited to a specific geographical location or have a time limit. For example, the processor can perform its functions "in real time", "offline", in "batch mode", etc. Parts of the processing can be performed by different (or the same) processing systems at different times and different locations. A computer can be any processor that communicates with a memory.
[0075] The memory 604 is typically provided in the device 600 for access by the processor 602 and can be any suitable processor-readable storage medium adapted to store instructions for execution by the processor and placed separately from and / or integrated with the processor 602, such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc. The memory 604 can store software operated by the processor 602 on the server device 600, including an operating system 608, a machine learning application 630, other applications 612, and application data 614. The other applications 612 can include applications such as a video cropping application, a data display engine, a web hosting engine, an image display engine, a notification engine, a social network engine, etc. In some implementations, the machine learning application 630 and / or the other applications 612 can include instructions that enable the processor 602 to perform the functions described herein. For example, Figure 4 and Figure 5 some or all of the methods of.
[0076] Other applications 612 may include, for example, video applications, media display applications, communication applications, web hosting engines or applications, mapping applications, media sharing applications, and the like. One or more methods disclosed herein may operate in several environments and platforms, such as as a stand-alone computer program that can run on any type of computing device, as a mobile application (“app”) that runs on a mobile computing device, and the like.
[0077] In various implementations, the machine learning application may utilize Bayesian classifiers, support vector machines, neural networks, or other learning techniques. In some implementations, the machine learning application 630 may include a trained model 634, an inference engine 636, and data 632. In some implementations, the data 632 may include training data, e.g., data used to generate the trained model 634. For example, the training data may include any type of data accessed with user permission, such as pictures or videos taken by the user on the user device, facial identification information of a person depicted in the pictures or videos on the user device, and the like. When the trained model 634 is a model that generates facial signals, the training data may include pictures, videos, and associated metadata.
[0078] The training data may be obtained from any source, such as, for example, a data repository specifically marked for training, licensed data provided as training data for machine learning, and the like. In one or more implementations where one or more users permit the use of their respective user data to train a machine learning model, e.g., the trained model 634, the training data may include such user data.
[0079] In some implementations, the training data may include synthetic data generated for training purposes, such as data not based on user input or activity in the context being trained, e.g., data generated from videos. In some implementations, the machine learning application 630 excludes the data 632. For example, in these implementations, the trained model 634 may be generated, for example, on a different device and provided as part of the machine learning application 630. In various implementations, the trained model 634 may be provided as a data file including the model structure or form and associated weights. The inference engine 636 may read the data file for the trained model 634 and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the trained model 634.
[0080] In some implementations, the trained model 634 can include one or more model forms or structures. For example, the model form or structure can include any type of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., "hidden layers" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that divides or segments input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results of the processing from each tile), a sequence-to-sequence neural network (e.g., a network that takes sequence data such as words in a sentence, frames in a video, etc. as input and produces a result sequence as output), and so on. The model form or structure can specify various nodes and the connectivity between the nodes organized into layers.
[0081] For example, the nodes in the first layer (e.g., the input layer) can receive data as input data 632 or application data 614. For example, when the trained model 634 generates a facial signal, the input data can include a photo or video captured by the user device. Subsequent intermediate layers can receive the output of the nodes in the previous layer as input according to each connectivity specified in the model form or structure. These layers can also be referred to as hidden layers or latent layers.
[0082] The final layer (e.g., the output layer) produces the output of the machine learning application. For example, the output can be an indication of whether a significant face exists in one or more video frames. In some implementations, the model form or structure also specifies the number and / or type of nodes in each layer.
[0083] In different implementations, the trained model 634 can include a plurality of nodes arranged in layers according to a model structure or form. In some implementations, the nodes can be computational nodes without memory, e.g., configured to process one input unit to produce one output unit. The computation performed by a node can include, for example, multiplying each of a plurality of node inputs by weights, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some implementations, the computation performed by a node can also include applying a step / activation function to the adjusted weighted sum. In some implementations, the step / activation function can be a non-linear function. In various implementations, such computations can include operations such as matrix multiplication. In some implementations, the computations performed by multiple nodes can be executed in parallel, e.g., using multiple processor cores of a multi-core processor, individual processing units of a GPU, or dedicated neural circuitry. In some implementations, a node can include memory, e.g., can be capable of storing and using one or more earlier inputs to process subsequent inputs. For example, a node with memory can include a long short-term memory (LSTM) node. An LSTM node can use the memory to maintain a "state" that allows the node to operate like a finite state machine (FSM). A model with such nodes can be useful in processing sequential data - such as words in a sentence or paragraph, frames in a video, speech, or other audio, etc.
[0084] In some implementations, the trained model 634 can include the weights of the respective nodes. For example, the model can be initiated as a plurality of nodes organized into layers specified by the model form or structure. At initialization, corresponding weights can be applied to the connections between each pair of nodes connected according to the model form, e.g., nodes in successive layers of a neural network. For example, the corresponding weights can be assigned randomly, or initialized to default values. Then, the model can be trained, e.g., using data 632, to produce results.
[0085] For example, training can include applying supervised learning techniques. In supervised learning, the training data can include a plurality of inputs (photos or videos) and corresponding expected outputs for each input (e.g., the presence of one or more important faces, etc.). Based on the comparison of the output of the model with the expected output, the values of the weights are automatically adjusted, e.g., in a manner that increases the probability that the model produces the expected output when provided with similar inputs.
[0086] In some implementations, training can include applying unsupervised learning techniques. In unsupervised learning, only input data can be provided, and a model can be trained to distinguish the data, e.g., to cluster the input data into multiple groups, where each group includes input data that is similar in some way, e.g., having similar prominent faces present in a photo or video frame. For example, a model can be trained to distinguish video frames or cropped rectangles that contain prominent faces from video frames or cropped rectangles that contain non-prominent faces or frames that do not contain faces.
[0087] In some implementations, unsupervised learning can be used to generate a knowledge representation, e.g., which can be used by a machine learning application 630. For example, unsupervised learning can be used to generate the personalized parameter signals utilized as described above with reference to Figure 4 and 5 In various implementations, a trained model includes a set of weights corresponding to the model structure. In implementations where data 632 is omitted, the machine learning application 630 can include a trained model 634 based on, e.g., prior training performed by a developer of the machine learning application 630, by a third party, etc. In some implementations, the trained model 634 can include a set of fixed weights downloaded, e.g., from a server that provides the weights.
[0088] The machine learning application 630 also includes an inference engine 636. The inference engine 636 is configured to apply the trained model 634 to data such as application data 614 to provide an inference. In some implementations, the inference engine 636 can include software code to be executed by the processor 602. In some implementations, the inference engine 636 can specify a circuit configuration (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.) that allows the processor 602 to apply the trained model. In some implementations, the inference engine 636 can include software instructions, hardware instructions, or a combination. In some implementations, the inference engine 636 can provide an application programming interface (API) that can be used by the operating system 608 and / or other applications 612 to call the inference engine 636, e.g., to apply the trained model 634 to the application data 614 to generate an inference. For example, an inference for a prominent face model can be, e.g., a classification of a video frame or cropped rectangle based on a comparison with previously captured photos or videos having one or more prominent faces.
[0089] The machine learning application 630 can offer several technical advantages. For example, when a trained model 634 is generated based on unsupervised learning, the trained model 634 can be applied by the inference engine 636 to generate a knowledge representation (e.g., a numerical representation) according to input data (e.g., application data 614). For example, a model trained to generate facial signals can generate a representation of a call with a smaller data size (e.g., 1 KB) than the input audio recording (e.g., 1 MB). In some implementations, these representations can help reduce the processing costs (e.g., computational costs, memory usage, etc.) for generating outputs (e.g., labels, classifications, etc.).
[0090] In some implementations, these representations can be provided as inputs to different machine learning applications that generate outputs from the output of the inference engine 636. In some implementations, the knowledge representations generated by the machine learning application 630 can be provided to different devices for further processing, e.g., via a network. For example, facial signals generated using the techniques described in reference Figure 4 or 5 can be provided to a client device for automatic video cropping using personalized parameters, as described in reference Figure 4 or 5. In such implementations, providing the knowledge representation instead of a photo or video of the important face can offer technical benefits, such as enabling faster data transmission at a reduced cost. In another example, a model trained to cluster important faces can generate clusters from an input photo or video. The clusters can be suitable for further processing (e.g., determining whether an important face exists in a video frame or a cropping rectangle, etc.) without accessing the original photo or video, and thus saving computational costs.
[0091] In some implementations, the machine learning application 630 can be implemented in an offline manner. In these implementations, the trained model 634 can be generated in a first stage and provided as part of the machine learning application 630. In some implementations, the machine learning application 630 can be implemented in an online manner. For example, in such an implementation, an application that invokes the machine learning application 630 (e.g., the operating system 608, one or more other applications 612) can utilize the inferences generated by the machine learning application 630, e.g., to provide inferences to the user, and can generate system logs (e.g., if the user permits, the user takes an action based on the inference; or if used as an input for further processing, the results of the further processing). The system logs can be generated periodically - e.g., hourly, monthly, quarterly, etc. - and can be used with the user's permission to update the trained model 634, e.g., to update the important face data of the trained model 634.
[0092] In some implementations, the machine learning application 630 may be implemented in a manner that can adapt to the specific configuration of the device 600 on which the machine learning application 630 executes. For example, the machine learning application 630 may determine a computation graph that utilizes the available computing resources such as the processor 602. For example, the machine learning application 630 may determine that the processor 602 includes a GPU having a specific number of GPU cores (e.g., 1000), and accordingly implement an inference engine (e.g., as 1000 separate processes or threads).
[0093] In some implementations, the machine learning application 630 may implement a collection of trained models. For example, the trained model 634 may include multiple trained models each applicable to the same input data. In these implementations, the machine learning application 630 may select a specific trained model, for example, based on the available computing resources, the success rate of previous inferences, and the like. In some implementations, the machine learning application 630 may execute the inference engine 636 such that multiple trained models are applied. In these implementations, the machine learning application 630 may, for example, use a voting technique to score the individual outputs from each applied trained model, or combine the outputs from the applied individual models by selecting one or more specific outputs. Further, in these implementations, the machine learning application may apply a time threshold (e.g., 0.5 ms) for applying an individual trained model, and utilize only those individual outputs that are available within the time threshold. Outputs not received within the time threshold may not be utilized, for example, discarded. For example, such an approach may be appropriate when there is a time limit specified, for example, by the operating system 608 or one or more applications 612 when invoking the machine learning application, such as automatically cropping a video based on whether one or more important faces are detected and other personalized criteria.
[0094] In different implementations, the machine learning application 630 may produce different types of outputs. For example, the machine learning application 630 may provide a representation or clustering (e.g., a numerical representation of the input data), a label (e.g., for input data including images, documents, audio recordings, etc.), and the like. In some implementations, the machine learning model 630 may produce an output based on a format specified by the calling application, such as the operating system 608 or one or more applications 612.
[0095] Any software in the memory 604 may alternatively be stored in any other suitable storage location or computer-readable medium. Additionally, the memory 604 (and / or other connected storage devices) may store one or more messages, one or more taxonomies, electronic encyclopedias, dictionaries, glossaries, knowledge bases, message data, grammars, facial identifiers (e.g., significant faces), and / or other instructions and data used in the features described herein. The memory 604 and any other type of storage (disk, optical disk, magnetic tape, or other tangible medium) may be considered "storage" or "storage devices".
[0096] The I / O interface 606 may provide functionality that enables the device 600 to interface with other systems and devices. The interfaced devices may be included as part of the device 600 or may be separate and communicate with the device 600. For example, network communication devices, storage devices (e.g., memory and / or database 106), and input / output devices may communicate via the I / O interface 606. In some implementations, the I / O interface may be connected to interface devices such as input devices (keyboard, pointing device, touch screen, microphone, camera, scanner, sensors, etc.) and / or output devices (display device, speaker device, printer, motor, etc.). The I / O interface 606 may also include a telephone interface, e.g., for coupling the device 600 to a cellular network or other telephone network.
[0097] Some examples of interface devices that may be connected to the I / O interface 606 may include one or more display devices 620 that may be used to display content - such as images, videos, and / or user interfaces of output applications as described herein. The display device 620 may be connected to the device 600 via a local connection (e.g., a display bus) and / or via a network connection, and may be any suitable display device. The display device 620 may include any suitable display device, such as an LCD, LED, or plasma display screen, CRT, television, monitor, touch screen, 3D display, or other visual display device. For example, the display device 620 may be a flat display screen provided on a mobile device, multiple display screens provided in goggles or a head-mounted device, or a monitor screen for a computer device.
[0098] For ease of illustration, Figure 6A box is shown for each of the processor 602, the memory 604, the I / O interface 606, and the software boxes 608, 612, and 630. These boxes may represent one or more processors or processing circuits, an operating system, memory, an I / O interface, an application, and / or a software module. In other implementations, device 600 may not have all of the components shown and / or may have other elements, including other types of elements instead of or in addition to those shown herein. Although some components are described as performing the boxes and operations as described in some implementations herein, any suitable component or combination of components of environment 100, device 600, a similar system, or any suitable one or more processors associated with such a system may perform the described boxes and operations.
[0099] The methods described herein may be implemented by computer program instructions or code that may be executed on a computer. For example, the code may be implemented by one or more digital processors (e.g., a microprocessor or other processing circuit) and may be stored on a computer program product that includes a non-transitory computer-readable medium (e.g., a storage medium), such as a magnetic, optical, electromagnetic, or semiconductor storage medium, including semiconductor or solid-state memory, magnetic tape, removable computer disk, random access memory (RAM), read-only memory (ROM), flash memory, hard disk, optical disk, solid-state memory drive, etc. The program instructions may also be included in an electronic signal and provided as an electronic signal, e.g., in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods may be implemented in hardware (e.g., logic gates, etc.) or in a combination of hardware and software. Example hardware may be a programmable processor (e.g., a field programmable gate array (FPGA), a complex programmable logic device), a general-purpose processor, a graphics processor, an application-specific integrated circuit (ASIC), etc. One or more methods may be performed as part of or a component of an application running on a system, or as an application or software running in combination with other applications and an operating system.
[0100] Although this specification has been described with respect to its particular implementations, these particular implementations are merely illustrative and not restrictive. The concepts illustrated in the examples may be applied to other examples and implementations.
[0101] In some implementations discussed herein, where personal information about a user may be collected or used (e.g., user data, facial recognition data, information about the user's social network, the user's location and time at that location, the user's biometric information, the user's activities and demographic information), the user is provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein specifically collect, store, and / or use a user's personal information to do so when receiving explicit authorization from the relevant user. For example, the user is provided with control over whether a program or feature collects user information about that particular user or other users associated with the program or feature. One or more options are presented to each user whose personal information is to be collected to allow control over the information collection associated with that user, providing permission or authorization regarding whether information is collected and regarding which portions of the information are to be collected. For example, one or more such control options may be provided to the user via a communication network. Additionally, certain data may be processed in one or more ways before it is stored or used such that personally identifiable information is removed. As an example, a user's identity may be processed such that personally identifiable information cannot be determined. As another example, the geographical location of a user device may be generalized to a larger area such that the user's specific location cannot be determined.
[0102] Note that, as is known to those skilled in the art, the functional blocks, operations, features, methods, devices, and systems described in this disclosure may be integrated or divided into different combinations of systems, devices, and functional blocks. Any suitable programming language and programming technique may be used to implement the routines of a particular implementation. Different programming techniques may be employed, such as procedural or object-oriented. The routines may be executed on a single processing device or multiple processors. Although steps, operations, or calculations may be presented in a particular order, the order may be changed in different particular implementations. In some implementations, multiple steps or operations shown as sequential in this specification may be executed simultaneously.
Claims
1. A computer-implemented method, comprising: obtaining an input video comprising a plurality of frames; determining a per-frame cropping score for each of one or more candidate cropping regions in each frame of the input video; using a trained machine learning model to generate a face signal for the one or more candidate cropping regions within each frame of the input video; adjusting each per-frame cropping score based on the face signal for the one or more candidate cropping regions; determining a minimum-cost path representing a cropping region location of the input video based on a motion cost and the adjusted per-frame cropping scores for the one or more candidate cropping regions; generating cropping key frames corresponding to the cropping region locations along the minimum-cost path, wherein the cropping key frames include a start frame, an end frame, and a cropping region location; and outputting a modified video having one or more of an output aspect ratio or an output orientation different from a corresponding input aspect ratio or an input orientation of the input video, wherein the input aspect ratio or input orientation is a parameter used during capture of the input video.
2. The computer-implemented method according to claim 1, wherein adjusting each per-frame cropping score comprises one of the following: if it is determined that a face exists in the candidate cropping region corresponding to the per-frame cropping score, increasing the per-frame cropping score by a first value; or if it is determined that at least one significant face exists in the candidate cropping region corresponding to the per-frame cropping score, increasing the per-frame cropping score by a second value, wherein the second value is greater than the first value.
3. The computer-implemented method according to claim 1, further comprising: determining a quality score for the cropping key frames; and performing automatic video cropping of the input video based on the quality score.
4. The computer-implemented method according to claim 1, further comprising: determining a confidence score for the cropping key frames; and performing automatic video cropping of the input video based on the confidence score.
5. The computer-implemented method according to claim 1, wherein determining the per-frame cropping score comprises: determining one or more of the following for each candidate cropping region: an aesthetic score, a face analysis score, or the presence of an active speaker.
6. The computer-implemented method according to claim 1, wherein generating the cropping key frames comprises: interpolating between two key frames.
7. The computer-implemented method according to claim 6, wherein the interpolation comprises: applying a B-spline.
8. The computer-implemented method according to claim 1, wherein generating the face signal comprises: accessing one or more personalized parameters.
9. The computer-implemented method according to claim 8, wherein the one or more personalized parameters include face recognition information for one or more significant faces.
10. The computer-implemented method according to claim 1, wherein outputting the modified video comprises: displaying the modified video on a display.
11. The computer-implemented method according to claim 1, further comprising: Before obtaining the input video, receive a video playback command at a device; In response to receiving the video playback command, detect the device orientation and display aspect ratio of the device; And Determine a cropping region based on the device orientation and the display aspect ratio of the device.
12. A non-transitory computer-readable medium having stored software instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations include: Obtain an input video including a plurality of frames; Determine a per-frame cropping score for each of one or more candidate cropping regions in each frame of the input video; Use a trained machine learning model to generate a face signal for each of the one or more candidate cropping regions within each frame of the input video; Adjust each per-frame cropping score based on the face signals of the one or more candidate cropping regions; Determine a minimum cost path representing the cropping region location of the input video based on a motion cost and the adjusted per-frame cropping scores of the one or more candidate cropping regions; Generate cropping key frames corresponding to the cropping region locations along the minimum cost path, where the cropping key frames include a start frame, an end frame, and a cropping region location; And Output a modified video having one or more of an output aspect ratio or an output orientation different from a corresponding input aspect ratio or input orientation of the input video, where the input aspect ratio or input orientation is a parameter used during capture of the input video.
13. The non-transitory computer-readable medium according to claim 12, wherein, Adjusting each per-frame cropping score includes one of the following: If it is determined that a face exists in the candidate cropping region corresponding to the per-frame cropping score, increase the per-frame cropping score by a first value; Or If it is determined that at least one important face exists in the candidate cropping region corresponding to the per-frame cropping score, increase the per-frame cropping score by a second value, where the second value is greater than the first value.
14. The non-transitory computer-readable medium according to claim 12, wherein, The operations further include: Determine a quality score for the cropping key frames; and Perform automatic video cropping of the input video based on the quality score.
15. The non-transitory computer-readable medium according to claim 12, wherein, The operations further include: Determine a confidence score for the cropping key frames; and Perform automatic video cropping of the input video based on the confidence score.
16. The non-transitory computer-readable medium according to claim 12, wherein, Determining the per-frame cropping score includes: Determine one or more of the following for each candidate cropping region: an aesthetic score, a face analysis score, or the presence of an active speaker.
17. The non-transitory computer-readable medium according to claim 12, wherein, Generating the cropping key frames includes: interpolating between two key frames.
18. The non-transitory computer-readable medium according to claim 17, wherein, The interpolation includes: applying a B-spline.
19. The non-transitory computer-readable medium according to claim 12, wherein, generating the face signal includes: accessing one or more personalized parameters.
20. A system, comprising: one or more processors, the one or more processors coupled to a non-transitory computer-readable medium having software instructions stored thereon, the software instructions, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: obtaining an input video including a plurality of frames; determining a per-frame cropping score for each of one or more candidate cropping regions in each frame of the input video; using a trained machine learning model to generate a face signal for each of the one or more candidate cropping regions within each frame of the input video; adjusting each per-frame cropping score based on the face signal of the one or more candidate cropping regions; determining a minimum-cost path representing a cropping region location of the input video based on a motion cost and the adjusted per-frame cropping scores of the one or more candidate cropping regions; generating cropping key frames corresponding to the cropping region locations along the minimum-cost path, wherein the cropping key frames include a start frame, an end frame, and a cropping region location; and outputting a modified video having one or more of an output aspect ratio or an output orientation different from a corresponding input aspect ratio or an input orientation of the input video, wherein the input aspect ratio or the input orientation is a parameter used during capture of the input video.
Citation Information
Patent Citations
Temporal occlusion costing applied to video editing
US20080260347A1
Dynamically cropping digital content for display in any aspect ratio
US20170249719A1