Space-time video saliency analysis

By using a unified spatiotemporal framework to predict the saliency of video frames through machine learning and artificial intelligence, the problem of analyzing salient moments in video clips is solved, enabling efficient utilization of video resources and improved playback efficiency.

CN121533029APending Publication Date: 2026-02-13APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380100520.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently analyze and classify significant moments in video clips, leading to wasted resources and unnecessary video playback time.

Method used

Using a unified spatiotemporal framework, machine learning and artificial intelligence methods are employed to predict the saliency of video frames and adjust the frame rate and compression rate to identify regions of interest in the video.

Benefits of technology

It reduces the computing resources and storage space required for video playback, improves the efficiency of video playback, and allows viewers to watch only the significant parts, saving time and resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121533029A_ABST
    Figure CN121533029A_ABST
Patent Text Reader

Abstract

A multi-stage training method is used to train a neural network to identify spatio-temporal saliency regions in a video frame using supervised eye gaze, action tags, and frame annotations. A first stage is used to train a neural network to detect spatially salient regions in video frames, and a second stage is used to train a neural network to detect temporally salient regions in these video frames. Events occurring within the video frames are predicted based on the identified regions. Frames of a video segment having salient regions may be played back at a lower compression rate and / or at an increased frame rate relative to frames of the video segment having non-salient regions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to video saliency, and more specifically to performing spatio-temporal saliency analysis in videos. BACKGROUND

[0002] Most modern smartphones include integrated digital camera technology. As technology has advanced, digital cameras integrated with smartphones rival the capabilities provided by standalone digital cameras. As more and more smartphones are used to capture video, the demand for resources increases, especially as the quality of the video produced by the integrated digital cameras increases. In today's era, it can be preferable to be able to more intelligently analyze and / or classify video segments of captured video. SUMMARY

[0003] The present embodiments can particularly relate to systems and methods for providing a unified spatio-temporal framework for capturing and predicting salient moments in media files, such as video clips. In some embodiments, the unified framework is used to predict spatio-temporal saliency of a media file, such as a video file, clip, segment, etc., during capture by an image sensor, such as a camera. The captured media file can be processed using file compression techniques based on the predictions made. For example, the predictions made by the unified framework can affect the video compression rate of different video frames within the same media file. Additionally, during playback of the media file, the frame rate can be adjusted based on the predictions made by the unified framework.

[0004] In some embodiments of the present disclosure, spatial saliency prediction and detection provides identification of regions of interest within each frame of a media file. For example, regions of interest in a video can include portions of a video frame in which action is occurring. Temporal saliency detection can provide identification of interesting video frames over time, such as when something particular occurs during a media file. Additionally, the unified spatio-temporal framework predicts salient actions occurring within a media file, as well as, for example, the type of salient action and / or the location within various frames of the media file that represents such salient actions. The described systems and methods can result in additional, less, or alternative actions, including those discussed elsewhere herein.

[0005] Various non-transitory computer-readable medium embodiments are disclosed herein. Such computer-readable media can be readable by one or more processors. Instructions can be stored on a computer-readable medium for causing one or more processors to perform any of the techniques disclosed herein.

[0006] According to the above-enumerated embodiments of program storage devices, various programmable electronic devices are also disclosed herein. Such electronic devices can include one or more image capture devices, such as optical image sensors / camera units; a display; a user interface; one or more processors; and a memory coupled to the one or more processors. Instructions can be stored in the memory that cause the one or more processors to perform the instructions according to the various techniques disclosed herein. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 A simplified network diagram is shown in block diagram form in accordance with one or more embodiments.

[0008] Figure 2 A simplified diagram of an electronic device is shown in block diagram form in accordance with one or more embodiments.

[0009] Figure 3 An example framework for determining spatiotemporal saliency for a video segment using a network trained to perform action recognition is shown in accordance with one or more embodiments.

[0010] Figure 4 An example framework including multiple components of a subnetwork for performing temporal saliency analysis of video data is shown in accordance with one or more embodiments.

[0011] Figure 5 An example framework for training a network to determine spatiotemporal saliency for a video segment using a multi-task learning approach is shown in accordance with one or more embodiments.

[0012] Figure 6 An example framework for training a network to determine spatiotemporal saliency for a video segment using a single-stage training strategy is shown in accordance with one or more embodiments.

[0013] Figure 7 An example framework for training a network to determine spatiotemporal saliency for a video segment using a multi-stage training strategy is shown in accordance with one or more embodiments.

[0014] Figure 8 An example network architecture and framework for determining spatiotemporal saliency for a video segment is shown in accordance with one or more embodiments.

[0015] Figure 9 An example method for performing spatiotemporal saliency analysis for a video segment is shown in flowchart form in accordance with one or more embodiments. DETAILED DESCRIPTION

[0016] The following disclosure relates to technical improvements for detecting salient portions of a media file (e.g., a video clip) using a unified framework that performs spatio-temporal saliency analysis via the use of machine learning (ML) and / or artificial intelligence (AI) based methods. According to aspects of the disclosure, systems, methods, and computer-readable media for performing spatio-temporal saliency analysis of a media file are provided. The media file can be captured, for example, by a mobile device of a user. For example, the media file can be captured using one or more camera inputs, one or more microphone inputs, or a combination thereof of the mobile device of the user. Additionally or alternatively, the media file can be transmitted, such as by another mobile device or a media server, to the mobile device of the user, for example, over a network. In some embodiments, the mobile device can be a user computing device, a tablet computing device, or the like.

[0017] According to one or more embodiments, the disclosed technology can include one or more software modules embodied on a mobile device of a user. The one or more software modules can include instructions for performing spatio-temporal saliency analysis of a media file. For example, a media file captured by the mobile device of the user can be received as input by the one or more software modules. Using a predictive model built with artificial intelligence and / or machine learning based techniques, the one or more software modules can provide an indication (e.g., a label, a tag, or the like) of portions (e.g., video frames) of the media file as output. In some embodiments, the predictive model can also predict a particular type of action that occurs within a particular portion of the media file. For example, when a video is captured during a swimming / diving competition, the unified spatio-temporal framework can not only predict when a salient action (e.g., a dive) occurs (e.g., a time indication, such as a timestamp, within the captured video when the dive occurs), but can also predict where a certain type of action occurs in the video clip or video frame. In some embodiments, a neural network can be trained to identify a particular type of action (e.g., a forward dive, a twist dive, a backward dive, or the like) that occurs within a salient portion of a video clip.

[0018] According to one or more embodiments, the disclosed technology addresses the need in the art to predict and identify interesting or salient portions of media files. For example, video frames of a media file that a viewer will find interesting are predicted. The interesting video frames can include, for example, action sequences (e.g., fight scenes, chase scenes, battle scenes), human expressions (e.g., changes in facial expressions, surprised actions), and / or interactions (e.g., conversations, kisses), animal activity (e.g., nature scenes, animal chases), etc. Predicting and identifying salient portions of media files provides a number of technical advantages, such as enabling a viewer to only playback the salient portions of a media file, thereby not only saving the viewer's time, but also reducing the resource usage of the viewer's device for providing the media file playback. For example, a viewer can choose to only watch the salient portions of a media file that are five minutes in length. In this example, the media file can be related to a track and field event, and the viewer can only want to watch the actual event (e.g., the race), and not the other aspects of the media file (e.g., introductions of the participants, pre-race commentary, etc.) for the entire duration.

[0019] In some embodiments, video playback of a media file is improved by compressing the media file and reducing the amount of computer resources needed to provide the video playback. In this example, portions of the media file that do not include salient portions as predicted and identified by the unified spatio-temporal framework described herein can undergo image data compression. Image data compression reduces the size of the media file, thereby making the media file require less space, such as on a storage device or memory of a viewer's mobile device or on another storage option (e.g., cloud storage, remote media server). Additionally or alternatively, compressing the media file will require less network resources (e.g., bandwidth) when streamed from the remote storage option to the viewer's mobile device. Continuing with this example, portions of the media file that do include salient portions will not undergo any type of image or video compression. By not compressing the video frames that include salient regions, playback of the uncompressed video frames can be provided at the original resolution. For example, if the media file was captured at 1080p, the video frames that include salient regions can be played back at 1080p during playback of the media file, and the video frames that do not include salient regions can be played back at a lower resolution, such as 720p or lower. Additionally or alternatively, the entire media file can undergo video compression, however, the video frames that include salient regions can be compressed at a lower rate when compared to the video frames that lack salient regions.

[0020] In some embodiments, video playback of a media file is improved by adjusting the frame rate during playback of the media file by a viewer's mobile device. Adjusting the frame rate during playback can reduce the amount of computer resources required by the viewer's mobile device to provide video playback. In this example, video frames of the media file that include salient portions as predicted and identified by the unified spatio-temporal framework described herein can be played back at a higher frame rate when compared to video frames of the media file that lack salient portions during playback. By reducing the frame rate of video frames that are deemed to be less salient during playback, the burden on system resources can be reduced. For example, during playback of a media file, less demand is placed on a processor, such as a video processor, when the frame rate is reduced. In this example, video frames containing salient regions can be displayed at a first frame rate, and video frames lacking salient regions can be displayed at a second frame rate, where the first frame rate is higher than the second frame rate. For example, the first frame rate can be 60 fps, which is suitable for 4K video resolution, and the second frame rate can be 24 fps, which is suitable for streaming video content. Other frame rates can be used, such as 30 fps, which is suitable for live TV broadcasts (e.g., sports, news), and 120 fps, which is deemed suitable for slow motion videos and video games.

[0021] In some embodiments, video playback of a media file is improved by adjusting the video compression rate of the media file during playback and adjusting the frame rate of the media file. For example, based on the results of the unified spatio-temporal framework described herein, video frames of the media file that are identified as including salient regions can be compressed at a first compression rate and played back at a first frame rate. Additionally, video frames of the media file that lack salient regions can be compressed at a second compression rate and played back at a second frame rate. In this example, the first compression rate is less than the second compression rate. Furthermore, the first frame rate is greater than the second frame rate. By adjusting the video compression rate and playback frame rate, computer and network resource usage can be reduced. Additionally, in this example, video frames that include salient regions (e.g., interesting moments in a video) can be played back to a viewer in an unaltered state.

[0022] Image data compression reduces the size of the media file, thereby requiring less space for the media file, such as on a storage device or memory of a viewer's mobile device or on another storage option (e.g., cloud storage, remote media server). Additionally or alternatively, when streamed from a remote storage option to a viewer's mobile device, compressed media files will require less network resources (e.g., bandwidth). Continuing with the example, portions of the media file that do include a significant portion of the salient area will not undergo any type of image or video compression. By not compressing video frames that include the salient area, playback of the uncompressed video frames can be provided at the original clarity. For example, if the media file was captured at 1080p, during playback of the media file, video frames that include the salient area can be played back at 1080p, and video frames that do not include the salient area can be played back at a lower clarity, such as 720p or lower.

[0023] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the disclosed concept. As part of this description, some of the diagrams presented herein are in block diagram form that shows architectures and devices in order to avoid obscuring the novel aspects of the disclosed implementations. In this context, reference to numbered drawing elements with no associated identifier (e.g., 100) refers to all instances of the drawing element with the identifier (e.g., 100a and 100b). Also, the diagrams of the disclosure can be presented in the form of flow diagrams. The blocks in any flow diagram can be presented in a particular order. However, the flow of any flow diagram is used merely for illustrative purposes as one implementation. In other implementations, any of the various components depicted in the flow diagram can be deleted, or can be performed in a different order, or even at the same time. Further, other implementations can include additional steps that are not depicted as part of the flow diagram. The language used in the disclosure has been principally selected for readability and instructional purposes and can not have been selected to

[0024] It is to be understood that in the development of any actual implementation (as in any development project), numerous decisions must be made to achieve the developer's specific goals (e.g., compliance with system- and business-related constraints), and that these goals can vary from one implementation to another. It is also to be understood that such development effort might be complex and time-consuming, but can nevertheless be a routine undertaking for those of ordinary skill in the art having the benefit of this disclosure on image capture.

[0025] For the purposes of this disclosure, images captured by a camera device are referred to as "media files." However, in one or more embodiments, the captured images referred to as "media files" can be any audiovisual media data, such as video clips, audio segments, movies, song files, music videos, etc.

[0026] Referring now to the drawings Figure 1 A simplified block diagram of a unified spatio-temporal service and framework 100 is depicted, including a spatio-temporal saliency service 102 connected to a client device 140 over a network (e.g., over a network 150). The client device 140 can be a personal computer or a multi-function device such as a mobile phone, tablet computer, personal digital assistant, portable music / video player, wearable device, or any other electronic device that includes a media playback system.

[0027] The spatio-temporal saliency service 102 can include one or more servers or other computing or storage devices on which various modules and storage devices can be included. Although the spatio-temporal saliency service 102 is depicted as including various components in an exemplary manner, in one or more embodiments, the various components and functions can be distributed across multiple network devices such as servers, network storage devices, etc. In addition, additional components can be used, and certain combinations of the functionality of any components can also be combined. In general, the spatio-temporal saliency service 102 can include one or more memory devices 112, one or more storage devices 114, and one or more processors 116 such as central processing units (CPUs) or graphics processing units (GPUs). In addition, the processors 116 can include multiple processors of the same or different types. The memories 112 can each include one or more different types of memory that can be used in connection with the processors 116 to perform device functions. For example, the memories 112 can include cache, ROM, and / or RAM. The memories 112 can store various programming modules during execution, including a training module 104 and an integrated spatio-temporal saliency module 106A.

[0028] The spatio-temporal saliency service 102 can store media files, media file data, saliency prediction data, neural network model training data, etc. Additional data can be stored by the spatio-temporal saliency service 102, including but not limited to media classification data, frame annotation data, and action label data. The spatio-temporal saliency service 102 can store this data in a media repository 118 within the storage devices 114. The storage devices 114 can include one or more physical storage devices. The physical storage devices can be located within a single location, or can be distributed across multiple locations such as multiple servers.

[0029] In another implementation, the media repository 118 can include model training data for creating datasets to train prediction models using the training module 104. The model training data can include, for example, labeled training data used by the training module 104 to train machine learning (ML) models, such as the integrated spatio-temporal saliency module 106A or the integrated saliency module 106B. As described herein, the trained ML models can then be used to predict salient regions, frames, or portions of media files. In some implementations, the training data can include data pairs of input data and output data. For example, the input data can include media items that have been processed with the output data, such as temporal annotations, action labels, etc. The temporal annotations and / or action labels can be objective or subjective. Additionally, the temporal annotations and / or action labels can be obtained through detailed experiments by human viewers. In another implementation, the output data can be obtained through detailed experiments based on eye gaze of a human viewer during playback of a media item. The objective labels can indicate, for example, quality measurements obtained by computing video frame quality metrics. The video frame quality metrics can be based at least in part or in combination on frame rate, resolution, etc. The subjective labels can indicate, for example, spatio-temporal saliency scores assigned by human annotators.

[0030] Returning to the spatio-temporal saliency service 102, the memory 112 includes modules that include computer-readable code executable by the processor 116 to cause the spatio-temporal saliency service 102 to perform various tasks. As depicted, the memory 112 can include the training module 104 and the integrated spatio-temporal saliency module 106A. According to one or more implementations, the training module 104 generates and maintains prediction models for the integrated spatio-temporal saliency module 106A of the spatio-temporal saliency service 102 and / or the integrated saliency module 106B of the client device 140. The prediction models, as well as additional data related to the models, including but not limited to training data (e.g., spatial saliency training data), classification data, label data, etc., can be stored on a local device, such as the storage device 114, on a remote storage device, such as the storage device 124 of the client device 140, or a combination or variation thereof.

[0031] The memory 112 also includes an integrated spatio-temporal saliency module 106A. In one or more embodiments, the integrated spatio-temporal saliency module 106A has access to a machine learning model (e.g., a predictive model generated by the training module 104) that includes a video saliency tool for detecting frames of interest or salient frames within a media file. The ML model can be accessed from the storage device 114 or from another storage location (e.g., a remote storage location) over a network, such as the network 150. For example, the video saliency tool can be used to detect spatio-temporal salient regions of video frames within a media file. The media file with the detected spatio-temporal salient regions can be categorized and tagged by the integrated spatio-temporal saliency module 106A. The media file processed by the integrated spatio-temporal saliency module 106A can be stored in the media repository 118 of the storage device 114. Additionally or alternatively, the processed media file can be transmitted or streamed to another device, such as the client device 140, over a network, such as the network 150.

[0032] In some embodiments, an integrated spatio-temporal saliency module can be provided on a mobile device of a spectator, such as the integrated saliency module 106B on the client device 140. In this example, the integrated saliency module 106B can include a video saliency tool that detects spatio-temporal salient regions within video frames of a media file in the same, if not similar, manner as described above with respect to the integrated spatio-temporal saliency module 106A. The integrated saliency module 106B can process a media file in real-time or near real-time, or a combination or variation thereof, as the event is captured by the client device 140. For example, when the client device 140 is used to capture and record an event or scene using sensor inputs (e.g., a microphone and / or a camera) of the client device 140. Alternatively, the media file can be obtained from a media repository, such as the media repository 130, or received from a remote storage location, such as the media repository 118, and provided as input to the integrated saliency module 106B. In some embodiments, a media file can be processed or provided as input to both the integrated spatio-temporal saliency module 106A and the integrated saliency module 106B. In this example, both modules 106A and 106B can process the media file in part or in whole. The captured and processed content can be stored within the media repository 130 of the storage device 124 and played back to the spectator using the media player 126. Alternatively, the processed media file can be stored in a remote storage location over a network, such as in the media repository 114 of the spatio-temporal saliency service 102 over the network 150.

[0033] Referring now to Figure 2 , a simplified functional block diagram of an illustrative multifunction device 200 is shown in accordance with one embodiment. The multifunction device 200 can be representative, for example, of the client device 140 Figure 1representative components of the device and client device 140 of the spatiotemporal saliency service 102. The multi-function electronic device can include a processor 205, a display 210, a user interface 215, graphics hardware 220, device sensors 225, communication circuitry 245, video codecs 255, memory 260, storage 265, and a communication bus 270. The multi-function electronic device can be, for example, a personal computing device such as a personal digital assistant (PDA), a personal music player, a mobile phone, a tablet computer, a laptop computer, etc.

[0034] The processor 205 can execute instructions necessary to control the operations performed by the multi-function device 200 (e.g., such as the spatiotemporal saliency prediction methods as disclosed herein). The processor 205 may, for example, drive the display 210 and receive user input from the user interface 215. The user interface 215 can allow a user to interact with the device 200. For example, the user interface 215 can take a variety of forms such as buttons, keypads, dials, click wheels, keyboards, display screens, and / or touch screens. The processor 205 can also be, for example, a system-on-a-chip such as those found in mobile devices and include a dedicated graphics processing unit (GPU). The processor 205 can be based on a reduced instruction set computer (RISC) or complex instruction set computer (CISC) architecture or any other suitable architecture, and can include one or more processing cores. The graphics hardware 220 can be a specialized computing hardware for processing graphics and / or assist the processor 205 in processing graphics information. In one embodiment, the graphics hardware 220 can include a programmable GPU.

[0035] The memory 260 can include one or more different types of media used by the processor 205 and the graphics hardware 220 to perform device functions. For example, the memory 260 can include memory cache, read-only memory (ROM), and / or random access memory (RAM). The storage 265 can store media (e.g., audio, image, and video files), computer program instructions or software, preference information, device profile information, model training data, and any other suitable data. The storage 265 can include one or more non-transitory computer-readable storage media including, for example, magnetic disks (fixed, floppy, and removable disks), and magnetic tapes, optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such as Electrically Programmable Read Only Memories (EPROM), and Electrically Erasable Programmable Read Only Memories (EEPROM). The memory 260 and the storage 265 can be used to tangibly retain computer program instructions or codes organized into one or more modules and written in any desired computer programming language. Such computer program codes, when executed by, for example, the processor 205, can implement one or more of the methods described herein.

[0036] Spatiotemporal saliency for video segments is determined using a network trained to perform action recognition.

[0037] Now go to Figure 3 This illustrates a framework 300 for determining the spatiotemporal saliency of a video segment using a network trained to perform action recognition, according to one or more embodiments. Framework 300 begins at box 302 with a signal generated by an electronic device (e.g., Figure 1 The client device 140 (within the context of training module 104 and integrated spatiotemporal saliency modules 106A / 106B) receives input frames, such as image frames from a media file described herein and above. In this example, the media file includes three video frames (302(a), 302(b), 302(c)); however, it should be understood that a media file may contain more or fewer video frames than the example provided. For example, a media file may contain thousands of video frames (e.g., a ten-minute video at 30fps would contain 18,000 frames). In this example, the first frame may be provided at time t, the second frame at time t+1, the third frame at time t+2, and so on. As described above, the media file may be captured by the client device 140 or received from another device via a network (such as network 150). As shown, each input frame in the input frames of the media file may be processed by a frame. At box 304, each input frame in the input frames is provided as input to the neural network. Using a neural network, a saliency mask at box 306 is used to mask portions of the input frame that will not be used to classify the spatial saliency of the input frame. For example, the saliency mask could be an attention mask in the form of a binary mask, which can be used to guide the extraction of spatial information from the frame. At box 308, a feature map of the input frame is generated. For example, the feature map can be generated by applying one or more filters or feature detectors to the input frame. Alternatively, the feature map can be generated based on previous video frames used as input. Action labels, such as annotation labels, can be assigned frame-by-frame using a spatial saliency network.

[0038] At box 310, a recurrent neural network such as ConvLSTM (Convolutional Long Short-Term Memory) is used to predict the spatiotemporal saliency portion based on each input frame and one or more previous frames in the input frame. The output from box 310 is fed to boxes 312 and 314. At box 312, the output from box 310 is used to perform an average pooling operation on the feature map generated by box 310. At box 314, the temporal saliency of the input frame is predicted (see...). Figure 4) and represented as a value ranging from zero (0) to one (1), where, for example, a value of zero indicates that a frame is extremely unlikely to contain salient content, and a value of one indicates that a particular frame is extremely likely to contain salient content. In some embodiments, input frames with predicted saliency values close to 0 (e.g., 0.1, 0.2) can contain irrelevant information and thus have little impact on action labels applied to a media file. Input frames with predicted saliency values closer to 1 (e.g., 0.8, 0.9) will have a higher impact on action labels applied to a media file.

[0039] At block 316, a feature map is generated by performing a weighted fusion operation on the feature maps generated by block 314 for the input frames. At block 318, a classifier subnetwork is used to predict probability values for one or more action types present in the media file (block 320). At block 322, an action classification label is determined for the entire media file. For example, the label can be a ground truth label for the entire media file. In some embodiments, training of the unified framework 300 can involve minimizing a cross-entropy loss between the determined label for the media file (block 322) and a“ground truth” action classification label for the media file. The generated output (e.g., action labels, classification labels) can be stored within a storage device as described herein and above for subsequent retrieval, additional post-processing, etc. In some embodiments, the output in the form of action labels can be used to train temporal saliency to identify what action is being performed in a media file.

[0040] Temporal saliency components and methods

[0041] Turning now to Figure 4 , a framework 400 is shown that includes example components for predicting temporal saliency on video data. In some embodiments, the components of framework 400 are part of the temporal saliency of block 314 of Figure 3 . Framework 400 includes, for example, a series of modules that can be used to predict the likelihood that a video frame contains a temporally salient region. Deep learning functions include, but are not limited to, a (FC) at blocks 402 and 406, a rectified linear unit (ReLU) activation function at block 404, and a Sigmoid differentiable function at block 408. Additional deep learning functions can be used by framework 400.

[0042] Temporal-spatial saliency from multi-task learning

[0043] Turning now to Figure 5 , a framework 500 is shown for training a network to determine temporal-spatial saliency for a video segment using a multi-task learning approach. Framework 500 begins at block 502 with the generation of a video segment by an electronic device (e.g., a smartphone, a camera, a video camera, etc.). The video segment is input to a network that is trained to determine temporal-spatial saliency for the video segment. The network is trained using a multi-task learning approach that includes a temporal saliency component and a spatial saliency component. Figure 1A client device 140 (in the context of the training module 104 and the integrated spatio-temporal saliency module 106A / 106B) receives an input frame, such as an image frame of a media file described herein and above. In this example, the media file includes three video frames (502(a), 502(b), 502(c)), however, it should be understood that the media file can contain more or less video frames than the example provided. The input frame can include, for example, an image frame of a media or video file described herein and above. As described above, the media file can be captured by the client device 140 or received from another device over a network, such as the network 150. Next, at block 504, the input frame can be provided as input to a neural network 504. Using the neural network, at block 506, a saliency mask is used to mask out portions of the input frame that will not be used to classify the spatial saliency of the input frame. In some embodiments, the saliency mask can be determined using supervision at block 508. For example, human eye gaze can be used to provide direct supervision to the spatial saliency. At block 510, a feature map of the input frame is generated. For example, the feature map can be generated by applying one or more filters or feature detectors to the input frame. Additionally or alternatively, the feature map can be generated based on previous video frames used as input.

[0044] At block 512, a recurrent neural network, ConvLSTM (Convolutional Long Short Term Memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 512 can be provided to block 514 and block 516. At block 514, the output from block 512 can be used to generate an average pool. At block 516, the temporal saliency of the input frame can be predicted (see Figure 4 ) and represented as a value ranging from zero (0) to one (1).

[0045] At block 518, a weighted fusion of the input frame can be predicted by multiplying the average pool value (block 514) by the temporal saliency at block 516. At block 522, a classifier can be used to predict the action recognition based on the probability (block 524), and a label can be applied (block 526) to the entire media file. For example, the label can be a true label for the entire media file. In some embodiments, training of the unified framework 500 can be used to minimize a cross-entropy loss. In some embodiments, supervision can also be applied at block 518 to predict the temporal saliency of the input frame. For example, a binary label can be applied to each input frame to indicate whether the current input frame is a key frame. The generated output (e.g., action label, classification label) can be stored within a storage device as described herein and above for subsequent retrieval, additional post-processing, etc. In some embodiments, the output in the form of an action label can be used to train the temporal saliency to recognize what action is being performed in the media file.

[0046] Next, at box 528, if the media file contains additional frames (yes), the process resumes at box 502 for the next available input frame of the media file. If the media file does not contain additional frames (no), the process ends. Once the process is complete, the results can be stored in a storage device as described herein and above for subsequent retrieval, post-processing, etc.

[0047] Based on the spatiotemporal saliency of multi-task learning using multi-stage training methods

[0048] Now go to Figure 6 A framework 600 is provided for training a network using a single-stage training strategy to determine the spatiotemporal saliency for video segments. The framework 600 may begin at box 602, where an electronic device (e.g., ...) is used. Figure 1 The client device 140 receives the input frame in the context of training module 104 and integrated spatiotemporal saliency modules 106A / 106B. In this example, the media file includes three video frames (602(a), 602(b), 602(c)); however, it should be understood that the media file may contain more or fewer video frames than the example provided. The input frame may include, for example, image frames from the media or video file described herein and above. The input frame may include, for example, image frames from the media file described herein and above. As described above, the media file may be captured by the client device 140 or received from another device via a network (such as network 150). Next, at box 604, the input frame may be provided as input to a neural network (such as a convolutional neural network (CNN)) at box 604. Using the neural network, a saliency mask (box 606) may be used to mask portions of the input frame that will not be used to classify the spatial saliency of the input frame. In some embodiments, at box 608, spatial saliency may be trained using supervision from eye gaze and action labels. In some implementations, Kullback-Leibler (KL) divergence and linear correlation coefficient can be used to measure the distance between the predicted saliency map and the eye gaze heatmap. The KL divergence and linear correlation coefficient terms can be minimized to make the distributions of the predicted saliency map and the eye gaze heatmap consistent.

[0049]

[0050] At block 610, a feature map of the input frame is generated. For example, the feature map can be generated by applying one or more filters or feature detectors to the input frame. Additionally or alternatively, the feature map can be generated based on previous video frames used as input. At block 612, a recurrent neural network, ConvLSTM (Convolutional Long Short Term Memory), can be used to predict spatio-temporal saliency based on the input frame and previous frames. The output from block 612 can be provided to block 612. At block 612, the output from block 612 can be used to generate a mean pool.

[0051] At block 616, a weighted fusion of the input frame can be predicted based on the mean pool value (block 614). At block 618, a classifier can be used to predict action recognition based on the probabilities (block 620), and labels can be applied (block 622) to the input frame. In some embodiments, training of the unified framework 300 can be used to minimize cross-entropy loss.

[0052] Next, at block 624, if the media file contains additional frames (YES), the process resumes at block 602 for the next available input frame of the media file. If the media file does not contain additional frames (NO), the process ends. Once the process ends, the results of the process can be stored within a storage device as described herein and above for subsequent retrieval, additional post-processing, etc.

[0053] Spatio-temporal saliency according to multi-task learning using a multi-stage training method

[0054] Now turning to Figure 7 , a framework 700 is provided for training a network to determine spatio-temporal saliency for a video segment using a multi-stage training strategy. In this example embodiment, the spatial saliency of the network can be fixed by freezing the network parameters used for spatial saliency. Additionally, the temporal saliency can be trained using L1 loss with supervision from frame-level annotations and action labels to regress the temporal saliency to the frame-level labels. The framework 700 illustrates an electronic device (e.g., Figure 1by the client device 140, in the context of the training module 104 and the integrated spatio-temporal saliency module 106A / 106B) receives an input frame, such as an image frame of a media file described herein and above. In this example, the media file includes three video frames (702(a), 702(b), 702(c)), however, it should be appreciated that the media file can contain more or less video frames than the provided example. The input frame can include, for example, an image frame of a media or video file described herein and above. As described above, the media file can be captured by the client device 140 or received from another device over a network, such as the network 150. Next, at block 704, the input frame can be provided as input to a neural network, such as a convolutional neural network (CNN). Using the neural network, a saliency mask at block 706 can be used to mask out portions of the input frame that will not be used to classify the spatial saliency of the input frame, as described herein. At block 710, a feature map of the input frame can be generated. For example, the feature map can be generated by applying one or more filters or feature detectors to the input frame. Additionally or alternatively, the feature map can be generated based on previous video frames used as input. In this example, as indicated by block 712, the network parameters used to predict spatial saliency can be frozen based on a first training phase, such as described above with reference to FIG. 6. As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels. Figure 6 As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels. Figure 7 As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels. Figure 4 As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels. As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels.

[0055] As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels. As shown by block 714, a second training phase can be performed based on temporal saliency, where the network is trained with supervision from frame-level annotations and action labels. At block 714, a recurrent neural network, such as ConvLSTM (convolutional long short-term memory), can be used to predict the spatio-temporal saliency based on the input frame and previous frames. The output from block 714 can be provided to block 716 and block 718. At block 716, the output from block 714 can be used to generate a mean pool. At block 718, the temporal saliency of the input frame can be predicted using a subnetwork (see FIG. 7B) and represented as a value ranging from zero (0) to one (1). The temporal saliency of the input frame can be trained with supervision from frame-level annotations and action labels. In some embodiments, an LI loss can be used to regress the temporal saliency to the frame-level labels.

[0056] Next, at block 728, if the media file contains additional frames (YES), the process resumes at block 702 for the next available input frame of the media file. If the media file does not contain additional frames (NO), the process ends. Once the process ends, the results of the process can be stored within a storage device as described herein and above for subsequent retrieval, additional post-processing, etc.

[0057] Detailed network architecture during inference

[0058] Turning now to Figure 8 , an example network architecture 800 for a neural network architecture for spatio-temporal saliency at an inference device is shown in accordance with one or more embodiments. As Figure 8 indicated, an input media file including one or more image frames at block 802 can be provided to a neural network at block 804. The neural network can include a unified framework for providing spatio-temporal saliency by combining an output at block 820 (spatial saliency output) with an output of the neural network at block 822 to a spatio-temporal saliency of the media file as described herein and above. The example network architecture 800 can include multiple layers for down-sampling and up-sampling at blocks 806(a)-806(d), layers for Post NN, US1, and US2 at blocks 810, 812, and 814, respectively, and a skip function at blocks 808(a) and 808(b), for example. The network architecture 800 can also include a Conv2D class at block 816 and a Sigmoid function at block 818 as part of the neural network. At block 824, the network architecture 800 includes a recurrent neural network ConvLSTM for providing a spatio-temporal prediction output. As described herein, additional functions for predicting salient regions of a media file can be provided including, but not limited to, a Post NN at block 826, an average pool at block 828, FCs at blocks 830 and 834, a ReLU at block 832, a Sigmoid at block 836, and an output at block 838.

[0059] Figure 9 An example method for predicting spatio-temporal saliency of a media item is shown in flowchart form. The method can be implemented by the integrated spatio-temporal saliency module 106A, the integrated spatio-temporal saliency module 106B, or a combination or variation thereof. The method can be implemented on a server device such as the spatio-temporal service 100, or on a client device such as the client device 140. Figure 1 For explanatory purposes, the following steps will be described in the context of the server device 100. However, various actions can be performed by alternative components. Additionally, various actions can be performed in different orders. Furthermore, some actions can be performed concurrently, can not be required, and other actions can be added. Figure 1 ​

[0060] The flowchart begins at block 905 with capturing image content by a device, such as a device equipped with one or more cameras, image sensors, etc. For example, the image content can include video data composed of a plurality of video frames. Additionally or alternatively, the image content can be retrieved from a storage device, such as a memory of the device, or a remote storage device, such as a memory of another device or a cloud computing device connected through a network. Next, at block 910, video frames of the image content having a spatiotemporal salient region can be identified. In some embodiments, a saliency score can be assigned to the image content to indicate a likelihood that the image content includes a salient region. Additionally, the image content can be analyzed in real-time or near real-time (e.g., during capture or shortly after capture).

[0061] The flowchart continues at block 915 with predicting an event occurring within the salient region of one or more frames of the video segment. As described herein, a neural network can be used to predict the event. At block 920, one or more action labels can be assigned to the video segment. The one or more action labels can be based at least in part on the event predicted by the neural network.

[0062] For playback of the image content, at block 925, video frames of the image content that show meaningful action changes can be encoded at a variable frame rate. For example, video frames that include a salient region can be played back at a lower compression rate and / or at an increased frame rate relative to frames of the image content that include non-salient regions.

[0063] According to some embodiments, a processor or processing element can be trained using supervised machine learning and / or unsupervised machine learning, and the machine learning can employ an artificial neural network, which can be, for example, a convolutional neural network, a recurrent neural network, a deep learning neural network, a reinforcement learning module or program, or a combined learning module or program that learns in two or more fields or areas of interest. Machine learning can involve identifying and recognizing patterns in existing data in order to facilitate predictions on subsequent data. Models can be created based on example inputs in order to make effective and reliable predictions on novel inputs.

[0064] According to certain embodiments, a machine learning program can be trained by inputting a sample data set or certain data, such as images, object statistics and information, historical estimates, and / or image / video / audio classification data, into the program. The machine learning program can utilize a deep learning algorithm, which can primarily focus on pattern recognition and can be trained after processing a plurality of examples. The machine learning program can include Bayesian program learning (BPL), speech recognition and synthesis, image or object recognition, optical character recognition, and / or natural language processing. The machine learning program can also include natural language processing, semantic analysis, automated reasoning, and / or other types of machine learning.

[0065] According to some embodiments, supervised machine learning techniques and / or unsupervised machine learning techniques can be used. In supervised machine learning, a processing element can be provided with example inputs and their associated outputs, and can seek to discover general rules that map inputs to outputs, such that when a novel input is provided later, the processing element can accurately predict the correct output based on the discovered rules. In unsupervised machine learning, a processing element can need to discover its own structure in unlabeled example inputs.

[0066] The scope of the subject disclosure should be determined with reference to the appended claims, along with the full range of equivalents to which such claims are entitled. In the appended claims, the terms "including," "includes," "having," "has," "with," and "containing," "contains" are used as equivalent terms to the respective terms "comprising," "comprises," "comprised of," and "comprising of."

Claims

1. A system for processing video, comprising: One or more processors; and One or more computer-readable media, the one or more computer-readable media including computer-readable code, the computer-readable code being executable by the one or more processors to: Receives a video segment containing multiple frames of video data; A neural network is used to identify spatiotemporally salient regions in one or more frames of the video segment; Predict at least one event occurring within the identified spatiotemporal saliency region of the one or more frames of the video segment; and One or more action tags are applied to the video segment based at least in part on at least one predicted event.

2. The system of claim 1, wherein the computer-readable code for predicting the at least one event occurring within an identified spatiotemporally salient region in the one or more frames of the video segment comprises computer-readable code for: Calculate the spatiotemporal saliency score of at least one spatiotemporal saliency region among the identified spatiotemporal saliency regions in the one or more frames of the video segment; and The at least one event is predicted based at least in part on the calculated spatiotemporal saliency score.

3. The system of claim 2, wherein the spatiotemporal saliency score predicts the degree to which one or more persons will be interested in corresponding regions of one or more frames of the video segment.

4. The system of claim 1, wherein the one or more computer-readable media further comprises computer-readable code for the following operations: During playback of the video segment, one or more frames with the identified spatiotemporally significant regions are provided at a lower compression rate.

5. The system of claim 1, wherein the one or more computer-readable media further comprises computer-readable code for the following operations: During playback of the video segment, one or more frames with the identified spatiotemporally significant regions are provided at an increased frame rate.

6. The system of claim 1, wherein the one or more computer-readable media further comprises computer-readable code for the following operations: During playback of the video segment, one or more frames with the identified spatiotemporally significant regions are provided at a lower compression rate and an increased frame rate.

7. The system of claim 1, wherein the one or more computer-readable media further comprises computer-readable code for the following operations: The video segments are classified, at least in part, based on one or more action tags applied.

8. The system of claim 1, wherein the one or more computer-readable media further comprises computer-readable code for the following operations: The video segment is stored, wherein the one or more frames having the identified spatiotemporally significant regions are stored at a first compression rate, and the one or more frames having the identified spatiotemporally significant regions are stored at a second compression rate.

9. The system of claim 8, wherein the first compression ratio is lower than the second compression ratio.

10. A system for processing video, comprising: One or more processors; and One or more computer-readable media, the one or more computer-readable media including computer-readable code, the computer-readable code being executable by the one or more processors to: Receive a video segment containing multiple frames of video data; and The neural network is trained to identify spatiotemporally salient regions in one or more frames of the video segment. The neural network is trained using multi-stage training.

11. The system of claim 10, wherein the multi-stage training includes at least one stage for training the neural network to identify spatially salient regions in the one or more frames of the video segment.

12. The system of claim 11, wherein the spatially salient regions in the one or more frames of the video segment are identified using supervised eye gaze, one or more motion tags, or both.

13. The system of claim 10, wherein the multi-stage training includes at least one stage for training the neural network to identify temporally salient regions in the one or more frames of the video segment.

14. The system of claim 13, wherein the temporally significant regions in one or more frames of the video segment are identified using supervision from frame annotations and action tags.

15. The system of claim 10, wherein the multi-stage training comprises a first stage and a second stage, the first stage being used to train the neural network to identify spatially salient regions in the one or more frames of the video segment, and the second stage being used to train the neural network to identify temporally salient regions in the one or more frames of the video segment.

16. A method for processing video for playback, comprising: Capture video containing multiple frames of image data; A neural network is used to identify spatiotemporally salient regions in one or more of the plurality of frames of the video; Predict at least one event occurring within an identified spatiotemporally significant region of one or more frames in the plurality of frames of the video; as well as One or more action tags are applied to the video based on at least one predicted event.

17. The method of claim 16, wherein predicting the at least one event occurring within an identified spatiotemporally saliency region in one or more of the plurality of frames of the video comprises: Calculate the spatiotemporal saliency score of at least one spatiotemporal saliency region among the identified spatiotemporal saliency regions in one or more of the plurality of frames of the video; and The at least one event is predicted based at least in part on the calculated spatiotemporal saliency score.

18. The method of claim 17, wherein the spatiotemporal saliency score predicts the degree to which one or more persons would be interested in corresponding regions of one or more frames of the plurality of frames of the video.

19. The method of claim 16, wherein the method further comprises: During playback of the video, one or more frames of the plurality of frames having the identified spatiotemporally significant regions are provided at a lower compression rate.

20. The method of claim 16, wherein the method further comprises: During playback of the video, one or more frames of the plurality of frames having the identified spatiotemporally significant regions are provided at an increased frame rate.