Method and electronic device for determining motion saliency and video playback style in a video

By detecting the object motion type and speed in the video frame through electronic devices and automatically adjusting the video playback style using machine learning models, the problem of classifying motion saliency and playback style in the video is solved, and the intelligence and beauty of video synthesis are improved.

CN116349233BActive Publication Date: 2025-09-26SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280007042.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-20
Filing Date
2022-01-18
Publication Date
2025-09-26
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

The existing technology lacks automated methods for classifying motion saliency and video playback style in videos, which results in the need for manual intervention in video editing and the inability to achieve intelligent video analysis and synthesis.

Method used

Detecting objects in video frames through an electronic device, determining their motion type and speed, and applying effects based on a machine learning model to automatically adjust the video playback style, including training a motion-based model to predict and apply effects.

Benefits of technology

It realizes the automatic processing of motion saliency and playback style in videos, improves the quality and aesthetics of video synthesis style, and supports real-time video analysis and generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116349233B_ABST
    Figure CN116349233B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method for applying an effect in a video by an electronic device. The method includes: detecting a first object and a second object in an image frame of the video; determining a motion type of the first object and the second object in the video; determining a motion speed of the first object and the second object in the video; determining a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion type and the motion speed of the first object and the second object; and applying the first effect to the first object and the second effect to the second object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a content analysis system, and for example, to a method and electronic device for determining motion saliency in a video and video playback style. Background Art

[0002] The electronic device includes an intelligent camera system that automatically processes user requests using video data. Video data may include frame visual details, cross-frame temporal information (e.g., motion), and audio stream information. In addition, motion information (cross-frame temporal information) is unique to a video and carries rich information about the action in the video. In addition, learning important motion information in a video will help the network learn strong spatiotemporal features that can be used in a large number of video analysis tasks, and understanding its patterns can help to better synthesize video styles, which can increase the overall video aesthetics. However, in traditional methods, there is no method for classifying different playback styles in a video. In addition, there are no competitors or third-party solutions for automatic intelligent video analysis. In addition, all available video editing methods require people to manually select the video playback type and speed to edit and generate interesting clips from the video. Summary of the Invention

[0003] Technical Solution

[0004] According to various exemplary embodiments of the present disclosure, a method for determining motion saliency and video playback style in a video by an electronic device is provided. The method includes: detecting, by the electronic device, a first object and a second object in at least one image frame of the video; determining, by the electronic device, a motion type of the first object and the second object in the video; determining, by the electronic device, a motion speed of the first object and the second object in the video; determining, by the electronic device, a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion type and the motion speed of the first object and the second object; and applying, by the electronic device, the first effect to the first object and the second effect to the second object.

[0005] Beneficial effects

[0006] The present disclosure provides a better video synthesis style and increases the overall video aesthetics by determining motion saliency and video playback style and applying individual effects determined based on motion type and object speed to each object in an image scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The embodiments are shown in the accompanying drawings, and the same reference numerals indicate corresponding parts throughout the drawings. In addition, the above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent through the following detailed description taken in conjunction with the accompanying drawings, in which:

[0008] Figure 1 is a block diagram illustrating an example configuration of an electronic device for determining motion saliency in a video and video playback style according to various embodiments;

[0009] Figure 2 is a block diagram illustrating an example configuration of a video motion controller included in an electronic device according to various embodiments;

[0010] Figure 3 is a block diagram illustrating various hardware components of a video playback controller included in an electronic device according to various embodiments;

[0011] Figure 4 is a flow chart illustrating an example method for determining motion saliency in a video and video playback style according to various embodiments;

[0012] Figure 5 is a diagram illustrating an example in which an electronic device determines motion saliency and video playback style in a video according to various embodiments;

[0013] Figure 6 is a diagram illustrating an example operation of an illustrative video feature extractor according to various embodiments;

[0014] Figure 7 is a diagram illustrating an example operation of a playback speed detector according to various embodiments;

[0015] Figure 8 is a diagram illustrating an example operation of a playback type detector according to various embodiments;

[0016] Figure 9 is a diagram illustrating example operations of a motion saliency detector according to various embodiments; and

[0017] Figure 10 is a diagram illustrating example integration in a single-shot photo pipeline in accordance with various embodiments. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure provide a method and an electronic device for determining motion saliency and video playback style in a video.

[0019] Embodiments of the present disclosure determine one or more motion patterns associated with a video to assist in generating an automatic video synthesis style.

[0020] Embodiments of the present disclosure predict saliency detection, video playback style, video playback type, and video playback speed associated with a video in real-time.

[0021] Embodiments of the present disclosure analyze and understand temporal motion patterns (e.g., linear, projectile) and spatial motion areas (e.g., waving a flag, etc.), and classify temporal motion patterns (e.g., linear, projectile) and spatial motion areas (e.g., waving a flag, etc.) into meaningful types for multiple applications (such as artificial intelligence (AI) playback style prediction, AI still-frame animation, and dynamic adaptive video generation).

[0022] Embodiments of the present disclosure determine the motion direction (e.g., linear, cyclic, trajectory, random, etc.) of subjects / objects in a video scene and use semantic information to select an appropriate playback style for the video.

[0023] Embodiments of the present disclosure determine the motion direction (eg, fast, slow, normal) of subjects / objects in a scene and use the semantic information to select an appropriate playback style for the video.

[0024] Embodiments of the present disclosure locate spatial regions of a video with stronger and patterned motion present in the video.

[0025] According to various exemplary embodiments of the present disclosure, a method for determining motion saliency and video playback style in a video by an electronic device is provided. The method includes: detecting, by the electronic device, a first object and a second object in at least one image frame of the video; determining, by the electronic device, a motion type of the first object and the second object in the video; determining, by the electronic device, a motion speed of the first object and the second object in the video; determining, by the electronic device, a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion type and the motion speed of the first object and the second object; and applying, by the electronic device, the first effect to the first object and the second effect to the second object.

[0026] According to an example embodiment, the steps of determining, by an electronic device, a first effect to be applied to a first object and a second effect to be applied to a second object include: classifying, by the electronic device, the motions of the first object and the second object into at least one specified category by applying at least one motion-based machine learning model to the motion types of the first object and the second object and the motion speeds of the first object and the second object; and predicting, by the electronic device, the first effect to be applied to the first object and the second effect to be applied to the second object based on at least one motion category of the first object and the second object.

[0027] According to an example embodiment, the method includes: training, by an electronic device, at least one motion-based training model to predict a first effect and a second effect, wherein the step of training the at least one motion-based training model includes: determining a plurality of motion cues in a video; determining a plurality of motion factors in the video; and training the at least one motion-based training model based on the plurality of motion cues and the plurality of motion factors to predict a first effect to be applied to a first object and a second effect to be applied to a second object.

[0028] According to an example embodiment, the motion cues include temporal motion features, spatial motion features, and spatiotemporal knowledge of the video.

[0029] According to an example embodiment, the motion factor includes a motion direction of each object in the video, a motion pattern of each object in the video, an energy level of motion in the video, and a saliency map of motion in the video.

[0030] According to an example embodiment, the step of applying, by an electronic device, a first effect to a first object and a second effect to a second object includes: determining, by the electronic device, a size of the first object and a duration for which the first object is in motion in the at least one image frame of the video; determining, by the electronic device, a size of the second object and a duration for which the second object is in motion in the at least one image frame of the video; determining, by the electronic device, a time period for which both the first object and the second object are in motion simultaneously in the at least one image frame of the video based on the size of the first object, the duration of the first object, the size of the second object, and the duration of the second object; and applying, by the electronic device, the first effect to the first object and the second effect to the second object during the determined time period.

[0031] According to an example embodiment, the step of determining, by an electronic device, the motion type of a first object and a second object in a video includes: determining, by the electronic device, the motion type of the first object at a first time interval; and determining, by the electronic device, the motion type of the second object at a second time interval, wherein the first time interval is different from the second time interval.

[0032] According to an example embodiment, the step of determining, by an electronic device, the movement speeds of a first object and a second object in a video includes: determining, by the electronic device, the movement speed of the first object in a first time interval; and determining, by the electronic device, the movement speed of the second object in a second time interval, wherein the first time interval is different from the second time interval.

[0033] According to an example embodiment, the first effect includes at least one of a first stylistic effect and a first frame rate speed effect; and wherein the second effect includes at least one of a second stylistic effect and a second frame rate speed effect.

[0034] Accordingly, various example embodiments of the present disclosure disclose an electronic device configured to determine motion saliency and video playback style in a video. The electronic device includes: a memory that stores a video; a processor; and a video motion controller that is communicatively coupled to the memory and the processor; and a video playback controller that is communicatively coupled to the memory and the processor. The video motion controller is configured to: detect a first object and a second object in at least one image frame of a video and determine a motion type of the first object and the second object in the video; determine a motion speed of the first object and the second object in the video; determine a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion type of the first object and the second object and the motion speed of the first object and the second object; and apply the first effect to the first object and the second effect to the second object.

[0035] These and other aspects of the various example embodiments of the present disclosure will be better appreciated and understood when considered in conjunction with the following description and accompanying drawings. However, it should be understood that the following description, while indicating example embodiments and many specific details thereof, is given by way of illustration and not limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope of the embodiments herein, and the embodiments herein include all such modifications.

[0036] The various example embodiments of the present invention and their various features and advantageous details are explained more fully with reference to the non-limiting embodiments shown in the accompanying drawings and described in detail in the following description. Descriptions of well-known components and processing technologies may be omitted to avoid unnecessary details that unnecessarily obscure the present disclosure. The various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. Unless otherwise indicated, the term "or" as used herein refers to a non-exclusive or. The examples used herein are intended only to facilitate understanding of the manner in which the embodiments of the present invention may be practiced, and further enable those skilled in the art to practice the embodiments of the present invention. Therefore, the examples should not be interpreted as limiting the scope of the embodiments of the present invention.

[0037] As traditional in the art, can be described and illustrated according to the block of one or more functions described for execution.These blocks (may be referred to as unit or module etc. in this article) are physically realized by analog or digital circuit (such as logic gate, integrated circuit, microprocessor, microcontroller, memory circuit, passive electronic component, active electronic component, optical component, hard-wired circuit etc.), and can be optionally driven by firmware and / or software.Circuit can be, for example, realized in one or more semiconductor chips, or be realized on substrate support member (such as printed circuit board etc.).The circuit of block can be realized by dedicated hardware, or realized by processor (such as, one or more programmed microprocessors and associated circuit), or realized by the combination of dedicated hardware of some functions of execution block and processor of other functions of execution block.Without departing from the scope of this disclosure, each block of embodiment can be physically divided into two or more interactive and discrete blocks.Similarly, without departing from the scope of this disclosure, the block of embodiment can be physically combined into more complicated block.

[0038] The accompanying drawings are used to help understand various technical features, and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. Therefore, the present disclosure should be interpreted as extending to any changes, equivalents and substitutes other than those specifically set forth in the accompanying drawings. Although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are usually only used to distinguish one element from another element.

[0039] Therefore, embodiments herein implement a method for determining motion saliency and video playback style in a video by an electronic device. The method includes detecting, by the electronic device, a first object and a second object in at least one image frame of a video. Furthermore, the method includes determining, by the electronic device, a motion type of the first object and the second object in the video. Furthermore, the method includes determining, by the electronic device, a motion speed of the first object and the second object in the video. Furthermore, the method includes determining, by the electronic device, a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion type of the first object and the second object and the motion speed of the first object and the second object. Furthermore, the method includes applying, by the electronic device, the first effect to the first object and the second effect to the second object.

[0040] Unlike conventional methods and systems, the disclosed method can be used to determine one or more motion patterns associated with a video to help generate an automatic video synthesis style. The method can be used to predict saliency detection, video playback style, video playback type, and video playback speed associated with a video in real time. In the proposed method, the video feature extractor can be a lightweight convolutional neural network (CNN) based on MobileNetV2 and extracts important spatiotemporal features from video clips to analyze multiple real-time videos in an electronic device in an efficient manner.

[0041] The method can be used to analyze and understand temporal motion patterns (e.g., linear, projectile) and spatial motion regions (e.g., waving a sign, etc.), and classify temporal motion patterns (e.g., linear, projectile) and spatial motion regions (e.g., waving a sign, etc.) into meaningful types for multiple applications (such as artificial intelligence (AI) playback style prediction, AI still frame animation, and dynamic adaptive video generation). The method can be used to determine the motion direction of the subject / object in the video scene (e.g., linear, loop, trajectory, random, etc.) and use semantic information to select the appropriate playback style for the video. The method can be used to determine the motion direction of the subject / object in the scene (e.g., fast, slow, normal) and use semantic information to select the appropriate playback style for the video. The method can be used to locate spatial regions of the video with stronger and patterned motion present in the video.

[0042] A user of an electronic device may point an imaging device (e.g., a camera, etc.) at a scene to capture one or more events, and an automatic short video / still frame animation will be created based on the frames captured by the camera in a single-shot camera mode and analysis of the captured frames to classify motion in real time.

[0043] This method can be used to play a video in an appropriate style (e.g., loop, slow motion, time-lapse, GIF loop, reverse video) based on the analysis results. This method can be used to generate still images, selfies, stories, or dynamic videos based on the detected saliency information. Based on this example method, the electronic device can retrieve similar videos from a gallery based on the motion patterns occurring in the video.

[0044] The network is designed to be lightweight to accommodate electronic device deployment and is capable of extracting better video feature descriptors that learn unique motion patterns. The electronic device is trained to extract important features about the motion present in the video, including the main motion direction, motion pattern, motion energy, and motion saliency. The method focuses on extracting only useful simple features from the video that are helpful for applications (e.g., camera applications, gallery applications, etc.). Better video feature descriptors that learn unique motion patterns are extracted, and the method requires less data to train. The proposed method can be used to generate time-varying adaptive video acceleration, which allows viewers to watch videos faster, but with less jittery, unnatural motion that is typical of uniformly accelerated videos.

[0045] In various example methods, an electronic device (e.g., a compact and lightweight mobile video photography system, etc.) can analyze and understand temporal motion patterns (e.g., linear, projectile, etc.) and spatial motion areas (e.g., waving a sign), and classify the temporal motion patterns (e.g., linear, projectile, etc.) and spatial motion areas (e.g., waving a sign) into meaningful types for multiple applications (such as AI playback style prediction, AI still-frame animation, dynamic adaptive video generation, etc.). In an example, for a single-shot photo application, AI understands the video to create meaningful dynamic videos, highlight videos, etc.

[0046] Referring now to the drawings, and more particularly to the Figures 1 to 10 , various example embodiments of the present disclosure are shown and described.

[0047] Figure 1 The present invention is a block diagram illustrating an example configuration of an electronic device (100) for determining motion saliency and video playback style in a video according to various embodiments. The electronic device (100) may include, for example, but not limited to, a foldable device, a cellular phone, a smartphone, a personal digital assistant (PDA), a tablet computer, a laptop computer, a smartwatch, an immersive device, a virtual reality device, a compact and lightweight mobile video system, a camera, and the Internet of Things (IoT).

[0048] In an embodiment, the electronic device (100) includes a processor (e.g., including a processing circuit) (110), a communicator (e.g., including a communication circuit) (120), a memory (130), a video motion controller (e.g., including a control circuit) (140), a video playback controller (e.g., including a control circuit) (150), and a machine learning model controller (e.g., including a control circuit) (160). The processor (110) is coupled to the communicator (120), the memory (130), the video motion controller (140), the video playback controller (150), and the machine learning model controller (160). The video playback controller (150) and / or the machine learning model controller (160) can be implemented as a hardware processor by combining the processor (110), the video playback controller (150), and the machine learning model controller 160.

[0049] The video motion controller (140) may include various control circuits and is configured to detect a first object and a second object in an image frame of a video. The video motion controller (140) is configured to determine a motion type of the first object at a first time interval and to determine a motion type of the second object at a second time interval. The first time interval is different from the second time interval.

[0050] The video motion controller (140) is configured to determine a motion speed of a first object at a first time interval and to determine a motion speed of a second object at a second time interval, wherein the first time interval is different from the second time interval.

[0051] The video playback controller (150) may include various control circuits and is configured to classify the motion of the first object and the second object into predefined (e.g., specified) categories by applying a motion-based machine learning model to the motion type of the first object and the second object and the motion speed of the first object and the second object using a machine learning model controller (160). The video playback controller (150) is configured to predict a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion category of the first object and the second object. The first effect may include, for example, but not limited to, a first style effect and a first frame rate speed effect. The second effect may include, for example, but not limited to, a second style effect and a second frame rate speed effect.

[0052] The video playback controller (150) is configured to train a motion-based training model to predict a first effect and a second effect. The motion-based training model is trained by determining a plurality of motion cues in a video, determining a plurality of motion factors in the video, and training the motion-based training model based on the plurality of motion cues and the plurality of motion factors to predict a first effect to be applied to a first object and a second effect to be applied to a second object.

[0053] The plurality of motion cues may include, for example, but not limited to, temporal motion features, spatial motion features, and spatiotemporal knowledge of the video. The plurality of motion factors may include, for example, but not limited to, the direction of motion of each object in the video, the motion pattern of each object in the video (e.g., spatial local, spatial global, repetitive), the energy level of motion in the video, and a saliency map of motion in the video. In the motion pattern of each object in the video, spatial local may indicate motion that occurs locally within a predetermined boundary in the scene of the video, spatial global may indicate motion that occurs globally outside a predetermined boundary in the scene of the video, and motion repetitive may indicate the same or substantially the same motion that occurs within a predetermined time period. The motion cues and the plurality of motion factors recognized by the electronic device (100) may be used to identify motion segments in the video. For portions of the video with a high amount of motion, there will be stronger motion factors and cues.

[0054] The electronic device (100) can learn the motion representation of a video, and thus can compare the extracted features of one video with the extracted features of other videos to find similar videos from various applications (e.g., gallery applications, etc.). In addition, the electronic device (100) can also use the weights learned from the motion feature extraction task to initialize a model for the action recognition task, and then fine-tune the model for the action recognition task. This eliminates and / or reduces the need to pre-train the action recognition task with a large action recognition dataset.

[0055] The video playback controller (150) is configured to determine the size of a first object and the duration of time the first object is in motion in an image frame of the video. The video playback controller (150) is configured to determine the size of a second object and the duration of time the second object is in motion in an image frame of the video. The video playback controller (150) is configured to determine the time period during which both the first object and the second object are in motion simultaneously in the image frame of the video based on the size of the first object, the duration of the first object, the size of the second object, and the duration of the second object.

[0056] The video playback controller (150) is configured to apply a first effect to a first object and a second effect during a determined time period. In an example, when two objects are in motion simultaneously in a video, the electronic device (100) may apply the effect based on the size of the objects and the presence of the objects over the duration of the video. In an example, a person is riding a bicycle and a thrown soccer ball enters the preview scene. Based on this method, the method can be used to apply the effect based on the detected bicycle event.

[0057] The processor (110) may include various processing circuits and is configured to execute instructions stored in the memory (130) and perform various processes. The communicator (120) may include various communication circuits and is configured to communicate internally between internal hardware components and to communicate with external devices via one or more networks. The memory (130) also stores instructions to be executed by the processor (110). The memory (130) may include a non-volatile storage element. Examples of such non-volatile storage elements may include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or an electrically programmable memory (EPROM) or an electrically erasable programmable memory (EEPROM) memory. Additionally, in some examples, the memory (130) may be considered a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be interpreted as meaning that the memory (130) is non-removable. In a specific instance, a non-transitory storage medium may store data that may change over time (e.g., in a random access memory (RAM) or a cache memory).

[0058] In addition, at least one of the plurality of modules / controllers may be implemented by an AI model. Functions associated with the AI ​​model may be executed by a non-volatile memory, a volatile memory, and a processor (110). The processor (110) may include one or more processors. In this case, the one or more processors may include, for example, but not limited to, a general-purpose processor (such as a central processing unit (CPU), an application processor (AP), etc.), a graphics processing unit (such as a graphics processing unit (GPU)), a visual processing unit (VPU), and / or an AI-specific processor (such as a neural processing unit (NPU)).

[0059] One or more processors control the processing of input data according to predefined operating rules or AI models stored in non-volatile memory and volatile memory. Predefined operating rules or AI models are provided through training or learning.

[0060] Providing an AI model through learning may refer to, for example, predefined operating rules or desired characteristics formed by applying a learning algorithm to a plurality of learning data. Learning may be performed in the device itself that executes the AI ​​according to the embodiment, and / or may be implemented by a separate server / system.

[0061] The AI ​​model may include multiple neural network layers. Each layer may have multiple weight values, and layer operations are performed by calculating the previous layer and operating on the multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0062] A learning algorithm may refer to a method for training a predetermined target device (e.g., a robot) using a plurality of learning data to enable, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0063] although Figure 1 Various hardware components of the electronic device (100) are shown, but it should be understood that various embodiments are not limited thereto. In other embodiments, the electronic device (100) may include fewer or greater numbers of components. In addition, the labels or names of the components are for illustrative purposes only and do not limit the scope of the present disclosure. One or more components may be combined to perform the same or substantially similar functions in the electronic device (100).

[0064] Figure 2 1 is a block diagram illustrating an example configuration of a video motion controller (140) included in an electronic device (100) according to various embodiments. In an embodiment, the video motion controller (140) includes a motion event detector (e.g., including various processing circuits and / or executable program instructions) (140a), a video feature extractor (e.g., including various processing circuits and / or executable program instructions) (140b), a spatiotemporal feature detector (e.g., including various processing circuits and / or executable program instructions) (140c), a playback speed detector (e.g., including various processing circuits and / or executable program instructions) (140d), and a playback type detector (e.g., including various processing circuits and / or executable program instructions) (140e).

[0065] The motion event detector (140a) receives an input video stream and detects motion in the input video stream. The video feature extractor (140b) can be a lightweight CNN based on MobileNetV2 and extracts important spatiotemporal features from the motion detected input video stream. The playback speed detector (140d) predicts the video playback speed of the motion detected input video stream, and the playback type detector (140e) predicts the video playback type of the motion detected input video stream. The spatiotemporal feature detector (140c) maps to predict the video playback speed and playback type on the motion detected input video stream.

[0066] although Figure 2 Various hardware components of the video motion controller (140) are shown, but it should be understood that other embodiments are not limited thereto. In other embodiments, the video motion controller (140) may include fewer or greater numbers of components. Furthermore, the labels or names of the components are for illustrative purposes only and do not limit the scope of the present disclosure. One or more components may be combined to perform the same or substantially similar functions in the video motion controller (140).

[0067] Figure 3 is a block diagram illustrating various components of a video playback controller (150) included in an electronic device (100) according to various embodiments. In an embodiment, the video playback controller (150) includes a motion saliency detector (e.g., including various processing circuits and / or executable program instructions) (150a) and a playback type application detector (e.g., including various processing circuits and / or executable program instructions) (150b). The motion saliency detector (150a) locates spatial regions of a video that have stronger and more patterned motion present in the video (assuming the background is nearly static). The motion saliency detector (150a) locates foreground motion regions or background motion regions in the video. The playback type application detector (150b) detects playback type applications.

[0068] although Figure 3 Various hardware components of the video playback controller (150) are shown, but it should be understood that other embodiments are not limited thereto. In other embodiments, the video playback controller (150) may include fewer or greater numbers of components. Furthermore, the labels or names of the components are for illustrative purposes only and do not limit the scope of the present disclosure. One or more components may be combined to perform the same or substantially similar functions in the video playback controller (150).

[0069] Figure 4 is a flow chart (400) illustrating an example method for determining motion saliency in a video and video playback style, according to various embodiments.

[0070] Operations (S402-S406) may be performed by a video motion controller (140). At S402, the method includes detecting a first object and a second object in an image frame of a video. At S404, the method includes determining a type of motion of the first object and the second object in the video. At S406, the method includes determining a speed of motion of the first object and the second object in the video.

[0071] Operations (S408-S418) may be performed by the video playback controller (150). At S408, the method includes classifying the motions of the first object and the second object into predefined (e.g., specified) categories by applying a motion-based machine learning model to the motion types of the first object and the second object and the motion speeds of the first object and the second object.

[0072] At S410, the method includes predicting a first effect to be applied to the first object and a second effect to be applied to the second object based on the motion categories of the first object and the second object. At S412, the method includes determining a size of the first object and a duration of time the first object is in motion in an image frame of the video. At S414, the method includes determining a size of the second object and a duration of time the second object is in motion in the image frame of the video.

[0073] At S416, the method includes determining a time period during which both the first object and the second object are in motion simultaneously in image frames of the video based on a size of the first object, a duration of the first object, a size of the second object, and a duration of the second object. At S418, the method includes applying a first effect to the first object and a second effect to the second object during the determined time period.

[0074] Unlike conventional methods and systems, the disclosed method can be used to determine one or more motion patterns associated with a video to help generate an automatic video synthesis style. The method can be used to predict saliency detection, video playback style, video playback type, and video playback speed associated with a video in real time. In the proposed method, the video feature extractor can be a lightweight CNN based on MobileNetV2 and extracts important spatiotemporal features from video clips to effectively analyze multiple real-time videos in electronic devices.

[0075] The method can be used to analyze and understand temporal motion patterns (e.g., linear, projectile) and spatial motion regions (e.g., waving a sign, etc.), and classify temporal motion patterns (e.g., linear, projectile) and spatial motion regions (e.g., waving a sign, etc.) into meaningful types for multiple applications (such as artificial intelligence (AI) playback style prediction, AI still frame animation, and dynamic adaptive video generation). The method can be used to determine the motion direction of the subject / object in the video scene (e.g., linear, loop, trajectory, random, etc.) and use semantic information to select the appropriate playback style for the video. The method can be used to determine the motion direction of the subject / object in the scene (e.g., fast, slow, normal) and use semantic information to select the appropriate playback style for the video. The method can be used to locate spatial regions of the video with stronger and patterned motion present in the video.

[0076] A user of the electronic device (100) may point an imaging device (e.g., a camera, etc.) at a scene to capture one or more events, and an automatic short video / still frame animation will be created based on the frames captured by the camera in a single-shot camera mode and analysis of the captured frames for classifying motion in real time.

[0077] This method can be used to play videos in an appropriate style (e.g., loop, slow motion, time-lapse, GIF loop, reverse video) based on the analysis results. This method can also be used to generate stills, selfies, stories, and dynamic videos from the detected saliency information. Based on the proposed method, an electronic device can retrieve similar videos from a gallery based on the motion patterns occurring in the video.

[0078] The network is designed to be lightweight to accommodate electronic device deployment and is capable of extracting better video feature descriptors that learn unique motion patterns. The electronic device is trained to extract important features about the motion present in the video, including the main motion direction, motion pattern, motion energy, and motion saliency. The approach focuses on extracting only useful, simple features from the video that are useful for applications (e.g., camera applications, gallery applications, etc.). This approach extracts better video feature descriptors that learn unique motion patterns and requires less data to train.

[0079] In the disclosed method, the electronic device (100) can analyze and understand temporal motion patterns (e.g., linear, projectile, etc.) and spatial motion areas (e.g., waving signs), and classify the temporal motion patterns (e.g., linear, projectile, etc.) and spatial motion areas (e.g., waving signs) into meaningful types for multiple applications (such as AI playback style prediction, AI still frame animation, dynamic adaptive video generation, etc.). In an example, for a single-shot photo application, AI understands the video to create meaningful dynamic videos, highlight videos, etc. The proposed method can be used to generate time-varying adaptive video acceleration, which can allow viewers to watch videos faster, but with less jitter, and unnatural motion typical of uniformly accelerated videos.

[0080] The various actions, motions, blocks, steps, etc. in the flowchart (400) may be performed in the order presented, in a different order, or simultaneously. In various embodiments, some of the actions, motions, blocks, steps, etc. may be omitted, added, modified, skipped, etc. without departing from the scope of the present disclosure.

[0081] Figure 5 is a diagram illustrating an example ( S500 ) of an electronic device ( 100 ) determining motion saliency and playback style in a video according to various embodiments.

[0082] In the example, a motion event detector (140a) receives an input video stream and detects motion in the input video stream. A video feature extractor (140b) may be a lightweight CNN based on MobileNetV2 and extracts important spatiotemporal features from the motion detected input video stream. A playback speed detector (140d) predicts the video playback speed of the motion detected input video stream and a playback type detector (140e) predicts the video playback type of the motion detected input video stream. The spatiotemporal feature detector (140c) maps to predict the video playback speed and playback type on the motion detected input video stream. In addition, a motion saliency detector (150a) locates spatial regions of the video with stronger and patterned motion present in the video. The motion saliency detector (150a) locates foreground motion regions or background motion regions in the video. A playback type application detector (150b) detects playback type applications.

[0083] Figure 6 is a diagram (S600) illustrating an example of the operation of a video feature extractor according to various embodiments. The method can be used to detect motion in an input video stream. In addition, the method allows a video feature extractor (140b) (e.g., a lightweight CNN based on MobileNet TV2) to extract important spatiotemporal features from the input video stream for motion detection. In addition, due to the efficiency of the Temporal Segmentation Module (TSM) network with a MobileNet TV2 backbone, the proposed method also uses the TSM network. The proposed method performs a multi-task training process, thereby enabling the learning of useful spatiotemporal features.

[0084] Figure 7 is a diagram illustrating an example of the operation of a playback speed detector (140d) according to various embodiments. The playback speed detector (140d) receives spatiotemporal features and applies the spatiotemporal features to a set of fully connected layers to predict video playback speed. The fully connected layers are trained together with the TSM backbone and saliency detection block, resulting in information sharing and joint learning of all tasks.

[0085] Figure 8 is a diagram illustrating an example of the operation of a playback type detector (140e) according to various embodiments. The playback type detector (140e) receives spatiotemporal features and applies them to a set of fully connected layers to predict the video playback type. The fully connected layers are trained together with the TSM backbone and saliency detection block, thereby enabling information sharing and joint learning for all tasks.

[0086] Figure 91 is a diagram (S900) illustrating an example of the operation of a motion saliency detector (150a) according to various embodiments. The motion saliency detector (150a) includes a combination of Conv+ReLU and upsampling+ReL. The learned spatiotemporal features are then converted into a 2D video saliency map by the motion saliency detector (150a). The saliency detector block is trained together with the TSM backbone and the playback style detection block, thereby achieving information sharing and joint learning for all tasks.

[0087] Figure 10 is a diagram illustrating an example of integration in a single-shot photo pipeline according to various embodiments. Figure 10 , F1, F2, F3 to F8 indicate different frames of a video session in a single shot photo session. A video feature extractor (140b) receives a first motion block and a second motion block in a single shot photo session. A playback speed detector (140d) predicts a video playback speed of the first motion block and the second motion block in the single shot photo session, and a playback type detector (140e) predicts a video playback type of the first motion block and the second motion block in the single shot photo session. The foregoing description of various example embodiments will disclose the general nature of the embodiments herein, which embodiments can be easily modified and / or adapted to various applications of these specific embodiments by applying current knowledge without departing from the general concept, and therefore, such adaptations and modifications should and are intended to be understood to be within the scope and range of equivalents of the disclosed embodiments. It should be understood that the phraseology or terminology employed herein is for the purpose of description and not limitation. Therefore, although the embodiments herein have been described in terms of various example embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modifications within the scope of the embodiments as described herein.

[0088] Although the present disclosure has been shown and described with reference to various exemplary embodiments, it should be understood that the various exemplary embodiments are intended to be illustrative rather than restrictive. Those skilled in the art will further understand that various changes in form and details may be made without departing from the true spirit and full scope of the present disclosure, including the appended claims and their equivalents. It should also be understood that any embodiment described herein may be used in combination with any other embodiment.

Claims

1. A method for determining motion saliency and video playback style in a video by an electronic device (100), wherein: The method comprises: detecting, by an electronic device, a first object and a second object in at least one image frame of the video; determining, by an electronic device, motion types of a first object and a second object in the video; determining, by an electronic device, movement speeds of a first object and a second object in the video; determining, by the electronic device, a first effect to be applied to the first object and a second effect to be applied to the second object based on motion types of the first object and the second object and motion speeds of the first object and the second object; and applying, by the electronic device, a first effect to a first object and a second effect to a second object, The step of determining, by the electronic device, a first effect to be applied to the first object and a second effect to be applied to the second object includes: classifying, by the electronic device, the motion of the first object and the second object into at least one designated category by applying at least one motion-based machine learning model to the motion types of the first object and the second object and the motion speeds of the first object and the second object; and A first effect to be applied to the first object and a second effect to be applied to the second object are predicted by the electronic device based on at least one motion category of the first object and the second object.

2. The method according to claim 1, wherein The method includes: training, by an electronic device, the at least one motion-based training model to predict a first effect and a second effect, wherein the step of training the at least one motion-based training model includes: Determining a plurality of motion cues in the video, wherein the plurality of motion cues include temporal motion features, spatial motion features, and spatiotemporal knowledge of the video; determining a plurality of motion factors in the video, wherein the plurality of motion factors include a direction of motion of each object in the video, a motion pattern of each object in the video, an energy level of motion in the video, and a saliency map of motion in the video; and The at least one motion-based training model is trained based on the plurality of motion cues and the plurality of motion factors to predict a first effect to be applied to a first object and a second effect to be applied to a second object.

3. The method according to claim 1, wherein The step of applying, by the electronic device, a first effect to a first object and a second effect to a second object includes: determining, by the electronic device, a size of the first object and a duration during which the first object is in motion in the at least one image frame of the video; determining, by the electronic device, a size of the second object and a duration that the second object is in motion in the at least one image frame of the video; determining, by the electronic device, a time period during which both the first object and the second object are in motion simultaneously in the at least one image frame of the video based on a size of the first object, a duration of the first object, a size of the second object, and a duration of the second object; and A first effect is applied to a first object and a second effect is applied to a second object by the electronic device during a determined time period.

4. The method according to claim 1, wherein The step of determining, by the electronic device, the motion types of the first object and the second object in the video includes: determining, by the electronic device, a motion type of the first object at a first time interval; and The electronic device determines a motion type of the second object at a second time interval, wherein the first time interval is different from the second time interval.

5. The method according to claim 1, wherein The step of determining, by an electronic device, the movement speeds of the first object and the second object in the video includes: determining, by the electronic device, a movement speed of the first object at a first time interval; and The electronic device determines a movement speed of the second object at a second time interval, wherein the first time interval is different from the second time interval.

6. The method of claim 1, wherein: The first effect includes at least one of a first style effect and a first frame rate speed effect; and The second effect includes at least one of a second style effect and a second frame rate speed effect.

7. An electronic device configured to determine motion saliency and video playback style in a video, wherein: The electronic device comprises: Memory, to store video; A processor is communicatively coupled to the memory, the processor being configured to: detecting a first object and a second object in at least one image frame of the video, determining a motion type of a first object and a second object in the video, and determining a motion speed of a first object and a second object in the video, determining a first effect to be applied to the first object and a second effect to be applied to the second object based on motion types of the first object and the second object and motion speeds of the first object and the second object; and applying a first effect to a first object and a second effect to a second object, The operation of determining a first effect to be applied to the first object and a second effect to be applied to the second object includes: classifying the motion of the first object and the second object into at least one specified category by applying at least one motion-based machine learning model to the motion type of the first object and the second object and the speed of the motion of the first object and the second object; and A first effect to be applied to the first object and a second effect to be applied to the second object are predicted based on at least one motion category of the first object and the second object.

8. The electronic device according to claim 7, wherein: The processor is configured to train the at least one motion-based training model to predict a first effect and a second effect, wherein the operation of training the at least one motion-based training model comprises: Determining a plurality of motion cues in the video, wherein the plurality of motion cues include temporal motion features, spatial motion features, and spatiotemporal knowledge of the video; determining a plurality of motion factors in the video, wherein the plurality of motion factors include a direction of motion of each object in the video, a motion pattern of each object in the video, an energy level of motion in the video, and a saliency map of motion in the video; and The at least one motion-based training model is trained based on the plurality of motion cues and the plurality of motion factors to predict a first effect to be applied to a first object and a second effect to be applied to a second object.

9. The electronic device according to claim 7, wherein: The operation of applying a first effect to a first object and applying a second effect to a second object includes: determining a size of a first object and a duration during which the first object is in motion in the at least one image frame of the video; determining a size of a second object and a duration that the second object is in motion in the at least one image frame of the video; determining a time period during which both the first object and the second object are in motion simultaneously in the at least one image frame of the video based on a size of the first object, a duration of the first object, a size of the second object, and a duration of the second object; and A first effect is applied to a first object and a second effect is applied to a second object during a determined time period.

10. The electronic device according to claim 7, wherein: The operation of determining the motion types of the first object and the second object in the video includes: determining a type of motion of the first object at a first time interval; and The motion type of the second object is determined at a second time interval, wherein the first time interval is different from the second time interval.

11. The electronic device according to claim 7, wherein: The operation of determining the motion speed of the first object and the second object in the video includes: determining a speed of motion of the first object at a first time interval; and The speed of movement of the second object is determined at a second time interval, wherein the first time interval is different from the second time interval.

12. The electronic device according to claim 7, wherein: The first effect includes at least one of a first style effect and a first frame rate speed effect; and The second effect includes at least one of a second style effect and a second frame rate speed effect.

13. The electronic device according to claim 8, wherein: The motion pattern of each object in the video includes spatial local motion indicating motion occurring within first predetermined boundaries in the scene of the video, spatial global motion indicating motion occurring outside second predetermined boundaries in the scene of the video, and motion repetition indicating the same or substantially the same motion occurring within a predetermined time period in the video.

Citation Information

Patent Citations

  • Image capturing apparatus and control method therefor

    CN105407266A

  • Method and system for real time synchronization of video playback with user motion

    CN112153468A