Enhanced video post-processing

By employing AI to automatically post-process video sequences, the method addresses the labor-intensive nature of current post-processing techniques, enhancing video quality and efficiency while preserving the original content.

WO2025131622A1PCT designated stage expired Publication Date: 2025-06-26VIDHANCE AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/084066
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-11-29
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current post-processing methods for video sequences are labor-intensive and often manual, requiring significant time and effort to enhance the visual quality and storytelling of videos.

Method used

A method utilizing artificial intelligence (AI) to automatically post-process video sequences by selecting key image frames, creating requests for AI analysis, and applying feedback for enhancements such as video stabilization, focus adjustment, and color correction.

Benefits of technology

This approach significantly reduces the manual labor required in post-processing, enabling faster and more efficient video enhancement while maintaining the original video unaltered, thus improving video quality and storytelling effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024084066_26062025_PF_FP_ABST
    Figure EP2024084066_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure may include a method performed in a computing device, for enhancing video output quality for a captured video sequence by applying post-processing modification on the video sequence by using an AI.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ENHANCED VIDEO POST-PROCESSING

[0002] TECHNICAL FIELD

[0003] The present disclosure relates post-processing of video sequences. More specifically, the proposed technique relates to using Artificial Intelligence (Al) for adapting the video sequence based on a number of image frames from the video sequence and request instructions. The disclosure also relates to corresponding devices and to a computer program for executing the proposed methods, and to a carrier containing said computer program.

[0004] BACKGROUND

[0005] Post-processing in video refers to the manipulation and enhancement of video footage after it has been filmed. It involves various steps and techniques to improve the visual quality, correct errors, add effects, and finalize the content before it's ready for distribution.

[0006] One aspect of post-processing is video editing. This involves trimming, rearranging, and combining clips, adjusting their timing, and adding transitions to create a coherent and engaging narrative. Editing software like Adobe Premiere Pro, Final Cut Pro, or DaVinci Resolve allows professionals to refine the footage, synchronize audio, and add text or graphics. Color correction and grading also play a significant role in post-processing. Color correction involves adjusting the exposure, contrast, and color balance to ensure consistency throughout the video. Color grading, on the other hand, is more artistic— it sets the overall tone and mood of the video by applying specific color palettes or styles.

[0007] Overall, post-processing in video production is a comprehensive and intricate process that involves multiple stages and tools to refine and enhance raw footage into a polished final product ready for distribution. These processes are typically often manual, or at least semimanual, and labor intensive. Thus, enhanced methods for post-processing video are needed.

[0008] SUMMARY An object of the present disclosure is to provide methods and devices which seek to mitigate, alleviate, or eliminate the above-identified deficiencies in the art and disadvantages singly or in any combination. This object is obtained by a method performed in a computing device, for post-processing a video sequence, the computing device comprising a processor and a communication interface operatively connected to an artificial intelligence, Al, the method comprising: obtaining an original video sequence comprising a first stream of image frames; selecting a plurality of image frames from the first stream of image frames; creating a request to the Al to adapt the original video sequence based on request information and the selected plurality of image frames; sending the selected plurality of image frames and the request including the request information to the Al; receiving feedback from the Al for post-processing the video sequence.

[0009] According to some aspects, the disclosure proposes a computing device configured to automatically post-process a video sequence, the device comprising: optionally a camera, optionally one or more sensors and / or software, a memory, a communication interface, processing circuitry, and optionally an artificial intelligence (or connected to a remote Al), the communication interface being operatively connected to the Al or a remote Al, the processing circuitry being configured to cause the computing device to execute the methods describes above and below.

[0010] According to some aspects, the disclosure proposes a computer program comprising computer program code which, when executed in a computing device, causes the computing device to execute the methods described below and above.

[0011] According to some aspects, the disclosure proposes a carrier containing the computer program, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0012] Other objects and advantages will become apparent to those skilled in the art from a review of the ensuing detailed description, which proceeds with reference to the following illustrative drawings, and the attendant claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 shows an example of video generation and modification.

[0014] Figure 2 is a flowchart of an exemplary method for automatic post-processing of a video sequence to obtain a modified version of the original video sequence of the present disclosure.

[0015] Figure 3 is a block diagram illustrating a computing device configured to generate an automatic post-processing of a video sequence.

[0016] Figure 4 illustrates different system setups using computing devices and local and remote Als for use in the methods of the present disclosure, where 4A shows an integrated Al and 4B and 4C show remote Als.

[0017] Figure 5 shows an example workflow of an embodiment of the present disclosure.

[0018] Figure 6 illustrates an example of modifying a video based on both stabilization and an object of interest, where 6A is an image frame representing an unmodified video, 6B representing a video modified for stabilization, and 6C a sequence of images used in a request to an Al model.

[0019] The figures are not necessarily to scale, and generally only show parts that are necessary in order to elucidate the inventive concept, wherein other parts may be omitted or merely suggested.

[0020] DETAILED DESCRIPTION

[0021] Aspects of the present disclosure will be described more fully hereinafter. The apparatus and method disclosed herein can, however, be realized in many different forms and should not be construed as being limited to the aspects set forth herein.

[0022] The terminology used herein is for the purpose of describing particular aspects of the disclosure only, and is not intended to limit the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0023] In some embodiments a non-limiting term "computing device" is used. This refers to any device having processing capabilities, such as a computer or a smartphone. The computing device may be a "mobile device" or "wireless device". The mobile or wireless device herein can be any type of device capable of communicating with a network node or another mobile device over radio signals or wired communication. The mobile device may be a wireless device, and may include a radio communication device, target device, device to device (D2D) wireless device, machine type wireless device or wireless device capable of machine to machine communication (M2 M), or a sensor equipped with a wireless device. The computing device may be a smartphone, stationary computer, laptop, headset, headmounted (head-worn) camera / device, bodycam, smart glasses, smartwatch, dashcam, action camera embedded equipped (LEE), laptop mounted equipment (LM E), iPad, Tablet, drone, or mobile terminal etc.

[0024] In some embodiments the computing device may be able to capture a video sequence, wherein the computing device comprises or is a "camera device". The camera device herein can be any type of mobile device comprising a camera, such as a smartphone, iPad, Tablet, headset, smart glasses or mobile terminal etc., as discussed above.

[0025] The computing device may comprise sensors, such as motion sensors and / or depth sensors. A motion sensor in the computing device, such as a smartphone, is a component that enables the device to detect and measure various types of motion. It typically consists of multiple sensors, including an accelerometer, gyroscope, and magnetometer. The accelerometer detects linear acceleration and tilt, allowing the smartphone to sense changes in orientation, shake, or movement in a particular direction. The gyroscope measures angular velocity and rotation, providing precise information about the device's orientation and rotational movements. The magnetometer, also known as a digital compass, detects magnetic fields and aids in determining the smartphone's absolute orientation relative to the Earth's magnetic field. A depth sensor, for example in a camera, is a specialized component designed to measure the distance between the camera and objects within its field of view, enabling the creation of 3D maps or depth maps. The depths sensors may be Time-of-Flight Sensors (ToF) that calculate distances by measuring the time it takes for a light signal to bounce off objects and return to the sensor, and are commonly used in smartphones for facial recognition and gesture control, or they may be Structured Light Sensors that project a pattern of light onto objects and measure distortions in the pattern to calculate depth. The depth sensors may also be Stereo Vision Systems utilizing two or more cameras with slightly different viewpoints, which compare the disparities between images to calculate depth. For example, utilizing various technologies like ToF, structured light, or stereo vision, these sensors emit signals or patterns of light and calculate the time it takes for the signals to return or analyze the disparities between multiple viewpoints. This data is then processed to generate depth information, allowing for precise recognition of distances and the creation of detailed depth maps.

[0026] Capturing motion of a camera while filming with a motion sensor, such as a gyroscope, attached to the camera device typically comprise attaching the motion sensor, like a gyroscope, securely to the camera or its support system and ensure that it is properly calibrated and aligned with the camera's orientation. It may also be integrated into the camera device. The sensor is enabled to capture data continuously during filming for data acquisition. The sensor will measure the camera's movements and rotations in real-time, providing information about its position and orientation changes. For data integration, to capture the video footage simultaneously while recording motion sensor data, it is essential to synchronize the timestamps of the video frames and the corresponding sensor readings, to align them accurately during analysis. After capturing the footage, it is possible to analyze the motion sensor data along with the video frames. The gyroscope readings can provide information about camera movements, rotations, and tilts. By examining this data, it is possible to obtain motion information regarding the camera's motion patterns during filming, even to assess the current motion of the camera during the capturing of a specific image frame, which may be incorporated into the motion parameters of each image frame. Hence, the motion parameter of each image frame may be used to assess the current motion of the image frame, including both rotational, translational movement. To convert the estimated motion into real-world units, a camera's intrinsic parameters (focal length, principal point, etc.) and potentially its extrinsic parameters (position and orientation) may be used. To estimate the scale of motion, additional information such as known object sizes or scene depth information may be used. An addition, any known camera motion estimation algorithm using optical flow may be implemented in the current device. Thus, the motion data may also comprise additional data, such as optical flow data and also depth map data. This data may be obtained using one or more depth sensors. A depth map of a video sequence is a representation of the scene's depth or the distance information for each pixel in the video frames. It provides an estimation of the 3D structure of the scene, indicating how far objects are from the camera. A depth map assigns depth values to each pixel based on its distance from the camera viewpoint. Typically, higher values indicate objects closer to the camera, while lower values represent objects farther away. The depth map can be represented as a grayscale image or a separate channel associated with each pixel in the video frames.

[0027] Video frames refer to individual images that compose a video sequence. Each frame is a static representation of the video at a particular point in time and consists of a grid of pixels. The resolution of video frames determines the number of pixels in each frame, such as 1920x1080 for Full HD or 3840x2160 for 4K Ultra HD. The frame rate of a video refers to the number of frames displayed per second (fps). Common frame rates include 24 fps, 30 fps, and 60 fps, which may also be measured in Hertz (Hz), where lfps equalizes 1 Hz. A higher frame rate provides smoother motion and is often used for fast-action content, while a lower frame rate is typically suitable for slower-paced videos. The frame rate, along with the resolution, affects the visual quality and smoothness of video playback, and they are essential parameters to consider when capturing, editing, and displaying video content.

[0028] A camera device may capture a "stream of image frames" constituting a "video sequence", where "image frame" and "video frame" may be used interchangeably herein. Also the terms "image" or "frame" may be used herein to refer to such image frames, and "video" may be used to refer to a "video sequence". The stream of image frames comprises a number of consecutive image frames in time. The computing device may obtain a video sequence, either via a camera device in the computing device, or it may obtain, such as receive, the video sequence from another source, such as another computing device.

[0029] Video generation, capturing and coding of video, may comprise several steps, where postprocessing in video production typically occurs after the footage has been captured. The initial phase of video generation involves capturing the raw video footage using cameras or recording devices. This raw footage might include various scenes, shots, or sequences. Once the raw footage is obtained, initially encoded and decoded, it goes through post-processing. This involves editing, color correction, adding effects, sound editing, and any necessary enhancements to improve the overall quality and storytelling of the video.

[0030] An example of video generation and modification is shown in Figure 1. Figure 1 shows a camera (camera app) of a computing device obtaining a video sequence at an application layer. Input from the one or more sensors from the hardware layer is submitted to the camera app (or other application in the computing device comprising processing circuitry), and the video is initially encoded to e.g., an mp4 encoded movie (unmodified original). In prior art solutions, the video is then optionally post-processed and decoded e.g., in the camera app, and then encoded for distribution into an mp4 encoded movie. However, in the present disclosure, instead the unmodified original is decoded and submitted to an Al for analysis, the Al feeding back a transform to the camera app of the computing device, where the post-processing of the transform is applied before the video is decoded into a mp4 encoded transformed version of the movie.

[0031] A video sequence of the present invention is captured in full resolution, which means that instead removing data, such as when addressing issues such as noise cancelling, white balance, and video stabilization while filming, all the captured unaltered data is kept at the extent possible. In some embodiments, some issues may be handled automatically while filming depending on the settings of the device, but all data should be kept to the extent possible. Thus, the original captured unaltered video (or a copy thereof) is also kept while making one or more modified copies of it, comprising non-destructive video editing, i.e., where alterations and edits made to a video file don't permanently change or modify the original source material. This method preserves the original video file while allowing editors (e.g., the Al) to make changes, apply effects, or perform edits without altering the original content irreversibly. By keeping the full original video, several different modifications can be made.

[0032] The present disclosure relates to post-processing of a captured video sequence. Postprocessing typically occurs between the capture / recording phase (including initial coding) and the final encoding phase, where the raw footage is refined, edited, enhanced, and prepared before it's compressed into a suitable format for distribution. Encoding and decoding happen before and after this post-processing phase, respectively. Post-processing of video typically includes manual work and is labor intensive. Thus, automating processes for post-processing video would be desirable. However, to mathematically define how to post-process a video e.g., using a tracker is very difficult. Thus, an aim of the current disclosure is to define methods based on capabilities of artificial intelligence, such as generative artificial intelligence, for automatic or semi-automatic post-processing of captured video. One aim is to use the Al in methods applied to "full data" unmodified video sequences for automatic determination of what the full data video is trying to tell the viewer, and then automatically process the video to follow the story told, using all the available data. The post-processing may for example relate to cropping the image frames of the video for video stabilization, adjusting focus or focus depth, adjusting white balance, adding artificial bokeh (simulated or artificially generated background blur effect), adjusting (perceived) exposure or exposure time, adjusting contrast or color, adjusting zoom level, such as crop based on an object or event of interest, depth map and audio adjustments.

[0033] Recently, a number of new Al tools / models have been developed, including generative artificial intelligence (Al) tools / models. Generative Al refers to a subset of artificial intelligence that involves systems capable of creating new content, such as images, text, audio, or even videos. Unlike traditional Al systems that rely on predefined rules and data, generative Al uses models that learn patterns from data and can generate new, original content that resembles the data it was trained on. One of the most popular types of generative Al is Generative Adversarial Networks (GANs). GANs consist of two neural networks, a generator, and a discriminator, which work in opposition. The generator creates new content, like images, by generating samples from random noise. The discriminator's role is to distinguish between real and generated content. Both networks improve iteratively as they compete against each other; the generator aims to create content that is realistic enough to fool the discriminator, while the discriminator aims to become better at distinguishing real from generated content. Other types of generative Als include GPTs (Generative Pre-trained Transformer), which models employ Transformer architecture, a neural network design known for its capacity to process and generate sequences of data. GPT models are pre-trained on vast amounts of text data from the internet, learning patterns, structures, and relationships within language, and is typically refined using reinforcement learning from human feedback. Language Vision Models, such as ChatGPT- Vision, may be a preferred model to use.

[0034] The Al, Al tool or model or generative Al-tool or model as referred to herein may be any Al capable of analyzing video and generating feedback regarding proposed modification. The generative Al tool may for example be based on a visual language model, a large language model having video generating properties, such as GPT-4. Several other generative Al tools for generating video also exist, such as DeepDream, DALL-E, Video-LLaVA, GANimation, and GANs for Video Generation. The Al may also be a separate and novel Al created for this purpose, such as handcrafted Al models for specific scenes. The Al used should be able to analyze video content to understand what is happening and where, i.e., such an event monitoring and object targeting, and based on this suggest how to modify the video to enhance the video, i.e., suggest appropriate modification to attain a better video in view default or specific requests. The suggestions may be general (how to get a better video in as many aspects as possible) as default, or it may relate to a certain defined feature, such as optimized cropping for stabilization, e.g., in view of an object or an event of interest. The suggestion may either be executed on the Al by creating the enhanced modified video, or by generating feedback, such as in the form of a transform to apply to the original video within the computing device. The Al may be present in / on the computing device, or may be remote, such as in a cloud. If being present on the device, the computing device can call upon the Al directly. When present remotely, the request and information needed may be transmitted to the Al, and the feedback received from the Al, such as via wireless transmissions.

[0035] Instead of creating content de novo, the present disclosure proposes to use the capabilities of generative Al to post-process unaltered video sequences automatically, which may be performed upon request by a user. Accordingly, it is an object of the invention to define a method of how to use a generative Al or generative Al tool to post-process a video sequence.

[0036] According to embodiments of the present invention, after obtaining a video sequence, the video sequence may be pre-processed, such as to find objects and label them, and prepared for sending to a cloud by reducing its size (selecting image frames). The image frames may then be sent to an Al which analyses the video and suggests edits for post-processing to obtain an altered (enhanced) video.

[0037] Thus, a method of post-processing a video sequence is proposed, comprising obtaining an original video sequence comprising a first stream of image frames, pre-processing said video sequence, which may include finding and labelling objects or events, and selecting a plurality of image frames from the first stream of image frames, creating a request to the Al to analyze and adapt / suggest how to adapt the original video sequence based on request information and the selected plurality of image frames, sending the selected plurality of image frames and the request including the request information to the Al and receiving feedback from the Al for post-processing the video sequence.

[0038] In some embodiments, a scaled down version of the video is inputted to the Al, or parts of the video is sent, such as a subset of the video frames of a video sequence is selected and inputted into (sent to) the Al. This may be advantageous if the Al is located outside of or remote from the computing device, and e.g., a wireless transmission comprising the selected frames needs to be sent. But selecting fewer frames may also be advantageous in a local Al as fewer frames is faster to process. As an example, one frame every x:th frame or every x seconds may be selected and sent, such as one frame per second may be selected and sent to the generative Al for analysis. If the video sequence to be modified shows a faster course of events, e.g., a fast-moving scene, more than one frame per second may be selected. The number of frames that are sent to the Al-tool (if not present in the computing device) may also be based on the available bandwidth, such that more frames are sent when the bandwidth is higher. This may be an automated process where the computing device decides, based on the available bandwidth, how many image frames of the video that are sent. In optimal conditions, the whole video (all frames) may be sent.

[0039] The selected image frames may be represented by thumbnails. Thumbnails of image frames selected from a video are essentially still images that represent specific points or scenes within the video. These thumbnails may serve as visual representations or previews of the video content. The thumbnails may be organized into a thumb map. A thumb map typically refers to a collection or grid of thumbnails representing various frames or scenes from a video.

[0040] The general concept of creating a thumb map of a video involves extracting frames, converting them into thumbnails, and arranging them systematically to create the thumb map representation of the video. Thus, to create the thumb map, or populate a grid with thumbnails, one may start by selecting and extracting frames from a video, e.g., using video editing software or specialized tools that allow frame extraction. Frames at regular intervals (e.g., a number of frames per second) or specific key moments may be selected to represent the video content. The thumbnails may be created by resizing and formatting the extracted frames as thumbnails. Thumbnails are usually smaller than the original frames and are often standardized in size for consistency in the thumb map. The thumbnails may be arranged in a grid format to create the thumb map. If necessary, it may be possible to label or annotate the thumbnails with relevant information, such as timestamps, titles, or descriptions, to provide context about the corresponding video segments. The process of creating a thumb map involves organizing and arranging these thumbnails visually, often in a grid or mosaic layout, to provide a quick overview or representation of the video's content. Thus, input to the Al tool may be a video or part of a video, a number of selected and extracted image frames, thumbnails or a thumb map thereof, besides the request information, which may comprise information of what is to be done, and additional data, such as sensor data. In some embodiments, multiple streams from different cameras (ultra-wide, main, tele, front) may be used as input. One may film an object with several cameras at the same time, which different cameras may be from different angles, or cameras (e.g. on the same device) using different focus or different zoom, etc., then it's possible to jump between different camera streams to tell the story as well as possible. For example, a close-up in a piece of wood with the telephoto camera.

[0041] Thus, the method includes a functional call, i.e., a specific instruction or command given to the Al tool to perform a certain action or function. It's a way to guide the tool to generate or complete a task based on the provided instructions within the prompt. The Al tool may either (if the whole video was uploaded / inputted / sent to the Al tool) respond with a modified video sequence. Alternatively, only the transform to be used is fed back to the computing device. In the context of video editing or processing, a "transform" generally refers to an operation or modification applied to alter certain characteristics of a video clip.

[0042] A transformative process applied to a video involves altering its visual and / or auditory characteristics to create a modified output. This transformation can take various forms, such as changing the video's color palette, applying filters or effects, adjusting the playback speed, adding overlays or graphics, or even manipulating the overall style or mood through editing techniques. For instance, a color grading transform might enhance or entirely change the video's color scheme, affecting its atmosphere or tone, while a speed adjustment transform can either slow down or speed up the video's sequence. Transformations are a fundamental part of video editing, allowing to modify and enhance video sequences / clips to achieve desired visual effects or alterations without permanently changing the original video content until the final export or rendering process. Transform effects can be added to a video sequence within a video editing software or platform. These effects might include transformations like scaling (resizing), rotation, cropping, flipping, or applying various visual effects like blurring, color adjustments, or distortions. These transformations are added as layers or effects onto the video clip but don't permanently alter the original video file until it's exported with the applied effects. The transforms may also be performed on the video, where the application of these transformations is to modify the appearance or characteristics of the video during the editing process. For example, to resize or rotate a video clip, apply color correction, or add visual effects within the editing software, is performing transforms on that video. However, until the edited video is rendered or exported, these transformations are typically non-destructive, meaning the original video file remains unchanged. Thus, in some embodiments, only the transform to be used is obtained from the Al tool, which transform only is applied to the video when playing back the video. Hence, no new video needs to be created and stored. Only if the video is to be shared, a new video based on the transform may be created.

[0043] The transform may be detailed, such as a matrix transformation, e.g., a mathematical operation applied to a matrix that results in a modified matrix output. This transformation may involve information regarding multiplying a given matrix (representing a set of data or geometric points) by another matrix called the transformation matrix. Each element in the output matrix is calculated based on specific rules defined by the transformation matrix. For instance, in 2D transformations, such as rotation, scaling, or translation, the transformation matrix would consist of values that dictate how the original points in the plane are rotated, resized, or moved. In alternative embodiments, the transform or feedback from the Al may be on a high level, and to be interpreted locally in the computing device. In some examples, the feedback may include high-level instructions such as "zoom on object 1", where 1 may be a dog that has been identified as an object of interest, where the processing circuitry / processor of the computing device interprets and translates these instructions to a transform to be applies on the original video sequence.

[0044] JSON (JavaScript Object Notation) is a lightweight data interchange format that is easy for humans to read and write and easy for machines to parse and generate. A JSONObject is a data structure used in programming. In e.g., JavaScript, Python, Java, etc., it is a collection of key-value pairs where keys are strings and values can be various data types including strings, numbers, arrays, other JSONobjects, boolean values, or null. JSONObjects are commonly used for exchanging data between a server and a client in web development, in configuration files, for API responses, and in many other scenarios where structured data needs to be transferred or stored. Thus, the feedback from the Al tool comprising the transform to be used may be in the form of a JSONObject.

[0045] The proposed methods may be carried out by a computing device, such as a computer or smartphone. In some embodiments, the methods may be carried out on an app (application) on the computing device, i.e., a software program developed for end-users to accomplish specific computing tasks. The app may be a dedicated app for video modification, or the processing capabilities may be built into a camera app or gallery app on the computing device. The prompt to the generative Al, i.e., the request or request information, may be built into the application, such that a user may not need to write any text, but may simply request that the original video be modified / enhanced. As a default "enhanced quality video" is generated. The Al analyses the video and optimes different features based on the content, such as based on events or objects, adapting the focus, lightning, contrast, white balance, framing / zoom and exposure, and also image / video stability. Video stabilization used in video editing and processing to reduce unwanted motion or shakiness in footage captured from handheld or unstable sources, typically involving algorithms and software that analyze the video frames and compensate for abrupt movements or vibrations by adjusting the positioning and orientation of the frames, to produce smoother and more visually steady video sequences, enhancing the overall quality and watchability of the footage by minimizing distracting or jarring movements that could otherwise detract from the viewing experience. According to the present disclosure, video stabilization may be performed using the Al by inputting the image frames, and also sensor data. The camera device may have an inertial measurement unit (I M U), a sensor unit with a combination of accelerometers, gyroscopes and magnetometer sensors, capable of easily calculating orientation, position, and velocity of a camera device. Thus, motion information (data) from the IMU may be sent to the Al tool, to be used for cropping for image stabilization purposes. Gyroscope data may e.g., be used for cropping the video.

[0046] Thus, in some embodiments, the Al receives image data (a number of image frames or thumbnails / thumb map), a request for video enhancement including request information, where the request information comprise information both about what should be edited, such a default or crop for image stabilization, crop based on an object, and additional data from sensors in the computing device, such as IMU data. The computing device may comprise e.g., motion sensors capturing motion data, and audio sensors capturing sound. The modification or enhancement of the video may relate to many different features, such as adjusting one or more of cropping the image frames of the video for video stabilization, focus or focus depth, white balance, exposure or exposure time, contrast or color, zoom level, depth map and audio.

[0047] Optionally, there may be different kinds or types of modifications that may be automatically requested (such as by selecting a button in the app), such as adapt focus or cropping of video, adjust colors, or similar. If the user is not pleased with the result, the user may again push the button (such as a "redo" button) to get a new version of the video. Thus, further modifications of the video are possible, or new versions of modifications on the original video, where a new alternative is generated (i.e., not the same modification as suggested in the first instance). In other embodiments, the user is also able to modify the result, e.g., by input prompt information (request information) to the request manually. The prompt information may be text or audio / speech based (talk). For example, the user may state that a specific event or object should be in focus, such as "focus on the clown", "make sure everyone is in the picture", or "zoom in when she hits the ball". The input may also be based on gestures, touch and / or speech (using e.g., a smartphone) or drawings, or similar. The video modification may be on demand by a user.

[0048] In an alternative embodiment, the app is running in the background and generates updated transformations of captured video sequences automatically. For example, cutting, transitions etc. may be automatically addressed to enhance the video sequence. In one embodiment, a default "enhanced" video copy (or transform thereof) may be attained automatically, without requiring the user to explicitly request the enhancement, this all happens automatically using a default request to the Al.

[0049] In a preferred embodiment, the original video / a copy of the original video is kept on the computing device, and a new version / copy is generated with the modifications. Thus, no information is lost by performing the methods of the present disclosure. The copies may be kept on the device, in the cloud, or be divided in between. The proposed modification may be applied in a non-destructive manner to the original video sequence, such that the original video sequence is kept unaltered, in the computing device or in a remote location such as a cloud, and the modification may only be shown during playback.

[0050] The main concept of the proposed methods is to use the capabilities of the generative Al to interpret visual data, suggest post-processing modifications to the video for automatic postprocessing of the video, thus replacing some or all of the manual labor performed when post-processing video. The original or source video is typically full resolution unprocessed video. The capabilities of target and event recognition and object tracking may be used for focusing the video on an object or event of interest, both regarding zoom (crop) and resolution. Blurring of selected objects may also be performed.

[0051] Regarding object tracking, the Al may analyze the video and determine that a certain object seems to be the object of interest, and thus modify, or propose to modify, the video based on said object, such as focusing the video on that object and cropping the video so that the object is e.g., centered. The Al may also focus on a certain event, and make sure that this event is put in the center of the modified video.

[0052] As illustrated in Figure 2, the present disclosure discloses methods performed in a computing device, for post-processing a video sequence, the computing device comprising a processor and a communication interface operatively connected to an artificial intelligence, Al, the method comprising: obtaining (SI) an original video sequence comprising a first stream of image frames; selecting (S2) a plurality of image frames from the first stream of image frames; creating (S3) a request to the Al to adapt the original video sequence based on request information and the selected plurality of image frames; sending (S4) the selected plurality of image frames and the request including the request information to the Al; and receiving (S5) feedback from the Al for post-processing the video sequence.

[0053] The video may be obtained by the computing device, e.g., by being captured in the computing device, such as using a camera in the computing device, or may be captured elsewhere in a remote camera device and received in the computing device, along with associated data (such as from sensors in the remote camera device). Preferably, as much data as possible is kept during the video capture, and the original video sequence is obtained as an unaltered video sequence, preferably yet not being post-processed or finally encoded for distribution. The video sequence may then be pre-processed, which may comprise selecting the image frames to send to the Al.

[0054] The selected plurality of image frames and the request is sent to the Al, wherein the Al may be a remote Al, present at a location distinct from the computing device, such as in a cloud. Alternatively, the Al is integral to the computing device, such that the Al is present on the computing device. The Al may be a generative Al, such as a generative Al tool or generative Al model capable of analyzing and / or generating video content.

[0055] In some embodiments, a subset of the images is selected and processed by the Al. Thus, selecting a plurality of image frames from the first stream of images may comprise selecting a subset of image frames from the first stream of images. In some embodiments, the selected image frames comprise a scaled down version of the original video sequence. The selection of the subset of image frames or scaled down version of the original video sequence may be automatic within the computing device, depending on the available bandwidth for transmitting the selected images frames to the remote Al.

[0056] The feedback from the Al may come in different forms. The feedback from the Al may comprise suggested modifications to be applied to the original video sequence. The suggested modifications from the Al may be a transform that may be applied to the original video sequence to obtain a modified video sequence. The transform may be a matrix transformation to be used, or may be more high-level instructions, to be interpreted locally in the computing device. These suggested or interpreted modifications are typically applied in a non-destructive manner to the video, and a copy of the original video sequence is kept, either being stored on the computing device or in a cloud. In some embodiments, altered versions of the video are stored on the computing device or in a cloud, or alternatively, the modifications are applied only during playback of the video, such that only the original video and the suggested modifications / transforms are stored.

[0057] In some embodiments, instead of selecting a subset of the image frames, selecting a plurality of image frames from the first stream of image frames includes selecting all the image frames of the first stream of image frames. This may depend on the available bandwidth for a remote Al, or may be applied when the Al is an integral Al. In such cases, the feedback may not only be suggestions or a transform, but may comprise a full modified version of the original video sequence.

[0058] When the computing device comprises a camera, obtaining the original video sequence may be performed by capturing the video sequence using a camera in the computing device. The computing device may also comprise one or more sensors, such as motion sensors, depth sensors and audio sensors, where the motion sensors may be selected from a gyroscope sensor, an accelerometer sensor and a magnetometer sensor, and the request information may comprise sensor data from the one or more sensors, such as IMU data, depth information etc. The request to the Al may be a command to perform video enhancement. The command may be explicit by the user, such as by pressing a button stating "enhance video", or may implicit, such as generated automatically when a video is captured, added to an app etc. The request will comprise request information, i.e., an instruction or prompt to the Al regarding what should be performed by the Al. The request information may be set to a default mode, wherein the Al receives the task to enhance the video based on its own analysis, without further guidance by the user. This may be used e.g., when the request is implicit (automatic), or when the user has no special request, but simply wants a generally enhanced video. Thus, the app or tool used for sending the request will have an in built default request, where the Al may analyze the video and suggest alterations based on its in built knowledge. In other embodiments, the user may indicate what the Al should focus on. This may for example be attained by using preset "buttons" in the computing device, or may be fed in manually by the user. Thus, the request information may comprise prompt information about what type of modifications and / or what features the post-processing of the original video sequence should relate to. The post processing, or the type of post processing modification, may comprise one or more of: cropping the image frames of the video for video stabilization, focus or focus depth, artificial bokeh, white balance, exposure or exposure time, contrast or color, zoom level, depth map and audio. Similarly, a feature that the post processing should relate to may be one or more of: object tacking, object recognition, event tracking, event recognition, desired perceived atmosphere or type of video to be attained.

[0059] If the user is not content with the result, or wants further modifications, the method may be repeated iteratively one or more times to obtain several versions of a modified video sequence, wherein the selected image frames and / or request information are changed in each repeated iteration. The selected image frames in the repeated method are selected from the original video sequence or a previously modified video sequence version. Thus, it Is possible to obtain several different versions or suggestions based on a same original video, or additional modification may be added on top of a previously modified video. The one or more predefined modification modes may be defined and selected by a user, which may be performed by using default or predefined options (modes) in the computing device, such that prompt information in the request information is predefined and just selected but not defined by the user, wherein the mode may be a default mode or a mode relating to a type of modification. In other embodiments, the user may more freely indicate what modifications to be made, wherein the user is able to add prompt information to the request information in the form of text, audio, touch gestures or motion of the computing device.

[0060] In some embodiments, the methods are performed by an application on the computing device, such as a distinct application for said purpose, or as an integral part of a video library or camera application.

[0061] Turning now to Figure 3, which is a schematic block diagram that illustrates some modules of an example embodiment of a computing device being configured for post-processing a video sequence. The computing device is configured to implement all aspects of the methods described in relation to Figure 2 and the methods defined above.

[0062] The computing device 10 may comprise a camera 11, configured for capturing a video sequence. It may further comprise one or more sensors 12, the sensors may be motion sensors for monitoring the motion of the camera device and / or software for estimating the motion, and thus the camera, during capturing of the video, or may be depth sensors for acquiring depth information. The camera device further comprises a memory 13 for storing the video sequence, and instructions for performing the current methods. The camera device further comprises a communication interface 14, such as a radio communication interface (i / f) configured for communication with another device, such as a recipient device. The radio communication interface 14 may be adapted to communicate over one or several radio access technologies. The communications interface may be configured to communicate with an Al being located remote from or in the computing device. Thus, optionally, the computing device may include an Al 17A, such as Al model / tool, to which the communication device is operatively connected, or it may be connected to a remote Al 17B.

[0063] The computing device 10 comprises a controller, CTL, or a processing circuitry 15 that may be constituted by any suitable Central Processing Unit, CPU, microcontroller, Digital Signal Processor, DSP, etc. capable of executing computer program code. The computer program may be stored in a memory, MEM 13. The memory 13 can be any combination of a Read And write Memory, RAM, and a Read Only Memory, ROM. The memory 13 may also comprise persistent storage, which, for example, can be any single one or combination of magnetic memory, optical memory, or solid state memory or even remotely mounted memory. The processing circuitry 15 may be configured to cause the computing device 10 to execute the methods described above and below.

[0064] The computing device may be a smartphone, stationary computer or laptop, iPad, Tablet, bodycam, head-mounted camera, drone, bodycam, smartwatch, dashcam, action camera or mobile terminal.

[0065] According to some aspects, the disclosure relates to a computer program comprising computer program code which, when executed, causes a computing device to execute the methods described above and below. According to some aspects the disclosure pertains to a computer program product or a computer readable medium holding said computer program. In some embodiments, the processing circuitry 15 may further comprise both a memory 13 storing a computer program and a processor 16 (not shown), the processor being configured to carry out the method of the computer program.

[0066] One embodiment includes computing device (10), configured to enhance video output quality for a captured video sequence, the device comprising: a camera (11), one or more sensors and / or software (12), a memory (13), a communication interface (14), and processing circuitry (15) configured to cause the camera device (10) to perform the methods above and below. In some aspects, the computing device is a personal computer, such as a stationary computer or a laptop, or a mobile device, such as a smartphone, iPad, Tablet, headset, smart glasses or mobile terminal.

[0067] Figure 4 illustrates different system setups using computing devices (10) and local and remote Als (17) for use in the methods of the present disclosure. Figure 4A depicts a computing device being a smartphone, where the Al (17A) is integrated into the smartphone. Figure 4B and 4C show remote Als (17B) present in a cloud, where 4B illustrates the computing device a smartphone and 4C as a computer or terminal. 1 imized zoom

[0068] In an example embodiment of the present disclosure, a video is shot using a smartphone. The shot video is unprocessed. The user enters an application on the smartphone designed for modifying a video sequence. The user selects the recently shot video using the app to be the video to be modified. In this example, the user also selects the feature "zoom" in the app as the feature that should be optimized. In other embodiments, this may happen automatically (by default).

[0069] The app then selects a number of image frames from the video sequence to be modified, such as one frame per second, e.g. selects an image frame every second in the sequence of consecutive image frames in time, and sends, via a communication interface in the smartphone, a request comprising request information that an optimized zoom is requested to a generative Al tool capable of editing videos, the generative Al tool being a cloud based application. Some Al processing may be performed on the device, and some in the cloud.

[0070] The generative Al tool analyzes the received image frames to determine which objects that seem most relevant and should be in the center of the video. The Al tool then determines how to zoom, and thus crop, the video frames to center on the most relevant objects. The information of how adapt the image frames, in this case how to crop the video frames, is then returned to the smartphone app, which uses said information to adapt the cropping of the full video to obtain the modified video sequence having an optimized zoom.

[0071] Alternatively, the user may, instead of using the automated button optimized zoom, write or speak the prompt (request information) to the Al tool. Such a request could thus include a selected number of image frames and a prompt (request information). In a basic example, the prompt could read for example: "This is a number of frames (1 per second) from a video.

[0072] I want to go back to the video and edit the cropping dynamically throughout the video, based on the content. Please review the image / video and first analyze what happens in it, then say what is important in it and finally suggest a zoom level from 1 (none) to 3 (high) and what to keep in the zoomed frame."

[0073] If the result is not satisfactory, the process may be reiterated to attain a better result. For example, a specific region may need more details to be handled properly. Thus, a new request may be sent, either using different selected images (either from the original video or from the first modified version of the video) or a different prompt (if manual input is selected), or both.

[0074] Other features that may be addressed similarly includes artificial blur (bokeh), where the Al is instructed to blur a certain object, or post-adjustments of exposure (altering perceived amount of light that has reached the camera sensor while filming) and white balance, i.e., the color temperature at which white objects on film look white. for sequences to generate a short video from thumbnails

[0075] In an example embodiment of the present disclosure, a number of thumbnails from an original source video is sent to an Al tool from a computing device, the thumbnails being numbered with a figure, an integer starting from 1 and ending at the total number of thumbnails, e.g., from 1-24 if 24 thumbnails are present. The Al tool also receives request information, either explicitly from the user or intrinsically build into the app, stating that the information shows thumbnails from a video where each figure shows the time in seconds, and that a video of x seconds, e.g., 6 seconds for 24 thumbnails, is desired, asking what sequences that the tool recommends to use. The tool then recommends thumbnails to best represent the video, based on analysis of content therein. In the present case, a video of people playing golf is used, and upon evaluating the thumbnails, the tool recommends for the sequence 1-2 s to start with image 3 or 4, to show the enthusiasm and the group dynamics, for the sequence of 2-3 s to use image 14 or 15 to show the actual golf activity, and for the sequence 3-4 to follow up with image 17 or 21, to capture interaction and communication amongst the participants, etc. Thus, the 6s video is created / suggested how to be created in the Al feedback, which instructions may be implemented in the computing device to obtain the (altered) video. Vision Model

[0076] In an example, as illustrated in Figure 5, a smartphone may obtain a video sequence by filming. To obtain a video sequence comprising as much data as possible, the device films with as wide angle as possible, and keep all video and motion data. Pre-processing of the video may then be performed, including finding objects and labeling them, and prepare for sending to the cloud (smaller size), e.g., selecting image frames to be included in the request. The Al, such as a Language Vision Model may then receive a request from the device to analyze the video, the model being used to understand the video and suggest several alternative edits. The suggestions are then included in the feedback from the Al for pos-processing the video. The device may then apply one of the edits given by the cloud engine (Al). In the last step, the edited video may be shown to the user, which may also offer to show alternative edits. The user may also prompt for specific edits. based on an object and free text from user

[0077] Video stabilization may be performed by cropping the video, as illustrated in Figure 6A and 6B, showing a first image frame in a video before and after stabilization where a little girl plays in a playhouse. The video sequences illustrated by 6B has been adjusted for stabilization, however, the object of interest is not the cottage, as may be achieved by default, but actually the little girl playing in the house. The girl exits the playhouse at one time in the video, which footage is not shown (cropped) in the altered video, which focuses on the playhouse. Hence, it is desirable to focus on the girl as the object of interest in the video.

[0078] In this case, the user may send a prompt to the Al model including the request information "Explain what is in this image", and the result / feedback from the model may be "This image features a snowy scene with a bright blue playhouse, distinguished by its wooden structure. You can observe snow on the roof (3), indicating recent snowfall or cold temperatures preventing the snow from melting. There is a decorative star (2) attached to the front of the playhouse. A child (4) is visible inside the playhouse, looking out through the window. In the foreground, snow covers the ground (1), and footprints indicate that people have been moving around the area. In the background, slightly out of focus, there is another person in the snow, which creates a sense of depth and activity in this wintery outdoor setting."

[0079] The user may then further prompt "The previous image is a high resolution image from a video. Here are low resolution frames of the video (1 per second). The original video was recorded at 30 fps, the images have their actual frame number printed at the top left of the image. Please look at the video and first analyze what happens in it, then describe what is important in it. In a list then give the following information about the video: 1. What is the video about? 2. What are the most important aspects of the video. 3. If the video were to be edited (cropped and cut) what is important to keep.4. Finally, in a table, give me 3 suggestions for possible ways this video could be edited to tell a story." and include the image frames in the request, as shown in Figure 6C.

[0080] The Al model may respond by "Based on the provided frames, it appears the video captures a sequence of a child playing in the snow near a blue playhouse. Video Analysis: The child initially appears inside the playhouse. The child then exits the playhouse and moves through the snow. The video ends with the child still outside the playhouse. What is the video about? The video is about a child playing near a blue playhouse in a snow-covered area, interacting with their snowy surroundings. What are the most important aspects of the video? The child's actions and interactions with the playhouse and the snowy environment. The snow- covered setting and winter ambiance. The emotions or movements of the child, which may provide narrative or emotional significance to the video. If the video were to be edited (cropped and cut) what is important to keep? Key moments where the child is actively engaging with the environment, such as exiting the playhouse or playing in the snow. Scenes that establish the setting, such as the snowy landscape and the exterior of the playhouse with the child inside or nearby. Any clear reactions or expressive moments from the child that add to the story or the emotion of the video."

[0081] The Al tool may then make a table of suggestions for the video edits to tell a story, as shown in table 1.

[0082] Table 1

[0083] In some embodiments, the Al model will feedback what modifications to be performed in detail in a computer readable format, such as JSON. When applying these suggestions to the video sequence, the edited video keeps the little girl in focus, while simultaneously addressing the stabilization, thus enhancing the video quality both in view of generic features for a viewer (minimize shake) and in view of a specific request, such as an object of interest.

[0084] Thus, the content of this disclosure thus enables a computing device to automatically postprocess a video sequence to enhance its video output quality, by requesting an Al model to suggest editions based on an original unaltered video sequence, to modify the video sequence in a non-destructive way.

[0085] Aspects of the disclosure are described with reference to the drawings, e.g., block diagram and / or flowchart. It is understood that several entities in the drawings, e.g., blocks of the block diagrams, and also combinations of entities in the drawings, can be implemented by computer program instructions, which instructions can be stored in a computer-readable memory, and also loaded onto a computer or other programmable data processing apparatus. Such computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer and / or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer and / or other programmable data processing apparatus, create means for implementing the functions / acts specified in the block diagrams and / or flowchart block or blocks.

[0086] In the drawings and specification, there have been disclosed exemplary aspects of the disclosure. However, many variations and modifications can be made to these aspects without substantially departing from the principles of the present disclosure. Thus, the disclosure should be regarded as illustrative rather than restrictive, and not as being limited to the particular aspects discussed above. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0087] The description of the example embodiments provided herein have been presented for purposes of illustration. The description is not intended to be exhaustive or to limit example embodiments to the precise form disclosed, and modifications and variations are possible in light of the above teachings or may be acquired from practice of various alternatives to the provided embodiments. The examples discussed herein were chosen and described in order to explain the principles and the nature of various example embodiments and its practical application to enable one skilled in the art to utilize the example embodiments in various manners and with various modifications as are suited to the particular use contemplated. The features of the embodiments described herein may be combined in all possible combinations of methods, apparatus, modules, systems, and computer program products. It should be appreciated that the example embodiments presented herein may be practiced in any combination with each other.

[0088] It should be noted that the word "comprising" does not necessarily exclude the presence of other elements or steps than those listed and the words "a" or "an" preceding an element do not exclude the presence of a plurality of such elements. It should further be noted that any reference signs do not limit the scope of the claims, that the example embodiments may be implemented at least in part by means of both hardware and software, and that several "means", "units" or "devices" may be represented by the same item of hardware.

[0089] The various example embodiments described herein are described in the general context of method steps or processes, which may be implemented in one aspect by a computer program product, embodied in a computer-readable medium, including computerexecutable instructions, such as program code, executed by computers in networked environments. A computer-readable medium may include removable and non-removable storage devices including, but not limited to, Read Only Memory (ROM), Random Access Memory (RAM), compact discs (CDs), digital versatile discs (DVD), etc. Generally, program modules may include routines, programs, objects, components, data structures, etc. that performs particular tasks or implement particular abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps or processes.

Claims

CLAIMS1. A method performed in a computing device, for post-processing a video sequence, the computing device comprising a processor and a communication interface operatively connected to an artificial intelligence, Al, the method comprising: obtaining (SI) an original video sequence comprising a first stream of image frames; selecting (S2) a plurality of image frames from the first stream of image frames; creating (S3) a request to the Al to adapt the original video sequence based on request information and the selected plurality of image frames; sending (S4) the selected plurality of image frames and the request including the request information to the Al; and receiving (S5) feedback from the Al for post-processing the video sequence.

2. The method according to claim 1, wherein the original video sequence is obtained as an unaltered video sequence.

3. The method according to claim 1, wherein the Al is a remote Al, present at a location distinct from the computing device, such as in a cloud.

4. The method according to claims 1-3, wherein selecting a plurality of image frames from the first stream of images comprises selecting a subset of image frames from the first stream of images.

5. The method according to claim 1-4, wherein the selected image frames comprise a scaled down version of the original video sequence.

6. The method according to claims 4-5, wherein the selection of the subset of image frames or scaled down version of the original video sequence is automatic within the computing device, depending on the available bandwidth for transmitting the selected images frames to the remote AL7. The method according to claim 1, wherein the feedback from the Al comprises suggested modifications to be applied to the original video sequence.

8. The method according to claim 4, wherein the suggested modifications from the Al is a transform that may be applied to the original video sequence to obtain a modified video sequence.

9. The method according to claims 4-5, wherein the modifications are applied in a nondestructive manner to the video, and a copy of the original video sequence is kept on the computing device.

10. The method according to claims 4-6, wherein the modifications are applied during playback of the video.

11. The method according to claims 1-2, wherein the Al is present on the computing device.

12. The method according to claim 11, wherein selecting a plurality of image frames from the first stream of image frames includes selecting all the image frames of the first stream of image frames.

13. The method according to claims 11-12, wherein the feedback from the Al comprises a modified version of the original video sequence.

14. The method of claims 1-13, wherein obtaining the original video sequence is performed by capturing the video sequence using a camera in the computing device.

15. The method of claim 14, wherein the computing device comprises one or more sensors, and wherein the request information comprises sensor data from the one or more sensors, such as IMU data.

16. The method according to claim 15, wherein the one or more sensors are motion sensors selected from a gyroscope sensor, an accelerometer sensor and a magnetometer sensor.

17. The method according to claims 1-16, wherein the request information may be set to a default mode, wherein the Al receives the task to enhance the video based on its own analysis, without further guidance.

18. The method according to claims 1-16, wherein the request information comprises prompt information about what type of modifications and / or what features the postprocessing of the original video sequence should relate to.

19. The method according to claim 1-18, wherein post processing, or the type of post processing modification, comprises one or more of: cropping the image frames of the video for video stabilization, focus or focus depth, artificial bokeh, white balance, exposure or exposure time, contrast or color, zoom level, depth map and audio.

20. The method according to claims 1-19, wherein a feature of the post processing is selected from one or more of: object tacking, object recognition, event tracking, event recognition, desired perceived atmosphere or type of video to be attained.

21. The method according to claims 1-20, wherein the method may be repeated iteratively one or more times to obtain several versions of a modified video sequence, wherein the selected image frames and / or request information are changed in each repeated iteration.

22. The method according to claim 21, wherein the selected image frames in the repeated method are selected from the original video sequence or a previously modified video sequence version.

23. The method according to claims 1-22, wherein one or more predefined modification modes may be defined and selected by a user, such that prompt information in the request information is predefined and not defined by the user, wherein the mode may be a default mode or a mode relating to a type of modification.

24. The method according to claims 1-23, wherein the user is able to add prompt information to the request information in the form of text, audio, touch gestures or motion of the computing device.

25. The method according to claims 1-24, wherein the method is performed by an application on the computing device.

26. The method according to claims 1-25, wherein the Al is a generative Al tool.

27. A computing device (10), configured to automatically post-process a video sequence, the device comprising: optionally a camera (11), optionally one or more sensors and / or software (12), a memory (13), a communication interface (14), processing circuitry (15), and optionally an artificial intelligence (17A), the communication interface (14) being operatively connected to the Al (17A) or a remote Al (17B), the processing circuitry being configured to cause the computing device (10) to execute the methods of claim 1-26.

28. The method according to claims 1-26, or the computing device according to claim 27, wherein the computing device is a smartphone, stationary computer or laptop, iPad, Tablet, bodycam, head-mounted camera, drone, bodycam, smartwatch, dashcam, action camera or mobile terminal.

29. A computer program comprising computer program code which, when executed in a computing device, causes the computing device to execute the methods according to any of the claims 1-26.

30. A carrier containing the computer program of claim 29, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

Citation Information

Patent Citations

  • VIDEO EDITING METHOD USING AUTOMATABLE ADAPTIVE MODELS

    FR3044816A1

  • Neural network for video editing

    US20150208023A1

  • Systems and methods for generating composite media using distributed networks

    US20230274545A1

  • Systems and methods for automating video editing

    US20230386520A1

  • Camera operable using natural language commands

    WO2018097889A1