Methods for the autonomous production of edited video streams

The use of a pre-trained machine learning model to adjust camera settings for automated video production addresses the inefficiencies of conventional systems, achieving high-quality, low-latency video production at a lower cost.

JP2026511257APending Publication Date: 2026-04-10STUDIO AUTOMATED SPORT & MEDIA HOLDING BV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Conventional automated video production systems struggle to capture context-dependent situations effectively, leading to inferior video quality compared to human-operated systems, particularly in medium-scale and small-scale events, and are often economically infeasible.

Method used

An autonomous video production method using a pre-trained machine learning model to determine navigation parameters for mechanically and/or optically adjustable cameras, allowing for the creation of an edited video stream by iteratively obtaining frames based on input frames, considering contextual information and minimizing computational load.

Benefits of technology

Enables high-quality, low-latency, and cost-effective automated video production by using machine learning to adjust camera settings in real-time, providing smoother transitions and reducing the need for extensive human labor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511257000001_ABST
    Figure 2026511257000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a computer implementation method for the autonomous production of an edited video stream of a subject region using a machine learning model. The machine learning model is specifically configured and trained to determine navigation parameters that represent corrections to frame settings for the input video frame, such as corrections to pan, tilt, or zoom settings, based on the input video frame. The method includes obtaining a first frame for an edited video stream of a first sub-region of the subject region and inputting the first frame into the machine learning model. Based on the first frame, navigation parameters relating to a second sub-region of the subject region are determined. Based on the determined navigation parameters, a second frame for an edited video stream of the second sub-region is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to autonomous video production of events, such as sports events.

Background Art

[0002] Popular events such as live sports events and concerts are often recorded by various cameras. Video footage recorded by various cameras is processed and edited into a finished video production, for example, to be broadcast to viewers of the event. Various cameras have conventionally been operated by humans in order to capture the event in an attractive way, for example, by dynamically pointing the camera at the area of interest of the event and zooming in. Such video production can consume relatively large labor and resources, and thus is often not economically feasible for medium-scale and small-scale events.

[0003] Automated video production systems can be used to reduce the overall video production cost. Known automated video production systems typically involve a fixed overview camera that records an overview video of the event, and an automated detection method is employed to automatically detect features of interest in the overview video and extract sub-frames from the overview video depicting the features of interest. Features within the area of interest can be automatically detected and tracked by dedicated detection and tracking algorithms. However, such conventional automated feature tracking systems typically struggle to evaluate context-dependent situations of an event that are likely to draw the interest of human viewers compared to human-operated systems. The quality of the video production resulting from such automated methods is often actually recognized as being inferior to that of conventionally human-edited video productions.

Summary of the Invention

Problems to be Solved by the Invention

[0004] One objective is to provide an improved method for automated video production for events. More generally, the objective is to overcome or improve upon some of the shortcomings of prior art, or at least to provide an alternative process that is more effective and can be used at a relatively lower cost than prior art. In any case, the present invention aims to at least provide a useful alternative and contribution to existing art. [Means for solving the problem]

[0005] According to a first embodiment, a computer implementation method for the autonomous production of an edited video stream of a subject area is provided. The edited video is produced from a video stream of the subject area, for example, still and / or dynamic video streams of the subject area. The video stream may be obtained, for example, from one or more cameras that depict only the subject area or a portion thereof. The subject area may be a real-world area, such as a scene of an event depicted by the video stream. The subject area may be, for example, a sports field, a podium, or a portion thereof.

[0006] This method involves obtaining a first frame from an edited video stream of a first sub-region of the subject area. The sub-region is a part of the subject area and may be smaller than the subject area. For example, the first frame of the edited video may be a crop from a video frame of the video stream of the subject area. Alternatively, the first frame may be obtained, for example, from an auto-controlled camera and therefore oriented and zoomed to depict the sub-region. Thus, the auto-controlled camera may be mechanically and / or optically adjusted to provide an adjustable field of view.

[0007] This method involves inputting a first frame into a machine learning model. The machine learning model may be pre-trained, for example, according to the training methods described herein.

[0008] This method involves using a machine learning model to determine navigation parameters related to a second subregion of the subject region based on a first frame. The second subregion may be the same as or different from the first subregion.

[0009] This method involves obtaining a second frame for the edited video stream of a second subregion based on determined navigation parameters. Similar to the first frame, the second frame for the edited video stream may also be a cutout from the overview video frame of the video stream of the subject region, e.g., a different cutout. Alternatively, the second frame may be obtained by making mechanically and / or optically adjustable physical adjustments to the camera, e.g., by changing the camera's field of view and / or optical zoom settings.

[0010] Therefore, this method enables the creation of an iterative edited video stream, where, for example, each frame of the edited video is taken from another frame, such as a preceding or following frame. For example, the second frame may be the next frame in the edited video stream, which can be determined from the first, current frame.

[0011] This method allows for the automatic production of edited video, for example, in real time, at a practically feasible computational cost, because, for example, the second frame can be determined based on a first frame, rather than on a static or non-static overview video stream of the subject area. The frames of the overview video stream are generally considerably larger than the frames for the edited video themselves and contain a lot of information for determining the appropriate sub-regions to focus on in the edited video. Also, in some situations, the overview video stream may not be available. However, the inventors have found that the frames for the edited video contain enough information to be used as input to a machine learning model for producing high-quality edited video. By using frames for the edited video, the computational load of this method can be kept very low. This can provide a low-latency system and method that enables the adjustment of a mechanically and / or optically adjustable camera in real time or near real time to acquire frames for the edited video.

[0012] Furthermore, machine learning models can take contextual information into account to infer the appropriate frame for edited footage, particularly compared to conventional automated feature tracking systems and methods.

[0013] The second frame for the edited video stream of the second subregion is determined based on navigation parameters, which are determined by a machine learning model. This method may optionally include determining navigation parameters that are adjusted based on the navigation parameters. The navigation parameters may be adjusted by considering secondary indicators, such as the past and / or future progression of the navigation parameters over time. This can provide a smoother edited video.

[0014] The first and second frames may be consecutive, but it will be understood that there may be other frames between them. The first frame may precede the second frame in time, but the first frame may follow the second frame in time. In one example, the first and second frames correspond to the same time.

[0015] A machine learning model can be pre-trained to determine navigation parameters that represent corrections to frame settings for an input video frame, based on the input video frame associated with frame settings such as tilt, pan, and / or zoom settings. Frame settings may include, in particular, one or more of tilt, pan, and zoom settings, which represent camera views of a real or virtual camera that capture the relevant video frame. The machine learning model can be trained accordingly to receive the input frame and infer the associated frame settings from it. The machine learning model may be configured to output corrections to the frame settings associated with the input frame that would make the input frame more appropriate for the scene depicted by the input frame.

[0016] Accordingly, the first embodiment may more specifically provide a computer implementation method for the autonomous production of an edited video stream of a subject region, which provides a pre-trained machine learning model configured to receive input video frames associated with frame settings such as pan settings, tilt settings, and / or zoom settings, the pre-trained machine learning model being pre-trained to determine navigation parameters representing corrections to frame settings for the input video frames, based on the input video frames; providing a first frame for an edited video stream of a first sub-region of the subject region; inputting the first frame into the pre-trained machine learning model; having the pre-trained machine learning model determine navigation parameters relating to a second sub-region of the subject region based on the first frame; and obtaining a second frame for an edited video stream of the second sub-region based on the determined navigation parameters. Accordingly, the first frame may be associated with a first frame setting, and the determined navigation parameters represent corrections to the first frame setting. The second frame setting may be obtained by applying corrections to the first frame setting. The second frame setting can be used to determine the input to the camera system in order to obtain the second frame. This process can be repeated iteratively to produce frames for the edited footage.

[0017] Optionally, the first and / or second frames of the edited video stream are obtained by taking subframes of frames from the video stream that represent at least a portion of the subject area. Subframes may be taken, for example, from frames of a static or non-static overview video stream of the subject area. Subframes may also be taken, for example, from frames of a non-static, mechanically and / or optically adjustable camera, such as a PTZ camera. The first and / or second frames may therefore be a virtual camera projection of the video stream, such as a cutout. Subframes may be obtained, for example, by cropping frames from the video stream and preferably adjusting the perspective of the crop, such as straightening it.

[0018] The edited footage may include a set of frames, such as a time series, including a first frame and a second frame. Each frame of the edited footage may be, for example, a subframe from each frame of the footage of the subject area, such as a cutout. The first frame may be, for example, a subframe of the first frame of the overview footage of the subject area, and the second frame may be, for example, a subframe of a second different frame of the overview footage of the subject area.

[0019] The first and second frames for the edited video stream can be obtained from the respective frames of the overview video stream of the subject area. The first frame for the edited video stream of the first sub-region can be obtained, for example, from the first frame of the overview video stream of the subject area. Similarly, the second frame for the edited video stream of the second sub-region can be obtained, for example, from the second frame of the overview video stream of the subject area.

[0020] Optionally, the first and / or second frames of the edited video stream are automatically captured by a mechanically and / or optically adjustable camera. In addition to, or alternatively to, taking subframes of still overview footage of the subject area, each frame of the edited video may be associated with the respective physical settings of the mechanically and / or optically adjustable camera, which may, for example, capture only a portion of the subject area.

[0021] Optionally, the first frame in the edited video stream of the first subregion is not equal to the first frame in the video stream of the subject region. The first frame in the edited video stream of the first subregion may be, for example, a smaller frame than the first frame in the video stream of the subject region, such as a cutout.

[0022] Optionally, the second frame in the edited video stream of the second subregion is not equal to the second frame in the video stream of the subject region. The second frame in the edited video stream of the second subregion may be, for example, a smaller frame than the second frame in the video stream of the subject region, such as a cutout.

[0023] Optionally, the first frame has a first frame setting associated therewith, the second frame has a second frame setting associated therewith, and the navigation parameter represents an adjustment from the first frame setting to the second frame setting. Thus, the navigation parameter can be, for example, a vector indicating the direction and speed or amount in which the line-of-sight direction and / or the field of view of the edited video stream is moved, starting from the first frame depicting a first sub-region and manipulating the camera virtually or physically in order to depict a second sub-region using the second frame. The first frame setting can include, for example, one or more of a first tilt setting, a first pan setting, and a first zoom setting of a real or virtual camera. Similarly, the second frame setting can include, for example, one or more of a second tilt setting, a second pan setting, and a second zoom setting of a virtual camera. The navigation parameter can thus provide adjustments, for example mappings, from the first pan setting to the second pan setting, from the first tilt setting to the second tilt setting, and from the first zoom setting to the second zoom setting.

[0024] Optionally, the navigation parameter includes one or more of a virtual pan adjustment, a virtual tilt adjustment, and a virtual zoom adjustment.

[0025] Optionally, the navigation parameter includes one or more of a mechanical pan adjustment, a mechanical tilt adjustment, and an optical zoom adjustment.

[0026] [[ID=第十二]]Optionally, the navigation parameter includes an adjustment speed and an adjustment direction. Including the adjustment speed and the adjustment direction in the navigation parameter can result in a smoother transition between frames of the edited video, for example, compared to navigation parameters that include an absolute adjustment amount or a relative adjustment ratio.

[0027] Optionally, the navigation parameters include one or more of pan adjustment, tilt adjustment, and zoom adjustment. In particular, the navigation parameters include one or more of relative pan adjustment, relative tilt adjustment, and relative zoom adjustment with respect to the first frame in order to adjust the first frame setting to match the second frame setting.

[0028] Optionally, one or more of the pan adjustment, tilt adjustment, and zoom adjustment are virtual adjustments. This adjustment is virtual in that it does not require physical movement of the camera, but instead may involve the selection of a different cutout from the overview of the subject area, i.e., it appears as if the pan, tilt, and / or zoom settings of the camera are (virtually) adjusted. Thus, the navigation parameters may include one or more of relative virtual pan adjustment, relative virtual tilt adjustment, and relative virtual zoom adjustment with respect to the first frame. Therefore, the first and second frames of the edited video may each be cutouts from respective frames of the video stream of the subject area.

[0029] [[ID=,8]]Optionally, one or more of the pan adjustment, tilt adjustment, and zoom adjustment are physical adjustments. This adjustment is physical in that it involves physical movement of the camera, such as mechanical pivoting of the camera about a base and / or mechanical movement of one or more lenses to change the optical zoom.

[0030] Optionally, the first frame is acquired at a first time, the navigation parameters are determined at the first time, and are related to a second time after the first time.

[0031] <, Optionally, this method involves, at a first time point, using a machine learning model to determine a time series of navigation parameters related to multiple time points after the first time point, based on the first frame, where the time series of navigation parameters specifically includes navigation parameters. Thus, the machine learning model can determine multiple navigation parameters for multiple future time points at the current time point. The trajectories of future frame settings, particularly tilt, pan, and / or zoom settings, can therefore be determined based on the current frame settings and / or past frame settings.

[0032] Optionally, a machine learning model is configured to determine a time series of navigation parameters that conform to a predefined progression characteristic. The progression characteristic may be used to provide smooth edited footage with minimal abrupt camera movements. The progression characteristic may, accordingly, constrain the progression of the navigation parameters in the time series. For example, the progression characteristic may impose a smooth acceleration and departure of the view from the current frame by gradually increasing the difference between consecutive navigation parameters in the time series from the current frame. The progression characteristic may further impose a smooth deceleration of the view towards the last frame setting in the time series by gradually decreasing the difference between consecutive navigation parameters in the time series towards the last navigation parameter in the time series.

[0033] Optionally, this method includes determining adjusted navigation parameters based on navigation parameters and at least one further navigation parameter related to a second time determined at a time preceding the first time. Thus, a smooth edited video can be obtained in which abrupt changes in the line of sight are effectively smoothed out.

[0034] Optionally, a second frame is retrieved for the edited video stream of the second subregion based on the adjusted navigation parameters.

[0035] Optionally, the adjusted navigation parameters are determined as a weighted average of multiple navigation parameters determined at time points preceding each second time point, each relating to a second time point. In particular, navigation parameters among those that are temporally closer to the second time point are given greater weight than navigation parameters among those that are temporally further away from the second time point. Navigation parameters predicted in the near past relating to future time points are generally more accurate in practice than navigation parameters predicted in the far past relating to future time points, and therefore can be given higher priority in determining the adjusted navigation parameters.

[0036] Optionally, the first and second time points are separated by a time period determined to account for method latency, particularly the time delay from the capture of the first frame to the finalization of the physical adjustment of the camera associated with the navigation parameters. In particular, with regard to controlling mechanically and / or optically adjustable cameras, time delays, e.g., system and method latency, can be substantial. Latency may result from one or more of several method steps, e.g., image frame acquisition time, data transmission time, processing time, and camera adjustment time. Therefore, latency can be predicted to optimize system performance. Latency can be determined by measuring or estimating the time delay between method steps. The navigation parameters determined at the first time point, e.g., the current time point, may be for camera settings related to a second time point in the future, separated from the first time point by a time period corresponding to the latency. It will be understood that the time delay can span multiple image frames, i.e., latency may be greater than the camera sampling time interval. Therefore, the navigation parameters may relate to a time instance, which is the number of future camera sampling time instances. Therefore, the time instance may be between the first time and the second time. In particular, when video frames for edited video are captured using a mechanically and / or optically adjustable camera, such as a PTZ camera, the first and second time may be temporally separated by a period of time at least substantially corresponding to the time delay from the capture of the first frame to the finalization of the physical adjustment of the camera associated with the navigation parameters.

[0037] Optionally, navigation parameters are determined as normalized navigation parameters, normalized with respect to the field of view of the first frame. For example, one or more of the pan, tilt, and zoom adjustments are determined in a normalized form, for example, with respect to the field of view of the first frame. In this way, for example, a machine learning model can be effectively trained to generate navigation parameters, particularly because it allows for the determination of relative pan and tilt adjustments of the second frame regardless of the field of view of the first frame, for example, the zoom setting. Thus, the result that certain pan and tilt adjustments may have different effects on different zoom settings of the first frame can be efficiently taken into account. The field of view of the first frame can be considered as the extent of a first subregion depicted by the first frame. The field of view can be defined as the dimensions of the first frame, such as the height and / or width of the first frame, and / or the field of view angle associated with the first frame. The field of view can be defined by the camera's optical or virtual zoom setting.

[0038] Optionally, this method includes denormalizing normalized navigation parameters, for example, based on the field of view of a first frame, and obtaining a second frame for an edited video stream of a second subregion based on the denormalized navigation parameters. For example, this method may include denormalizing one or more of normalized pan adjustments, normalized tilt adjustments, and normalized zoom adjustments, and obtaining a second frame for an edited video stream of a second subregion based on the denormalized pan, tilt, and zoom adjustments.

[0039] Optionally, the video stream is recorded from a stationary viewpoint by one or more cameras. Recordings from multiple cameras can be, for example, stitched together to form a video stream.

[0040] Optionally, this method includes calibrating one or more still or non-still cameras. Calibration may include, in particular, determining the mapping between the coordinate systems of one or more cameras and the coordinate system of the subject area. Calibration may include determining the orientation of one or more still cameras with respect to the subject area, such as determining the orientation of the pan axis and / or tilt axis with respect to the horizontal plane and / or the vertical and horizontal spacing between the subject area and one or more cameras. Calibration may further include determining a correction mapping to compensate for lens distortion of one or more cameras. Calibration may include correlating the image frames from each camera with respect to each other in order to stitch the image frames together to form a single image stream.

[0041] Optionally, this method includes obtaining an expanded first frame of an expanded first sub-region of the subject area that is wider than the first sub-region by a margin, for example, a predetermined margin, and determining navigation parameters based on the expanded first frame. Thus, the expanded first frame may contain additional information compared to the first frame, which can be used to improve the determination of navigation parameters and, consequently, the second frame. The expanded first sub-region may be smaller than the subject area to manage the computational load. The expanded first frame may be, for example, a cutout from the frame of the image of the subject area. The expanded first frame may be obtained, for example, by zooming out from the first frame by, for example, a predetermined amount.

[0042] Optionally, the margins are asymmetrical with respect to the first sub-region. Certain areas of the subject region may not generally provide highly useful additional information, while other areas generally do. For example, areas above the first frame may generally provide more useful information than areas below the first frame. Therefore, the margins may be configured asymmetrically around the first sub-region, for example, having a relatively wide margin above the first frame and a relatively narrow margin below the first frame. Asymmetrical margins can be obtained, for example, by zooming out from the first frame and, in addition, by tilting and / or panning the view.

[0043] Optionally, the margins depend on the subject area.

[0044] Optionally, the margins depend on navigation parameters, such as the zoom value, and / or on the first frame relative to the edited footage, such as the size of the first frame relative to the frame of the subject area video stream. The margins may be, for example, relatively small when the first frame is a relatively large cutout from the frame of the subject area video stream, e.g., when zoomed out, and relatively large when the first frame is a relatively small cutout from the frame of the subject area video stream, e.g., when zoomed in.

[0045] Optionally, this method includes labeling the data points of the extended first frame to distinguish between data points of the extended first frame corresponding to data points of the first frame and data points of the extended first frame corresponding to data points of the margin. Thus, appropriate navigation parameters can be obtained, where it can be automatically explained which parts of the extended first frame represent the first frame and which parts do not. Each extended frame may include a label channel dedicated to labeling data points within the margin and / or data points outside the margin. For example, each frame may include one or more color channels, such as a red channel, a green channel, and a blue channel, and even additional label channels. Labeling channels can be used to automatically distinguish between parts of the extended frame that form the margin and are therefore not visible to the viewer of the edited video stream and parts of the extended frame that form the frame of the edited video and should be visible to the viewer of the edited video stream.

[0046] Optionally, this method includes, for example, resampling the augmented first frame and determining navigation parameters based on the resampled, for example, augmented first frame. The augmented first frame may be resampled to, for example, a predetermined resolution. The augmented first frame with a relatively high resolution may be downsampled to, for example, improve computational efficiency. The augmented first frame with a relatively low resolution may be upsampled.

[0047] Optionally, this resampling is such that non-uniform sampling across the entire frame is used to obtain the output frame. Regions of the first frame that generally contain useful information, such as the central region, may be given a higher resolution than regions that contain less useful information, such as the edge regions of the first frame.

[0048] Optionally, navigation parameters are determined based on a set of frames in an edited video stream that includes only the first frame. It has been found that a high-quality edited video stream can only be obtained by using the first frame to determine the navigation parameters and then using the second frame. This provides a particularly computationally efficient method.

[0049] Optionally, navigation parameters are determined based on a set of frames from an edited video stream, including at least two frames. Therefore, in addition to the first frame, additional frames may be used to determine the navigation parameters, and then the second frame. Thus, the data on which navigation parameters are determined can be improved by considering multiple frames, such as preceding and / or succeeding frames. It will be understood that the preceding and succeeding frames do not need to be temporally consecutive, and intermediate frames are possible.

[0050] Optionally, the set of frames in the edited video stream includes frames that temporally precede the second frame. Optionally, the set of frames in the edited video stream includes frames that temporally precede the first frame.

[0051] Optionally, the set of frames in the edited video stream may include frames that temporally follow the second frame. Optionally, the set of frames in the edited video stream may include frames that temporally follow the first frame. The edited video stream may be buffered, for example, to allow the use of subsequent frames when determining navigation parameters. The overview video of the subject area may also be buffered to the same extent. It will be understood that the set of frames in the edited video stream may include frames that temporally precede and / or follow and / or coincide with the first frame. Thus, navigation parameters, and then the second frame, may be determined based on the first frame and in addition to one or more other frames.

[0052] Optionally, the first frame is acquired at a first time, the navigation parameters are determined at the first time and are related to a second time that follows the first time, and this method involves using a machine learning model at the first time to determine a time series of navigation parameters related to multiple time points after the first time, based on the first frame, the time series of navigation parameters specifically includes the navigation parameters, and the time series of navigation parameters is determined based on a set of frames of an edited video stream that includes subsequent frames that temporally follow the first frame, and optionally, the time series of navigation parameters is determined based on a set of frames of an edited video stream that includes preceding frames that are temporally before the first frame. In particular, the time series of navigation parameters may include navigation parameters related to future time points with respect to the current time. Thus, the edited video stream may be buffered so that subsequent frames of the edited video are considered in determining the time series of navigation parameters, and the predicted range of navigation parameters may extend into the future beyond the current time. It will be understood that as time progresses from one point in time to the next, and a new time series for the navigation parameters is determined, subsequent and / or preceding frames in the buffer may be modified. Therefore, frames within the buffer time window can be considered estimated frames. Those frames of the edited video stream that fall out of the buffer as time progresses can be part of the actual edited video that is broadcast directly to the viewer, for example.

[0053] Optionally, this method includes determining the navigation parameters, adjusting the navigation parameters according to a smoothing criterion, and obtaining a second frame for the edited video stream based on the adjusted navigation parameters. Post-processing steps may be provided, for example, in which the determined navigation parameters are compared to preceding and / or succeeding navigation parameters and adjusted accordingly. For example, some navigation parameters may be smoothed. Alternatively, the navigation parameters may be adjusted according to consecutive frames and / or navigation parameters to provide smooth video.

[0054] Optionally, the determined navigation parameters are adjusted based on a set of navigation parameters that include navigation parameters that temporally precede and / or temporally follow the determined navigation parameters. The determined navigation parameters may be adjusted based on a set of navigation parameters that include temporally following navigation parameters, in particular, to allow for the prediction of large changes in navigation parameters. If a large subsequent navigation parameter is observed for a subsequent frame, the navigation parameters may be adjusted in anticipation of this, toward the large subsequent navigation parameter. Thus, large changes that would have been caused by a large subsequent navigation parameter are mitigated by adjusting the navigation parameters accordingly, and a smooth edited image is created.

[0055] Optionally, this method includes determining a navigation trend based on the aforementioned set of navigation parameters and adjusting the navigation parameters according to the determined navigation trend.

[0056] Optionally, the navigation parameters are determined by a trained machine learning model, particularly an end-to-end artificial neural network. The machine learning model may be, for example, a convolutional neural network or a transformer deep learning architecture. The machine learning model can be end-to-end in that it directly determines the navigation parameters based on the first frame without requiring additional computational steps.

[0057] According to a second aspect, a computer implementation method is provided for the autonomous production of an edited video stream from a video stream of a subject region, which includes inputting a first frame of edited video of a first sub-region of the subject region into a model, in particular a machine learning model; causing the model to determine navigation parameters relating to a second sub-region of the subject region based on the input first frame; and obtaining a second frame for the edited video stream of the second sub-region. This method may, in particular, follow any system or method described herein.

[0058] The term "model" as used herein has a broad meaning, and it will be understood that functions, equations, algorithms, correspondences, mappings, and similar things can be considered models.

[0059] According to a third aspect, a machine learning model is provided for the autonomous production of an edited video stream of a subject region for use in the manner of the first and / or second aspects. The machine learning model may include a convolutional neural network or a transformer architecture. The machine learning model may be arranged and configured to take frames of the edited video, e.g., images, e.g., a first frame, as input and generate navigation parameters as output. The navigation parameters may include, for example, pan, tilt, and zoom values.

[0060] According to a fourth aspect, a computer implementation method is provided for generating datasets, in particular training datasets for machine learning models, such as those in the third aspect. This method includes providing an image of a subject region, having a human or a trained auxiliary machine learning algorithm navigate the subject region in the image, obtaining a curated dataset of navigation parameters from there, and / or obtaining a curated dataset of frames from there. The curated data may be human-curated or machine-curated. Navigation of the subject region in the image may involve virtual zoom, virtual pan, and / or virtual tilt within the image of the subject region to obtain a curated edited image represented by the curated dataset of frames. Each frame in the curated dataset of frames may be associated with a navigation parameter in the curated dataset of navigation parameters. The curated dataset of navigation parameters may be obtained from commands of a human-controlled input device, such as a joystick or other similar device, which is operated by a human to adjust pan, tilt, and / or zoom settings. Optionally, this method involves selecting a subset of frames from the video and having a human navigate through the subject area within the selected subset of video frames. The human only needs to navigate through the keyframes of the video, for example, for efficient annotation of the data. The curated dataset of navigation parameters may optionally include interpolated navigation parameters associated with the curated video frames between the video frames annotated by the human.

[0061] Optionally, this method includes having a trained auxiliary machine learning model detect the presence of a given entity of interest within one or more frames of video, having the trained auxiliary machine learning model determine the placement of the detected entity of interest within one or more frames of video, and obtaining a curated dataset of navigation parameters from there. For some applications, an auxiliary machine learning model may be more accurate than a human operator in detecting and tracking a given entity of interest, such as a person or object, within video. Therefore, it may be desirable to have an auxiliary machine learning model generate the training data instead of a human. Automated object detection and tracking by an auxiliary machine learning model can be used in itself to generate training data for the machine learning model, but it will be understood that the final production of edited video by the machine learning model may not involve entity detection and tracking as such. Entity detection and tracking using an auxiliary machine learning model can indeed be error-prone, but this can be mitigated during the training phase of the machine learning model. An auxiliary machine learning model, for example, configured to detect and track a player on a pitch, may mistakenly deviate when it also detects other people, such as a sign displaying a person, or a spectator, such as another player. In particular, machine learning models for producing edited video are preferably not dedicated feature tracking systems, because, above all, they are too error-prone for live streaming of events. Using a feature tracking system during the training phase allows for tuning the training data to eliminate errors, but this is generally undesirable or impractical during the editing phase. Instead of mere object detection and tracking, the machine learning model is further trained to exhibit desirable frame transition behavior, preferably similar to or exceeding the performance of a human camera operator.

[0062] Optionally, the method described in the claims includes augmenting a curated dataset of navigation parameters with an augmentation dataset, generating a dataset of frames corresponding to the augmented dataset of navigation parameters, and labeling the dataset of frames with the augmentation dataset. The augmentation dataset includes, in particular, frame-setting augments such as tilt augmentation, pan augmentation, and zoom augmentation. Thus, the augmentation dataset offsets navigation parameters such as those curated by a human or auxiliary machine learning model. By augmenting the curated dataset of navigation parameters in this way, the machine learning model is trained to relate input frames to corrections applied to the frame settings of the input frames. Thus, the machine learning model can be trained on the frames themselves to correct the frame settings of the input frames. The estimated corrections learned to be generated by the machine learning model can then be used to obtain a second frame.

[0063] The curated dataset of navigation parameters can be augmented, for example, by randomly generated augments. The augmentation of the dataset may follow a normal distribution, for example.

[0064] Optionally, this method includes generating a dataset of frames from video to an edited video, either according to the method described herein or otherwise, obtaining a dataset of differences representing the difference between the generated dataset of frames and the curated dataset of frames, and labeling the dataset of frames with the dataset of differences. The dataset of frames from video to an edited video may, in particular, be generated by a pre-trained machine learning model such as the one described herein. This results in particularly efficient training.

[0065] According to the fifth aspect, a computer implementation method is provided for training a machine learning model for the autonomous production of an edited video stream from a video stream of a subject region using a dataset obtained by the method according to the fourth aspect.

[0066] Optionally, the machine learning model can be configured to output a time series of navigation parameters.

[0067] Optionally, the machine learning model is configured to output a time series of navigation parameters that conform to a predefined progression characteristic. The progression characteristic can be used to train the machine learning model to provide smooth edited footage with minimal abrupt camera movements. The progression characteristic may constrain the navigation parameters accordingly, and the constraints are particularly time-dependent. The progression characteristic, therefore, forces the machine learning model to learn smooth transitions of navigation parameters over time, as opposed to any sequence of navigation parameters that could lead to jerky edited footage. Thus, the machine learning model can be trained to learn how to smoothly correct the frame settings of a real or virtual camera. The progression characteristic can, accordingly, be used as a tool to control the frame setting adjustment characteristics during the training phase of the machine learning model, which in turn allows control over the smoothing behavior of the machine learning model. After training the machine learning model, it will be understood that no post-processing is required, or only minimal post-processing, is needed to smooth the edited footage, as the machine learning model can be trained to provide smooth edited footage. The progression characteristic may, for example, impose a smooth acceleration and departure of the view of the current frame by gradually increasing the difference between consecutive navigation parameters in the time series from the current frame. The progression characteristic may further impose a smooth deceleration of the view towards the last frame setting in the time series by gradually decreasing the difference between consecutive navigation parameters in the time series towards the last navigation parameter in the time series.

[0068] Optionally, this method includes generating a first training dataset and pre-training a machine learning model using the first dataset. This method further includes causing the pre-trained machine learning model to generate a second training dataset and training the pre-trained machine learning model using the second dataset.

[0069] The first dataset is generated by various methods, particularly by providing video of the subject region, having a human navigate the subject region within the video, obtaining a curated dataset of navigation parameters from it, extending the curated dataset of navigation parameters with an extension dataset, generating a dataset of frames corresponding to the extended dataset of navigation parameters, and labeling the dataset of frames with the extension dataset.

[0070] The second dataset can be generated by a machine learning model trained on the first dataset. This can be achieved by having the pre-trained model generate a dataset of frames from the subject region footage to the edited footage, obtaining a dataset of differences representing the difference between the generated dataset of frames and the curated dataset of frames, and labeling the dataset of frames with the dataset of differences. This procedure can be repeated, allowing the pre-trained model to generate further generations of datasets or its own training. Datasets from different generations can optionally be merged to minimize bias in the training data.

[0071] Optionally, this method includes training a machine learning model using a second dataset, then having the machine learning model generate a third dataset, and training the machine learning model using at least a portion of the third dataset. The third dataset may be generated by a machine learning model trained with the first and second datasets, and this includes having the pre-trained machine learning model generate a dataset of frames from the video to an edited video, obtaining a dataset of differences representing the difference between the generated dataset of frames and the curated dataset of frames, and labeling the dataset of frames with the dataset of differences.

[0072] According to a sixth aspect, edited video obtained by a method as described herein is provided.

[0073] According to a sixth aspect, a non-temporary computer-readable medium is provided for storing instructions that, when executed by one or more processors, cause a device to perform a method as described herein.

[0074] According to a seventh aspect, a system is provided for the autonomous production of an edited video stream of a subject area. The system comprises a camera system and a processing unit operably connected to the camera, which is configured to perform methods such as those described herein. The camera system may include a fixed overview camera configured to record an overview video stream of a subject area from a stationary line of sight. In addition, or alternatively, the camera system may include a mechanically and / or optically adjustable camera, such as a pan-tilt-zoom or PTZ camera, which is configured to record an overview video stream of a subject area from an adjustable line of sight. The system may also include a camera controller operably configured between the processing unit and the mechanically and / or optically adjustable camera, which transmits one or more video frames acquired by the camera to the processing unit, receives navigation parameters from the processing unit, and is configured to mechanically and / or optically adjust the camera based thereon.

[0075] For example, this embodiment may provide a system for the autonomous production of an edited video stream of a subject area, comprising a camera system comprising, for example, a fixed overview camera for acquiring overview video frames of the subject, and a mechanically and / or optically adjustable camera for acquiring video frames of the edited video, and a processing unit operably connected to the overview camera and the mechanically and / or optically adjustable camera, configured to perform methods as described herein.

[0076] Optionally, the processing unit is configured to acquire a first frame of a first sub-region of the subject area, the first frame being acquired from the overview camera, the first frame being input to a machine learning model, the machine learning model determining navigation parameters related to a second sub-region of the subject area based on the first frame, and transmitting control signals representing the navigation parameters to the mechanically and / or optically tunable camera to adjust the mechanically and / or optically tunable camera. The overview camera and the mechanically and / or optically tunable camera may be operated independently, and video frames from the overview camera, e.g., cutouts therefrom, may be used as input to the machine learning model for determining the navigation parameters, while frames from the mechanically and / or optically tunable camera may be used for edited video. The mechanically and / or optically tunable camera may therefore be controlled based on the navigation parameters output by the machine learning model.

[0077] It will be understood that any combination of the embodiments, features, and options described herein may be used.

[0078] Embodiments of the present invention will be described in detail with reference to the accompanying drawings. [Brief explanation of the drawing]

[0079] [Figure 1] This figure shows a schematic example of a system and method for producing edited video streams. [Figure 2] This figure shows a schematic example of a system and method for producing edited video streams. [Figure 3] This figure shows a schematic example of a system and method for producing edited video streams. [Figure 4A] This figure shows a schematic example of a method for creating an edited video stream. [Figure 4B]This figure shows a schematic example of a method for creating an edited video stream. [Modes for carrying out the invention]

[0080] Figures 1 to 3 show schematic examples of a system and method 100 for producing edited video streams of a subject area. System 100 here comprises a video production unit 10 positioned and configured to receive video streams of the subject area acquired from the camera system 20 of system 100. In the example of Figure 1, the camera system comprises two fixed cameras 21 and 22. In this example, the two cameras 21 and 22 each transmit their respective video streams 1a and 1b to the video production module 10. In this example, cameras 21 and 22 are actually stationary and directed to acquire their respective video streams 1a and 1b from fixed, stationary viewpoints of at least a portion of the subject area, such as an event site, for example, a sports pitch or a music concert. The respective video streams 1a and 1b can be combined to render an overview video stream of the subject area. The video production unit 10 receives the video streams 1a and 1b and, based on them, produces an edited video stream 2 for broadcast, for example, to an audience of the event.

[0081] In alternative configurations, or in addition to the fixed cameras 21, 22, the camera system 20 may include a mechanically and / or optically adjustable camera 23, such as a pan-tilt-zoom or PTZ camera, which is configured to provide an adjustable view of the subject area. A mechanically and / or optically adjustable camera may, for example, have a limited field of view and may not capture the entire subject area in a single frame. Figure 2 shows an example of system 100 having a mechanically and optically adjustable camera 23 in addition to the fixed overview cameras 21, 22. Figure 3 shows an example of system 100 having only the mechanically and optically adjustable camera 23, excluding the fixed overview camera.

[0082] The video production unit 10 in the example shown in Figures 1 and 2 includes a processing module 30. The processing module 30 is configured to receive and process video streams 1a and 1b received from cameras 21 and 22 in order to obtain video streams of the subject area. The two video streams 1a and 1b may, for example, be appropriately merged or spliced ​​together to form a single stream of video frames of the subject area. The video production unit 10 is configured to produce an edited video stream 2 from video stream 1, particularly in real time. The edited video stream 2 may, for example, be a time series of frames streamed directly to the audience of an event, or it may be stored in memory as edited video for later viewing or further editing. The edited video may, for example, be used for automated event detection, which may then be used to automatically compile a summary video of highlights from the edited video.

[0083] Frames for the edited video stream 2 are determined using a generator module 40 of the video production unit 10. The generator module 40 is configured to generate navigation parameters 5.j based on an input video frame, for example, an image. For example, including a machine learning model such as a convolutional neural network or a transformer, the generator module 40 is configured to have a first video frame 2.1 for the edited video depicting a first subregion of the subject area as input for transmission to the processing module 30, and to generate navigation parameters 5.2 based on it. The processing module 30 receives the navigation parameters 5.2 generated by the generator and then provides a second frame 2.2 for the edited video depicting a second subregion of the subject area associated with the navigation parameters 5.2. The second frame 2.2 for the edited video can then be input to the generator module 40 to determine further navigation parameters 5.3 and further frames 2.3 for the edited video. This process can be repeated. Thus, this method can be an iterative method in which frames 2.i for the edited video stream 2 are determined iteratively.

[0084] Navigation parameter 5.j may indicate the coordinates of the subject area, for example, to obtain a cutout from a video frame of the subject area's video stream. In this example, navigation parameter 5.2 indicates a relative adjustment, such as a vector, from a first frame depicting a first sub-region of the subject area to a second frame depicting a second sub-region of the subject area, indicating the direction and amount or speed at which the view of the edited video stream should change, starting from the first frame and reaching the second frame. Navigation parameters may, in particular, indicate one or more of relative pan adjustments, relative tilt adjustments, and relative zoom adjustments.

[0085] In the example in Figure 1, the relative adjustments are relative, virtual adjustments with respect to the first frame and the overview video streams acquired from fixed cameras 21, 22, in particular, relative virtual tilt values, relative virtual pan values, and relative virtual zoom values. It will be understood that tilt may correspond to vertical adjustment and pan to horizontal adjustment, or vice versa. In practice, the tilt and pan axes may be inclined with respect to the horizontal plane of the real world. Calibration may be employed to provide a mapping between the camera's tilt and pan and the real-world tilt and pan. Zoom may change the field of view, e.g., viewing angle. It will be understood that navigation parameters may, alternatively, be relative to the frames of the video stream of the subject area.

[0086] The processing module 30 may, for example, extract subframes from the frames of the video stream 1 of the subject area and output these subframes as frames 2.i of the edited video stream 2. The subframes may, for example, only show a portion of the subject area, particularly a portion of the subject area of ​​most interest to the viewer. The subframes may be obtained, for example, by cropping the frame and may optionally be adjusted to accommodate changes in perspective.

[0087] The first frame of the edited video may be a subframe from the first frame of the still overview video stream of the subject area, and the second frame of the edited video may be a subframe from the second frame of the still overview video stream of the subject area. The first and second frames of the video stream of the subject area do not need to be temporally consecutive, but it will be understood that other frames may be interspersed between them. Similarly, the first and second frames of the edited video stream do not need to be temporally consecutive, but it will be understood that other frames may be interspersed between them.

[0088] It will also be understood that navigation parameters do not necessarily have to represent virtual adjustments. Instead, as shown in the examples in Figures 2 and 3, the camera may be mechanically tilted, panned, and / or optically zoomed in or out. It will also be understood that, instead of taking subframes of the overview image to generate image frames for the edited image, frames for the edited image stream may alternatively be obtained by physical adjustments of the camera, such as optical zoom, mechanical tilt, and mechanical pan adjustments of the camera. Such a system may provide higher image quality compared to taking subframes of the overview image frames, but at the cost of increased delay due to physical adjustments of the camera compared to virtual cropping of the overview image. Image frames acquired by a physically movable camera may also be cropped to some extent, but frames for edited image acquired from a physically movable camera may similarly be considered cutouts.

[0089] In the exemplary system shown in Figure 2, video frames acquired from overview cameras 21 and 22, particularly cutouts from them, are used to determine navigation parameters 5.j, and video frames acquired from mechanically and optically adjustable camera 23 are used to generate an edited video stream 2. The mechanically and optically adjustable camera 23 is therefore controlled, i.e., mechanically and / or optically adjusted, based on the determined navigation parameters 5.j. Here, the navigation parameters 5.j are converted into appropriate control signals 6.j for controlling camera 23 by module 41, which in this example may actually be integrated with generator module 40. Camera 23 in this example transmits a feedback signal 7.j to mapping module 41, which indicates its settings, e.g., current pan, tilt, and zoom settings.

[0090] In the example in Figure 3, system 100 does not have a fixed overview camera. In this example, system 100 includes only a mechanically and optically adjustable camera 23, and the acquired video frames are used directly as input to the generator module 40. Based on the video frames from the mechanically and / or optically adjustable camera 23, the generator module 40 generates navigation parameters 5.j, which are transmitted to the camera 23 to mechanically and / or optically adjust the camera 23 accordingly. Here, the navigation parameters 5.j are used directly as control signals to the camera 23. In this example, the video stream 2 acquired by the camera 23 directly represents the edited video stream 2. Frames of the edited video stream 2 are transmitted from the camera 23 to the generator module 40. In this example, the camera 23 also transmits state information to the generator module 40, including, in particular, the zoom setting used when normalizing the navigation parameters.

[0091] Navigation parameter 5.j can be generated by the generator module 40 in a normalized form, for example, with respect to frame 2.i input to the generator module 40. For example, a tilt value of 0.5 may indicate tilting upward by an amount corresponding to half the height of the first frame, and a tilt value of -0.5 may indicate tilting downward by an amount corresponding to half the height of the first frame. Zoom values ​​can optionally be generated on a logarithmic scale. For example, a zoom value of 1 may indicate a 2x zoom in, and a zoom value of -1 may indicate a 1 / 2 zoom out.

[0092] The navigation parameter 5.j can be denormalized, for example, by the processing module 30, so that adjustments relative to the virtual or real camera 23 are obtained to appropriately depict a desired sub-region of the subject area.

[0093] Figures 4A and 4B illustrate schematic examples of the method, particularly in conjunction with an exemplary system as shown in Figure 1. Figure 4A shows a schematic example of the first frame 1.1 of the video stream of the subject area. Here, the first frame is acquired from the video stream 1.1 of fixed overview cameras 21, 22 and provides an overview of the subject area, e.g., a real-world scene. It will be understood that the first frame may alternatively be acquired from a mechanically and optically movable camera 23, such as by a system as shown in Figures 2 and 3. The first frame 2.1 for the edited video stream is acquired from the first frame 1.1 of the video stream. Here, the first frame 2.1 for the edited video stream is taken as a subframe of the first frame 1.1 of the video stream, e.g., a perspective-adjusted cutout. Thus, the first frame 2.1 for the edited video is associated in this example with specific coordinates within the first frame 1.1 of the video stream, e.g., virtual pan, virtual tilt, and virtual zoom settings. The first frame 2.1 of the edited footage may, alternatively, be associated with mechanical pan, mechanical tilt, and optical zoom settings associated with a mechanically and optically movable camera 23. The first frame may also be associated with navigation parameters 5.1 generated by the generator module 40.

[0094] In the example in Figure 4A, the first frame 2.1 of the edited video stream is expanded by a margin of 3.1, thereby obtaining the expanded first frame 2.1' of the edited video. The expanded first frame 2.1' is larger than the first frame 2.1 by a margin of 3.1. Here, the margin 3.1 is asymmetric with respect to the first frame 2.1 of the edited video. In particular, here, the margin 3.1 is wider below the first frame 2.1 than above it. Therefore, with respect to the first frame 2.1, additional visual information can be assumed to reside in the expanded first frame 2.1', which can be used to generate appropriate navigation parameters. Here, more information is contained below the first frame 2.1 than above it by the margin 3.1, which can be more data-efficient for certain applications, as it is generally expected that most useful visual information will be below the first frame 2.1. The margin 3.1 is preferably such that the extended first frame 2.1' is smaller than the first frame 1.1 of the image of the subject area.

[0095] Labeling is assigned to the data points of the extended first frame 2.1', thereby enabling an automated distinction between the data points of the extended first frame 2.1' that form the margin 3.1 and the data points that form the first frame 2.1 for the edited video intended to be seen by the viewer. The extended first frame 2.1' is input to the generation module 40, which outputs the navigation parameter 5.2. In this example, the navigation parameter 5.2 indicates relative adjustments from the first frame 2.1 of the edited video stream to obtain the second frame 2.2 for the edited video stream. Here, the navigation parameter 5.2 indicates relative virtual zoom-out adjustments, such as a relative virtual tilt adjustment upwards, a relative virtual pan adjustment to the right, and a ratio. However, it will be understood that the navigation parameter may also indicate global coordinates, for example, with respect to the first or second frame 1.1, 1.2 of the video stream of the subject area.

[0096] The second frame 2.2 for the edited video stream may be obtained from the second frame 1.2 of the video stream of the subject area, for example, by the processing module 30, as if virtually adjusting the camera settings according to the determined navigation parameters 5.2, which is schematically shown in Figure 4B. For comparison, the first frame 2.1 is shown as a dashed line in Figure 4B.

[0097] The second frame 2.2 for the edited video, like the first frame 2.1, may be a cutout from the second frame 1.2 of the video stream of the subject region. The second frame 2.2 for the edited video may then have a margin 3.2 added to obtain an extended second frame 2.2', which may be used to determine further navigation parameters, and a third frame for the edited video stream, etc. The extended second frame 2.2' may be downsampled to a predetermined resolution, for example, before being input to the generator module 40.

[0098] Frames in an edited video stream can be determined based on preceding frames in the edited video stream. Therefore, the next frame in an edited video stream can be determined based on the current frame and / or one or more previous frames of the edited video. Frames in an edited video stream can also be determined based on subsequent frames in the edited video stream, for example, by buffering frames in the edited video stream. Therefore, the next frame in an edited video stream can be determined based on one or more future frames of the buffered edited video. Frames in an edited video stream can also be determined based on a combination of preceding and following frames in the edited video stream.

[0099] In these examples, the generator module 40 includes a machine learning model, such as a deep learning end-to-end convolutional neural network, and / or transformer architecture, which is configured to take frames, e.g., digital images, from an edited video stream as input and output navigation parameters based on them. However, other machine learning models may be used, such as support vector machines, decision tree-based learning systems, random forests, regression models, autoencoder clustering, and nearest neighbor machine learning algorithms. In some examples, an alternative regression model may be used instead of an artificial neural network.

[0100] Deep learning in a neural network environment can involve a large number of interconnected nodes called neurons. Input neurons activated from an external source activate other neurons based on their connections with other neurons governed by neural network parameters. The neural network can exhibit a specific pattern of behavior based on each parameter. By training a deep learning model, the model parameters representing the connections between neurons in the network are refined so that the neural network exhibits desired behavior (better behavior in the task for which it is intended, for example, classifying components in a material stream).

[0101] Deep learning operates on the understanding that many datasets contain a hierarchy of features—from low-level features (e.g., edges) to high-level features (e.g., patterns, objects). For example, while examining an image, the model begins by looking for edges that form motifs that form parts of the object being explored. Learned observable features include objects and quantifiable regularities learned by the machine learning model. Given a large set of well-classified data, a machine learning model is well-equipped to distinguish and extract features relevant to the successful classification of new data.

[0102] Optionally, the machine learning model utilizes a vision transformer architecture (ViT). The vision transformer divides the input image into a series of patches, serializes each patch into a vector, and maps these vectors to smaller dimensions, for example, by matrix multiplication. These vector embeddings can then be processed by a transformer encoder.

[0103] Optionally, machine learning models utilize convolutional neural networks. In some examples, deep learning can leverage neural network segmentation to find and identify learned observable features within data. Each filter or layer in a convolutional neural network architecture can transform the input data to improve the model's selectivity and robustness (of features) to the data. This abstraction of the data allows the machine to focus on the features in the data it is trying to classify and ignore irrelevant background information. Deep learning machine models using convolutional neural networks can be used for image analysis.

[0104] The machine learning model is trained on a first dataset. In this example, the first dataset is obtained from a human-curated dataset of frames of human-curated edited footage. Alternatively, or in addition, the first dataset may be obtained from a machine-curated dataset, for example, from an auxiliary machine learning model. The human-curated dataset of frames is obtained here by, for example, providing footage of the subject region by cameras 21, 22, and by allowing a human to navigate the subject region within the footage to appropriately depict the subject region by virtual tilt, virtual pan, and virtual zoom within the footage of the subject region, for example, by using a human-controllable input device such as a joystick or similar. The human-curated edited footage may be recorded as a human-curated dataset of frames. The human may navigate the entire footage of the subject region or only a selected subset thereof, for example, only "key" moments within the footage. The human-curated edited footage may include, for example, interpolated frames between those key moments. Human-curated datasets of navigation parameters are also recorded (optionally interpolated), which correspond to human-applied navigation, such as virtual tilt, virtual pan, and virtual zoom, allowing for the appropriate depiction of interesting parts of the subject area over time.

[0105] The human-curated dataset of navigation parameters is augmented by the augmentation dataset. In this example, augmentation includes pan augmentation, tilt augmentation, and zoom augmentation. Thus, here, each navigation parameter includes pan, tilt, and zoom settings, and each navigation parameter in the navigation parameter dataset is offset by pan augmentation, tilt augmentation, and zoom augmentation. The augmentation can be randomly generated, for example, by subtraction from a normal distribution. The augmented dataset of frames is transformed into different augmented edited videos in a known style from the human-curated edited videos by applying the augmentation. The augmented edited videos, i.e., the dataset of frames, are labeled with the augmentation dataset to create a first dataset for training a machine learning model.

[0106] The machine learning model may be trained on a second dataset, either in addition to or alternatively. The second dataset may be obtained by creating edited footage from video of a subject area. The edited footage may be generated in various ways. In this example, the edited footage is produced, for example, using the first dataset and a pre-trained machine learning model. This provides a particular efficient training method. The difference between the edited footage produced by the model, i.e., the dataset generated by the model of frames, and the edited footage curated by a human, i.e., the human-curated dataset of frames, may be determined. Each frame in the dataset generated by the model of frames is compared, for example, to the corresponding frame in the human-curated dataset of frames, and the differences between them, such as relative tilt values, relative pan values, and relative zoom values, may be stored in the difference dataset. The dataset generated by the model of frames is labeled in the difference dataset to create a second dataset for training the machine learning model. Further generations of datasets may be generated accordingly by the pre-trained machine learning model. Different generations of datasets can be combined to prevent undesirable biases in the training data.

[0107] It will be understood that the methods described herein may include computer implementation steps. All of the steps described above can be computer implementation steps. Embodiments may include a computer device, and the process is performed in the computer device. The present invention is also extended to computer programs, in particular computer programs on or within a carrier made to carry out the present invention. The program may be in the form of source code or object code, or any other form suitable for use in implementing the process according to the present invention. The carrier may be any entity or device capable of carrying the program. For example, the carrier may include a storage medium such as ROM, e.g., semiconductor ROM, or a hard disk. Furthermore, the carrier may be a transmittable carrier such as an electrical signal or optical signal that can be transmitted via an electrical cable or optical cable, or wirelessly or by other means, e.g., via the internet or the cloud.

[0108] Some embodiments may be implemented, for example, using a machine or a tangible computer-readable medium or article capable of storing instructions or sets of instructions that, when performed by a machine, can cause a machine to perform the methods and / or operations according to the embodiment.

[0109] Various embodiments may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include processors, microprocessors, circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, microchips, chipsets, etc. Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, mobile apps, middleware, firmware, software modules, routines, subroutines, functions, computer implementation methods, procedures, software interfaces, application programming interfaces (APIs), methods, instruction sets, computing code, computer code, etc.

[0110] Herein, the present invention is described with reference to specific examples of embodiments of the invention. However, it will be apparent that various modifications and changes can be made without departing from the essence of the invention. For the purpose of clarity and brevity, features are described herein as part of the same or distinct embodiments, but alternative embodiments having all or some combinations of features described in these distinct embodiments are also contemplated.

[0111] However, other modifications, alterations, and alternative forms are also possible. Therefore, this specification, drawings, and examples should be considered illustrative rather than restrictive.

[0112] In the claims, reference symbols in parentheses shall not be construed as limiting the claims. The phrases “including” and “equipped with” shall not exclude the existence of features or steps other than those described in the claims. Furthermore, the words “a” and “an” shall not be construed as limiting to “only one,” but rather as meaning “at least one,” and not as excluding the plural. The mere fact that certain measurements are cited in different claims shall not indicate that combinations of these measurements cannot be used to one's advantage. [Explanation of symbols]

[0113] 1. Video Stream 1.1, 1.2 First or second frame 1a, 1b Video Streams 2 Edited video stream 2.1 First video frame 2.1' First frame 2.2 Second Frame 2.2' Extended second frame 2.3 Frame 2. i-frame 3.1 Margin 5.2 Navigation Parameters 5.3 Navigation Parameters 5.j Navigation Parameters 7.j Feedback signal 10 Video Production Units 20 Camera Systems 21, 22 Fixed cameras 23 Mechanically and / or optically adjustable cameras 30 Processing Modules 40 Generator Modules 41 Mapping Module 100 Systems

Claims

1. A computer implementation method for autonomous production of an edited video stream of a subject area, A step of obtaining a first frame from the edited video stream of a first sub-region of the subject region, The first step involves inputting the aforementioned first frame into a machine learning model, The machine learning model determines navigation parameters related to a second sub-region of the subject region based on the first frame, A step of obtaining a second frame for the edited video stream of the second subregion based on the determined navigation parameters. Computer implementation methods, including those mentioned above.

2. The method according to claim 1, wherein the first frame has a first frame setting associated with it, the second frame has a second frame setting associated with it, and the navigation parameter represents a relative or absolute adjustment from the first frame setting to the second frame setting.

3. The method according to claim 1 or 2, wherein the navigation parameter includes one or more of pan adjustment, tilt adjustment, and zoom adjustment.

4. The method according to any one of claims 1 to 3, wherein the first frame and / or the second frame of the edited video stream are obtained by taking subframes of frames of the video stream that are at least a portion of the subject area.

5. The method according to any one of claims 1 to 4, wherein the navigation parameter includes one or more of virtual pan adjustment, virtual tilt adjustment, and virtual zoom adjustment.

6. The method according to any one of claims 1 to 5, wherein the first frame and / or the second frame of the edited video stream are acquired by a camera that is automatically, mechanically, and / or optically adjustable.

7. The method according to any one of claims 1 to 6, wherein the navigation parameter includes one or more of mechanical pan adjustment, mechanical tilt adjustment, and optical zoom adjustment.

8. The method according to any one of claims 1 to 7, wherein the navigation parameters include adjustment speed and adjustment direction.

9. The method according to any one of claims 1 to 8, wherein the first frame is acquired at a first time, and the navigation parameters are determined at the first time and are related to a second time after the first time.

10. The method according to claim 9, comprising the step of using the machine learning model at a first time to determine a time series of navigation parameters relating to a plurality of time periods after the first time, based on the first frame, wherein the time series of navigation parameters particularly includes the navigation parameters.

11. The method according to claim 10, further comprising the step of determining adjusted navigation parameters based on the aforementioned navigation parameters and at least one further navigation parameter related to a second time determined at a time preceding the first time.

12. The method according to claim 11, wherein the second frame of the edited video stream for the second subregion is obtained based on the adjusted navigation parameters.

13. The method according to claim 11 or 12, wherein the adjusted navigation parameters are determined as a weighted average of a plurality of navigation parameters determined at a time that is related to the second time and is prior to the second time, and in particular, the navigation parameters among the plurality of navigation parameters that are temporally closer to the second time are given a greater weight than the navigation parameters among the plurality of navigation parameters that are temporally further from the second time.

14. The method according to any one of claims 9 to 13, wherein the first time and the second time are separated by a time period determined to take into account the waiting time of the method, in particular between the capture of the first frame and the finalization of the physical adjustment of the camera corresponding to the navigation parameters.

15. The method according to any one of claims 1 to 14, wherein the navigation parameter is determined as a normalized navigation parameter normalized with respect to the field of view of the first frame.

16. The method according to claim 15, comprising the steps of denormalizing the normalized navigation parameters and obtaining the second frame for the edited video stream of the second subregion based on the denormalized navigation parameters.

17. The method according to any one of claims 1 to 16, comprising the steps of: obtaining an extended first frame of the extended first subregion of the subject region which is wider by a predetermined margin than the first subregion; and determining the navigation parameters based on the extended first frame.

18. The method according to claim 17, wherein the margin is asymmetric with respect to the first subregion.

19. The method according to claim 17 or 18, comprising the step of labeling the data points of the extended first frame to distinguish between data points of the extended first frame corresponding to data points of the first frame and data points of the extended first frame corresponding to data points of the margin.

20. The method according to any one of claims 1 to 19, comprising the steps of resampling, for example, an extended first frame, and determining the navigation parameters based on the resampled, for example, extended first frame.

21. The method according to claim 20, wherein the resampling is such that a non-uniform sample is obtained across the entire frame.

22. The method according to any one of claims 1 to 21, wherein the navigation parameter is determined based on a set of frames of the edited video stream, which includes only the first frame.

23. The method according to any one of claims 1 to 21, wherein the navigation parameter is determined based on a set of frames of the edited video stream, which includes at least two frames.

24. The method according to claim 23, wherein the set of frames of the edited video stream includes frames that temporally precede the second frame.

25. The method according to claim 23 or 24, wherein the set of frames of the edited video stream includes frames that temporally follow the second frame.

26. The method according to any one of claims 1 to 25, comprising the steps of: determining the navigation parameters, adjusting the navigation parameters according to a smoothing criterion; and obtaining the second frame for the edited video stream based on the adjusted navigation parameters.

27. The method according to claim 26, wherein the determined navigation parameter is adjusted based on a set of navigation parameters including navigation parameters that precede the determined navigation parameter in time and / or navigation parameters that follow the determined navigation parameter in time.

28. The method according to claim 27, comprising the steps of determining a navigation trend based on the set of navigation parameters, and adjusting the navigation parameters in accordance with the determined navigation trend.

29. The method according to any one of claims 1 to 28, wherein the machine learning model includes an end-to-end artificial neural network.

30. A computer implementation method for generating a dataset for a machine learning model to be used in any one of claims 1 to 29, particularly a training dataset, comprising the steps of providing images of a subject region; A step of having a human and / or auxiliary machine learning model navigate the subject area in the video, From there, the step of obtaining a curated dataset of navigation parameters, and / or The next step is to obtain a curated dataset of frames from there. Computer implementation methods, including those mentioned above.

31. The steps include causing the trained auxiliary machine learning model to detect the presence of a predetermined entity of interest within one or more frames of the video, The steps include causing the trained auxiliary machine learning model to determine the placement of the detected entity of interest within one or more frames of the video, From there, the step of obtaining the curated dataset of navigation parameters and The method according to claim 30, including the method described in claim 30.

32. The steps include extending the curated dataset of navigation parameters with an extended dataset, The steps include generating a dataset of frames corresponding to the dataset of extended navigation parameters, The steps include labeling the frame dataset with the extended dataset and The method according to claim 30 or 31, including the method described in claim 30 or 31.

33. For example, the method according to any one of claims 1 to 29, comprising the steps of generating a dataset of frames for the edited video from the video, A step of obtaining a difference dataset representing the difference between the generated dataset of the frame and the curated dataset of the frame, The steps include labeling the frame dataset with the difference dataset and The method according to claim 30, 31, or 32, including the following:

34. A computer implementation method for training a machine learning model for the autonomous production of an edited video stream from a video stream of a subject region using a dataset obtained by any one of claims 30 to 33.

35. The method according to claim 34, wherein the machine learning model is configured to output a time series of navigation parameters.

36. The method according to claim 35, wherein the machine learning model is configured to output the time series of navigation parameters that fit predefined progress characteristics.

37. For example, the method according to claim 32 or claim 33, comprising the steps of generating a first training dataset and pre-training the machine learning model using the first dataset, For example, the method according to claim 33, comprising the steps of causing the pre-trained machine learning model to generate a second training dataset, and training the pre-trained machine learning model using the second dataset, The method according to any one of claims 34 to 36, including the method described in any one of claims 34 to 36.

38. After training the machine learning model using the second dataset, the machine learning model is made to generate a third dataset, for example, by the method described in claim 33. The steps include training the machine learning model using at least a portion of the third dataset and The method according to claim 37, including the method described in claim 37.

39. A machine learning model for autonomous production of edited video streams of a subject region, trained by the method described in any one of claims 34 to 38.

40. Edited video obtained by the method described in any one of claims 1 to 38.

41. A non-temporary computer-readable medium for storing instructions that cause a device to perform the method according to any one of claims 1 to 38 when executed by one or more processors.

42. A system for the autonomous production of edited video streams of a subject area, Camera system and, A processing unit operably connected to the camera and configured to perform the method described in any one of claims 1 to 38, and A system equipped with these features.

43. The system according to claim 42, further comprising a fixed panoramic camera configured to record a panoramic video stream of the subject area from a stationary line of sight.

44. The camera system according to claim 42 or 43, comprising a mechanically and / or optically adjustable camera configured to record a video stream of the subject area from an adjustable line of sight.

45. The system according to claim 44, comprising a camera controller configured to be operablely positioned between the processing unit and the mechanically and / or optically adjustable camera, wherein the camera controller is configured to transmit one or more video frames acquired by the camera to the processing unit, receive navigation parameters from the processing unit, and mechanically and / or optically adjust the camera based thereon.