Method for autonomous production of an edited video stream
Patent Information
- Application Number
- BR112025020696
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Publication Date
- 2026-08-25
Smart Images

Figure 00000000_0000_ABST
Description
1 / 51 METHOD FOR AUTONOMOUS PRODUCTION OF AN EDITED VIDEO STREAM FIELD
[0001] The invention relates to a stand-alone video event production system, for example, of sporting events. STATE OF THE ART
[0002] Popular events such as live sporting events and music concerts are usually recorded by multiple cameras. The video footage recorded by the various cameras is processed into an edited video production, for example, to be broadcast to viewers of the event. The multiple cameras are conventionally operated by humans to capture the event in an attention-grabbing way, for example, by dynamically directing and zooming the camera into areas of interest at the event. Such video production can be relatively laborious and expensive and is therefore generally not financially feasible for small and medium-scale events.
[0003] Automated video production systems can be used to reduce overall video production costs. Known automated video production systems typically involve a stationary overview camera that records an overview video of the event, where automated detection methods are employed to automatically detect features of interest within the overview video, and extracting subframes from the overview video that depict the features of interest. Features in the area of interest can be automatically detected and tracked. Petition 870250101914, dated 06 / 11 / 2025, page 9 / 63 2 / 51 by dedicated tracking and detection algorithms. However, such conventional automated functionality tracking systems are, when compared to human-operated systems, poor mainly at appreciating situations that depend on the context of events that may be of interest to a human audience. The quality of video production resulting from such automated methods is therefore generally perceived as inferior to conventional human-edited video productions. SUMMARY
[0004] It is an objective to provide an improved method for the automated production of event video. In a more general sense, it is an objective to overcome or improve some of the disadvantages of the previous technique, or at least to provide alternative processes that are more effective than the previous technique and that can be used relatively cheaply. At any rate, the present invention at least aims to offer a useful alternative and contribution to existing technique.
[0005] According to a first aspect, a computer-implemented method is provided for the autonomous production of an edited video stream of a subject area. The edited video can be produced from a video stream of the subject area, for example, a stationary and / or dynamic video stream of the subject area. The video stream, for example, can be acquired from one or more cameras depicting the subject area or only parts thereof. The subject area can be a real-world area, such as a scene from an event as depicted by the video stream. The area Petition 870250101914, dated 06 / 11 / 2025, page 10 / 63 3 / 51 The theme, for example, could be a sports field, a podium, or parts thereof.
[0006] The method comprises obtaining a first frame for the edited video transmission of a first sub-area of the subject area. The sub-area is a part of the subject area, and may be smaller than the subject area. For example, the first frame for the edited video may be a cropped frame from the video transmission of the subject area. Alternatively, the first frame may have been obtained from a camera, for example, automatically controlled and thus directed and zoomed in to depict the sub-area. The automatically controlled camera may thus be mechanically and / or optically adjusted to provide an adjustable viewing direction.
[0007] The method involves entering the first frame into a machine learning model. The machine learning model can be pre-trained, for example, according to training methods described here.
[0008] The method involves determining, using the machine learning model, and based on the first frame, a navigation parameter that refers to a second sub-area of the subject area. The second sub-area may be the same as or different from the first sub-area.
[0009] The method involves obtaining a second frame for the edited video transmission from the second sub-area, based on the determined navigation parameter. Like the first frame, the second frame for the edited video transmission can also be a clipping, for example, a different clipping, from an overview video frame of the Petition 870250101914, dated 06 / 11 / 2025, page 11 / 63 4 / 51 video transmission of the subject area. Additionally or alternatively, the second frame can be obtained by a physical adjustment of the camera, mechanically and / or optically adjustable, for example, by changing the viewing direction and / or the optical zoom setting of the camera.
[00010] The method thus allows the production of an edited video stream in an iterative manner, for example, where each frame of the edited video is obtained from another, for example, a previous or subsequent frame. For example, the second frame could be a next frame for the edited video stream, which could be determined from the current first frame.
[00011] The method allows for the automatic production of edited videos, for example, in real time, with virtually negligible computational costs, as the second frame can be determined based on the first frame and, for example, not based on a stationary or non-stationary overview video stream of the subject area. Frames from the overview video stream are generally much larger and generally contain more information to determine the appropriate sub-area that may be of interest for the edited video than frames for the edited video itself. Yet, in some situations, an overview video stream may not be available. The inventors, however, found that the frames for the edited video contain enough information to be used as an input to the machine learning model for the production of a high-quality edited video.By using frames from the edited video, the computational load of the method can be kept very low. This can particularly provide a system that... Petition 870250101914, dated 06 / 11 / 2025, page 12 / 63 5 / 51 low-latency method, which allows real-time or near real-time adjustment of a mechanically and / or optically adjustable camera to obtain the frames for the edited video.
[00012] Additionally, the machine learning model can take contextual information into account to infer the appropriate frame for the edited video, particularly when compared with conventional automated functionality tracking systems and methods.
[00013] The second frame for the edited video transmission of the second subarea is determined based on the navigation parameter, which was determined by the machine learning model. The method may optionally include the determination of an adjusted navigation parameter based on the navigation parameter. The navigation parameter can be adjusted, for example, by taking into account secondary indices, such as a historical and / or future progression of the navigation parameters over time. This can provide a smooth edited video.
[00014] It will be noticed that the first frame and the second frame can be consecutive frames, but that there may alternatively be other frames between the first frame and the second frame. The first frame may precede the second frame in time, but the first frame may also succeed the second frame in time. In one example, the first frame and the second frame correspond to the same instant in time.
[00015] The machine learning model can be pre-trained specifically for, based on an input video frame that is associated with a Petition 870250101914, dated 06 / 11 / 2025, page 13 / 63 6 / 51 frame configuration such as a tilt configuration, a pan configuration and / or a zoom configuration, determine a selected navigation parameter of a correction of said frame configuration for said input video frame. The frame configuration may include in particular one or more of a tilt configuration, a pan configuration and a zoom configuration, depicting a camera view from a real or virtual camera that acquires the associated video frame. The machine learning model may be trained appropriately to receive the input frame, and to infer from it the associated frame configuration. The machine learning model may be configured to issue a correction to the frame configuration associated with the input frame, which may make the input frame more appropriate for the scene depicted by the input frame.
[00016] Thus, the first aspect may provide in a more particular way a computer-implemented method for the autonomous production of an edited video stream of a thematic area, comprising: providing a pre-trained machine learning model configured to receive an input video frame that is associated with a frame setting such as a pan setting, a tilt setting and / or a zoom setting, wherein the pre-trained machine learning model is pre-trained to determine, based on the input video frame, a navigation parameter selected from a frame setting correction for said input video frame; obtaining a first frame for the edited video stream of a Petition 870250101914, dated 06 / 11 / 2025, page 14 / 63 7 / 51 first sub-area of the thematic area; input the first frame into the pre-trained machine learning model; determine, using the pre-trained machine learning model, based on the first frame, a navigation parameter that refers to a second sub-area of the thematic area; and obtain a second frame for the edited video transmission of the second sub-area based on the determined navigation parameter. Thus, the first frame can be associated with a first frame configuration, where the determined navigation parameter is selected from a correction for the first frame configuration. A second frame configuration can be obtained by applying the correction to the first frame configuration. The second frame configuration can be used to determine an input for a camera system, in order to obtain the second frame. This process can be repeated iteratively to produce frames for the edited video.
[00017] Optionally, the first and / or second frame of the edited video stream is obtained by taking a subframe from a frame of a video stream of at least part of the subject area. The subframe, for example, can be taken from a frame of a stationary or non-stationary video stream offering an overview of the subject area. The subframe, for example, can also be taken from a frame of a camera, for example, a mechanically and / or optically adjustable non-stationary PTZ camera. The first and / or second frame can thus be virtual camera projections, for example, clippings, from the video stream. The subframes, for example, can be obtained by cutting frames from the stream. Petition 870250101914, dated 06 / 11 / 2025, page 15 / 63 8 / 51 of video and preferably adjusting, for example, by straightening, a perspective of the cut.
[00018] The edited video may include a set, for example, a time series, of frames, including the first frame and the second frame. Each frame of the edited video, for example, may be a subframe, for example, a clipping, from a respective frame of the thematic area video. The first frame, for example, may be a subframe of a first frame of an overview video of the thematic area, and the second frame, for example, may be a subframe of a different second frame of the overview video of the thematic area.
[00019] The first frame and second frame for the edited video stream can be obtained from respective frames of the thematic area overview video stream. The first frame for the edited video stream of the first sub-area, for example, can be obtained from a first frame of the thematic area overview video stream. Similarly, the second frame for the edited video stream of the second sub-area, for example, can be obtained from a second frame of the thematic area overview video stream.
[00020] Optionally, the first and / or second frame of the edited video transmission is obtained by a mechanically and / or optically adjustable camera. Additionally or alternatively, by taking subframes from the stationary overview video of the subject area, each frame of the edited video can be associated with a respective physical configuration of a mechanically and / or optically adjustable camera, wherein the Petition 870250101914, dated 06 / 11 / 2025, page 16 / 63 A 9 / 51 camera that is mechanically and / or optically adjustable, for example, can capture only part of the subject area.
[00021] Optionally, the first frame for the edited video stream of the first subarea is not the same as the first frame of the video stream of the thematic area. The first frame for the edited video stream of the first subarea, for example, may be smaller than, for example, a crop from, the first frame of the video stream of the thematic area.
[00022] Optionally, the second frame for the edited video transmission of the second sub-area is not the same as a second frame of the video transmission of the thematic area. The second frame for the edited video transmission of the second sub-area, for example, may be smaller than, for example, a crop from, the second frame of the video transmission of the thematic area.
[00023] Optionally, the first frame has a first frame configuration associated with it, and the second frame has a second frame configuration associated with it, where the navigation parameter is selected from an adjustment from the first frame configuration to the second frame configuration. Thus, the navigation parameter can be a vector that indicates a direction and speed or an amount by which the direction of observation and / or field of view of the edited video transmission should be moved, for example, virtually or physically directing a camera, starting from the first frame depicting the first sub-area, to depict the second sub-area using the second frame. The first frame configuration, for example, may include one or more first tilt configurations, a Petition 870250101914, dated 06 / 11 / 2025, page 17 / 63 10 / 51 first pan setting and a first zoom setting, for example, from a real or virtual camera. Similarly, the second frame setting may include one or more of a second tilt setting, a second pan setting, and a second zoom setting, for example, from the virtual camera. The navigation parameter can thus provide an adjustment, for example, a mapping, from the first pan setting to the second pan setting, from the first tilt setting to the second tilt setting, and from the first zoom setting to the second zoom setting.
[00024] Optionally, the navigation parameter includes one or more of the virtual pan adjustment, virtual tilt adjustment, and virtual zoom adjustment.
[00025] Optionally, the navigation parameter includes one or more mechanical pan adjustments, a mechanical tilt adjustment, and an optical zoom adjustment.
[00026] Optionally, the navigation parameter includes an adjustment speed and an adjustment direction, having the navigation parameter include an adjustment speed and direction can provide for a smoother transition between frames of the edited video, compared with, for example, a navigation parameter that includes absolute adjustment amounts or relative adjustment ratios.
[00027] Optionally, the navigation parameter includes one or more pan adjustments, a tilt adjustment, and a zoom adjustment. In particular, the navigation parameter includes one or more relative pan adjustments, a relative tilt adjustment, and an adjustment Petition 870250101914, dated 06 / 11 / 2025, page 18 / 63 11 / 51 relative zoom, with respect to the first frame, to adjust the first frame setting to the second frame setting.
[00028] Optionally, one or more of the pan, tilt, and zoom adjustments is a virtual adjustment. The adjustment is virtual in the sense that no physical movement of a camera is required, but the adjustment may involve selecting another crop from an overview of the subject area instead of, that is, adjusting (virtually) a pan, tilt, and / or zoom setting of a camera. Thus, the navigation parameter may include one or more of a relative virtual pan adjustment, a relative virtual tilt adjustment, and a relative virtual zoom adjustment, with respect to the first frame. Thus, each of the first and second frames of the edited video may be a crop from a respective frame of the subject area video stream.
[00029] Optionally, one or more of the pan adjustment, the tilt adjustment, and the zoom adjustment is a physical adjustment. The adjustment is physical in the sense that it involves a physical movement of a camera, for example, a mechanical pivoting of the camera relative to a base and / or a mechanical movement of one or more lenses to change the optical zoom.
[00030] Optionally, the first frame is obtained at a first time instant and the navigation parameter is determined at the first time instant and belongs to a second time instant subsequent in time to the first time instant. Petition 870250101914, dated 06 / 11 / 2025, page 19 / 63 12 / 51
[00031] Optionally, the method comprises, at the first time instant, determining, using the machine learning model and based on the first frame, a time sequence of navigation parameters belonging to multiple subsequent time instants in time to the first time instant, the time sequence of navigation parameters particularly including the navigation parameter. The machine learning model can thus, at a current time instant, determine multiple navigation parameters for multiple future time instants. A future frame configuration trajectory, in particular tilt, pan and / or zoom configurations, can thus be determined, based on a current frame configuration and / or a past frame configuration.
[00032] Optionally, the machine learning model is configured to determine the timing sequence of navigation parameters conforming to a predefined progression feature. The progression feature can be used to provide smooth edited video with minimal abrupt camera movement. The progression feature can appropriately constrain the progression of the time sequence navigation parameters. The progression feature, for example, can enforce that a view of a current frame is smoothly accelerated from, for example, by making a difference between successive time sequence navigation parameters progressively increase from the current frame. The progression feature, for example, can additionally enforce that a view is smoothly decelerated to a final setting of Petition 870250101914, dated 06 / 11 / 2025, page 20 / 63 13 / 51 time sequence frame, for example, making a difference between successive time sequence navigation parameters progressively decrease towards a final time sequence navigation parameter.
[00033] Optionally, the method comprises determining an adjusted navigation parameter based on the navigation parameter, and additionally based on at least one additional navigation parameter belonging to the second time instant that was determined at a time instant prior to the first time instant. Thus, a smooth edited video can be obtained in which sudden changes in the viewing direction are effectively smoothed.
[00034] Optionally, the second frame for the edited video transmission of the second sub-area is obtained based on the adjusted navigation parameter.
[00035] Optionally, the adjusted navigation parameter is determined as a weighted average of a plurality of navigation parameters, each belonging to the second time instant and each determined at a time instant prior to the second time instant, particularly where navigation parameters from the plurality of navigation parameters closer in time to the second time instant are given more weight than navigation parameters from the plurality of navigation parameters further in time from the second time instant. Navigation parameters that were predicted in the recent past belonging to a future time instant are generally in practice more accurate than navigation parameters that were predicted in the more distant past belonging to a future time instant and therefore, Petition 870250101914, dated 06 / 11 / 2025, page 21 / 63 14 / 51 may be prioritized when determining the adjusted navigation parameter.
[00036] Optionally, the first time instant and the second time instant are spaced in time by a period of time that is determined to account for a method latency, particularly a time delay from the capture of the first frame to the completion of a physical adjustment of a camera associated with the navigation parameter. Particularly for the control of a mechanically and / or optically adjustable camera, the time delay, for example, the system and method latency, can be substantial. The latency, for example, can be attributed to several method steps, for example, one or more of a video frame acquisition time, a data transmission time, a processing time, and a camera adjustment time. The latency can thus be anticipated to optimize system performance. The latency can be determined by measuring or estimating a time delay between method steps.The navigation parameter determined at a first point in time, for example, the current one, may belong to a camera configuration at a second point in time in the future, which is separated in time from the first point in time by a period of time that corresponds to the latency. It will be noticed that the time delay can span several video frames, that is, the latency can be greater than a camera sampling time interval. The navigation parameter may thus belong to a point in time that is a number of camera sampling time instances in the future. There may appropriately be time instances between the first and second points in time. Petition 870250101914, dated 06 / 11 / 2025, page 22 / 63 15 / 51 of time. Particularly when the video frames for the edited video are captured using a mechanically and / or optically adjustable camera, for example, a PTZ camera, the first time instant and the second time instant may be spaced in time by a period of time that at least substantially corresponds with a time delay from the capture of the first frame to the completion of a physical adjustment of the camera associated with the navigation parameter.
[00037] Optionally, the navigation parameter is determined as a normalized navigation parameter, which is normalized with respect to a field of view of the first frame. For example, one or more of the pan adjustment, the tilt adjustment, and the zoom adjustment are determined in normalized form, for example, with respect to a field of view of the first frame. In this way, for example, a machine learning model can be effectively trained to generate the navigation parameter, particularly since it allows determining a relative pan and tilt adjustment for the second frame independently of the field of view, for example, the zoom setting, of the first frame. The effect that a certain pan and tilt adjustment can have different effects for different zoom settings of the first frame can thus be taken into account efficiently.The field of view of the first frame can be thought of as an extension of the first sub-area represented by the first frame. The field of view, for example, can be defined as a dimension of the first frame, such as a height and / or width of the first frame, and / or an angle of observation associated with it. Petition 870250101914, dated 06 / 11 / 2025, p. 23 / 63 16 / 51 first frame. The field of view can be defined by an optical or virtual zoom setting on a camera.
[00038] Optionally, the method comprises denormalizing the normalized navigation parameter, for example, based on the field of view of the first frame, and obtaining the second frame for the edited video transmission of the second subarea based on the denormalized navigation parameter. For example, the method may comprise denormalizing one or more of the normalized pan adjustment, the normalized tilt adjustment, and the normalized zoom adjustment, and obtaining the second frame for the edited video transmission of the second subarea based on the denormalized pan, tilt, and zoom adjustments.
[00039] Optionally, the video transmission is recorded from a stationary viewpoint by one or more cameras. Recordings from multiple cameras, for example, can be combined to form the video transmission.
[00040] Optionally, the method comprises calibrating one or more stationary or non-stationary cameras. Calibration may include, in particular, determining a mapping between a coordinate system of one or more cameras and a coordinate system of the subject area. Calibration, for example, may include determining an orientation of one or more stationary cameras with respect to the subject area, such as determining the orientation of a pan axis and / or a tilt axis with respect to a horizontal plane and / or a vertical and horizontal spacing between the subject area and the one or more cameras. Calibration may additionally include Petition 870250101914, dated 06 / 11 / 2025, page 24 / 63 17 / 51 determination of a correction mapping to correct a lens distortion of one or more cameras. Calibration may include correlated video frames from respective cameras to join the video frames to form a single video transmission.
[00041] Optionally, the method comprises obtaining a first extended frame from a first extended subarea of the subject area that is larger than the first subarea by a margin, for example, a predetermined margin, and determining the navigation parameter based on the first extended frame. Thus, the first extended frame may include additional information compared to the first frame, which can be used to improve the determination of the navigation parameter and, in turn, the second frame. The first extended subarea may be smaller than the subject area to manage the computational burden. The first extended frame, for example, may be a crop from a frame of the subject area video. The first extended frame, for example, may be obtained by zooming out from the first frame, for example, by a predetermined amount.
[00042] Optionally, the margin is asymmetrical with respect to the first subarea. Certain regions of the subject area generally may not provide very useful additional information, while other regions generally do. For example, regions above the first frame may generally provide more useful information than regions below the first frame. The margin can thus be arranged asymmetrically around the first subarea, for example, to have a relatively wide margin above the first frame and a relatively thin margin below the first frame. Petition 870250101914, dated 06 / 11 / 2025, page 25 / 63 18 / 51 An asymmetrical margin, for example, can be achieved by zooming out from the first frame and additionally tilting and / or panning the view.
[00043] Optionally, the margin is dependent on the subject area.
[00044] Optionally, the margin is dependent on the navigation parameter, for example, dependent on a zoom value, and / or dependent on the first frame for the edited video, for example, the size of the first frame relative to a frame of the video stream in the subject area. The margin, for example, can be relatively small if the first frame is a relatively large crop of a frame from the video stream in the subject area, for example, zoomed out, and relatively large if the first frame is a relatively small crop from a frame of the video stream in the subject area, for example, zoomed in.
[00045] Optionally, the method includes marking the data points of the first extended frame to distinguish between data points of the first extended frame that correspond with data points of the first frame and data points of the first extended frame that correspond with data points of the margin. Thus, an appropriate navigation parameter can be obtained, in which it has been automatically taken into account which part of the first extended frame represents the first frame and which part does not. Each extended frame, for example, may include a dedicated marker channel to mark data points that are in the margin and / or mark data points that are not in the margin. For example, each frame may include Petition 870250101914, dated 06 / 11 / 2025, p. 26 / 63 19 / 51 one or more colored channels, for example, a red channel, a green channel, and a blue channel, as well as an additional marker channel. With the marker channel, it can be automatically differentiated between parts of the extended frame that form the border and are appropriately not visible to viewers of the edited video stream, and parts of the extended frame that form the frame for the edited video and should be visible to viewers of the edited video stream.
[00046] Optionally, the method comprises resampling the first frame, for example, extended, and determining the navigation parameter based on the first frame, for example, extended, resample. The first frame, for example, extended, can be resampled to a predetermined resolution. A first frame, for example, extended, of a relatively high resolution can be undersampled, for example, to improve computational efficiency. A first frame, for example, extended, of a relatively low resolution can be oversampled.
[00047] Optionally, the act of resampling is such that uniform sampling across the frame is used to obtain the output frame. Regions of the first frame that generally include information, such as a central region, for example, may be given a higher resolution than regions that include less useful information, such as edge regions of the first frame.
[00048] Optionally, the navigation parameter is determined based on a set of frames from the edited video stream that contains only the first frame. It has been found that edited video streams Petition 870250101914, dated 06 / 11 / 2025, page 27 / 63 High-quality 20 / 51 can only be obtained by using the first frame to determine the navigation parameter, and then the second frame. This provides a particularly computationally efficient method.
[00049] Optionally, the navigation parameter is determined based on a set of frames from the edited video stream containing at least two frames. Thus, in addition to the first frame, additional frames can be used to determine the navigation parameter, and in turn the second frame. The data on which the navigation parameter is based can be enriched by taking into account multiple frames, for example, previous and / or subsequent frames. It will be noted that the previous and subsequent frames do not need to be consecutive in time, but that there may be intermediate frames.
[00050] Optionally, the edited video stream frame set includes frames that precede the second frame in time. Optionally, the edited video stream frame set includes frames that precede the first frame in time.
[00051] Optionally, the edited video stream frame set includes frames that follow the second frame in time. Optionally, the edited video stream frame set includes frames that follow the first frame in time. The edited video stream, for example, can be temporarily stored to allow the use of subsequent frames in determining the navigation parameter. The overview video of the subject area can be temporarily stored for the same amount of time. It will be noticed that the frame set of Petition 870250101914, dated 06 / 11 / 2025, page 28 / 63 21 / 51 edited video transmission may include frames that precede and / or follow and / or coincide with the first frame in time. The navigation parameter, and the second frame in turn, can thus be determined based on the first frame, and additionally for one or more other frames.
[00052] Optionally, the first frame is obtained at a first time instant and the navigation parameter is determined at the first time instant and belongs to a second time instant that succeeds the first time instant, wherein the method comprises at the first time instant determining, using the machine learning model, and based on the first frame, a time sequence of navigation parameters belonging to multiple time instants that succeed the first time instant, the time sequence of navigation parameters particularly including the navigation parameter; wherein the time sequence of navigation parameters is determined based on a set of frames from the edited video transmission including subsequent frames that succeed the first frame in time;and optionally where the navigation parameter timeline is determined based on a set of frames from the edited video stream, including previous frames that precede the first frame in time. In particular, the navigation parameter timeline may include a navigation parameter that belongs to a future time instant relative to a current time instant. Thus, the edited video stream may be temporarily stored to allow subsequent frames of the edited video to be taken into account in determining the navigation parameter timeline, wherein the; Petition 870250101914, dated 06 / 11 / 2025, page 29 / 63 22 / 51 The navigation parameter forecast horizon can extend beyond the current time instant into the future. It will be noticed that subsequent and / or previous frames within temporary storage can be stitched together as time progresses from one time instant to the next, and a new time sequence of navigation parameters is determined. Frames within the temporary storage time window can thus be considered as estimated frames. These frames for the edited video transmission that are outside temporary storage when time passes can be made part of the actual edited video that is, for example, transmitted directly to a viewer.
[00053] Optionally, the method comprises, after determining the navigation parameter, adjusting the navigation parameter according to a smoothing criterion, and obtaining the second frame for the edited video transmission based on the adjusted navigation parameter. A subsequent processing step, for example, may be provided, in which the determined navigation parameter is compared with previous and / or subsequent navigation parameters, and adjusted appropriately. For example, some navigation parameters may be smoothed. Furthermore, the navigation parameter may be adjusted according to consecutive frames and / or navigation parameters to provide a smooth video.
[00054] Optionally, the determined navigation parameter is adjusted based on a set of navigation parameters that includes navigation parameters that precede the determined navigation parameter in time and / or navigation parameters that follow the parameter. Petition 870250101914, dated 06 / 11 / 2025, page 30 / 63 23 / 51 Time-determined navigation. The time-determined navigation parameter can be adjusted specifically based on a set of navigation parameters that includes navigation parameters that follow the time-determined navigation parameter, for example, to allow anticipation of large changes in the navigation parameters. If a large following navigation parameter is observed for a subsequent frame, the navigation parameter can be adjusted to said large following navigation parameter in anticipation of it. Thus, a large change that may have been caused by the large following navigation parameter can be mitigated by adjusting the navigation parameter appropriately, thus creating a smooth edited video.
[00055] Optionally, the method comprises determining a navigation trend based on said set of navigation parameters, and adjusting the navigation parameter according to the determined navigation trend.
[00056] Optionally, the navigation parameter is determined by a trained machine learning model, particularly an end-to-end artificial neural network. The machine learning model, for example, could be a convolutional neural network or a transformer deep learning architecture. The machine learning model can be end-to-end because it determines the navigation parameter directly based on the first frame, without requiring additional computational steps.
[00057] According to a second aspect, a computer-implemented method is provided for production Petition 870250101914, dated 06 / 11 / 2025, page 31 / 63 24 / 51 autonomously from an edited video stream from a video stream of a thematic area, comprising entering a first frame of the edited video from a first sub-area of the thematic area into a model, particularly a machine learning model; having the model determine, based on the first frame entered, a navigation parameter that refers to a second sub-area of the thematic area; and obtaining a second frame for the edited video stream of said second sub-area. The method may be particularly suited to any system or methods as described herein.
[00058] It will be noticed that the term model as used here has a broad meaning, and that functions, equations, algorithms, correspondences, mappings, and the like can be considered as models.
[00059] According to a third aspect, a machine learning model is provided for the autonomous production of an edited video stream of a thematic area, for use in a method according to the first aspect and / or the second aspect. The machine learning model may comprise a convolutional neural network or transformer architecture. The machine learning model may be arranged to receive a frame for the edited video, for example, an image, for example, the first frame, as an input, and generate the navigation parameter as an output. The navigation parameter, for example, may include a pan, tilt, and zoom value.
[00060] According to a fourth aspect, a computer-implemented method for generating a dataset, particularly a dataset of Petition 870250101914, dated 06 / 11 / 2025, page 32 / 63 25 / 51 Training for a machine learning model as per the third aspect. The method comprises providing a video of a subject area; having a trained or human auxiliary machine learning algorithm navigate the subject area within the video; and obtaining a selected dataset of navigation parameters from the same, and / or obtaining a selected dataset of frames from the same. The selected data may be selected by humans or selected by machine. Navigation of the subject area within the video may involve virtual zoom, virtual pan, and / or virtual tilt within the subject area video to obtain the selected edited video which is represented by the selected dataset of frames. Each frame of the selected dataset of frames may be associated with a navigation parameter from the selected dataset of navigation parameters.The selected set of navigation parameters, for example, can be obtained from commands of a human-controlled input device, such as a joystick or similar control device, which is operated by the human to adjust a pan, tilt, and / or zoom setting. Optionally, the method involves selecting a subset of frames from the video and having the human navigate the thematic area within the selected subset of video frames. The human, for example, can navigate only key frames of the video for efficient data annotation. The selected set of navigation parameters may optionally include interpolated navigation parameters that are associated with these video frames. Petition 870250101914, dated 06 / 11 / 2025, page 33 / 63 26 / 51 healed which are among the video frames that were annotated by the human.
[00061] Optionally, the method comprises having the trained auxiliary machine learning model detect the presence of a predetermined entity of interest in one or more frames of the video; having the trained auxiliary machine learning model determine the location of the detected entity of interest within one or more frames of the video; and obtaining the selected dataset of navigation parameters from it. For some applications, an auxiliary machine learning model may be more accurate in detecting and tracking a predetermined entity of interest, such as a person or an object, in a video than a human operator. Thus, it may be desirable to have the auxiliary machine learning model generate the training data instead of a human.It will be observed that the automated object tracking and detection of an object by the auxiliary machine learning model can, by itself, be used to generate training data for the machine learning model, but that the eventual production of the edited video by the machine learning model may not involve entity tracking and detection in this way. Entity tracking and detection using the auxiliary machine learning model may be prone to errors in practice, but this can be mitigated in the training stage of the machine learning model. The auxiliary machine learning model, for example, may be configured to detect and track a player on a field, and may erroneously deviate when it also detects one. Petition 870250101914, dated 06 / 11 / 2025, page 34 / 63 27 / 51 other people, such as another player, a sign displaying a human, the audience, etc. In particular, the machine learning model for producing the edited video is preferably not a dedicated feature tracking system, as this can be very error-prone for live event broadcasting. The use of feature tracking systems in the training stage allows for the adjustment of training data to filter out errors, which is generally undesirable or impractical to do in the edited video production stage. Instead of simple object tracking and detection, the machine learning model is additionally trained to exhibit a desired frame transition behavior, preferably resembling or exceeding the performance of a human camera operator.
[00062] Optionally, the method according to the claim comprises augmenting the selected dataset of navigation parameters by an augmentation dataset; generating a frame dataset corresponding with the augmented navigation parameter dataset; and tagging the frame dataset with the augmentation dataset. The augmentation dataset particularly includes frame configuration augmentations, such as tilt augmentations, pan augmentations, and zoom augmentations. The augmentation dataset thus particularly shifts the navigation parameters as selected by a human or auxiliary machine learning model. By augmenting the selected dataset of navigation parameters in this way, the machine learning model is trained. Petition 870250101914, dated 06 / 11 / 2025, page 35 / 63 28 / 51 refers to an input frame for a correction applied to the frame configuration of the input frame. The machine learning model can thus be trained to correct the frame configuration of its input frame, based on the frame itself. The estimated correction that the machine learning model learns to generate can subsequently be used to obtain the second frame.
[00063] The selected dataset of navigation parameters, for example, can be augmented by increments that are generated randomly. The increments of the dataset of increments, for example, can be subject to a normal distribution.
[00064] Optionally, the method comprises generating a frame dataset for the edited video from the video, for example, according to a method as described herein or by the other method; obtaining a selected difference dataset from a difference between the generated frame dataset and the selected frame dataset; and tagging the frame dataset with the difference dataset. The frame dataset for the edited video from the video may be generated particularly by a machine learning model as described herein that has been pre-trained. This provides particularly efficient training.
[00065] According to a fifth aspect, a computer-implemented method is provided with training of a machine learning model for the autonomous production of an edited video stream from a video stream of a thematic area, using a Petition 870250101914, dated 06 / 11 / 2025, page 36 / 63 29 / 51 dataset obtained by a method in accordance with the fourth aspect.
[00066] Optionally, the machine learning model is configured to output a time sequence of navigation parameters.
[00067] Optionally, the machine learning model is configured to output the timing sequence of navigation parameters conforming to a predefined progression feature. The progression feature can be used to train the machine learning model to provide a smooth edited video with minimal abrupt camera movement. The progression feature can appropriately constrain the navigation parameters, where the constraint is particularly time-dependent. The progression feature thus forces the machine learning model to learn a smooth transition of navigation parameters over time, as opposed to just any sequence of navigation parameters which can lead to a choppy edited video. The machine learning model can thus be trained to learn how to smoothly correct the frame rate of the real or virtual camera.The progression feature can be appropriately used as a tool to control frame setup adjustment features during the machine learning model training stage, which in turn allows control over the smoothing behavior of the machine learning model. It will be observed that after training the machine learning model, little or no further processing may be necessary to smooth the edited video, since the model... Petition 870250101914, dated 06 / 11 / 2025, page 37 / 63 30 / 51 machine learning can be trained to provide a smooth edited video. The progression feature, for example, can enforce that a view of a current frame is smoothly accelerated away, for example, by making a difference between successive time sequence navigation parameters progressively increase from the current frame. The progression feature, for example, can further enforce that a view is smoothly decelerated to a final time sequence frame setting, for example, by making a difference between successive time sequence navigation parameters progressively decrease to a final time sequence navigation parameter.
[00068] Optionally, the method comprises generating a first training dataset and pretraining the machine learning model using the first dataset. The method further comprises having the pretrained machine learning model generate a second training dataset, and training the pretrained machine learning model using the second dataset.
[00069] The first dataset can be generated in several ways, and particularly by providing a video of a thematic area; having a person navigate the thematic area within the video; obtaining a selected dataset of navigation parameters from them; augmenting the selected dataset of navigation parameters by an augmentation dataset; generating a frame dataset corresponding with the augmented navigation parameter dataset; and tagging the frame dataset with the augmentation dataset. Petition 870250101914, dated 06 / 11 / 2025, p. 38 / 63 31 / 51
[00070] The second dataset can be generated by the machine learning model that was trained with the first dataset, by having the pre-trained machine learning model generate a frame dataset for the edited video from the subject area video; obtain a selected difference dataset from a difference between the generated frame dataset and the selected frame dataset; and tag the frame dataset with the difference dataset. This procedure can be repeated, such that the pre-trained machine learning model can generate additional generations of datasets or its own training. Datasets from different generations can optionally be merged to minimize biases in the training data.
[00071] Optionally, the method comprises, after training the machine learning model using the second dataset, having the machine learning model generate a third dataset; and training the machine learning model using at least part of the third dataset. The third dataset may be generated by the machine learning model that was trained with the first dataset and with the second dataset, comprising having the pre-trained machine learning model generate a dataset of frames for the edited video from the video; obtaining a dataset of selected differences from a difference between the generated dataset of frames and the selected dataset of Petition 870250101914, dated 06 / 11 / 2025, pp. 39 / 63 32 / 51 frames; and mark the set of frames with the set of differences.
[00072] According to a sixth aspect, an edited video is provided obtained by a method as described herein.
[00073] According to a sixth aspect, a non-transient, computer-readable medium is provided that stores instructions which, when executed by one or more processors, cause a device to perform the method as described herein.
[00074] According to a seventh aspect, a system is provided for the autonomous production of an edited video feed of a subject area. The system comprises a camera system; and a processing unit, operationally connected to the camera, configured to perform a method as described herein. The camera system may include a stationary overview camera arranged to record an overview video feed of the subject area from the stationary observation direction. Additionally or alternatively, the camera system may include a mechanically and / or optically adjustable camera, for example, a PTZ or pan-tilt-zoom camera, arranged to record a video feed of the subject area from an adjustable observation direction.The system may comprise a camera controller arranged operationally between the processing unit and the mechanically and / or optically adjustable camera, the camera controller being arranged to transmit one or more video frames acquired by the camera to the processing unit, and receive a navigation parameter from the unit. Petition 870250101914, dated 06 / 11 / 2025, pp. 40 / 63 33 / 51 processing, and based on the same, mechanically and / or optically adjusting the camera.
[00075] For example, the aspect may provide a system for the autonomous production of an edited video transmission of a thematic area, comprising a camera system including a general-view camera, for example, stationary to obtain general-view video frames of the subject matter and additionally a mechanically and / or optically adjustable camera to obtain video frames for the edited video; and a processing unit, operatively connected to the general-view camera and the mechanically and / or optically adjustable camera, configured to perform a method as described herein.
[00076] Optionally, the processing unit is configured to obtain a first frame of a first sub-area of the subject area, the first frame being operated from the overview camera; input the first frame into a machine learning model; determine, by the machine learning model, based on the first frame, a navigation parameter that refers to a second sub-area of the subject area; and transmit a selected control signal of the navigation parameter to the mechanically and / or optically adjustable camera, to adjust the mechanically and / or optically adjustable camera.The overview camera and the mechanically and / or optically adjustable camera can be operated independently, wherein the video frames from the overview camera, for example, clips from it, can be used as inputs to the machine learning model to determine the navigation parameter, while the frames from the mechanically and / or optically adjustable camera... Petition 870250101914, dated 06 / 11 / 2025, pp. 41 / 63 34 / 51 adjustable settings can be used for edited video. The mechanically and / or optically adjustable camera can thus be controlled based on the navigation parameter issued by the machine learning model.
[00077] It will be noticed that any of the aspects, functionalities and options described here can be combined. BRIEF DESCRIPTION OF THE DRAWINGS
[00078] Embodiments of the present invention will now be described in detail with reference to the accompanying drawings in which:
[00079] Figures 1 to 3 show schematic examples of a system and method for producing an edited video transmission.
[00080] Figures 4A and 4B show a schematic example of a method for producing an edited video transmission. DETAILED DESCRIPTION
[00081] Figures 1 to 3 show schematic examples of a system and method 100 for producing an edited video transmission of a thematic area. The system 100 comprises a video production unit 10, arranged here to receive a video transmission of the thematic area, obtained from a camera system 20 of the system 100. In the example of Figure 1, the camera system comprises two stationary cameras 21 and 22. In this example, each of the two cameras 21, 22 transmits respective video transmissions 1a, 1b to the video production module 10. The cameras 21, 22 are in this example stationary in practice, and directed in such a way as to acquire a respective video transmission 1a, 1b to Petition 870250101914, dated 06 / 11 / 2025, pp. 42 / 63 35 / 51 from a fixed, stationary viewpoint of at least part of the subject area, such as from an event venue, for example, a sports field or music concert. The respective video streams 1a, 1b can be combined to render an overview video stream of the subject area. The video production unit 10 receives the video stream(s) 1a, 1b, and produces an edited video stream 2 based on it, for example, for broadcast to event spectators.
[00082] In an alternative arrangement, or in addition to the stationary cameras 21, 22, the camera system 20 may include a mechanically and / or optically adjustable camera 23, for example, a pan-tilt-zoom camera or PTZ camera, arranged to provide an adjustable view of the subject area. The mechanically and / or optically adjustable camera, for example, may have a limited field of view and not capture the entire subject area in one frame. Figure 2 shows an example of a system 100 having a mechanically and optically adjustable camera 23, in addition to the stationary overview cameras 21, 22. Figure 3 shows an example of a system 100 having only a mechanically and optically adjustable camera 23, excluding a stationary overview camera.
[00083] The video production unit 10 in the example of Figures 1 and 2 includes a processing module 30. The processing module 30 is arranged here to receive and process the video transmissions 1a, 1b, received here from cameras 21, 22, to obtain a video transmission of the subject area. The two video transmissions 1a, 1b, for example, can be appropriately merged or joined to form a single frame transmission of Petition 870250101914, dated 06 / 11 / 2025, pp. 43 / 63 36 / 51 video of the thematic area. Video production unit 10 is arranged for the production of an edited video stream 2 from video stream 1, particularly in real time. The edited video stream 2 can be a time series of frames, which, for example, are transmitted directly to the event viewers, or stored in memory as an edited video for later viewing or further editing. The edited video, for example, can be used for automated event detection, which in turn can be used to automatically compile a summary video of highlights from the edited video.
[00084] The frames for the edited video transmission 2 are determined using a generator module 40 of the video production unit 10. The generator module 40 is configured to generate a navigation parameter 5.j based on an input video frame, for example, an image. The generator module 40, for example, including a machine learning model, such as a convolutional neural network or transformer, is specifically configured to have a first video frame 2.1 for the edited video depicting a first sub-area of the subject area as an input and generate a navigation parameter 5.2 based on the same, for transmission to the processing module 30. The processing module 30 receives the navigation parameter 5.2 generated by the generator, and in turn provides a second frame 2.2 for the edited video depicting a second sub-area of the subject area, associated with the navigation parameter 5.2. The second frame 2.2. For the edited video, in turn, it can be entered into the generator module 40, to determine one. Petition 870250101914, dated 06 / 11 / 2025, pp. 44 / 63 37 / 51 additional navigation parameter 5.3, to determine an additional frame 2.3 for the edited video. This process can be repeated. The method can thus be an iterative method, in which frames 2.i for the edited video transmission 2 are determined iteratively.
[00085] The navigation parameter 5.j, for example, can indicate a coordinate of the subject area, to obtain a crop from a video frame of the video stream of the subject area. In this example, the navigation parameter 5.2 indicates a relative adjustment from the first frame depicting the first sub-area of the subject area to the second frame depicting the second sub-area of the subject area, like a vector, indicating a direction and a quantity or speed at which the view of the edited video stream should be changed, starting from the first frame, to reach the second frame. The navigation parameter can indicate in particular one or more of a relative pan adjustment, a relative tilt adjustment, and a relative zoom adjustment.
[00086] In the example in Figure 1, the relative adjustment is a relative virtual adjustment, particularly a relative virtual tilt value, relative virtual pan value, and relative virtual zoom value, relative to the first frame and with respect to the overview video feed obtained from the stationary cameras 21, 22. It will be noticed that a tilt can correspond with an adjustment in the vertical direction, while a pan can correspond with an adjustment in the horizontal direction, or vice versa. In practice, a tilt axis and pan axis can be at an angle relative to a real-world horizontal plane. Thus, a calibration can Petition 870250101914, dated 06 / 11 / 2025, pages 45 / 63 38 / 51 can be used to provide a mapping between the camera's tilt and pan and the tilt and pan in the real world. A zoom can change a field of view, for example, an angle of observation. It will be noticed that the navigation parameter can alternatively be relative to a frame of the video feed of the subject area.
[00087] Processing module 30 can, for example, take a subframe from a frame of the video stream 1 of the subject area and output the subframe as a frame 2.i to the edited video stream 2. The subframe, for example, can show only a part of the subject area, particularly a part of the subject area that is more interesting to viewers. The subframe, for example, can be obtained by cropping a frame, optionally adjusted for a change in perspective.
[00088] The first frame for the edited video can be a subframe from a first frame of the stationary overview video stream of the subject area, and the second frame for the edited video can be a subframe from a second frame of the stationary overview video stream of the subject area. It will be noticed that the first and second frames of the subject area video stream do not need to be successive in time, but that other frames can be interposed between them. Similarly, it will be noticed that the first and second frames of the edited video stream do not need to be successive in time, but that other frames can be interposed between them.
[00089] It will also be noticed that the navigation parameter does not need to represent a virtual adjustment. Instead Petition 870250101914, dated 06 / 11 / 2025, pp. 46 / 63 39 / 51 In addition, a camera can be mechanically tilted and widened, and / or optically zoomed in or out, as shown for the examples in Figure 2 and Figure 3. It will also be noticed that instead of taking subframes from the overview video to generate video frames for the edited video, the frames for the edited video transmission can alternatively be obtained by physically adjusting a camera, for example, involving an optical zoom, mechanical tilt, and mechanical pan adjustment of the camera. Such a system can provide increased video quality compared to taking subframes from an overview video frame, but at the cost of increased delays due to the physical adjustment of the camera, compared to virtually cropping an overview video. Video frames acquired by the physically mobile camera can also be cropped to some extent, and the frames for the edited video obtained from the physically mobile camera can also be cropped.
[00090] In the example system of Figure 2, the video frames obtained from the overview cameras 21 and 22, particularly clippings from them, are used to determine the navigation parameter 5.j, while the video frames obtained from the mechanically and optically adjustable camera 23 are used to generate the edited video transmission 2. The mechanically and optically adjustable camera 23 is appropriately controlled here, i.e., mechanically and / or optically adjusted, based on the determined navigation parameter 5.j. Here, the navigation parameter 5.j is in this example converted by module 41, module 41 which in practice can be integrated with the generator module 40, to Petition 870250101914, dated 06 / 11 / 2025, pp. 47 / 63 40 / 51 an appropriate control signal 6.j for controlling camera 23. Camera 23 in this example transmits response signals 7.j back to the mapping module 41, indicating its settings, for example, its current pan, tilt and zoom settings.
[00091] In the example in Figure 3, system 100 does not include a stationary overview camera. In this example, system 100 includes only a mechanically and optically adjustable camera 23, whose acquired video frames are used directly as input to the generator module 40. Based on the video frames from the mechanically and / or optically adjustable camera 23, the generator module 40 generates a navigation parameter 5.j, which is transmitted to the camera 23 for the camera 23 to be mechanically and / or optically adjusted appropriately. Here, the navigation parameter 5.j is used directly as a control signal for the camera 23. The video transmission 2 acquired by the camera 23 in this example also directly represents the edited video transmission 2. The frames of the edited video transmission 2 are transmitted from the camera 23 to the generator module 40.Furthermore, in this example, camera 23 transmits status information to generator module 40, specifically including the zoom setting for use in normalizing the navigation parameter.
[00092] The navigation parameter 5.j can be generated by generator module 40 in normalized form, for example, with respect to frame 2.1 entered into generator module 40. For example, a slope value of 0.5 can indicate an upward slope by an amount corresponding to half the height of the first frame, and Petition 870250101914, dated 06 / 11 / 2025, pp. 48 / 63 41 / 51 A tilt value of -0.5 may indicate a downward tilt by an amount corresponding to half the height of the first frame. A zoom value may optionally be generated on a logarithmic scale. For example, a zoom value of 1 may indicate a zoom in by a factor of 2, and a zoom value of -1 may indicate a zoom out by a factor of 2.
[00093] The navigation parameter 5.j can be denormalized, for example, by processing module 30, to obtain a relative adjustment for the virtual or real camera 23 that appropriately represents the desired subarea of the subject area.
[00094] Figures 4A and 4B show a schematic example of a method, particularly in conjunction with the example system as shown in Figure 1. Figure 4A shows a schematic example of a first frame 1.1 of the video feed of the subject area. Here, the first frame is obtained from a video feed 1.1 of the stationary overview cameras 21, 22, and provides an overview of the subject area, for example, a real-world scene. It will be noticed that the first frame can alternatively be obtained from a mechanically and optically movable camera 23, such as by a system as shown in Figures 2 and 3. A first frame 2.1 for the edited video feed is obtained from the first frame 1.1 of the video feed. Here, the first frame 2.1 for the edited video feed is taken as a subframe of the first frame 1.1 of the video feed, for example, a perspective-adjusted crop. Thus, the first frame 2.1 for the edited video is in this example associated with a certain coordinate within the first one. Petition 870250101914, dated 06 / 11 / 2025, pp. 49 / 63 42 / 51 frame 1.1 of the video transmission, for example, a virtual pan, virtual tilt, virtual zoom configuration. The first frame 2.1 of the edited video may alternatively be associated with a mechanical pan, mechanical tilt, optical zoom configuration, associated with a mechanically and optically movable camera 23. The first frame may be associated with navigation parameter 5.1, which was generated by the generator module 40.
[00095] In the example in Figure 4A, the first frame 2.1 for the edited video transmission is extended by a margin of 3.1, thus obtaining an extended first frame 2.1' for the edited video. The extended first frame 2.1' is larger than the first frame 2.1 by the margin of 3.1. Here, the margin of 3.1 is asymmetrical with respect to the first frame 2.1 for the edited video. In particular, here, the margin of 3.1 is wider below the first frame 2.1 than above the first frame 2.1. Thus, with respect to the first frame 2.1, additional visual information may be present in the extended first frame 2.1', which can be used to generate an appropriate navigation parameter. Here, more information was included by the margin of 3.1 below the first frame 2.1 than above the first frame 2.1, which may be more data-efficient for certain applications, since useful visual information is generally expected to be below the first frame 2.1. Margin 3.Preferably, the first extended frame 2.1' is smaller than the first frame 1.1 of the thematic area video.
[00096] A marker is assigned to the first extended frame data points 2.1' to allow automated differentiation between data points of Petition 870250101914, dated 06 / 11 / 2025, pp. 50 / 63 43 / 51 first extended frame 2.1' that form margin 3.1, and data points that form the first frame 2.1 for the edited video that is intended to be viewed by viewers. The first extended frame 2.1' is entered into the generator module 40, and a navigation parameter 5.2 is output. The navigation parameter 5.2 in this example indicates a relative adjustment from the first frame 2.1 of the edited video stream to obtain a second frame 2.2 for the edited video stream. Here, the navigation parameter 5.2 indicates a relative virtual tilt adjustment in the upward direction, a relative virtual pan adjustment in the right direction, and a relative virtual zoom-out adjustment, such as a ratio. However, it will be understood that the navigation parameter can also be indicative of a global coordinate, for example, with respect to the first or second frame 1.1, 1.2 of the video stream of the subject area.
[00097] The second frame 2.2 for the edited video transmission can be obtained from a second frame 1.2 of the video transmission of the thematic area, for example, by the processing module 30, as if it were virtually adjusting a camera configuration according to the determined navigation parameter 5.2, shown schematically in figure 4B. For comparison, the first frame 2.1 is shown in figure 4B by the dashed lines.
[00098] The second frame 2.2 for the edited video, similar to the first frame 2.1, can be a cut from a second frame 1.2 of the video stream of the thematic area. The second frame 2.2 for the edited video, in turn, can be appended with a margin 3.2 to obtain Petition 870250101914, dated 06 / 11 / 2025, pages 51 / 63 44 / 51 a second extended frame 2.2', which can be used to determine an additional navigation parameter and a third frame for the transmission of edited video, etc. The second extended frame 2.2', for example, can be resampled to a predetermined resolution before being entered into the generator module 40.
[00099] Frames for the edited video stream can be determined based on previous frames of the edited video stream. Thus, a next frame for the edited video stream can be determined based on a current frame and / or one or more previous frames of the edited video. Frames for the edited video stream can also be determined based on subsequent frames of the edited video stream, for example, by temporarily dampening frames for the edited video stream. Thus, a next frame for the edited video stream can be determined based on one or more future frames of the edited video that were temporarily stored. Frames for the edited video stream can also be determined based on a combination of previous and subsequent frames of the edited video stream. [000100] The generator module 40 in these examples comprises a machine learning model, such as an end-to-end deep learning convolutional neural network, and / or a transformer architecture, arranged to receive a frame for the transmission of edited video, for example, a digital image, as an input, and based on it output a navigation parameter. However, other machine learning models can also be used, for example, such as Petition 870250101914, dated 06 / 11 / 2025, pages 52 / 63 45 / 51 support vector machines, decision tree-based learning systems, random forests, regression models, autoencoding clustering, nearest neighbor machine learning algorithm, etc. In some examples, an alternative regression model may be used instead of an artificial neural network. [000101] Deep learning in a neural network environment can include multiple interconnected nodes referred to as neurons. Input neurons, activated from an external source, activate other neurons based on connections with these other neurons that are governed by neural network parameters. A neural network can behave in a certain way based on its own parameters. Training a deep learning model refines the model's parameters, representing the connections between neurons in the network, such that the neural network behaves in a desired way (better at the task for which it is intended, for example, classifying components in a material stream). [000102] Deep learning operates on the understanding that many datasets include a hierarchy of functionalities – from low-level functionalities (e.g., edges) to high-level functionalities (e.g., patterns, objects, etc.). While examining an image, for example, a model begins searching for edges that form motifs that form parts, which form the object being considered. Observable functionalities learned include objects and quantifiable regularities learned by the machine learning model. A machine learning model provided with a large, well-sorted dataset is well-suited to this. Petition 870250101914, dated 06 / 11 / 2025, pages 53 / 63 46 / 51 equipped to distinguish and extract the functionalities relevant to the successful classification of new data. [000103] Optionally, the machine learning model uses a view transformer (ViT) architecture. A view transformer can split an input image into a series of packets, serialize each packet into a vector, and map the vector down to a lower dimension, for example, by matrix multiplication. These vector embeddings can then be processed by a transformer encoder. [000104] Optionally, the machine learning model uses a convolutional neural network. In some examples, deep learning can use neural network segmentation to locate and identify observable learned features in the data. Each filter or layer of the convolutional neural network architecture can transform the input data to increase the selectivity and robustness (of functionality) of the model for the data. This abstraction of the data allows the machine to focus on the features in the data it is trying to classify and ignore irrelevant background information. Deep learning machine models using convolutional neural networks can be used for image analysis. [000105] The machine learning model is trained based on a first dataset. In this example, the first dataset is obtained from a human-selected dataset of frames for a human-selected edited video. Alternatively or additionally, the first dataset can be obtained from a machine-selected dataset, for example, from a model. Petition 870250101914, dated 06 / 11 / 2025, pages 54 / 63 47 / 51 of auxiliary machine learning. The human-selected dataset of frames is obtained here by providing a video of the subject area, for example, by cameras 21, 22, and having a person navigate the subject area within said video, for example, by virtual tilt, virtual pan and virtual zoom within the subject area video, such as using a human-controllable input device, such as a controller or similar, to appropriately represent the subject area. The human-curated edited video can be recorded as a human-selected dataset of frames. The human can navigate the entire subject area video, or only a selection subset thereof, for example, only key moments in the video. The human-curated edited video, for example, may include frames interspersed between these key moments.A human-selected dataset of navigation parameters can also be recorded (and optionally interpolated), corresponding with the navigation applied by the human, for example, virtual tilt, virtual pan, and virtual zoom, to appropriately represent interesting parts of the subject area over time. [000106] The human-selected dataset of navigation parameters is augmented with a dataset of augmentations. The augmentations in this example include pan augmentations, tilt augmentations, and zoom augmentations. Thus, here, each navigation parameter includes a pan, tilt, and zoom setting, where each navigation parameter in the navigation parameter dataset is shifted by a pan augmentation, a tilt augmentation, and a zoom augmentation. Petition 870250101914, dated 06 / 11 / 2025, pages 55 / 63 48 / 51 The augmentations can, for example, be generated randomly, such as drawn from a normal distribution. The augmented frame dataset translates into an augmented edited video that differs in a known way from the human-curated edited video by the applied augmentation. The augmented edited video, that is, a frame dataset, is tagged with the augmentation dataset, creating the first dataset to train the machine learning model. [000107] The machine learning model can be additionally or alternatively trained based on a second dataset. The second dataset can be obtained by producing edited video from the subject area video. The edited video can be edited in various ways. In this example, the edited video is produced with the machine learning model that has been pre-trained, for example, using the first dataset. This provides a particularly efficient training method. Differences between the edited video produced by the method, i.e., a dataset of frames generated by the model, and the human-curated edited video, i.e., a human-selected dataset of frames, can be determined.Each frame in the model-generated frame dataset, for example, can be compared to a corresponding frame in the human-selected frame dataset, and a difference between them, for example, a relative tilt value, a relative pan value, and a relative zoom value, can be stored in the difference dataset. The model-generated frame dataset is tagged with the difference dataset, creating the second dataset. Petition 870250101914, dated 06 / 11 / 2025, pp. 56 / 63 49 / 51 of data to train the machine learning model. Additional generations of datasets can be generated by the pre-trained machine learning model as appropriate. Datasets from different generations can be combined to avoid unwanted biases in the training data. [000108] It will be understood that the methods described herein may include computer-implemented steps. All the steps mentioned above may be computer-implemented steps. Embodiments may comprise computer apparatus, in which processes are performed on the computer apparatus. The invention also extends to computer programs, particularly computer programs on or in a carrier, adapted to put the invention into practice. The program may be in the form of source or object code or any other form suitable for use in implementing the processes according to the invention. The carrier may be any entity or device capable of carrying the program. For example, the carrier may comprise a storage medium, such as a ROM, for example, a semiconductor ROM or hard disk.Additionally, the carrier can be a transmissible carrier such as an electrical or optical signal that can be transported via electrical or optical cable or by radio or other means, for example, via the internet or the cloud. [000109] Some embodiments can be implemented, for example, using a machine or tangible computer-readable medium or article that can store an instruction or a set of instructions that, if executed by a machine, can cause the machine to perform a method and / or operations according to the embodiments. Petition 870250101914, dated 06 / 11 / 2025, pp. 57 / 63 50 / 51 [000110] Various modalities can be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include processors, microprocessors, circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, microchips, chip assemblies, etc.Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, mobile apps, middleware, firmware, software modules, routines, subroutines, functions, computer-implemented methods, procedures, software interfaces, application program interfaces (APIs), methods, instruction sets, computer code, computer code, etc. [000111] Here, the invention is described with reference to specific examples of embodiments of the invention. However, it will be evident that various modifications and alterations can be made herein without departing from the essence of the invention. For the purpose of clarity and a concise description, functionalities are described herein as part of the same embodiments or of separate embodiments; however, alternative embodiments having combinations of all or some of the functionalities described in these separate embodiments are also contemplated. [000112] However, other modifications, variations, and alternatives are also possible. The specifications, drawings, and examples should be provided appropriately. Petition 870250101914, dated 06 / 11 / 2025, pp. 58 / 63 51 / 51 considered in an illustrative rather than a restrictive sense. [000113] In claims, any reference marks placed in parentheses should not be interpreted as limiting the claim. The word 'including' does not exclude the presence of other functionalities or steps different from those listed in a claim. Additionally, the words 'a' and 'an' should not be interpreted as limited to only one, but should be used to mean 'at least one', and do not exclude a plurality. The mere fact that certain measures are cited in mutually different claims does not indicate that a combination of these measures cannot be used to an advantage. Petition 870250101914, dated 06 / 11 / 2025, pp. 59 / 63
Claims
1 / 10 CLAIMS 1. A computer-implemented method for autonomously producing an edited video stream of a subject area, characterized in that it comprises: obtaining a first frame for the edited video stream of a first sub-area of the subject area; inserting the first frame into a machine learning model; determining, by the machine learning model, based on the first frame, a navigation parameter related to a second sub-area of the subject area; and obtaining a second frame for the edited video stream of the second sub-area based on the determined navigation parameter.
2. A method according to claim 1, characterized in that the first frame has a first frame configuration associated with it and the second frame has a second frame configuration associated with it, and wherein the navigation parameter is representative of a relative or absolute adjustment of the first frame configuration to the second frame configuration.
3. A method according to claim 1 or 2, characterized in that the navigation parameter includes one or more of a pan adjustment, a tilt adjustment, and a zoom adjustment.
4. A method, according to any of the preceding claims, characterized in that the first frame and / or the second frame of the edited video stream are obtained by capturing a subframe of a frame from a video stream of at least part of the subject area. Petition 870250087343, dated 09 / 26 / 2025, p. 21 / 33 2 / 10 5. A method, according to any of the preceding claims, characterized in that the navigation parameter includes one or more virtual pan adjustments, virtual tilt adjustments, and virtual zoom adjustments.
6. A method, according to any of the preceding claims, characterized in that the first frame and / or the second frame of the edited video stream are obtained by a camera that is automatically adjustable mechanically and / or optically.
7. A method, according to any of the preceding claims, characterized in that the navigation parameter includes one or more mechanical pan adjustments, mechanical tilt adjustments, and optical zoom adjustments.
8. A method, according to any of the preceding claims, characterized in that the navigation parameter includes an adjustment speed and an adjustment direction.
9. A method, according to any of the preceding claims, characterized in that the first frame is obtained at a first instant of time and the navigation parameter is determined at the first instant of time and belongs to a second instant of time subsequent to the first instant of time.
10. Method, according to claim 9, characterized in that it comprises, at the first instant of time, determining, using the machine learning model and based on the first frame, a temporal sequence of navigation parameters belonging to multiple instants of time subsequent to the first instant of time, the temporal sequence of navigation parameters including in particular the navigation parameter.
11. A method according to claim 10, characterized in that it comprises determining an adjusted navigation parameter based on the navigation parameter and, additionally, based on at least one additional navigation parameter belonging to the second time instant that was determined at a time instant prior to the first time instant.
12. Method, according to claim 11, characterized in that the second frame for the edited video stream of the second subarea is obtained based on the adjusted navigation parameter.
13. A method according to claim 11 or 12, characterized in that the adjusted navigation parameter is determined as a weighted average of a plurality of navigation parameters, each belonging to the second time instant and which were determined at a time instant prior to the second time instant, particularly in that the navigation parameters of the plurality of navigation parameters closer in time to the second time instant are given more weight than the navigation parameters of the plurality of navigation parameters farther in time from the second time instant.
14. Method, according to any one of claims 9 to 13, characterized in that the first time instant and the second time instant are spaced in time by a period of time determined to take into account the latency of the method, particularly the delay between the capture of the first frame and the completion of a physical adjustment of the camera that corresponds to the navigation parameter.
15. A method, according to any of the preceding claims, characterized in that the navigation parameter is determined as a normalized navigation parameter, which is normalized with respect to a field of view of the first frame.
16. Method, according to claim 15, characterized in that it comprises denormalizing the normalized navigation parameter and obtaining the second frame for the edited video stream from the second subarea based on the denormalized navigation parameter.
17. A method, according to any of the preceding claims, characterized in that it comprises obtaining a first extended frame of a first extended subarea of the study area that is larger than the first subarea by a predetermined margin and determining the navigation parameter based on the first extended frame.
18. Method, according to claim 17, characterized in that the margin is asymmetrical with respect to the first subarea.
19. Method, according to claim 17 or 18, characterized in that it comprises labeling the data points of the first extended frame to distinguish between data points of the first extended frame that correspond to data points of the first frame and data points of the first extended frame that correspond to margin data points.
20. Method, according to any of the preceding claims, characterized in that Petition 870250087343, dated 09 / 26 / 2025, page 24 / 33 5 / 10 comprises resampling the first frame, for example, extended, and determining the navigation parameter based on the first resampled frame, for example, extended.
21. Method according to claim 20, characterized in that the resampling is such that a non-uniform sampling across the frame is obtained.
22. A method, according to any of the preceding claims, characterized in that the navigation parameter is determined based on a set of frames from the edited video stream that contains only the first frame.
23. A method, according to any one of claims 1 to 21, characterized in that the navigation parameter is determined based on a set of frames from the edited video stream that contains at least two frames.
24. Method according to claim 23, characterized in that the set of frames of the edited video stream includes frames that precede the second frame in time.
25. Method, according to claim 23 or 24, characterized in that the set of frames of the edited video stream includes frames that succeed the second frame in time.
26. A method, according to any of the preceding claims, characterized in that it comprises, after determining the navigation parameter, adjusting the navigation parameter according to a smoothing criterion and obtaining the second frame for the edited video stream based on the adjusted navigation parameter. Petition 870250087343, dated 09 / 26 / 2025, p. 25 / 33 6 / 10 27. Method according to claim 26, characterized in that the determined navigation parameter is adjusted based on a set of navigation parameters that includes navigation parameters that precede the determined navigation parameter in terms of time and / or navigation parameters that succeed the determined navigation parameter in terms of time.
28. Method according to claim 27, characterized in that it comprises determining a navigation trend based on said set of navigation parameters and adjusting the navigation parameter according to the determined navigation trend.
29. A method, according to any of the preceding claims, characterized in that the machine learning model includes an end-to-end artificial neural network.
30. A computer-implemented method for generating a dataset, particularly a training dataset for a machine learning model for use in the method, according to any of the preceding claims, characterized in that it comprises: providing a video of a subject area; having a human model and / or an auxiliary machine learning model navigating the subject area within the video; and obtaining a curated dataset of navigation parameters from it, and / or obtaining a curated dataset of frames from it.
31. Method, according to claim 30, characterized in that it comprises Petition 870250087343, dated 09 / 26 / 2025, page 26 / 33 7 / 10 causing the trained auxiliary machine learning model to detect the presence of a predetermined entity of interest in one or more frames of the video; causing the trained auxiliary machine learning model to determine a location of the detected entity of interest within one or more frames of the video; and obtaining the curated dataset of navigation parameters from it.
32. A method according to claim 30 or 31, characterized in that it comprises: augmenting the curated dataset of navigation parameters by an augmentation dataset; generating a frame dataset corresponding to the augmented navigation parameter dataset; and labeling the frame dataset with the augmentation dataset.
33. A method according to claim 30, 31 or 32, characterized in that it comprises: generating a frame dataset for the edited video from the video, for example, according to a method as defined in any one of claims 1 to 2; obtaining a difference dataset representative of a difference between the generated frame dataset and the curated frame dataset; and labeling the frame dataset with the difference dataset.
34. Computer-implemented method for training a machine learning model for autonomous production of a video stream edited from a video stream of a thematic area, characterized in that it uses a dataset obtained by the method, as defined in any one of claims 30 to 33.
35. Method according to claim 34, characterized in that the machine learning model is configured to emit a temporal sequence of navigation parameters.
36. Method, according to claim 35, characterized in that the machine learning model is configured to emit the temporal sequence of navigation parameters in accordance with a predefined progression characteristic.
37. A method according to any one of claims 34 to 36, characterized in that it comprises: generating a first training dataset, for example, according to the method in claim 32 or 33, and pretraining the machine learning model using the first dataset; causing the pretrained machine learning model to generate a second training dataset, for example, according to the method in claim 33, and training the pretrained machine learning model using the second dataset.
38. Method according to claim 37, characterized in that it comprises: after training the machine learning model using the second dataset, causing the machine learning model to generate a third dataset, for example, according to the method of claim 33; Petition 870250087343, dated 09 / 26 / 2025, p. 28 / 33 9 / 10 training the machine learning model using at least part of the third dataset.
39. Machine learning model for autonomous production of an edited video stream of a thematic area, characterized in that it is trained according to a method as defined in any one of claims 34 to 38.
40. Edited video characterized in that it is obtained by a method as defined in the preceding claim.
41. A non-transient, computer-readable medium characterized in that it stores instructions which, when executed by one or more processors, cause a device to execute the method as defined in any of the preceding claims.
42. System for autonomous production of an edited video stream of a subject area, characterized in that it comprises: a camera system; and a processing unit, operatively connected to the camera and configured to execute a method, as defined in any one of claims 1 to 38.
43. System according to claim 42, characterized in that the camera system includes a stationary panoramic camera arranged to record a panoramic video stream of the subject area from a stationary viewing direction.
44. System, according to claim 42 or 43, characterized in that the camera system includes a mechanically and / or optically adjustable camera arranged to record a video stream of the subject area from an adjustable viewing direction.
45. System, according to claim 44, characterized in that it comprises a camera controller operatively disposed between the processing unit and the mechanically and / or optically adjustable camera, the camera controller being disposed to transmit one or more video frames acquired by the camera to the processing unit, receiving a navigation parameter from the processing unit and, based on this, mechanically and / or optically adjusting the camera. Petition 870250087343, dated 09 / 26 / 2025, pp. 30 / 33