Camera Control
Patent Information
- Application Number
- JP2024549684
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-25
- Filing Date
- 2023-02-27
- Publication Date
- 2026-03-05
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to remote control of a camera. [Background technology]
[0002] Cameras for capturing video from live events, such as sporting events, are conventionally mounted on heads that allow the camera to tilt and pan. Tilting involves rotating the camera about a roughly horizontal axis to raise or lower the camera's field of view. Panning involves rotating the camera about a roughly vertical axis to move the camera's field of view left or right. Such cameras conventionally have variable zoom lenses. By adjusting the zoom of the lens, the camera's field of view can be narrowed or widened. As an object moves in front of the camera, it is desirable to control the pan, tilt, and zoom of the camera to capture the best view of the object. Typically, this is done by an operator positioned next to the camera.
[0003] The camera may be equipped with a motorized head that allows its pan and tilt to be controlled remotely. Similarly, the camera may be equipped with a motorized zoom lens that allows the zoom of the camera to be controlled remotely. In principle, this type of equipment can avoid the need for an operator to be present at the camera's location. This allows the camera to save considerable travel costs when the camera is located far from the studio's base, and allows the camera to be located in places that are too dangerous for people to be present (e.g., motor racing circuits). With this in mind, an efficient way for a studio to capture video of a live event at a remote location would be to transport a large number of remotely controllable cameras to the location, arrange a video feed from this location to the studio's central production facility, and then remotely control the cameras from the central facility. However, this arrangement is problematic in that the time delay of the signals between the central facility and the remote location is too long to effectively control the remote cameras. There is a first delay in the video feed transmitted between the camera and the production facility. There is a second delay as the signal travels to the camera after the operator has seen the feed and generated a control signal to the camera. If the camera is filming a fast-moving subject, such as a downhill skier or a racing car, these overall delays can result in the camera not being able to react fast enough to the subject's movement, causing the subject to disappear from the camera's field of view, resulting in a poorly captured video sequence. Also, if the camera is remotely controlled, once the subject disappears from the camera's field of view, it can be difficult for the camera operator to recapture the subject, as they cannot see the subject directly.
[0004] There is a need for improved methods of remotely controlling cameras. Summary of the Invention
[0005] According to one aspect, a method of creating a processed video stream is provided, the method including capturing a first portion of an input video stream using a camera, storing the first portion of the input video stream at a first bandwidth, transmitting the first portion of the input video stream from the camera to a remote processing facility at a second bandwidth lower than the first bandwidth, designating a first sub-region of the first portion of the input video stream for further processing at the processing facility, creating a first cropped video stream by cropping the stored first portion of the input video stream to the sub-region in response to the designation of the first sub-region, and creating a processed video stream incorporating the first cropped video stream.
[0006] The step of specifying the first sub-region may be performed simultaneously with the step of capturing the input video stream.
[0007] The processed video stream may be a live broadcast stream depicting a live event.
[0008] Capturing the input video stream may include capturing a video stream of a live event.
[0009] The camera may be adjustable to vary its field of view. The method may include generating a camera control signal in response to the designation of the first sub-region, and adjusting the field of view of the camera in response to the camera control signal.
[0010] After the step of adjusting the field of view of the camera, the following may be performed: capturing a second portion of the input video stream using the camera; transmitting the second portion of the input video stream to processing equipment; designating a second sub-region of the second portion of the input video stream for further processing at the processing equipment; creating a second cropped video stream by cropping the second portion of the input video stream to the sub-region in response to the designation of the second sub-region; and creating a processed video stream incorporating the second cropped video stream.
[0011] The method may include transmitting a camera control signal to a camera device including a camera, the camera device configured to automatically adjust a field of view of the camera in response to the camera control signal.
[0012] The step of generating a camera control signal may include automatically analyzing a position and / or size of the first sub-region relative to the entire field of the first portion of the input video stream, and automatically applying a predetermined algorithm in response to the determination to generate the camera control signal.
[0013] The camera control signal may be a signal that causes the field of view of the camera to be changed such that the location of the center of the first sub-region in the first portion of the input video stream is closer to the center of the field of view of the camera.
[0014] The step of adjusting the camera may include adjusting the pan or tilt of the camera or translating the camera, which may include rotating and / or translating the camera.
[0015] The camera control signal may be a signal that causes a camera's field of view to change the size of a region of the size of a first sub-region in a first portion of the input video stream to approach a predetermined target size.
[0016] The step of adjusting the camera may include adjusting the zoom of the camera.
[0017] The method may include estimating a responsiveness of the camera to a previous camera control signal, and generating a camera control signal in response to the estimated responsiveness.
[0018] The step of designating a first sub-region of the first portion of the input video stream for further processing may be performed by a human designating the first sub-region, and the method may include displaying a boundary of the first sub-region to a user on a display.
[0019] The step of designating a first sub-region of the first portion of the input video stream for further processing may be performed automatically by analyzing the first portion of the input video stream to identify an object of interest therein and designating the first sub-region to surround the object.
[0020] The processed video stream may have a lower resolution than the input video stream.
[0021] According to a second aspect, there is provided a method of creating a processed video stream, the method comprising: receiving at a processing facility a first portion of an input video stream captured using a camera; designating at the processing facility a first sub-region of the first portion of the input video stream for further processing; in response to the designation of the first sub-region, (i) creating a first cropped video stream by cropping the first portion of the input video stream to the sub-region; and (ii) generating a camera control signal, creating a processed video stream incorporating the first cropped video stream; and transmitting the camera control signal to a camera to adjust a field of view of the camera.
[0022] According to a third aspect, there is provided a use of the above method for reducing communication delays between a camera and a production facility.
[0023] According to a fourth aspect, there is provided a video processing device including an input for receiving a captured video stream, a memory, and a controller configured to: store in the memory the captured video stream at a first bandwidth, compress the captured video stream to create a compressed video stream having a second bandwidth lower than the first bandwidth, transmit the compressed video stream in a form such that points are identifiable during a progression of the compressed video stream, receive a designation of a sub-frame region and an associated point during the progression of the compressed video stream, and create an output video stream by cropping the stored video stream to the designated region at a point in the stored video stream that corresponds to the designated point.
[0024] The controller may be configured to generate output signals for controlling a camera orientation and / or zoom in response to one or more of: (i) a position of the specified sub-frame region relative to the overall frame; and (ii) a size of the specified sub-frame region relative to the overall frame.
[0025] The controller may be configured to perform video processing at a point in the video stream during a period between saving the point in the video stream and cropping the point in the video stream.
[0026] Video processing may include performing processing to improve the visual quality of that portion of the video stream.
[0027] The controller and camera may be configured such that, in response to a command indicating movement of a designated sub-region in a certain direction, (a) the controller moves the position of the sub-region selected for output in that direction, and (b) the camera adjusts the camera's field of view in that direction. Depending on how the sub-regions are designated to the controller, one of these movements may cancel out the other.
[0028] The low bandwidth video and the high bandwidth video may have a common timestamp, and when a command is sent to the controller, the command may be time-stamped with the time of the low bandwidth video to which the command pertains, and the sub-region specified by the command may be selected in the high bandwidth video of the corresponding timestamp.
[0029] The camera may have a motorized head that allows the camera to pan and / or tilt.The camera may have a motorized lens unit that allows the zoom to be adjusted.
[0030] The invention will now be described by way of example with reference to the drawings, in which: [Brief description of the drawings]
[0031] [Figure 1] Camera equipment and production facilities are shown. [Diagram 2] The control station is shown. [Diagram 3] A process for manipulating the image and providing camera control signals is described. [Figure 4] 1 shows a system for remotely controlled image cropping. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0032] FIG. 1 shows camera equipment, generally designated "1," and production facilities, generally designated "2." The production facilities are at a remote location away from the camera equipment. They are connected by a communications network, generally designated "3." This may be a publicly accessible network such as the Internet. In reality, the camera equipment is at the scene of a live event, for example, a sporting event, an art performance, or a news event. The event may take place miles or even thousands of miles away from the production facilities.
[0033] The camera arrangement includes a camera 10. The camera is capable of capturing video. The camera has a variable zoom lens 11 with a motor 12 that allows the lens to be zoomed. The camera is mounted on a pan and tilt head 13. The pan and tilt head allows the direction in which the camera is pointed to be adjusted. The pan and tilt head is provided with motors 14, 15 for adjusting the pan and tilt. The camera can also be mounted on other types of motion devices, for example a guide track, an articulated support arm, or a drone. These motion devices may be capable of translating the camera in one or more axes. For example, the camera may be capable of translating the camera in a first axis that is horizontal or has a horizontal component and / or a second axis that is vertical or has a vertical component. The first and second axes may be orthogonal. Each motion device may adjust the camera's view by adjusting the camera's position. Each motion device may be actuated by one or more motors, linear actuators, hydraulic actuators, pneumatic actuators, propellers, etc. to adjust the camera's position. The camera is attached to a tripod 16 or other mounting mechanism.
[0034] The cameras are connected to an on-site camera controller 17 by a communication cable or a wireless data link. The on-site camera controller receives the video signals from the cameras. The on-site camera controller includes a buffer memory 70, a video processor 71, and a program memory 72. The buffer memory stores the video data received from the cameras. The video processor executes code that processes the video data. The program memory stores the code executable by the video processor in a non-transient form. The video processor also controls the transmission of the video data to the production facility. The video data from the cameras is received by the camera controller at a particular resolution, hereafter referred to as full resolution. Full resolution may have data representing a capture of each pixel of the camera's image sensor, or full resolution data may have been compressed somewhat from the raw data captured by the camera. One function the video processor can perform is to compress the captured video data. Compressing the video data results in a video stream that represents the same scene but with reduced bandwidth. This can be the result of, for example, pixel culling, combining multiple pixels into tiles that represent larger portions of the image, or applying algorithms that use the same data to represent similar portions of the image. The reduced bandwidth representation of the full resolution video stream is called the compressed stream. The full resolution stream, or a high resolution stream close to full resolution, is stored in a buffer memory 70. In one mode of operation, the video processor transmits the full / high resolution stream to the production facility 2. In another mode of operation, the video processor transmits the compressed stream to the production facility 2 instead. The camera controller also receives motion control signals from the production facility to control the operation of the camera and zoom signals to control the zoom of the camera, and sends these signals to the appropriate motors and other motion units to control the camera operation. The on-site camera controller serves as a communications interface between the network 3 and the cameras.The local camera controller is connected to the network 3 by a cable (e.g., an Ethernet cable) or a wireless link (e.g., a cellular data link). The camera controller may be integrated into the camera or may be a separate unit.
[0035] The production facility 2 includes a production control unit 20, an image subset control terminal 21, and a production control terminal 22. The production control unit manages the video feeds received from the cameras 10 and possibly other cameras, and generates control signals to the cameras. The image subset control terminal 21 allows a user to select portions of the images received from the cameras. The production control terminal 22 allows a user to combine video feeds received from multiple cameras, computer generated sources, and stored video data to create a video output stream, indicated generally at "23".
[0036] The production control unit is coupled to network 3 for receiving video from camera 10 and possibly other cameras as well. The production control unit includes a processor 24 and a memory 25. The memory stores in a non-transitory form program code executable by the processor to cause the processor to perform the functions of the production control unit as described herein. The production control unit is also coupled to terminals 21, 22.
[0037] The image subset control terminal 21 is shown in more detail in Figure 2. The image subset control terminal 21 includes a display 40, a position user interface input 41, a zoom user interface input 42, and a camera selection user interface input 43.
[0038] An important function of the terminal 21 is to allow the user of the terminal to select a sub-region of the video stream generated by the camera. The camera may generate a video stream with a higher resolution than is required for the output stream 23. In this case, the number of pixels in the height and / or width of the image captured by the camera is greater than the number of pixels in each dimension of the image transmitted in the output stream. As an example, the camera may generate video in 4K resolution (3840×2160 pixels or 4096×2160 pixels), while the resolution required for the output stream is 1080i resolution (1920×1080 pixels). To create the output stream, the video from the camera is downscaled and / or cropped. In one example, the entire field of view of the camera's video stream can be downscaled to the output resolution. In another example, a portion of the camera's video stream, which has the same resolution as the output stream, is cropped from the camera stream. In another example, a portion of the camera's video stream, which is less than the entire field of view of the camera's video stream but greater than the output video stream, is cropped and downscaled from the camera stream.
[0039] Figure 2 shows the terminal 21 displaying frames of a video stream received from a camera. The video stream is transmitted to the production control unit (controller) 20 and from there to the terminal 21 as described above. The camera video stream includes a subject 44, in this case a skier. The display 40 of the terminal 21 displays a bounding box 45. The bounding box 45 encloses (delineates) a sub-region of the camera video stream that can potentially be included in the output video stream and should be selected. The bounding box has the same ratio, e.g. pixel aspect ratio, as the output stream. A user of the terminal can change the size of the bounding box using the zoom input 42 by scaling it larger or smaller relative to the boundaries of the camera stream. A user can change the position of the bounding box using the position input 41 by moving it left, right, up or down relative to the boundaries of the camera stream. In this way, the user can select a sub-region of the camera stream by delineating it with a bounding box. The part of the camera image that is within the bounding box is considered to be selected.
[0040] In Figure 2, the bounding box is indicated by an outline. The area specified by the boundary may be indicated in other ways, such as by a highlighted area of the display (e.g., as an area having a higher brightness or contrast than the rest of the display, or as an area having a color cast).
[0041] Once a bounding box is specified by the user at terminal 21 while on the video stream from the camera, terminal 21 sends the size and position of the bounding box to production control unit 20. This can be done simply by sending the pixel positions of the two diagonal corners of the bounding box in the camera stream, or in any other suitable manner. Production control unit 20 then processes the camera stream, downscaling and / or cropping it to create an intermediate video stream that represents only the portion of the camera stream delineated by the bounding box and has the resolution of the intended output stream. For example, if the camera stream has a resolution of 4096x2160 while the output stream has a resolution of 1920x1080, and the pixel locations of the two diagonal corners of the currently specified bounding box relative to the camera stream are [400,900] and [2511,2087], production control unit 20 creates an intermediate video stream by cropping the camera stream to a rectangular 2112x1188 window with [400,900] as one corner and downscaling the rectangle to 1920x1080. The selected rectangle could also be upscaled rather than downscaled, but this would result in a lower quality output image. The output stream is derived from this intermediate image. Alternatively, the output stream can be created from one or more selected intermediate streams, as described below.
[0042] In practice, there are several cameras at the location whose streams, together with other video feeds such as computer-generated images and overlays, must be combined and staggered to create an output video stream. In the case of several cameras 10 at remote locations, it is advantageous to have one terminal 21 for each of them. In this case, the boundaries of the desired area of the video can be simply specified in real time for each camera, and the respective intermediate streams are created accordingly. Alternatively, it is also possible to use one terminal to specify the areas of interest of several streams simultaneously. The areas of interest of the streams can be specified manually as described above, or automatically by image recognition software, which is arranged to specify areas with certain characteristics, for example, relatively high contrast or an appearance similar to a certain object, such as a person.
[0043] The terminal 22 receives the available video (preferably intermediate streams) and other inputs and provides a user interface that lets the user select which video should be used to create the output image. For example, the terminal 22 allows the user to switch between different cameras or intermediate feeds as a subject moves from one camera's field of view to another. The production control unit 20 provides the available video (preferably intermediate streams) and other inputs to the terminal 22. The user of the terminal 22 indicates, using a user input device (e.g., keyboard, touch screen, or pointing device) at the terminal 22, which of the available content should create the output stream. This indication is sent to the production control unit 20. The production control unit 20 then creates the output stream 23 by selecting the appropriate content.
[0044] One or both of the terminals 21, 22 may be implemented automatically, using algorithms to select the desired parts of the video stream or to select the desired content stream, in which case they may be integrated into the production control unit 20.
[0045] Thus, in a first mode of operation, a video stream is transmitted from the camera's locale to a production facility, which may be remotely located away from the camera, where a region of the video stream can be selected for further transmission or storage, and depending on where that region is relative to the boundaries of the video stream, the camera can be automatically controlled to pan to bring objects in that region closer to the center of the captured frame.
[0046] Figure 3 shows some of the process steps when creating an output image using this mode of operation. The video feed from the camera is received at step 50 and passed to the image subset control terminal 21. At the terminal 21, the desired region of the video stream is selected (step 51). This step can be done manually or automatically. The terminal 21 sends the position and size of a bounding box to the production control unit 20 (step 52). The production control unit 20 crops and / or downscales the selected region of the camera stream to create an intermediate stream at the desired output resolution (step 53). This is passed to the terminal 22 for selection of the stream to be incorporated into the output stream (step 54). In practice there may be multiple output streams representing each part of the race, each participant, each viewpoint etc.
[0047] The size and position of the bounding box is an indication of the area of the camera stream that is of most interest. In the present system, that information is used to trigger camera movement and / or zoom. This can avoid the need for separate control of the camera, resulting in the camera remaining automatically focused on the area of interest. This can be done in a way that reduces the latency between the camera and where you control it. The mechanism for this is now described.
[0048] As shown in Fig. 3, the size and position of the bounding box (window) are passed as input to step 56, which determines the deviation of said size and position from a predefined reference (step 55). Depending on the deviation, control signals for movement (e.g. pan and tilt, or movement of an arm or camera dolly) and / or zoom are generated (step 57). Steps 56 and 57 may be performed in the production control unit 20. These control signals are sent to the interface 17 and used to control the movement and / or zoom of the camera. This signal may be sent to the camera via the same link used to send the video from the camera, or via a separate link 60, either of which may or may not go through the network 3, respectively. In this way, the movement and / or zoom of the camera is remotely controlled in response to the selection of an area of the video stream from the camera designated for further processing. This avoids the need to separately select a portion of the captured video stream of interest to control the camera. If the camera control signals are generated in a way that results in maintaining a position corresponding to the designated area away from the edge of the captured video stream, then there is room to move or zoom the designated area beyond its current position even when delays in the camera's operation occur, thereby reducing transmission delays between the camera and the location where the video is analyzed.
[0049] Below are some examples of how the camera can be controlled in response to bounding box movement.
[0050] 1. The center of the bounding box is determined. If the center is above a predefined point in the camera image frame, preferably above the center of the camera image frame, the camera is signaled to tilt up or move up. If the center of the bounding box is below the center of the camera image, the camera is signaled to tilt down or move down. If the center of the bounding box is to the left of the center of the camera image, the camera is signaled to pan left or move left. If the center of the bounding box is to the right of the center of the camera image, the camera is signaled to pan right or move right. In each case, the center may be the geometric center, i.e. the intersection of the diagonals of the box or image frame. In this way, the camera is controlled to move its field of view to move the location in the camera field of view that is the center of the bounding box closer to the center of the camera field of view. The speed of movement of the camera field of view depends on the distance from the center of the camera field of view to the center of the bounding box. This speed of movement may be controlled to be faster as the distance increases.
[0051] 2. A predefined preferred size for the bounding box is determined. This can be expressed as a percentage of the size of the camera image frame. This percentage can be, for example, 60%. This percentage is preferably between 80% and 50%. If the percentage is too high, the tolerance for delays during camera control is too low. If the percentage is too low, the bounding box is too small and the output image has to be upscaled from time to time, which reduces the output quality. If the bounding box is larger than the preferred size, the camera will be zoomed out. If the bounding box is smaller than the preferred size, the camera will be zoomed in. Thus, the camera is controlled to scale its field of view to bring the size of the bounding box to the predefined preferred size. The zoom speed depends on the difference between the size of the bounding box and the preferred size. The zoom speed is controlled to be faster the larger this difference is.
[0052] In other words, the camera captures a video stream at a first resolution. This video stream is transmitted to a control location. The control location is remote to the camera. Thus, there is a substantial signal delay in (i) transmitting the video from the camera to the control location and / or (ii) transmitting the control signal from the control location to the camera. The camera is configured such that its field of view is adjusted in response to a signal received from the control location. This includes one or more of: (i) rotating the camera (e.g., panning or tilting); (ii) adjusting the zoom of the camera; (iii) translating the camera, e.g., with a track, articulated arm, or drone. At the control location, a sub-area of the video stream captured by the camera is specified. This specification can be done manually or automatically. The video stream captured by the camera is displayed on a terminal with a user interface of the control center at the same time as it is received at the control center (control unit), thereby allowing a user of the terminal to specify the area of the displayed video. The user interface allows the specified region to be (i) moved vertically and / or horizontally relative to the captured video and / or (ii) resized relative to the full frame size of the captured video. In response to specifying a region of the video, two steps are performed.
[0053] 1. A second video stream is created by cropping the video stream from the camera to a specified region. This second video stream is output for viewing elsewhere.
[0054] 2. A control signal is generated according to the size and / or position of the specified region relative to the full frame of the video, and the control signal is sent to the camera to control the camera, the control signal being generated to move the field of view of the camera such that the object position occupying the specified region (i) approaches the center of the field of view of the camera and / or (ii) approaches a predefined size within the field of view of the camera.
[0055] Taken together, these steps allow the camera to be controlled automatically with reduced perception of lag at the point of control when compared to systems where the operator cannot select subregions of the captured video, and also allow the operator to control the camera's field of view so that it is easier to maintain the subject within the camera's field of view.
[0056] While the operator is designating the designated area, the operator's control station (terminal) may display the entire field of view of the captured video and highlight the area designated by the operator, or it may display only the area designated by the operator. This second approach is beneficial in that it gives the operator the feeling that he is controlling the camera with minimal lag. This arises because short-term adjustments made by the operator can be responded to by adjusting the designated sub-area of the captured video, while longer-term adjustments can be responded to by camera movement. That is, the system can be viewed as operating in two feedback loops, one within the other.
[0057] It may be desirable to automatically adjust the designated region in the opposite direction when the camera's field of view moves or changes size, reducing the likelihood that the operator will perceive the system as overreacting to a control input.
[0058] In these approaches, the camera follows the guide given by the position and size of the bounding box to focus on the areas of the image that are of most interest.
[0059] The camera control signal is generated automatically and / or algorithmically by the processor 24. The camera control signal is generated by a computer configured to generate the signal according to pre-stored instructions for executing an algorithm whereby the signal is generated in response to the position and / or size of a bounding box relative to the video stream captured by the camera. The camera control signal may be generated periodically, for example every frame of the image, every time the bounding box moves, or at a predefined interval, for example every 5 ms or 10 ms.
[0060] The nature of the camera control signals will depend on the interface 17 and the motors or other devices used to control the camera, and may indicate, for example, a target state or commanded movement for each of the camera's pan, tilt, position, and / or zoom.
[0061] The responsiveness of the camera is altered to take into account the signal delay between the camera 1 and the control station (production facility) 2. The speed or magnitude of the camera adjustment is altered according to the signal delay. This helps to avoid the camera reacting too quickly or too slowly to bounding box movement or resizing. In one example, the control station has access to a measurement of the signal delay between the camera device and the control unit 20. This measurement is made by the control station or by the camera device and signaled to the control unit. This timing measurement can be made by any known measurement technique, for example by synchronizing both ends of the data link to a common clock and measuring the signal propagation time with respect to that clock. After knowing the delay, the speed of the camera adjustment in response to a particular deviation of the box center or size from a reference point or size is adjusted according to the delay. If the delay increases, the speed of the adjustment is increased. In an alternative approach, the preferred speed of adjustment can be learned automatically by treating an increase in the frequency of reversals in the bounding box movement or sizing, which indicates an overshoot of the camera movement. The frequency of reversals of bounding box movement (top to bottom or left to right) or sizing (increase to decrease or vice versa) is detected over a period of time, e.g. 10s or 30s. When this frequency is above a first predefined value, the responsiveness of the camera is altered, e.g. by tailoring the control signal to command a smaller rate or magnitude of adjustment. When the frequency is below a second predefined value (less than the first predefined value), the responsiveness of the camera is altered, e.g. by tailoring the control signal to command a larger rate or magnitude of adjustment. The camera control signal is generated such that the rate of change of the camera's field of view is responsive to deviation of the designated area from a predefined position (typically this is centered with respect to the captured video stream) and / or size. The more the designated area deviates from the predefined position and / or size, the more the rate of change is increased. This helps the operator to keep the subject within the camera's field of view.In this way, the responsiveness of the control mechanism is automatically matched to the level of delay through the link. Other control loop specifications can also be used to match the responsiveness through the link.
[0062] The units (terminals) 20, 21, 22 may be combined together in any suitable manner or may be divided into several physical devices, such as terminals and computer servers.
[0063] In the above example, the camera physically moves in response to the selection of different regions in the video stream. In another system, the transmission of relatively high resolution video from a camera can be used to give a remote operator, remote from the camera, the sensation of actually moving the camera. The camera captures video at a higher resolution than the intended output resolution. The video is transmitted from the camera to a remote location, remote from the camera, by an operator at the remote location. The operator is provided with a user interface device that simulates a camera. The device is mounted, for example, on a pan / tilt head, and has a physical type of handle or other device typically used to move a camera. An example is a pan bar. The device has a physical type of zoom control typically used to zoom a camera, for example, a twist grip, slider, or rocker on a pan / tilt handle. The user interface device has a video display. The display is a video display configured to simulate the type typically used as part of a camera so that the operator can see what the camera is capturing. In other words, the operator is provided with a simulated (pseudo) camera. Sensors are provided to sense the positions of various user input devices. A processing device, either integrated into the simulated camera or a separate unit, selects a portion of the video stream received from the remote camera in response to the sensed position of the user input device and displays that portion on the video screen of the operator's device, which portion is also sent for further processing as detailed above. The processing device selects the portion in a manner that gives the operator the sensation that he is operating a real camera. In this way, when the pointing direction of the operator's device is moved in a given direction by a given angle, the selected portion is moved as if the remote camera had moved in the same direction by the same angle. When the zoom control of the operator's device is changed to make the camera zoom in or out by a certain amount, the selected portion is zoomed in or out by that amount.This gives the user the feeling that he or she is operating a real camera, making selection of a portion of the video captured by the remote camera a more intuitive operation for a trained camera operator. Optionally, the remote camera moves in response to movement of the selected portion, in a manner described in more detail above.
[0064] In the above examples, the entire region of the captured video stream (full frames) was typically transmitted to the production facility, and the reduction of that stream to selected regions (partial frames) was performed at the production facility. A second mode of operation will now be described.
[0065] In an example of the second mode of operation, the reduction of the video stream to a selected region (partial frame) is performed close to the camera, for example in the camera controller. Close to the camera means a location sufficiently convenient for the camera to be able to immediately send the full resolution video stream captured by the camera. The full resolution video stream, or a high resolution representation thereof, is buffered. The buffering may use a memory 70. A low resolution version of the full frame video stream is transmitted to the production facility. The low resolution version is of lower resolution and / or has lower bandwidth than the buffered stream. At the production facility, an operator or an automated system selects a sub-region of the stream from time to time. This can be done by moving the region within the full frame (similar to pan and tilt) and by changing the size of the region (similar to a zoom adjustment), as described above. Information defining the position and size of the selected region is sent to the camera controller 17. This information indicates the timestamp of the part of the video stream where the selection of the indicated region was made. This is done by indicating a timestamp, frame number, or other similar criterion in the stream transmitted to the production facility, and including the same criterion, or a criterion that can be correlated to it, in the information returned to the camera controller, the criterion being a reference to the portion of the video stream being played or processed at the production facility when the region was selected. (i) Select a specified sub-frame region within a portion of the buffered video stream at a specified time point and output this as an output stream, for example for playback to a viewer. (ii) using the logic described above, instruct the camera to adjust the camera zoom to move the subject of the selected region closer to the center of the captured video stream and / or so that the selected region occupies a predetermined percentage of the full frame of the captured video stream (e.g., an area that is 30% of the full frame); This system eliminates the need to transmit high definition video streams to the production facility, reducing delays and costs.
[0066] Due to delays in transmitting the stream to and receiving commands from the production facility, the high resolution stream is typically buffered in the camera controller for some time, e.g., several seconds. During that time, the camera controller can effectively perform post-processing on the buffered stream, e.g., to improve its visual quality. Examples of such processes include sharpening, noise reduction, blur reduction, color balance adjustment, and color gamut adjustment. By performing these processes on a portion of the captured video stream before a command is received as to which region of the portion of the stream should be selected, the time in which a properly post-processed version of the stream can be output is reduced. Once a portion of the captured video stream has been output for viewing or storage, the corresponding data in the buffer 70 can be deleted, freeing up memory.
[0067] FIG. 4 shows a system for carrying out this sequence of steps. Similar components are designated as in FIG. 1. At the video capture site, the camera 10 is provided with physical actuators (e.g., motors or linear actuators) or software settings to adjust one or more of its pan, tilt, zoom, frame rate, sensitivity, and other video capture parameters. An automated production facility 17 is located at the camera 10. The facility is next to the camera 10 or close enough to the camera 10 to allow immediate data exchange between them at a suitably high bandwidth. High resolution video data from a full frame of the camera's optical sensor is cached in temporary storage at the production facility. The temporary storage may cache, for example, two seconds of full frame data. Meanwhile, a compressed version of the captured data is transmitted from the camera to a remote control facility. The compressed data is of lower resolution than the cached data. The compressed data, or a portion of it, is played back to an operator at the control facility. An operator operates a simulated camera device 21, which simulates what a camera operator would see if the simulated camera device were controlled to control camera 10. A subregion of the full frame captured by camera 10 is displayed to the user at device 21, or the entire full frame is displayed overlaid with a bounding box that specifies a subregion of the frame that is selected for output.
[0068] A user of the control facility selects a sub-region of the full frame captured by the camera, the user selects the size of the sub-region which corresponds to zoom, and the user selects the position of the sub-region on the full frame which corresponds to pan and tilt.
[0069] The operation of the production facility 17 and the camera 10 responds to input sent from the control facility according to the size and position of the selected sub-region for a given frame of the video or time position within the video. The size and position can vary from frame to frame or from time to time. There is a delay in the command depending on the size and position of the sub-frame arriving at the production facility 17, as there is time taken for the compressed video to be transmitted to the control center and for the command to be returned. This is why the full frame full resolution video is buffered at the production facility. As time progresses, a command is received at the production facility for the selected size / position and the frame / time to which that selection applies. The production facility processor selects a sub-portion of a frame or other unit of the cached video data for that frame / time. The selected sub-portion was specified by a user of the control facility, and that sub-portion is output as a frame in the output video delivery. Portions of the full frame around this sub-portion are discarded and not included in the output video delivery. In this way the output video delivery corresponds to what was selected as a sub-frame at the control facility.
[0070] Furthermore, the zoom, pan and / or tilt of the camera is adjusted according to the size and / or position of the subframe in the full frame. The adjustment is such as to pull the selected subframe to a predefined size centered in the full frame. Thus, if the selected subframe is made larger, the camera zooms out and vice versa. If the selected subframe is moved to the left in the full frame, the camera pans to the left and vice versa. If the selected subframe is moved up in the full frame, the camera tilts up and vice versa. In this way, a continuous range of further zoom, pan or tilt is possible for the user of the control facility.
[0071] With a user interface on the control facility that mimics the camera, the user can have the feeling that they are operating the camera locally.
[0072] The applicant discloses each of the individual features described herein and combinations of any two or more of the features to the extent that such features or combinations can be implemented based on the present specification as a whole in light of the general knowledge common to those skilled in the art, regardless of whether such features or combinations solve some of the problems disclosed herein and without limiting the scope of the claims. The applicant indicates that aspects of the present invention may consist of such individual features or combinations of features. In view of the above description, it will be apparent to those skilled in the art that various modifications may be made within the scope of the present invention.
Claims
1. 1. A method for creating a processed video stream, comprising: capturing a first portion of an input video stream using a camera, the camera being adjustable to vary its field of view; storing the first portion of the input video stream at a first bandwidth; transmitting the first portion of the input video stream from the camera to a remote processing facility at a second bandwidth lower than the first bandwidth; designating, at the processing facility, a first sub-region of the first portion of the input video stream for further processing; creating a first cropped video stream by cropping the first portion of the stored input video stream to the first sub-region in response to the designation of the first sub-region; generating a camera control signal in response to the designation of the first sub-region; creating a processed video stream incorporating the first cropped video stream; adjusting the field of view of the camera in response to the camera control signal.
2. The method of claim 1 , wherein the step of specifying the first sub-region is performed simultaneously with the step of capturing the input video stream.
3. The method of claim 1 , wherein the processed video stream is a live broadcast stream depicting a live event.
4. The method of claim 1 , wherein capturing the input video stream comprises capturing a video stream of a live event.
5. After the step of adjusting the field of view of the camera, capturing a second portion of the input video stream using the camera; transmitting the second portion of the input video stream to the processing facility; designating, at the processing facility, a second sub-region of the second portion of the input video stream for further processing; creating a second cropped video stream by cropping the second portion of the input video stream to the second sub-region in response to the designation of the second sub-region; The method of claim 1 , further comprising creating a processed video stream incorporating the second cropped video stream.
6. transmitting the camera control signal to a camera device including the camera; The method of claim 1 , wherein the camera device is configured to automatically adjust the field of view of the camera in response to the camera control signal.
7. The step of generating a camera control signal includes: automatically analyzing the position and / or size of the first sub-region relative to the entire field of the first portion of the input video stream; The method of claim 1 , further comprising automatically applying a predetermined algorithm in response to the analysis to generate the camera control signal.
8. 2. The method of claim 1, wherein the camera control signal is a signal that changes the field of view of the camera so that the position of the center of the first sub-region in the first portion of the input video stream is closer to the center of the field of view of the camera.
9. The method of claim 1 , wherein adjusting the field of view of the camera comprises adjusting the pan or tilt of the camera or translating the camera.
10. 2. The method of claim 1, wherein the camera control signal is a signal that changes the field of view of the camera so that a region of the size of the first subregion in the first portion of the input video stream approaches a predetermined target size.
11. The method of claim 10 , wherein adjusting the field of view of the camera comprises adjusting the zoom of the camera.
12. estimating the responsiveness of the camera to previous camera control signals; The method of claim 1 , further comprising generating the camera control signal in response to the estimated responsiveness.
13. wherein the step of designating the first sub-region of the first portion of the input video stream for further processing is performed by a human designating the first sub-region; The method of claim 1 , further comprising displaying the boundary of the first sub-region to a user on a display.
14. 2. The method of claim 1, wherein the step of designating the first sub-region of the first portion of the input video stream for further processing is performed automatically by analyzing the first portion of the input video stream to identify an object of interest therein and designating the first sub-region to surround the object.
15. The method of claim 1 , wherein the processed video stream has a lower resolution than the input video stream.
16. 1. A method for creating a processed video stream, comprising: receiving at a processing facility a first portion of an input video stream captured using a camera, the camera being adjustable to vary its field of view; designating, at the processing facility, a first sub-region of the first portion of the input video stream for further processing; In response to the designation of the first sub-region, (i) creating a first cropped video stream by cropping the first portion of the input video stream to the first sub-region, and (ii) generating a camera control signal; creating a processed video stream incorporating the first cropped video stream; transmitting the camera control signal to the camera to adjust the field of view of the camera.
17. Use of the method according to any one of claims 1 to 16 for reducing communication delays between cameras and production facilities.
18. 1. A video processing device comprising: an input for receiving a captured video stream; a memory; and a controller, The controller storing the captured video stream at a first bandwidth in the memory; compressing the captured video stream to create a compressed video stream having a second bandwidth lower than the first bandwidth; transmitting the compressed video stream in a manner such that points in the progression of the compressed video stream are identifiable; receiving a designation of a sub-frame region and an associated point during progression of the compressed video stream; generating an output signal for controlling a camera direction and / or zoom in response to one or more of: (i) a position of the designated sub-frame region relative to the overall frame; and (ii) a size of the designated sub-frame region relative to the overall frame; a video processing device configured to create an output video stream by cropping the stored video stream to the specified sub-frame region at a point in the stored video stream that corresponds to the specified point.
19. 20. The video processing device of claim 18, wherein the controller is configured to perform video processing at a point in the video stream during a period between saving the point in the video stream and cropping the point in the video stream.
20. 20. The video processing device of claim 19, wherein the video processing comprises performing processing to improve a visual quality of the point in the video stream.