Dynamic HDR camera capture adjustments

The method for controlling video camera capture modes addresses exposure inconsistencies in HDR imaging by adjusting iris, shutter, and gain settings, facilitating automatic exposure control and enabling focused artistic and storytelling efforts.

JP2025538965APending Publication Date: 2025-12-03KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525226
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-03
Filing Date
2023-10-23
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Existing video capture technologies struggle to efficiently manage dynamic range adjustments in non-uniformly lit environments, leading to issues like clipping and inconsistent exposure across different lighting conditions, particularly in high dynamic range (HDR) imaging, which complicates the production process and requires manual adjustments that divert resources from artistic and storytelling aspects.

Method used

A method for controlling a video camera to set capture modes that adjust output pixel intensities based on scene illumination, involving iris, shutter, and gain settings, and storing brightness allocation functions for different positions, allowing for automatic exposure control and graded output video sequences.

Benefits of technology

Enables rapid, near-automatic exposure management across varying lighting conditions, allowing camera operators to focus on artistic and storytelling elements while ensuring consistent and visually appealing HDR captures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538965000001_ABST
    Figure 2025538965000001_ABST
Patent Text Reader

Abstract

To meet future demands for high quality yet economically feasible high dynamic range video programming, the inventors propose a method for setting a video camera capture mode in a video camera 201 that specifies output pixel brightness in one or more graded versions of an output video sequence of images. The method includes the steps of a video camera operator moving to at least two positions Pos1, Pos2 in a scene having different illumination from one another and capturing at least one high dynamic range image o_ImHDR for each of the at least two positions in the scene, a color composition director analyzing the at least one captured high dynamic range image for each of the at least two positions to determine areas of maximum brightness and to determine at least one of a camera iris setting, a shutter time, and an analog gain setting, and generating a master capture image o_ImHDR for each of the at least two positions Pos1, Pos2 using at least one of a camera iris setting, a shutter time, and an analog gain setting. The method includes the steps of capturing a digital number of a master capture and holding constant an iris setting, a shutter time setting, and an analog gain setting for at least a subsequent capture at a corresponding position; determining an image ODR of at least a first grade of the master capture, which comprises mapping the digital number of the master capture to unit values ​​of the image ODR of the first grade by a luminance allocation function FL_M by a color composition director, the color composition director establishing the shape of such luminance allocation function; and storing the luminance allocation functions Fs1, Fs2 for the at least two positions, or parameters uniquely defining the luminance allocation functions, in respective memory locations 221, 222 of the camera.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to methods and apparatus for coordinating the capture of images in differently illuminated areas of a scene by one or more cameras, in particular by cameras that generate image signals including a primary high dynamic range (HDR) image and a luminance mapping function for calculating a secondary grade image having a different (usually lower) dynamic range than the primary HDR image based on pixel colors of the primary HDR image. [Background technology]

[0002] Optimal camera exposure is especially important in non-uniformly lit environments. Single View This is a challenging problem (when there is no direct line of sight and the camera moves freely through various areas of different lighting and object brightness). Historically, much video production was done under controlled lighting (e.g., in a theater or news studio capture), with many overhead lighting fixtures to create a uniform base lighting. Today, there is a desire to shoot on location, and sometimes, due to cost, with a small team (e.g., one presenter, one cameraman, and one audio guy). Also, the line between professional and "amateur" producers is becoming somewhat blurred, as can be seen, for example, from internet video bloggers (who may know a little about lighting techniques, but not necessarily enough to avoid creating difficult lighting conditions), or even amateur captures with a cell phone can be the most interesting news item about a local event. Further technical assistance on the technical matters of obtaining (good) YcbCr pixel color codes is always helpful. Especially for small production teams, if the technical components can relieve some of the burden of making a good-looking video, creators can focus on other aspects of video creation, such as composition and story (although they don't necessarily want an automated system that gives them no say in the color or brightness of various image objects).

[0003] SDR (Standard Dynamic Range, also known as Low Dynamic Range LDR) already works ideally under well-lit conditions (e.g. studio 1000 lux), but can also capture in ~300 lux (daytime, indoors), ~10 times lower (low light), or even below 1 lux (dark). [The "~" symbol indicates "approximately"] When using "amateur" cameras, autoexposure is very automatic (more professional systems usually control some aspects such as the maximum amount of noise, but in some cases it is better to have a lot of noise but at least something visible than almost nothing). In real-world environments, the different areas of the scene to be captured are often different. average Not only can the brightness, light levels, or illuminance be very different, but there can also be a significant spread or ratio between the brightness of the object areas projected onto different sensor pixels (especially if there are luminous objects in the scene). For example, just before sunrise, the nearby environment where the videographer is standing may not yet be in direct sunlight, while the distant sky may already be illuminated by the sun. As such, the sky brightness may be measured in hundreds of nits (which is an engineering nomenclature, and is expressed in the physical unit cd / m). 2While the brightness of surrounding objects may be a few nits or less, at night, the average street brightness is a few nits, but looking in the direction of the light source can reveal tens of thousands of nits. During the day with sunlight, objects may reflect thousands of nits, but again, it all depends on whether there are diffuse white objects (or even specular objects) in the sunlight or black objects in the shadow areas. Indoor objects fall into the 100 nit range, while again, it depends on whether the object is lying, for example, in sunlight, near a window, or in an adjacent unlit room, which can be captured from the same shooting location when the door to the unlit room is open. In fact, it is precisely for this reason that engineers wanted to move towards HDR imaging chains (another factor being the viewer's visual impact). Because SDR cameras have a small dynamic range between capturing the luminance of dark objects at their darkest (noise-free) and the full pixel well where there aren't many photoelectrons, actors or speakers in a room are usually properly exposed, while objects outside a window are blown out (or clipped to pastel colors). This can lead to some pretty strange results: for example, if a presenter is standing against a bookshelf at the far end of a room, away from a window, in lighting that decreases by a factor of two, half of the bookshelf facing the window will appear to disappear (clipped to maximum white). In fact, with such a limited dynamic range of the camera, codec, and / or display, lighting and exposure must be very careful, and at least something will often be blown out or crushed. In the worst case, even half a face may be blown out (unless there's an artistic reason for doing so), though this is usually only done by amateurs.

[0004] But it's not just a matter of the camera (its maximum capture capacity and optimally controlled use of that capacity), but also of the way the captured image is coded in standard ways. Even if a camera can accurately capture very deep blacks—that is, captures with a noise floor of, say, 1 / 10,000 full pixel well or less—a typical representation of the output image, as a percentage of peak luminance (which typically corresponds to 128 lumas out of 255 in approximately square-root luma coding), is to make the main people and objects in the scene about 25% white (the white you see is typically produced by driving the display to its maximum brightness, e.g., with the backlight turned all the way on and the LCD pixels fully open). This means that under these productions, the standard is left with only two stops (i.e., a multiplicative factor of 2) of brightness above the face color, and of course anything a little brighter than the local diffuse white in the scene will be clipped (as is also possible with the raw camera capture of the sensor, if the sensor is exposed according to the same technical standards). You may wish to adjust this using the camera controls, but this is not the current scene's Current location from Current View The various scene shots may be slightly tweaked in post-production, but in reality they are simply mixed together. Even though automatic camera exposure control can be applied several times to different shots in different positions, there is no relationship between those settings.

[0005] The human eye adapts almost perfectly to all of the above, by altering the chemical nature of signaling pathways in cone photoreceptors to differentiate their sensitivity and locally emphasize neuronal signals for bright and dark objects in the visual field as desired (e.g., staring at a bright red square for a while will result in the appearance of an anti-red cyan square in the visual field roughly where the original square was imaged, though this, too, is quickly corrected). The brain is most interested in what kind of object we see—a ripe yellow banana, for example—but less interested in the exact sunlight that illuminated it. The brain converges on a final representation that allows us to see a tiger hiding in the bushes, whether it's daytime or nighttime. If we want to represent an original scene in its original conditions (e.g., the original luminance of the visual field) through a (moderately faithful or crude) simulation, such as a displayed version of that scene under the lights of a dark living room at night, ideally the simulation will contain information that is more or less relevant to human vision, so that the brain can attempt to create the illusion of seeing a more or less accurate scene. However, cameras count photons, converting each group of N incoming photons into measured photoelectrons. In that respect, cameras are simple devices, but they lack significant processing power for advanced applications. While such counting is good when precise measurements need to be made, such as determining whether a section of a building is well lit, this characteristic makes them less suitable for video display chains. For example, driving into a tunnel and then exiting it back into a clearer environment can be discernible, even with a single camera. The lighting design in a tunnel is optimized for human vision and not necessarily camera-friendly. When driving out of a tunnel, humans first see, with some delay, an image of the external environment that is almost completely white (at the time of capture), and then an auto-exposure algorithm based on average luminance adjusts the image to a technically adequately exposed image. As humans, we see the usual impression of a dimly lit tunnel, and outside we see a brighter, but usually still well-visible, environment.That is, human vision appears to be perfectly adapted to most of the world's lighting conditions, even though much of it is artificial. Only under the worst conditions, such as when driving into a shadowed area on a sunny day, can there be visibility problems, primarily due to sunlight reflecting off dirt on a car windshield.

[0006] In many scenarios, multiple cameras and possibly multiple mobile cameramen are involved in the production of a video, and it may be desirable to adjust the brightness look of these cameras. In some productions, this is done in a different location or facility, such as an Outside Broadcast (OB) truck (or a grading booth if not broadcasting in real time). Typically, today, everything related to the production of good video must be done in real time, and therefore by a team of specialists with different focuses. For example, the director of a sports broadcast is too busy to have any say in the capture, except to be able to choose which cameramen should capture roughly what, and therefore which primary camera feeds will end up in the final broadcast signal at what time. Indeed, whereas a film is to some extent planned during shooting, the perfect artistic synchronization of plot, geometric capture composition, emotion, music, etc., is thoroughly planned out before the film (storyboards, etc.), and in post-production (to the extent that as night shooting days, the applicable look is graded for the captured content), a real-time producer must be able to control all the technology. On the FlyIt has to be captured in a certain way. Live broadcast producers can only capture what they have the ability and experience to capture. For example, they have watched and produced many soccer matches in the past and know when viewers prefer to watch the scoreboard over someone's boring speech. However, technically, the primary camera feed must be "just right, or more," so the director relies on it to select. Cameras can have several capture settings, such as roll-off or knee points for black control, and they usually set these to the same standard value to get, for example, blacks that look the same. When cameras have different black behavior, this can be noticeable because you get milky blacks. Therefore, they use standard video test signals (color bars and gray bars) and make some adjustments as needed. Under fast-paced conditions, if the colorimetry is really wrong, the director will simply discard the feed, such as the cameraman's feed who is still struggling to get the right frame of a zoomed, fast-moving action. However, there are cases where a feed is considered important, and you get what you get, regardless of any colorimetry artifacts. Therefore, things are standardized. simple It is preferable that the primary camera feeds meet at least the minimum requirements of uniformity (e.g., if there is an identical set of cameras, set all controllable parameters of all cameras to the same value). Therefore, any system that meets the requirements of such a scenario should be simple enough to be practical and feasible.

[0007] Because the complexity of liberal high dynamic range productions increases compared to standard SDR productions, we can expect to see a greater reliance on tailored technical solutions. However, that complexity also allows for a uniform (though typically customizable) approach to some approaches. For example, any future systems adapted for consumer capture may have a similar need for relative simplicity of operation, while simultaneously providing robust coverage for many different capture situations. While high dynamic range creates new opportunities, it also creates challenges. Indeed, if we had a camera with infinite dynamic range (which we don't yet), we could shoot with any settings and correct it all in post as desired. However, complete freedom isn't desirable in all scenarios, as there are scenarios in which we want to limit post-production or have already resolved key aspects during shooting.

[0008] For example, when producing a movie consisting of several shots (e.g., daytime and nighttime), adjustments to these shots can be made after capture (in post), e.g., a human color grader would typically modify the final brightness of the video's master HDR image (e.g., a maximum 1000 nit image) so that a dark scene appears dark enough after the previous daytime scene in the movie, or conversely, not too dark compared to the following explosion scene, etc. Without limitation, the selected maximum brightness of any (relevant) pixel in the image is the maximum brightness of this particular grading (a.k.a., graded image), e.g., if a scene is to be represented by only one graded video, we consider it to be the 1000 nit master grading. In principle, a grader can optimize the master image brightness for each and every pixel in the color grading software (e.g., define a YCbCr color code for the pixels of an explosion so that the brightest pixel in the fireball has a maximum video brightness (ML_V) of 1000 nits, but the master HDR video still does not exceed 600 nits. Correlations between brightness and color codes, such as brightness coding luma codes, can be made by selecting a primary EQTF such as PQ, but the details of the actual coding are possible variables and are not necessary for explaining this invention, so we will not discuss pixel brightness). However, this Real WorldWe haven't said much yet about the relationship between the brightness of the fireball in the scene. Not only do these brightnesses depend on the amount of heat generated by the explosion (which isn't usually what pyrotechnicians aim for), but it's also important to note that a camera isn't a brightness meter (it's simply a photon counter), so the amount of photoelectrons accumulated at the pixel corresponding to the brightest point of the fireball depends on the physics of the sensor, as well as the iris aperture setting and exposure time of the shutter (and possibly a neutral density filter in the light path). This bridges the gap between the world of the scene (and camera) and the world of the ultimate display. What is captured isn't the most important aspect, but rather what is seen on the display (this is one reason why some HDR coding formats refer to the display, and in some cases, prefer to work in-camera). In filmmaking, the color grader is a watershed entity: choosing the appropriate specifications for what exactly will be displayed for the captured image. Ultimately, what is displayed in the image is what matters, not the mathematical number of pixel colors. However, opponents of the display-referential approach may argue that there is no certainty about how the result will look and that some respect should be given to the original image (this, for example, is for full quality in the future). Nevertheless, the present applicant has shown that the requirements can be met by defining (at least one) graded image for an ideally envisioned target display. It should be noted that at least a high maximum luminance master grading has no real quality loss compared to the camera raw and, for most intents and purposes, can form the preserved original (since raw camera data is not that important for human consumption). Even for low-bit (8-bit) LDR, the corresponding grading has proven to be largely lossless, but in any case, the display-referential portion is satisfied by a secondary grading, making it an absolute HDR codec approach. For example, the reference LDR grading corresponds to the master HDR grading.

[0009] Grading (or color grading, where the main aspects of fine-tuning are color brightness and luminance, the latter being the specific formulation of the video brightness channel) refers to a human or automaton (or semi-autonomous combination) specifying pixel brightness for various objects in a captured image along a selected luminance range as needed. This may be necessary in several display scenarios. For example, as shown in Figure 6, it may be desired to create a master HDR video with a white-point brightness of 2000 nits, also known as the maximum brightness grade, as the (only) output video of one or more cameras recording a video program in a scene (or one of them). The raw capture from the camera sensor is not grading because it is not optimized for human consumption on a display (e.g., objects in shadow areas may appear darker than desired, whether in HDR or LDR grading). While devices or users further down the video communications chain can use any kind of secondary grading (see brief discussion in Figure 10), it is advantageous if a camera outputs one (master) graded version of the captured content (called the HDR master) but also outputs a second (different) graded version of the captured video (e.g., usefully, the LDR video (called the LDR master)). Some content consumers use the first version, some the second, and some both. Grading assigns selected (appealing) brightness values ​​to various objects, which can be captured as digital numbers from the camera's analog-to-digital converter (ADC 206), which is a digital capture of each sensor pixel well's fill state with photoelectrons. The nomenclature "digital number" indicates that the value, a relative indication of how bright an object is relatively in the captured scene, is a 16-bit value, e.g., 011011011 11110000.For example, a ~5000 nit flame in an open fireplace in a scene may have a digital number of 50000 (or normalized 50000 / 65536 = 0.76), which depends on the camera's selected capture settings, such as the iris aperture, and the grader may choose to look good at, say, 500 nits (average pixel brightness in the flame object image area) in a master grade of 2000 nits video. Therefore, if an end consumer purchases a television with a display maximum luminance (ML_D) equal to a 2000 nit (ML_M) value, the end consumer will typically display all of the luminance formulated in the video signal, i.e., in the master HDR output image. This means that the end consumer will see the image almost exactly as the creator intended. If the television is only capable of, say, ML_D = 700 nits, it will not be able to perfectly display the movie as intended, but the receiving device, such as a television display, will downgrade the 2000 nit master HDR image by applying display optimization mapping as described in the previous patent, which is unrelated to the camera capture and production techniques currently being described. Display optimization can be coarse, or attempt to preserve as much of the original intended look as possible within limited luminance dynamic range capabilities. In this patent, we refer to the raw digital number image that measures the scene as a "capture," to distinguish it from an "image" (e.g., an output image, such as from a camera to an OB truck, or a broadcast image, which may be defined differently but elegantly similar if the camera is designed to shift some color mass). Images are typically considered to be graded, that is, to have optimal lightness values, typically luminance, or indeed luma value coding for them, according to some reference (e.g., a 2000 nit target display). A common choice of such a reference facilitates camera calibration, but is still not trivial, at least a commonly referenceable one.It will be clear from the context if "grading" refers to a method of modifying brightness, but when referring to some graded image (whether that grading is the result of mapping the original camera's digital numbers to the capture, or a re-grading or further grading from a previously created grade), in fact it is almost always referring to the resulting graded image (rather than the process).

[0010] Furthermore, even if we had the freedom to rely on perfect grading in post-production (i.e., any "errors" in pixel brightness or color would, at least in theory, affect pixel color values ​​during post-production grading). Complete There are still challenges and opportunities for wanting more freedom in the live capture (exposure) of a single scene (consisting of following someone on a factory tour, or a bike going through a tunnel), which can be compensated for by redefining it (this is a freedom that is not available in real-time productions like capturing a Tour de France on a bike). In fact, we want to link freedom closely with consistency, to get a kind of consistent freedom. This may become important in the future, as not only do different kinds of video productions emerge and continue to spread, but also new technologies (for example, regarding the capture of different views and aspects of a scene) emerge and continue to spread.

[0011] Additionally, capturing several videos with different dynamic ranges (e.g., HDR) from the camera to the distribution network or offloading to memory Version For example, they may want to produce an LDR version (and a corresponding LDR version) "instantly" so that the LDR version is broadcast immediately to customers (e.g., via cable or satellite systems) and the HDR version is stored in the cloud for later use (e.g., future pay-per-view) and reused as a rerun or video snippet 10 years later.

[0012] Several types of video productions could benefit from the following new insights and embodiments. In classic film and series productions, the same scene might be shot multiple times in succession from different angles (e.g., a fight scene, once from behind the attacker looking down at the victim on the ground, and once from the side closer to the ground), and in such applications, the innovations presented below can improve production speed, simplify capture, facilitate complex capture situations, or increase or mitigate post-processing possibilities, as well as benefit some situations where you want to capture the entire action throughout a complexly lit scene in one or more coordinated shots. Not only are professional programs shot from several static television cameras installed on-site, but even semi-professional corporate communications videos (e.g., a company's private network broadcast of an employee's visit report from another company or hospital to a division) have found a desire to move away from the static presenter and become someone who moves around everywhere, leaving the presentation room and walking into a hallway, even getting in a car and continuing the presentation while driving. Especially if technology can help them take on some of the complex tasks, even "amateurs" can begin to create more professional videos. This isn't difficult when simply creating any capture "as is" - that is, when fixing exposure settings or relying on the camera's auto-exposure (which typically suffers from clipping and arbitrary results such as areas being the wrong color (e.g., parts that are too dark to see well or uglier than optimal) - but it is considerably more difficult for high-quality HDR production, especially considering that the technology is relatively new and many people are still working on optimizing the fundamentals. In fact, the technology helps avoid what can sometimes become a somewhat chaotic state.

[0013] Figure 1A shows a typical important dynamic An illustrative example of a capture is shown. One-Level ExposureEven with exposure (i.e. being a constant integral measure of the current scene lighting level that requires corresponding values ​​for camera settings, iris, etc. to fill the pixel wells to a certain level, i.e. leading to the setting of such values ​​for further captures), it is known that the exposure of a speaker's face when e.g. near a window where the light level falls off quadratically with distance can result in overexposure of objects near the window. This is a common issue in low dynamic range LDR (aka standard DR) capture, and certainly present when there is sunlight, resulting in some pixels being clipped to the maximum captureable level (white, e.g. luma code 255).

[0014] We want the photographer to be able to walk freely through a hallway (102) with several lighting fixtures (109) that gradually change the local illumination (illuminance) from outdoors (101) to indoors (103), possibly through a camera. Obtaining the "correct" exposure with a high-dynamic-range camera (i.e., a camera with a contrast ratio of, say, 14 or more stops, or 16,000:1, a full well of, say, 100,000 photoelectrons, and noise of 5 electrons, a sensor design that could be considered "equivalent full well" today, and a dynamic range of over 2,000,000:1 likely to be achievable for mainstream cameras in the not-too-distant future) is not particularly difficult with such a large dynamic range, in the sense of "good" capture of all information. Since you can go much darker (with good capture accuracy) than with a bad camera, you focus on capturing the brightest objects well enough, for example, without clipping or desaturating the color channels. This may not result in a satisfactory picture if you (directly) map the whites of the brightest image to the maximum displayable white of, say, a 450-nit display. However, at least in principle it is always possible to brighten it later by grading (as there is still information in dark pixel colors without too much degradation, especially if they were captured in low noise conditions). Note that even 14 stops of capture should be taken care in very difficult HDR scenes / environments (e.g. a light bulb filament that can exceed 1 million nits should not be exposed if noisy exposure for dark corners is not needed).

[0015] For high dynamic range cameras ( / video use chains) exposure is less of a problem, since it can record objects as relatively dark percentages of white (which can then be brightened as desired in post-capture image processing), and even HDR captures can be subject to lower dynamic range (usually SecondaryThe problem of correct exposure arises again when the image is desired. The HDR scene needs to be compressed in the available low dynamic range, which requires more reflections (which can partly be seen as virtual re-exposure). However, there are more possibilities digitally than simple control, especially with iris aperture. Still, it is not always easy to create a good LDR image for a difficult HDR scene, or more specifically for a master HDR grading.

[0016] In general, iris or exposure control can be either manual or automatic. For example, as shown in EP 0152698, analysis of the captured image can determine the capture of red, green, and blue components. Maximum and setting the iris and / or shutter (possibly in conjunction with neutral density filter selection and electronic gain) so that this maximum value does not clip (or is not over-clipped).

[0017] A common control method, at least in SDR capture scenarios, is to control the brightness (or relative photon concentration) of the scene. average (or more precisely, a smart averaging algorithm, e.g. giving less weight to bright sky pixels in the sum) is used as a very reasonable measure. This is also particularly relevant to how SDR luma captured in an SDR display scenario is easily displayed: the brightest luma in the image (e.g. 255) acts as the maximum control range value for the display drive signal, which typically drives the display to display the maximum possible output (e.g. for a fixed backlight LCD, the LCD pixels are driven to maximum transparency), and the displayed color appears visually white. A reasonable assumption of the "display as painting" approach is that Single Light(That is, under fairly uniform lighting conditions (e.g., by using controlled baseline lighting, e.g., by a matrix of ceiling lighting fixtures in a television studio production), diffusely reflecting objects in the physical world reflect approximately 1% to 95% of their current luminous intensity. Specular reflection can simply be white. This creates a histogram of luminance (or luma) centered around the ~25% level (or the middle of the luma code). To the human eye, a painting that is primarily interesting for its object color, whether lit by a strong or weak lighting fixture, will look roughly the same when displayed on a 200-nit ML_D display (i.e., with white pixels having a displayed luminance of 200 nits) as on a 100-nit display. This is because the eye compensates for the difference in luminance, and the brain only cares about seeing the relative differences between the object points in the painting on the LDR display, respectively. Also, by mapping the luminance range of this scene to that which fills the camera sensor's pixels by ~25%, we can achieve the same results.) Relatively It makes sense to characterize it as a good sampling of the information available in such a scene, even for an LDR sensor. But even in an SDR imaging chain, i.e., for example, an SDR camera capture produces an SDR image as output for display on a conventional SDR display, the average value is still Bimodal It doesn't measure up so well in scenes (e.g., an interior room and the outside world through a window) or in very narrow modal scenes (e.g., containing only a few shades of black), such as a coal mine. Still, with a little fine-tuning and grading by the camera operator, SDR has seemed to work surprisingly well in real life for, say, a century, showing viewers everything from the coronation of a queen to the depths of the ocean. There are only a few desiderata that will drive the move to HDR, and ideally, an improved treatment of black, such as capturing and rendering objects brighter than white, and a more professionally controlled image description framework. Common From the base, for example, it has a maximum brightness of about 2000 nits ( Video imageML_V) We are building layers of technology such as HDR images and associated target displays.

[0018] However, this averaging method is a widespread (and in fact ad hoc / de facto standardized) video production practice dating back to the SDR era, and it is one example of many technological approaches that warrant a complete rethinking and redefined approach in the HDR ecosystem. In a chicken-and-egg problem, much of the early HDR technology, or at least the standardized technology, focused on the definition and coding of HDR images (e.g., via hybrid log-gamma or SL-HDR) in addition to the technical ability to create brighter displays. How to connect this new coding and display capability with image creation was not necessarily developed. One way of looking at things is to continue creating as usual, viewing codecs simply as "translators" capable of recording whatever is created. Another way is to want to take advantage of the enhanced capabilities, such as in 3D movies where objects are thrown at the viewer's head and are now coordinated with HDR effects such as intensely bright and colorful explosions. A third way is to want additional production techniques to control the vast new dynamic range capabilities we have so that things don't get too messy. This appears to be the rationale behind the 1000-nit bridge point for content creation, which seeks to closely marry the relative HDR paradigm of HLG with absolute paradigms like Dolby Vision and SL-HDR.

[0019] U.S. Patent Application Publication No. 2017 / 0180759 is an example of a general-purpose system for obtaining master HDR video to a recipient (e.g., end consumer) via a transmission method. This allows an original master-graded video with a target display maximum luminance (a.k.a., white point luminance) of 5000 nits to be represented as a proxy video for communication with a target display maximum luminance of, say, only 2000 nits (i.e., pixel luminance is at most 2000 nits luminance). At the receiving end, the TDML video can be lower than 2000 nits (e.g., 400 nits), but the original master HDR video can be reconstructed from the received proxy (referred to as intermediate dynamic range video) or a brighter output video can be created.

[0020] To help the reader quickly recall the various parts of the video communication and usage chain (which should not be confused), for the purposes of illustrating Figure 10, we have summarized the situation in general terms (to present a variety of situations, consumer productions where, of course, some of the components are only vestigially present at best).

[0021] Exactly what happens can depend on whether there is a live production, such as a sporting event or news report, or whether there is a movie, such as one that has been shot over several days, edited together, and distributed, but generally, several components are present, at least as far as the present innovation is concerned. At the filming location 3000, there is a first actor 3001 under controlled lighting 3003, for example, hanging from a crane, and a second actor 3002 in a darkened area of ​​the scene (this may be either a constructed scene or a natural scene located in situ). There may be one or more cameras. In this example, there is a first camera 3004 attached to a person, for example, as a Steadicam, and a second camera 3005 that is statically mounted on a tripod on location, at least for the moment. With better video compression and widespread availability of the Internet, even raw shots (e.g., dailies) can be communicated to a production studio 3010. The production studio 3010 no longer needs to be located near the filming location (e.g., employing an Internet Protocol video communications connection 3009). Assume a "director" is viewing the available raw camera footage (either in real time or offline). While classic LDR filming was primarily about shot timing, logical placement, and so on, compositional decisions can now be made regarding the grading of various objects in one or more scene shots. With the relatively new HDR production, this is relatively simple. For example, boosting the LDR with a fixed LUT associated with HDR might be simply straight from the camera, with a few tweaks. For this discussion, we'll assume a person (or automaton) takes on the role of "grader," whose role is to have a say in the brightness and color of objects in the image (leaving the pace of the story to others, such as the director or editor). For example, different graded versions of a shot (or part of a film, or program, etc.) with different TDMLs can be viewed simultaneously, typically on a first reference display 3011 and a second reference display 3012. Grade changes can be made via the color console 3013.Graphic elements such as robots and monsters are also mixed in via a graphics unit 3015, which typically has a connection 3016 (e.g., again via the Internet) to a graphics supplier (these techniques are known and do not require detailed description in this patent application). The final product (cut) is, for example, a 5000-nit future-proof master HDR grading M_5000. This is stored in memory 3018 for later use and transmitted via a distribution medium 3019 (e.g., again via the Internet) to, for example, a content broadcaster 3020 (e.g., the British Broadcasting Corporation). The broadcaster can then distribute the video to end customers via various communication channels. For example (excluding various possible intermediate units such as a cable headend), from the 5000-nit master, a first intermediate version (proxy) IM_2000 for transmission via satellite with 2000-nit TDML is created in the broadcaster's equipment and communicated via television satellite 3021 to a satellite dish 3021 connected to a satellite television set-top box 3051. Assume that the STB is responsible for computing a display-adapted version of the received video (IMDA_550), which is coordinated with the end-user display 3052 and communicated, for example, via an HDMI® cable. It is also common today for broadcasters to provide at least part of their content via a web portal. For example, a secondary proxy video (IM_1000, which is only 1000 nits TDML) is communicated, for example, via the content delivery network 3022, and end consumers can access that version, for example, on their mobile phones 3055. It is clear that many different setups can work with and benefit from the following innovative embodiments. Indeed, video distribution is becoming more hybrid (“standard”), and the distinction between live “broadcast,” VOD, user-generated content, etc., is to some extent disappearing, as, for example, IP packages can be delivered not only via classic communication systems like cable or satellite, but also via communication standards like 5G. Therefore, again, the video production example is not intended to be limiting in any way, but merely conceptually illustrative.

[0022] This video production and communication requires a variety of HDR and LDR images. High-dynamic-range images are generally understood to have a wider dynamic range than current low-dynamic-range (also known as standard-dynamic-range) images (well-understood by those skilled in the video and television arts). These images possess a sufficient brightness range, ranging from deep black to a Lambertian-reflected white (e.g., a piece of white paper, for clarity) under uniform illumination. The darkness of the black depended on various technical characteristics of the capture (e.g., camera noise) and display (e.g., ambient light reflected off the display screen). There was no pixel brightness significantly above white. If luminance values ​​in nits were related to pixel brightness (as in current, more advanced HDR systems), LDR images would be characterized as having a typical white brightness of 100 nits (i.e., an associated TDML of 100 nits). The minimum reasonable black for LDR is 0.1 nits. Because some HDR coding does not desire deeper blacks, HDR images or videos have the potential to code brighter pixel brightness (typically at least twice as bright) than SDR currents, i.e., 200 nits or higher TDML. The camera has a means to accurately capture scene brightness above Lambertian white (e.g., multiple pixels for different exposures, multiple exposures of different lengths, or LOFIC and dual-gain converted pixels), and the display has a means to display extra-bright pixels (e.g., a 2D LED backlight matrix with separately controllable intensity, where normally bright parts of the display get local LED illumination so that they appear 100 nits or less for those pixels, and LEDs for, say, self-luminous objects in the image are driven, say, 10 times brighter, so that those pixels appear as 1000 nits on the display screen).

[0023] US Patent Application No. 2017 / 0180759 describes various methods for creating HDR videos starting from a pre-made master HDR video. Secondary HDR video (especially various communication systemsWhile the present invention relates to the creation of HDR-grade video (which is useful for various technical constraints of a video system), it does not specifically teach how to create various grades (at least a master HDR-grade video) from camera captures, much less how various captures of different shooting localities with different local lighting of the environment can be adjusted by one or more cameras to easily arrive at a specific HDR grade as output (whether to be used directly for, e.g., consumer display, or further handled, e.g., further re-graded, stored, mixed with other content, etc.). Summary of the Invention [Problem to be solved by the invention]

[0024] Given the complexity of HDR video technology (and the fact that this technology field is only recently emerging), there is a need for practical, rapid, and near-automatic handling of the primary grading and often also the secondary grading of proper exposure for one or more cameras' on-the-fly dynamic lighting environment capture that automatically emerges from the camera as a graded output video sequence, allowing the camera operator and possibly the director to focus on artistic aspects (e.g., geometric composition) and storytelling aspects (e.g., acting) (such as from which angles should actors be filmed). [Means for solving the problem]

[0025] The above needs are met by a method of controlling a video camera (201) that includes setting a video camera capture mode that specifies output pixel intensities at corresponding locations within a captured scene in one or more graded output video sequences of images, the graded output video sequences being characterized by having different maximum pixel intensities. The method comprises: a video camera operator moving to at least two positions (Pos1, Pos2) in a scene having different illumination from each other and capturing at least one high dynamic range image (o_ImHDR) for each of the at least two positions in the scene; a color composition director analyzing at least one captured high dynamic range image for each of the at least two positions to determine a respective region of maximum brightness, and determining at least one of an iris setting, a shutter time, and an analog gain setting of the camera in response to the maximum brightness; for each of the at least two positions (Pos1, Pos2), capturing a respective master capture (RW) using the determined at least one of the camera's iris setting, shutter time setting, and analog gain setting, and storing the determined iris setting, shutter time setting, and analog gain setting for subsequent image capture when the video camera is in the corresponding position; determining at least each first grade image (ODR) for a master capture, including determining an adjustable brightness allocation function (FL_M) for each position that maps a digital number of the master capture to a nit value of the first grade image (ODR); Storing the brightness allocation functions (Fs1, Fs2) for the at least two positions, or parameters that uniquely define the brightness allocation functions, in respective memory locations (221, 222) of the camera.

[0026] A camera's behavior mode refers to its capture methodology, specifically the way it outputs video; more precisely, in this context, how the camera outputs a particular luminance for pixels in the output image that correspond to points in the scene focused by the camera lens onto the image sensor. The camera outputs different versions of its video, i.e., different grades (i.e., different gradings) of video with different maximum luminance values ​​(ML_V) (e.g., 3000 nits and 150 nits). The mode (and its characteristics) is set by at least determining, for each location, a luminance allocation function (typically by a color synthesis director, either human or automaton) that maps the captured digital number to a graded luminance as desired for that location, and then adjusting the mapping to a common range that typically ends at a selected common maximum luminance value (ML_V) for the various locations. This occurs after at least one set of capture settings (iris, shutter speed, etc.) has been determined or is suitable for adequate capture of various positions in the environment where filming will occur (except for potential clipping of very high scene luminances above some selectable maximum, which can be done iteratively by looking at what gets clipped in the output image for each position). The faithfully captured maximum luminance will usually be less than the final maximum luminance in the scene, but typically only a few pixels are allowed to clip, or at least those that are clipped are not very important for following the program or story, for example. The region may be as small as a single pixel; for example, the color composition director may select it by clicking on it (in one of the images captured during the initial environment discovery phase until appropriate values ​​for the basic capture parameters and functions have been established). For example, a different function that darkens the darkest scene elements in at least one output graded video more than the initially tested function corresponds to a different mode of the camera's colorimetric behavior (i.e., specifying the luminance of the output pixels). Thus, the behavior mode is determined by the functions for at least the various positions.The color composition director knows how the maximum luminance varies in the region of maximum brightness, so for example, closing the iris by one stop will result in a linear drop by a factor of two in digital numbers, and some drop in the output grading depending on the initially set mapping function (e.g., the default function or one loaded into the processor from a previous operation) (the exact value is not important, since we only need positioning at or near the absolute maximum of the output range). As the iris opens, the maximum will be clipped higher, which is noticeable because neighboring slightly lower luminances will also start to be clipped. The iris value (and possibly other parameters, i.e., all of those that determine the sensor pixel exposure) can be chosen to be a value that satisfies the captured pixel clip given the capabilities of this camera (e.g., possibly looking at noise behavior for black), but no more, as chosen for this shot. The master capture is the digital number output by the ADC when making a capture using the capture parameters just established and now set and fixed for any subsequent captures, at least for that corresponding position (and possibly a second position, if that works well by faithfully capturing substantially all scene objects there as well). This then serves as a stable starting point for optimizing, what is crucial, the correct grading functions for various positions. The shapes of these functions are adjustable (i.e., you can, for example, brighten the darkest pixels in the scene and darken the brightest pixels at will), and the shapes used are sufficient if they produce at least one graded output video that has an appearance acceptable to the color composite director. If a subrange of luminance is, for example, too dark, the color composite director can locally raise the function within that subrange relative to the current function shape (thus producing a higher output).

[0027] A video camera has roughly two internal technical processes. The first is the optimal sampling of the optical signal that actually represents the physical world, i.e., the correct recording and usually linear quantification of the colors of a small part of the scene that is imaged by the lens, usually onto a quadruple of subpixels (e.g., red, green / green-blue Bayer, or cyan-magenta-yellow based sampling) of the sensor (240). The controllable aperture area of ​​the iris (202) and shutter (203) (and possibly a neutral density filter) determines the number of photons that flow into each pixel well (a linear multiplier of the brightness of the local scene object); thus, these settings can be controlled so that, for example, the darkest pixel of the scene falls above the noise floor, e.g., a 20 photon-electron measurement above 10 photon noise (and a noise of ±X photons for the value 20), and the brightest object of the scene fills the pixel to, for example, 95% (i.e., 95% of what the pixel well can measure, which the ADC expresses as a so-called digital number DN). Using the analog gain (205), we can make it appear as though more photons are entering the scene (as though the scene is brighter, since we want to select the iris and shutter for other visual characteristics of the captured image, such as depth of field and subject motion), by boosting the voltage representing, say, 25% pixel fill to a 50% level (which can also be amplified in the digital domain, but under any more general class of image enhancement processes, we think of it as a multiplicative, not necessarily linear, boost). The analog-to-digital converter (206) represents the spatial signal (e.g., the image of the red pixel) as a matrix of digital numbers. Assuming we have a good quality sensor with a good ADC, a 16-bit digital number representation, for example, would give values ​​between 0 and 65535. For several reasons, these values ​​are not perfectly usable, especially in systems requiring typical video, such as Rec. 709 SDR video. Therefore, SecondThe internal process of the image processing circuit (207) can perform all the necessary transformations in the digital domain, for example by applying an OETF (Opto-Electronic Transfer Function) that is approximately square root shaped to convert the digital numbers, which ultimately result in the Y'CbCr color coding of the pixel in the output image, where Y' is the luma that describes the lightness of the pixel, and Cb and Cr are the chrominance (also known as chroma), i.e., the hue and saturation, that describe the color (in the case of HDR, these are the nonlinear components defined in the OETF version of the perceptual quantization function standardized in, for example, SMPTE ST.2084).

[0028] The image processing circuitry (207) (e.g., including a color pixel processing pipeline with configurable processing of input pixel color triplets) functions in a novel manner in the technical insights, aspects, and embodiments described below, for example, by applying a configurable function to the luminance or luma components of pixel colors to obtain a primary-grade image (ImHDR) and / or a secondary-grade image (ImRDR) (e.g., a standard dynamic range (SDR) image) (thus providing at least one graded image with appropriate image object pixel luminance or perceived lightness as output). These are typically recorded in a memory 208 built into the camera. The camera also has a communication circuitry (218) for outputting images via cable (SDI, USB, etc.), Wi-Fi, 5G, etc. (possibly by connecting the camera to an add-on device). Regarding the communication system (209) that the camera can use, for example, for future-oriented professional cameras, an Internet Protocol communication system is envisioned. Also, for example, a consumer taking a picture with a camera embodied in a mobile phone can upload it directly to the cloud using IP over 5G, although of course many other communication systems are possible and this is not the core of this technical contribution.

[0029] A simple camera or a simple configuration of a more versatile camera might provide, for example, one single grading as output image (e.g., a 1000 nit ML_V HDR image with appropriately assigned luminance for any particular scene (e.g., a dark room with a small window outside or a dimly lit souk with sunlight shining on some objects through a crack in the roof)). That is, a series of time-sequential images are captured while at least one cameraman walks through the scene capturing the action and other scene content. One An output video sequence of 1000 frames per second is generated, typically with different luminance allocations for various differently illuminated scene regions, which Basic CaptureTherefore, on the one hand, most of the scene information should be present in some form in the video codification (i.e., the brightness of the various scene objects in the rendered output video; this is the "technical rendering standard"). This may entail, for example, some clipping of the brightest sun-reflecting clouds next to the sun. On the other hand, it is preferable to somehow encode the scene brightness (i.e., have a brightness that looks obviously wrong on the display) but also have these object brightnesses with brightness that looks correct (even if one wishes to further adjust this initial allocation of pixel brightness according to artistic preference to give a scene in a film the final look). An example of a mismatch between the initial (e.g., automatic) grading and the artistic final grading is a high-key look with clipping. The creator intentionally (usually not in the camera capture, but it is possible) clips the colors of the brightly lit half of a face and pushes the remaining colors into the upper region of the RGB gamut, resulting in bright, desaturated colors. Display-ready pixel brightness for a high TDML grading (e.g., with ML_V=5000 nits) should typically not be around 5000 nits, as this would make the person's face look like a light source. For the final master HDR grading (i.e., film for release), you might want the clipped color of the face to be at a level of, say, 1000 nits. All of this can also be achieved directly from the camera by defining an appropriate mapping function, as explained in Figure 6, for example.

[0030] this Primary Grading The idea is that the image quality is established based on well-structured captures from a sensor, i.e., most objects in a scene (even bright ones) are well represented by a precise spread of pixel colors (e.g., the different bright gray values ​​of sunlit clouds).

[0031] In addition, the basic configuration is Basic Capture Settings(iris setting, shutter time setting, possibly analog gain setting greater than 1.0). Any further grading decisions can be stably built on these digital number captures. In fact, the camera doesn't even need to output the raw capture, but can simply output a primary grade, say a 1000 nit (master) HDR image. As mentioned before, the capture itself is only a technical image and not very useful to humans (all psychovisually wrong brightness), so no decision needs to be made anywhere, but the primary luminance allocation function is invertible and can be jointly stored or jointly output.

[0032] The optimal configuration of the basic capture settings may be determined by a human (e.g., a color composition director or a camera operator who can also take on the role of color composition director during this initialization phase, e.g., a consumer, or the only technician in a two-person off-site production team, the other being a presenter), but this may also be determined by an automaton (i.e., some firmware, for example).

[0033] Figure 6 shows an example of a user interface representation, which is displayed on a display (depending on the application and embodiment, this display may be in an OB truck; in a broadcaster's or producer's studio, for example, in a production using a REMI (remote integrated model); or somewhere on the set with a lighting fixture enclosure that interacts with the camera; a computer used by a one-person production team, i.e., a camera operator, to perform visual checks to better control the camera; or it may be attached to the camera itself (e.g., viewer). In principle, only the camera operator needs to be on-site (e.g., producing a semi-professional high school program). Graphical diagrams for determining mapping functions (between various representations of pixel brightness) are also shown; one of these diagrams, the image diagram (610), shows a captured image (possibly mapped with a function to create a basic impression of the look of the display being used, and also to provide less visibility of various scene objects, since accurate colorimetry is not necessarily required to determine all the various settings in any given application). Assuming it is a touch screen (those skilled in the art can see for themselves how an equivalent version can be created from the basic example, for example, by using a mouse-like sensor), the human controller (i.e., what we call the role of the color composition director in the claims) can click on the object of interest (OOI) using his finger 611. For example, a double tap indicates that the user has clicked on these (bright) color(Quickly) we indicate that we want all three color components to be captured well by the sensor, i.e., below pixel well overflow for at least one of the three color components. This means that the flames will always be well-represented, at least basically according to technical rendering standards, and there will be no uncurable color errors. An image analysis software program interacting with user interface software (e.g., running on a computer in an OB (Out of Studio Broadcast) truck) can then determine which colors are present in the areas where there are fireplace flames. For example, suppose a camera operator (in cooperation with a color composite director) captures an initial raw capture, i.e., a derived grade HDR image (or a derived grading for a display where the color composite director determines the settings). In the representation of the present teachings, (at least one) original high dynamic range image (o_imHDR) is determined. Note that the original HDR image does not need to be, and usually is not, a digital number from, for example, an ADC (or indeed a linear rescaling of them). This can be useful if we map these DNs using a function that lies within the value range of some representation (e.g., using the highest possible DN). Highest Luma Code(When mapping to o_imHDR) (not necessarily to a power of (2;N)-1, where N is the number of bits representing luma, for example, if some luma code is reserved for administrative purposes such as timing codes), to see which objects are well captured from the secondary image representation. For example, if all patches of flame have a pixel value of 940 (maximum luma in a narrow range) without any change, this may indicate that the iris should be closed to reduce incoming light and all numerical values ​​in o_imHDR should be lowered to see all spatial pattern details of the flame. Depending on whether the color compositing director wants to see these flame details, the corresponding iris setting becomes the final setting (for example, they may want to open the iris more to avoid too many capture noise issues at the bottom end of the base capture and resulting o_imHDR). Once the color compositing director clicks or taps at least one location, the software can check whether the colors are already well represented in this capture. For example, suppose the red and green components of all pixels of this flame fall between 990 and 1020 luma, and the blue component is low relative to the well-captured bright yellow. This already serves as a (graded or coded) representation of a good captured image (e.g. OB track software can also receive the capture and check if all digital numbers are below the power(2; ADC bit count)). If there is clearly nothing more, this is a good capture if the maximum pixel luma of o_imHDR is not far from the absolute maximum, even if it is roughly close to the maximum. In this case you can check again by opening the iris by e.g. 1 / 3 stop. The capture is considered good because on the one hand the pixels of this flame are not clipped, as one would want, and on the other hand the flame is a bright scene object, so it does not capture too few photons of other objects in the scene, which would leave darker objects noisy in the scene capture / image.Note that the sun and some sunlit clouds may be brighter than the flames, but the Human Interaction UI allows you to specifically select these as not the maximum values ​​to capture (i.e., they will be poorly captured and clipped (see RW range in Figure 3)) if the initial original HDR image and the settings used for it are not yet snug and accurate. At least one additional original HDR image is taken. This too is done either by software operating automatically or, advantageously, under human supervision. For example, software knowing that this tapping indicates the selection of a near image gamut top image, if that color component (at least the largest color component) is below the value corresponding to half-pixel filling, the software will choose, for example, to double the shutter time, i.e., increase it by one stop (provided that this is still possible given the required image repetition rate of the camera). If there is clipping, a secondary image can be taken at, for example, 0.75% of the previous exposure. Finally, the original image is captured, in which the (selected) brightest object is Near overflow If at least one color subpixel (i.e., any corresponding image pixel close to clipping but not yet clipped) is actually captured (whether DN, luma, or luminance, when the full range of luminance is related to the DN vs. luma range), the HDR capture situation is considered optimal, and the optimal values ​​of the base capture settings are loaded into the camera (i.e., the primary camera that handles system color blending initialization in a multi-camera system). This is the case for the scene. Because of this positionIn principle, these settings are valid only for this position in the scene. There are scenarios where you can choose one set of capture values ​​(i.e., one or more of iris, shutter time, etc.) for all positions used in this shoot (either when the scene does not have a very large dynamic range compared to the camera's capabilities, or when using a camera with a very high sensor dynamic range), but in other situations you may want to adjust the various capture settings. Preferably, use different optimal capture settings for different positions, rather than partially optimizing either the darkest or the brightest scene colors. For example, if you need to get a good capture of a criminal hiding in the darkest shadows, clip some of the flames (keeping in mind that the captured values ​​will need to be brightened in post-processing). However, if possible, it would be useful if at least these basic capture settings were kept the same for the entire shoot, i.e., for all significantly differently lit positions throughout the scene (this is usually done using brightness allocation / remapping). function , at least for those that decided on a lower dynamic range grading for the secondary. do not ). So there is a stage of integration of the basic capture settings. For example, if for the first part of the scene 1 / 100th of a second is considered to be a good setting for the shutter, and in another area with brighter objects 1 / 200th is decided, the camera will load 1 / 200th of a second for all captures of all positions in the scene (if there is not too much unwanted noise, this situation will be optimized for the brightest object of interest across all shooting positions). This is because the overall shooting environment Overall desired brightest objectThis is because the grading factor determines: the shorter the shutter time, the lower all the digital numbers in the capture. Also, depending on the luminance and luma assignment, these values ​​will also vary, even if in a different, possibly non-linear way, but this is not an issue since the dynamic range was faithfully captured by a high-quality HDR camera (ease of manipulation is a preferred characteristic over having the best capture of each individual position, which is anyway too high for many uses in many situations). Optimization of this graded version of the capture resides in optimizing the mapping function.

[0034] In the explanation (as explained, without wishing to be limiting), a single optimal set of base capture settings can be determined after initialization (so that the adjusted mapping function pre-settings can be focused on), i.e., by walking to various positions in at least difficult lighting. The camera operator discovered the scene , assumed to be determined after applying the present technical principle, those skilled in the art can also understand how several sets of basic capture settings can exist for various corresponding positions and how these are loaded, just like the functions that map luminance to obtain various (at least one) output gradings are loaded for calculation from respective storage locations when a camera operator walks in and starts shooting at the corresponding positions during actual video capture, for example, in a movie, so that when walking in (or starting shooting at) a slightly darker room in the overall shooting location, the iris is immediately reset to the determined value for that location.

[0035] The different illuminations include: Typically, illumination from at least one light source onto a scene object; How muchIt starts with the illumination and assigns them a brightness value. For example, in outdoor shooting, the sun may have a large contribution and the sky a small contribution, which can give the illumination level of all diffuse objects from the sun (in which case brightness depends on the object's reflectivity (whether the object is, for example, black or white)). While indoor position-dependent illumination depends on the number of lighting fixtures, their position, orientation (light fixtures), etc., HDR capture also requires including local illumination, or more precisely, "outliers," in determining the lighting situation. As explained, all or most of the pixels of, for example, a bright flame or other bright object with thousands of nits should be considered when determining the basic capture settings, but in large areas of bright areas, especially outdoors in sunlight seen through a window, it is desirable to capture them at least below clipping and perhaps with a lower sensor pixel fill percentage. Small specular specks (e.g., a light bulb reflecting off metal) may be clipped in the capture. For simplicity, we will describe it as if only position is important, but those skilled in the art will understand that in more specialized embodiments, the camera's directionIt is understood that the color / brightness of the light can also be taken into account. As illustrated with FIG. 4, a first camera 401 in a first position can see different color / brightness configurations when angled toward the interior of a room (in this example, the brightest object is a flame 420, but it could also be a dim object that is much dimmer than the outdoor object), but when facing forward, it can see the outdoor world through a window (410). The light bulb 411 is a small object that will typically be clipped in any image (and sensor capture). The same applies if the oval luminaire 421 were large, although with such a large luminaire, it may still be desirable to have gray values ​​that vary from the outside to the center (so that there are no ugly "holes" or "blotches" in the image), at least for the camera's highest dynamic range grade image output. An acceptable decision would be not to clip all pixels within the ellipse, but to clip, for example, the brightest 10% in the center, so that the rest of the lighting fixture still shows the gradation (the viewer is not usually focused on the lighting fixture, but it may be good to include this information in at least one captured graded image with action in the shot, e.g., for later image processing). Normal objects with "medium" lighting, such as a kitchen 412, a portrait 422, or a plant 423, are automatically fine in the basic capture if the camera(-system) is set up according to the procedure described (these objects (e.g., portraits) may be more important in the primary and secondary grading output from the camera). That is, they may be fine in the technical sense, in that a moderate subset of codes represents the colors of various objects (because the camera is a linear capture, i.e., the higher lightness range is not numerically worse than the darkest color range), but In a visual sense (visual impression on the viewer)However, this is not necessarily without problems in that a "simple" assignment will not necessarily produce the desired intra-object contrast in all derived grade images. A more important object for a human operator (or automaton) to see is the black fire poker 424, which is in a shadowy area of ​​the room (where the light from the elliptical light fixture is blocked by the fireplace, which similarly blocks the direct light from the flame). This capture may be too noisy, in which case it may be decided to open the iris and shutter more, potentially losing the gradient of the elliptical light fixture but at least resulting in a better quality capture of the poker (which would require light boosting, for example, if SDR output video is desired).

[0036] Master Capture (RW) is an image with the correct basic capture settings (the most recent of at least one high dynamic range image (o_imHDR) that led to such settings). From this image of the current scene position and / or orientation, at least one grading is determined to produce, for example, an HDR video image as output, which is not simply a scaled copy (by a linear multiplier) of the capture, but typically has a better luminance position (value) along the luminance range of at least one scene region (e.g., slightly darkening the brightest objects, or making important objects at a fixed level, e.g., 200 nits, or slightly brightening the darkest captured digital number, corresponding to what it would have with pure scaling, such as linearly mapping the maximum digital number to the maximum image luminance and all lower values). Typically, optimal representative values ​​(e.g., intermediate subrange values) are determined for several objects or regions, at least the important ones (portrait, fireplace, kitchen sink, etc.), and the shape of the function is obtained by mapping a subrange of input values ​​to a subrange of output values ​​(e.g., luminance to luminance, or luma to luma accordingly).

[0037] Although the grading is usually not an image with an object luminance pixel ratio equal to the corresponding digital number ratio (i.e., even after subtracting a black offset to compensate for the different starting points for each selected set of two pixels L1 / L2=DN1 / DN2 from the image), there are different schools or application designs that the present technical solution must be able to accommodate.

[0038] The first application creates a primary HDR grading that slightly remaps the luminance location that the digital numbers get by simple max-to-max scaling (i.e., mapping the ADC max or any first range max to the second range max, e.g., 1000 nits). For this redistribution / remapping from pure scaling, the color blend director can use a simple shape function, e.g., a power function, for which the color blend director may still want to adjust the power value. In some scenarios, this may be sufficient (at least for one of the possible graded image outputs).

[0039] However, FIG. 7 illustrates (non-limiting) a typical simple grading control example for quickly establishing the luminance mapping function for primary HDR grading: Fs1 and Fs2 for each of two typical exemplary locations (i.e., assuming the basic capture settings are determined the same at all locations, e.g., when an elliptical luminaire just begins clipping to sensor and ADC maximums). For truly good grading, the color composition director, upon receiving this 2000 nit ML_V-defined HDR output image (ImHDR in FIG. 2; in the first part of the scene discovery phase, this image may have initially been used as image o_imHDR to establish as the basis for the correct camera capture settings, but now, given the camera shots under the determined capture settings, one or more HDR images will be generated for final grading with the optimal mapping function per location), will want to see relatively accurately how all luminances will appear on a 2000 nit display in a typical viewing environment. Note that the same principles generally apply when only a primary SDR grade is output, especially when longer luma and chroma word lengths (e.g. 10-bit) or representations of other three color components are output, with different subranges of different scene image regions mapped to different relative positions (e.g. brighter dark pixels), but ideally, at least in the future, an HDR grade would always be generated as the first grade in an HDR camera.

[0040] Indoors are relatively complex environments, with several different light sources (outdoor kitchen lighting through windows, elliptical light fixtures, additional lighting from flames, shadowy areas, etc.) In contrast to SDR photography, which only considers basic average light levels (and adds filler lighting in the darkest areas given that general base lighting capture), good quality HDR photography requires: Various differently lit sub-parts of the scene (e.g., an object on a desk under a strong spotlight, an object seen through an open door in an unlit adjacent room, a self-illuminating object, etc.) should be consideredAll of these should have the correct visual impact on the end viewer in the final grade, or at least a reasonable impact (so that, for example, objects of little importance to the story don't attract all attention). To create beautiful HDR movies, much more attention is paid to the lighting configuration of the scene (rather than just making the lighting "suitable for capture"), and consequently to the primary and further grading. However, to be sufficiently accurate yet relatively fast (because there may be time before the actual shoot, but not too much, or because the average consumer may not care about too many manipulations), the indoor position function shown in the graph above is controlled with three control points in the example described. Advantageously, the director first establishes some good floor values. The guideline used here, as mentioned above, is to avoid mapping the brightest objects in the scene, i.e., 2000 nits, close to 65000 digital numbers, and to see where all other luminances "fall off" (linearly) below this. The idea is to give the darker objects in the scene in a 2000 nit ML_V grading a luminance that is roughly what they would be in a 100 nit SDR grading, and slightly brighter (say by a multiplicative factor of 1.2), so that the brighter subset of dark objects (as determined by the color blending director) end up at, say, some 100 nits.

[0041] Indeed, in this purely illustrative example, the director decides to select the first control point CP1 to determine the painting's luminance value on the HDR luminance axis (shown vertically) for grading. If the portrait isn't heavily lit by an elliptical lighting fixture (a level of intensity that the viewer wants to see clearly in this 2000-nit HDR video, but not in an excessive way, otherwise the portrait might distract from the action of the actors or presenter), a value of 200 nits would be appropriate. Under normal lighting levels (average indoor lighting for this configuration of lighting fixtures present in this scene), the portrait's pixels are given a luminance of ~50 nits. For this HDR grading, the director decides to map the portrait's average color (or pixel or set of pixels to be clicked) to, say, 200 nits to create the impression of additional brightness. The second control point CP2 is used to determine dark black (poker). A good black value is 5 nits. These two points already determine the first part of the first luminance mapping function Fs1 for this indoor location: the brightening segment F_Bri. For example, such a multi-linear mapping can be determined, or a smoother one (e.g., a parabolic smoother section with connecting straight line segments) that passes through or substantially follows the multi-linear curves (e.g., tangents). As mentioned above, the fireplace flame is also an object of interest. In the previous sub-process, we have already ensured that it has been captured with good quality. Next, in the grading stage (the image processing circuit 207 Actual shooting It always runs on the fly for each video image captured in time sequence during Standard GradingIn this example (i.e., the function to be used is established), the color blending director's criterion is to ensure that the flames have a good HDR impact, but are not overly bright. The contrast between objects depends on the selection of other areas of the scene. Therefore, an additional third control point CP3 can be introduced (e.g., by clicking on the displayed view showing the function and (representative) brightness on the vertical / output axis and a digital number on the horizontal / input axis), and the director can move it to output the desired flame brightness, setting it to, for example, 600 nits in the HDR primary grading (i.e., ImHDR). This establishes a second segment (F_diboos), which allows darkening or boosting a second selectable pixel brightness / color subrange (typically, color processing is usually 3D, but hue and saturation are largely maintained between input and output; that is, only the color component ratios, brightness, and the common amplitude factor of the three color components are changed). To keep the specification simple, the rest of the function is determined automatically. For example, the darkest color segment, F_zero, can be established by connecting the first control point to (0,0). For the top segment, two options can be chosen: this segment can continue the slope of the F_diboos segment to generate the F_cont segment, or an additional relative boost by F_boos can be applied to the brightest colors by connecting the ADC output maximum (65535) to the HDR image maximum. (In this example, the color composition director or camera operator considers the 2000 nit ML_V HDR image to be a good representation of the filmed scene.) The choice of the brightest segment of this mapping function depends on the luminance of objects present at other locations throughout the filming location to obtain a better alignment of various objects throughout the movie or program (e.g., the contrast between the luminance of a sunlit pixel outdoors and the luminance of a selected flame, especially an elliptical lighting fixture). The decision also depends on what luminance is present or absent in various environments, such as the choice of clipping point (or, for example, may be temporarily present if you walk into the right part of a room with a mirror reflecting the outdoor environment).For example, in this example, although not absolutely necessary, it would be a good idea to make the brightest objects (i.e., the brightest parts of the lighting fixtures, even if some of these are street lights in night scenes) at various locations equal at all locations (2000 nits in the example) (note that typically there is a fixed ML_V for the entire movie or program as a graded version (e.g., HDR output movie)). However, the function specification also allows for the opposite desideratum: if for some reason you want the street lights in an outdoor night scene to be only 1000 nits (e.g., to reduce glare to the viewer in the darkest areas), this could equally be done by specifying, for example, a function Fs3 for a nighttime exterior shot that ends with a small vertical segment that maps the brightest digital number (in the output video of the first graded version with an ML_V of 2000 nits) to an HDR output luminance of 1000 nits, saving it to memory, and loading it for shooting. It is up to the video creator to decide how "perfect" the HDR look should be, how much effort should be spent pre-setting all the cameras accordingly, how good the grade(s) coming out of the cameras are already (e.g., ready for viewing), or how much post-work should be done to get a better-looking grade according to the creator's designator. The basic techniques presented allow for a range from such a simple scenario of a function that depends on only two positions with relatively simple (i.e., coarse grading) shapes, to many different functions for many different positions and shooting scenarios. However, two examples are sufficient to illustrate the principles, as those skilled in the art will understand, mutatis mutandis, other variations.

[0042] Moving to the next outdoor location (again, using the same basic capture settings in this example), and this time assuming a daytime outdoor shot in the illustration of FIG. 7 , determining the optimal or desired second luminance allocation function Fs2 for this simple, shadowed, primarily sunlit environment can be done faster and easier with only two control points (the third is implicit, i.e., requires no user interaction, and is located at the maximum value of both axes). As shown in the bottom graph of FIG. 7 , the house has a somewhat lower scene luminance, and therefore a lower digital number, due, for example, to clouds moving in front of the sun. In principle, the technique could determine separate (additional) luminance allocation functions (and typically a secondary grading function, if desired) for these lighting situations (despite being in their existing positions), but the idea with this approach is that this is not necessarily necessary because, when properly configured, low values ​​will scale appropriately (representing the actual darkening in the scene and its final appearance (e.g., after standard display adaptation algorithms) as a reasonably corresponding darkening in the output image). That is, by selecting an appropriate mapping function for a smaller subrange of the displayed image's luminance, an emerging sun will result in the appearance of a brighter object—for example, twice as bright in the HDR grading (depending on the slope of the mapping function for that subrange of the DN, which in turn typically depends on the selected ML_V of the output-graded video). And vice versa, a returning cloud will appropriately dim the house, now in shadow, in at least one output grading. In this scene, the color composition director focuses on two aspects of the grading: first, they want a good value for the house in shadow (even if they can't actually measure both situations, the camera operator and director know that there is a fixed relationship between the sunlit house and the shadowed house, since people usually don't go into as much cloud cover as possible when shooting the same sunny day, so a slope of 1 / 5, for example, can turn a 10x scene luminance change into a 2x grade image luminance change).Because they are outdoor objects, these houses are selected to be brighter in the 2000 nits master HDR grading output than the indoor, strongly lit (e.g., 300 nits portraits (on average)). They therefore appear about 200 nits darker than the houses when lit by sunlight and viewed from an indoor environment. Of course, in principle, a director could decide to place more emphasis on the lightness of the portraits and even grade that part of the scene with a locally different function. However, while not precluded in itself, such complex grading is not feasible (at least in the primary grading output) as it is fast and easy (single or multiple). Camera Capture Optimization System (for example, 2000 nits) is anomalous. The master (e.g., 2000 nits) HDR output video is already graded in terms of placing all object luminances in positions that make sense for HDR home viewing, but it usually still has a relatively simple relationship to the relative brightness captured in the camera's digital numbers (though not simple enough to be a fixed function). Thus, if a scene object point measurement of 1000 DV is less than 25000 DV, then no matter what function you choose, the function will be strictly increasing, and so the output nit value of the grading will also be lower. Secondary grading could in principle involve image location-dependent mapping, but in practice, with single image source material (no mixing of different images from different contributors), one function for all pixels in the image, regardless of their location, has proven to be sufficient for multiple usage scenarios.

[0043] While the director can independently select the function for the second position, it is advantageous if the image color composition analysis circuit 250 includes a circuit that presents several, at least two-position based, grading situations for display. For example, it can send a split view to the display 235, placing a rectangle half the image width on the master capture RW captured at the previous position to compare the brightness of various objects in the current position's grading. For example, the director can drag a selection of the portion of the image containing the oval light fixture and portrait and move it next to, say, a street light (which can also be moved) in another grading at another position to easily compare those objects side by side and see how their interior brightness has been adjusted. For example, the director can first select the right side of the indoor scene in Figure 4 and compare the brightness appearance of the fireplace, portrait, plant, and wall indoor objects with the ground, house, and house in the outdoor second position capture to determine whether, for example, the viewer will not be surprised if the video is later recut by quickly switching from the first position shot to the second position shot. Sometimes, due to the fast pace of modern production, there's a constant switching between two positions, say a speaker, which shouldn't look like a disco strobe. Then, by shifting and swapping, you can compare the brightness of an outdoor house seen through a window in an indoor scene with the outdoor house and objects in an outdoor shot. Note that if you also want to compare the brightness levels of indoor objects in both indoor and outdoor shots, if you're using only one master capture RW per position (rather than several positions), you should ideally make sure to get some of the indoor objects in view as well (or you could mix different objects from multiple outdoor captures in one half of the screen to compare the outdoor objects with the outdoor objects in the other half). In this example, there's a stool somewhere in the hallway. It will have a brightness that depends on the type of lighting in the hallway and the time of year it's outside (e.g., it's sunny in summer).So in the discovery phase, you can choose to either look only at indoor objects and treat outdoor objects as secondary locations, or you can have some of these already captured in the first capture of the primary location and determine the function accordingly. A primary location for calibration might be, for example, where most of the shots are taken, e.g., an auditorium for a business or educational film with only a few outdoor scenes interspersed.

[0044] For example, the director can see the indoor objects from the outside. Looks dark but Sufficiently visibleSuppose we want to achieve a certain brightness. This can be achieved by placing the object at, say, 15 nits. Here again, we understand the importance of the grading function compared to purely physical measurements. We can say that we give indoor objects the same digital number, using a single set of pre-defined capture settings, regardless of whether they are in view in an indoor or outdoor shot. However, it matters to the viewer whether the majority of the pixels are indoor pixels (establishing the basic brightness look of the shot) or whether they view only a few of these objects through an open window, with the remaining image pixels imaging sunny outdoor pixels. This modifies the selected F_diboos segment of the second brightness mapping function Fs2. The other two segments, F_zer2 and F_boos2, can be obtained automatically by connecting the respective control points to the ends of their respective ranges. If the function works well enough for the director—that is, it creates a good-looking graded image—it does not need to be further tweaked (e.g., by adding a third control point) and can be sent to the camera and stored in the function memory in the memory portion of the second function. Using these examples, you won't get the most perfect graded output you can imagine, but the point is that you can get much better results than using a simple fixed allocation function for every shot, but it's still simple enough to be practical in many shooting scenarios (unless you know what environments you'll encounter, like an unpredictable race of different people through different parts of the world). You can still do some pre-discovery settings for environments that look similar enough to the ones you'll encounter to get one or more graded outputs that look good enough and are optimized (e.g. 2000 nits ML_V HDR and 100 nits SDR).

[0045] Often two or three control points are sufficient to re-grade (both the primary / master grading and the secondary grading, if applicable (i.e., selected as the parallel output video of the camera, for example)), but the method or system allows the director to choose as many functions as he wishes by continuing to add more control points to change the shape of the function. However, this also depends on the type of shooting, for example, how many cameras have static angle looks in a somewhat complexly lit location, etc. Also, often only a few positions, even just two positions, are sufficient. For example, create a "coarse" grade situation that is well suited to all brightly lit environments (e.g., for all outdoor locations, whether in full sun in the open or, for example, in a shadowy location between tall buildings), and a secondary grade situation for all locations that are virtually dark (about 100 times darker in the real scene, but which will require some adjusted grading anyway), such as all indoor locations (which will convey some relative darkening so that the end viewer can see the difference between ideally bright outdoor captures in all gradings, but on the other hand, both locations will produce graded images in which all or most objects are visible (i.e., not too dark, not obscured, not clipped), and also ideally convey well-adjusted brightness and color). When you want to create a perfect HDR impression, for example when there is a specially lit decoration created intentionally for the HDR look of a movie, but even when there is a variety of different lighting environments "by chance", such as in a show about escape rooms, you will want to fine-tune the various lighting situations in many different shooting positions in order to capture this technical situation in the best way, even without considering a specific light impression for the final grading.

[0046] This initialization approach creates a camera or camera-based capture system that performs much better technically: when starting a shoot, the user can focus on aspects other than color and brightness distribution and composition, but the colorimetry analysis still doesn't translate to a very simple technical formulation, but this time it allows the human video creator to work with a sophisticated formulation that allows for a simple, short initialization pass.

[0047] Advantageously, the method for setting a video camera capture mode in a video camera (201) further comprises: The color synthesis director calculates the corresponding second grade image (ImRDR) from the luminance of the master capture (RW) or the first grade image for at least two positions. Secondary determining grading functions (FsL1, FsL2); storing the secondary grading functions (FsL1, FsL2) or parameters uniquely defining the secondary grading functions in the camera's memory (220); The second grade image (ImRDR) has a lower maximum brightness than the first grade image.

[0048] Those skilled in the art will understand what the relationship between functions will be, for example, if there is a function relating primary luminance, i.e., luma, and digital number, and a function relating primary luminance, i.e., luma, and secondary luminance, i.e., luma, then what the relationship will be between secondary luminance, i.e., luma (via the selected EOTF) and digital number. This is because simply relating input and output values ​​along that range determines the shape of the function. Those skilled in the art will also understand how such a set of duplets can be parameterized, for example, by the slope value of a polylinear approximation. Sometimes, it may be technically advantageous to use a function relating secondary luminance, i.e., luma, to primary luminance, i.e., luma, rather than DN. This is because this function is advantageously communicated jointly as metadata, for example, in SL-HDR2 or even SL-HDR1 format (after conversion). However, other applications may desire only SDR output images, i.e., pixel color matrices, without the need to communicate the functions that generated these images. Therefore, functions FsL1 and FsL2 are typically only used internally in the camera when output from the current capture is required.

[0049] Some cameras, for example, need to output only one grading for dedicated broadcast (e.g., an HDR grading (e.g., with a 1000-nit ML_V target display associated with the video image) or a classic standard dynamic range (a.k.a., LDR) output). It is often useful if the camera can output two gradings, which are already readily correct. It is even more useful if these are already in a format associated with these two gradings (e.g., the applicant's SL_HDR format (standardized in ETSI TS 103 433)). This format can output video of SDR images, e.g., as images to be transmitted for consumer broadcast or narrowcast, and can output functions to calculate, e.g., a 1000-nit HDR image from the corresponding SDR image. These functions correspond to the functionality of the present method / system, as described below. One advantage is the ability to serve two categories of cable operator customers: those with legacy SDR televisions and those purchasing new HDR displays.

[0050] So if your primary grade video is, say, 1000 nits HDR, you can't get it to, say, SDR video. SecondaryThe secondary grading functions (FsL1, FsL2) can operate directly from the master capture RW, i.e., from the digital number, or, more advantageously, map the primary grading luminance to the secondary grading luminance. In a second alternative, during the setup phase, for each location (and possibly several orientations), a first-grade image representative of the luminance of all image objects and a second image representative of the luminance (i.e., pixel array) of these image objects are determined, and for each luminance in one image, the corresponding luminance in the other image is determined. This function is then used to calculate on-the-fly a luminance-to-luminance mapping of the input pixels of the first-grade image captured during filming, resulting in the secondary-grade image as the secondary output image. This is all done by the image processing circuit 207 almost simultaneously with filming. While some believe that secondary grading should simply be a simple derivative of primary grading, it can also be argued that both grading methods are equally important and attractive.

[0051] The existence of such a configuration system for (in) the camera will be discussed later. During the actual shoot , meaning that the camera can be operated given this configuration. This is made possible by a method for capturing high dynamic range video in a video camera. The video camera outputs at least one (and possibly more than one) graded high dynamic range image in which pixels are assigned a luminance value. The method: applying the method for setting the video camera capture mode as described above in a video camera (201); determining a corresponding position among at least two positions (Pos1, Pos2) for a current capture position; loading a brightness allocation function (Fs1) corresponding to said location from the camera memory (220); and applying a brightness allocation function (Fs1) to map the digital numbers of successive images captured while capturing at the current position to corresponding first grade images (ODR; ImHDR), and storing or outputting the first grade images.

[0052] handlePosition means that the camera operator (or another operator operating a second camera) is not standing in exactly the same position (or orientation) in the scene selected during initialization. The shoot must be free to take place. This means that the person is near the surveyed and saved position, which usually means that the person is, for example, in the same part of the scene and in the same lighting conditions. This can mean different things. Outdoors under the same natural lighting, the whole world is a set of corresponding positions, at least when the system is operated in such a way that, for example, the color composition director does not choose to distinguish between different outdoor positions (for example, if this technical role is actually played by the cameraman, e.g., if the camera is operated and guided by a single amateur videographer). Indoors, the lighting can be more complex because lighting fixtures are ubiquitous, room shapes vary, and there are shading objects that cast different shadows. A good example to illustrate this is a shoot where the director intentionally wants to shoot in one well-lit room, one averagely lit room, and one dimly lit room (which has only indirect lighting, for example, through a half-open door from an adjacent averagely lit room). To have a maximum HDR appearance. With the present camera (or a system of devices including the camera), you can adjust the brightness in one or more grades of all these rooms as professionally and accurately as desired, and then freely walk through them (e.g., following an actor). Even if the door is opened a little more, you will still get both a good quality SDR grade and an HDR grade, as long as the system is properly initialized for the dark corner of the darkest room. Various embodiments have additional technical elements in the camera that allow it to quickly and reliably determine what position, i.e., what lighting situation, you are in at each moment of shooting. Essentially, the present system aims not at the lighting situation (i.e., not necessarily the lighting itself, let alone the average light level), but rather at the aspects of lighting that are important for the representation of the image, i.e., the light moving towards the camera from multiple areas and objects of the scene, i.e., the scene brightness at various object points of interest.It's not so important to ensure that the shape of the object is exactly the same in any given capture—that is, for example, that the circular set of pixels is the same. Rather, it's important that the same, or roughly similar, intensities occur to form the configuration of the region of the image that is captured and later graded. Thus, the geometric patch of pixels whose intensities are distributed around L_obj (while appearing in the same position) may shift to the right in a later capture due to a shift in the camera operator's position to the left (moving partially out of the captured image region, potentially halving its area). Also, the camera operator may move back or undergo a perspective transformation, such that the camera operator suddenly appears behind and is partially occluded by an object in front of them, reducing the imaged region. Therefore, what's important is that the shape of the intensities around L_obj is present in all of these captures, and not suddenly move farther and closer together, as in typical studio captures, where the region grows from nearly the entire image at one end to just a few pixels at the other (taking into account the minimum focal length of the lens as well). While these scenarios are not necessarily problematic with the present technology, the viewer's impression may not be fully preserved in the grading, which is usually not a matter of sufficient concern. Various technical factors can help determine this similarity. For example, image and / or recognition methods can be used, which should naturally be of the type that can reasonably recognize location / location situations characterized by, for example, five major distinct luminance regions (e.g., flames, normal objects, dark objects in shadow areas, lighting fixtures, and sunny outdoor scenes). In some scenarios, the mere presence of such subranges of luminance may already identify at least two significantly different locations, but more robustness can be achieved if, for example, objects in the bright range are also verified to be between dark and mid-luminance objects.Additionally or alternatively, other technical means of determining position can be used that imply the same luminance configuration as long as the scene does not change significantly (even explosions usually work within a set of ambient luminances, either precisely or roughly; even in such cases, our approach of course does not need to be able to perfectly handle 100% of all possible shooting situations in order to improve upon a fixed SDR approach, since an explosion, for example, would clip an inordinate number of pixels even for an impressionistic HDR grading).

[0053] The functions can be stored in a data structure together with the position information (and possibly other information related to the function, such as the maximum luminance of the luminance output range or the luminance input range), which can take several coding forms. For example, they can be enumerated and labeled (position_1, position_2), or they can be absolute, related to, for example, GPS coordinates or other coordinates of a positioning system, or have semantic information useful to the operator (e.g., "basement," "center of the music stage facing the hall in performance conditions," etc.). Although not necessarily required, it is also useful to have precision, such as centimeter or decimeter accuracy (although aggregate characteristics as precise as a meter are often sufficient), for example, for indoor or studio use with a contiguous subset of differential global navigation satellite systems, or any positioning system with similar capabilities. This data format can be communicated to various devices and stored in memory locations for use in various user interface applications, etc.

[0054] Advantageously, the method is used in association with a camera (201) that includes a user interaction device, such as a double-throw switch, for toggling a function memory location of saved positions of different illuminations in the shooting environment, which when pressed in one direction selects the previous position in a chain of linearly linked positions, and when pressed in the opposite direction selects the next position. A double-throw switch is a switch that can be moved in (at least) two directions and activates (different) functions for these two directions. For example, it could actually be a small joystick or whatever the camera manufacturer deems easy to implement, typically located on the side or back of the camera.

[0055] This is a simple and very fast method that relies on the user to quickly select a position. For example, and especially in some locations, the sequence is easy to remember, and the camera operator can flick a switch just before entering an indoor area. In some embodiments, for example, the device uses aggregating brightness scales that begin applying a new function from the moment the device sees the first capture where the number of photons drops (or rises) significantly, signifying that the operator walked into a less lit position during that capture (e.g., walking into a door, having cover from the ceiling or sidewall, etc., or, in a musical performance, going from facing the stage to facing the audience behind, which is the earliest need for a change in the secondary brightness mapping function, e.g., to create an SDR output feed). Where smoother transitions are desirable (e.g., taking into account the rate at which outdoor light fades due to entrance geometry, or simply where sudden changes are generally acceptable, but where no jarring changes are generally required (advantageously, the system does not need to do more than the adjustment time of a classical auto-exposure algorithm), see below for transitions over a longer range of lighting conditions, such as in a hallway), a delay of several images can allow for advanced temporal adjustment of brightness.

[0056] Advantageously, the camera (201) includes a voice recognition system to select a stored luminance allocation function or secondary grading function based on a location-associated description, such as "living room." This frees up the camera operator's hands, which is useful for configurations such as changing the viewing angle of a zoom lens. Talking to the camera uses a different part of the brain and therefore interferes less with important tasks. In the case of (slightly) recut capture, the camera operator can freeze for a moment while selecting a new location, and those images with audio are cut out from the final video production. However, real-time measures may be in place so that the camera operator's whispers are barely picked up by the main camera microphone. This can be achieved, for example, by having a set of beamforming microphones with the main receiving lobe pointed toward the camera operator, i.e., the audio capture lobe is behind the camera (while the main microphone 523 is focused on the scene where the presenter or performer is taking place, i.e., capturing from approximately the other side). Also, other cameras in the scene are placed far enough away from the whispering camera operator so that the camera operator's voice is barely recorded, or at least not perceptible, and does not need to be removed in audio processing. The camera operator can train whispered names of locations while capturing one or more high dynamic range images (o_ImHDR) and use names that are easy to distinguish (e.g., "shadows under the trees in the forest" is the longest name you might want to use for quick and easy navigation; "shadows of the trees" is more appropriate when there are not many locations that require complex descriptions to distinguish them, and "tree line" or "edge of the forest" are other possible locations, for example, when one half of the hemisphere is dark and the other half is brightly lit).

[0057] In systems or situations where this is not possible, other techniques (embodiments) can be used.

[0058] Advantageously, this method / system uses location beacons that are fixed in commonly used locations (such as studios) or can be suspended before filming (for example, in a person's home that has been scouted for interesting decoration). This can be a simple beacon that provides, for example, three different ultrasonic or microwave electromagnetic pulse sequences starting from the second, and the camera (201) includes location determination circuitry based on triangulation or the like. There can be one beacon per location, and when properly positioned, the camera can detect which beacon is nearby based on the arrival time one second after the clock. Alternatively, the camera can emit its own signal to the beacon and wait for a return signal or the like.

[0059] You could also stick a rapidly flashing pattern of LEDs on the ceiling, and if they're infrared and the camera has separate infrared detection, it doesn't even matter if they're in view (if the video sensor blocks IR), and the pattern can identify a room, sub-area of ​​a room, etc.

[0060] Alternatively (or in addition), the camera (201) also includes a localization system based on analyzing images captured near each of at least two locations. Monitoring the amount of light at each location is highly useful. The automaton itself can detect when the light scale has changed sufficiently to summarize the situation or approach the situation at a certain location. In some situations (e.g., in a disco or explosion, where the lighting changes in a complex manner), the system may rely on geometric recognition of color (texture) patches. For example, a red sofa against green wallpaper or a rectangular shape may be recognized as being in one room (this facilitates identification if the wallpaper contains a specific object, such as a printed flower), but not as being outside, for example. In areas where photography is frequent, this can be accurately trained. Even with fast initialization, this information is quickly collected by the camera (see below). This often allows for robust room identification. The advantage of this technique is that high-quality imaging sensors and, in some cases, image processing capabilities (which can be reused for other purposes) are already available. The downside is that more complex algorithms may require a dedicated processor and additional processing capabilities and power over existing typical camera image processing functions. However, as integrated circuits (ICs) become more powerful every year, this may be an option for future cameras (e.g., in mobile phones, which are becoming powerful computers anyway). Future devices like phones are expected to have more processing power and possibly even neural processors. Also, because cameras are already more regularly equipped with Wi-Fi and 5G connectivity, these stations can also be identified (e.g., triangulation, time of flight). If image analysis processing is performed by external means, such as the cloud or on-set computers, the camera's capabilities and upgradability are less important. Low-quality, low-resolution communication images may be sufficient to identify the shooting location.

[0061] While this innovative concept is already very useful for on-the-fly single-camera capture, it has the potential to become even more useful and powerful for multi-camera filming.It would be advantageous if a method of setting, at a secondary video camera (402), a video camera capture mode that specifies output pixel intensities in one or more graded versions of an output video sequence of images includes the steps of setting, at a first video camera (401), a video camera capture mode that specifies output pixel intensities in one or more graded versions of an output video sequence of images, communicating between the cameras to copy a group of settings including iris settings, shutter time settings, and analog gain settings, and also communicating any of the determined intensity allocation functions from the memory of the first camera to the memory of the second camera.

[0062] For example, the first cameraman discovers the scene with his camera and generates typical functions for several interesting lighting positions in the scene. Then, just before starting the actual shoot, he downloads these settings to the other cameras. While it is advantageous if all cameras are of the same type (i.e., same manufacturer and version), this approach can also be used with cameras that behave differently, ideally with some additional measurements. For example, if the second camera has a sensor with a smaller dynamic range, say a 20,000 pixel full well, and already has 50 pixels of noise, its behavior for a fire in a room can be configured, for example, in relation to pixel overflow. Naturally, this camera will produce noisy blacks (although the video brightness is already well-balanced), but this can be solved by using an additional post-processing brightness mapping function and / or noise reduction that somewhat darkens the darkest brightnesses. If you have to work with cameras that deviate greatly (e.g., cheap cameras that are discarded during filming), you can always use this method twice, with two camera operators independently discovering the scene with the two most different cameras (and other cameras copying functions and basic capture settings based on how close they are to the best and worst cameras, respectively).

[0063] Advantageously, the method for setting the video camera capture mode in the secondary video camera (402) comprises either a step in which one of the first video camera and the second camera is a static camera having a fixed position in a part of the shooting environment and the other camera is a mobile camera, and a step in which a brightness allocation function for the position of the static camera is copied to a corresponding function memory of the mobile camera, or a step in which a brightness allocation function in the mobile camera for the position of the static camera is copied from the corresponding function memory of the mobile camera to the memory of the static camera.

[0064] This is a classic setup such as in studio broadcasts (i.e., studio cameras, but it could also be static cameras in, for example, field productions of sporting events). In a classic setup, several cameras can be placed at convenient locations on the studio floor in front of decorations (sets, scene design) to capture, for example, parts of the set or speakers from different directions, but fixed cameras can also be attached to the ceiling of a room for a panoramic view, for example, during a film shoot. If that static camera sees all or most of the relevant objects in its view, for example, of a living room adjacent to the kitchen (which is considered a single free-range environment for the actors), the function of the static camera can be copied to a dynamic camera that is also involved in the shoot (or at least part of the function of the static camera is copied, for example, where everything except the kitchen window is determined by the static camera, and for example, that part of the secondary grading curve can already form the first part of the secondary grading curve of the dynamic camera, but the dynamic camera determines by itself from its own scene discovery the upper part of the luminance mapping function corresponding to the world outside the window 410). Conversely, the dynamic camera operator (whose role is played by the color composition director when loading one of the determined functions into another camera, or by the camera operator when copying at least one function from his camera) passes by the static camera and copies to it at least one appropriate function (and usually also the basic capture settings, such as the iris setting for that position). In more advanced systems, the static camera is rotated and two functions are copied to it, for example: one useful for shooting in the direction of the kitchen (which may or may not include outdoor pixels), and another for shooting in the direction in front of the fireplace.This would involve adding a universal orientation code (e.g., based on a compass) to the function data structure, after which the static camera can decide by itself what to use in which situation (e.g., automatically by dividing the angle based on which side of the midpoint direction of two reference angles the static camera is currently pointing on), or the static camera can be instructed by the camera operator via user interface software what to use under which conditions specifically (e.g., a standard camera would send its operating menu to a dynamic camera, so the operator could program the static camera by looking at the options on the dynamic camera's display).

[0065] One advantageous way of embodying the innovative concept is in a multi-device system (200) for configuring a video camera. The multi-device system comprises: setting a capture mode in the video camera (201) that specifies output pixel intensities in one or more graded versions of an output video sequence of images output by the video camera to a memory (208) or a communication system (209); a video camera (201), the camera including a location capture user interface (209) configured to enable an operator of the video camera to navigate to at least two positions (Pos1, Pos2) within a scene having different illumination from one another and to capture at least one high dynamic range image (o_ImHDR) for each position selected via the location capture user interface (209) to be a representative master HDR capture for each location; receiving at least one high dynamic range image (o_ImHDR); and a color synthesis director analyzing the at least one high dynamic range image (o_ImHDR); a) an area of ​​maximum brightness in the image and at least one of an iris setting, a shutter time, and an analog gain setting of the camera based on the area of ​​maximum brightness; b) an image color synthesis analysis circuit (250) that allows determining, via a function determination circuit (251), for at least a first grade image (ODR) corresponding to a master capture, a brightness allocation function (FL_M) of the digital number of the master capture to the nit value of the first grade image (ODR) for at least two positions; The camera includes a function memory for storing parameters that uniquely define the luminance allocation functions (Fs1, Fs2) or luminance allocation functions (Fs1, Fs2) for at least two positions determined by and received from the image color synthesis analysis circuit (250).

[0066] Rather than having all components present in a single camera (which is good for ultimate portability, especially if a one-person team wants to explore an environment, for example, in an urban exploration (Urbex) shoot), it is advantageous if some of the functionality resides on, for example, a personal computer. This has the advantage of, on the one hand, a large amount of computing power and the possibility of installing various software components, but on the other hand, it allows for connecting a larger, higher-quality monitor and more easily shielding it from ambient light (or even putting it in a dedicated, darkened grading booth or OB truck that can be quickly erected on-site). While the capture mode generally refers to how an image is captured, this patent application also specifies how the capture is output, i.e., what kind of, for example, 1000-nit HDR video is output (whether the darkest objects in the scene are rendered slightly brighter or, conversely, kept appropriately dark). This includes the possibility to roughly or precisely specify the optimized brightness of the corresponding grade in at least one output graded video for the brightness of all objects that may occur in the captured scene. Naturally, it is desirable to output several different grades of video (and typically coded differently, e.g., perceptual quantizer vs. Rec. 709) for different dynamic range usage. Additionally, the camera includes new circuitry that allows the operator to walk into a representative lighting environment and specify this by using a capture user interface 210 to initialize capture and specify all data from the camera side (e.g., the image color synthesis analysis circuit 250 present in a personal computer can operate with a third user interface, a mapping selection user interface 252, using which the color synthesis director can specify various mapping functions, i.e., shifting control points, as described in particular in FIG. 7).On the one hand, the color composition director captures a representative image there (a good capture is typically the master capture RW), and on the other hand, it records the corresponding position, albeit with minimal data such as at least a sequence number (e.g., location number 3). Selection User Interface 230 In various embodiments, the camera is specified with more information, such as semantic information, which can be more easily used later. This needs to be managed by the capture UI operation software, which, on the other hand, allows the camera to operate in this lighting environment later. Basic Capture Settings (iris, etc.), and on the other hand, essentially starting from a raw capture of digital numbers, allowing us to calculate different grades of video and their pixel brightness. function This must be managed to be saved in the corresponding location in the function memory 220 (note that the basic capture settings may also be saved there or in another memory). After that, the images first captured in the initialization phase become irrelevant, and all the (usually adjusted) information about optimal shooting at various positions is in the saved basic capture settings and functions. The camera is ready for the actual shoot (i.e. recording an actual talk show, filming a segment of a movie, etc.) An important user interface is the selection user interface (230), which allows the camera operator to quickly indicate which location-dependent setting the camera should switch to.

[0067] A useful embodiment of the system for configuring at least one video camera (200) includes a function determination circuit (251) that determines whether the color composition director calculates, from a master capture or first grade image (ODR), a corresponding second grade image (ImRDR) for at least two positions, two functions. SecondaryThe camera (201) determines the grading functions (FsL1, FsL2), and the camera (201) stores the secondary grading functions (FsL1, FsL2) in memory for future captures. When the camera operator toggles to a new position, both the primary function (FsL1) for calculating the first grade output video from the captured DN (e.g., 3000 nits HDR master output) and the secondary function (FsL1) for calculating the secondary (e.g., SDR) output video for that position are selected for camera operation for the shot at that position. For example, toggling up is the first position, toggling right is the second position, toggling down is the third position, and toggling left is the fourth position. This is user-friendly enough for many shooting scenarios, but if more positions are needed, smarter selections can be assigned to user actions. For example, toggling up could be the next position depending on which position the camera operator was filming at, and toggling down could mean, for example, walking into a room on the other side of a hallway, or a more advanced automatic or semi-automatic system, for example with audio, could be used (e.g., if the camera operator has to physically walk to a far room before continuing filming, the camera operator has more time to change to the preset data for the new position than if they had to be continuously changed by, for example, a running cameraman running after the actors along a staircase).

[0068] A typical secondary grading of any high dynamic range primary grade video is SDR grade video, although secondary HDR video with lower or higher ML_V is also possible. The innovative camera has memory for at least these various functions and their management, particularly the selection of the appropriate functions to produce high quality grade video output during operation. The innovative part within the computer, i.e., the innovative part operating in a separate window such as a mobile phone (which may also function primarily as a camera or secondarily as a camera), has configuration functionality including the usual user interface for the appropriate settings and functions of the camera, as well as correct communication with the camera for various positions (unless the system operates fully automatically).

[0069] Thus, the novel camera itself includes a system for setting a video camera capture mode that specifies output pixel intensities in one or more graded versions of an output video sequence of images, or is configured to operate in such a system by, for example, communicating several HDR captures to a personal computer, receiving corresponding intensity mapping functions, and storing them in respective memory locations. The camera has a selection user interface (230) for selecting from memory a secondary grading function or intensity mapping function corresponding to the capture location.

[0070] Various useful embodiments of the novel camera include, among others (possibly combined with various interaction devices in high-end cameras for selectable operation or higher reliability): A camera including a user interaction device, such as a double-throw switch (249), for toggling between function memory locations of saved positions of different illumination in a photographed environment, the user interaction device selecting the previous position in a chain of linearly linked positions when pressed in one direction and selecting the next position when pressed in the opposite direction. A camera (201) including a voice recognition system and preferably a multi-microphone beamformer system directed at the camera operator to select a stored luminance allocation function or secondary grading function based on a description associated with a location, such as "living room." A camera (201) including a location and / or orientation determination circuit, wherein the location determination is based on triangulation using a positioning system positioned in a spatial region around at least two locations, and the orientation determination circuit is connectable to a compass. 16. A camera (201) according to claim 12, 13, 14 or 15, comprising a location determination system based on analysis of images captured near each of the at least two locations.

[0071] The camera typically identifies shapes of various colors at different locations based on basic image filtering operations such as edge detection and feature merging into distinct higher level patterns. Various image analysis versions are possible, some of which are described in the detailed illustration-based instruction section below.

[0072] While sampling positions, the camera operator can scan the environment, for example capturing images at angles around the main direction and grouping them together as if captured with a wide-angle lens. [Brief explanation of the drawings]

[0073] These and other aspects of the method and apparatus according to the present invention will be apparent from and will be described with reference to the implementations and embodiments described below and the accompanying drawings, which merely serve as non-limiting, specific examples illustrating more general concepts, and which use dashed lines to indicate that a component is optional, while a component not using dashed lines is not necessarily essential. Although dashed lines are described as essential, they can also be used to indicate elements hidden inside an object, or for intangible things such as object / region selection (and how they are represented on a display).

[0074] [Figure 1A] FIG. 1 (in FIG. 1A) shows, on the one hand, a schematic illustration of how one or more camera operators can take pictures in positions with various lighting conditions, and how, in accordance with the present invention, these positions can be examined to obtain corresponding suitable shooting conditions for the camera, i.e., modes that specify, as desired, for each location, the optimal values ​​of the various brightnesses that objects may have in the scene to be represented in at least one output graded video. [Figure 1B] FIG. 1B also shows how a better secondary grade of output video (RDR) can be derived compared to a simple technical formulation of the luma or luminance of the output video. [Figure 1C] FIG. 1C, for example, illustrates the same concept of desirable brightening of dark scenes and resulting dark image objects in a two-dimensional graph to better illustrate the concept. [Figure 2] Figure 2 illustrates, with typical general components (relating to the roles of the humans operating the various devices) what the entire system generally does, resulting in a technologically improved camera (201) that is automatically ready for use in the very different lighting environments illustrated in Figure 1. Although two separate devices are shown, the functionality of both devices may reside in a single camera. [Figure 3]FIG. 3 shows in more detail how there are several ways to create an output video graded to, for example, 700 nits of video maximum brightness, some ways better, some ways less suitable, and the better ways will meet the requirements of the present methods, systems, and devices, particularly the new camera. [Figure 4] FIG. 4 is an example of a complex indoor lighting environment to illustrate some concepts regarding one camera shooting in multiple positions or multiple cameras shooting in multiple positions, and also the potential impact of orientation at any position, which is also considered in more advanced embodiments of the present innovation (the cameras shown may be the same camera operated at different times or different cameras scanned simultaneously). [Figure 5] FIG. 5 shows an embodiment of a more advanced camera that incorporates location-dependent function generation circuitry and / or software and additional circuitry for selecting appropriate location-dependent functions during actual capture, as well as a display in the glasses that allows intelligent on-the-fly function selection. [Figure 6] FIG. 6 shows an example of a user interface for defining a luminance mapping function for creating a primary grading from the digital number of any raw captured video image. [Figure 7] Figure 7 provides further insight into the example shape of the function for a two-position capture example (indoor vs. outdoor), expressed as a continuous video output that is 2000 nits ML_V HDR-grade output video. [Figure 8A-8B] Figure 8 shows how a secondary grading, say an SDR grading, can be graded if you create a primary 2000 nit HDR grading as a starting point. [Figure 9A-9B] FIG. 9 shows a schematic example of how a camera may detect locations by identifying specific color patterns due to the presence of specific distinguishable objects at one or more locations. [Figure 10] FIG. 10 shows an example of a video production from scene (or capture) to end of display. [Figures 11A-11B] FIG. 11 shows another example of adjusting the mapping function to create, for example, a master HDR video version from time-series merged (i.e., edit-cut) shots taken at different times from two locations, taking into account that the capture settings saved for the two locations are different (e.g., the second location has an iris aperture that is one stop more, while other exposure-determining parameters are identical for both locations). DETAILED DESCRIPTION OF THE INVENTION

[0075] FIG. 1A shows an example where a (conceptual) first cameraman 150 and a second cameraman 151 can be shooting in different locations (this could be the same cameraman shooting at different times, or two cameramen shooting in parallel, with cameras initialized and settings copied in accordance with the present innovation). The outdoor environment 101 may be completely different for various reasons. Differently litThat is, they are lit quite differently for a variety of reasons, such as lighting levels and spread (i.e., non-uniformity) of lighting (e.g., the sun shining everywhere on an object at the same angle given distance, or uniform lighting from an overcast sky), or equidistant lighting poles (113), resulting in a lighting profile with lighting dimming somewhat in the middle between the poles. At night, outdoors are typically much darker (and have stronger contrast, i.e., higher dynamic range) than indoor shots, and the opposite is usually true during the day. A representative outdoor object in this scene (which must achieve adequate brightness in at least one graded output video) is, for example, the house 110. Because the grading is intended to be optimized for at least one viewing situation, it is undesirable for a house to appear overly bright, let alone a glowing house, for example, unless it is specifically intended by sunlight, i.e., unless it is intended to be perceived that way by the viewer. Another important object for monitoring output brightness is bushes in shadow 114. During the day the oval area of ​​a street light may have a similar brightness to a house due to sunlight reflection on the cover, but at night it will be the brightest object in the scene (so bright that we want to dim its relative excess brightness (compared to the raw camera capture ADC digital numbers) in the graded output video, e.g. as a ratio to the brightness of an averagely lit object (like a house)). This will result in it being less noticeable or distracting in the viewing-ready grade.

[0076] Indoor objects, such as a plant 111 (or a stool 112), have varying brightness levels, depending not only on the number of lighting fixtures illuminating the room, but also on where the lighting fixtures hang and where the object is placed, but generally the lighting levels are about 100 times less than outdoors (at least in sunny summer weather outdoors; in a stormy winter shoot, some indoor objects may have higher brightness than some outdoor objects).

[0077] An advanced embodiment of the system comprises: Location-Dependent FunctionsThis utilizes variable definitions of the camera settings (and location-dependent camera settings). In a basic embodiment, the idea is that having one settings data set (at least one luminance mapping function, such as iris) for each location is sufficient. In some cases, the same value (such as iris opening) can be selected for capture for both capture locations, but it is generally advantageous to use different values ​​that are optimal for each location and advantageously optimal in relation to each other. In some advanced situations, the director may select, for example, two functions to perform various possible tasks. For example, in the first location, the camera operator may select either Function 1 or Alternate Function 2 and decide on the fly which function is best. This is useful both when the functions realize small changes (i.e., slightly different shapes) or when they realize large changes. This is used to account for further variability in the location of the shooting environment. For example, a steam bath may have more or less mist. Also, having alternate functions at a location can override the chosen situation. For example, the color composition director may select two possible functions for an outdoor location, but at initialization it is not yet known which will work well during shooting. The camera operator, color compositing (CC) director, or camera operator working with the CC director can, for example, decide to replace the initial version currently loaded in the primary memory for that position, a feature selected using a toggle switch for this shoot position, which is replaced with an alternative feature that, moving forward, becomes the primary feature for this position in the selection UI. Or, the CC director can even tweak a feature for one position and load it into the primary position for the remainder of the shoot, making this the new, fine-tuned, on-the-fly grading behavior for this position (typically, this should be done for small changes and moderation). Even with various features to choose from depending on the capture director, the ability to tweak features in various locations and quickly associate one or more features per location allows for very quick, yet powerful, capture in the sense that it already produces a fairly good output grading (i.e., an output image with correctly graded brightness).Another typical example of a (generalized) location-by-location classification under (at least) two functional categories is a hallway 102. In such a long, narrow environment, there are different lighting fixtures at various positions along the hallway. Naturally, we would treat these as just three different capture locations according to the basic system (without worrying about any particular relationship), but it is better to group them into a general location or group of locations. For example, if the hallway is lit only by exterior lighting from the front, it will gradually darken, but at one location there is a lighting fixture 109 on the ceiling. The lighting fixture 109 locally brightens again (and is in view, and therefore a distinct object with pixel brightness that needs to be taken into account in functionality, possibly in the basic capture settings). The camera operator can quickly toggle from one situation to another, or, more preferably, the system (e.g., the camera itself) can do this automatically on the fly (usually after an exploratory testing phase before setting the optimal capture settings and functionality for the shot).

[0078] For example, the CC director, together with the camera operator, may decide that a good first position for a first representative lighting fixture would be near the entrance of a hallway (e.g., 1 meter behind the door, facing inwards, if the shot is to follow an actor walking in), and that a second representative position would be slightly in front of where the lighting fixture is hanging (thus getting some lighting from the fixture, but not full illumination). During filming, the camera could, for example, behave like this: the camera operator flicks a switch to indicate moving / walking from the entrance position to a position lit by the lighting fixture in the hallway (the type of position, or function, could be saved together for such advanced behavior, such as "graded lighting" or "moving"). During the creation of at least one output graded video, such as an HDR version, the camera could use a function to continuously adjust between the two functions. The amount of adjustment, i.e., how far the function used deviates from the entrance position function to the lit position function, determines, for example, where exactly the operator will stand in a hallway, if the positioning embodiment allows this (another possibility is that a delay would allow this, although for life production a delay of a second or less is often desired, although this can be done in offline production, using the first function for too many images at first, but then, when arriving at the second position, correcting half of the previous images with a function that gradually changes between the first and second place representation shapes).

[0079] These concepts can be better explained in Figures 1B and 1C.

[0080] Figure 1B shows a rough representation of how a first representation (PQ) of an image (e.g., a first grading, e.g., HDR grading) maps to a second (usually lower) dynamic range grading (RDR). Typically, in a relatively simple HDR shoot, you obtain some distinct positions in the first representation, e.g., relative to the raw capture. For example, you can set the luminance in this first representation equal to a digital number multiplied by a constant. The constant might be such that the maximum digital number (e.g., power(2;14)) maps to 4000 nits. All objects then occupy a luminance position (bright or dark along the vertical axis of all possible luminances from 0 to 4000 nits) depending on the luminance they had in the scene (so, for example, if an ADC images a luminance of 7000 nits to its maximum digital number, a scene luminance output of 3000 nits, or 43%, will be 1714 nits in the HDR image, which was initially assigned automatically). This is often reasonably good for a video camera's primary output ("master HDR video"), but not necessarily the best primary-grade output video (i.e., ImHDR). The problem with the following simple video production rules can be more problematic when creating low dynamic range video: The dotted luminance mapping line represents a simple function (F1), such as a gamma-log function (e.g., a function that starts out as a power law for dark HDR input luminances, and becomes logarithmic in shape to map bright input luminances to fit into a small output luminance range). The problem with such mappings is that they generally do not map well to small dynamic ranges (e.g., SDR). Some objects will appear too dark when displayed on, for example, an LDR display. An optimal looking image is humanThe key to generating the optimal shape of the luminance mapping function Fopt (or at least an automaton that can calculate a more sophisticated function for each shot and lighting scenario, following good principles of colorimetric image optimization rather than just a single fixed, averagely good mapping function) is to generate the optimal shape of the luminance mapping function Fopt. For example, a plant photographed indoors would be mapped too darkly with a logarithmic gamma function, so a shape that brightens the darkest image objects is needed. This can be seen in Figure 1C, which shows the same thing on a 2D plot instead of two 1D luminance axes (with luminance normalized to a maximum of 1.0), where the solid curve is higher than the dotted curve, resulting in more pixel luminance in the output image, especially for the darkest objects. Further fine-tuning of the shape can be done, for example, by dimming the luminance of the house to achieve the desired inter-object contrast DEL.

[0081] Figure 1C illustrates the general principle, but also explains how the camera calculates the gradual change in function. If the dotted curve is appropriate for a first location in a hallway and the solid curve is appropriate for a second location, for intermediate positions, the camera uses a function shape that lies between these two functions (i.e., it moves gradually from the first shape to the second shape). Using some algorithm, the amount of deviation can be controlled as a function of the distance traveled to the second location (often, the determination of full brightness is secondary to a visually smooth appearance). As can be seen from the description of the technology here, all gradings produced by cameras, not just SDR grading, benefit from the proper allocation of brightness to various objects, especially practical adjustments for the brightness of objects at various locations in the captured scene throughout the video (indoors, outdoors, natural lighting, additional lighting fixtures, etc.).

[0082] FIG. 2 shows conceptual portions of a camera and the remaining possible devices in an initialization / mode setting system to illustrate aspects of the new approach (those skilled in the art will understand which elements function in which combinations or separately, or may be realized by other equivalent embodiments).

[0083] The first basic parts of the camera have already been described above, so we will now describe some further typical elements for the new technological approach.

[0084] The capture user interface 210 is associated with a further control algorithm, which runs for example on a control processor 241 (which processor depends on the type of camera, e.g., professional cameras which are slowly being replaced, or rapidly evolving mobile phones, etc. Within the camera there may be a general-purpose processor capable of running this algorithm on a GPU, or another camera may have a dedicated ASIC or FPGA). It manages at least which positions are being captured, what must be communicated to external devices, including the image color synthesis analysis circuit 250 (external in this embodiment), and what is expected to be received (e.g., the luminance mapping function Fs2 communicated in signal S_Fs, and in which memory location this function, e.g., for the second position, should be saved). We assume that the camera has a dedicated communication circuit 240, although at least some or all of the functionality may be integrated into another similar circuit.

[0085] For example, assume that at least one high dynamic range image (o_ImHDR) is output via this communication circuitry 240, and that basic capture settings and mapping functions are received (i.e., input) via this circuitry (the camera may further communicate via dedicated cables to the lens, etc., but such details are irrelevant to understanding the present invention).

[0086] We further assume that the connection is IP-based and via Wi-Fi®, either with a MIMO antenna 242 connected to the camera or a USB to Wi-Fi® adapter (other similar technologies are understood, e.g., using 5G cellular, cable-based LAN, etc.). When transmitting more than one image, an error-resilient communication protocol such as Secure Reliable Transport (SRT) or Zixi is not required, although it would be useful to have the functionality doubled from Wi-Fi® communication to transmit all images of the actual shoot.

[0087] Note that in order to determine the luminance mapping function for grading, the received images do not need to be of the highest quality, such as resolution, and may have compression artifacts. The actual shooting video output is provided by the image processor 207 as ImHDR video images (and possibly also ImRDR video images), and in many applications is already compressed directly to a sink that desires AV1 coding, for example using HEVC, VVC, or AV1, although some applications / users may desire an uncompressed (but graded) video output, for example output to an SD card embodiment of the video memory 208, or output directly via some communication system (NETW).

[0088] Thus, with the help of the image color synthesis analysis circuit 250 and via a mapping selection user interface 252, the CC director sees on a monitoring display 253 how the grading looks either approximately (with the wrong colors) or in grade. For example, by changing the shape of the secondary grading function FsL1 on the fly via control points, there is one view that displays the LDR colors, and a second view that displays either the brighter HDR image or just the LDR image. Further figures will illustrate some examples.

[0089] The function determination circuit (251) already provides a first automatic suggestion of a luminance mapping function or secondary grading function by performing an automatic image analysis of the scene. The CC director can fine-tune this function via the UI or do it all by themselves, starting from a master capture RW or at least one HDR image. The applicant has developed an autometa algorithm for mapping, for example, any HDR input image (e.g., with an ML_V equal to 1000 nits or 4000 nits) to, for example, a typically SDR output (RDR embodiment) image. The resulting luminance mapping function (here acting as a secondary regrading function) is scene-dependent. For camera capture, the function shape essentially depends on the lighting conditions at any location. The final result is communicated in a signal format S_Fs, which formulates the function with several parameters that uniquely define its shape (e.g., a parabolic function can be characterized by values ​​a, b, and c if its equation is Y_out=a*Y_in*Y_in+b*Y_in+c). Optimized Functions (e.g., Fs1) is the output from the function determination circuit 251. Naturally, the camera and the external device know this agreed-upon function specification, for example, because the computer runs the camera manufacturer's app.

[0090] Figure 6 shows an example with a well-working function that works to establish both a primary (e.g., 2000 nits ML_V HDR output video) and a secondary grade (e.g., 200 nits video).

[0091] These images illustrate the underlying technical principles, but also show the actual view the director sees as he moves parts of the shape to specify the function in the three subwindows on display 253. How the director changes the shape is also merely an implementation detail. For example, if the slope of a segment needs to be changed, the director does so by turning a wheel to increase the slope value, or by tapping with a finger more rapidly to indicate the end points of the control line for a linear segment or linear approximation or segment. Some users may prefer the first option and other users the second, so both options are available in the same system.

[0092] The mapping from the digital number (DIG_IN) (again normalized to 1.0 for simplicity (using this normalization, the number of bits the ADC has becomes irrelevant, which is usually only a secondary concern when defining optimal grading and optimal grading functions)) to the 2000 nits output HDR video consists of two successive mappings. First, a coarse mapping RC is set (in the coarse mapping view 620), e.g., with two linear segments at the outer ends (the light and dark input digital number) and a smooth segment in between. The locations of the three segments are determined by setting arrows 628 and 629. This is done depending on the device being used, e.g., by dragging a mouse on a computer or by clicking with a pen on a touch-sensitive screen connected to a camera. The arrows can also be set (at least initially, before any human fine-tuning) by, e.g., clicking with the user's finger 611 on an object of interest OOI (e.g., a flame) in, e.g., the master capture's view in the image view 610. The span of the digital number (or luminance, if the same algorithm is used to map from the input luminance of the primary-grade video ImHDR to the output luminance of the secondary-grade video ImRDR) is represented by the up and down arrows (i.e., the positioning of arrows 628 and 629). This may also be located in part of the middle segment, for example, if Autometa has determined three segments. The coarse mapping view 620 also shows a small copy of the selected area (i.e., the fireplace) as a copied object of interest OOIC in the view of the object of interest 625 with the values ​​correctly placed. Starting from segment 621, which still grades these objects as relatively dark, the CC Director toggles or continuously moves through several possible slopes of the dark linear segment (B1, B2) until it reaches the CC Director's optimal segment 622, which grades them brighter in the primary HDR output (ImHDR).While this is good for the darkest objects, perhaps other important objects (the fireplace) may still be suboptimal with such a coarse grading strategy when ending up at a particular offset OF_i and a particular span of luminance DCON_i or inter-object contrast.

[0093] Thus, in the tertiary view window 630, the CC director displays a range of luminances in the 2000 nit range resulting from the coarse grading. Fine adjustment and get a better grade of 2000 nits brightness for the final output (the function we load into the camera is the composite function F2(F1(DN))).

[0094] The UI already includes arrows (638 and 639, copied to correct their new positions; the horizontal position in view 630 corresponds to the vertical axis position in view 620), allowing the user to place the second copied object of interest, e.g., OOIC2, in its correct new position on the graph. For example, a simple algorithm could be used to adjust the contrast of the flame, fixing the brightness of the upper coarse grade of the flame's luminance range (this becomes the anchor, Anch), and then repeatedly flicking a button or dragging the mouse to increase the slope of the lower segment to a higher angle than the diagonal, until the brightness of the lower part of the flame's luminance range reaches the offset DCO from the diagonal. This creates a customizable second grading curve (CC), resulting in a larger contrast range DCON_fi for the final output 2000 nit grading (oHDR_fi) than DCON_i for the intermediate 2000 nit grading (oHDR_im).

[0095] Finally (returning to Figure 2), the actual Shooting operationNow, the image processing circuit 207 fetches from memory the appropriate function F_SEL (e.g., Fs1) and, if necessary, the corresponding FsL1 of the secondary RDR grading, and begins applying it to the captured image as long as filming is occurring at that position (or indeed near that position as determined by the camera operator or an automatic algorithm) until filming reaches a new position. The iris and shutter settings need only be done once, perhaps just before filming begins, by iris signal S_ir and shutter signal S_sh coming from, for example, the camera's control processor or passing through communication circuitry, etc. Otherwise (and the same values ​​are reset with each position change decision, even if the values ​​themselves do not need to be changed), the correct values ​​are commanded via S_ir and S_sh when the camera operator toggles to a new position, and at substantially the same time that new functions are loaded to calculate one or more graded versions of the captured video as output.

[0096] Regarding the signal format of the HDR output, and possibly the RDR output, a typical useful format is a perceptual quantizer EOTF (standardized in SMPTE 2084) for determining the nonlinear R'G'B' color components and, for example, a Rec. 2020-based Y'CbCR matrixing, which in turn holds, for example, VVC (MPEG Generic Video Coding) compression or uncompressed signal coding. If the secondary output is considered conventional SDR, the Rec. 709 format can be used. There may be further conversion to a specific format for communication, such as narrower ranges, packetization, metadata specification, and possibly encryption.

[0097] Therefore, the output video signal of any camera embodiment according to the present technical teachings typically has a first luminance mapping function (Fs1, Fs2, ...) applied to obtain a first grading actual image (along a luminance range up to a selected maximum ML_V of the target display associated with the video). That is, each pixel has a luminance, which is typically coded via an EOTF or OETF (typically a perceptual quantizer or Rec. 709). A secondary grading can also be added to the video output signal if desired, but is typically coded as a function (e.g., secondary grading functions FsL1, FsL2 that calculate a secondary video image from a primary video image). In many cases, the primary grading is HDR grading, and the secondary grading is, for example, SDR grading. However, for example, in the case of backward-compatible HDR broadcasting, the primary grading can be SDR video, and the jointly coded function can be a luminance upgrade function to obtain HDR grading from SDR-grade video. In this scenario, with full backward compatibility, SDR luminance is coded according to the Rec. 709 OETF, but with partial backward compatibility, SDR luminance up to 100 nits is also coded as luma according to a perceptual quantizer such as the EOTF.

[0098] Figure 3 further illustrates what is typically different, i.e., what can be achieved with the present innovation and working method, compared to some simpler approaches that are applicable but have lower visual quality.

[0099] Assume the camera operator has already established a proper capture (i.e. good iris opening, etc.) of the scene shown on the leftmost vertical axis. Here, an ML_V primary grading of, say, 700 nits can be created three different ways (three different technical philosophies).

[0100] A first representation of the 700 nits image (NDR) can be formed by mapping the maximum possible digital number of the camera (i.e., the ADC maximum) to the maximum grading value, which in this example selection is 700 nits. In this case, all other image intensities are scaled proportionally (i.e., linearly, s*DN+b). This may be fine if the capture is to serve as some version of the raw capture for offline post-grading, such as in the film production industry (though some prefer non-linearity), but usually a good 700 nits direct from the camera is needed. Grading (usually because you are capturing very bright scene objects, some objects will be unpleasantly dark). This means that you need to set all camera settings, i.e. the basic capture settings, and the mapping function all at once, i.e. for the whole shoot, Same for all positions This is a situation that can be achieved if the

[0101] Another possibility is the second representation UDR. This other possibility occurs when something changes, for example in the lighting situation. A new optimal exposure every time This is typically obtained when using a variant (potentially an improvement) of the classical auto-exposure algorithm that determines the dominant luminance of most pixels that come out in an average-based exposure calculation for a normal object, placing all Lambertian reflective objects (under primary or base illumination) at the same output image luminance location (looking at the luminance axis of all possible image luminances in a UDR image, an illuminated portrait comes out at about the same brightness as an outdoor house).

[0102] This is not ideal. A certain degree of brightness difference (i.e., appearance to the viewer) between the average object brightness in a more strongly lit outdoor environment and the average object brightness in an indoor environment (DL_env) is desirable, and a technique that allows human control over this is desirable. That is, typically, a highly controllable adjustment between the brightness captured at various positions is desired. This is shown in the third representation (optimized ODR), which serves as the primary HDR grading (imHDR), with a selected master grading maximum brightness ML_M equal to 700 nits (and different shots from different positions optimally adjusted along that range). The dotted arrow indicates mapping one brightness by the optimal brightness mapping function FL_M, which is the optimal brightness function stored in the camera for the indoor position capturing the flame, as described elsewhere in this patent application.

[0103] FIG. 5 shows a configuration of a device having one or more position determining circuits. High-performance cameraThe basic components (lens, sensor, image processing circuitry) are similar to other cameras. Here, a viewfinder 550 is shown, allowing the camera operator, acting in the role of color composition (CC) director, to see several views. This may not be an ideal view, such as in a separate grading booth or production studio adjacent to the set, but sometimes one has to live with constraints, such as a solo shoot in Africa where the final customer is not yet present. Some early adopters prefer to work this way, and this embodiment can accommodate that. Alternatively, for better resolution, surround shielding, etc., the operator / CC director can briefly wear glasses 557, e.g., with projection means 558 and a light shield 559. For example, a visor, such as those used in virtual reality viewing, could be used. Also shown is a voice recognition circuit 520 or software connected to at least two microphones (521, 522) that forms an audio beamformer. Voice recognition does not need to be as complex as full voice recognition, since only a few location descriptions (e.g., "fireplace") need to be correctly and quickly recognized. Whether the camera uses an in-camera recognition algorithm or uses IP communication capabilities to have it performed by a cloud service or a computer in the production studio is a detail beyond the need for explanation in this application.

[0104] An external beacon 510 is also shown. This can be a small IC with an antenna in a small box that can be glued to a wall or the like. The beacon can provide triangulation, or localization, if it broadcasts a specific signal sequence or the like. It interacts with location detection circuitry 511 within the camera. This circuitry performs, for example, triangulation calculations. Alternatively, for a coarser location determination, it may simply determine whether it is in a room, for example, based on signal pattern recognition or signal timing. All of these location systems can operate similarly to actual capture during the initial discovery phase of control, particularly during pre-configuration of the camera or camera system (i.e., the camera in the system with other devices, such as a computer). Therefore, during actual capture, there is the same data path, or at least part of the same data path, for obtaining basic or processed measurement data (e.g., a location estimate) to the capture UI 210 and selection UI 230.

[0105] The video communication to the outside world over the network may be for example a contribution to a final production studio (where the video feed is mixed with, for example, other video content, and then to a broadcaster), or the video communication may be streamed to a cloud service, for example cloud storage for later use, a YouTube live channel, etc. This high performance camera also has circuitry for determining the included functionality.

[0106] The image analysis circuit is shown in Figure 9. The idea of ​​all these technologies is that during life shooting, the position-dependent behavior of the camera still facilitates operation. In some shoots, there is a focus puller who timely selects the shooting location just before changing the focus, for example via a miniature display, but in some situations the cameraman has to do it all himself (he is already quite busy following, for example, fast-moving people or action with proper framing and geometric composition), so it is good to be able to rely on or at least be helped by some technical circuit that determines the position information (a large amount of work is done in the initialization phase of scene detection).

[0107] But before that, FIG. 8 shows an example of (again, simply and quickly) creating a secondary grade version of the image (eg, SDR output).

[0108] The consideration for grading an HDR primary grading is to ensure that all scene objects are visually plausible or impressive (i.e., not too dark to be easily seen, the brightness impact of one object is not excessive compared to another, the intensity of appearance of bright objects, etc.) on a high-quality image representation, typically an archived master grading, that serves to derive the secondary grading. Thus, most objects are already more or less accurately specified in terms of luminance, which can be accounted for by the darkest object (e.g., keeping the darkest object at the same luminance across all grading and re-grading approaches).

[0109] Secondary grading primarily involves various technical considerations regarding how to best fit the object luminance range into the primary grading so that it fits nicely into a smaller dynamic range. Fitting nicely means trying to preserve as much of the original look of the primary grading as possible. For example, balancing intra-object contrast to preserve enough visual detail in a flame with inter-object contrast to ensure the flame appears bright enough above the rest of the room and not adjacent in brightness. This involves moving away from concepts of equalizing luminance and somewhat darkening the darkest objects to create visual contrast. In any case, even if the technical and artistic details of curve construction differ, the technical user interface and the calculations behind it are the same or similar (and cameras similarly use such features for parallel calculation and output of position-dependent secondary grading RDRs).

[0110] For example, a typical scene object, such as a painted portrait, is given a normal luminance LuN in the RDR grading. Because the painting is illuminated, and because the RDR grading's 400-nit ML_V maximum luminance easily allows for this, we want the painting pixels to have a luminance of approximately 75 nits, rather than ~50 nits. The circle in Figure 8B is not a control point here, because in this strategy it is either a point along the luminance range or the corresponding regrading function segment (the main segment F_mainL). While using this segment for many colors may be sufficient, compression requirements generally dictate the desire for a multi-segment curve for ImHDR to ImRDR regrading. At least some of these are so smoothly connected that they appear like a single, nonlinear segment. However, in this example, we also want to give importance to the RDR grading result of the flame, thus increasing its contrast somewhat. Therefore, the actual second control point Cp22 can be placed some distance PHW from the average or minimum luminance of the flame object, for example, halfway between it and the painting. Then, a selected, e.g., linear, segment of the main normal object luminance is continued at its lower end up to this Cp22. In this example, a boost from Cp22 to (max_ImHDR, max_ImRDR), i.e., (2000, 400), was considered a good regrading not only for the flame, but also for all other bright objects at this position in the scene (establishing the boost segment F_boostL of this regrading function FsL1). Note that in this example, the regrading is shown in an absolute axis system ending with the respective ML_V values ​​in units of nits, rather than the normalized 1.0. This is not because one type of grading must be done in this area and another type in another, but just to make it clear that all variations are equally possible.The darker segments of the scene (F_drkL) are again automatically shifted by the CC Director to the lower control point Cp21, and if it is deemed acceptable, no separate control points are created for these objects (e.g., instead of continuing to (0,0), consider vertically raising the starting point to (0,Xnit) to brighten the darkest pixels). Figure 8A roughly illustrates the desiderata for RDR grading by projecting several important objects and their representative luminance or brightness values, and Figure 8B illustrates the determination of the actual curves, i.e., the actual secondary grading function FsL1 for outdoor and adjusted indoor use. While linear segments are shown for both the primary and secondary grading from the raw digital numbers, one or both of these segments may also be curved (e.g., have a slight curvature compared to a linear function). While linear functions are simple and work well enough, for example, when applied to the luminance channel only (while typically preserving hue and saturation or substantially leaving the corresponding Cb and Cr unchanged), curved segments may be preferable for the grading curve. Grading also applies to transformations of the three color components, such as matrixing, i.e., non-linear R'G'B' coefficients or any color representation. This is generally an aspect of camera color science, but the details are left to the knowledge of those skilled in the art of colorimetry to understand the principles of the new teachings.

[0111] FIG. 9 shows examples of how various embodiments of the camera location determination circuit (540) can determine roughly or more precisely where (and in what orientation) the camera operator is currently shooting. Image analysis technology is vast, resulting from decades of research, and several alternative algorithms are available. Therefore, this technical element will be described with only a few examples. FIG. 9A shows that in addition to simply determining that one is shooting in a room (basic capture parameters and functions are often determined to be appropriate for any shooting method in that room (and possibly adjacent rooms, but not the outside)), advanced embodiments can also use 3D scene estimation techniques to determine where in the room and in what orientation the camera is shooting. Of course, the accuracy of this measurement does not need to be as high as, for example, depth map estimation, so many techniques and less computationally intensive and inexpensive techniques can be used. It is not necessary to know where the center of the lens is to centimeter accuracy. This is because the mapping function and basic capture parameters should work the same whether 10% of the window is in view or 100% (e.g., zoomed to 100%) (i.e., every pixel of the current indoor image is imaging a sunlight-lit outdoor object). Here again we see a major difference in our approach compared to classical auto-exposure techniques, which naturally result in a completely different setup when the window is zoomed (indeed, an outdoor setting, rather than an outdoor setting viewed from an indoor setting). Thus, for example, after basic object feature extraction and / or analysis, the geometry can be determined by calculating the distance DP between objects on the sensor, the shift of the object on the sensor relative to the camera rotation, etc.

[0112] In many cases, it is only necessary to recognize a room by recognizing a few typical objects, so Figures 9B and 9C focus on that.

[0113] Figure 9B shows an example of an interesting, standout feature (i.e., one whose discovery can be considered interesting despite the "blandness" or chaos" of other features): the red bricks of a chimney. For example, if red were a rare color in this room, it would already count as a standout feature, at least as a starting feature. These bricks are size-independent, i.e., position-independent, and can even be determined by looking for red corners on gray mortar. (If size-dependent features, such as rectangles or the entire shape of a fireplace (encoded as distances and angles, such as linear boundary segments), are desired, the algorithm can zoom in and out of the image or a portion thereof several times, or apply other techniques.) Thus, for example, two adjacent bricks are aggregated as such a neighboring pattern, as determined by (non-limiting) the G-criterion or generalized G-criterion (see, e.g., Sahli and Mertens, "Model-based car tracking through the integration of search and estimation," Proceedings of the SPIE Conference on Enhanced and Synthetic Vision, 1998, pp. 160-160).

[0114] The idea behind the G-criterion is that there are elements (typically pixels in image processing) that have some properties (e.g., a red color component in the simple example of a fireplace), but also more complex aggregate properties that result from precomputing other properties. The element properties typically have a distribution of possible values ​​(e.g., the red color component value depends on the lighting), and they are geometrically distributed in the image: there are locations where there are red brick pixels and locations where there are not. Figure 9C illustrates the principle concept.

[0115] For example, suppose we use a scale that is high for red brick (a saturated color) and low for unsaturated gray mortar. For mortar, R=G=B is roughly true, and for brick, R>G,B is true.

[0116] Therefore, as the discrimination characteristic P, for example, take the ratio of the function R - (G + B) / 2 (or R / (R + G + B)). A "dual lobe" histogram is expected, where here, one type of "object" is around one value (e.g., 1 / 3), and another type is around another value (e.g., 1 / 2 < P ≤ 1).

[0117] Next, for a moving sampling filter that checks several positions in the image, two sampling regions R1 and R2 are selected.

[0118] Here, the G criterion is calculated as follows: G = sum of all possible values Pi[abs_value_of(number_occurences_Pi_in_R1 - number_occurences_Pi_in_R2)] / Equation 1

[0119] This idea is also, as expected, to well - shape the regions separately. Thus, R1 is an L - shaped mortar region around the brick, and R2 is a fragment of the brick within it.

[0120] What happens when a G - criterion detector is placed on the boundary of such a brick?

[0121] Running the possible P values Pi from low to high, it can be seen that low values around 1 / 3 occur more frequently in the mortar. So, for example (theoretically), there is A_R1, which is the area or number of pixels of the L-shaped region R1 when all pixels are completely achromatic. In R2, there are no such achromatic pixels. Therefore, the first term of the sum is A_R1. The size of the second region (in pixels, but not in shape) is typically interpreted as the same. Assume the brick has only maximum red pixels (R = 255, G = B = 0).

[0122] In this case, there are many Pi values ​​(Np) that have zero occurrence in either region and therefore do not contribute to the sum. Finally, there are pixels that are only present in red brick region R2; that is, pixels with a value of Pimax=1. Again, of these, there are A_R2=A_R1. Therefore, the sum is 2*A_R1. If the normalization factor is also 2*A_R1, then the G criterion for detection is 1.0. If this analysis filter is placed on all brick pixels, both regions will only contain P values=1. Therefore, only one bin will count A_R1-A_R2=0.

[0123] Thus, the G-criterion detects what is there and where it is by obtaining a value close to 1.0 if it exists; otherwise, it is 0. While the G-criterion's statistics are somewhat complex, its power lies in its ability to input any (or multiple) characteristics P as desired. For example, if a room is characterized by black wallpaper with yellow stripes, we can calculate the cumulative sum or derivative of the striped pattern, leaving a uniformly painted wall aside. For example, if we classify yellow as +1 and blue as -1, we can calculate the P feature by measuring the pixel colors at intermediate positions in the multicolored wallpaper bands, obtaining a local representative feature P = M1 + (-1) * M2 + M3 + (-1) * M4, where the binarized measurements M1 through M4 depend on the underlying pixel color. That is, for wallpaper, we obtain P = 1 + (-1) * (-1) + 1 + (-1) * (-1) = 4. Any texture, color, or geometric scale can be constructed and used in the G-criterion. Also, the shape of the sampling area can be freely determined (only the number of sampled pixels should be the same for ease of comparison and normalization).

[0124] The generalized G criterion does not compare the characteristic situation present in a location with its neighbors, e.g., adjacent locations in the image, but with general characteristic patterns. For example, if a red patch is known to have a P value much greater than 1 / 3, or if the subtraction exceeds zero, then the various red bins that occur can be compared with a reference bin of 0. That is, regardless of whether that color patch is present in the image or not, we can compare the unique red region R1 anywhere in the image with a virtual region R2 consisting of all Pimin=0 values ​​(itself), and obtain a rectangle.

[0125] The G-criterion is only the first step in selecting candidates, and more detailed algorithms can be implemented if more certainty is needed.

[0126] So, typically, upon initialization, the location circuit (540) captures one or more images from this location, e.g., a Master Capture RW, typically with the base capture settings for this location. Salient features, such as rare colors or corners, begin to be determined. For these salient objects, slightly more structured low-level computer vision features are constructed, e.g., a brick detector with a G-criterion. Several representations for these representative objects are saved, e.g., small portions of the image that are correlated, descriptions of shape boundaries, etc. Various mid-level computer vision descriptions of the location are constructed and saved.

[0127] During the actual location determination stage, the location determination circuit (540) performs one or more such calculations to estimate the camera's location. Additional calculations can be performed to cross-validate the location, for example by checking whether this indoor location is somewhere in an outdoor scene, or by checking for color texture patterns typical of outdoors on a captured copy of some of the currently captured image.

[0128] Note that this identification of typical structures in anything (here, images of various locations) is typical of what gets trained in the hidden layers of a neural network. Indeed, as NN processors become more common and cheaper, such ICs can be used, at least in situations where sufficient training can be afforded (e.g., if you're a corporate communications specialist and frequently visit this location among other locations because it's on the company campus). Another variant, however, relies on the user simply capturing a few images of scenes they consider representative for identification (something humans are good at: capturing images of portraits and of fireplaces, for example, which are unlikely to be visible in the woods outside). Then, as you start sipping your coffee before shooting, the circuit's algorithm begins simple (untrained) statistical image analysis.

[0129] FIG. 11 shows another example of how all capture (mode) parameters can be adjusted to facilitate on-the-fly shooting possibilities and excellent HDR video output. The basic capture settings and various functions are typically adjusted in the sense that, at least when configuring the camera's behavior at each position (i.e., setting the iris and mapping functions to obtain, for example, graded HDR luma or luminance from the digital numbers being captured), the results look good when viewing the entire shoot, i.e., all sequential shots at various positions, directly. This is illustrated using another example in which a single camera operator can freely walk between a well-lit (strongly lit) first room (Ro_1) and a darker second room (Ro_2). Objects in this second room are typically indirectly illuminated through an opening (e.g., an open door) in the wall (1105). This opening also serves as a viewport to part of the other room when the camera is in the first room and facing that direction (allowing the cameraman to move between rooms). In the future, with an ultimate, nearly infinite capture dynamic range, the basic capture will be unproblematic (i.e., there will be no significant choices, such as iris opening values), and only the mapping function will be predefined (because, given one or more basic capture settings, it will still be necessary to create, for example, 800 nit ML_V output video, 1500 nit video, and / or 3000 nit video from the digital number captured at each moment). However, in this example, we consider that there will be at least one near-future camera whose capture dynamic range will not be perfect, and there will be no opportunity or explicit choice regarding the non-use of fill light in the dark second room. In the first room, during the early stages of discovery, the CC director and / or camera operator naturally assumed that the sun and at least some of the sunlit clouds outside the small window (1104) would be clipped beyond the full pixel well and the maximum digital number DN (here, considered from an 18-bit ADC).Therefore, the area of ​​maximum brightness is determined as a set of pixels (each with a luminance) slightly darker than the pixels whose luminances are all clipped to the same maximum value (in this example, a section of clouds, where we still want to encode the gray value variations in the cloud luminance). This cloud luminance forms a good aggregate maximum because there are no other significant scene areas or objects with high luminance elsewhere. In various gradings, or at least in the only HDR video output, this aggregate maximum is mapped to the TDML value ML_V. There is a bright lighting fixture (1103) that brightly illuminates several important objects 1102 (e.g., silverware) on a table. The bright lighting fixture is captured moderately without clipping, as are all other objects in the illuminated room down to the averagely illuminated object 1101. In a darker second room, using the same basic capture settings, e.g., the first iris value Ir1, the hollow object 1111 is captured with a small digital number, but still with sufficient fidelity (i.e., usually with little noise). However, the monster 1110 hiding in the shadow of something would be captured too darkly with this capture setting, which is not ideal. Therefore, whenever the cameraman begins filming in the second room (not necessarily facing the camera room, since it may not be a problem if much detail is not visible from the dark room in such a partial view surrounded by much brighter pixels from the first room), the camera memory values ​​can be set so that a larger iris opening Ir2 (e.g., twice as open) is used. Thus, the mapping to the 1000-nit master HDR video output of something in the first room is done using a function of the first position (denoted by the function designated p1) and its iris setting, i.e., the first iris value Ir1. That is, the function F_p1_Ir1 is used. For this, one input / output pair of values ​​is represented from the range of digital numbers to the range of possible brightnesses for a 1000-nit HDR grading. Similarly, other digital number values ​​are mapped to other output brightnesses, together forming a strictly increasing mapping function.When shooting in a dark room, we can define a function that starts with the digital number generated by the first iris setting, but knows to use DN when captured with the second iris setting Ir2, which is chosen to be optimal for that room because it captures the darkest objects without noise and the brightest objects are absent. (If we point the camera at an object in the first room through a door opening, we usually don't see the outside world through the window and don't see much of the bright room, so it might not be too much of a problem if something clips at a very high scene luminance, but otherwise we want to optimize the entire luminance set for the shooting environment together; only when shooting in a position like "second room rotated toward the opening" can we lose some of the quality of the darkest objects due to noise, but this is usually fine because the viewer doesn't need to see the details anyway when there's no monster in view and the darkest and brightest pixels are displayed together on the display screen.) Although the shape of the second room position function F_p2_Ir2 differs depending on whether you start with an Ir2-based digital number or an Ir1-based digital number, the relationship between them is the same as the CC director's initial specification: that dark room objects, such as hollow objects 1111 (e.g., a tub), should have pixels with a desired constant value, particularly a luminance value around a certain dark luminance L_drk, which is an amount or percentage darker than the bright luminance L_bri of the pixels of bright room objects. Thus, from this explanation, we can see how the discovery stage adjusts all the values ​​required for two or more shooting positions to ultimately obtain a luminance grade that allows for liberal actual shooting and eliminates concerns about colorimetric issues in the output video or video version. Various subranges of the characteristic scene region for the first position are adjusted by placing them at specific distances from the subranges of the characteristic region for the second position. For at least some subranges, this is actually achieved via the determination of the function shape.

[0130] The algorithmic components disclosed in this document may actually be implemented (in whole or in part) in hardware (e.g., as part of an application-specific IC) or as software running on a specialized digital signal processor or general-purpose processor.

[0131] It should be possible for a person skilled in the art to understand from this summary description which components are optional improvements and can be implemented in combination with other components, and how the (optional) steps of the method correspond to the respective means of the apparatus (and vice versa). The word "apparatus" in this application is used in the broadest sense, i.e., a group of means for achieving a specific purpose, and may therefore be, for example, (a small circuit part of) an IC, a dedicated device (such as an electronic appliance with a display), or part of a networked system. The term "arrangement" is also intended to be used in the broadest sense and may include, in particular, a single apparatus, part of an apparatus, a collection of (parts of) cooperating apparatus, etc.

[0132] The explicit meaning of computer program product should be understood to encompass any physical realization of a set of commands which, after a series of loading steps (which may include intermediate conversion steps such as translation into an intermediate language or a final processor language), enables a general-purpose or special-purpose processor to input the commands into the processor and perform any of the characteristic functions of the invention. In particular, a computer program product may be realized as data on a carrier such as a disk or tape, data residing in a memory, data traveling via a wired or wireless network connection, or program code on paper. Apart from the program code, characteristic data required for the program may also be embodied as a computer program product.

[0133] Some of the steps required for the operation of the method, such as data input and output steps, may not be written in the computer program product but may already be present in the functionality of the processor.

[0134] It should be noted that the above embodiments are illustrative rather than limiting of the present invention. Those skilled in the art can easily realize mapping of the presented examples to other areas of the claims, but for the sake of brevity, not all these options are mentioned in detail. Apart from the combinations of elements of the present invention combined in the claims, other combinations of elements are possible. Any combination of elements can be realized by a single dedicated element.

[0135] Any reference signs placed between parentheses in the claims are not intended to limit the claim. The word "comprising" does not exclude the presence of elements or aspects not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.

Claims

1. 1. A method of controlling a video camera, comprising setting a video camera capture mode that specifies output pixel intensities at corresponding locations within a captured scene in one or more graded output video sequences of images, the graded output video sequences being characterized as having different maximum pixel intensities, the method comprising: an operator of the video camera moving to at least two locations within a scene having different illumination from one another and capturing at least one high dynamic range image for each of the at least two locations within the scene; a color composition director analyzing the at least one high dynamic range image captured for each of the at least two positions to determine a respective region of maximum brightness, and determining at least one of an iris setting, a shutter time, and an analog gain setting of the video camera in response to the maximum brightness; for each of the at least two positions, capturing a respective master capture using the determined at least one of an iris setting, a shutter time setting, and an analog gain setting of the video camera, and storing the determined iris setting, shutter time setting, and analog gain setting for subsequent image capture when the video camera is in the corresponding position; determining at least a respective first grade image for each of the master captures, the determining step including determining an adjustable brightness allocation function for each position that maps a digital number of the master capture to a nit value of the first grade image; storing the brightness allocation functions for the at least two locations, or parameters uniquely defining these brightness allocation functions, in respective memory locations of the video camera; A method for controlling a video camera, including:

2. determining, for the at least two locations, two corresponding secondary grading functions for calculating corresponding second-grade images from the luminance of the master capture or the first-grade image, respectively; storing said secondary grading functions or parameters that uniquely define these secondary grading functions in a memory of said video camera; Further comprising:

2. The method of claim 1, wherein the second grade image has a lower maximum brightness than the first grade image.

3. 1. A method of capturing high dynamic range video in a video camera, the video camera outputting at least one graded high dynamic range image in which pixels are assigned a luminance value, the method comprising: The method includes initially applying the method of setting the video camera capture mode of claim 1 or 2, and then during video capture, the method includes: determining a corresponding one of the at least two positions for a current capture position; loading a brightness assignment function corresponding to the corresponding location from a memory of the video camera; applying the brightness assignment function in an image processing circuit to map digital numbers of successive images captured during capturing at the current capture position to corresponding first grade images, and storing or outputting the first grade images; A method comprising:

4. 4. The method of claim 3, wherein the video camera includes a user interaction device such as a double-throw switch for toggling between function memory locations of saved positions of different lighting in the shooting environment, the user interaction device selecting the previous position in a chain of linearly linked positions when pressed in one direction and selecting the next position when pressed in the opposite direction, the method further comprising the step of fetching the brightness assignment function corresponding to the selected position from a memory of the image processing circuit.

5. The method of claim 3 , wherein the video camera includes a voice recognition system for selecting a stored luminance allocation function or secondary grading function based on a description associated with a location, such as “living room.”

6. The method of claim 3 , wherein the video camera includes a location and / or orientation determination circuit based on triangulation using a positioning system at least temporarily positioned in a spatial region surrounding the at least two locations.

7. The method of claim 3 , wherein the video camera includes a location determination system based on analysis of images captured near each of the at least two locations.

8. 3. A method for controlling a secondary video camera to set a video camera capture mode that specifies output pixel brightness in one or more graded versions of an output video sequence of images, the method comprising: in a first video camera, setting a video camera capture mode that specifies output pixel brightness in one or more graded versions of an output video sequence of images according to the method of claim 1 or 2; communicating between the cameras to copy a group of settings including the iris setting, shutter time setting, and analog gain setting, and communicating any of the determined brightness allocation functions from the memory of the first video camera to the memory of the second camera.

9. 9. The method of claim 8, wherein one of the first video camera and the second camera is a static camera having a fixed position in a portion of the shooting environment and the other camera is a moving camera, and the method further comprises either a step of copying the brightness allocation function for the position of the static camera to a corresponding function memory of the moving camera, or a step of copying the brightness allocation function in the moving camera for the position of the static camera from the corresponding function memory of the moving camera to the memory of the static camera.

10. 1. A system for configuring a video camera, the system comprising: setting a capture mode in a video camera that specifies output pixel intensities in one or more graded versions of an output video sequence of images that are output by the video camera to a memory or a communication system; the video camera including a location capture user interface configured to enable an operator of the video camera to navigate to at least two locations within a scene having different lighting from one another and capture at least one high dynamic range image for each location selected via the location capture user interface to be a respective representative master HDR capture for each location; receiving each of the at least one high dynamic range image, and a color synthesis director analyzing the at least one high dynamic range image; a) a region of maximum brightness in the image, and at least one of an iris setting, a shutter time, and an analog gain setting of the video camera based on the region of maximum brightness; b) via a function determination circuit, for at least each first grade image corresponding to each of the master captures, a brightness allocation function of each digital number of the master capture to a nit value of the first grade image for the at least two positions; an image color synthesis analysis circuit that enables determining Including, A system for configuring a video camera, wherein the video camera includes a function memory for storing the luminance allocation functions for the at least two locations determined by the image color synthesis analysis circuit and received from the image color synthesis analysis circuit, or parameters that uniquely define these luminance allocation functions.

11. 11. The system for configuring a video camera of claim 10, wherein the function determination circuitry enables the color synthesis director to determine two respective secondary grading functions for the at least two positions in order to calculate a corresponding second grade image from the master capture or the first grade image, and the video camera stores these secondary grading functions in memory for future captures.

12. 12. A camera including or operating in a system according to claim 10 or 11, the camera having a selection user interface for selecting from a memory a secondary grading function or brightness mapping function corresponding to a capture position.

13. 13. The camera of claim 12, including a user interaction device such as a double-throw switch for toggling between function memory locations of saved positions of different illumination in the shooting environment, wherein the user interaction device selects the previous position in a chain of linearly linked positions when pressed in one direction and the next position when pressed in the opposite direction.

14. 13. The camera of claim 12, comprising a voice recognition system and preferably a multi-microphone beamformer system directed at the camera operator to select a stored luminance allocation function or secondary grading function based on a description associated with a location, e.g., "living room".

15. 15. A camera as claimed in claim 12, 13 or 14, comprising a location and / or orientation determination circuit, wherein location determination is based on triangulation using a positioning system arranged in a spatial region around the at least two positions, and wherein the location and / or orientation determination circuit is connectable to a compass.

16. 16. A camera as claimed in claim 12, 13, 14 or 15, including location determination circuitry based on analysis of images captured near each of said at least two locations.