Configuration of visual appearance resulting from post-processing of images captured by a camera
By extracting the visual differences between the image captured by the camera and the image desired by the user, and automatically applying post-processing operations, the problem of cumbersome camera configuration is solved, and simple and accurate visual appearance adjustment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AXIS
- Filing Date
- 2025-12-11
- Publication Date
- 2026-06-16
AI Technical Summary
The configuration parameters and post-processing operations of modern cameras are cumbersome, making it difficult for users to understand and intuitively adjust them to achieve the desired visual appearance. Existing technologies cannot provide a simple and effective configuration solution.
By obtaining a first image captured by the camera and a second image modified to the user's expectations, visual differences are extracted, and image post-processing operations are automatically applied to generate a visual appearance that meets the user's expectations, reducing reliance on complex interfaces.
Users can modify images using any software to meet their expectations, reducing the maintenance of complex interfaces and computational workload, and improving the convenience and accuracy of configuration.
Smart Images

Figure CN122227089A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of post-processing of images of a scene captured by a camera. More specifically, this disclosure relates to the configuration of the visual appearance of an image resulting from such image post-processing. Background Technology
[0002] Modern cameras, such as surveillance cameras, typically provide an interface through which users can (remotely) (re)configure various camera settings and image post-processing operations. The interface can be, for example, graphical and accessible via a webpage hosted on a server running on the camera, and may also provide a live view of the video stream captured by the camera. With this live view, users can observe the resulting (visual) changes to the output video stream and (if deemed necessary) make further changes until the desired visual appearance of the output video stream is achieved.
[0003] As cameras become more capable, the number of configuration parameters and post-processing operations may increase, making the maintenance of such an interface potentially cumbersome. Furthermore, users may find it difficult to understand the purpose of parameters and operations, and to intuitively understand which parameters and operations should be changed and / or used, and how to change and / or use them to ensure the visual appearance of the output video stream corresponds to the user's expectations. US 11,290,638 B1 discloses a technique for automatically ensuring consistency between images capturing the same or similar types of objects. US 8,300,933 B2 discloses a technique for generating a color correction matrix (CCM) for an image sensor. Summary of the Invention
[0004] This disclosure aims to further develop contemporary technology and provide technical solutions that at least partially overcome the aforementioned problems. The technical solutions include methods, apparatus, systems, computer programs, computer-readable storage media, and computer program products as defined in the appended independent claims, and embodiments of these entities are defined in the appended dependent claims.
[0005] According to a first aspect of this disclosure, a method is provided for configuring the visual appearance of an output image from a camera capturing a scene. The method is implemented by a computer and, for example, executed by processing circuitry of a computer system. The method includes: obtaining a first image depicting the scene captured by the camera, and providing the first image or a copy of the first image to one or more users. The method further includes: obtaining a second image depicting the scene, wherein the second image corresponds to the first image after being visually modified according to the expectations of one or more users. The method further includes: extracting a set of one or more visual differences between the first image and the second image. The one or more visual differences are caused by changes in at least one of contrast, saturation, brightness, sharpness, white balance, global tone mapping, and local tone mapping, and / or by the addition of at least one of text overlay, image overlay, vector overlay, and privacy masking. The method further includes: determining one or more image post-processing operations based on the extracted set of one or more visual differences, such that a third image depicting the scene, captured by the camera and subjected to one or more post-processing operations, is visually more equivalent to the second image compared to the first image.
[0006] The proposed technical solution offers the following advantages: it allows, for example, a user to obtain a first image from a camera and modify that image to generate a second image according to the user's expectations. The second image is then compared (by a computer) to identify the changes made by the user, and these changes can then be automatically applied to visually alter the output stream from the camera, making it more visually similar to the first image (e.g., making it equivalent to the visual appearance of the second image, rather than the first). Thus, the user is not constrained to use, for example, an interface provided by the camera for configuring settings, but can use any available software to modify the first image to better match the user's visual expectations. The envisioned technical solution also reduces the need to create and maintain complex interfaces for modifying all possible camera settings and post-processing operations, and reduces the computational workload that the camera would otherwise have to perform to provide such an interface.
[0007] In one or more embodiments of the method, the method may include: capturing a third image of the scene using a camera, and (then) performing one or more image post-processing operations on the third image. This is the opposite of, for example, attempting to simulate the appearance of an image of a scene captured by a camera and then subjected to one or more post-processing operations; however, this is also conceived as a possible technical solution.
[0008] In one or more embodiments of the method, the method may further include providing a preview of a third image, after one or more image processing operations, to one or more users responsible for providing a second image by modifying a first image. One or more users may then confirm that one or more post-processing operations determined based on the extracted visual differences between the first and second images meet the user's expectations, and, for example, send confirmation to a camera that the camera should use the one or more post-processing operations when capturing future images of the scene. If the user does not confirm, the method may include at least repeating the operation of determining one or more post-processing operations, which may be based on further input from the user as part of a third image for which the first suggestion has not been confirmed. In other examples, one or more users may continue by sending an updated second image, and the method may include: extracting one or more visual differences again, and determining one or more (new or updated) image post-processing operations based on this. In some examples, generating the preview may include, for example, generating the third image in parallel with an image of an output video stream for which one or more post-processing operations have not been applied. This may be particularly suitable for cameras capable of such parallel processing, for example, because the output video / image stream is unaffected until the user is satisfied with the one or more post-processing operations determined based on the preview.
[0009] In one or more embodiments of the method, one or more image post-processing operations may result in (e.g., causing) a change in at least one of contrast, saturation, brightness, sharpness, white balance, global tone mapping, and local tone mapping.
[0010] In one or more embodiments of the method, one or more image post-processing operations may include the addition of at least one of text overlay, image overlay, vector overlay, and privacy masking.
[0011] In one or more embodiments of the method, the method may include: providing a first image to one or more users, and then obtaining both the first image and a second image from one or more users. By doing so, the device responsible for performing the method does not need to track, for example, which image captured by a camera from which the second image is derived, because it receives both the first image and the second image and continues by extracting the visual differences between them.
[0012] In one or more embodiments of the method, the method may include: providing a first image and an identifier of the first image (such as an identifier number or string) to one or more users, and then obtaining a second image and the identifier of the first image from the one or more users. By doing so, an apparatus (such as a camera) performing the method may locally store the first image and use a reference provided by the one or more users to identify which locally stored image is the first image on which the second image is based.
[0013] In one or more embodiments of the method, determining one or more image post-processing operations may include: calculating at least one of a gain and a color matrix based on one or more extracted (visual) differences, and then applying at least one of the gain and color matrices as part of one or more image post-processing operations (i.e., one or more post-processing operations may include applying at least one of the gain and color matrices).
[0014] In one or more embodiments of the method, determining one or more image post-processing operations may include: (e.g., using any conventional tool for this purpose) vectorizing at least some of the extracted one or more differences to generate one or more vector commands, and executing one or more vector commands as part of one or more image post-processing operations (i.e., one or more post-processing operations may include executing one or more vector commands).
[0015] According to a second aspect of this disclosure, an apparatus is provided for configuring the visual appearance of an output image from a camera capturing a scene. The apparatus includes processing circuitry configured to perform operations of the method of the first aspect. For example, the processing circuitry is configured to: obtain a first image depicting the scene captured by the camera; obtain a second image depicting the scene; extract a set of one or more visual differences between the first image and the second image; and determine one or more image post-processing operations as described above based on the extracted set of one or more visual differences.
[0016] In one or more embodiments of the device, the device may be a camera, such as a surveillance camera.
[0017] According to a third aspect of this disclosure, a system is provided for configuring the visual appearance of an output image from a camera capturing a scene. The system includes a first means, wherein the first means is the means of the first aspect (or any example implementation thereof as described herein). The system includes a camera for capturing a first image and a second means external to the camera. The second means includes processing circuitry, and the processing circuitry of the second means is configured to receive the first image (from the first means) and provide image manipulation and / or image creation software that allows a user to generate a second image by modifying the first image (e.g., by manipulating existing features of the first image and / or by adding one or more additional features of the first image, such as one or more overlays, privacy masking, etc.). The second means and its processing circuitry are further configured to provide the second image to the processing circuitry of the first means.
[0018] In one or more embodiments of the system, the first device of the system may be, for example, a surveillance camera as described above.
[0019] According to a fourth aspect of this disclosure, a computer program is provided, comprising computer code that, when executed on the processing circuitry of a device, causes the device to perform the methods of the first aspect (or any of its exemplary embodiments as described herein). For example, when executed on the processing circuitry, the computer code may cause the device to: acquire a first image depicting a scene captured by a camera; acquire a second image depicting the scene; extract a set of one or more visual differences between the first and second images; and determine one or more image post-processing operations as described above based on the extracted set of one or more visual differences.
[0020] According to the fifth aspect of this disclosure, a computer-readable storage medium is provided on which the computer program of the fourth aspect (or any example implementation thereof as described herein) is stored.
[0021] According to the sixth aspect of this disclosure, a computer program product is provided, the computer program product including the computer-readable storage medium of the fifth aspect.
[0022] As used herein, a computer-readable storage medium may be, for example, non-transitory and may be, for example, a hard disk drive (HDD), a solid-state drive (SSD), a USB flash drive, an SD card, a CD / DVD, and / or any other storage medium capable of non-transitory data storage. In other embodiments, a computer-readable storage medium may be transient and, for example, correspond to a signal (electrical signal, optical signal, mechanical signal, or similar signal) present on, for example, a communication link, line, or similar signal transmission device, in which case the computer-readable storage medium is more like a data carrier than a data storage entity.
[0023] Other objects and advantages of this disclosure will become apparent from the following detailed description, drawings, and claims. Within the scope of this disclosure and the appended claims, it is contemplated that all the features and advantages described with reference to, for example, the method of the first aspect, are relevant to, applicable to, and can be used in conjunction with the apparatus of the second aspect, the system of the third aspect, the computer program of the fourth aspect, the computer-readable storage medium of the fifth aspect, and the computer software product of the sixth aspect, and vice versa. Attached Figure Description
[0024] Exemplary embodiments will now be described with reference to the accompanying drawings, in which:
[0025] Figure 1 The illustrations schematically depict one or more exemplary embodiments of the apparatus according to the present disclosure;
[0026] Figure 2A The flowcharts schematically illustrate one or more exemplary embodiments of the method according to the present disclosure;
[0027] Figure 2B Schematic illustration for, for example Figure 2A The signaling diagram of the method;
[0028] Figure 3 Schematic illustration as such Figure 2A An example of a method that is part of a process for obtaining one or more image post-processing operations;
[0029] Figure 4A and Figure 4B The schematic illustrations depict components and functional blocks of various exemplary embodiments of the apparatus contemplated herein, as well as
[0030] Figure 5 The illustrations schematically depict example implementations of computer programs, computer program products, and computer-readable storage media as envisioned herein.
[0031] In the accompanying drawings and figures, unless otherwise stated, the same reference numerals will be used for the same elements. Unless explicitly stated to the contrary, the drawings show only elements necessary to illustrate exemplary embodiments, while other elements may be omitted or only suggested for clarity. As illustrated in the drawings, for illustrative purposes, the (absolute or relative) dimensions of elements and regions may be exaggerated or reduced relative to their true values, and therefore these dimensions are provided to illustrate the general structure of the embodiments. Detailed Implementation
[0032] First refer to Figure 1 and Figure 2A The present disclosure will now describe in more detail how it envisions providing a more user-friendly and flexible way to configure the visual appearance of images output from a camera.
[0033] Figure 1 The illustration schematically depicts an example of a device 100 for such a visual configuration according to the present disclosure, while Figure 2A The illustration schematically depicts an example of a method 200 for such a visual configuration performed by device 100 according to the present disclosure. Device 100 is configured to acquire (e.g., as part of operation S210 of method 200) a first image 121 of the scene. The first image 121 is captured, for example, by camera 120 and provided to device 100. Camera 120 may be, for example, a surveillance camera. Device 100 and camera 120 may be separate entities, or device 100 may be camera 120 itself.
[0034] The device 100 is further configured to acquire (e.g., as part of operation S220 of method 200) a second image 122 of the scene, wherein the second image 122 is expected to show the same scene, but has been visually modified according to the expectations of one or more users (such as user 130). For example, as contemplated herein, "same scene" means a scene captured using the same camera field of view, although the content of the scene may of course have changed due to, for example, movement of one or more objects in the scene and changes in lighting conditions.
[0035] For example, device 100 may be configured to send a copy of first image 121 to user 130. Optionally, device 100 may also send an identifier # such as a text string or number of the first image 121, wherein device 100 has associated the identifier # with a specific first image 121. User 130 may receive the first image 121 from device 100 and then perform image manipulation and / or image creation using any suitable software technology, such as provided by / on another device 140 such as a personal computer, tablet, smartphone, workstation, or server. Importantly, user 130 is not bound to use any specific software or interface to visually modify the first image 121, but can choose to use any software suitable for modifying the first image as desired by user 130. The result of modifying the first image 121 using the other device 140 is a second image 122. User 130 may use an interface such as an API, which may be provided by device 100 for this purpose, to send the second image 122 to device 100. An interface may also be provided to allow a user to receive the first image 121 from device 100, for example, upon a request from user 130. In some examples, user 130 may also send a copy of the first image 121 to the device. In other examples, if the identifier # of the first image 121 has been received, user 130 may instead (or additionally) send the identifier # of the first image 121. In the latter case, it is assumed that device 100 has stored a copy of the first image 121, so that it can retrieve the first image 121 based on the identifier # provided by user 130. In any case, it is envisioned that device 100 will possess both the first image 121 and a copy of the modified second image 122.
[0036] The apparatus 100 is further configured to extract (e.g., as part of operation S230 of method 200) the visual differences between the first image 121 and the second image 122. For example, the apparatus may begin by subtracting the first image 121 from the second image 122 to obtain a difference (“delta”) image that indicates to the user 130 how to modify the first image 121 to obtain the second image 122. Such a difference image may be created, for example, by subtracting the value of each pixel of the first image 122 (or multiple values if the image has more than one channel, such as in a color image) from the values of corresponding pixels in the second image 122. Other examples of how to identify the differences between the two images 121 and 122 are also contemplated.
[0037] After identifying the visual differences between the first image 121 and the second image 122, the device 100 is further configured to determine (as part of, for example, operation S240 of method 200) one or more image post-processing operations that should be applied to the first image 121 to make the first image 122 visually more similar (or equivalent) to the second image 122.
[0038] As envisioned herein, the term "visual difference" includes any modification to the first image 121 visible to an observer when comparing the first image 121 and the second image 122. To create such a visual difference, the user 130 modifies the first image 121 by changing at least one of the following: image contrast, image saturation, image brightness, image sharpness, image white balance, global tone mapping, and local tone mapping; and / or by adding one or more overlays to the first image 121, such as text overlay, image overlay, and vector overlay; and / or by adding one or more privacy masking elements, i.e., areas in the first image 121 that should not be visible to the observer of the image (or at least make objects in the image more difficult to identify). As envisioned herein, overlays and masking can be defined by drawing shapes on the first image 121, for example, using appropriate tools for drawing geometry such as circles, triangles, rectangles, or any type of polygonal object, or, for example, by defining such shapes freehand. Text overlays can be created, for example, by selecting one or more target regions in the image and subsequently typing one or more numbers and letters. Changes can also be created, for example, by selecting one or more areas in an image where the tone mapping should be altered, such as making certain colors visually sharper and / or darker than others. This tone mapping change can be local or global, such as applying only to a portion of the image or the entire image. Changes can also be made, for example, by inserting one or more geometric objects filled with one or more colors, where the fill can be opaque or at least partially transparent, and / or filled with one or more predefined patterns such as lines, dots, circles, and shadows.
[0039] Thus, as envisioned herein, one or more image post-processing operations could be operations that, if applied to the first image, at least partially recreate the changes made by user 130 to the first image 121 when creating the second image 122. For example, device 100 could identify a region in the image where a text overlay has been inserted based on a difference image, and the corresponding post-processing operation could include inserting such a text overlay at that location. As another example, device 100 could identify a specific region in the image where a privacy mask has been inserted, and the corresponding post-processing operation could include recreating such a privacy mask, and similarly for any conceivable visual modification. If the changes made involve the addition of one or more vector overlays, device 100 could vectorize the difference image (e.g., an increment buffer) to capture such vector graphics. The vector graphics could then be converted into a series of vector commands that could be executed by, for example, a suitable vector graphics library. If the changes made involve, for example, color or similar changes, i.e., parameters among one or more imaging parameters used by camera 120 that need to be changed, device 100 could calculate, for example, gain and / or color matrices based on a comparison of the first image 121 and the second image 122. This color matrix can then be used to determine the desired (global or local) tone mapping and, for example, white balance required to recreate the second image 122 from the first image 121.
[0040] As envisioned herein, user 130 may, for example, add overlays and / or masking (such as privacy masking). Text overlays can be defined by position, font, font size, font color, orientation, letter spacing, and shadow effects, etc. Image overlays can be defined by position, size, and image content, etc. Vector overlays or masking may include, for example, bounding boxes or polygons with defined line widths, line colors, sizes, fill colors, fill opacities, and fill patterns, etc. Freeform masking, etc., can be defined using tools like the "Magic Pen Tool" or similar tools. User 130 may, for example, change imaging parameters to alter contrast, saturation, brightness, and sharpness, etc., and / or, for example, change colors (e.g., making grass greener and the sky and / or water bluer, etc.), all depending on the visual result desired by user 130. Tools suitable for such modifications may include, for example, graphic tools, photo / image manipulation tools, CAD software, tools for vector image manipulation / creation, or any other suitable tools, including one or more tools implementing machine learning / artificial intelligence to modify images (such as first image 121), for example, based on user input such as prompts and voice commands, etc.
[0041] In some examples, device 100 can vectorize the difference image using any suitable software / library for this purpose. As envisioned herein, "vectorization" means converting raster image data into vector data, i.e., how the image is recreated is defined by one or more mathematical expressions rather than the original pixel data. For example, vectorizing a raster image of a line can result in a vector command drawing a line between corresponding points in the image, with line width and color corresponding to the line width and color of the rasterized line. Vectorizing a raster image of a circle can result in a vector command drawing a circle at a specific corresponding center point, with line color, fill color, and radius corresponding to the line color, fill color, and radius of the rasterized circle, and similarly for any other conceivable shape / object suitable for such rasterization.
[0042] After determining one or more image post-processing operations (such as those captured by camera 120) that the first image 121 needs to undergo in order to recreate an image that is visually more similar to (or equivalent to) the second image 122, the apparatus 100 can be configured to apply one or more post-processing operations as part of an image post-processing chain to generate a third image 123 that is visually more similar to the second image 122 than the first image 121 as a result of applying one or more post-processing operations. The third image 123 may, for example, be captured by camera 120 (as part of operation S250 of method 200), and then undergo post-processing operations and be output by apparatus 100.
[0043] As used herein, visual equivalence of two images does not preclude, for example, one image being scaled and / or rotated relative to the other. Rather, “visually equivalent” (or similar) means that the two images depict the same scene with a similar visual appearance (although not necessarily the same instantaneous image of the scene), i.e., such that the two images have, for example, the same white balance, color / tone mapping, the same overlay (such as graphics and text), and the same privacy masking. In other words, if one image can be transformed into another image by operations such as scaling, rotating, and perhaps translating, but without changing any of the properties listed above, then the two images can be considered visually equivalent / similar. Thus, performing the methods envisioned herein, but rescaling, translating, or rotating a third image, for example, still constitutes the third image being visually equivalent / similar to the second image.
[0044] In some examples, device 100 may be configured to generate a preview of a third image 123 by, for example, simulating the effects of operation and post-processing of camera 120, or by actually capturing another image of the scene using camera 120 and then post-processing that image. As part of, for example, operation S260, device 100 may provide a preview of the third image 123 to user 130, such that (when applied to the image captured by camera 120) user 130 can determine whether one or more specific post-processing operations produce the desired visual appearance of the output image 123 from device 100.
[0045] Figure 2B The signaling diagram schematically illustrates an example of how information can be exchanged between camera 120, device 100, user 130, and another device 140 used by user 130 to create a second image 122. Figure 2B The arrow and rectangle diagram method 200 provides various example operations, where dashed arrows and outer lines indicate optional operations.
[0046] For example, as part of operation S210a, device 100 may receive a first image 121 from camera 120, and as part of operation S212, device 100 may send the first image 121 to user 130. As part of operation S214, user 130 may provide the first image 121 to another device, using the software of the other device 140 to generate a second image 122 (as part of operation S216). For example, operation S214 may include user 130 uploading (or transmitting) the first image 121 to the other device 140, and as part of operation S218, user 130 receiving the second image 122 from the other device 140. Of course, user 130 may not directly receive either the first image 121 or the second image 122, but may coordinate the transmission of the first image from device 100 to the other device 140, and the transmission of the second image from the other device 140 to device 100. In any case, the result is that device 100 simultaneously possesses both the first image 121 and the second image 122, i.e., as part of operation S220, and optionally (receiving a copy of the first image 121 and / or the identifier # of the first image 121) operation S210b. As part of operation S230, device 100 extracts visual differences, and as part of operation S240, device 100 determines one or more image post-processing operations required to make the output image from device 100 / camera 120 visually more similar to / equivalent to the second image 122, such as the first image 121 described herein. As part of operation S250, device 100 may actually generate and output the third image 123, for example, by modifying the first image 121 using one or more post-processing operations, or by capturing a new image of the scene using camera 120 and instead applying one or more post-processing operations to that new image.
[0047] In some examples, device 100 may send a preview of the third image 123 to user 130 (as part of operation S260), and user 130 may check the preview and receive an indication (as part of operation S262) that user 130 is satisfied with the result. If user 130 is satisfied, device 100 may reconfigure camera 120 and perform image post-processing operations on further images captured by the camera, such that the visual appearance of subsequent output images from device 100 and / or camera 120 is consistent with the visual appearance of the second image 122. For example, device 100 may receive one or more further images from camera 120 (as part of operation S270) and post-process / generate corresponding further third images (as part of, for example, another operation S250). If user 130 is not satisfied with the determined / suggested post-processing operations, partial processing may be repeated (e.g., operations S214, S216, S218, S220, S230, S240) until satisfaction is achieved.
[0048] Figure 3This schematically illustrates an example 300 of how user 130 can modify a first image 121 to create a second image 122, and how device 100 can continue to generate / determine one or more image post-processing operations. The first image 121, captured and, for example, post-processed by device 100 using previously set parameters, includes multiple objects. User 130 uses additional device 140 to generate the second image 122 (as part of operation S260). In this example, the user adds a privacy mask 310 around the triangular object. The user also adjusts the local tone mapping in region 312, i.e., the circle is darkened slightly, and a text overlay 314 is inserted in the upper right corner of the image. As part of operation S230, device 100 compares the first image 121 and the second image 122 to find differences and determines (as part of operation S240) a set of image post-processing operations 320 required to make the further image captured by camera 120 visually more similar to / equivalent to the second image 122 than the first image 121. In some examples, operation S240 may include operation S242 of vectorizing the differential image, for example, finding a set of one or more vector commands to be included / performed as part of post-processing operation 320. In this particular example, determining the desired post-processing operation includes operation 322: adding privacy masking in a specific area of the image, performing local tone mapping 324 in another specific area of the image, and creating a text overlay by adding a bounding box (e.g., by appropriate vector commands) and a text string 328 to the upper right corner of the image. The device 100 then captures a new image, for example using camera 120, performs image post-processing operation 320 on the new image, and the resulting third image 123 is envisioned to be visually more similar to / equivalent to the second image 122 than the first image 121, thereby achieving the goal of allowing user 130 to more easily configure the visual appearance of the output image stream from camera 120 by only providing the second image 122, which indicates how user 130 wants to modify the current settings / post-processing operation.
[0049] This document also envisions that graphic overlays, such as those providing text and / or images on an image, do not need to be static, but can include, for example, placeholders for information that will change over time. For instance, a user could create a second image 122 by inserting a text overlay with labels indicating dynamic content such as date, timestamp, temperature, or any other parameter value, and the device 100 could be configured to detect such dynamic content (e.g., based on the user-inserted labels) and take it into account when generating the third image 123. For example, a text overlay including the label #datetime could instruct that "#datetime" be replaced with the current date and time when generating the third image 123. In other words, even if two images do not display exactly the same text or graphics, they can be considered visually similar / equivalent as long as they display the same type of text (such as date, time, temperature, etc.) and / or graphics, preferably in the same or similar locations in each image.
[0050] This article, and as Figure 2B As illustrated in the diagram, a system 150 is also envisioned, which includes a (first) device 100, a camera 120, and a further (second) device 140, namely a system for configuring the visual appearance of the output image from the camera 140 capturing the scene. In some examples, the camera 140 may be a surveillance camera. In some examples, the first device 100 may be the camera 140.
[0051] Figure 4AThe illustration schematically depicts one or more examples of apparatus 400 (such as apparatus 100) for performing a method 200 of configuring a visual appearance as contemplated herein. Apparatus 400 may be, for example, a camera (such as a surveillance camera), or some other apparatus external to such a camera but capable of acquiring images captured by the camera. Apparatus 400 includes at least a processor (or “processing circuitry”) 410 and optionally includes memory 412. As used herein, a “processor” or “processing circuitry” may be, for example, any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller (µC), digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), etc., capable of executing software instructions stored in memory 412. Memory 412 may be located external to processor 410 or may be located internal to processor 410. As used herein, “memory” may be any combination of random access memory (RAM) and read-only memory (ROM), or any other type of memory capable of storing instructions. Memory 412 contains (i.e., stores) instructions that, when executed by processor 410, cause device 400 to perform the methods described herein (i.e., method 200 or any implementation thereof). Device 400 may further include one or more additional items 414, which in some cases may be useful for performing the method. If device 400 is a camera, the additional items 414 may then include, for example, an image sensor and, for example, one or more lenses for focusing light from the scene onto the image sensor, such that the camera can capture an image of the scene as part of performing the envisioned method. The camera may be a video camera configured to capture a stream of images. The additional items 414 may also include, for example, various other electronic components required for capturing the scene, such as to properly operate the image sensor and / or lenses as desired. Performing the method in a surveillance camera can be useful because processing is moved to the “edge,” i.e., closer to the location where the actual scene is captured compared to performing image analysis elsewhere, such as at a more centralized processing server or similar location.
[0052] Device 400 may be connected to a network, for example, so that the results of performing method 200 can be transmitted to, for example, a user / operator and / or another device such as a server. For this purpose, device 400 may include a network interface 416, which may be, for example, a wireless network interface (as defined in any standard that supports, for example, Wi-Fi, IEEE 802.11 or subsequent standards) or a wired network interface (as defined in any standard that supports, for example, Ethernet, IEEE 802.3 or subsequent standards). Network interface 416 may also support, for example, any other wireless standard capable of transmitting encoded video, such as Bluetooth. Various components 410, 412, 414, and 416 (if present) may be connected via one or more communication buses 420, enabling these components to communicate with each other and exchange data as needed.
[0053] Device 400 may be, for example, a surveillance camera mounted or potentially mounted on a building or other supporting structure, such as in the form of a PTZ camera or, for example, a fisheye camera capable of providing a wider field of view, or any other type of surveillance / monitoring camera. In any such example of device 400, in addition to the components explained herein, device 4600 is envisioned to include all necessary components (if any), provided that device 400 is still capable of performing method 200 or any of its implementations as envisioned herein. In some examples, the various components of device 400 may be further configured to implement the method operations described herein (such as at least S210, S220, S230, and S240). In other examples, device 400 may be distributed across multiple physical and / or logical entities to form, for example, a computer system, where two or more operations (and / or two or more different sub-operations of the same operation) may be performed on / by different physical and / or logical entities, for example, as part of a distributed computing process, etc.
[0054] Figure 4B One or more embodiments of the apparatus 400 are schematically illustrated in the form of multiple functional / computing blocks 410a, 410b, 410c, and 410d. Each of these blocks 410a to 410d is responsible for performing the following tasks: Figure 2A and / or Figure 2BThe flowchart illustrates specific operational functions of method 200. For example, one such functional block 410a can be configured to acquire a first image of the scene. Block 410a can be referred to as a first image acquisition block / module, etc. Another block 410b can be configured to acquire a second image of the scene (e.g., from user 130) (as in operation S220). Block 410b can be referred to as a second image acquisition block / module, etc. Another block 410c can be configured to extract visual differences between the first and second images (as in operation S230), and can be referred to as, for example, a difference extraction block / module or a differencer, etc. Yet another block 410d can be configured to determine one or more image post-processing operations required to make the image captured by the camera more similar to / equivalent to the second image than the first image. Block 410d can be referred to as a post-processing operation determination block / module, etc. The apparatus 400 may optionally include one or more additional functional blocks, such as one or more blocks 410e, such as blocks for capturing a third image and subjecting the third image to post-processing operations determined by block 410d (as in operation S250), to provide a preview of the third image (as in operation S260), and / or to provide / perform any operation of method 200 as contemplated herein but not yet covered by any of the blocks 410a to 410d.
[0055] Generally, each functional block 410a to 410e can be implemented in hardware or software. Preferably, one or more or all functional blocks 410a to 410e can be implemented by processing circuitry 410, which can cooperate with storage medium / memory 412 and / or communication interface 416. Processing circuitry 410 can therefore be arranged to retrieve instructions as provided by functional blocks 410a to 410e from memory 412, execute those instructions, and thereby perform any operation by method 200 or any implementation thereof as disclosed herein, performed by / in device 400 / device 100.
[0056] Figure 5 The illustration schematically depicts a computer program product 510 including a computer-readable device / storage medium 530. On the computer storage medium 530, a computer program 520 (including computer code) may be stored, which, when executed, can cause the processor 410 of the device 400 and entities and devices operatively coupled thereto, such as a communication interface 416 and a memory 412, to perform method 200. The computer program 520 and / or the computer program product 510 can therefore provide means for performing any operation of method 200 (or any implementation thereof) as disclosed herein by the device 400.
[0057] exist Figure 5In the example, computer program product 510 and computer-readable storage medium 530 are illustrated as an optical disc such as a CD (optical disc), DVD (digital versatile optical disc), or Blu-ray disc. Computer program product 510 and computer-readable storage medium 530 may also be embodied as a memory such as random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM), and more specifically, as a non-volatile storage medium such as USB (Universal Serial Bus) memory or flash memory such as compact flash memory in an external memory device. Therefore, although computer program 520 is schematically shown herein as a track on the depicted optical disc, computer program 520 may be stored in any manner suitable for computer program product 510 and computer-readable storage medium 530. Storage medium 530 may be non-transitory or transient.
[0058] In summary, this disclosure improves upon contemporary technology by providing an improved way to allow users to (re)configure the visual appearance of images captured by a camera. Instead of relying on a complex interface that exposes multiple configuration parameters and / or a visually cluttered interface that risks confusing the user, the envisioned solution allows the user to obtain an image of the scene captured by the camera and then perform post-processing using the camera's current settings. The user can then use any preferred software to visually modify the image to meet their expectations, and the envisioned solution then compares the original and modified images to automatically determine at least how the post-processing operation should be modified so that the future image captured by the camera is visually more similar to / equivalent to the modified image provided by the user. Therefore, the user does not need to use any complex interface, nor does such an interface need to be provided and maintained by the camera provider.
[0059] Although features and elements may be described above in specific combinations, each feature or element may be used alone without other features and elements, or in various combinations with or without other features and elements. Furthermore, those skilled in the art, in practicing the claimed invention, can understand and implement variations of the disclosed embodiments through study of the drawings, this disclosure, and the appended claims.
[0060] In the claims, the words "comprising" and "including" do not exclude other elements, and the indefinite articles "a" or "an" do not exclude a plural. The fact that certain features are set forth in mutually different dependent claims does not imply that combinations of these features cannot be used advantageously. List of reference numerals in the drawings.
[0061] 100 devices
[0062] 120 camera
[0063] 121 First Image
[0064] 122 Second Image
[0065] 123 Third Image
[0066] 130 users
[0067] 140 Device for Image Modification
[0068] 150 system
[0069] 200 methods
[0070] S210, S220, S230, S240 Method Operation
[0071] S250 and S260 optional operation methods
[0072] Example of 300 image modification / setting reconfiguration
[0073] 310 Privacy Coverage
[0074] 312 Local Tone Mapping
[0075] 314 Text / Graphic Overlay
[0076] Post-processing operations for images 320, 322, 324, 326, and 328.
[0077] 400 device
[0078] 410 Processor / Processing Circuit
[0079] 412 memory
[0080] 414 Optional additional items
[0081] 416 communication interface
[0082] 420 data bus
[0083] Functional blocks 410a to 410e
[0084] 510 Computer Program Products
[0085] 520 computer program
[0086] 530 Computer-readable storage media
Claims
1. A computer-implemented method (200) for configuring the visual appearance of an output image from a camera (120) capturing a scene, comprising: (S210) Obtain (121) a first image (S21) depicting the scene captured by the camera; The first image or a copy of the first image is provided to one or more users (130) to allow the one or more users to visually modify the first image or a copy of the first image; (S220) Obtain (122) a second image (122) depicting the scene from the one or more users, wherein the second image corresponds to the first image after being visually modified according to the expectations of the one or more users; Extract (S230) a set of one or more visual differences between the first image and the second image, wherein the one or more visual differences are caused by changes in at least one of contrast, saturation, brightness, sharpness, white balance, global tone mapping, and local tone mapping, and / or by the addition of at least one overlay or privacy masking. Based on the extracted set of one or more visual differences, one or more image post-processing operations (320) are determined (S240) such that a third image (123) depicting the scene, captured by the camera and processed by the one or more image post-processing operations, is visually more equivalent to the second image than the first image.
2. The method according to claim 1, further comprising: The camera is used to capture (S250) the third image of the scene, and the third image is subjected to the one or more image post-processing operations.
3. The method according to claim 2, further comprising: (S260) Provide (S130) a preview of the third image after post-processing of the one or more images to one or more users (130) who are responsible for providing the second image by modifying the first image.
4. The method according to claim 1, wherein, The method includes: providing the first image to the one or more users, and then obtaining both the first image and the second image from the one or more users.
5. The method according to claim 1, wherein, The method includes: providing the first image and an identifier of the first image to the one or more users, and then obtaining the second image and the identifier of the first image from the one or more users.
6. The method according to claim 1, wherein, The one or more visual differences are caused by changes in at least one of contrast, saturation, brightness, sharpness, white balance, global tone mapping, and local tone mapping, and wherein the determination includes: Calculate at least one of the gain and / or color matrices based on one or more extracted differences, and include applying the at least one of the gain and color matrices as part of the one or more image post-processing operations.
7. The method according to claim 1, wherein, The one or more visual differences are caused by the addition of at least one vector superposition, and wherein the determination includes: At least some of the extracted differences are vectorized to generate one or more vector commands for generating the vector overlay in the third image, and the one or more vector commands are executed as part of the one or more image post-processing operations.
8. An apparatus (100) for configuring the visual appearance of an output image from a camera (120) capturing a scene, the apparatus (100) including processing circuitry configured to: Obtain a first image (121) depicting the scene captured by the camera; The first image or a copy of the first image is provided to one or more users (130) to allow the one or more users to visually modify the first image or a copy of the first image; Obtain a second image (122) depicting the scene, wherein the second image corresponds to the first image after it has been visually modified according to the expectations of the one or more users; Extract a set of one or more visual differences between the first image and the second image, wherein the one or more visual differences are caused by changes in at least one of contrast, saturation, brightness, sharpness, white balance, global tone mapping, and local tone mapping, and / or by the addition of at least one overlay or privacy masking. Based on the extracted set of one or more visual differences, one or more image post-processing operations (320) are determined such that a third image (123) depicting the scene, captured by the camera and processed by the one or more image post-processing operations, is visually more equivalent to the second image than the first image.
9. The apparatus according to claim 8, wherein, The device is the camera (120).
10. A system (150) for configuring the visual appearance of an output image from a camera capturing a scene, comprising: The first device (100, 400) according to claim 8; A camera (120) is used to capture the first image of the scene; The second device (140) external to the camera includes a processing circuit configured to: receive the first image, provide image manipulation and / or image creation software that allows the one or more users to generate the second image by visually modifying the first image, and provide the second image to the processing circuit of the first device.
11. The system according to claim 10, wherein, The camera is a surveillance camera, and the first device is the surveillance camera.
12. A non-transitory computer-readable storage medium (530) storing a computer program (520), said computer program comprising computer code, said computer code causing the device (400) to: Obtain a first image (121) depicting the scene captured by the camera; The first image or a copy of the first image is provided to one or more users (130) to allow the one or more users to visually modify the first image or a copy of the first image; Obtain a second image (122) depicting the scene, wherein the second image corresponds to the first image after it has been visually modified according to the expectations of the one or more users; Extract a set of one or more visual differences between the first image and the second image, wherein the one or more visual differences are caused by changes in at least one of contrast, saturation, brightness, sharpness, white balance, global tone mapping, and local tone mapping, and / or by the addition of at least one overlay or privacy masking. Based on the extracted set of one or more visual differences, one or more image post-processing operations (320) are determined such that a third image (123) depicting the scene, captured by the camera and processed by the one or more image post-processing operations, is visually more equivalent to the second image than the first image.