Optimal body or face protection using adaptive dewarping based on contextual segmentation layer
Through the adaptive de-distortion method, combined with image segmentation and depth analysis, the problem of image distortion in wide-angle images is solved, and the visual effect of keeping straight lines straight and conserving proportions is achieved.
Patent Information
- Application Number
- CN202080056969.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-11
- Filing Date
- 2020-06-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-06-11
AI Technical Summary
The prior art is difficult to effectively correct image distortion in wide-angle images in photography, especially while maintaining the proportion of the object and straightness of the straightness of the straightness of the line, resulting in the image being visually unsatisfactory.
Adaptive de-distortion method is adopted, through image segmentation and depth analysis, different de-distortion algorithms are applied according to the type and depth of the object to ensure that the straight lines in the image remain straight and avoid undesirable proportional deformation.
It realizes the visual attractiveness of the image while maintaining the conservation of the straightness and proportion of the image, and avoids the undesirable consequences of image distortion.
Smart Images

Figure CN114175091B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of currently pending U.S. Provisional Patent Application No. 62 / 859,861, filed on June 11, 2019, entitled “Method for Adaptive Dewarping Based on Context Segmentation Layers,” the entire contents of which are incorporated herein by reference. Background Art
[0003] Embodiments of the present invention relate to the field of photography, and more particularly to how image distortion in wide-angle images may be corrected differently depending on image context, segmentation layers, and / or depth of objects visible in the image.
[0004] In photography, in the case of narrow angle lenses with a full field of view below 60°, it is often desirable to have an image in which straight lines in the object remain straight. This is achieved by making the image follow as closely as possible the straight line H=f*tan(θ) relationship between the image height H and the field angle θ, which is still feasible in narrow angle lenses. In the case of a very limited full field of view below 60°, this type of straight line H=f*tan(θ) relationship does not significantly affect the object scale on the periphery of the image. Images that fully follow this relationship are considered to have no optical distortion. For optical lenses that do not fully follow this relationship for all field angles θ, the images obtained from these lenses are considered to have some optical distortion. This optical distortion is particularly present in wide angle images with a full field of view exceeding 60°. Correcting residual image distortion of wide angle images or purposefully modifying them are known techniques in image processing, which are often used when the optical lens itself cannot be designed to produce the desired projection for the desired application.
[0005] Although rectilinear projection is ideal for keeping straight lines of objects straight in an image, it is sometimes not the projection that creates the most visually pleasing images in photography. One such example is a group selfie, or group photo, with a wide-angle lens, where people are positioned at various locations in the field of view. People in the center appear to have normal proportions, but people toward the edges appear stretched and deformed because of the rapidly increasing number of pixels / degrees of the projection. This unsatisfactory visual effect on human faces is visible not only in lenses with rectilinear projection, but also in every lens that is not specifically designed to keep the proportions visually pleasing.
[0006] Some image processing algorithms or some lenses are specifically designed to limit this undesirable effect towards the edges by limiting the rapid increase in the number of pixels / degrees towards the edges at the expense of creating curved lines. In other words, even with a perfect calibration of the lens and a dewarping algorithm, the dewarping algorithm can only correct either straight lines or facial proportions because these two corrections require different dewarping projections. If the correction algorithms are optimized to provide more visually pleasing images of humans positioned towards the edges of wide-angle images, a process known as body and face protection, they will have the undesirable consequence of doing so, since geometric distortion is added to the image, and the resulting image will have curved lines even if the original object scene consists of straight lines. Conversely, if the correction algorithm is optimized to straighten the lines in the image, it will worsen the proportions of humans positioned towards the edges.
[0007] For example, the image distortion transformation method proposed in U.S. Pat. No. 10,204,398 B2 is used to transform a distorted image of an original image from an imager into a transformed image, wherein the distortion of the image is modified according to a preconfigured or selected target image distortion profile. Even though the target distortion profile may be asymmetric, for example to maintain a similar field of view, the target distortion profile is nevertheless applied to the entire image without regard to the location of objects in the image or their depth. Thus, when using this method, the appearance of straight lines may be improved while the appearance of people in the image may be deteriorated, or vice versa.
[0008] There are some other existing geometric correction methods for distorted images, such as perspective tilt correction using the horizon proposed in US Pat. No. 10,356,316 B2. However, these methods can only correct the perspective of the entire image, but cannot apply correction to specific elements.
[0009] Another problem to overcome is the fact that real lenses from a mass production batch are all slightly different from each other due to tolerance errors in the shape, position or orientation of each optical element. These tolerance errors produce slightly different distortion profiles for each lens of the mass production batch, and therefore, even after de-warping the image based on the theoretical distortion curve for the mass-produced lens, there may still be residual geometric distortion in the image.
[0010] To achieve body and face preservation, i.e., having the most visually appealing human proportions while still making straight lines in the object look like straight lines in the image, there exist some more advanced image processing algorithms that apply specific image dewarping depending on the content of the image. However, these algorithms have the undesirable consequence of destroying the perspective in the background when applying corrections for foreground objects or people. New methods are needed to overcome all of these problems. Summary of the invention
[0011] To overcome all of the problems previously mentioned, embodiments of the present invention propose an adaptive dewarping method to apply different dewarping algorithms based on scene context, the location of objects, and / or the depth of objects present in the image. Starting from an initial input image, the method first applies an image segmentation process to segment the various objects visible in the original image. Then, the depth of each object is used to sort all objects by layer according to their object type and their depth relative to the camera. The depth values used to segment these layers are either absolute depth measurements or relative depths between objects, and can be calculated using an artificial intelligence neural network that is trained to infer depth from 2D images, from parallax calculations using a pair of stereo images, from depth measurement specific devices (such as time-of-flight sensors, structured light systems, lidar, radar, 3D capture sensors, etc.). A specific dewarping algorithm or projection will be used to dewarp each category of objects identified in the original image. For example, if a human face is detected in the original image, a specific dewarping algorithm will be used on the human to avoid stretching the human face and make them more visually attractive, a process known as face protection. If the same original image also contains a building, the adaptive dewarping algorithm will apply a different dewarping on the building to keep the line straight. According to the present invention, there is no restriction on the type of objects that can be identified by the adaptive dewarping method and the dewarping applied thereto. The dewarping to be applied can be predefined as a preset for a specific object type (e.g., a human face, a building, etc.), or can be calculated to respect well-known characteristics of the object, such as facial proportions, human body proportions, etc. The method is applied on each segmentation layer and depth layer, including applying the method to the background layer. The background layer consists of objects that are far away from the camera and do not have a preset distortion dewarping. In a preferred embodiment according to the present invention, the background is dewarped so as to keep the perspective of the scene undistorted. Compared with the prior art, the adaptive method allows different dewarping to be applied to each type of object based on type, layer, depth, size, texture, etc.
[0012] In some cases, when a given layer is deformed, the adaptive dewarping may create areas of missing information in the later layers because different dewarping is applied to the later layers. Only when this happens can an additional image completion step be applied on the resulting image to further make it more visually appealing. The image completion step consists of the following operations: segmenting the object after adaptive dewarping based on the context in several depth layers. Some depth layers have missing information, and some other depth layers do not have any missing information. Then, a completion algorithm is used on the layer with missing information to fill in the area with missing information. This can be done by applying blur based on the texture and color around the missing information area, applying a gradient line that gradually changes the color from the color on one side of the missing information area to the color on the other side, using an artificial intelligence network trained to complete the missing information of the picture, etc. The completion algorithm outputs completed depth layers, which can then be merged back into a single image with the filled information, where perspective correction is applied, the shape of the person is corrected to avoid unsatisfactory stretching, and the missing background information is filled using the completion algorithm.
[0013] In some embodiments according to the invention, the dewarped projection for the background layer depends on the context identified in the foreground object.
[0014] In some embodiments according to the present invention, the adaptive dewarping method is used to: maximize the straightness of lines in the image compared to the original lines in the object scene, maximize the full field of view of the output image compared to the full field of view of the original image, and / or maximize the conservation of proportions in the output image compared to the true proportions in the object scene.
[0015] In some embodiments according to the present invention, instead of using a completion algorithm to complete the missing information, the relative magnification of the previous layer is increased to cover the missing information area in the later layer. In some cases, this technique may make the front objects appear larger or closer than they were in the original image.
[0016] In some embodiments according to the invention, the selection of the dewarped projection for the background layer depends on the detected context of the original wide-angle image.
[0017] In some embodiments according to the invention, the processing includes creating a virtual camera centered on the element having the distorted geometry, applying a rectilinear correction on the virtual camera, and translating the result to the correct position in the final image.
[0018] In some embodiments according to the invention, the context and segmentation layer based adaptive dewarping method includes processing by a processor inside a physical device that also utilizes an imager to create the original wide angle image and displays the final image to a display screen. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The foregoing summary of the invention and the following detailed description of the preferred embodiments of the invention will be better understood when read in conjunction with the accompanying drawings. For illustrative purposes, presently preferred embodiments are shown in the accompanying drawings. However, it should be understood that the invention is not limited to the precise arrangements and means shown.
[0020] In the attached picture:
[0021] Figure 1 The resolution curve for a rectilinear image is shown;
[0022] Figure 2 shows how existing wide-angle cameras with no or small deviations from rectilinear projection create visually unsatisfactory views;
[0023] Figure 3 shows the resolution curve of an image from a wide-angle camera with deviations from rectilinear projection;
[0024] Figure 4 shows how existing wide-angle cameras with large deviations from rectilinear projection also create visually unsatisfactory views;
[0025] Figure 5 Basic methods for applying corrections to images to make them more visually pleasing but affecting perspective are shown;
[0026] Figure 6 An adaptive dewarping method in its previous form is shown;
[0027] Figure 7 An adaptive dewarping method based on context segmentation and depth layer is shown;
[0028] Figure 8 A method for filling in missing image information after an adaptive dewarping method has been applied is shown;
[0029] Fig. 9 A method for hiding missing image information after an adaptive dewarping method has been applied is shown;
[0030] Fig.10 shows a method in which the dewarped projection of a background layer depends on the context of objects in the foreground;
[0031] Fig.11The steps of an algorithm according to a depth and segmentation layer based context-based adaptive dewarping method for ideal face preservation are shown; and
[0032] Fig.12 An example embodiment of a physical device that captures a raw wide-angle image, processes it, and outputs a final image on a display screen is shown. DETAILED DESCRIPTION
[0033] Certain terms used in the following description are for convenience only and are not limiting. The words "right", "left", "bottom" and "top" indicate directions in the drawings to which reference is made. The terminology includes the above words, their derivatives and words of similar meanings. Additionally, the words "a" and "an" used in the claims and corresponding parts of the specification mean "at least one".
[0034] It should also be understood that the terms "about," "approximately," "generally," "substantially," and similar terms used herein when referring to a dimension or characteristic of a component indicate that the dimension / characteristic described is not a strict boundary or parameter and does not exclude minor variations thereof that are functionally similar. At a minimum, such references, including numerical parameters that will include variations using mathematical and industrial principles accepted in the art (e.g., rounding, measurement or other systematic errors, manufacturing tolerances, etc.), will not change the least significant digit.
[0035] Figure 1Theoretical resolution curves for a perfectly rectilinear image with a 40° half field of view at 100 and a 70° half field of view at 150 are shown, corresponding to full fields of view of 80° and 140°, respectively. Most narrow-angle imaging lenses designed for use in photographic applications with a full field of view of less than 60° aim to have as low image distortion as possible by following a rectilinear image projection as closely as possible. In a rectilinear lens, the relationship between the image height H on the image sensor and the half field of view angle θ in the object plane follows the equation H=f*tan(θ) as closely as possible. Narrow-angle lenses with a full field of view of less than 60° typically follow this projection. However, wide-angle lenses (also known as panoramic lenses) with a full field of view greater than 60° typically do not follow this H=f*tan(θ) equation exactly. An image that perfectly follows the H=f*tan(θ) equation - either directly from capturing the image from the imaging lens on the image sensor or after any hardware or software distortion correction or dewarping - is considered to have no image distortion. Any deviation from this equation is called image distortion, geometric distortion, or optical distortion, and is generally avoided in photography. Deviations from this equation are also related to TV distortion, in which the corners of rectangular objects appear to be extended or compressed in the image rather than perfectly rectangular. Diagrams 100 and 150 show resolution curves as a function of half field of view. The resolution curves are obtained by taking the mathematical derivative of the position curve as a function of the half field of view angle θ. In diagram 100 for the case of a full field of view of 80°, the value 110 represents a 1x magnification at the center of the field of view at a half field of view angle θ of 0°. Alternatively, instead of calculating the resolution as a magnification ratio relative to the center, the resolution can also be calculated in mm / degree, mm / radian, pixels / degree, or pixels / radian, etc. The value of pixels / degree is particularly useful when the image sensor consists of pixels of constant size, which is the most common case. For a half field of view angle θ value of 40°, the resolution value 112 is 1.7 times the resolution at the center 110 for the theoretical straight line projection, and the resulting image appears to have been stretched. On the diagram 150 for the case of a full field of view of 140°, the value 160 represents a 1x magnification at the center of the field of view at a half field of view angle θ of 0°. For a half field of view θ value of 45°, the resolution value 162 is 2 times the resolution at the center 160 for the theoretical straight line projection, and the resulting image appears to be even more stretched. For wider half field of view angles θ, the difference in resolution from the center to the edge becomes increasingly larger, and the image becomes even more stretched and is unsatisfactory for some photographic applications. For example, using theoretical rectilinear projection, at a half field of view value of 60°, resolution 164 is 4 times greater than resolution 160. At a half field of view value of 70°, resolution 166 is 8.55 times greater than resolution 160.At larger half field of view values, the resolution keeps increasing until it reaches infinity at a half field of view angle of 90°.
[0036] Figure 2 Example images of a group selfie or group photo are shown as it would look when captured by a theoretically perfect rectilinear lens, or as it would look after hardware or software correction of image distortion to obtain an image with a perfectly rectilinear projection. In example image 200, the full field of view in the diagonal direction of the image is 80°, while in example image 250, the full field of view in the diagonal direction of the image is 140°. In example image 200 with an 80° diagonal field of view, we can see that person 212 with his head in the center of the image looks normal because the resolution is almost constant in the central area of the image. However, for person 214 with his head toward the edge, the face is stretched in the direction away from the center and it looks deformed. This phenomenon of rectilinear projection is visually unsatisfactory for consumer photography applications, but this stretching is needed to keep the lines in the object scene looking straight like horizontal lines 220, vertical lines 230, or vanishing lines 240. Similarly, in the example image 250 with a 140° diagonal field of view, we can see that the person 262 whose head is in the center of the image looks normal because the resolution is almost constant in the central area of the image. However, for the person 264 whose head is toward the edge, the face is stretched in the direction away from the center and it looks deformed. For the person 266 whose head is at the half field of view angle closer to the 70° corner, the stretching is even more obvious. Again, this phenomenon of rectilinear projection is visually unsatisfactory for consumer photography applications, but this stretching is needed to keep the lines in the object scene looking straight like horizontal lines 270, vertical lines 280, or vanishing lines 290.
[0037] Figure 3 The resolution curves for more satisfactory wide-angle images are shown, which are either directly obtained from a wide-angle lens with a maximum in the area between the center and the edge followed by a resolution drop-off toward the edge, or after the image distortion has been purposely modified using hardware or software dewarping or correction algorithms to avoid Figure 2, which is obtained after the undesirable effect of . This resolution curve 300 - which has compression zones at the center and edges of the field of view and expansion zones in the middle region of the image located between the center and the edges of the image - is merely an example resolution curve that creates a more visually appealing image, but there are other resolution curves that create more visually appealing images. This kind of resolution curve 300 - the curve 300 has a given resolution value 310 at the center, which increases smoothly until a maximum value 312 and then drops back to an edge value 314 - is typical of some wide-angle lenses or ultra-wide-angle lenses, which create expansion zones in the middle region of half the field of view and compression zones at the edges to make the image visually pleasing. In this example, the maximum resolution is at a half field angle θ of about 45°, with a magnification value of about 2x, such as Figure 1 As is the case with a straight line curve, many similar resolution curves with different maximum magnification values and positions can be used to create a more visually pleasing image. In addition, in this example, the average magnification value is about 1.5x, which is also the magnification of the equidistant H=f*θ projection that creates an image of the same size for the same total field of view. In this example, in order to create a more visually pleasing image, the magnification 310 at the center is lower than the average, the magnification 312 at the maximum magnification is higher than the average, and the magnification 314 at the edge of the field of view is lower than the average magnification value. However, in some other embodiments according to the present invention, the magnification 314 at the edge of the field of view may also be higher than the average magnification value.
[0038] Figure 4 Two example images 400 and 450 of a group selfie or group photo are shown, as it is respectively shown in a video frame having a similar Figure 3 The resolution curve of the lens will look like when captured, as well as the image distortion correction for scale-saving projection (also known as body and face protection) in order to avoid or minimize Figure 2 In the top image 400, the head of the person 422 standing at the center still looks normal. The heads of the persons 424 and 426 standing at the half-viewing angles of 45° and 70°, respectively, also look smaller than those in the top image 400. Figure 2 This is more normal because Figure 3 The selected projection of the resolution curve does not have Figure 1430 . In the case of the curve of , there is a large increase in resolution towards the edges. In photography, this result for the face is more visually pleasing. However, because the lens does not follow the straight line projection mapping equation H=f*tan(θ), there are geometric distortions in the image, and straight lines in the object scene do not appear straight in the image, as seen in the case of the curved horizontal line 430 and the vertical line 435. However, in this example, the vanishing lines 440 remain straight because they are oriented in the radial direction from the center of the image. In the bottom image 450, additional image processing has been performed to obtain an image with a perfect ratio-saving projection (also known as face and body protection). With this projection, the heads of people 472, 474, and 476, respectively standing at the center, at a half field of view angle of 45°, and at a half field of view angle of 70°, all have similar proportions due to the body and face protection correction, which is also visually pleasing for group selfie pictures. However, because this face-preserving projection that maintains scale does not follow the rectilinear projection mapping equation H = f*tan(θ), there are geometric distortions in the image and straight lines in the object do not appear straight in the image, as seen in the case of the curved horizontal lines 480 and vanishing lines 490. In this example projection, the vertical lines 485 remain straight in the image, but this is not always the case.
[0039] Figure 5 A simple method according to the present invention is shown to keep the original straight lines of objects straight in the image and to ensure that some objects (such as people's faces) are not over-stretched when they are close to the edge of the image. The method allows the original wide-angle image to be enhanced based on the image context. In the original image 500 with a rectilinear projection - the image 500 either comes from using a lens with a H=f*tan(θ) distribution function or after the distortion is corrected using image processing, the person 522 in the center looks normal, and the people 524 and 526 towards the edges appear to be increasingly stretched. In the case of this original image, all horizontal, vertical and vanishing lines 530, 535 and 540 are straight. The original image 500 is similar to Figure 2A simple method for improving the appearance of the image is to locally correct the shape around the stretched object (such as heads 524 and 526) while keeping the lines straight, resulting in example image 550. The method begins by receiving an original wide-angle image having at least one element, the at least one element having a distorted geometry. Here, the distorted geometry can be of any kind, including uneven or stretched proportions, curved lines when the original object was straight, unsatisfactory optical distortion, or any other unsatisfactory artifacts visible in the original wide-angle image. The method then creates at least one classified element by classifying the at least one element having a distorted geometry from the original wide-angle image. Here, the classification of the element can be based on various methods, including but not limited to the shape of the at least one element, the position of the at least one element in the original wide-angle image, the depth of the at least one element compared to other elements in the original wide-angle image, etc. The method then allows the distorted geometry to be de-distorted by processing the original wide-angle image, thereby creating a final image. This type of processing using a correction algorithm can be done by an AI algorithm that uses deep learning to correct the shape of the object, or this type of processing can be done by a traditional image deformation algorithm that corrects the unsatisfactory shape by knowing where the unsatisfactory shape is in the image and creating its resolution curve. In this final image, the correction is performed on the entire image without distinguishing the previous layer from the background. The correction can be performed by transforming the texture mesh or display mesh. Alternatively, the correction can also be performed by dewarping the image pixel by pixel. The correction of the image makes people 574 and 576 towards the edge look more normal, just like the person 572 standing in the center, that is, body and face protection is applied. However, since the local deformation is applied to the entire image, the undesirable result from the correction foreground destroys the perspective in the background. In this example, straight lines that are not touched or hidden by the foreground object remain straight, such as lines 580 and 590. This is not the case for straight lines hidden behind foreground objects (such as people 574 or 576). Lines that were continuous in the real object scene and in the resulting original wide-angle image are now discontinuous in the final image due to the correction applied on the foreground. This is particularly evident where line segments 581 and 582 form a continuous horizontal line in object space but a discontinuous line in the corrected image, or where line segments 585 and 586 form a continuous vertical line in object space but a discontinuous line in the corrected image. The figure shows an example of a simple correction method based on classification and without a segmentation layer, but the method is not limited to humans and can be applied to a variety of other objects to correct their distorted geometry due to the non-linear magnification of the image across the entire field of view.
[0040] Figure 6 Another example of an image context-based adaptive dewarping method in its simple form using element classification is shown. Figure 6 In this example, the original image 600 has several distorted geometries, including curved lines 610, unequal proportions between human faces as seen by the size difference between the face at the center 612 and the face at the edge 614, and the image has a diagonal field of view of 140°. The original image 600 is merely an example original image that was either captured directly from a wide angle imager or after some processing has been applied, but the methods according to the present invention are not limited to any value of the scene content or diagonal full field of view. Some of the existing methods completely correct these proportions by using a scale-saving projection. This is like Figure 4 As in the previous example of , represented here by image 620, where line 630 is even more curved, but the face proportions are equal, as seen in the equal size of face 632 at the center and face 634 at the edge. In this example, the scale-saving projection can be, but is not limited to, an equirectangular projection, a cylindrical projection, or any other custom projection. Some other methods in the existing methods completely straighten the lines in the image, as in Figure 2As in the example of , represented here by image 640, where the line 650 is straight, but the facial proportions are even worse than in the original image, as seen by the larger size difference between the face 652 at the center and the face 654 at the edge than in the original image. In both of these existing methods, it is not possible to conserve the original full field of view when modifying the image projection. In the case of the method of the present invention, the original wide-angle image is processed using a simple form of an adaptive dewarping method based on image context to equally maximize the straightness of the lines in the final image compared to the original lines in the object scene, maximize the final image full field of view compared to the original wide-angle image full field of view, and maximize the conservation of proportions in the final image compared to the true proportions in the object scene. Again, the method begins by receiving an original wide-angle image having an element, the element having a distorted geometry. The method then creates at least one classified element by classifying the at least one element having a distorted geometry from the original wide-angle image. Here, the classification of the elements can be based on various methods, including but not limited to the shape of the at least one element, the position of the at least one element in the original wide-angle image, the depth of the at least one element compared to other elements in the original wide-angle image, etc. Then, the method allows the distorted geometry to be dedistorted by processing the original wide-angle image to create a final image, and maximize the field of view of the final image. Correction can be performed by transforming the texture grid or the display grid. Alternatively, correction can also be performed by dedistorting the image pixel by pixel. The resulting image 660 has a line 670 that is straighter than in the original image 600 but not as straight as in the image 640. The resulting image 660 also has a facial ratio that is more equal than in the original image 600 but not as equal as in the image 620, as seen by comparing the ratio of the face 672 at the center to the face 674 at the edge. Finally, the diagonal full field of view in image 660 remains as close as possible to the value 140° from the original image 600 to avoid creating areas without information in the corners or on the sides of the image, or to avoid having to crop the image to avoid such areas without information. In the case of the adaptive dewarping method using this simple form, the level of the ideal balance between these three items depends on which compromise is acceptable for the desired application. In some embodiments, each of the three items to be maximized can be assigned an adjustable correction weight to adjust the level of processing performed on the original wide-angle image. These adjustable correction weights are either predefined, for example, by the requirements of the application, or selected by the user according to their preferences. Depending on the input original image content, context, and application, the level at which the curved lines are straightened, the conservation of the field of view, and the conservation of the object proportions applied can vary according to the adaptive dewarping method according to the simple form of the present invention.In some embodiments, the processing steps of the method may be performed by an artificial intelligence algorithm.
[0041] Figure 7 A preferred method for adaptive dewarping based on image context segmentation and segmentation layers according to the present invention is shown. The method receives as input an original wide-angle image 700 having a plurality of elements, each element being in one of the foreground or background of the original wide-angle image, one or more of which have distorted geometric shapes. These distorted geometric shapes can be of any kind, including uneven or stretched proportions, curved lines when the original object was straight, unsatisfactory optical distortions, or any other unsatisfactory artifacts visible in the original wide-angle image. The wide-angle image can have any field of view, but is typically Figure 2 and Figure 4The unsatisfactory effect shown in is most obvious in wide-angle images with a full field of view of more than 60°. In a preferred embodiment according to the present invention, the original wide-angle image 700 is captured directly by an imager having an optical system that includes at least a camera module and a wide-angle optical lens with or without a deviation from a rectilinear projection. The wide-angle lens typically has an angular field of view of at least 60°. In other embodiments, the optical system is composed of any combination of refractive lens elements, reflector elements, diffractive elements, metasurfaces, or any other optical elements that help form an image in the image plane of the camera. In some other embodiments according to the present invention, the original image 700 has been processed by a processor to correct the original distortion from the camera module, improve image quality, or apply any other image processing. Alternatively, the original wide-angle image 700 can be created by combining multiple narrow-angle images inside an imager with a processor, or is completely generated by a computer. The original wide-angle image 700 has elements that are visually unsatisfactory for human observers. In this example diagram, without limiting the scope of the invention in any way, these elements are a normal-looking human 702 at the center, a human 703 at the edge with an unsatisfactory stretched face, a tree 704 at the deformed edge, a building 706 on the edge, a building 707, which appears curved due to image distortion even though the building 706 is straight in the object scene, and a background 708, which appears normal and straight due to its position at the center, and a background 708 composed of various distant objects such as mountains or the sun. After receiving the original wide-angle image, the method performs an object segmentation and depth analysis step 710, which is based on element depth and image context. This first processing step for segmenting the original wide-angle image into multiple segmentation layers is performed via a software algorithm, a hardware algorithm, or a trained artificial intelligence algorithm, or is not performed via a neural network. This first processing step can be performed inside a processor, a CPU, a GPU, an ASIC, an FPGA, or any other device configured to perform image segmentation or execute an algorithm. In some embodiments, the processing step is performed inside the same physical device where the imager with the wide angle camera module is located. In other embodiments, the processing step is performed inside a different device where adaptive dewarping is required to improve the image. The segmentation processing step analyzes the original wide angle image content and segments its various elements in various segmentation layers, each segmentation layer including at least one of these elements. The segmentation can be performed depending on the element classification and optionally also depending on the depth analysis. The depth analysis step segments the various segmentation layers based on the distance of the various elements in the original wide angle image. The segmentation step can also be based on the position or shape of the various elements in the image.The depth of an element, especially the depth of an element in the foreground of an original wide-angle image, can be estimated using a depth estimation algorithm, which includes: an AI neural network that is trained to derive the relative depth of an element compared to other elements from a single image; an algorithm that analyzes the difference between consecutive frames of a video sequence when the camera is in autonomous or involuntary motion by combining gyroscope information from the device to reconstruct the 3D structure of the scene; or any other algorithm that is used to estimate, calculate or measure the depth of an element in the scene. When depth estimation is performed by a neural network, the network can have any shape, including but not limited to a neural network with a convolutional layer. The network can also be composed of sub-networks or sub-layers, each of which performs a separate task, including but not limited to convolution, pooling (maximum pooling, average pooling or other types of pooling), striding, padding, downsampling, upsampling, multi-feature fusion, rectified linear units, cascades, fully connected layers, flatten layers, etc. In other embodiments of the present invention, the depth of each object can be calculated based on a stereo image pair captured from different positions to calculate the difference due to parallax, and the depth of each object can be calculated from a time-of-flight hardware module, from a structured light system, from a lidar system, from a radar system, from a 3D capture, or by any other means for estimating, measuring or calculating the distance of an object visible in an image. In all the above examples of methods or systems for evaluating depth, the depth information obtained can be either absolute or relative. In the case of relative depth, the depth does not have to be accurate, and it may only be information that distinguishes the relative position of each layer. Without limiting the possible methods for ranking the depth of the layer, one such example of relative depth measurement is a relative depth measurement based on superposition. In image 700, due to the superposition, the head of person 703 partially hides tree 704, allowing the depth estimation algorithm to rank the relative depth of tree 704 as farther than person 703 even if the absolute distance is not available. In some embodiments of the present invention, both the segmentation algorithm based on image context and element classification and the depth analysis are performed together, and they help each other to improve the results of their analysis. In. Figure 7In the example of , the segmentation and depth analysis algorithm creates five different layers based on the depth and context of the object. The context analysis can be the result of a classification algorithm executed simultaneously with the segmentation algorithm. The classification algorithm is used to classify each segmented element into the identified category. In this example, the first layer 720 and the second layer 725 are for people. Each layer from the segmentation algorithm corresponds to a predefined distance range. For this reason, even if the two humans 702 and 703 from the original image 700 are the closest objects to the wide-angle camera, their distance from each other is greater than the predetermined minimum step size, and therefore they form two different layers 720 and 725. In layer 720, person 722 still looks stretched and visually unsatisfactory, just like person 703 in the original image 700. In layer 725, person 727 still looks correct, just like person 702 from the original image 700. In this example, the third layer 730 includes objects that are not recognized by the segmentation algorithm, or recognized objects for which no specific adaptive dewarping is required, such as tree 734. The fourth layer 740 is for buildings, where buildings 742 and 744 are still distorted like buildings 706 and 707 in the original image 700. Here, the two buildings 742 and 744 from layer 740 are considered to be at the same distance from the camera compared to the predetermined minimum distance step. Because they are from the same classification type and at the same depth, the segmentation and depth analysis algorithm 710 classifies them in the same layer 740. The tree 734 is also considered to be at the same distance as the two buildings 742 and 744, but because the segmentation algorithm finds that they are from two different categories, they are in different layers 730 and 740. Finally, the last layer 750 is the background, which consists of all objects such as the distant mountain 755 in the image, which will not be affected by the perspective correction. The method then processes at least one of the segmented layers to at least partially de-distort any of the one or more elements with distorted geometry, thereby creating a de-distorted layer. This is done by adaptive de-warping 760 that depends on the depth of the objects in the original wide-angle image and the image context. In a preferred embodiment, the specific dewarping process to be applied on the segmented layer depends either on the location of the segmented layer in the original wide-angle image or on the classification of the elements in the segmented layer. The context of the original wide-angle image depends on these and is usually determined automatically by analyzing the elements from the original wide-angle image 700. In other cases, the context can also be determined using information from the segmentation and depth of each layer obtained from the algorithm 710. Alternatively, the exact information and parameters of the adaptive dewarping to be applied can be transmitted to the adaptive dewarping algorithm via metadata, or tags in the image, user input, selection from a list of adaptive dewarping algorithms, or automatically selected according to the application.In this example, since the original image has a segmented layer with people, context-based custom dewarping 760 with body and face protection specifically for people is applied on layers 720 and 725 to obtain dewarped layers 765 and 770, respectively. This custom dewarping for people does not try to maintain perspective or straight lines of objects, but rather keeps the shapes of humans visually pleasing no matter where they are in the field of view. Next, custom dewarping for unknown or unidentified objects is applied to layer 730 to obtain dewarped layer 775. This custom dewarping improves the shape of objects towards the edges of the image based on the difference in magnification of one edge of the object to the other edge, as if they were imaged at the center of the image, but no specific correction is made for known objects (buildings, people) that require specific corrections. Next, adaptive dewarping is applied on layer 740 to obtain a dewarped view of a building 780. For buildings, keeping straight lines is important for the image to be visually pleasing, and therefore the projection applied on this layer keeps the lines straight. Finally, if it is necessary to obtain the desired projection, the background layer 750 can also be optionally de-warped to obtain a de-warped layer 785. The last step of the method according to the present invention is to merge at least one de-warped layer back with the other segmented layers via a merging algorithm to form a final image 790. In this example, the first layer of the final image is the background, and then all layers are superimposed in descending order of distance from the camera calculated by the depth analysis algorithm 710 to form a complete image 790 with adaptive de-warping. In some embodiments according to the present invention, the merging of at least one de-warped layer with the other layers is performed by adjusting the texture grid or the display grid. Alternatively, the merging can also be performed pixel by pixel. As can be seen from the example figure, the merged final image 790 has some dotted parts 792 on the tree, in which the camera did not initially capture information. Will be utilized. Figure 8 This missing information exists in this example, but in some other examples according to the invention, if the layer on top is increased in size by an adaptive dewarping algorithm, there may be an output image without any missing background information, such as using Fig. 9As explained. In addition, in some embodiments according to the present invention, at least one of the multiple layers after adaptive dewarping based on image context can be further processed before merging them together. An example is when the processing step further includes the following operation: adding some active blur on at least one of the depth layers so as to add a bokeh effect that depends on the context and depth, rather than a traditional bokeh effect based only on depth. For example, in the context of an image of a human face in front of a distant background, such a context-based bokeh effect can be automatically added to blur the background and keep the human face well focused. In other applications of the present invention, when the background is more important than the foreground, the opposite can also be done, where for a reverse bokeh effect, the background is sharply focused and the foreground object is blurred. In addition, in some embodiments according to the present invention, the multiple segmented layers after context-based adaptive dewarping can be further processed before merging them together, so as to purposefully add some translation, some rotation, or some scaling to at least one of the dewarped layers to create an active perspective or 3D effect. Furthermore, in some other embodiments according to the present invention, further processing of the multiple segmented layers before merging them together may further include: performing perspective tilt correction on at least one element in the segmented layers so as to correct the perspective with respect to the horizon or any target direction in the scene. This perspective tilt correction is particularly useful when the element is a building that was captured at an unsatisfactory-looking tilt angle in the original wide-angle image, so as to correct its shape so that it appears as if it was not captured at such a tilt angle, but this perspective tilt correction may also be applied to any kind of element. Furthermore, in some other embodiments according to the present invention, further processing of the multiple segmented layers before merging them together may further include: stabilizing at least one segmented layer so as to avoid unsatisfactory movement of one or more segmented layers between frames in the video sequence. Furthermore, in some embodiments of the method of the present invention, some unwanted object layers may be removed before merging these layers to create a final merged image. In some embodiments according to the present invention, when more than one object is touching or close to each other in the original wide-angle image, such as Figure 7 In the example of humans 703 and trees 704 in FIG. 1 , the specific dewarping process for the segmented layers depends on adjustable correction weights that can be added to the context-based adaptive dewarping. These correction weights can adapt the level of dewarping performed on these objects by increasing or decreasing the level of dewarping to ensure that less important objects do not interfere with more important objects nearby. Figure 7In the example of , the tree 704 may have a lower correction weight to avoid interfering with the dewarping layer of the human 703. These correction weights may be preset in the device running the adaptive dewarping, or manually adjusted or selected according to user preferences. These correction weights may also be automatically adjusted by an algorithm, including an algorithm trained via artificial intelligence methods, which automatically interprets the user's intent based on how the picture was captured. This adjustment of the correction weights is particularly useful when the user can see a preview of the dewarped image before capturing the image and he adjusts the camera parameters accordingly to obtain a better looking final dewarped image.
[0042] Figure 8 An alternative method for filling in missing image information during application of a context- and depth-layer-based adaptive dewarping method, sometimes referred to as inpainting, is shown. This inpainting technique is used to complete at least a portion of the missing information in the original wide-angle image. Figure 7 In the possible case where the final image of the explained method has missing information, this further step allows to improve the final image to make it more visually pleasing to a human observer. Image 800 is an example original wide angle image either from a lens with a rectilinear H=f*tan(θ) projection or from a wide angle lens in which the distortion has been de-warped to obtain a rectilinear projection. In this example image 800, the setting is indoors with five people in a group selfie (or group photo) setting. This image 800 is just an example, but this method for filling in missing information is not limited to any setting and can be used with any image on which an adaptive dewarping algorithm is used. Image 800 contains a line 802 from a wall, which remains straight because the image follows a rectilinear projection. Image 800 also has a background wall texture 804 that is partially hidden by people 805, 806, 807, 808 and 809. As in Figure 7 As in the case of the adaptive dewarping method of , the first step is the image segmentation and depth analysis step 810. For simplicity, in this example, the algorithm creates two layers: a layer 820 where all people are standing at relatively the same distance from the camera; and a layer 825 with the background. The segmentation and depth layer 820 has no missing information because the object is in the foreground, and the segmentation and depth layer 825 has missing information because it is in the background. As in Figure 7 As in the case of the adaptive dewarping method of , the next step is context-based adaptive dewarping 830 to dewarp the distorted geometry. The layer 820 has a person in it, and thus the adaptive dewarping with body and face protection is used to obtain the dewarped layer 840 to make the shape of the person visually attractive. After adaptive dewarping, if the layer is as Figure 7If the layers 840 and 850 are directly merged together as in the method of , we will obtain an image 860 with missing information. Compared to the original image 800, the person 867 in the center has not been moved or deformed by the adaptive dewarping process, and therefore there is no area of missing information behind it. However, the person 866 is moved to the left by the adaptive dewarping process, and he creates an area of missing information 862 in the background. Similarly, the person 868 is moved to the right by the adaptive dewarping process, and he creates an area of missing information in the background. Because the people 865 and 869 are closer to the camera, compared to their corresponding images 805 and 809, the people 865 and 869 are enlarged by the adaptive dewarping that corrects the perspective, and therefore no area of missing information is created behind them. Instead of merging the layers 840 and 850, an optional additional step 870 is to fill in the missing information using a completion algorithm or an image restoration algorithm. The completion algorithm can use various methods to complete the image, including but not limited to: applying blur based on the texture and color around the missing information area; applying gradient lines that gradually change the color from the color on one side of the missing information area to the color on the other side, usually in a direction that minimizes the length of these generated gradient lines; using an artificial intelligence network trained to complete the missing information of the picture, etc. According to the method of the present invention, the completion algorithm can be executed on a hardware processor consisting of a processing unit (CPU or GPU), which is either located in the same device as the camera or in a separate device that receives the output merged image 795. Alternatively, the completion algorithm can also be executed in parallel with the context-based adaptive dewarping step 830. The output of the completion algorithm 870 is: a completed layer 885, because the layer 850 has missing information; and an unmodified layer 875, because the layer 840 does not have missing information. Figure 8 The example of shows one completed layer 885, since only the background has missing information, but if there are many layers with missing information, many completed layers can be output. The layer 875 without missing information is then merged with the completed layer 885 in depth order from the farthest layer to the nearest layer. The result is a final image 890 with filled information, in which perspective correction is applied, the shape of the person is corrected to avoid unsatisfactory stretching, and the missing background information is filled using the completion algorithm compared to Figure 860, resulting in a visually pleasing filled background 892. In some embodiments of the present invention, the algorithm can also optionally use information from any previous frame or even from multiple previous frames of the video sequence to complete the missing information of the current frame. This is possible when the missing information of the current frame is visible in any previous frame before some movement in the scene creates the missing information area.
[0043] Fig. 9 Shows Figure 8 An alternative method to the method of is for hiding rather than filling missing image information during application of a context- and depth-layer-based adaptive dewarping method. In this example, at least a portion of the missing information in the original wide-angle image is hidden by scaling at least one dewarping layer. Figure 7 In the event that the output image of the interpreted method has missing information, the alternative method allows improving the output image to make it more visually pleasing to a human observer. Figure 8 Starting with image 900, which is the same as image 800, this image is also an example original wide angle image either from a lens with a rectilinear H=f*tan(θ) projection or from a wide angle lens in which the distortion has been de-warped to obtain a rectilinear projection. In this example image 900, the setting is indoors with five people in a group selfie (or group photo) setting. This image 900 is just an example, but this method for hiding missing information is not limited to any setting and can be used with any image on which an adaptive dewarping algorithm is used. Image 900 contains a line 902 from the wall, which remains straight because the image follows a rectilinear projection. Image 900 also has a background wall texture 904 that is partially hidden by people 905, 906, 907, 908, and 909. As shown in Figure 7 As in the case of the adaptive dewarping method of , the first step is the image segmentation and depth analysis step 910. For simplicity, in this example, the algorithm creates two layers: a layer 920 where all people are standing at relatively the same distance from the camera; and a layer 925 with the background. The segmentation and depth layer 920 has no missing information because the object is in the foreground, and the segmentation and depth layer 925 has missing information because it is in the background. As in Figure 7 As in the case of the adaptive dewarping method, the next step is context-based adaptive dewarping 930. The layer 920 has a person in it, and thus adaptive dewarping with body and face protection is used to obtain a dewarped layer 940 to make the shape of the person visually attractive. After adaptive dewarping, if the layer is Figure 7, we will obtain image 960 with missing information. Compared to the original image 900, person 967 in the center has not been moved or deformed by the adaptive dewarping process, and therefore there is no area of missing information behind him. However, person 966 is moved to the left by the adaptive dewarping process, and he creates an area of missing information 962 in the background. Similarly, person 968 is moved to the right by the adaptive dewarping process, and he creates an area of missing information in the background. Because persons 965 and 969 are closer to the camera, persons 965 and 969 are enlarged by the perspective-corrected adaptive dewarping compared to their corresponding images 905 and 909, and therefore no area of missing information is created behind them. Instead of merging layers 940 and 950 or as Figure 8 In order to fill the missing information as in the method of , an optional additional step 970 is to hide the area of missing information using an algorithm for adjusting relative magnification. The method adjusts the relative magnification of the objects in the previous layers and enlarges them, which is just enough so that there will be no area of missing information in the background when these layers are combined together. According to the method of the present invention, the algorithm for adjusting the relative magnification on some layers can be executed on a hardware processor consisting of a processing unit (CPU or GPU), which is either located in the same device as the camera or in a separate device that receives the output merged image 795. Alternatively, the hiding of missing information by the algorithm for adjusting relative magnification can also be performed in parallel with the context-based adaptive dewarping step 930. The output of the algorithm 970 for adjusting relative magnification is: an enlarged layer 975, because layer 940 has objects that are moved to create areas of missing information in the merged image 960; and an unmodified layer 985, because layer 950 is in the background and does not need to change the magnification. In this example, person 943 at the center is not adjusted by the adaptive dewarping algorithm 930, and therefore no zone of missing information is created. For this reason, person 978 remains unchanged in layer 975. For persons 942 and 944 to be moved to the left and right, respectively, by the adaptive dewarping algorithm 930, they need to be enlarged by the algorithm 970 to hide the zone of missing information created when they are moved, and therefore persons 977 and 979 in the resulting layer are enlarged. Furthermore, in this example, persons 941 and 945 are also enlarged to persons 976 and 980, respectively, even though there is no zone of missing information behind them. This particular case is to illustrate that the algorithm 970 for adjusting relative magnification can even adjust objects or layers that have no missing information near them in order to keep the overall proportions of the image being respected. Fig. 9The example shows one layer 975 with adjusted magnification, but if there are many layers with objects that need to be enlarged to hide the missing information area, many layers with adjusted magnification can be output. The layer 975 with adjusted magnification is then merged with the background layer 985 in depth order from the farthest layer to the nearest layer. The result is an image 990 with some foreground objects enlarged, in which perspective correction is applied, the shape of the person is corrected to avoid unsatisfactory stretching, and the missing background information is hidden by the algorithm that adjusts the relative magnification compared to Figure 960. In the resulting image 990, there is no missing background around the area 992 compared to the area 962 of missing information, which is visually pleasing.
[0044] In some other embodiments according to the present invention, Figure 8 How to complete and Fig. 9 The hiding methods can be used together by adjusting the relative magnifications to minimize their relative impact on the image.
[0045] Fig.10 1000 and 1050, as they are shown in the image with a dewarped projection of the background layer having a similar texture to that of the foreground. Figure 3 Due to the distortion profile, the background lines 1010 and 1060 have visible curved lines for both horizontal and vertical lines, as previously described in Figure 4As explained above. For the specific example of image 1000, since there are humans 1015, 1016, 1017, 1018, and 1019 inside the picture, the algorithm 1020 for context-based background dewarping will detect that these humans are in a group selfie or group photo scene, and the ideal background dewarping would be a cylindrical projection. The output of the dewarping algorithm 1020 is a background layer 1030, where the background lines 1040 show that the vertical lines in the objects are straight in the image, but the horizontal lines in the objects are curved in the image, as in a cylindrical projection. In the case when the background layer has some missing information because it is hidden by the foreground objects, an optional image restoration technique can be used to complete the image if required for the final output, as represented by the dotted line 1045. Next, in the example image 1050, a background 1060 that is the same as the previous background 1010 is visible, but this time there are no humans in front in the foreground. In this case, the algorithm 1070 for context-based background dewarping will detect that because it is an indoor scene, it is preferred to keep the straight lines in the scene as straight lines in this image, and the output projection should be straight. The output of the dewarping algorithm 1070 is a background layer 1080, in which the background lines 1090 are straight in the image, as in a rectilinear projection. The ideal outputs of the cylindrical and rectilinear projections in this figure are merely examples of background projections that may be ideal for a given context, but the method according to the present invention is not limited to any particular projection and can be any of a stereographic projection, an equidistant projection, an equisolid projection, an orthographic projection, a Mercator projection, or any other custom projection.
[0046] Fig.11 An example implementation of an algorithm according to a depth and segmentation layer based context-based adaptive dewarping method for ideal face preservation is shown. In this example implementation of the algorithm, processing at least one segmentation layer to create at least one dewarping layer includes: creating a virtual camera centered on an element having a distorted geometry; applying a rectilinear correction on the virtual camera; and translating the result to the correct position in the final image. The example algorithm starts with an original wide angle image 1100. The original wide angle image can have any field of view, but is typically Figure 2 and Figure 4The distorted geometry shown in is most apparent in wide angle images with a full field of view of more than 60°. In embodiments according to the invention, the original wide angle image 1100 is captured directly using a camera module with a wide angle lens, with or without deviations from rectilinear projection. In some other embodiments according to the invention, the original image 1100 has been processed by a processor to correct the original distortion from the camera module, improve image quality, or apply any other image processing. Alternatively, the original wide angle image 1100 may be combined from multiple narrow angle images within the processor, or entirely computer generated. In this example, without limiting the scope of the invention in any way, the original wide angle image has a background 1110 consisting of a mountain landscape, and a human face 1115 as an object in the foreground. Because the human face is close to the corner, the face is stretched and the original object proportions are not maintained. As Figure 4 As already explained in , the original wide angle image 1100 is visually unsatisfactory. A context-based adaptive dewarping method performs segmentation and classification of the human face 1115. The next step 1120 in the example algorithm is to create a virtual camera 1130 with a human face 1135 centered therein. The virtual camera 1130 has a narrow field of view and is rotated as if it were at the center of the original image, thereby fixing the stretching because the original proportions are maintained in the narrow field of view at the center of the image. This is represented by a circular head shape that represents the ideal proportions in this example. Mathematically, this example step 1120 of rotating the virtual camera is described as follows. Each point Pin in the segmentation layer of the original wide angle image is assigned a coordinate (x, y) in the original wide angle image.
[0047]
[0048] The center position of the face in the original wide-angle image is Pin 0 , which has coordinates (x 0 , y 0 ).
[0049]
[0050] Use function F to pin from the center position according to the optical distortion 0 To calculate the Euler angles .
[0051]
[0052] For each input point Pin, use the so-called P 相机 The conversion function is used to calculate the virtual camera projection position Pin in 3D space with coordinates (x', y', z') 3d .
[0053]
[0054] .
[0055] Next, in the example algorithm, the virtual camera is rotated by multiplying it by the rotation matrix M, which transforms each input point Pin in 3D space 3d The Euler angles on the y-axis are inverted to obtain the position Pout with coordinates (u', v', w') 3d .
[0056]
[0057] .
[0058] Then, using the inverse function P -1 显示器 The position Pout in 3D space 3d Transformed into a position Pout in 2D space with coordinates (u, v).
[0059]
[0060] .
[0061] The next step 1140 of the example algorithm is to translate the result from the virtual camera 1130 back to the original position in the image before the virtual camera was rotated, giving a frame 1150 in which the human face 1155 has ideal proportions, but may still be rotated or the wrong size to match the background perfectly. Mathematically, we use the translation vector T = (t x , t y ) to calculate the position Pout' of each point in the segmentation layer.
[0062]
[0063] The next step 1160 of the example algorithm is optional and includes any further processing by the adaptive dewarping algorithm of the virtual camera 1170 to improve the final projection of the human face 1175, including rotation, scaling, or any other transformation for ideal face preservation. Mathematically, we apply the optional rotation matrix R and scaling matrix S to Pout' in order to calculate the final position Pout'' of each point of the segmentation layer.
[0064]
[0065] The final step 1180 of the example algorithm is to merge the segmented human face layer 1195 back into the other layers, represented here by the background layer with mountains 1190. When the layers are merged back together, the texture mesh or display mesh can be adjusted as needed for the best fit between the merged layers. This method is shown as an example for face protection, but in accordance with the present invention, the method can be applied to any kind of object protection. Fig.11 The algorithm described in is only an example implementation and is not restrictive. Other algorithms can be used to achieve the same result while being consistent with the spirit of the present invention.
[0066] Fig.12An example embodiment of a physical device 1230 is shown that captures a raw wide-angle image, processes it according to the method of the present invention to enhance the image based on the image context, and outputs the final image on a display screen. An object scene 1200 is visible to the physical device 1230, which means that some light rays from the object scene (here shown by two extreme light rays 1210 and 1212 that define the field of view of the imager) are reaching the imager 1220 of the physical device 1230. In this example, the imager 1220 is a wide-angle lens with a field of view that is typically greater than 60°, and forms an optical image in an image plane using an image sensor that is typically located at the image plane of the wide-angle lens, and transforms the light rays from the optical image into a digital image file representing a raw wide-angle image 1240. The raw wide-angle image file has multiple elements in the foreground or background, at least one of which has a distorted geometry that is visible, for example, through the stretched face of a person 1245. In other embodiments, the imager may be composed of any other means of creating a digital image, including other optical systems with lenses, mirrors, diffractive elements, metasurfaces, etc., or any processor that creates or generates a digital image file from any source. This embodiment is merely an example of such a physical device according to the present invention, and the example does not limit the scope of the present invention. The physical device may be any device that includes a means for receiving the original wide-angle image 1240, processing it, and displaying it, such as a smart phone, a tablet, a laptop or desktop personal computer, a portable camera, etc. In this example, the physical device 1230 also includes a processor 1250, which is capable of executing an algorithm for processing the original wide-angle image 1240 into a final image, including segmentation, classification, at least partial dewarping, merging layers, other various image quality processing and enhancement, etc. In this example, the processor 1250 is a central processing unit (CPU), but in other embodiments, the processing can be done by any kind of processor, including a CPU, a GPU, a TPU, an ASIC, an FPGA, or any other hardware processor configured to execute a software algorithm for implementing the functions described or capable of processing a digital image file. Processor 1250 then outputs a final image 1270 to display 1260. Final image 1270 has de-warped geometry, which is visible, for example, by the correct proportions of the face of person 1275. In this example, display 1260 is part of physical device 1230, similar to the screen of a smartphone or the like, but in other embodiments, the final image file may instead be transmitted to any other device for display or analysis by another algorithm.This example of a single physical device 1230 including an imager 1220, a processor 1250, and a display 1260 is merely an example embodiment according to the present invention, but these three features may also be part of multiple physical devices, with digital image files exchanged between them via any communication link to share digital image files, including but not limited to: a computer host bus, a hard drive, a solid state drive, a USB drive, transferring over the air via Wi-Fi, or any other means of transferring digital image files between multiple physical devices.
[0067] In some embodiments according to the present invention, an adaptive dewarping method based on segmentation and depth layers is used to maximize the straightness of lines in an image compared to the original lines in the object scene, maximize the full field of view of the output image compared to the full field of view of the original image, and / or maximize the conservation of proportions in the output image compared to the true proportions in the object scene. When the method is used to maximize the straightness of lines in an image compared to the original image file, the context-based adaptive dewarping method 760 transforms various segmentation and depth layers while giving priority to making the straight lines in the object scene as straight as possible in the merged image 790. When the method is used to maximize the full field of view of the output image compared to the full field of view of the original image, a special dewarping target is used for the corners of the original image to ensure that the diagonal full field of view of the output merged image remains as close as possible to the diagonal full field of view of the original image. This is done to avoid losing information in the image by reducing the field of view that forces the corners to be cropped, or avoiding the creation of black corners without information or black sides of the output image without information. In order to keep the field of view as close as possible between the original image and the output image, the special dewarping in the corner can ignore the segmentation or depth layer and not apply specific dewarping based on the context of the area in the image corner. According to the present invention, it is also possible to choose not to apply specific adaptive dewarping based on the context and depth layer in the area or another area of the image, or not to apply specific adaptive dewarping on a specific layer for any other reason. When the method is used to maximize the conservation of scale in the output image compared to the real scale in the object scene, the context-based adaptive dewarping method 760 transforms various segmentations and depth layers while giving priority to scale. In this case, the scales in the output merged image 790 all look similar to the scales in the real object scene that would be needed when humans are visible in the image. In some embodiments according to the present invention, all three cases are maximized together.
[0068] In all embodiments according to the invention, the adaptive dewarping algorithm may optionally use information from any previous frame in the video sequence for temporal filtering. With temporal filtering, the final output from the adaptive dewarping may be smoother by removing potential artifacts that may be produced by poor interpretation of the algorithm on a particular frame by favoring temporal consistency over results that have large deviations from previous frames. Temporal filtering is also useful in situations where some jitter of the camera or part of the object scene would otherwise produce artifacts.
[0069] All of the above are diagrams and examples showing the adaptive dewarping method. In all of these examples, the imager, camera or lens can have any field of view, from a very narrow angle to an extremely wide angle. These examples are not intended to be an exhaustive list or to limit the scope and spirit of the invention. Those skilled in the art will appreciate that the embodiments described above may be changed without departing from the broad inventive concept thereof. Therefore, it is to be understood that the invention is not limited to the specific embodiments disclosed, but is intended to cover modifications within the spirit and scope of the invention as defined by the appended claims.
Claims
1. A method for enhancing wide-angle images based on image context and segmentation layers, the method include: a. receiving, by a processor, an original wide-angle image, the original wide-angle image being created by an imager and having a plurality of elements, each element being in one of a foreground or a background of the original wide-angle image, one or more of the elements having a distorted geometric shape; b. segmenting, by the processor, the original wide-angle image into a plurality of segmented layers, each of the segmented layers comprising at least one of the elements, the segmentation being based on at least one of the following: a shape of one or more of the elements, a position of one or more of the elements in the original wide-angle image, or a depth of one or more of the elements compared to other elements; c. processing, by the processor, at least one of the segmented layers to at least partially de-distort any of the one or more elements having the distorted geometry in the at least one segmented layer, thereby creating at least one de-distorted layer; as well as d. The processor combines the at least one dewarped layer with other segmented layers to form a final image. 2 . The method of claim 1 , wherein the original wide-angle image is captured by the imager, which is an optical system including at least a camera and a wide-angle lens.
3. The method of claim 2, wherein the wide-angle lens has a diagonal field of view of at least more than 60°.
4. The method of claim 1 , wherein at least one of the elements is in the foreground of the original wide-angle image, and segmentation of at least one foreground element is based on a relative depth compared to other elements, the relative depth being calculated by an artificial intelligence neural network.
5. The method of claim 1, wherein a specific dewarping process for the at least one segmented layer depends either on a position of the at least one segmented layer in the original wide-angle image or on a classification of at least one of the elements in the at least one segmented layer. 6 . The method of claim 1 , wherein the distorted geometric shape comprises at least one of: a stretched proportion, a curved line, or an image with optical distortion.
7. The method of claim 1 , wherein the specific dewarping process for the at least one segmented layer depends on adjustable correction weights, wherein each of the elements to be dewarped is assigned an adjustable correction weight to adapt the level of dewarping for the element.
8. The method of claim 7, wherein the adjustable correction weights are either selected according to user preference or are preset.
9. The method of claim 1 , wherein the processing step further comprises at least one of blurring the at least one dewarping layer, rotating the at least one dewarping layer, translating the at least one dewarping layer, scaling the at least one dewarping layer, correcting a perspective tilt of the at least one dewarping layer, or stabilizing the at least one dewarping layer.
10. The method of claim 1, wherein an image restoration technique is used to complete at least a portion of the missing information in the original wide-angle image.
11. The method of claim 1, wherein at least a portion of missing information in the original wide-angle image is hidden by scaling the at least one dewarping layer.
12. The method of claim 1, wherein processing of the at least one segmented layer uses a dewarped projection for a background layer, the dewarped projection being dependent on a detected context of the original wide-angle image.
13. The method of claim 1, wherein the at least one segmented layer is processed to create the at least one dewarped layer include: creating a virtual camera centered on the element having the distorted geometry; applying a rectilinear correction on the virtual camera; and translating the result to the correct position in the final image.
14. The method of claim 1, wherein the merging of the at least one dewarped layer with other layers is performed by adjusting a texture or a display mesh.
15. A method for enhancing a wide-angle image based on image context, the method include: a. receiving, by a processor, an original wide-angle image, the original wide-angle image being created by an imager and having at least one element, the at least one element having a distorted geometry; b. creating, by the processor, at least one classified element by classifying the at least one element having the distorted geometric shape from the original wide-angle image, the classification being based on at least one of: a shape of the at least one element, a location of the at least one element in the original wide-angle image, or a depth of the at least one element compared to other elements in the original wide-angle image; as well as c. creating a final image by the processor by processing the original wide-angle image to balance between the following three items: i) maximizing the straightness of lines in the final image; ii) maximizing the full field of view of the final image; and iii) maximizing conservation of scale in the final image.
16. The method of claim 15, wherein the processing of the original wide-angle image is performed according to an adjustable correction weight, wherein each of the three items is assigned the adjustable correction weight to adjust the level of processing of the original wide-angle image.
17. The method of claim 16, wherein the adjustable correction weight is adjusted either according to application or according to user preference.
18. The method of claim 15, wherein the processing of the original wide-angle image is performed using an artificial intelligence algorithm.
19. The method according to claim 15, wherein continuous lines in the object scene of the original wide-angle image are discontinuous in the final image.
20. The method of claim 15, wherein processing the original wide angle image comprises transforming a texture or a display mesh.
21. A device for enhancing a wide-angle image based on image context and segmentation layers, the device include: a. An imager that creates an original wide-angle image having a plurality of elements, each element being in one of a foreground or a background of the original wide-angle image, one or more of the elements having a distorted geometric shape; as well as b. a processor configured to: i. segmenting the original wide-angle image into a plurality of segmentation layers, each of the segmentation layers comprising at least one of the elements, the segmentation being based on at least one of the following: a shape of one or more of the elements, a position of one or more of the elements in the original wide-angle image, or a depth of one or more of the elements compared to other elements, ii. processing at least one of said segmented layers to at least partially de-distort any of said one or more elements having said distorted geometry in said at least one segmented layer, thereby creating at least one de-distorted layer, and iii. merging the at least one dewarped layer with other segmented layers to form a final image.
Citation Information
Patent Citations
Image distortion transformation method and apparatus
US10204398B2
Panoramic camera
US10356316B2
Omnidirectional image processing device and omnidirectional image processing method
CN102395994A
Method of detecting and correcting digital images of books in the book spine area
CN102790841A