Inference or training of an artificial neural network using an imager with deliberately controlled distortion

By generating images using an imager with intentionally controlled distortion and inputting these images into specially trained convolutional neural networks, the problem of existing convolutional neural networks being poorly effective when processing controlled distortion images is achieved, achieving higher image analysis and processing resolution.

CN114787828BActive Publication Date: 2025-06-27IMMERVISION INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202080079552.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-18
Filing Date
2020-11-17
Publication Date
2025-06-27
Estimated Expiration
2040-11-17

AI Technical Summary

Technical Problem

Existing convolutional neural networks are poor in processing images with controlled distortions and cannot effectively utilize high-resolution images, resulting in limited processing capabilities in applications requiring global image analysis.

Method used

By using an imager with intentionally controlled distortions to generate images and input these images into specially trained convolutional neural networks, the neural network learns to process controlled distortion images, thereby improving the processing effect.

Benefits of technology

By increasing the number of pixels in the region of interest, the accuracy and effectiveness of neural networks when processing controlled distorted images are improved, and the resolution of image analysis and processing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787828B_ABST
    Figure CN114787828B_ABST
Patent Text Reader

Abstract

A method for training and using a convolutional neural network with deliberately distorted images is disclosed. By deliberately distorting the images to create regions of interest with a higher number of pixels, the result output of the neural network is improved. An imager device is used to create the distorted images, which includes an optical system specifically designed to output distorted images or includes software or hardware image distortion processing algorithms to create distorted images from normal images. A method for training a neural network using a distorted image generator from various existing data sets is also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 936,647, filed on November 18, 2019, entitled "Using imager with on-purpose controlled distortion for inference or training of an artificial intelligence neural network", which is currently pending, the entire content of which is incorporated herein by reference. BACKGROUND OF THE INVENTION

[0003] Embodiments of the present invention relate to the field of artificial intelligence convolutional neural networks and their use, and more particularly to how to use an imager with controlled distortion to correctly use these neural networks.

[0004] Due to the increasing processing power of personal computers, mobile devices, or large server clusters of large companies, the use of artificial intelligence to process or analyze digital image files has become increasingly popular. The rise in the use of artificial intelligence can also be explained by the new functions it may bring in a wide range of applications.

[0005] When analyzing digital image files, the most common type of neural network is the convolutional neural network, which means that certain convolutional operations are performed in certain layers of the network. In the past, the idea of using neural networks (NNs) to process digital image files for general applications has been proposed, including the use of convolutional neural networks (CNNs) in U.S. Patent Nos. 9,715,642, 9,754,351, or 10,360,494. The use of convolutional neural networks for certain specific applications has also been widely proposed in the past, including object recognition in U.S. Patent Application Publication No. 2018 / 0032844, face recognition in U.S. Patent No. 10,268,947, depth estimation in U.S. Patent No. 10,353,271, age and gender estimation in U.S. Patent Application Publication No. 2018 / 0150684, and so on.

[0006] However, existing convolutional neural networks for processing images are severely limited in terms of the input image resolution. Especially for applications that require global image analysis, which cannot be sequentially applied to smaller sub - parts of the image, such as depth estimation of a single image. Using a modern computer with a GPU having approximately 10GB of RAM memory, these neural networks are currently limited to analyzing and processing images with a resolution of approximately 512x512 (about 0.25MPx), which is significantly lower than the approximately 20 - 50MPx images available in modern mobile devices or cameras. As a result of the resolution limitation of digital image files that can be effectively processed for certain applications compared to the case where the full resolution of the input image is used, the processing or analysis of the neural network is poor. This limitation is even more critical in embedded system applications where the processing power is highly restricted.

[0007] One way to increase the number of pixels on an object of interest without increasing the total number of pixels in the image is to use deliberately controlled distortion. The idea of deliberately modifying the image resolution at the optical system, hardware, or software level has been proposed in the past, such as in U.S. Patent Nos. 6,844,990, 6,865,028, 9,829,700, or 10,204,398. However, in existing convolutional neural networks, these distorted images from the imager cannot be well - analyzed or processed, and new types of networks or training methods are needed to use images with deliberately controlled distortion. Another way to have a high - resolution input in a neural network is to crop a sub - region of the complete image and only analyze that sub - region inside the neural network. However, cropping a sub - region or region of interest of the complete image loses the complete scene information and continuity, which is important in applications where the neural network requires global information from the complete scene to provide the best output.

[0008] One type of digital image file that typically has controlled distortion is a wide - angle image, whose total field of view is usually greater than about 80°. However, compared to narrow - angle images without controlled distortion, such wide - angle images with associated ground truth data are rare. Most existing large - scale image datasets for training existing neural networks are based on narrow - angle images without distortion, so a new training method is needed to use wide - angle images or narrow - angle images with deliberately controlled distortion to train neural networks. Summary of the Invention

[0009] To overcome all the problems mentioned above, embodiments of the present invention propose a method for training and using a convolutional neural network using images with deliberately targeted distortion.

[0010] In a preferred embodiment according to the present invention, the method starts with creating a digital image file with controlled distortion from an imager. The imager can be any device that creates a distorted image, including a virtual image generator, image distortion transformation software or hardware, or a device with an optical system that directly captures an image with controlled distortion using an image sensor in the focal plane of the optical system. The imager can output an image with a unique static distortion profile or a dynamic distortion profile that can vary over time. For this preferred embodiment, the image with controlled distortion output from the imager has at least one region of interest where the resolution (calculated as pixels per degree of the object field of view) is at least 10% higher than a normal digital image file without controlled distortion. Then the image with controlled distortion is input into any type of neural network. The neural network generally includes at least one layer of convolutional operations, but this is not always required according to the present invention. The neural network can run on any physical device capable of executing algorithms. When the neural network has been specifically trained using images with controlled distortion, it can process the input distorted image. Inputting the distorted image into a neural network specifically trained with a set of distorted images results in a more precise output of interpretation data in the region of interest with an increased number of pixels, which also helps to obtain improved results in other parts of the image outside the region of interest. This improved result of the interpretation data can be anything, depending on the application of the neural network, including image depth information, object recognition, object classification, object segmentation, optical flow estimation, connecting edges and lines, Simultaneous Localization and Mapping (SLAM), super-resolution image construction, etc. In some embodiments of the present invention, the interpretation data from the output of the neural network can still be an image with controlled distortion. In that case, depending on whether the image is used by a human observer, optional steps of image distortion correction and dewarping can obtain a final output image without distortion. If the output from the neural network is to be directly used by another algorithm unit, computer, or any other automated process, this optional step is generally not required.

[0011] To use a convolutional neural network with input digital image files having deliberately controlled distortion, the neural network must be specifically trained for these image files. The method according to the present invention includes a distorted image dataset generator from a large existing image dataset without controlled distortion. Since the existing image datasets include various objects captured using ordinary lenses without deliberate distortion, they cannot be directly used to train the network we proposed. The distorted image dataset generator processes the original images from the existing datasets to add any type of deliberate distortion, including radially symmetric distortion, free-form distortion centered or not centered on a specific object, or stretching distortion in the corners of the image. Then, the resulting distorted image dataset can optionally be extended by using data augmentation techniques or operations such as rotation, translation, scaling, homothety, and mirroring to increase the number of cases for training the neural network. The dataset can also be extended by using projection techniques such as planar projection, line projection, perspective skew correction projection, or any type of projection technique. Then, the neural network is trained using the new dataset generated with images having controlled distortion to learn to use these images with controlled distortion by any type of supervised or unsupervised learning technique.

[0012] In some alternative embodiments according to the present invention, the original image from the imager (with or without distortion) is first transformed into a well-defined normalized view having deliberately controlled normalized distortion in order to use a neural network specifically trained with this normalized distortion distribution, thus avoiding long re-training of the neural network for each new distortion distribution. The normalized view may or may not have some regions with missing texture information, depending on how the original image was captured and the requirements of the normalized view.

[0013] In some alternative embodiments according to the present invention, the original image from the imager is first processed to remove or minimize image distortion in order to use the processed image with an existing neural network that has been trained with images without controlled distortion, thus avoiding training a new neural network for the specific distortion distribution generated by the imager. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above-described invention content, as well as the following detailed description of the preferred embodiments of the present invention, will be better understood when read in conjunction with the accompanying drawings. For purposes of illustration, the presently preferred embodiments are shown in the drawings. However, it should be understood that the present invention is not limited to the precise arrangements and instrumentalities shown.

[0015] In the drawings:

[0016] FIG. 1 shows the inference processing of normal images using a neural network according to the prior art;

[0017] Figure 2 Disclosed is a method for performing inference processing on an image with controlled distortion using a neural network to improve the output of the network;

[0018] Figure 3 Disclosed is how to train an artificial neural network through deep learning to improve its performance on images with controlled distortion;

[0019] Figure 4 Disclosed is how to create a distorted data set from an original data set using a software or hardware image transformation algorithm;

[0020] Figure 5 Compares the performance of an inference neural network trained without images with controlled distortion and a neural network trained with images with controlled distortion in processing images with distortion;

[0021] Figure 6 Disclosed is an example of the change over time of controlled distortion in an image output from an imager before inference processing inside a neural network.

[0022] Figure 7 Disclosed is an example of transforming distortion into a standardized distortion distribution before inputting an image into a neural network for inference processing; and

[0023] Figure 8 Disclosed is an example of performing distortion correction on distortion before inputting an image into a neural network for inference processing. DETAILED DESCRIPTION

[0024] The words "a" and "an" as used in the corresponding parts of the claims and the specification mean "at least one".

[0025] Figure 1 illustrates the inference processing of a normal image using an artificial intelligence neural network according to the prior art. An artificial intelligence neural network 100 performs image processing on a normal image 110 to output a result at 140. The neural network can be of any type. In some embodiments, the neural network can be a convolutional neural network (CNN) trained via deep machine learning techniques or the like, but this is not always the case according to the present invention, and some other neural networks with or without any image convolution can also be used. In some embodiments, the network can be a generative adversarial network (GAN). The input normal image 110 is input into the network for inference processing via the input nodes of the input layer 120. The exact number of nodes depends on the application, and the figure with three input nodes is just an example network and in no way limits the types of networks that can be used to process input digital images. The network can also consist of an unknown number of hidden layers, such as layer 125 and layer 130 in this example figure, with each layer having an arbitrary number of nodes. It can also consist of several sub-networks or several sub-layers, each of which performs a separate task, including but not limited to convolution, pooling (max pooling, average pooling, or other types of pooling), striding, padding, downsampling, upsampling, multi-feature fusion, rectified linear unit, concatenate, fully connected layer, flattened layer, etc. The network can also include a final output layer 135, which includes an arbitrary number of output nodes. In this example figure, the dashed lines around the nodes indicate that the neural network is not trained using images with controlled distortion. The output interpretation data 140 of the network is the result of the original input digital image and can be of various types, including but not limited to image depth information, object recognition, object classification, object segmentation, optical flow estimation, connection edges, simultaneous localization and mapping (SLAM), super-resolution image creation, etc. Since the input digital image 110 does not have controlled distortion to create a region of interest in the image and there is no part where the number of pixels increases in the image, the result of the neural network conforms to the prior art. Specifically, for the example of Figure 1, the application shown is generating a depth map from the input image. The resulting depth map has low resolution anywhere in the image, including Figure 2 the car that will be the object of interest in the example of

[0026] Figure 2A method for performing inference processing on an image with controlled distortion using an artificial intelligence neural network according to the present invention to improve the output of the neural network is shown. The method starts with creating an image with controlled distortion from an imager 205. The imager 205 can be any type of device that creates a digital image file with controlled distortion to increase the number of pixels in the region of interest, including (but in no way limiting the scope of the present invention) a virtual image generator, a software or hardware image distortion transformation algorithm, or a device that intentionally changes the distortion of a digital image file. The device that intentionally changes the distortion of a digital image file can be of any type, including but not limited to a computer including a central processing unit (CPU), some memory units, and some means of receiving and transmitting digital image files, such as a personal computer (PC), a smartphone, a tablet, an embedded system, or any other device capable of transforming the distortion of a digital image file. The imager can also be mainly a hardware algorithm or executed on an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc. The imager can also be a device including at least one camera system, the camera system including at least one or more optical systems that form an image with controlled distortion, etc. Here, the optical system can be made of any combination of refractive optical elements, reflective optical elements, diffractive optical elements, metamaterial optical elements, or any other optical elements. The optical system can also be an optical system including source optical elements (such as a deformable mirror, a liquid lens, a spatial light modulator, etc.) to change and adjust the resulting distortion distribution of the optical system in real time. The optical system can use any number of aspherical or freeform optical elements to create at least one resolution-enhanced region to better control the distortion. In some embodiments according to the present invention, the optical system is preferably a wide-angle lens with a diagonal field of view greater than 60°, where the wide-angle lens includes a plurality of optical elements sequentially divided into a front group of elements, an aperture stop, and a rear group of elements, and the wide-angle lens forms an image on the imaging plane.

[0027] The output of the imager device 205 is an image 210 with intentionally controlled distortion. In Figure 2In this example, for simplicity, only one image file is shown, but the method according to the present invention can also be compatible with multiple image files, which are synthesized or not synthesized into a digital video file. The digital image file 210 has a controlled distortion defined as at least one region of interest, wherein the resolution (or magnification) calculated in pixels per degree of the object field of view is at least 10% higher than that of a normal digital image file 110. In some other embodiments according to the present invention, the controlled distortion is defined as having a region of interest with at least 20%, 30%, 40% or 50% more pixels per degree than in an image without distortion. By creating the at least one region of interest, the imager can maintain the same total field of view as the image without this region of interest or can change the total field of view.

[0028] Then, the digital image file 210 with deliberately controlled distortion is input into the artificial intelligence neural network 200. The neural network 200 can be of any type, including machine learning neural networks trained via deep learning techniques, including but not limited to convolutional neural networks (CNNs), etc. The neural network 200 includes algorithms, software code, etc. running on a physical computing device to interpret any type of input data, and is trained to process images with controlled distortion. The physical computing device can be any hardware capable of running such algorithms, including but not limited to personal computers, mobile phones, tablets, cars, robots, embedded systems, etc. The physical computing device can include any of the following: an electronic motherboard (or main board), at least one processor, a partial central processing unit (CPU) or no central processing unit, some memory (RAM, ROM, etc.), drives (hard disk drives, SSD drives, etc.), a graphics processing unit (GPU), a tensor processing unit (TPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other component that allows the neural network to run and transform the input digital image file into an output interpreted data result.

[0029] In Figure 2 this embodiment, the artificial intelligence neural network 200 has been specifically trained with distorted images having controlled distortion in order to better process them, as will be referred to Figure 3As explained. The input digital image file 210 with controlled distortion is received by the network 200 via the input nodes of the input layer 220. The number of nodes depends on the application, and the figure with three input nodes is just an example network and in no way limits the types of networks that can be used to process the input digital image. The network can also consist of an unknown number of hidden layers, such as layer 225 and layer 230 in this example figure, with each layer having an arbitrary number of nodes. It can also consist of several sub-networks or several sub-layers, each performing a separate task, including but not limited to convolution, pooling (max pooling, average pooling or other types of pooling), striding, padding, downsampling, upsampling, multi-feature fusion, rectified linear unit, concatenation, fully connected layer, flattened layer, etc. The network can also include an output layer 235, which includes an arbitrary number of output nodes. In this example figure, the solid lines around the nodes indicate that a neural network trained with an image with controlled distortion outputs improved results, and the arrows in this neural network are from left to right or from the input to the output layer of this neural network, indicating the inference processing of the network. Then, the network continues to perform inference processing on the input digital image file to output the interpreted data. The output interpreted data 240 of the network is derived from the input digital image file 210 with controlled distortion and can be of various types, including but not limited to image depth information, object recognition, object classification, object segmentation, optical flow estimation, connection edges, simultaneous localization and mapping (SLAM), super-resolution image creation, etc.

[0030] Because the input digital image file 210 has controlled distortion to create a region of interest in the image, which at least partially has an increased number of pixels, and thus the result of the output interpretation data from the artificial intelligence neural network is improved compared to the result (such as the prior art output 140) from an input digital image file without controlled distortion. For example, when the application of the artificial intelligence algorithm is to estimate a depth map from a single image shown in the figure, this improvement can be a more accurate depth map with a higher resolution. Compared to a network using normal images without controlled distortion in the prior art, due to the larger number of pixels of the object of interest, better object classification or recognition may be obtained, or any other improved result from the neural network. The improvement on at least a single image can be measured in various ways depending on whether the output of the neural network is qualitative or quantitative, including but not limited to the reduction of the relative (calculated in %) or absolute (calculated in units suitable for the application of the network) difference between the output and the ground truth, the root mean square (RMS) error, the average relative error, the average log10 error, the threshold accuracy, etc. The improvement can also be calculated based on the scores of true positives, false negatives, true negatives, and false positives in the output as precision P-score, recall R-score, F-score, etc. The improvement can also be measured as an increase in the probability output or confidence output from the neural network, especially when the output is qualitative, such as in a classification neural network. In some embodiments, the improvement between the original image with controlled distortion and the original image without controlled distortion is also measured as a percentage increase in accuracy by using a large dataset of input digital image files with controlled distortion and comparing its results with the results of a similar large dataset of input digital images without controlled distortion.

[0031] In Figure 2 the example, the output of the neural network is a digital image file, but this is not always the case, and the output can also be a text output, an optical signal, a tactile feedback, or any other output generated by inputting an image with controlled distortion into the neural network. In the case where the output is a digital image file 240, if the output digital image is to be displayed to a human observer, the image can also optionally be further processed by image distortion correction to at least partially remove the controlled distortion, thereby obtaining a digital image file 250 with less or no controlled distortion. This optional additional distortion correction step is performed using a software algorithm running on a computer composed of a processor, etc., or directly on a hardware device configured to process the output digital image file 240 to remove, correct, modify, or process the distortion.

[0032] This optional step may not be required if the output image will be used by software or hardware algorithms or any other computer that uses the image without human intervention. In some embodiments of the present invention, the full neural network 200 is composed of a number of sub-networks that are configured to analyze the global image and local sub-parts of the image and combine the results. For the global image, the sub-network may include a number of downsampling layers followed by a number of upsampling layers to restore the original image resolution, and these layers may or may not use convolution. For the local sub-parts of the image, the sub-network may directly process, for example (without limiting the scope of the present invention in any way), several cropped parts of the original image, or use the intermediate layers from the downsampling or upsampling sub-networks applied to the global image as input. Then, the results from the global image sub-network and the local image sub-network can be combined using an average layer, concatenation, and convolutional layers, etc. to produce the final output of the entire network.

[0033] Figure 3 A method for training an artificial neural network by deep learning to improve its performance on images with controlled distortion is shown. In Figure 3 this example, for simplicity, only image files are shown, but the method according to the present invention will also be compatible with digital video files. The method of training a neural network by supervised learning, semi-supervised learning, or unsupervised learning starts with a large database of images 310 that have no intentionally added controlled distortion. These image databases are also often referred to as data sets. In Figure 3 this example, the original image 310 without controlled distortion is an original image of a cat with normal proportions. To be able to use these large existing data sets to train a neural network with supervised, semi-supervised, or unsupervised learning, the method according to the present invention processes the original images from the data set into a software or hardware image transformation algorithm 320 that transforms the original image into a target controlled distortion distribution in a manner similar to the digital file 210 output from the imager 205 in Figure 2 . The software or hardware image transformation algorithm 320 is executed on an image transformation device and will be in Figure 4This is further explained in. In some embodiments according to the present invention, in addition to processing the image itself, their corresponding desired result images (commonly referred to as ground truth images) can also be processed in the same way to add deliberately controlled distortions. The controlled distortion target can be of any type, including but not limited to: radial barrel distortion with rotational symmetry often found in wide-angle images such as Example 330, free-form distortion with or without rotational symmetry and centered or not centered on a specific object such as Example 340, stretching distortion or pincushion distortion visible only in the corners of the image or any other part of the image such as Example 350, stretching distortion or pincushion distortion throughout the image such as Example 360, and any other type of distortion that produces at least one region of interest and has at least 10% more pixels per degree in that region than a perfect image. Here, a perfect image can be an image with uniform pixel density and scale following a straight projection, or any other ideal image for a given neural network. In some other embodiments according to the present invention, controlled distortion is defined as having a region of interest with at least 20%, 30%, 40%, or 50% more pixels per degree than an image without distortion.

[0034] Any new image generated with controlled distortion can have the same field of view as the original image without controlled distortion or a different field of view. When the field of view of the generated new image is larger than the field of view of the original image, the remaining part of the image can be filled with anything, including computer-generated background images, backgrounds extracted from another image, multiple copies of the original image, multiple images from the original dataset, image extrapolation, blank, or any other type of image completion to fill the missing part of the field of view as needed.

[0035] Then a new dataset generated using images with controlled distortion (such as 330, 340, 350, and / or 360) is used to train the neural network 370 to learn to use these images with controlled distortion. In Figure 3In this example, the arrows in the illustrated neural network are from right to left or from the output of the neural network to the input layer, representing the training of the network through backpropagation, rather than the inference process from input to output represented by arrows from left to right in other diagrams. This learning can be supervised learning, where the input image and the desired output ground truth result of the neural network form a known pair. It can also be used in unsupervised learning, where the input image is associated with an unknown ground truth output result from the network. The new dataset can also be used for any other type of learning or reinforcement of deep learning neural networks, including a hybrid mode between supervised and unsupervised called semi-supervised, or any other way of training artificial intelligence using an image dataset. When training the network, any optimization technique can be used to optimize the weights between each node of each layer, including (but in no way limiting the scope of the present invention) gradient descent, backpropagation, genetic algorithms, simulated annealing, randomized optimization algorithms, etc. The loss function (also known as the cost function or energy function) used during neural network optimization can be of any type according to the present invention, depending on the desired application of the neural network. In some embodiments of the present invention, when training a neural network to analyze or process wide-angle images that typically have a total field of view greater than about 60°, since existing wide-angle datasets are very rare and usually do not have ground truth results for the desired application, wide-angle images generated using a virtual 3D environment can be used to train the neural network. In some cases, when there is a small wide-angle dataset but a larger one is needed for precise training, a combination of existing real wide-angle images and virtual generated wide-angle images is used.

[0036] Figure 4 Illustrates how to create a distorted dataset from an original dataset using a software or hardware image transformation algorithm running on an image transformation device. The method starts with an original image dataset 410. There are some publicly available datasets on the Internet, including images of real natural objects, images of artificial or virtual objects, or hybrid datasets of real and virtual objects. The objects in these image or video datasets can be of various types, including text, faces, animals, buildings, street scenes, etc., to help train various types of artificial intelligence neural networks. These existing datasets are captured from ordinary lenses without intentional distortion or generated using normal views without controlled distortion. Then, in step 420, the method selects an image from the dataset to adapt it. Figure 4The example method shown transforms only one image from the original dataset, but in the practical situation of generating a new dataset, the same method can be continuously applied to multiple desired original images. Additionally, the method according to the present invention is also compatible with creating a dataset from multiple image files, whether or not the multiple image files are synthesized into a digital video file. In some embodiments according to the present invention, in addition to processing the original image 410, their corresponding desired result images (commonly referred to as ground truth images) are also processed in the same way to add deliberately controlled distortions to both the original images and the ground truth images.

[0037] The next step of the method is to select the target of the desired deliberately controlled distortion and the image field of view in step 430. The target of the deliberately added controlled distortion depends on the specific application required by the neural network trained using the new dataset and can be of any type, including but not limited to radial barrel distortion with rotational symmetry commonly found in wide-angle images, free-form distortion with or without rotational symmetry and centered or not centered on a specific object, stretching or pincushion distortion visible only in the corners of the image or any other part of the image, stretching or pincushion distortion of the entire image, or any other type of distortion that produces at least one region of interest, where the number of pixels per degree in the region of interest is at least 10% more than in the original image from the original dataset 410, which typically has a uniform pixel density or follows a linear projection. In some other embodiments according to the present invention, the controlled distortion is defined as having a region of interest where the number of pixels per degree is at least 20%, 30%, 40%, or 50% more than in the undistorted original image from the original dataset 410. For the field of view, its selection also depends on the specific application required by the neural network to be trained using the new dataset and can be any value different from an ultra-narrow field of view to an ultra-wide field of view. The field of view of the transformed image can be different from or the same as the field of view of the original image.

[0038] Once the target and field of view of the desired controlled distortion are selected, the next step is the image transformation step 440. The transformation includes an image transformation device that is configured to execute a software or hardware transformation algorithm. This device can perform certain image processing, including but not limited to distortion transformation. This processing is performed at the hardware or software level by any device capable of executing an image distortion change algorithm or any image processing algorithm. The image transformation device that changes the distortion of a digital image file can be of any type, including but by no means limited to a computer that includes a central processing unit (CPU), some memory units, and some means of receiving and transmitting digital image files. It can be a personal computer (PC), a smartphone, a tablet, an embedded system, or any other device capable of transforming the distortion of a digital image file. Such a device for transforming distortion can also consist mainly of a hardware algorithm or be executed on an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.

[0039] In step 440, the image transformation device receives the original digital image file without controlled distortion and, before transforming this original input digital image file into an output transformed digital image file with a controlled distortion target, receives the selection of the controlled distortion target. The output of step 440 is step 450, where a single digital image with the desired distortion and field of view is stored in memory or on a storage drive. The associated ground truth information or classification of this new digital image is either known from the information already available in the original dataset or determined by any other means, including any general nearest neighbor algorithm from near set theory or topological similarity to compare the original image and the new image. Then, in step 460, optionally, the single image 450 with distortion is used to create multiple similar transformed digital images by operating using data augmentation techniques such as rotation, translation, scaling, similarity enlargement, mirroring, or any other image transformation operation to increase the number of cases, orientations, sizes, or positions in the complete images used for training the neural network. The dataset can also be extended by using projection techniques, such as planar projection, line projection, perspective tilt correction, or any type of projection. Then, all the resulting images from the data augmentation step 460 are added to the new dataset 470 with intentionally distorted images as the final step in the method of creating this new image dataset for training the neural network. Then, this new transformed digital image file is used to train the neural network for inference processing of digital image files with controlled distortion.

[0040] Figure 5 The performance of neural networks trained without images with controlled distortion was compared with that of neural networks trained with images with controlled distortion in processing images with distortion. In Figure 5In this example, for simplicity, only one image file is shown, but the method according to the present invention is also compatible with multiple image files, which are combined or not combined into a digital video file. The original distorted image 510 is an example group picture of 5 people, which comes from Figure 2 the imager device described in and has an increase in pixels per degree towards the corners of the image. Such images are common in wide-angle imagers with a diagonal field of view exceeding 60°, where the corners stretch the image scale so as to keep the straight lines in the object as straight as possible in the image, with an increase in pixels per degree from the center to the corners. This stretching of the image makes it more difficult for automatic analysis using classical image processing or artificial intelligence-based image processing algorithms to output the best results, because the proportions of the human face are different from those used by the algorithms. Therefore, when the distorted image 510 is input into a neural network not trained with the distorted image 520, the output 530 is poor. In Figure 5 this example, the output of the network is the classification and recognition of people, but this is only an example output according to the present invention, and any image processing or image analysis output from the neural network can be used according to the present invention. In the result window 530, people A and E are stretched such that the algorithm cannot even classify their shapes as people. For people B and D, they are not stretched like this. The algorithm 520 can classify them as humans but cannot recognize them. The algorithm can only recognize the person C located in the center, because at the center of the image, the number of pixels per degree is basically unchanged and the proportion of the human face remains unchanged. When the same distorted image 510 is input into Figure 3 the neural network described using the distorted image 540 for training, the output result 550 is improved. This time, since the network is used to recognize people with distorted proportions, it can correctly recognize all five people. Figure 5 This example of

[0041] Figure 6 has classification and recognition applications, but when the input digital image file has controlled distortion, the convolutional neural network trained with the distorted image according to the method of the present invention provides improved performance for any application. Figure 6In the example, the imager output represents three images 610, 620, and 630 of a moving cat captured or generated at three different times representing different times from a video sequence, thus allowing an object of interest to be followed with an area of increased resolution. The intentionally controlled distortion added to image 610 is represented by a deformation grid 605. The circular regions 607 in the grid and 612 in the image represent regions where local magnification is applied as needed to distort the image to provide more imaging pixels to the neural network. If the total field of view remains the same, the regions of increased magnification are surrounded by regions of decreased magnification to compensate for the region of interest, and still have the same total field of view within the same total number of pixels. However, this is not always required, and in some other embodiments, the regions of increased magnification can be compensated by a smaller total field of view rather than by regions of decreased magnification.

[0042] After that, as represented by the vertical axis in the figure, the same kind of local magnification is applied to images 620 and 630 to which deformation grids 615 and 625 are respectively applied. The circular regions 617 and 627 in the grid and the circular regions 622 and 632 in the image represent the regions of this local magnification. In Figure 6 this example of, each image has only 1 region of local magnification, but this in no way limits the scope of the present invention, which can also be used in conjunction with multiple of these regions of local magnification in an image. The present invention can also apply multiple of these regions of local magnification in an image simultaneously. Then, the images with intentionally controlled distortion are input into an artificial intelligence neural network 645, which is trained by learning techniques using the distorted images, as Figure 3 explained. Due to the magnified view around the walking cat, the input to the neural network has more information pixels in that region. Since the network has an object with increased resolution as input, the output of the neural network 645 is an improved result 650.

[0043] In Figure 6 this example of, in all three cases, the neural network is able to identify the moving cat, but the application of the neural network 645 is not limited to identification and can be any other application with any type of output 650 according to the present invention. As a comparison, Figure 6A fourth output from the imager is also shown at 640, but this time without the real-time controlled distortion following the object of interest. The lack of intentional controlled distortion added to the image 640 is represented by the uniform grid 635. The image 640 is then processed in the neural network 655 and the result 660 is output, where the neural network 655 can be the same as or different from the network 645. In this example, this time, due to the resolution of the object of interest not being high enough, the neural network is unable to identify the cat in the image. In some embodiments of the present information, the neural network is configured to combine the inputs or outputs of at least two image frames captured or generated at different times in order to improve the result by giving a certain weight to the temporal consistency between consecutive image frames in the video. This video processing can optionally be done by using a recurrent neural network.

[0044] Figure 7 An example is shown where, before inputting an image with standardized controlled distortion into a neural network, the distortion of the image is transformed into a transformed digital image file with a standardized controlled distortion distribution format. In this example, the object of interest is a human face, but the method according to the present invention is not limited to any type of object and can be applied to any other object. The example starts with the original image 710. The original image 710 may or may not already have some controlled distortion. The source of the image can be any imager, including a device with an optical system or any device capable of generating virtual images or image transformations. In Figure 7 the example, each detected human face can be individually transformed into a unified and standard image format with controlled distortion. Three human faces in the image 710 are transformed into transformed digital image files 730, 740, and 750 with standardized controlled distortion using a software or hardware image transformation algorithm 720. The transformation applied can be the same or different for any human face, depending on, for example, the position of the human face in the image or the direction it is looking. The transformation algorithm 720 can be executed by any hardware device configured to transform the distortion distribution of the image, including a computer that contains a processor for executing software algorithms, an ASIC, an FPGA, etc.

[0045] In example images 730 and 750 with standardized controlled distortion, since the face is not directly looking at the image capture system, part of the face is not imaged by the camera, and thus black areas appear when transformed to this standardized view. The distorted image 740 is looking directly at the face, and there are no black areas with missing information after transformation to the standardized distorted view. Since the image format is standard, the neural network 760 only needs to be trained once, rather than being trained separately for each type of distorted image it can receive, which is the main advantage of using the standardized distortion format. Using the same standardized distortion format can avoid the costs and time required to generate new distorted datasets and retrain the neural network. In this example, the result output 770 of the neural network 760 is that all faces are well recognized, and better performance is obtained due to the standardized images with intentional distortion, but the output can be of any type, depending on the application using the neural network. The method of this example provides an improvement because a standardized controlled distortion distribution is selected to maximize the pixel coverage of the face in the MxN pixel input area, where M and N are the number of rows and columns in the input digital image, respectively. However, Figure 7 The views shown are only one example of a standardized projection for transforming digital image files, and any other standardized projection can be used according to the method of the present invention, including but not limited to images with equirectangular distortion projections, images with preset circular, rectangular, or free-form magnifications, etc.

[0046] In Figure 8 the example shown, the image transformation device at least partially removes controlled distortion from the input digital image file. This is done by processing or rectifying the input digital image file into a transformed digital image file before inputting the transformed digital image file into the neural network. In this example, the object of interest is the face, but the method according to the present invention is not limited to any type of object and can be applied to any other object. This example starts with the original image 810 with distortion. The source of this image can be any imager device with an optical system or any device capable of generating virtual images or image transformations. In Figure 8In an example embodiment, all detected human faces are processed in a software or hardware image transformation algorithm 820 to at least partially remove distortions. The transformation algorithm 820 can be accomplished by any hardware device configured to transform the distortion distribution of an image, including a computer that includes a processor for executing a software algorithm, an ASIC, an FPGA, etc. The transformation algorithm 820 performs distortion correction on the original image 810 to remove, correct, modify, or process the distortion, thereby obtaining human face images 830, 840, and 850 without distortion. Then, the images 830, 840, and 850 are input into a general neural network 860 that is trained using images without intentionally controlled distortion, and the output is result 870. In this example, the result output 870 from the neural network 860 is that all human faces are well recognized, which may be because the distortion in the original image 810 is distortion-corrected before being input into the neural network. The example output is not limited to human face recognition and can be of any type, depending on the application using the neural network.

[0047] In some other embodiments according to the present invention, the original image before input into the neural network includes additional information or parameters, whether written within the digital image file metadata, within visible or invisible markers or watermarks in the image, or transmitted to the neural network via another source. These additional information or parameters can be used to assist the image transformation algorithm or the neural network itself to further improve the result.

[0048] All of the above figures and examples show methods of using intentionally controlled distortion to improve the result output from a neural network. In all of these examples, the imager, camera, or lens can have any field of view from very narrow to extremely wide-angle. The neural network having at least an input and an output can be of any type. These examples are not intended to be exhaustive or limit the scope and spirit of the present invention. Those skilled in the art will understand that changes can be made to the above examples and embodiments without departing from their broad inventive concept. Therefore, it should be understood that the present invention is not limited to the specific examples or embodiments disclosed, but is intended to cover modifications within the spirit and scope of the present invention as defined by the appended claims.

Claims

1. A method for performing inference processing on at least one input digital image file with controlled distortion using an artificial intelligence neural network to improve the output of the neural network, the method comprising: a. Receiving, by the neural network, an input digital image file with controlled distortion created by an imager; b. Performing inference processing on the input digital image file by the neural network, wherein the neural network is formed by an algorithm or software code running on a computing device and has been specifically trained to process images with controlled distortion, and wherein the inference processing is performed without removing the controlled distortion from the input digital image file; c. Outputting, by the neural network, interpretation data derived from the input digital image file through the inference processing, the interpretation data output being an output digital image file with controlled distortion; and d. Performing distortion correction on the output digital image file to at least partially remove the controlled distortion.

2. The method according to claim 1, wherein The imager is a device that intentionally changes the distortion of a digital image file into controlled distortion.

3. The method according to claim 1, wherein The imager includes at least one camera system, and the at least one camera system includes at least one optical system.

4. The method according to claim 1, wherein The neural network is a machine learning neural network trained via deep learning techniques.

5. The method according to claim 1, wherein The neural network is trained using digital image files with controlled distortion.

6. The method according to claim 1, wherein The controlled distortion in the input digital image file from the imager varies over time.

7. The method according to claim 1, characterized in that The input digital image file includes ground truth information or classification information.

8. The method according to claim 1, characterized in that, The input digital image file with controlled distortion has at least one region of interest, wherein the resolution of the input digital image file in the at least one region of interest is at least 10% higher than that of a digital image file without controlled distortion.

9. A method for performing inference processing on at least one input digital image file with controlled distortion using an artificial intelligence neural network to improve the output of the neural network, the method comprising: a. Receiving, by an image transformation device, an input digital image file with controlled distortion created by an imager; b. Transforming, by the image transformation device, the input digital image file into a transformed digital image file, wherein the transformed digital image file has standardized controlled distortion; c. Performing inference processing on the transformed digital image file by the neural network, wherein the neural network is formed by an algorithm or software code running on a computing device and has been specifically trained to process images with standardized controlled distortion, and wherein the inference processing is performed without removing the standardized controlled distortion from the transformed digital image file; and d. Outputting, by the neural network, interpretation data derived from the transformed digital image file through the inference processing, the interpretation data output being an output digital image file with standardized controlled distortion; and e. Performing distortion correction on the output digital image file to at least partially remove the standardized controlled distortion.

10. The method according to claim 9, wherein The imager is a device that intentionally changes the distortion of a digital image file into controlled distortion.

11. The method according to claim 9, wherein The imager includes at least one camera system, and the at least one camera system includes at least one optical system.

12. The method according to claim 9, wherein The neural network is a machine learning neural network trained via deep learning techniques.

13. The method according to claim 9, wherein The neural network is trained using digital image files with controlled distortion that is standardized.

14. The method according to claim 9, wherein The controlled distortion in the input digital image files from the imager varies over time.

15. The method according to claim 9, characterized in that, The input digital image files include ground truth information or classification information.

Citation Information

Patent Citations

  • Image distortion transformation method and apparatus

    US10204398B2

  • Face detection using small-scale convolutional neural network (CNN) modules for embedded systems

    US10268947B2

  • Depth estimation method for monocular image based on multi-scale CNN and continuous CRF

    US10353271B2

  • Object recognition based on boosting binary convolutional neural network features

    US20180032844A1

  • Age and gender estimation using small-scale convolutional neural network (CNN) modules for embedded systems

    US20180150684A1