Use of an imaging device that deliberately controls distortion for inference or training of an artificial intelligence neural network

By training convolutional neural networks to process deliberately distorted images with higher resolution regions of interest, the method addresses the limitations of existing networks in handling high-resolution and distorted images, achieving improved processing accuracy and detail.

JP7698937B2Active Publication Date: 2025-06-26IMMERVISION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024003426
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-18
Filing Date
2024-01-12
Publication Date
2025-06-26
Estimated Expiration
2040-11-17

AI Technical Summary

Technical Problem

Existing convolutional neural networks are limited in their ability to process high-resolution images due to their restricted input resolution, which is typically around 512x512 pixels, and they struggle with images containing controlled distortion, such as those with wide-angle views.

Method used

A method is proposed to train and utilize convolutional neural networks using deliberately selected distorted images, where an imaging device creates a file of a digitally distorted image with a region of interest having at least 10% higher resolution than normal images. This distorted image is then input into a neural network that has been specially trained to process such images, resulting in more detailed interpretation data for the region of interest and potentially improving the processing of the entire image.

Benefits of technology

The method allows for improved processing and analysis of high-resolution images with controlled distortion, enhancing the accuracy and detail of outputs such as image depth information, object recognition, and super-resolution images, while also reducing the need for retraining the neural network for different distortion contours.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698937000001
    Figure 0007698937000001
  • Figure 0007698937000002
    Figure 0007698937000002
  • Figure 0007698937000003
    Figure 0007698937000003
Patent Text Reader

Abstract

To provide a method of training and using a convolutional neural network with images having on-purpose distortion, and a method of training the neural network using a distorted image generator from various existing datasets.SOLUTION: By distorting an image on purpose to create a region of interest with higher number of pixels than other regions, a resulting output from a neural network is improved. The distorted image is created using an imager device comprising either an optical system specifically designed to output distorted images, or software or hardware image distortion manipulation algorithm for creating distorted images from normal images.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to U.S. Provisional Patent Application No. 62 / 936,647, filed November 18, 2019 (currently under examination), the entire disclosure of which is incorporated herein by reference.

[0002] Embodiments of the present invention relate to convolutional neural networks of artificial intelligence and their use, and more particularly to a method of appropriately using such neural networks using an imaging device that controls distortion.

Background Art

[0003] The use of artificial intelligence for the processing or analysis of digital images is becoming increasingly popular. This is due to the increasing available processing power in personal computers or mobile devices, or in large server farms provided by large companies. The increasing use of artificial intelligence can also be explained by the fact that its new capabilities are applicable in a wide range of applications.

[0004] The most common type of neural network used for the analysis of digital images is the convolutional neural network. This means that a convolutional operation is performed in multiple layers of the neural network. The idea of using neural networks (NNs) for their processing in general applications of digital images has already been seen in past cases such as U.S. Patents 9,715,642, 10,360,494, etc., including the use of convolutional neural networks (CNNs). The use of convolutional neural networks for specific applications has also been seen in past cases such as object recognition like U.S. Patent Application Publication 2018 / 0032844, face recognition like U.S. Patent 10,268,947, depth estimation like U.S. Patent 10,353,271, age / gender estimation like U.S. Patent Application Publication 2018 / 0150684, etc.

[0005] However, the images processed by existing convolutional neural networks have a significantly limited resolution at the time of input. This is particularly true in applications that require global image analysis, such as depth estimation from a single image, where the image cannot be divided into small parts and each part processed sequentially. Even when these neural networks utilize the latest computers equipped with GPUs containing approximately 10 gigabytes of RAM, the resolution of the images that can currently be analyzed and processed is limited to approximately 512×512. This resolution is about 250,000 pixels, which is significantly lower than the 20 - 50 million pixels available on the latest mobile devices or cameras. Thus, in certain applications of digital images, the resolution at which effective processing is possible is limited. As a result, the processing or analysis by neural networks is cruder than what would be achieved if the maximum resolution of the input image were utilized. This limitation is even more severe in applications where embedded systems with severely limited processing capabilities are used.

[0006] One way to increase the number of pixels of the object of interest without increasing the total number of pixels of the image is to deliberately use controlled distortion. The idea of deliberately changing the resolution of an image by an optical system, hardware, or software has already been seen in past cases such as U.S. Patents 6,844,990, 6,865,028, 9,829,700, 10,204,398, etc. However, the distorted images obtained from these imaging devices cannot be well - analyzed or processed by existing convolutional neural networks. The use of deliberately controlled distorted images requires a new type of neural network or training method. As another method of giving high - resolution input to a neural network, a small area may be cut out from the entire image and only that area analyzed inside the neural network. However, when a small area or area of interest is cut out from the entire image, the connection and overall information of the entire scene are lost. These are important in applications where the neural network needs to extract global information from the entire scene to give the best output.

[0007] Among digital images, a frequently encountered type of controlled distortion is in wide-angle images, especially those with an overall field of view generally wider than about 80°. However, it is rare for such wide-angle images to contain relevant ground truth data compared to narrow-angle images without controlled distortion. Most of the existing datasets of large images used for training existing neural networks are based on narrow-angle images without distortion. Therefore, a new training method is required to train neural networks using wide-angle images with deliberately controlled distortion or narrow-angle images.

Summary of the Invention

[0008] To solve all of the above problems, embodiments of the present invention provide a method of training and using a convolutional neural network using deliberately selected distorted images.

[0009] In a preferred embodiment according to the present invention, the method first causes an imaging device to create a file of a digitally distorted image. This imaging device can be anything that creates a distorted image, such as a virtual image generator, a device that executes software or hardware for image distortion processing, or a device that directly captures a controlled distorted image with an optical system and an image sensor installed on its focal plane. The images that this imaging device can output include static distortion with a constant contour or dynamic distortion whose contour can change over time. In a preferred embodiment, the controlled distorted image output from the imaging device includes at least one region of interest. The region of interest has a resolution (calculated as the number of pixels per degree of angle of view) that is at least 10% higher than that of a normal digital image without controlled distortion. The controlled distorted image is then input into a neural network (any type is acceptable). This neural network includes at least one convolutional layer. This is common but not essential for the present invention. This neural network can operate on any physical device with the function of executing an algorithm and can process the above distorted image as long as it has received special training using the controlled distorted image. As a result of inputting this image into a neural network that has received special training using distorted images, more detailed interpretation data for the region of interest with an increased number of pixels is output. The interpretation data can then also be used to improve the results for the image portion outside the region of interest. The interpretation data to be improved can be anything depending on the use of the neural network, such as image depth information, object recognition results, object classification results, object segmentation results, optical flow estimation results, results of connecting edges and lines, results of SLAM (a technique that simultaneously performs self-position estimation and environmental map creation), or an image obtained by super-resolution, etc. In some of the embodiments of the present invention, the interpretation data output from the neural network can also be a controlled distorted image. In this case, depending on whether this image is observed by a human or not, if necessary, a final output image without distortion can be obtained by correcting the distortion of the image and restoring it to its original form.This process, which is selected as necessary, is not necessary when the output of the neural network is directly utilized by units of another algorithm, computers, or other automated processes.

[0010] To use input files of digital images with deliberately controlled distortion in a convolutional neural network, the neural network must be trained with special training provided for those images. The above method according to the present invention utilizes a device that generates a dataset of distorted images from an existing dataset of large images without controlled distortion. The existing image dataset includes images of various types of objects captured with a normal lens without deliberately causing distortion, so it cannot be directly used for training the neural network proposed by the inventor. The device that generates the dataset of distorted images processes the original images from the existing dataset to deliberately add distortion. The distortion can be of any type, such as rotationally symmetric distortion, free-form distortion (which may or may not have a center for a specific object), or elongation of the corners of the image. The resulting dataset of distorted images can then be expanded, if necessary, by operations such as data augmentation, i.e., rotation, translation, scaling, similarity transformation, and mirroring, etc., to increase the number of states of the images used for training the neural network. Any type of projection method, such as orthographic projection, spherical aberration correction, perspective view tilt correction, etc., can be used for expanding the dataset. In this way, a new dataset generated from images with controlled distortion is used for training the neural network. The learning of how to use images with controlled distortion by the neural network can be of any type, with or without a teacher.

[0011] In some of the other methods according to the present invention, the original image output from the imaging device is first converted into a standardized display with clear boundaries, regardless of the presence or absence of distortion. This display includes deliberately controlled distortion that is standardized. The standardization of the distortion aims to use a neural network that has received special training using standardized distortion with standardized contours. This avoids having to retrain the neural network for a long time each time the distortion contour is updated. This standardized display may or may not have an area of information about the lost texture, depending on the method of capturing the original image and the requirements of its display.

[0012] In some of the other embodiments according to the present invention, the original image output from the imaging device first undergoes a process of removing or minimizing the distortion of the image. The processed image is utilized by an existing neural network that has already been trained to utilize an image without controlled distortion. This eliminates the need to train a new neural network to be prepared for the specific distortion contour obtained as a result of the output from the imaging device.

Brief Description of the Drawings

[0013] The above summary will be better understood when read in conjunction with the detailed description of the preferred embodiments of the invention described below, in association with the accompanying drawings. For the purpose of explaining the invention, the drawings show the preferred embodiments at the present time. However, it should be understood that the invention is not limited to the details of the arrangements and means shown.

[0014]

Figure 1

[0015]

Figure 2

[0016]

Figure 3

[0017]

Figure 4

[0018]

Figure 5

[0019]

Figure 6

[0020]

Figure 7

[0021]

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0022] When the words "a" and "an" are used in the corresponding parts of the claims and the specification, they mean "at least one".

[0023] Figure 1 shows the inference process for a normal image using an artificial intelligence neural network according to the prior art. The artificial intelligence neural network 100 performs image processing on the normal image 110 and outputs the result 140. The neural network can be of any type. In some embodiments, the neural network may be a convolutional neural network (CNN) trained in deep machine learning or the like. However, embodiments according to the present invention are not always so, and other neural networks may be used. Further, the neural network may or may not perform image convolution. In some embodiments, the neural network may be an adversarial generative network (GAN). The normal image 110 to be input is input to the neural network through the input nodes of the input layer 120 for the purpose of inference processing. Since the exact number of nodes depends on the application, the figure with 3 input nodes is only an example of a neural network and does not limit the types of neural networks that can be used for processing the input digital image. The neural network may also include hidden layers of unknown number such as the layers 125 and 130 shown in the figure of this example. Each hidden layer may have any number of nodes. The neural network may also include a plurality of sub-networks or sub-layers, each processing a different task. Those tasks include, but are not limited to, convolution, pooling (max pooling, average pooling, or other types of pooling), striding, padding, downsampling, upsampling, multi-functional fusion, rectified linear transformation, concatenation, fully connected, or flattening, etc. The neural network may also include a final output layer 135. The output layer may be composed of any number of output nodes. In the figure of this example, the dashed line represents the nodes of the neural network that have not been trained using images with controlled distortion. The interpretation data 140 output from the neural network is the result of the input of the original digital image.The types are diverse and include, but are not limited to, image depth information, object recognition results, object classification results, object segmentation results, optical flow estimation results, edge and line connection results, SLAM results, or images obtained by super-resolution, etc. Since the input digital image 110 has no controlled distortion that creates a region of interest, there is no part where the number of pixels has been increased. Therefore, the output of the neural network follows the existing prior art. Specifically, the application shown in the example of FIG. 1 is the generation of a depth map from an input image. The resulting depth map has low resolution everywhere in the image. It also includes, in the example of FIG. 2, the car that is the object of interest.

[0024] FIG. 2 shows an inference process using a neural network for an image with controlled distortion for the purpose of improving the output of an artificial intelligence neural network. This method first causes the imaging device 205 to create an image with controlled distortion. The imaging device 205 may be of any type as long as it is a device that creates a file of a digital image with controlled distortion so as to increase the number of pixels in the region of interest. The imaging device 205 includes a virtual image generator, a device that executes an algorithm for distorting an image by software or hardware, or a device that deliberately changes the distortion of a digital image, but the scope of the present invention is not limited thereto. The device that deliberately changes the distortion of a digital image may be of any type, such as a personal computer (PC), a smartphone, a tablet, or a computer equipped with a central processing unit (CPU), a memory device, and means for transmitting and receiving a digital image file, or any other device having a function of converting the distortion of a digital image. However, it is not limited to these. The imaging device 205 may operate according to an algorithm mainly executed by hardware such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). The imaging device 205 may be a device equipped with at least one camera system. This camera system includes at least one optical system that forms an image including controlled distortion and the like. This optical system may be formed by any combination of refractive optical elements, reflective optical elements, diffractive optical elements, optical elements made of metamaterials, or other optical elements. This optical system may also include an active optical element, such as a deformable mirror, a liquid lens, or a spatial light modulator, for the purpose of changing and adapting the contour of the distortion added to the image in real time. This optical system can further control the distortion better by using aspherical or free-form optical elements to increase the resolution of at least one region. In some of the embodiments according to the present invention, the optical system is preferably a wide-angle lens with a diagonal angle of view wider than 60°. This wide-angle lens includes a plurality of optical elements, which are divided in the order of a front group, a diaphragm, and a rear group.The wide-angle lens forms an image on the image plane.

[0025] The output of the imaging device 205 is an image 210 with deliberately controlled distortion. In the example of FIG. 2, only one image is shown for simplicity, but the method according to the present invention can also handle multiple images. These images may or may not be incorporated into a single digital video. This digital image 210 has controlled distortion and defines at least one region of interest. In the region of interest, the resolution (i.e., magnification), calculated as the number of pixels per degree of the angle of view, is at least 10% higher than that of a normal digital image 110. In some other embodiments according to the present invention, the controlled distortion is defined such that the number of pixels per degree of the angle of view in the region of interest is at least 20%, 30%, 40%, or 50% more than that of an undistorted image. By creating at least one region of interest with such a resolution, the imaging device 205 may maintain or change the overall angle of view to be equal to that of an image without a region of interest.

[0026] A file of a digital image 210 with deliberately controlled distortion is input into an artificial intelligence neural network 200. The neural network 200 can be of any type and includes a machine learning neural network trained by deep learning. However, it is not limited to this, and may include a convolutional neural network (CNN) or the like. The neural network 200 includes an algorithm or software code, etc. that is executed on a physical computing device to interpret input data (which can be of any type), and is trained to process an image with controlled distortion. This physical computing device can be any hardware that has the function of executing its algorithm, etc., and can be a personal computer, mobile phone, tablet, automobile, robot, or embedded system, etc., but is not limited to these. This physical computing device may be equipped with any of the following: an electronic main board (motherboard), at least one processor, part or all of a central processing unit (CPU), memory (RAM, ROM, etc.), a drive (hard disk drive, SSD, etc.), an image processing unit (GPU), a tensor processing unit (TPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other elements that cause the neural network to convert the input file of the digital image into output data indicating the result of the interpretation.

[0027] In the example of FIG. 2, the artificial intelligence neural network 200 has received special training using images with controlled distortion. This training aims to process those images even better, as will be described later based on FIG. 3. The input file of the digital image 210 with controlled distortion is received by the neural network 200 through the input nodes of the input layer 220. The number of nodes depends on the application. The three input nodes shown in the figure are just an example of a neural network and do not limit the types of neural networks available for processing the input digital image. The neural network 200 may also include hidden layers of unknown number, such as layers 225 and 230 shown in the figure of this example. Each hidden layer may have any number of nodes. The neural network 200 may also include multiple sub-networks or sub-layers, each processing different tasks. Those tasks include, but are not limited to, convolution, pooling (max pooling, average pooling, or other types of pooling), striding, padding, downsampling, upsampling, multi-functional fusion, rectified linear transformation, concatenation, fully connected, or flattening, etc. The neural network 200 may also include an output layer 235. The output layer may be composed of any number of output nodes. In the figure of this example, the solid lines represent the nodes of the neural network trained using images with controlled distortion, and the arrows from left to right within the neural network 200, that is, from the input layer to the output layer of the neural network 200, represent the flow of the inference process of the neural network 200. The neural network 200 then performs an inference process on the input file of the digital image and outputs interpretation data. The interpretation data 240 output from the neural network 200 is obtained from the input file of the digital image 210 with controlled distortion. Its types are diverse and include, but are not limited to, image depth information, object recognition results, object classification results, object segmentation results, optical flow estimation results, results of edge and line connections, SLAM results, or images obtained by super-resolution, etc.

[0028] The digital image 210 of the input file has controlled distortion that creates a region of interest within the image, so there is at least one part where the number of pixels has increased. Therefore, the result of the interpretation output from the artificial intelligence neural network 200 is improved compared to the result obtained from an input file of a digital image without controlled distortion, such as the output 140 of the prior art. This improvement may be, for example, an increase in the accuracy of the depth map, accompanied by an increase in the number of pixels representing the resolution of the depth map, when the use of the artificial intelligence algorithm is the estimation of the depth map from a single image schematically shown in FIG. 2. The above improvement may be an improvement in object classification or object recognition performance accompanied by an increase in the number of pixels of the object of interest, or any other improvement in results. The improvement for at least one image can be evaluated in different ways depending on whether the output of the neural network is qualitative or quantitative. The evaluation targets include, but are not limited to, the reduction amount of the relative value (calculated in % units) or the absolute value (calculated in units suitable for the use of the neural network) of the difference between the output and the ground truth data, the root mean square (RMS) error, the mean relative error, the mean logarithmic error (with base 10), or the accuracy of the threshold, etc. The degree of the above improvement can also be calculated as a score based on the number of true positives, false negatives, true negatives, and false positives included in the output, such as the precision (P-score), recall (R-score), F-score, etc. The degree of the above improvement can also be measured as the degree of increase in the probability of obtaining the output, that is, the reliability of the output, especially when the output of the neural network is qualitative, such as when the neural network performs classification. In some embodiments, the degree of improvement between the case where the original image has controlled distortion and the case where it does not is measured as the rate of increase in accuracy between them by comparison between the result obtained from a large dataset of the input file of the digital image with controlled distortion and the result obtained from a similarly large dataset of the input file of the digital image without controlled distortion.

[0029] In the example of FIG. 2, the output of the neural network is a digital image file. However, this is not always the case, and the output may be any output obtained by inputting characters, optical signals, tactile feedback, or other images with controlled distortion into the neural network. If the output is a file of digital image 240, and if the image is to be observed by a human, the image may be further processed by image distortion correction as needed, so that at least a part of the controlled distortion is removed, thereby reducing or completely removing the controlled distortion, and a file of digital image 250 with reduced or completely removed controlled distortion may be provided. This distortion correction added as needed is performed directly by hardware configured to process the output file of digital image 240 to remove, correct, or process its distortion, in accordance with software algorithms operating on a computer formed by a processor, or.

[0030] This distortion correction, which is added as necessary, may be unnecessary when the output image is to be used by software or hardware algorithms or other computers without going through a person. In some embodiments of the present invention, the entire neural network 200 consists of several sub-networks. These sub-networks are configured to analyze the overall structure of the image and the local structure of each part thereof, and combine the results thereof. For the purpose of analyzing the overall structure of the image, the sub-network includes several downsampling layers and subsequent layers, and an upsampling layer that returns the resolution to the value in the original image. These layers may or may not use convolution. For the purpose of analyzing the local structure of each part of the image, the sub-network may be able to process those parts, for example, by directly cutting out several parts from the original image, or by inputting them from the sub-network for downsampling or upsampling used for analyzing the overall structure of the image to an intermediate layer. However, it is not limited to that configuration. The results obtained from the sub-network used for analyzing the overall structure of the image and the sub-network used for analyzing the local structure of the image are then combined in an averaging layer or a combination / convolution layer, etc., so that the final output of the entire neural network can be generated.

[0031] Figure 3 shows a method for training an artificial intelligence neural network by deep learning aimed at improving the processing ability for an image with controlled distortion. In the example of Figure 3, only one image is shown for simplicity, but the method according to the present invention can also be applied to digital videos. This training method of the neural network can be by any of supervised learning, semi-supervised learning, and unsupervised learning. First, in a large image database, an image 310 without deliberately added controlled distortion is prepared. The image data stored in this database is often also called a dataset. In the example of Figure 3, the original image 310 without controlled distortion is an image of a cat, and the balance (proportion) of its body dimensions is normal. For the purpose of enabling the neural network to be trained by supervised learning, semi-supervised learning, or unsupervised learning using a given large dataset of images, the method according to the present invention processes the original image 310 represented by the dataset with an image conversion algorithm 320 by software or hardware, and gives the original image 310 a controlled distortion with the desired contour. The way of giving it is the same as the way of giving it to the digital image 210 output from the imaging device 205 in Figure 2. The image conversion algorithm 320 is executed by an image conversion device. This will be further described with reference to Figure 4. In some of the embodiments according to the present invention, in addition to the images themselves, the same method may also be used to deliberately add controlled distortion to the corresponding images (well-known as ground truth images) required as the results of their processing. The desired controlled distortion can be of any type. It includes, but is not limited to, the following. Barrel distortion in the radial direction with rotational symmetry. This is as shown in Example 330 and often appears in wide-angle images. Free-form distortion. This may or may not have rotational symmetry, and may or may not have a center for a specific object as shown in Example 340. Elongation of a part of the image, i.e., spool-type distortion. This may be seen only at the corners of the image as shown in Example 350, or in any other part of the image. Elongation of the entire image as shown in Example 360, i.e., spool-type distortion.In addition to these, any type of distortion may be used as long as it produces at least one region of interest in which the number of pixels per degree of angular field is at least 10% more than that of a complete image. The complete image may mean an image with corrected spherical aberration and uniform pixel density and size, or any other image that is ideal for a given neural network. In some of the other embodiments according to the present invention, controlled distortion is defined such that the number of pixels per degree of angular field in the region of interest is at least 20%, 30%, 40%, or 50% more than that of an undistorted image.

[0032] Images with controlled distortion are newly generated. All of them may be the same or different in angular field compared to the original undistorted image. If the angular field of the newly generated image is wider than that of the original image, the remaining part of the image may be filled with any image. The image may be a background image generated by a computer, a background image extracted from other images, a multiple replicated original image, a plurality of images drawn from the original dataset, an image by extrapolation, or any image necessary to fill the lost part from the angular field, such as a blank.

[0033] A new dataset generated from images with controlled distortion, such as images 330, 340, 350, and / or 360, is then used for training the neural network 370. The neural network 370 learns how to use these images with controlled distortion. In the example of FIG. 3, the arrows in the neural network 370 shown schematically go from right to left, i.e., from the output layer to the input layer of the neural network 370. These represent the flow of training of the neural network 370 by the error backpropagation method, which is different from the flow of inference processing from the input layer to the output layer represented by the arrows going from left to right in other figures. The learning of the neural network 370 may be supervised learning (pairs of images input to the neural network and the ground truth images that should be output from the neural network as a result are known), or unsupervised learning (the images input to the neural network are associated with ground truth images whose output from the neural network as a result is unknown). The new dataset of images can also be used for any type of deep learning that trains and further enhances a neural network, such as a hybrid type of supervised and unsupervised learning (known as semi-supervised learning), or any method of training artificial intelligence using a dataset of images. When training the neural network, any technique may be used for optimizing the weights between the nodes of each layer. Such techniques include gradient descent, error backpropagation, genetic algorithms, annealing methods, random optimization algorithms, etc. However, these do not limit the scope of the present invention. The loss function (also known as the cost function or energy function) used for optimizing the neural network may be of any type depending on the application required of the neural network in the method according to the present invention. In some embodiments of the present invention, when training a neural network to analyze or process wide-angle images whose viewing angle is generally wider than about 60°, wide-angle images generated from a virtual three-dimensional space may be used for the training.This is because it is quite rare for existing wide-angle images to accompany the ground truth images required in the desired applications, and it is often the case that they do not exist. If there is only a small dataset of wide-angle images and the dataset needs to be enlarged for accurate training, virtual wide-angle images may be combined with the existing real wide-angle images for use.

[0034] Figure 4 shows a method of creating a dataset of distorted images from an original dataset using an image conversion algorithm implemented in software or hardware in an image conversion device. This method first prepares a dataset of original images (step 410). There are multiple such datasets publicly available on the Internet, which include images of natural real objects, artifacts, virtual objects, or mixtures thereof. The objects represented by the still images or videos included in the above dataset can be selected from various types useful for training various types of artificial intelligence neural networks, such as characters, human faces, animals, buildings, street views, etc. The images of the existing dataset are captured through a normal lens without being deliberately distorted, or generated from a normal scene without controlled distortion being added. This method then selects one of the images from the above dataset as the target (step 420). In the example of the method shown in Figure 4, only one image is converted from the original dataset. However, in the actual case of generating a new dataset, the same method can be continuously applied to the desired number of original images. Also, the method according to the present invention can also handle the creation of a dataset from multiple image files. Whether the dataset is assembled as a file of a single digital video or not is acceptable. In some of the embodiments according to the present invention, in addition to step 410 of processing the original images, a step of deliberately adding controlled distortion in the same way to both the original images and the resulting necessary images (well-known as ground truth images) corresponding to those images is provided.

[0035] In the next step 430 of the above method, what is required as deliberately controlled distortion, that is, the target controlled distortion and the required angle of view, are selected. The target controlled distortion depends on the specific application required for the neural network to be trained using the new dataset, and can be of any kind. Such distortion can be barrel distortion in the radial direction with rotational symmetry (often seen in wide-angle images), free-form distortion (which may or may not have rotational symmetry, and may or may not have a center for a specific object), stretching of a part of the image, that is, spool-type distortion (which can be seen in only the corners of the image or any other part of the image), stretching of the entire image, that is, spool-type distortion, or any other distortion that generates at least one region of interest where the number of pixels per degree of angle of view is at least 10% more than that of the original image included in the original dataset (prepared in step 410). However, it is not limited to these. The original image may generally have a uniform pixel density or may have spherical aberration corrected. In some of the other embodiments according to the present invention, the controlled distortion is defined such that the number of pixels per degree of angle of view in the region of interest is at least 20%, 30%, 40%, or 50% more than that of the distortion-free original image included in the original dataset (prepared in step 410). The selected angle of view also depends on the specific application required for the neural network to be trained using the new dataset, and can be any different value between an ultra-narrow angle and an ultra-wide angle. The converted image may or may not have a different angle of view from the original image.

[0036] After the desired controlled distortion and the required field of view are selected, the next step is the image conversion step 440. The image conversion device provided for this conversion is configured to execute an image conversion algorithm by software or hardware. In this image conversion device, several image processes are possible, such as, but not limited to, distortion processing. Image processing is performed by a device having a function of executing an image conversion algorithm for distortion processing or other image processing algorithms, operating on hardware or executing software. The image conversion device for changing the distortion of a digital image may be of any type, and may be, but not limited to, a computer equipped with a central processing unit (CPU), a memory device, and means for transmitting and receiving digital image files. The image conversion device may be any other device having a function of converting the distortion of a digital image, such as a personal computer (PC), a smartphone, a tablet, an embedded system, etc. The device for converting the distortion may operate according to an algorithm mainly executed by hardware such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0037] In step 440, the image conversion device receives a file of the original digital image without distortion and accepts the selection of the desired controlled distortion. Then, it converts the input file of the original digital image into an output file of the digital image with the desired controlled distortion. The output of step 440 is processed in step 450. In step 450, a digital image with the desired angle of view and the desired distortion is stored in a memory or a storage drive. The relevant ground truth information, or the classification of the new digital image after conversion, is known from the information already available in the original dataset or is determined by some other means. Such means include general approximation algorithms based on the theory of near sets or the comparison of topological similarities between the original image and the new image. In the subsequent step 460, one distorted image handled in step 450 is used, if necessary, to create a plurality of similar digital images that have undergone conversion processing. In this creation, data augmentation, i.e., rotation, translation, scaling, similarity transformation, mirroring, and other image conversion operations, are used to increase the number of overall states, orientations, sizes, or positions of the images used for training the neural network. For the expansion of the image dataset, any type of projection method such as orthographic projection, spherical aberration correction, perspective tilt correction, etc. may be used. All the images obtained as a result of the data augmentation step 460 are added, in the next step 470 which is the final step of this method, to a new dataset of images containing intentional distortion. This new dataset is used for training the neural network. In particular, the converted digital images of this new dataset are used for training the neural network prepared for the inference process on digital images.

[0038] Figure 5 shows a comparison of the capabilities between a neural network trained without using images with controlled distortion and a neural network trained using images with controlled distortion. In the example of Figure 5, only one image is shown for simplicity purposes, but the method according to the present invention can also handle multiple images. These images may or may not be incorporated into a single digital video. The original distorted image 510 is, for example, a group photo of five people, obtained from an imaging device as shown in Figure 2, with the number of pixels per degree of the field of view increasing towards the four corners of the image. This type of distortion in the image is common in wide-angle imaging devices with a diagonal field of view exceeding 60°, where the shape of the image stretches at the four corners and the number of pixels per degree of the field of view increases from the center of the image towards the four corners, and the straight lines in the object are kept as straight as possible in the image. Such stretching of the image makes it difficult to optimize the output by automatic analysis using classical image processing or image processing algorithms based on artificial intelligence. This is because the balance of the facial dimensions is not the target of the image processing algorithm. For this reason, when the distorted image 510 is input into a neural network 520 that has not been trained using distorted images, its output 530 is inferior. In the example of Figure 5, the output of the neural network is the classification and recognition of people. However, this is only an example of the output according to the present invention, and the present invention can be applied even if the output from the neural network is the result of other image processing or image analysis. As shown by the output results within the window 530, the images of people A and E are stretched, so the neural network 520 was unable to classify their shapes as human beings even by the image processing algorithm. The images of people B and D are not stretched that far. The neural network 520 was able to classify their shapes as human beings by the image processing algorithm but was unable to recognize them. Only the person C in the center was recognized by the image processing algorithm of the neural network 520. This is because the number of pixels per degree of the field of view is almost constant in the center of the image, and the balance of the facial dimensions of the person is maintained.When an image 510 with the same distortion is input into a neural network 540 that has been trained using distorted images, as shown in FIG. 3, its output result 550 is improved. In this case, since the neural network 540 is used for recognizing people with distorted proportions, all five people could be recognized correctly. The example in FIG. 5 is for classification and recognition applications. However, a convolutional neural network trained using distorted images according to the method of the present invention exhibits improved capabilities for any application when the input file has deliberately controlled distortion in the digital image.

[0039] FIG. 6 shows an example where the deliberately controlled distortion included in the digital image output from the imaging device changes over time, such as in a plurality of frames forming one sequence of a video. In this example, the imaging device may be a camera system equipped with an active optical element that changes the distortion over time, hardware capable of directly converting the image distortion, or a device (computer, mobile phone, tablet, embedded system, ASIC, FPGA, etc.) capable of executing an image conversion algorithm by software. In the example of FIG. 6, the output of the imaging device is three images 610, 620, 630 of a moving cat. These are images representing different times, captured or generated at three different times from one sequence of the video, and the object of interest can be tracked using the area where the resolution is increased. The deliberately controlled distortion added to image 610 is represented by a distorted mesh 605. The circular region 607 within this mesh and the circular region 612 within image 610 represent the regions where distortion has been added to the image by local magnification. This distortion is of the degree necessary to provide more pixels to the neural network. If the overall viewing angle is the same, the magnified area is surrounded by the reduced magnification area. This cancels out the increase in the number of pixels in the region of interest and the expansion of the viewing angle, so that the whole maintains the same viewing angle within the same total number of pixels. However, this is not always necessary, and in some other embodiments, the increase in magnification in some regions may be canceled out not by a decrease in magnification in other regions, but by a reduction in the overall viewing angle.

[0040] In the subsequent time domain represented by the vertical axis in the figure, a similar local enlargement is performed on images 620 and 630. Distorted meshes 615 and 625 are applied to each of these images. The circular regions 617 and 627 of each mesh and the circular regions 622 and 632 of each image represent the enlarged local regions. In the example of FIG. 6, there is only one enlarged local region per image. However, this does not limit the scope of the invention. The present invention can also be implemented to simultaneously enlarge a plurality of local regions within one image. The images with deliberately controlled distortion are then input into the artificial intelligence neural network 645. This neural network 645 is trained by learning using the distorted images, as described with reference to FIG. 3. Since the area around the walking cat is enlarged, the input to the neural network 645 contains a larger amount of pixel information around the cat. Since an image of an object with increased resolution is input into the neural network 645, the output result 650 is improved.

[0041] In the example of FIG. 6, for all three images, the neural network 645 was able to recognize the moving cat. However, the use of the neural network 645 is not limited to recognition. Depending on any other use, the result 650 according to the present invention may be of any kind. FIG. 6 also shows, as one comparison, a fourth image 640 output from the imaging device. However, at the time of its output, there is no controlled distortion for real-time tracking of the object of interest. The fact that no deliberately controlled distortion has been added to the image 640 is represented by the uniform mesh 635. The image 640 is then processed by the neural network 655, and its output is the result 650. The neural network 655 may be the same as or different from the neural network 645. In this example, since the resolution of the object of interest is not high enough, the neural network 655 was unable to identify the cat in the image. In some embodiments of the present invention, the neural network is configured to combine at least two image frames captured or generated at different times among its input or output. Thereby, consistency in weight and chronological order occurs between consecutive image frames in a single video, so that the result is improved. Such video processing may be performed using a regression neural network if necessary.

[0042] Figure 7 shows an example in which the distortion of an image is converted into a controlled distortion with a standardized contour before the image is input into a neural network. In this example, the object of interest is a human face. However, the target of the method according to the present invention is not limited to any type of object, and it may be any other object. This example first obtains the original image 710. The original image 710 may or may not already have controlled distortion. The source of this image may be any imaging device, such as a device including an optical system, or any device having a function of generating or converting a virtual image. In the example of Figure 7, the detected human faces are individually converted into a standard format unified for the images with controlled distortion. Three faces in the image 710 are converted using an image conversion algorithm 720 by software or hardware, and changed into digital images 730, 740, 750 with controlled distortion that is standardized. The conversion applied may be the same for each face, depending on, for example, the position or orientation within the face image, or may be different for each face. The image conversion algorithm 720 may be executed on any hardware configured to convert the contour of the image distortion, such as a computer equipped with a processor that executes a software algorithm, an ASIC, an FPGA, etc.

[0043] In examples 730 and 750 of the standardized and controlled distorted images, since the human faces were not directly facing the imaging device, some parts of those faces were not captured by the camera, resulting in black regions appearing when converted to the standardized display. In the distorted image 740, since the face was directly facing the imaging device, no black region representing information loss appears even when converted to the standardized display. Since the image format is standard, the neural network 760 only needs to be trained once for the distorted images and does not need to be trained for each receivable type. This is the main advantage of using the standard format of distortion. That is, by using the same format, there is no need to spend cost and time generating a new dataset of distorted images to retrain the neural network. In this example, the output 770 obtained as a result from the neural network 760 indicates successful recognition for all faces. That is, since the intentional distortion included in the image is standard, the processing ability is improved. However, depending on the application of the neural network, its output can be of any type. The method of this example is improved because the standardized contour of the controlled distortion is selected so that the M×N pixels (where M is the number of rows and N is the number of columns in the input digital image) covering the human face are maximized. However, what is schematically shown in FIG. 7 is only an example of the projection method standardized for the conversion of digital images, and any other projection method may be used in accordance with the method of the present invention. Such projection methods include, but are not limited to, orthographic cylindrical projection, or the expansion of a predetermined region in a circular shape, rectangular shape, or free shape.

[0044] FIG. 8 shows an example in which an image conversion device removes at least a part of controlled distortion from a digital image of an input file. In this removal, the image is processed before being input into a neural network to correct the distortion, and the converted image is input into the neural network. In this example, the object of interest is a human face. However, the method according to the present invention is not limited to any particular type of object and is applicable to any object. This example first prepares an original distorted image 810. The source of this image may be any imaging device including an optical system or any device having a function of generating or converting a virtual image. In the example of FIG. 8, all the detected human faces are processed according to an image conversion algorithm 820 by software or hardware, and at least a part of the distortion is removed. The image conversion algorithm 820 may be executed by any hardware configured to convert the contour of the image distortion, such as a computer equipped with a processor for executing a software algorithm, an ASIC, an FPGA, etc. By removing, correcting, changing, or processing the distortion of the original image 810 by the image conversion algorithm 820, face images 830, 840, 850 without distortion are obtained. The images 830, 840, 850 are then input into a normal neural network 860 trained to use images without deliberately controlled distortion. The output is the result 870. In this example, the result 870 from the neural network 860 indicates the success of recognizing all the faces. The reason is probably that the distortion in the original image 810 was removed before being input into the neural network. The output of this example is not limited to the recognition result of a human face and may be of any type according to the use of the neural network.

[0045] In some of the other embodiments according to the present invention, the original image before being input into the neural network contains additional information or parameters. These may be written in the metadata of the digital image file, or in visible marks, invisible marks, or watermarks within the image, or may be sent from another source to the neural network. It is also possible to utilize these additional information or parameters to assist the image conversion algorithm or the neural network itself in further improving the results.

[0046] The above drawings and examples all show methods of improving the output results from a neural network by utilizing deliberately controlled distortion. In all of these examples, the imaging device, camera, or lens may have any angle of view between ultra-narrow angle and ultra-wide angle. The neural network may be of any type as long as it has at least an input layer and an output layer. The listing of these examples is not intended to create an exhaustive list nor to limit the scope and spirit of the present invention. It will be understood by those skilled in the art that changes can be made to the above examples and embodiments without departing from the broad concept of the invention. Therefore, it is understood by those skilled in the art that the present invention is not limited to the specific examples or embodiments disclosed, and is intended to cover changes made within the spirit and scope of the present invention as defined in the appended claims.

Claims

1. 1. A method for performing inference processing using an artificial intelligence neural network on at least one input file of digital images with controlled distortion in order to improve the output of said neural network, comprising: a. receiving, by an image conversion device, an input file of a first controlled distorted digital image created by an imaging device; b. converting, by said image conversion device, said input file of digital images into a file of transformed digital images having a second controlled distortion; c) performing inference processing on the transformed digital image file using a neural network formed by an algorithm or software code executed on a computing device; d. outputting, by said neural network, interpretation data derived from said transformed digital image file by said inference process; Including, The first controlled distortion comprises: a region of interest that is a region having at least 10% higher resolution than if the digital image of the input file did not have the first controlled distortion; At least one of the neural network has not been trained using the first controlled distorted image data, but has been trained using the second controlled distorted image data; The inference process is performed without removing the second controlled distortion from the transformed digital image. A method comprising:

2. The interpretation data output from the neural network is an output file of the second controlled distorted digital image; e. performing distortion correction on the digital image of the output file to remove at least a portion of the second controlled distortion. The method of claim 1 further comprising:

3. The method of claim 1, wherein the imaging device is a device that intentionally changes the distortion of a digital image to the first controlled distortion.

4. The imaging device includes at least one camera system, The at least one camera system is comprised of at least one optical system. The method of claim 1.

5. The method of claim 1, wherein the neural network is a machine learning neural network trained by deep learning.

6. The method of claim 1, wherein the neural network is trained using a file of the second controlled distorted digital image.

7. The method of claim 1, wherein the interpretation data output from the neural network is image depth information, object recognition results, object classification results, object segmentation results, optical flow estimation results, connection relationships between edges and lines, SLAM results, or images obtained by super-resolution.

8. The method of claim 1, wherein the first controlled distortion in the digital image of the input file obtained from the imaging device varies with time.

9. The method of claim 1, wherein the input file of the digital image includes ground truth information or information regarding classification.

Citation Information

Patent Citations

  • Program, learning processing method, learning model, data structure, learning device and object recognition device

    JP2019117577A

  • Information processing device, learning processing method, learning device, and object recognition device

    US20190197669A1