SUPER-RESOLUTION SYSTEM TRAINED ON LOW-RESOLUTION VEILED DATA
The super-resolution system addresses the limitations of existing techniques by training neural networks to enhance image resolution and detect obscured objects, improving accuracy and reducing bandwidth requirements.
Patent Information
- Application Number
- DE102024128682
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2024-10-04
- Publication Date
- 2026-02-12
AI Technical Summary
Existing super-resolution techniques fail to enhance image resolution sufficiently for applications like crowdsourced mapping and cannot effectively handle obscured objects in vehicle camera data.
A super-resolution system utilizing super-resolution neural networks trained on paired low-resolution and high-resolution image data, with focused loss models to prioritize and reconstruct obscured objects, improving resolution and detection accuracy.
Enhances image resolution and improves object detection accuracy, particularly for obscured objects, reducing the need for oversampling and wider communication bandwidth.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
INTRODUCTION
[0001] The present disclosure relates to a super-resolution system for increasing the resolution of image data acquired by one or more cameras. The super-resolution model comprises one or more super-resolution neural networks trained on the basis of paired training data containing obfuscated low-resolution image data and high-resolution image data.
[0002] A vehicle can use various types of perceptual sensors to collect perceptual data about its environment. One particular type of perceptual sensor commonly used in vehicles is a camera, which collects image data of the surroundings. This image data, representing the environment, can be used in a variety of vehicle systems and applications, such as crowdsourced mapping. Crowdsourced mapping involves collecting perceptual data from numerous connected vehicles, which is then used to create and update maps.
[0003] The image data captured by many cameras in vehicles typically has a relatively low resolution and limited color information. This lower resolution can cause problems when an object detection system attempts to capture and interpret certain types of objects, such as traffic signs. This problem can be exacerbated if the traffic sign is located at a distance or is obscured. For example, a traffic sign might be obscured by vegetation, weather conditions like rain and fog, graffiti, or objects near the sign, such as poles and surrounding vehicles. When the object detection system is used for crowdsourcing mapping, the camera may perform oversampling in certain geographic areas to ensure sufficient image data is available for accurate extraction of traffic sign features.However, oversampling data requires longer data campaigns and a wider communication bandwidth.
[0004] Currently, super-resolution imaging techniques exist to improve or increase the resolution and / or frame rate of image data. However, existing super-resolution techniques may not be able to achieve the resolution improvements required for some applications, such as crowdsourced mapping. Furthermore, existing super-resolution imaging techniques are unable to compensate for obscured objects in the environment.
[0005] While current object detection systems fulfill their purpose, there is a need for an improved approach to increasing the resolution of image data captured by a camera. SUMMARY
[0006] According to several aspects, a super-resolution system is disclosed that increases the resolution of image data captured by one or more cameras. The super-resolution system comprises one or more controllers with one or more super-resolution neural networks containing a model for obfuscated image data. The one or more controllers contain one or more processors that execute instructions for receiving paired training data by the model for obfuscated image data during a training phase. The paired training data is representative of the image data captured by the one or more cameras representing an environment and includes low-resolution and high-resolution obfuscated image data.The low-resolution and high-resolution obfuscated image data represent identical images, and the low-resolution obfuscated image data contains an object of interest located in the environment that has been obfuscated using an obfuscation technique. One or more controllers upscale the resolution of the low-resolution obfuscated image data using the obfuscated image model to produce a reconstructed high-resolution image. The one or more controllers then calculate the total loss associated with the reconstructed high-resolution image using the high-resolution image data from the paired training data as trusted input. The obfuscated image model is trained iteratively to minimize this total loss.One or more controllers receive low-resolution real image data through the obfuscated image data model during a test phase. The one or more controllers then upscale the resolution of this low-resolution real image data using the obfuscated image data model to generate high-resolution real image data.
[0007] In another aspect, the total loss associated with the reconstructed high-resolution image is a sum of a loss of mean squared error, a loss of perception, an opponent loss, and a total variance loss.
[0008] In another aspect, one or more super-resolution neural networks include a model for focused loss.
[0009] In one aspect, one or more controllers execute instructions to: receive the paired training data through the focused loss model during the training phase; increase the resolution of the obfuscated, low-resolution image data through the focused loss model to generate a focused, high-resolution reconstructed image; calculate a focused loss associated with the focused, high-resolution reconstructed image through the focused loss model, where the high-resolution image data of the paired training data serves as trusted input data and the focused loss is a sum of a focused mean squared error loss, a focused perception loss, and a focused total variance loss; and receive real, low-resolution image data through the focused loss model in a test phase.and increasing the resolution of the real low-resolution image data through the focused loss model to generate real high-resolution image data.
[0010] In another aspect, one or more controllers execute instructions to determine a bounding box that defines a limited area within an image frame of the obfuscated, low-resolution image data, where the bounding box contains the object of interest.
[0011] In another aspect, the focused loss associated with the focused reconstructed high-resolution image has a higher value for a limited weighting factor corresponding to the limited area of the image frame than for a total weighting factor corresponding to the entirety of the image frame.
[0012] In one aspect, the one or more controllers determine the focused mean squared loss by: determining the mean squared loss associated with the limited area of the image frame, and determining the mean squared loss associated with the entirety of the image frame, wherein the focused mean squared loss is the sum of a weighted mean squared loss associated with the limited area within the image frame and a weighted mean squared loss associated with the entirety of the image frame.
[0013] In another aspect, the one or more controllers determine the focused perceptual loss by: determining a focused perceptual loss associated with the limited area of the image frame, and determining a focused perceptual loss associated with the entirety of the image frame, where the focused perceptual loss is the sum of a weighted perceptual loss associated with the limited area within the image frame and a weighted perceptual loss associated with the entirety of the image frame.
[0014] In another aspect, the one or more controllers determine the focused total variance loss by: determining a focused total variance loss associated with the limited area of the image frame, and determining a focused total variance loss associated with the entirety of the image frame. The focused total variance loss is the sum of a weighted focused total variance loss associated with the limited area within the image frame and a weighted focused total variance loss associated with the entirety of the image frame.
[0015] In one aspect, the focused loss of mean squared error, the focused loss of perception, and the focused loss of total variance each contain different values for the limited weighting factor and the total weighting factor.
[0016] In another aspect, the object of interest is one of the following: a traffic sign, a pedestrian, a cyclist, an animal, a road sign, an advertising board, a commercial sign, a vehicle in the vicinity, and an infrastructure element.
[0017] In another aspect, the obfuscated image data with low resolution has a resolution of 480 x 640 pixels or less, and the image data with high resolution has a resolution of more than 480 x 640 pixels.
[0018] In one aspect, the obfuscation technique includes one of the following possibilities: deleting part of the object of interest, randomly removing pixels representing the object of interest, blurring the object of interest, and darkening the image data associated with the object of interest.
[0019] In another aspect, a super-resolution system is disclosed that increases the resolution of image data captured by one or more cameras. The super-resolution system comprises one or more controllers with one or more super-resolution neural networks, each containing a focused loss model. The controller(s) contain one or more processors that execute instructions for the focused loss model to receive paired training data during a training phase. This paired training data is representative of the image data captured by the one or more cameras, representing an environment, and includes obfuscated low-resolution and high-resolution image data.The low-resolution obfuscated image data and the high-resolution image data both represent identical images, and the low-resolution obfuscated image data contains an object of interest located in the environment that has been obfuscated using an obfuscation technique. One or more controllers upscale the resolution of the low-resolution obfuscated image data using the focused loss model to produce a focused, high-resolution reconstructed image. One or more controllers then compute the focused loss associated with the focused, high-resolution reconstructed image using the focused loss model.The high-resolution image data of the paired training data serve as reliable input data, and the focused loss is the sum of a focused loss of mean squared error, a focused loss of perception, and a focused total variance loss. One or more controllers receive low-resolution real-world image data through the obfuscated image data model in a test phase. The one or more controllers then upscale the resolution of the low-resolution real-world image data using the obfuscated image data model to generate high-resolution real-world image data.
[0020] In another aspect, one or more controllers execute instructions to determine a bounding box that defines a limited area within an image frame of the obfuscated, low-resolution image data, where the bounding box contains the object of interest.
[0021] In another aspect, the focused loss associated with the focused reconstructed high-resolution image has a higher value for a limited weighting factor corresponding to the limited area of the image frame than for a total weighting factor corresponding to the entirety of the image frame.
[0022] In one aspect, the one or more controllers determine the focused mean squared loss by: determining the mean squared loss associated with the limited area of the image frame, and determining the mean squared loss associated with the entirety of the image frame, wherein the focused mean squared loss is the sum of a weighted mean squared loss associated with the limited area within the image frame and a weighted mean squared loss associated with the entirety of the image frame.
[0023] In another aspect, the one or more controllers determine the focused perceptual loss by: determining a focused perceptual loss associated with the limited area of the image frame, and determining a focused perceptual loss associated with the entirety of the image frame, where the focused perceptual loss is the sum of a weighted perceptual loss associated with the limited area within the image frame and a weighted perceptual loss associated with the entirety of the image frame.
[0024] In another aspect, the one or more controllers determine the focused total variance loss by: determining a focused total variance loss associated with the limited area of the image frame, and determining a focused total variance loss associated with the entirety of the image frame, where the focused total variance loss is the sum of a weighted focused total variance loss associated with the limited area within the image frame and a weighted focused total variance loss associated with the entirety of the image frame.
[0025] In one aspect, the focused loss of mean squared error, the focused loss of perception, and the focused loss of total variance each contain different values for the limited weighting factor and the total weighting factor.
[0026] Further areas of application will become apparent from the description given here. It is understood that the description and the specific examples serve only for illustration and are not intended to limit the scope of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings described here are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. Fig. Figure 1 is a schematic diagram of the disclosed super-resolution system in a vehicle with one or more controllers that are in electronic communication with one or more cameras, according to an exemplary embodiment; Fig. Figures 2A - 2C show examples of an object of interest that is part of obfuscated low-resolution image data according to an exemplary embodiment, wherein the object of interest is obfuscated according to an obfuscation technique; Fig. 3 is a block diagram of the software architecture of one or more in Fig. 1 controller shown according to an exemplary embodiment; and Fig. Figure 4 shows an exemplary image frame of the obfuscated, low-resolution image data, according to an exemplary embodiment. DETAILED DESCRIPTION
[0028] The following description is for illustrative purposes only and is not intended to limit the present disclosure, its application or uses.
[0029] In Fig. Figure 1 shows an exemplary vehicle 10 with the disclosed super-resolution system 12, which increases the resolution of the image data captured by one or more cameras 24. The vehicle 10 can be any type of vehicle, such as a sedan, a truck, a sport utility vehicle (SUV), a van, or a motorhome. In the Fig. In the embodiment shown in Figure 1, the super-resolution system 12 comprises one or more controllers 20 that are in electronic communication with a plurality of perception sensors 22 that collect perception data representative of an environment. In the embodiment shown in Fig. In the non-limiting embodiment shown in Figure 1, the plurality of perception sensors 22 includes one or more cameras 24 for capturing image data, an inertial measurement unit (IMU) 26, a global positioning system (GPS) 28, radar 30, and LiDAR 32; however, other or additional perception sensors may also be used. The one or more cameras 24 are positioned to capture image data depicting the environment outside the vehicle 10.
[0030] In one embodiment, the one or more controllers 20 communicate wirelessly with one or more computers 38 in a back-end office 40, where the one or more computers 38 receive the perception data acquired by the one or more perception sensors 22. In a non-limiting embodiment, the one or more computers 38 are part of a crowdsourcing mapping system that collects perception data from numerous connected vehicles to create and update maps. In the described embodiment, the super-resolution system 12 is part of an object detection system for a vehicle. However, it is clear that the super-resolution system 12 is not limited to object detection systems for a vehicle, and the disclosed super-resolution system 12 can also be used in other applications, such as improving image quality in various types of media (e.g.,Photos, videos, and digital applications), in medical imaging (e.g., magnetic resonance imaging (MRI), computed tomography (CT), and microscopy), in satellite imaging, in the restoration of archival videos, and in fine art, computer vision, and industrial inspection systems. Although the super-resolution system 12 is shown as part of an object detection system in the vehicle 10, the super-resolution system 12 can also be part of other vehicle systems, such as automated driving systems (ADS), advanced driver assistance systems (ADAS), navigation systems, and dashcam systems.
[0031] The image data captured by the one or more cameras 24 contain an object 34 of interest located in the environment, which is identified by the object detection system. In the Fig. In the embodiment shown in Figure 1, the object 34 of interest is a traffic sign 36, in particular a stop sign. Although in Fig. 1. If the object of interest (34) is represented as a stop sign, then the object of interest (34) can be any type of object identified by the object detection system, such as a pedestrian, a cyclist, an animal, a road sign, a billboard, a commercial sign, a vehicle in the vicinity, and an infrastructure feature. Examples of infrastructure features include mailboxes, traffic lights, and traffic cones used in road construction.
[0032] It is evident that in some cases the object 34 of interest, located in the vicinity of vehicle 10, can be obscured or concealed. For example, the object 34 of interest can be obscured by vegetation, weather conditions such as rain and fog, graffiti covering part or all of the object 34 of interest, or by surrounding objects near the object 34 of interest, such as masts and nearby vehicles. As explained below, the disclosed super-resolution system 12 is based on paired training data 56 ( Fig. 3) trained to process the obfuscated image data 60 with low resolution ( Fig. 3) included. The low-resolution obfuscated image data is a simulation of the real event when the object 34 of interest is obscured or covered. In particular, the low-resolution obfuscated image data obscures part of the object 34 of interest based on an obfuscation technique.
[0033] Examples of obfuscation techniques (see Fig. 2A - 2C) include, among other things, blocking or deleting part of the object of interest 34 (see Fig. 2A), the random removal of pixels representing the object of interest 34 (see Fig. 2B), blurring the object of interest 34 (see Fig. 2C) and darkening the image data associated with the object 34 of interest (not shown). Blocking or deleting part of the object 34 of interest reproduces real-world cases where the object 34 of interest is obscured by objects such as other vehicles in the vicinity, vegetation such as tree branches, and poles such as light poles commonly found in parking lots. Randomly removing pixels from the object 34 of interest reproduces real-world cases where the object 34 of interest is obscured by elements such as vegetation like bushes, graffiti, stickers affixed to the sign, damage such as cracks or bullet holes, fog, and poles. In one embodiment, randomly removing pixels representing the object 34 of interest may involve removing approximately 25 to approximately 75 percent of the pixels representing the object 34 of interest.Blurring the object of interest (OFI) reproduces real-world cases where OFI is obscured by factors such as bad or harsh weather like rain or fog. Furthermore, darkening the image data can also represent real-world cases where OFI is obscured by bad weather.
[0034] Fig. 3 is a block diagram of the software architecture of one or more in Fig. 1 Controller 20 shown, wherein the one or more controllers 20 comprise one or more super-resolution neural networks 50 and an object detection module 58. In the Fig. In the embodiment shown in Figure 3, the one or more super-resolution networks 50 comprise a model 52 for obfuscated image data and a model 54 for focused loss. While in Fig. 3 where two super-resolution neural networks 50 are shown, in an alternative embodiment the one or more controllers 20 may instead contain only the model 52 for obfuscated image data or the model 54 for focused loss.
[0035] The one or more controllers 20 first undergo a training phase in which they receive the paired training data 56. The paired training data 56 is representative of the image data depicting the environment of the vehicle 10, including the object of interest 34, which is captured by the one or more cameras 24. It comprises the low-resolution obfuscated image data 60 and the high-resolution image data 62, where the low-resolution obfuscated image data 60 and the high-resolution image data 62 represent identical images. The low-resolution obfuscated image data 60 has a resolution of 480 x 640 pixels or less, and the high-resolution image data 62 has a resolution of more than 480 x 640 pixels. It should be noted that the object 34 of interest within the low-resolution obfuscated image data 60 is identified based on one of the parameters specified in the Fig. The obfuscation techniques shown in Figures 2A-2C are used to simulate real-world processes in which the object 34 of interest is obfuscated. In contrast to the obfuscated low-resolution image data 60, the object 34 of interest is visible and unobfuscated in the high-resolution image data 62. Accordingly, one or more super-resolution neural networks 50 can map the obfuscated low-resolution image data 60 to the high-resolution image data 62 to reconstruct the object 34 of interest when a reconstructed high-resolution image is generated. It should also be noted that the object 34 of interest contained in the reconstructed high-resolution image determined by the super-resolution neural networks 50 is unobfuscated and fully visible.
[0036] The model 52 for obfuscated image data is any super-resolution neural network that increases the resolution of the image data from a low resolution to a high resolution, e.g., a super-resolution generative adversarial network (SRGAN) and a fast super-resolution convolutional neural network (FSRCNN). During the training phase, the model 52 for obfuscated image data receives the paired training data 56 as input and increases the resolution of the low-resolution obfuscated image data 60 to generate a reconstructed high-resolution image. The model 52 for obfuscated image data is specifically trained to increase the resolution of the object of interest 34 ( Fig. 1) That is, the model 52 for obfuscated image data is specifically trained to increase the resolution of objects that are of the same type or classification as the object 34 of interest. For example, if the object 34 of interest is classified as a stop sign, the model 52 for obfuscated image data is specifically trained to increase the resolution of objects classified as stop signs. In addition to increasing the resolution, the model 52 for obfuscated image data is specifically trained to reconstruct the object 34 of interest, which is obfuscated in the low-resolution obfuscated image data 60, in the reconstructed high-resolution image. As mentioned earlier, the object 34 of interest contained in the reconstructed high-resolution image determined by the model 52 for obfuscated image data is not obfuscated and is fully visible.
[0037] The model 52 for obfuscated image data calculates a total loss associated with the reconstructed high-resolution image, where the high-resolution image data 62 of the paired training data 56 serve as trusted input data. The total loss associated with the reconstructed high-resolution image is determined by calculating a loss (L2) of mean squared error, a perceptual loss, an adversarial loss, and a total variance loss associated with the reconstructed high-resolution image. The total loss associated with the reconstructed high-resolution image is the sum of the loss of mean squared error, a perceptual loss, an adversarial loss, and a total variance loss, or total loss = loss of mean squared error + perceptual loss + adversarial loss + total variance loss.The model of the obfuscated image data 52 is then trained in an iterative process to minimize the overall loss associated with the reconstructed high-resolution image.
[0038] The focused loss model 54 is any type of super-resolution neural network that increases the resolution of image data from low resolution to high resolution, such as an SRGAN or an FSRCNN. During the training phase, the focused loss model 54 receives the paired training data 56 as input and increases the resolution of the obfuscated, low-resolution image data 60 to produce a focused, high-resolution reconstructed image. The focused loss model 54 is specifically trained to achieve the resolution of the object 34 of interest ( Fig. 1) to increase and reconstruct the object 34 of interest, which is obscured in the obscured low-resolution image data 60, in the reconstructed high-resolution image. The focused loss model 54 computes a focused loss associated with the focused reconstructed high-resolution image, using the high-resolution image data 62 of the paired training data 56 as trusted input data. The focused loss model 54 is trained on the basis of an iterative process to minimize the focused loss associated with the focused reconstructed high-resolution image.
[0039] Fig. Figure 4 shows an example image frame 70 of the obfuscated low-resolution image data 60. As in the Fig. 3 and Fig. As shown in Figure 4, the model 54 for focused loss determines a bounding frame 72 that defines a bounded area 74 within the image frame 70 containing the object of interest 34, which has been obscured based on one of the obfuscation techniques described above. The focused loss associated with the focused reconstructed high-resolution image assigns a greater weight to the bounded area 74 within the bounding frame 72 than to the entirety of the image frame 70 containing the obfuscated low-resolution image data 60. In particular, the focused loss associated with the focused reconstructed high-resolution image assigns a higher value to a bounded weighting factor corresponding to the bounded area 74 of the image frame 70 than to an overall weighting factor corresponding to the entirety of the image frame 70 containing the obfuscated low-resolution image data 60.The model 54 for focused loss determines the focused loss associated with the focused reconstructed high-resolution image by calculating a focused loss of mean squared error, a focused perception loss, and a focused total variance loss, where the focused loss is the sum of the focused loss of mean squared error, the focused perception loss, and the focused total variance loss.
[0040] The focused mean squared loss is determined by calculating the mean squared loss associated with the entirety of image frame 70, masking the limited area 74 containing the object of interest 34 within image frame 70 of the low-resolution masked image data 60, and calculating the mean squared loss associated with the limited area 74 of image frame 70. The focused mean squared loss is the sum of a weighted mean squared loss associated with the limited area 74 within image frame 70 and a weighted mean squared loss associated with the entirety of image frame 70.
[0041] The weighted loss of mean square error associated with the limited area 74 within image frame 70 is determined by multiplying the loss of mean square error associated with the limited area 74 within image frame 70 by the limited weighting factor, or (A * loss of mean square error associated with the limited area 74 within image frame 70), where A represents the limited weighting factor. The weighted loss of mean square error associated with the entirety of image frame 70 is determined by multiplying the loss of mean square error associated with the entirety of image frame 70 by the total weighting factor, or (B * loss of mean square error associated with the entirety of image frame 70), where B represents the total weighting factor.
[0042] The limited weighting factor A is greater than the total weighting factor B, or A > B, and the sum of the limited weighting factor A and the total weighting factor B is equal to one, or A + B = 1. By way of example only, in one embodiment the limited weighting factor A is equal to 0.8 and the total weighting factor is equal to 0.2. Therefore, the limited area 74 within image frame 70 containing the object 34 of interest is given more weight compared to the entire image frame 70, which improves the focused loss ability of the model 54 to reconstruct the object 34 of interest within the focused reconstructed high-resolution image.
[0043] The focused perceptual loss is determined by calculating the perceptual loss associated with the entirety of image frame 70, masking the limited area 74 of image frame 70 from the obfuscated, low-resolution image data 60, and determining the perceptual loss associated with this limited area 74. The focused perceptual loss is the sum of the weighted perceptual loss associated with the limited area 74 within image frame 70 and the weighted perceptual loss associated with the entirety of image frame 70.
[0044] The weighted perceptual loss associated with the limited area 74 within image frame 70 is determined by multiplying the perceptual loss associated with the limited area 74 within image frame 70 by the limited weighting factor, or (A * perceptual loss associated with the limited area 74 within image frame 70). The weighted perceptual loss associated with the entirety of image frame 70 is determined by multiplying the perceptual loss associated with the entirety of image frame 70 by the total weighting factor, or (B * perceptual loss associated with the entirety of image frame 70).In some embodiments, the limited weighting factor used to determine the weighted perceptual loss can be a different value than the limited weighting factor used to determine the weighted loss of mean squared error. Likewise, the total weighting factor used to determine the weighted perceptual loss can be a different value than the total weighting factor used to determine the weighted loss of mean squared error. Accordingly, the focused loss of mean squared error, the focused perceptual loss, and the focused total variance loss can be prioritized by assigning different values to the limited weighting factor and the total weighting factor for each distinct type of loss.
[0045] The focused total variance loss is determined by calculating the total variance loss associated with the entirety of image frame 70, masking the limited area 74 of image frame 70 from the obfuscated, low-resolution image data 60, and determining the total variance loss associated with the limited area 74 of image frame 70. The focused total variance loss is the sum of the weighted total variance loss associated with the limited area 74 within image frame 70 and the weighted total variance loss associated with the entirety of image frame 70.
[0046] The weighted total variance loss associated with the limited area 74 within image frame 70 is determined by multiplying the total variance loss associated with the limited area 74 within image frame 70 by the limited weighting factor, or (A * total variance loss associated with the limited area 74 within image frame 70). The weighted total variance loss associated with the entirety of image frame 70 is determined by multiplying the total variance loss associated with the entirety of image frame 70 by the total weighting factor, or (B * total variance loss associated with the entirety of image frame 70).In embodiments, the limited weighting factor used to determine the weighted total variance loss can be a different value than the limited weighting factor used to determine the weighted loss of mean squared error and the weighted loss of perception. Likewise, the total weighting factor used to determine the weighted total variance loss can be a different value than the total weighting factor used to determine the weighted loss of mean squared error and the weighted loss of perception. Accordingly, the focused loss of mean squared error, the focused loss of perception, and the focused total variance loss can be prioritized by assigning different values to the limited weighting factor and the total weighting factor for each distinct type of loss.
[0047] As in Fig. As shown in Figure 3, after training the one or more super-resolution neural networks 50, the one or more controllers 20 can be subjected to a test phase. During the test phase, the one or more controllers 20 receive real, low-resolution image data 64 representing the environment, including the object of interest 34, which was captured by the one or more cameras 24. During the test phase, the model 52 for obfuscated image data, the model 54 for focused loss, or both the model 52 for obfuscated image data and the model 54 for focused loss, receive the real, low-resolution image data 64 as input and upscale the resolution of the real, low-resolution image data 64 to generate real, high-resolution image data 66.
[0048] The object detection module 58 receives the real high-resolution image data 66 as input and executes one or more object detection algorithms to determine an instance of the object of interest 34 within the real high-resolution image data 66. The object detection module 58 can execute any type of object detection algorithm, such as the YOLO (You-Only-Look-Once) algorithm, without being limited to it. It should be noted that the real high-resolution image data 66, determined by the model 52 for obfuscated image data and the model 54 for focused loss, result in improved object detection accuracy of the object of interest 34 compared to high-resolution images determined by a standard image data model that is not trained on obfuscated low-resolution image data that obscures the object of interest 34.In a non-restrictive example, the real high-resolution image data 66 determined by the model 52 for obfuscated image data result in an object detection accuracy of the object of interest 34 of approximately 59%, and the real high-resolution image data 66 determined by the model 54 for focused loss result in an object detection accuracy of the object of interest 34 of approximately 64%. In contrast, a standard image data model not trained on low-resolution obfuscated image data may result in an object detection accuracy of only about 40%. In the present example, all test data were obtained from the same dataset of 1225 images of stop signs.
[0049] With general reference to the figures, the disclosed super-resolution system offers various technical effects and advantages. The super-resolution system employs a tailored approach to train super-resolution neural networks based on low-resolution image data that obscures the object of interest. The super-resolution neural networks are specifically trained to increase the resolution of the object of interest. The obscured objects in the low-resolution image data simulate real-world processes where the object of interest is disguised within its environment. Therefore, the super-resolution system improves the detectability of objects in both low-resolution images and images containing obscured or disguised objects. Furthermore, the disclosed super-resolution system also maximizes the performance of cameras capturing lower-resolution image data.
[0050] Controllers can refer to or be part of an electronic circuit, a combinational logic circuit, a field-programmable gate array (FPGA), a (shared, dedicated, or grouped) processor that executes code, or a combination of some or all of the above, such as in a system-on-a-chip. Furthermore, controllers can be microprocessor-controlled, such as a computer with at least one processor, memory (RAM and / or ROM), and associated input and output buses. The processor can operate under the control of an operating system residing in memory. The operating system can manage computer resources so that the computer program code, embodied as one or more computer software applications (e.g., an application residing in memory), can direct instructions from the processor to be executed.In an alternative embodiment, the processor can execute the application directly; in this case, the operating system can be omitted.
[0051] The description of the present revelation is merely exemplary, and variations that do not deviate from the core of the present revelation shall fall within its scope of protection. Such variations are not to be considered a deviation from the spirit and scope of the present revelation.
Claims
[1] Super-resolution system that increases the resolution of image data captured by one or more cameras, the super-resolution system comprising: one or more controllers with one or more super-resolution neural networks containing a model for obfuscated image data, wherein the one or more controllers contain one or more processors that execute instructions to: Receiving paired training data by the obfuscated image data model during a training phase, wherein the paired training data are representative of the image data captured by the one or more cameras representing an environment and include low-resolution obfuscated image data and high-resolution image data, and wherein the low-resolution obfuscated image data and the high-resolution image data both represent identical images, and the low-resolution obfuscated image data includes an object of interest located in the environment and obfuscated based on an obfuscation technique; Increasing the resolution of the low-resolution obfuscated image data through the model of the obfuscated image data to generate a reconstructed high-resolution image; Calculating the total loss associated with the reconstructed high-resolution image by the obfuscated image data model, where the high-resolution image data of the paired training data serve as trusted input data and the obfuscated image data model is trained in an iterative process to minimize the total loss; Receiving real-world, low-resolution image data through the obfuscated image data model in a test phase; and Increasing the resolution of real low-resolution image data through the obfuscated image data model to generate real high-resolution image data. [2] Super-resolution system according to claim 1, wherein the total loss associated with the reconstructed high-resolution image is a sum of a mean squared error loss, a perception loss, an opponent loss and a total variance loss. [3] Super-resolution system according to claim 1, wherein the one or more super-resolution neural networks comprise a focused loss model. [4] Super-resolution system according to claim 3, wherein one or more controllers execute instructions to: Receiving the paired training data by the focused loss model during the training phase; Increasing the resolution of the obscured, low-resolution image data through the focused loss model to produce a focused, high-resolution reconstructed image; Calculating a focused loss associated with the focused high-resolution reconstructed image using the focused loss model, where the high-resolution image data of the paired training data serve as trusted input data and the focused loss is a sum of a focused mean squared error loss, a focused perception loss, and a focused total variance loss; Receiving real-world, low-resolution image data through the focused loss model during a test phase; and Increasing the resolution of real low-resolution image data through the focused loss model to generate real high-resolution image data. [5] Super-resolution system according to claim 4, wherein one or more controllers execute instructions to: Determining a bounding box that defines a limited area within an image frame of the obfuscated, low-resolution image data, where the bounding box contains the object of interest. [6] Super-resolution system according to claim 5, wherein the focused loss associated with the focused reconstructed high-resolution image assigns a higher value to a limited weighting factor corresponding to the limited area of the image frame than to a total weighting factor corresponding to the entirety of the image frame. [7] Super-resolution system according to claim 6, wherein the one or more controllers determine the focused loss of the mean squared error by: Determining the mean squared error loss associated with the limited area of the image frame; and Determining the mean squared error loss associated with the entirety of the image frame, wherein the focused mean squared error loss is the sum of a weighted mean squared error loss associated with the limited area within the image frame and a weighted mean squared error loss associated with the entirety of the image frame. [8] Super-resolution system according to claim 6, wherein the one or more controllers determine the focused perceptual loss by: Determining a focused perceptual loss associated with the limited area of the image frame; and Determining a focused perceptual loss associated with the entirety of the image frame, wherein the focused perceptual loss is the sum of a weighted perceptual loss associated with the limited area within the image frame and a weighted perceptual loss associated with the entirety of the image frame. [9] Superresolution system according to claim 6, wherein the one or more controllers determine the focused total variance loss by: Determining a focused total variance loss associated with the limited area of the image frame; and Determining a focused total variance loss associated with the entirety of the image frame, wherein the focused total variance loss is the sum of a weighted focused total variance loss associated with the bounded area within the image frame and a weighted focused total variance loss associated with the entirety of the image frame. [10] Super-resolution system according to claim 6, wherein the focused loss of mean squared error, the focused perception loss and the focused total variance loss each contain different values for the limited weighting factor and the total weighting factor.
Citation Information
Patent Citations
Information processing apparatus
US20210327028A1