Super-resolution system based on confusion low-resolution data training

By training a super-resolution neural network and utilizing a confused image data model and a focus loss model, the problem of insufficient resolution of camera image data was solved, thereby improving the accuracy of object detection and the precision of crowdsourced maps.

CN121504729APending Publication Date: 2026-02-10GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411257282.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-07
Filing Date
2024-09-09
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing super-resolution technologies cannot effectively improve the resolution of image data captured by cameras, especially when objects in the environment are occluded or obscured, and cannot meet the needs of applications such as crowdsourced maps.

Method used

A super-resolution neural network is trained using a scrambled image data model and a focus loss model. By pairing scrambled low-resolution image data with high-resolution image data for training, the image resolution is improved. Loss functions such as mean squared error loss, perceptual loss, and total variance loss are used to optimize the training process, specifically training the resolution of the object of interest.

Benefits of technology

It significantly improves the accuracy of object detection, especially when objects are occluded or obfuscated, enhances the resolution of image data and the reliability of object detection, and improves the accuracy of crowdsourced maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504729A_ABST
    Figure CN121504729A_ABST
Patent Text Reader

Abstract

A super-resolution system for improving resolution of image data captured by one or more cameras includes one or more controllers including one or more super-resolution neural networks including at least one of an obfuscated image data model and a focus loss model. The one or more super-resolution neural networks receive pairing training data during a training phase, where the pairing training data represents image data captured by the one or more cameras, the image data representing an ambient environment and including low-resolution image data and high-resolution image data. The obfuscated low-resolution image data and high-resolution image data each represent the same image, the obfuscated low-resolution image data including an object of interest located in the surrounding environment, the object of interest being obfuscated based on an obfuscation technique.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a super-resolution system for improving the resolution of image data captured by one or more cameras. The super-resolution model includes one or more super-resolution neural networks trained on paired training data, which includes mixed low-resolution image data and high-resolution image data. Background Technology

[0002] Vehicles can use various types of perception sensors to collect perception data related to their surroundings. One specific type of perception sensor commonly used in vehicles is the camera, which collects image data about the surrounding environment. This image data representing the surrounding environment can be used in various vehicle systems and applications, such as crowdsourced mapping. Crowdsourced mapping involves collecting perception data from numerous connected vehicles, where the perception data is used to generate and update maps.

[0003] It should be understood that image data collected by numerous cameras within a vehicle often has relatively low resolution and limited color information. This low-resolution image data can pose challenges when object detection systems attempt to detect and interpret certain types of objects, such as traffic signs. This problem can be further exacerbated when traffic signs are distant or occluded. For example, traffic signs can be confused with vegetation, weather conditions such as rain or fog, graffiti, or objects like utility poles and surrounding vehicles near the traffic sign. If the object detection system is used for crowdsourced mapmaking, the cameras can oversample certain geographic areas to ensure sufficient image data for accurate traffic sign feature extraction. However, oversampled data requires longer data transmission times and greater communication bandwidth.

[0004] Super-resolution imaging techniques exist to enhance or improve the resolution and / or frame rate of image data. However, existing super-resolution techniques may not provide the resolution improvements required for certain applications, such as crowdsourced maps. Furthermore, existing super-resolution imaging techniques cannot compensate for occluded objects in the environment.

[0005] Therefore, while current object detection systems have achieved their intended purpose, there is still a need in the art for an improved method to increase the resolution of image data captured by cameras. Summary of the Invention

[0006] According to several aspects, a system for a super-resolution system is disclosed that improves the resolution of image data captured by one or more cameras. The super-resolution system includes one or more controllers, each controller including one or more super-resolution neural networks, the super-resolution neural networks including a scrambled image data model. The one or more controllers include one or more processors that execute instructions to receive paired training data via the scrambled image data model during a training phase. The paired training data represents image data captured by the one or more cameras, representing a surrounding environment and including scrambled low-resolution image data and high-resolution image data. Both the scrambled low-resolution image data and the high-resolution image data represent the same image, and the scrambled low-resolution image data includes objects of interest located in the surrounding environment, which are scrambled based on scrambling techniques. The one or more controllers improve the resolution of the scrambled low-resolution image data via the scrambled image data model to create a reconstructed high-resolution image. The one or more controllers compute a total loss associated with the reconstructed high-resolution image via the scrambled image data model, wherein the high-resolution image data of the paired training data is used as ground truth data, and the scrambled image data model is trained based on an iterative process to minimize the total loss. During a testing phase, the one or more controllers receive real low-resolution image data via the scrambled image data model. One or more controllers increase the resolution of real low-resolution image data by obfuscating the image data model to create real high-resolution image data.

[0007] On the other hand, the total loss associated with the reconstructed high-resolution image is the sum of mean squared error loss, perceptual loss, adversarial loss, and total variance loss.

[0008] On the other hand, one or more super-resolution neural networks include a focusing loss model.

[0009] On one hand, one or more controllers execute instructions to receive paired training data via a focus loss model during the training phase, to improve the resolution of confused low-resolution image data via the focus loss model to create a focus-reconstructed high-resolution image, and to calculate the focus loss associated with the focus-reconstructed high-resolution image via the focus loss model, wherein the high-resolution image data of the paired training data is used as ground reality data, and the focus loss is the sum of the focus mean square error loss, the focus perception loss, and the focus total variance loss. On the testing phase, real low-resolution image data is received via the focus loss model, and the resolution of the real low-resolution image data is improved via the focus loss model to create real high-resolution image data.

[0010] On the other hand, one or more controllers execute instructions to determine bounding boxes of bounded regions within an image frame that defines obfuscated low-resolution image data, wherein the bounding boxes contain objects of interest.

[0011] On the other hand, the focusing loss associated with the high-resolution image of the focused reconstruction assigns a higher value to the bounded weighting factor corresponding to the bounded region of the image frame compared to the overall weighting factor corresponding to the entire image frame.

[0012] On one hand, one or more controllers determine the focus mean square error loss by determining the mean square error loss associated with a bounded region of the image frame and the mean square error loss associated with the entire image frame, wherein the focus mean square error loss is the sum of the weighted mean square error loss associated with the bounded region within the image frame and the weighted mean square error loss associated with the entire image frame.

[0013] On the other hand, one or more controllers determine the focus perception loss by determining the focus perception loss associated with a bounded region of the image frame and the focus perception loss associated with the entire image frame, wherein the focus perception loss is the sum of the weighted perception loss associated with the bounded region within the image frame and the weighted perception loss associated with the entire image frame.

[0014] In another aspect, one or more controllers determine the total focus variance loss by determining the total focus variance loss associated with a bounded region within the image frame and by determining the total focus variance loss associated with the entire image frame. The total focus variance loss is the sum of the weighted total focus variance loss associated with the bounded region within the image frame and the weighted total focus variance loss associated with the entire image frame.

[0015] On the one hand, focusing on mean squared error loss, focusing on perceived loss, and focusing on total variance loss all include different values ​​of bounded weighting factor and overall weighting factor.

[0016] On the other hand, objects of interest include one of the following: traffic signs, pedestrians, cyclists, animals, street signs, billboards, commercial signs, surrounding vehicles, and infrastructure assets.

[0017] On the other hand, obfuscated low-resolution image data includes resolutions of 480x640 pixels or less, while high-resolution image data includes resolutions greater than 480x640 pixels.

[0018] On one hand, obfuscation techniques include one of the following: deleting a portion of an object of interest, randomly removing pixels representing the object of interest, obfuscating the object of interest, and darkening image data associated with the object of interest.

[0019] In another aspect, a super-resolution system for improving the resolution of image data captured by one or more cameras is disclosed. The super-resolution system includes one or more controllers, each including one or more super-resolution neural networks, which in turn include a focus loss model. Each controller includes one or more processors that execute instructions to receive paired training data via the focus loss model during a training phase. The paired training data represents image data captured by the one or more cameras, representing the surrounding environment and including obfuscated low-resolution image data and high-resolution image data. Both the obfuscated low-resolution and high-resolution image data represent the same image, and the obfuscated low-resolution image data includes objects of interest located in the surrounding environment, which are obfuscated using obfuscation techniques. The one or more controllers improve the resolution of the obfuscated low-resolution image data via the focus loss model to create a focus-reconstructed high-resolution image. The one or more controllers compute a focus loss associated with the focus-reconstructed high-resolution image via the focus loss model. The high-resolution image data of the paired training data is used as ground-based real-world data, and the focus loss is the sum of the focus mean square error loss, the focus perception loss, and the total focus variance loss. During a testing phase, the one or more controllers receive real low-resolution image data via the obfuscated image data model. One or more controllers increase the resolution of real low-resolution image data by obfuscating the image data model to create real high-resolution image data.

[0020] On the other hand, one or more controllers execute instructions to determine bounding boxes of bounded regions within an image frame that defines obfuscated low-resolution image data, wherein the bounding boxes contain objects of interest.

[0021] On the other hand, the focusing loss associated with the high-resolution image of the focused reconstruction assigns a higher value to the bounded weighting factor corresponding to the bounded region of the image frame compared to the overall weighting factor corresponding to the entire image frame.

[0022] On one hand, one or more controllers determine the focus mean square error loss by determining the mean square error loss associated with a bounded region of the image frame and the mean square error loss associated with the entire image frame, wherein the focus mean square error loss is the sum of the weighted mean square error loss associated with the bounded region within the image frame and the weighted mean square error loss associated with the entire image frame.

[0023] On the other hand, one or more controllers determine the focus perception loss by determining the focus perception loss associated with a bounded region of the image frame and the focus perception loss associated with the entire image frame, wherein the focus perception loss is the sum of the weighted perception loss associated with the bounded region within the image frame and the weighted perception loss associated with the entire image frame.

[0024] In another aspect, one or more controllers determine the total focus variance loss by determining the total focus variance loss associated with a bounded region of the image frame and the total focus variance loss associated with the entire image frame, wherein the total focus variance loss is the sum of the weighted total focus variance loss associated with the bounded region within the image frame and the weighted total focus variance loss associated with the entire image frame.

[0025] On the one hand, focusing on mean squared error loss, focusing on perceived loss, and focusing on total variance loss all include different values ​​of bounded weighting factor and overall weighting factor.

[0026] Further areas of application will become apparent from the description provided herein. It should be understood that these descriptions and specific examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0027] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure in any way.

[0028] Figure 1 This is a schematic diagram of a super-resolution system in a vehicle disclosed according to an exemplary embodiment, the system including one or more controllers that communicate electronically with one or more cameras;

[0029] Figure 2A-2C An example of an object of interest as part of obfuscated low-resolution image data is shown according to an exemplary embodiment, wherein the object of interest is obfuscated according to an obfuscation technique;

[0030] Figure 3 yes Figure 1 The diagram shown is a block diagram of the software architecture of one or more controllers according to an exemplary embodiment; and

[0031] Figure 4 An exemplary image frame of obfuscated low-resolution image data is shown according to an exemplary embodiment. Detailed Implementation

[0032] The following description is merely exemplary in nature and is not intended to limit this disclosure, application, or use.

[0033] refer to Figure 1 An exemplary vehicle 10 is shown, including the disclosed super-resolution system 12, which enhances the resolution of image data captured by one or more cameras 24. It should be understood that vehicle 10 can be any type of vehicle, such as, but not limited to, a sedan, truck, SUV, van, or motorhome. Figure 1In the illustrated embodiment, the super-resolution system 12 includes one or more controllers 20 that electronically communicate with a plurality of sensing sensors 22 that collect perceived data representing the surrounding environment. Figure 1 In the non-limiting embodiment shown, the plurality of sensing sensors 22 include one or more cameras 24, an inertial measurement unit (IMU) 26, a global positioning system (GPS) 28, a radar 30, and a lidar (LiDAR) 32 for collecting image data; however, it should be understood that different or additional sensing sensors may also be used. One or more cameras 24 are positioned to capture image data representing the external surrounding environment of the vehicle 10.

[0034] In one embodiment, one or more controllers 20 wirelessly communicate with one or more computers 38 in a back-end office 40, wherein one or more computers 38 receive perception data collected by one or more perception sensors 22. In a non-limiting embodiment, one or more computers 38 are part of a crowdsourced mapping system that collects perception data from numerous connected vehicles to generate and update maps. In the described embodiment, the super-resolution system 12 is part of a vehicle's object detection system. However, it should be understood that the super-resolution system 12 is not limited to a vehicle's object detection system, and the disclosed super-resolution system 12 can be used for other applications, such as improving the image quality of various types of media (e.g., photographs, videos, and digital applications), medical imaging (e.g., magnetic resonance imaging (MRI), computed tomography (CT), and microscopy), satellite imaging, restoration of archival videos and works of art, computer vision, and industrial inspection systems. Furthermore, although the super-resolution system 12 is shown as part of an object detection system in vehicle 10, the super-resolution system 12 can also be part of other vehicle systems, such as automated driving systems (ADS), advanced driver assistance systems (ADAS), navigation systems, and dashcam systems.

[0035] It should be understood that the image data captured by one or more cameras 24 includes objects of interest 34 located in the surrounding environment, wherein the objects of interest 34 are identified by an object detection system. Figure 1 In the illustrated embodiment, the object of interest 34 is the traffic sign 36, particularly the stop sign. Although Figure 1 Object of Interest 34 is shown as a stop sign; however, it should be understood that Object of Interest 34 can be any type of object identified by the object detection system, such as, but not limited to, pedestrians, cyclists, animals, street signs, billboards, commercial signs, surrounding vehicles, and infrastructure assets. Some examples of infrastructure assets include, but are not limited to, mailboxes, traffic lights such as traffic lights, and traffic cones used for road construction.

[0036] It should be understood that, in some cases, the object of interest 34 in the environment surrounding vehicle 10 may become blurred. For example, the object of interest 34 may be obscured by vegetation, weather conditions such as rain or fog, graffiti covering part or all of the object of interest 34, or surrounding objects such as utility poles and surrounding vehicles located near the object of interest 34. As explained below, the disclosed super-resolution system 12 is based on low-resolution image data 60, including obscured low-resolution image data 60. Figure 3 ) paired training data 56 ( Figure 3 The obfuscated low-resolution image data is used for training. It replicates the true state of the object of interest 34 when it is obfuscated. Specifically, the obfuscated low-resolution image data uses obfuscation techniques to obfuscate a portion of the object of interest 34.

[0037] refer to Figures 2A to 2C Some examples of obfuscation techniques include, but are not limited to, blocking or deleting a portion of an object of interest 34 (such as...). Figure 2A As shown), randomly remove pixels representing the object of interest 34 (e.g.) Figure 2B As shown), obfuscate object of interest 34 ( Figure 2C As shown in the figure, the image data associated with the object of interest 34 is darkened (not shown). Occluding or removing a portion of the object of interest 34 can recreate the realistic situation where the object of interest 34 is obscured by other objects in the surrounding environment, such as vehicles, vegetation (e.g., tree branches), and utility poles (e.g., lampposts commonly found in parking lots). When the object of interest 34 is obscured by vegetation such as bushes, graffiti, stickers on signs, damage such as cracks or bullet holes, fog, and utility poles, random pixel removal of the object of interest 34 can recreate the realistic situation. In one embodiment, random removal of pixels representing the object of interest 34 may involve removing approximately 25% to approximately 75% of the pixels representing the object of interest 34. Burying the object of interest 34 can recreate the realistic situation where the object of interest 34 is obscured by inclement weather such as rain or fog. In addition, darkening the image data can also recreate the realistic situation where the object of interest 34 is obscured by inclement weather.

[0038] Figure 3 yes Figure 1 The diagram shows a block diagram of the software architecture of one or more controllers 20, wherein the one or more controllers 20 include one or more super-resolution neural networks 50 and object detection modules 58. Figure 3 In the illustrated embodiment, one or more super-resolution networks 50 include a scrambled image data model 52 and a focus loss model 54. Although Figure 3 Two super-resolution neural networks 50 are shown, but it should be understood that in another embodiment, one or more controllers 20 may include only the obfuscated image data model 52 or alternatively include the focus loss model 54.

[0039] One or more controllers 20 first undergo a training phase, during which they receive paired training data 56. The paired training data 56 represents image data depicting the environment surrounding the vehicle 10, including objects of interest 34 captured by one or more cameras 24, and includes obfuscated low-resolution image data 60 and high-resolution image data 62, where both represent the same image. The obfuscated low-resolution image data 60 includes a resolution of 480x640 pixels or less, and the high-resolution image data 62 includes a resolution greater than 480x640 pixels. It should be understood that the objects of interest 34 within the obfuscated low-resolution image data 60 are based on… Figure 2A-2C One of the obfuscation techniques shown is used to perform obfuscation in order to replicate the actual obfuscation of the object of interest 34. It should be understood that, unlike the obfuscated low-resolution image data 60, the object of interest 34 is visible and not obfuscated within the high-resolution image data 62. Therefore, one or more super-resolution neural networks 50 can map the obfuscated low-resolution image data 60 to the high-resolution image data 62 to reconstruct the object of interest 34 when generating the reconstructed high-resolution image. It should also be understood that the object of interest 34 included in the reconstructed high-resolution image determined by the super-resolution neural network 50 is not obfuscated and is fully visible.

[0040] The image data obfuscation model 52 is any type of super-resolution neural network that upscales image data from low to high resolution, such as, but not limited to, Super-Resolution Generative Adversarial Networks (SRGAN) and Fast Super-Resolution Convolutional Neural Networks (FSRCNN). During the training phase, the image data obfuscation model 52 receives paired training data 56 as input and upscales the low-resolution image data 60 to create a reconstructed high-resolution image. The image data obfuscation model 52 is specifically trained to upscale the object of interest 34 (…). Figure 1 The resolution of the obfuscated image data model 52 is improved. In other words, the obfuscated image data model 52 is specifically trained to improve the resolution of objects of the same type or category as the object of interest 34. For example, when the object of interest 34 is classified as a stop sign, the obfuscated image data model 52 is specifically trained to improve the resolution of objects classified as stop signs. In addition to improving resolution, the obfuscated image data model 52 is also specifically trained to reconstruct the obfuscated object of interest 34 in the obfuscated low-resolution image data 60 within the reconstructed high-resolution image. As described above, the object of interest 34 included in the reconstructed high-resolution image determined by the obfuscated image data model 52 is not obfuscated and is fully visible.

[0041] The confused image data model 52 calculates the total loss associated with the reconstructed high-resolution image, where the high-resolution image data 62 of the paired training data 56 is used as ground reality data. The total loss associated with the reconstructed high-resolution image is determined by calculating the mean squared error (L2) loss, perceptual loss, adversarial loss, and total variance loss associated with the reconstructed high-resolution image. The total loss associated with the reconstructed high-resolution image is the sum of the mean squared error loss, perceptual loss, adversarial loss, and total variance loss, or total loss = mean squared error loss + perceptual loss + adversarial loss + total variance loss. The confused image data model 52 is then trained based on an iterative process to minimize the total loss associated with the reconstructed high-resolution image.

[0042] The focus loss model 54 is any type of super-resolution neural network that upscales image data from low to high resolution, such as, but not limited to, SRGAN or FSRCNN. During the training phase, the focus loss model 54 receives paired training data 56 as input and upscales the resolution of the confused low-resolution image data 60 to create a high-resolution image with focused reconstruction. The focus loss model 54 is specifically trained to upscale the object of interest 34 ( Figure 1 The resolution of the image is determined, and the object of interest 34 that was obfuscated in the obfuscated low-resolution image data 60 is reconstructed within the reconstructed high-resolution image data. A focus loss model 54 calculates the focus loss associated with the reconstructed high-resolution image, where the high-resolution image data 62 paired with training data 56 is used as ground truth data. The focus loss model 54 is trained based on an iterative process to minimize the focus loss associated with the reconstructed high-resolution image.

[0043] Figure 4 An exemplary image frame 70 of obfuscated low-resolution image data 60 is shown. (Reference) Figure 3 and Figure 4The focus loss model 54 determines a bounding box 72 that defines a bounded region 74 within an image frame 70 containing an object of interest 34, which has been obfuscated based on one of the aforementioned obfuscation techniques. The focus loss associated with the focus-reconstructed high-resolution image assigns a greater weight to the bounded region 74 within the bounding box 72 compared to the entire image frame 70 of the obfuscated low-resolution image data 60. Specifically, the focus loss associated with the focus-reconstructed high-resolution image assigns a higher value to the bounded weighting factor corresponding to the bounded region 74 of the image frame 70 compared to the overall weighting factor corresponding to the entire image frame 70 of the obfuscated low-resolution image data 60. The focus loss model 54 determines the focus loss associated with the focus-reconstructed high-resolution image by calculating the focus mean square error loss, the focus perception loss, and the total focus variance loss, where the focus loss is the sum of the focus mean square error loss, the focus perception loss, and the total focus variance loss.

[0044] The focus mean square error loss is determined by identifying the mean square error loss associated with the entire image frame 70, the bounded region 74 within the image frame 70 containing the object of interest 34, which masks the obfuscated low-resolution image data 60, and the mean square error loss associated with the bounded region 74 of the image frame 70. The focus mean square error loss is the sum of the weighted mean square error loss associated with the bounded region 74 within the image frame 70 and the weighted mean square error loss associated with the entire image frame 70.

[0045] The weighted mean square error loss associated with the bounded region 74 within image frame 70 is determined by multiplying the mean square error loss associated with the bounded region 74 within image frame 70 by a bounded weighting factor, or (A * mean square error loss associated with the bounded region 74 within image frame 70), where A represents the bounded weighting factor. The weighted mean square error loss associated with the entire image frame 70 is determined by multiplying the mean square error loss associated with the entire image frame 70 by an overall weighting factor, or (B * mean square error loss associated with the entire image frame 70), where B represents the overall weighting factor.

[0046] It is worth noting that the bounded weighting factor A is greater than the overall weighting factor B, or A < B, and the sum of the bounded weighting factor A and the overall weighting factor B equals 1, or A + B = 1. As an example only, in one embodiment, the bounded weighting factor A equals 0.8, and the overall weighting factor equals 0.2. Therefore, compared to the entire image frame 70, the bounded region 74 containing the object of interest 34 within image frame 70 is given a higher weight, thereby improving the ability of the focusing loss model 54 to reconstruct the object of interest 34 in the focused reconstructed high-resolution image.

[0047] The focus perception loss is determined by identifying the perceptual loss associated with the entire image frame 70, the bounded region 74 of the image frame 70 that masks the obfuscated low-resolution image data 60, and the perceptual loss associated with the bounded region 74 of the image frame 70. The focus perception loss is the sum of the weighted perceptual loss associated with the bounded region 74 within the image frame 70 and the weighted perceptual loss associated with the entire image frame 70.

[0048] The weighted perceptual loss associated with a bounded region 74 within image frame 70 is determined by multiplying the perceptual loss associated with the bounded region 74 within image frame 70 by a bounded weighting factor, or (A * perceptual loss associated with the bounded region 74 within image frame 70). The weighted perceptual loss associated with the entire image frame 70 is determined by multiplying the perceptual loss associated with the entire image frame 70 by an overall weighting factor, or (B * perceptual loss associated with the entire image frame 70). In an embodiment, the bounded weighting factor used to determine the weighted perceptual loss may be a different value than the bounded weighting factor used to determine the weighted mean squared error loss. Similarly, the overall weighting factor used to determine the weighted perceptual loss may also be a different value than the overall weighting factor used to determine the weighted mean squared error loss. Therefore, the priority of the key mean squared error loss, key perceptual loss, and key total variance loss can be determined by assigning different values ​​to the bounded weighting factor and the overall weighting factor for each type of loss.

[0049] The total focus variance loss is determined by identifying the total variance loss associated with the entire image frame 70, the bounded region 74 of the image frame 70 that masks the obfuscated low-resolution image data 60, and the total variance loss associated with the bounded region 74 of the image frame 70. The total focus variance loss is the sum of the weighted total variance loss associated with the bounded region 74 within the image frame 70 and the weighted total variance loss associated with the entire image frame 70.

[0050] The weighted total variance loss associated with the bounded region 74 within image frame 70 is determined by multiplying the total variance loss associated with the bounded region 74 within image frame 70 by a bounded weighting factor, or (A * the total variance loss associated with the bounded region 74 within image frame 70). The weighted total variance loss associated with the entire image frame 70 is determined by multiplying the total variance loss associated with the entire image frame 70 by an overall weighting factor, or (B * the total variance loss associated with the entire image frame 70). In an embodiment, the bounded weighting factor used to determine the weighted total variance loss may be a different value than the bounded weighting factors used to determine the weighted mean squared error loss and the weighted perceptual loss. Similarly, the overall weighting factor used to determine the weighted total variance loss may be a different value than the overall weighting factor used to determine the weighted mean squared error loss and the weighted perceptual loss. Therefore, the priority of the key mean squared error loss, the key perceptual loss, and the key total variance loss can be determined by assigning different values ​​to the bounded weighting factor and the overall weighting factor for each different type of loss.

[0051] refer to Figure 3 Once one or more super-resolution neural networks 50 are trained, one or more controllers 20 can enter the testing phase. During the testing phase, one or more controllers 20 receive real low-resolution image data 64 representing the surrounding environment of an object of interest 34 captured by one or more cameras 24. During the testing phase, a scrambled image data model 52, a focus loss model 54, or a scrambled image data model 52 and a focus loss model 54 receive real low-resolution image data 64 as input and improve the resolution of the real low-resolution image data 64 to create real high-resolution image data 66.

[0052] Object detection module 58 receives real-life high-resolution image data 66 as input and executes one or more object detection algorithms to identify instances of objects of interest 34 within the real-life high-resolution image data 66. It should be understood that object detection module 58 can execute any type of object detection algorithm, such as, but not limited to, the "One-Look-Only" (YOLO) algorithm. It should be understood that the real high-resolution image data 66 determined by the obfuscated image data model 52 and the focus loss model 54 improves the object detection accuracy of objects of interest 34 compared to high-resolution images determined by a standard image data model that is not trained on obfuscated low-resolution image data of objects of interest 34. In a non-limiting example, the real high-resolution image data 66 determined by the obfuscated image data model 52 achieves an object detection accuracy of approximately 59% for objects of interest 34, while the real high-resolution image data 66 determined by the focus loss model 54 achieves an object detection accuracy of approximately 64% for objects of interest 34. In contrast, a standard image data model not trained on obfuscated low-resolution image data may result in an object detection accuracy of only approximately 40%. In this example, all test data were obtained based on the same dataset of 1225 parking sign images.

[0053] Referring generally to the accompanying drawings, the disclosed super-resolution system has several technical effects and advantages. The super-resolution system employs a customized method to train a super-resolution neural network based on low-resolution image data containing obfuscated objects of interest. This super-resolution neural network is specifically trained to improve the resolution of the objects of interest. The obfuscated objects in the low-resolution image data replicate the real-world situation where the objects of interest become blurred in the surrounding environment. Therefore, the super-resolution system improves the detectability of objects in both low-resolution images and images containing obfuscated objects. Furthermore, the disclosed super-resolution system can maximize the capabilities of cameras acquiring low-resolution image data.

[0054] A controller can refer to electronic circuitry, combinational logic circuitry, a field-programmable gate array (FPGA), a processor (shared, dedicated, or grouped) that executes code, or a combination of some or all of the above. For example, in a system-on-a-chip, the controller can be part of electronic circuitry, combinational logic circuitry, or an FPGA. Alternatively, the controller can be microprocessor-based, such as a computer having at least one processor, memory (RAM and / or ROM), and associated input and output buses. The processor can operate under the control of an operating system in memory. The operating system can manage computer resources so that computer program code embodied as one or more computer software applications (e.g., applications in memory) can have instructions that are executed by the processor. In alternative embodiments, the processor can directly execute the application, in which case the operating system can be omitted.

[0055] The descriptions in this disclosure are merely exemplary in nature, and changes that do not depart from the spirit and scope of this disclosure are intended to fall within its scope. Such changes should not be considered as departing from the spirit and scope of this disclosure.

Claims

1. A super-resolution system for improving the resolution of image data captured by one or more cameras, the super-resolution system comprising: One or more controllers, including one or more super-resolution neural networks, the one or more super-resolution neural networks including a scrambled image data model, wherein the one or more controllers include one or more processors, the one or more processors executing instructions to: During the training phase, paired training data is received through the obfuscated image data model, wherein the paired training data represents image data captured by the one or more cameras, the image data represents the surrounding environment and includes obfuscated low-resolution image data and high-resolution image data, wherein the obfuscated low-resolution image data and high-resolution image data both represent the same image, the obfuscated low-resolution image data includes objects of interest located in the surrounding environment, the objects of interest being obfuscated based on obfuscation techniques; The resolution of the obfuscated low-resolution image data is increased by the obfuscated image data model to create a reconstructed high-resolution image; The total loss associated with the reconstructed high-resolution image is calculated using the obfuscated image data model, wherein the high-resolution image data of the paired training data is used as ground reality data, and the obfuscated image data model is trained based on an iterative process to minimize the total loss. During the testing phase, real low-resolution image data is received through the obfuscated image data model. as well as The resolution of the real low-resolution image data is improved by the obfuscated image data model to create real high-resolution image data.

2. The super-resolution system of claim 1, wherein the total loss associated with the reconstructed high-resolution image is the sum of mean squared error loss, perceptual loss, adversarial loss, and total variance loss.

3. The super-resolution system according to claim 1, wherein the one or more super-resolution neural networks include a focusing loss model.

4. The super-resolution system of claim 3, wherein the one or more controllers execute instructions to: During the training phase, the paired training data is received through the focusing loss model; The resolution of the confused low-resolution image data is improved by the focusing loss model to create a high-resolution image with focused reconstruction. The focus loss associated with the reconstructed high-resolution image is calculated using the focus loss model, wherein the high-resolution image data of the paired training data is used as ground reality data, and the focus loss is the sum of the focus mean square error loss, the focus perception loss, and the total focus variance loss. During the testing phase, real low-resolution image data is received through the aforementioned focus loss model; as well as The resolution of the real low-resolution image data is improved by the focus loss model to create real high-resolution image data.

5. The super-resolution system of claim 4, wherein the one or more controllers execute instructions to: Determine a bounding box that defines a bounded region within an image frame that contains the obfuscated low-resolution image data, wherein the bounding box contains the object of interest.

6. The super-resolution system of claim 5, wherein the focusing loss associated with the focused reconstructed high-resolution image is assigned a higher value to the bounded weighting factor corresponding to the bounded region of the image frame compared to the overall weighting factor corresponding to the entire image frame.

7. The super-resolution system of claim 6, wherein the one or more controllers determine the focusing mean square error loss in the following manner: Determine the mean square error loss associated with the bounded region of the image frame; and Determine the mean square error loss associated with the entire image frame, wherein the focus mean square error loss is the sum of the weighted mean square error loss associated with the bounded region within the image frame and the weighted mean square error loss associated with the entire image frame.

8. The super-resolution system of claim 6, wherein the one or more controllers determine the focus sensing loss in the following manner: Determine the focus-perceived loss associated with the bounded region of the image frame; and Determine the focus perception loss associated with the entire image frame, wherein the focus perception loss is the sum of the weighted perception loss associated with the bounded region within the image frame and the weighted perception loss associated with the entire image frame.

9. The super-resolution system of claim 6, wherein the one or more controllers determine the total focusing variance loss in the following manner: Determine the total focus variance loss associated with the bounded region of the image frame; and Determine the total focus variance loss associated with the entire image frame, wherein the total focus variance loss is the sum of the weighted total focus variance loss associated with the bounded region within the image frame and the weighted total focus variance loss associated with the entire image frame.

10. The super-resolution system according to claim 6, wherein, The focusing mean square error loss, the focusing perception loss, and the focusing total variance loss all include different values ​​of the bounded weighting factor and the overall weighting factor.