Multi-camera image display

By combining RGB and NIR imaging sensors with an ambient light sensor and using a convolutional neural network to process vehicle camera images, the problem of oversaturation or undersaturation under different lighting conditions is solved, generating clear single-view images and improving the image quality of the vehicle camera system.

CN122179533APending Publication Date: 2026-06-09GM GLOBAL TECHNOLOGY OPERATIONS LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GM GLOBAL TECHNOLOGY OPERATIONS LLC
Filing Date
2025-01-17
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Camera images used in vehicles are prone to oversaturation or undersaturation under different lighting conditions, resulting in poor image quality and difficulty in distinguishing elements in the image, thus limiting their usefulness as a replacement or supplement to mirrors.

Method used

By combining an RGB imaging sensor and a NIR imaging sensor with an ambient light sensor, the system performs image preprocessing and fusion through a controller, and uses a convolutional neural network to adjust the image under different lighting conditions to generate a clear single viewing image.

Benefits of technology

It provides clear images under various lighting conditions, improving image usability and making it more effective as a mirror replacement or complement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122179533A_ABST
    Figure CN122179533A_ABST
Patent Text Reader

Abstract

A vehicle includes at least one outward-facing camera and a viewing screen in communication with a controller. The camera includes a red-green-blue (RGB) imaging sensor and a near-infrared (NIR) imaging sensor. An ambient light sensor is disposed on the vehicle and detects a magnitude of ambient lighting in an external environment. The controller stores instructions to receive an RGB image from the RGB imaging sensor at a time t, receive a NIR image from the NIR imaging sensor at the time t, and receive a lumen value of ambient light at the time t. The RGB image is pre-processed into a pre-processed RGB image using one of a plurality of pre-processing techniques that depend on the lumen magnitude of the ambient lighting. The pre-processed RGB image and the NIR image are fused into a single viewing image using a neural network. The viewing image is displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This subject matter disclosure relates to vehicles, and more specifically to a system for generating a single image using multiple cameras in a vehicle. Background Technology

[0002] The vehicle includes observation mirrors, such as rear-view side mirrors, to provide the vehicle operator with a field of vision that would otherwise be impossible or impractical from the vehicle operator's position.

[0003] Some vehicles include cameras, such as those used in vehicle vision systems, which are adjacent to or replace conventionally placed mirrors. In this case, the view generated by the camera is provided to the driver.

[0004] However, cameras are limited by lighting conditions. When the lighting is too bright, the resulting image may be oversaturated. Similarly, when the lighting is too dim, the resulting image may be undersaturated. When an image is oversaturated or undersaturated, it may be difficult to distinguish the different elements present in the image. This, in turn, limits the usefulness of the image as a mirror replacement or mirror complement.

[0005] The goal is to provide a system that compensates for oversaturation and / or undersaturation and provides the user with a clear image produced by the camera without requiring the user to manually adjust the saturation of the camera image. Summary of the Invention

[0006] In one exemplary embodiment, the vehicle includes at least one externally facing camera communicating with a controller. The camera includes a red-green-blue (RGB) imaging sensor and a near-infrared (NIR) imaging sensor. A viewing screen communicates with the controller. An ambient light sensor is disposed on the vehicle and configured to detect the magnitude of ambient illumination in the external environment. The ambient light sensor communicates with the controller. The controller includes a processor and a memory. The memory stores instructions for causing the processor to perform the following operations: receiving an RGB image from the RGB imaging sensor at time t, receiving an NIR image from the NIR imaging sensor at time t, and receiving the lumen value of the ambient light at time t. Preprocessing the RGB image into a processed RGB image using one of several preprocessing techniques, wherein the preprocessing technique used depends on the lumen value of the ambient illumination. Fusing the processed RGB image and the NIR image into a single viewing image using a neural network. Displaying the viewing image on the viewing screen.

[0007] In addition to one or more features described herein, the neural network is a convolutional neural network trained via a training dataset that includes a first set of NIR and RGB images captured under low-light conditions, a second set of NIR and RGB images captured under optimal lighting conditions, and a third set of NIR and RGB images captured under high-light conditions.

[0008] In addition to one or more features described herein, the first set of NIR images and RGB images were captured under ambient lighting conditions below a first threshold, and an enhancement function was used to process the RGB images in the first set of NIR images and RGB images.

[0009] In addition to one or more features described herein, the second set of NIR and RGB images were captured under ambient lighting conditions above a first threshold and below a second threshold, and the RGB images in the second set of NIR and RGB images were not processed.

[0010] In addition to one or more features described herein, the third set of NIR and RGB images were captured under ambient lighting conditions above the first and second thresholds, and the RGB images in the third set of NIR and RGB images were processed using a tone mapping function.

[0011] In addition to one or more features described herein, the training dataset also includes color features extracted from RGB images in the first, second, and third training datasets, and contrast features extracted from NIR images in the first, second, and third training datasets.

[0012] In addition to one or more features described in this paper, color features are also extracted via a first loss function according to the following formula:

[0013] Where L is the extracted feature, Y2 is the final fused image output, Y is the ground reality image from the RGB imaging sensor, C is a set of color channels defining each image, H is the height of each image, and W is the width of each image; and contrast features are extracted via a second loss function according to the following formula:

[0014] Where N is the raw NIR image from the NIR imaging sensor.

[0015] In addition to one or more features described in this paper, the first and second loss functions are combined into a third loss function according to the following formula: + Where 0 < λ < 1, λ is a weighted parameter between 0 and 1, and λ depends on the lumen value of the ambient light detected by the ambient light sensor.

[0016] In addition to one or more features described in this paper, λ has increasing values ​​at both high and low lumen values.

[0017] In addition to one or more of the features described herein, the controller also includes mutual camera functionality testing.

[0018] In addition to one or more features described herein, the NIR imaging sensor and the RGB imaging sensor are configured to be close to each other.

[0019] In addition to one or more features described herein, the viewing screen is located near the side mirror, and the viewing image is a side view image.

[0020] In addition to one or more features described herein, the NIR imaging sensor and the RGB imaging sensor define the field of view on the side of the vehicle, and the observed image is a replacement image for the side mirror.

[0021] In another exemplary embodiment, a method for providing an observation image to a vehicle driver includes: receiving an RGB image from an RGB imaging sensor within a camera at time t, receiving an NIR image from an NIR imaging sensor within the camera at time t, and receiving a lumen value of ambient light at time t; preprocessing the RGB image into a processed RGB image using one of a variety of preprocessing techniques, wherein the preprocessing technique used depends on the lumen value of the ambient light; fusing the processed RGB image and the NIR image into a single observation image using a neural network; and displaying the observation image on an observation screen.

[0022] In addition to one or more features described herein, the neural network is a convolutional neural network trained via a training dataset that includes a first set of NIR and RGB images captured under low-light conditions, a second set of NIR and RGB images captured under optimal lighting conditions, and a third set of NIR and RGB images captured under high-light conditions.

[0023] In addition to one or more features described herein, the first set of NIR images and RGB images were captured under ambient lighting conditions below a first threshold, and an enhancement function was used to process the RGB images in the first set of NIR images and RGB images.

[0024] In addition to one or more features described herein, the second set of NIR and RGB images were captured under ambient lighting conditions above a first threshold and below a second threshold, and the RGB images in the second set of NIR and RGB images were not processed.

[0025] In addition to one or more features described herein, the training dataset also includes color features extracted from RGB images in the first, second, and third training datasets, and contrast features extracted from NIR images in the first, second, and third training datasets.

[0026] In addition to one or more features described herein, the observed image is a side-view mirror replacement image.

[0027] In addition to one or more features described herein, the observed image is a side mirror complement, and the observation screen is positioned near the side mirror complemented by the observed image.

[0028] The above-described features and advantages, as well as other features and advantages, of this disclosure will become apparent when taken in conjunction with the accompanying drawings and the following detailed description. Attached Figure Description

[0029] Other features, advantages, and details appear by way of example only in the following detailed description, which is described in detail with reference to the accompanying drawings, in which:

[0030] Figure 1 It is a vehicle that includes a vehicle vision system for generating camera mirror views;

[0031] Figure 2 It is a preprocessing procedure used to prepare images from multiple imaging sensors; and

[0032] Figure 3 It is used to transfer from Figure 2 The process of combining preprocessed images into a single image. Detailed Implementation

[0033] The following description is exemplary in nature only and is not intended to limit this disclosure, its application, or use. It should be understood that throughout the drawings, corresponding reference numerals denote the same or corresponding parts and features. As used herein, the term "module" refers to processing circuitry that may include application-specific integrated circuits (ASICs), electronic circuitry, processor (shared, dedicated, or group) and memory executing one or more software or firmware programs, combinational logic circuitry, and / or other suitable components that provide the described functionality.

[0034] As used herein, the term controller refers to a dedicated control system including a processor and memory configured to implement a control scheme, a general-purpose control system including a processor and memory (wherein the memory stores instructions for enabling the processor to implement the control scheme), a set of distributed processors and memories configured to cooperate in implementing a control scheme, or any similar configuration of elements configured to implement a control scheme.

[0035] According to an exemplary embodiment, Figure 1 A vehicle 10, comprising a body 12 and a passenger compartment 14, is shown. The vehicle 10 includes a wing-suit camera 20 defining a rearward field of view 22. Figure 1 In one example, the wing-mounted camera 20 is positioned in the conventional location on the vehicle body 12 of the side-view mirror. In an alternative example, the camera 20 can be mounted on the vehicle body 12 at any relevant location that provides the driver of the vehicle 10 with the desired field of vision.

[0036] Each of the cameras 20 is connected to the vehicle vision system controller (controller 30). Cameras 20 include an imaging sensor 24 for capturing red-green-blue (RGB) images and an imaging sensor 26 for capturing near-infrared (NIR) images. When undersaturated or oversaturated, RGB images retain the color characteristics of the image, but the edges of objects in the image are blurred. Due to the blurred edges, a vehicle operator may have difficulty distinguishing features in the image. Similarly, when oversaturated or undersaturated, NIR images retain the edges of objects in the image. In some practical examples, the RGB imaging sensor 24 and the NIR imaging sensor 26 are positioned close to each other and can be overlapped, such that the images present the same scene by shifting one of the images by a set amount along one or both axes. Since the RGB imaging sensor 24 and the NIR imaging sensor 26 are mechanically fixed relative to each other, the offset can be determined during calibration according to any known technique and stored in the controller 30.

[0037] Connected to the controller 30 are one or more ambient light sensors 40. The ambient light sensors 40 are positioned on the top of the vehicle 10 and configured to detect the lumen level of the ambient lighting conditions (typically referred to as ambient lighting) in which the vehicle 10 is operating. In an alternative example, the ambient light sensors 40 may be positioned at any other location on the vehicle body 12 exposed to ambient lighting. In one example, the ambient light sensors 40 also sense the directionality of the ambient lighting. As an example, the ambient light sensors may detect bright light in front of the vehicle 10 and darkness behind the vehicle 10.

[0038] Multiple viewing screens 50 are connected to controller 30. Each viewing screen 50 is configured to display a video feed provided by controller 30, wherein the video feed provides a real-time display of a video feed captured by a corresponding camera 20. Each viewing screen 50 is oriented such that a vehicle operator located in the driver's seat can see the displayed video feed. Typically, the displayed video feed is the video feed of the nearest rear-view camera 20. However, in some examples, the corresponding camera 20 may be positioned remotely from the viewing screen 50. An isometric partial view of region 60 shows a rear-view camera 20 and a corresponding viewing screen 50 in an exemplary configuration.

[0039] During periods of oversaturation (excessive illumination) and undersaturation (insufficient illumination), the RGB imaging sensor 24 in camera 20 may lack the sensitivity to provide a complete image that can be displayed to the user on viewing screen 50. Furthermore, due to the nature of the NIR imaging sensor 26, no color is provided in the NIR image. To correct this deficiency, controller 30 includes a processing module 32 configured to combine images from the RGB imaging sensor 24 and the NIR imaging sensor 26 into a single coherent image, regardless of current ambient light conditions. Processing module 32 uses a machine learning (ML) system, such as a convolutional neural network, to learn weighted values ​​for the images from each of the RGB and NIR imaging sensors 24 and to apply these weighted values ​​when performing the combination. The weighted values ​​used in a particular combination depend on the environmental conditions detected at that particular time.

[0040] Continue to refer to Figure 1 , Figure 2 An image preprocessing procedure is described for preparing images from imaging sensors 24, 26 for fusion into a final viewing image for display on viewing screen 50. Initially, imaging sensors 24, 26 capture images and provide them to controller 30, wherein the images are preprocessed in image capture step 210.

[0041] The image from image capture step 210 is provided to denoising step 232. Within denoising step 232, a weighted least-squares filter is applied to the image from RGB imaging sensor 24 to denoise the RGB image. Denoising step 232 uses the RGB image and NIR image as training inputs, and images used to learn weights for projecting the input image onto the desired output image. Denoising operates according to any established denoising process.

[0042] After the RGB image has been denoised, the preprocessing is divided into three different processing paths 204, 206, and 208, where specific paths 204, 206, and 208 are selected based on the lumen level of the ambient lighting detected by the ambient lighting sensor 40.

[0043] When the ambient lighting sensor 40 detects ambient lighting below a first threshold, process 200 proceeds along a first path 204 corresponding to the underexposure (night) condition, and in the contrast-increasing step 234, increases the contrast of the image from the RGB imaging sensor 24. After increasing the contrast of the image from the RGB imaging sensor 24, the first path 204 proceeds to the fusion network 238.

[0044] When the ambient lighting sensor 40 detects that the lumen level of the ambient lighting is between a first threshold and a second threshold that is higher than the first threshold, process 200 determines that no preprocessing is required for the image from the RGB sensor 24, and proceeds directly to the fusion network 238 along the second neural network path 206.

[0045] When the ambient light sensor 40 detects that the ambient light lumen level is higher than a second threshold (indicating overexposure conditions), the tone mapping network is used to adjust the dynamic range of the overexposed image from the RGB imaging sensor 24 in tone mapping step 236. After tone mapping, process 200 proceeds to the fusion network 238.

[0046] The specific values ​​of the first and second thresholds vary depending on the fidelity of the RGB imaging sensor 24 and the NIR imaging sensor 26, and can be determined empirically during the calibration process for any given component or design.

[0047] The fusion network 238 uses a convolutional neural network to extract a set of features from the RGB image and a second set of features from the NIR image. The fusion network 238 then combines the features into a single image, applies learned weights corresponding to the measured lumen level of the ambient lighting, and outputs the fused image to the corresponding viewing screen 50. For example, the features extracted from the RGB image could be color features, while the features extracted from the NIR image could be edge and contrast features.

[0048] Continue to refer to Figure 1 and Figure 2 , Figure 3 An exemplary process 300 is shown for combining images from a single camera's RGB imaging sensor 24 and NIR imaging sensor 26 using a convolutional neural network 330. Figure 2 Step 238 (operation).

[0049] The processed RGB image 310 from the RGB imaging sensor 24 and the NIR image 320 from the NIR imaging sensor 26 are fed to the convolutional neural network 330, which fuses the images into a single image output 340.

[0050] The convolutional neural network 330 was trained using training images generated under three different conditions: low light, best light, and high light, where the boundaries between the lighting conditions were related to a first threshold and a second threshold for preprocessing recognition.

[0051] In low-light conditions, RGB images from RGB imaging sensor 24 suffer from loss of detail and color, while NIR images from NIR imaging sensor 26 provide high resolution and clear texture, but lack color information. A convolutional neural network 30 is trained for low-light conditions.

[0052] The subnetwork is used to denoise the RGB images in the training set according to the following formula:

[0053]

[0054] Where Y1 is the denoised image, X is the noisy RGB image, and E is any regular enhancement function.

[0055] Under optimal lighting conditions, RGB images from an RGB imaging sensor provide a high level of detail, where color information is preserved, while NIR images from an NIR imaging sensor provide high resolution, sharp texture, but lack color information.

[0056] The subnetwork is used to denoise the RGB images in the training set according to the following formula:

[0057]

[0058] Under bright light conditions, the RGB image from the RGB imaging sensor 24 loses detail and color information due to saturation, while the NIR image from the NIR imaging sensor 26 provides high resolution and clear texture without color information. Under bright light conditions, the tone mapping process is applied to the RGB image according to the following formula:

[0059]

[0060] Where T is any regular tone mapping function, and the tone mapping function is applied to the image for dynamic range adjustment.

[0061] After the images in the training data are augmented, features are extracted from the RGB images in all three sets according to the following loss function:

[0062]

[0063] Where L represents the extracted features, Y2 is the final fused image output, and Y is the ground reality image from the RGB image sensor. C x H x W is the image size, where C represents the color channels, H represents the image height, and W represents the image width.

[0064] Similarly, features from the NIR imaging sensor 26 are extracted using a loss function based on the following loss function:

[0065]

[0066] Where N is the raw NIR image from NIR imaging sensor 26.

[0067] The result is the combined loss function used to train the convolutional network under all lighting conditions, where the combined loss function is:

[0068] +

[0069] in It is a weighted parameter whose value varies according to ambient lighting conditions. It has higher values ​​under low-light conditions, thus providing more weight to features extracted from NIR images in low-light conditions. Similarly... It has a higher value under high illumination conditions and a lower value under standard illumination conditions.

[0070] The final result is a training dataset that includes RGB images, NIR images, and features extracted under three different conditions. The training dataset is used to train a convolutional neural network 330, and the resulting trained convolutional neural network is used to combine RGB and NIR images during the operation of the vehicle 10.

[0071] In some implementations, one or more images used to generate the training dataset may have regions with varying lighting conditions. For example, a vehicle leaving a tunnel and entering a sunlit area would have a first region with high lighting conditions (the tunnel opening) and a second region with low lighting conditions (the tunnel walls and interior). Similar examples could include tree lines, city skylines, geographical features (e.g., hills or mountains), or any similar features that occlude a portion of the image. In such examples, the image is divided into regions corresponding to lighting conditions, and a separate tone mapping process is applied to each region of the image.

[0072] In some further embodiments, controller 30 may include a mutual camera functionality test as part of the combination process. In one example, the mutual camera functionality test is performed by identifying the most prominent object in each of the images from RGB imaging sensor 24 and NIR imaging sensor 26 using a conventional object detection process. The edges of the most prominent object are compared, and if the edges of the most prominent object match, RGB imaging sensor 24 and NIR imaging sensor 26 are determined to be functional. When one or both of NIR imaging sensor 26 and RGB imaging sensor 24 are in a faulty state (e.g., the output from one or both imaging sensors is frozen), the edges of the most prominent object will not match, and controller 30 will determine that at least one imaging sensor is not functioning properly.

[0073] Upon determining that one of the imaging sensors is malfunctioning, controller 30 provides a warning to the vehicle operator indicating an error in the camera function. In some examples, a dynamic image testing process can then be performed to identify which imaging sensor is faulty, and images from the non-faulty imaging sensors are used to generate a display on the corresponding viewing screen 50, without combining the images.

[0074] The terms “a” and “an” do not indicate a limitation of quantity, but rather that at least one of the referenced items is present. Unless the context clearly indicates otherwise, the term “or” means “and / or”. Throughout the specification, the reference to “aspect” means that a particular element described in connection with that aspect (e.g., a feature, structure, step, or characteristic) is included in at least one aspect described herein, and may or may not be present in other aspects. Furthermore, it should be understood that the described elements may be combined in any suitable manner in the aspects.

[0075] When an element, such as a layer, film, region, or substrate, is referred to as being “on” another element, it can be directly on the other element, or there may be intermediate elements present. Conversely, when an element is referred to as being “directly” on another element, there are no intermediate elements present.

[0076] Unless otherwise stated herein, all test standards are the most recent standards in force as of the filing date of this application, or, if priority is claimed, the filing date of the earliest priority application in which a test standard appears.

[0077] Unless otherwise defined, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0078] While the foregoing disclosure has been described with reference to exemplary embodiments, those skilled in the art will understand that various changes can be made and elements can be substituted with equivalents without departing from its scope. Furthermore, many modifications can be made to adapt particular situations or materials to the teachings of this disclosure without departing from the basic scope of this disclosure. Therefore, it is intended that this disclosure be limited to the specific embodiments disclosed, but will include all embodiments falling within its scope.

Claims

1. A vehicle comprising: At least one externally facing camera that communicates with the controller, wherein the externally facing camera includes a red-green-blue (RGB) imaging sensor and a near-infrared (NIR) imaging sensor; A viewing screen that communicates with the controller; An ambient light sensor is disposed on the vehicle and configured to detect the magnitude of ambient lighting in the external environment, the ambient light sensor communicating with the controller; The controller includes a processor and a memory, the memory storing instructions for causing the processor to perform operations, including: At time t, an RGB image from the RGB imaging sensor is received, an NIR image from the NIR imaging sensor is received at time t, and the lumen value of ambient light is received at time t. An RGB image is preprocessed into a processed RGB image using one of a variety of preprocessing techniques, the preprocessing technique used depending on the amount of lumen value of the ambient lighting; A neural network is used to fuse the processed RGB image and the NIR image into a single viewing image; and The viewing image is displayed on the viewing screen.

2. The vehicle of claim 1, wherein the neural network is a convolutional neural network trained via a training dataset comprising a first set of NIR and RGB images captured under low light conditions, a second set of NIR and RGB images captured under intermediate lighting conditions, and a third set of NIR and RGB images captured under high light conditions.

3. The vehicle of claim 2, wherein the first set of NIR images and RGB images are captured under ambient lighting conditions below a first threshold, and wherein the RGB images in the first set of NIR images and RGB images are processed using an enhancement function; a second set of NIR images and RGB images are captured under ambient lighting conditions above a first threshold and below a second threshold, and wherein the RGB images in the second set of NIR images and RGB images are not processed; a third set of NIR images and RGB images are captured under ambient lighting conditions above a first threshold and above a second threshold, and wherein the RGB images in the third set of NIR images and RGB images are processed using a tone mapping function.

4. The vehicle according to claim 2, wherein the training dataset includes color features extracted from RGB images in the first training dataset, the second training dataset, and the third training dataset, and contrast features extracted from NIR images in the first training dataset, the second training dataset, and the third training dataset.

5. The vehicle of claim 4, wherein the color features are extracted via a first loss function according to the following formula: Where L represents the extracted features, Y2 is the final fused image output, Y is the ground reality image from the RGB imaging sensor, C is a set of color channels defining each image, H is the height of each image, and W is the width of each image; and The contrast features are extracted via a second loss function according to the following formula: Where N is the raw NIR image from the NIR imaging sensor.

6. The vehicle according to claim 5, wherein the first loss function and the second loss function are combined into a third loss function according to the following formula: + in , It is a weighted parameter between 0 and 1, and among them Depending on the lumen value of the ambient light detected by the ambient light sensor, and It has an increasing value at both high and low lumen values.

7. The vehicle of claim 1, wherein the controller further comprises mutual camera function testing.

8. The vehicle of claim 1, wherein the NIR imaging sensor and the RGB imaging sensor are disposed close to each other.

9. The vehicle according to claim 1, wherein the observation screen is located near the side mirror, and wherein the observation image is a side view image.

10. The vehicle of claim 1, wherein the NIR imaging sensor and the RGB imaging sensor define the field of view of the side of the vehicle, and wherein the viewed image is a side mirror replacement image.