Design for processing infrared images

By processing infrared images through generative adversarial networks, the problem of unnatural driver images in motor vehicles is solved, and color or grayscale images suitable for video calls are generated, saving costs and optimizing space utilization.

CN112740264BActive Publication Date: 2025-10-21VOLKSWAGEN AG
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN201980063468.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-28
Filing Date
2019-09-20
Publication Date
2025-10-21
Estimated Expiration
2039-09-20

AI Technical Summary

Technical Problem

Existing infrared cameras produce unnatural shadows and shading when imaging drivers in motor vehicles and can only provide grayscale images, while videophone customers expect color images, resulting in additional cost and space issues.

Method used

Generative adversarial networks are used to process infrared images, filter out shadows and active lighting effects, and generate color or grayscale images that can be used for video calls.

Benefits of technology

By generating adversarial network processing, the impact of unnatural lighting is reduced, the naturalness and color effects of the image are improved, the cost of additional cameras is saved, and the utilization of space inside the vehicle is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112740264B_ABST
    Figure CN112740264B_ABST
Patent Text Reader

Abstract

A method, device and computer readable storage medium having instructions for processing an infrared image and a method and computer readable storage medium having instructions for training a neural network. In a first step (10), an infrared image of an infrared camera is read in. Next, by applying a neural network to the infrared image, an output image is generated (11). For this application, the neural network has been trained for filtering shadows or for reducing lighting effects. Finally, the output image is output (12) for further use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, a device, and a computer-readable storage medium with instructions for processing infrared images, as well as a motor vehicle using such a method or such a device. The present invention also relates to a method for training a neural network, a computer-readable storage medium with instructions, and a neural network trained using such a method. Background Art

[0002] Video conferencing systems have been around since the days of ISDN (Integrated Services Digital Network). However, these systems have only become more widespread with the expansion of mobile radio capacity for mobile devices such as smartphones. Such systems are also becoming increasingly common in PC environments, particularly in the business sector. All of these systems typically feature a color camera that captures the user from the front.

[0003] Cameras are now also being used in motor vehicles for interior monitoring, for example, to film the driver. The primary application for such cameras is autonomous driving, where they check whether the driver can take over driving again. If such cameras are present in a motor vehicle, it would be advantageous to also use them for video telephony applications. This is a particularly interesting application for the driver, especially during autonomous driving. However, so-called driver observation cameras are designed to function even in dim light. To avoid dazzling the driver, illumination in the near-infrared range is used, with a wavelength of approximately 940 nm, which is invisible to the human eye. The cameras are accordingly designed specifically for this wavelength.

[0004] However, due to the near-infrared image capture and the active illumination of the face from below, the image is not particularly suitable for video conferencing. It results in a "ghost face" with unusual, unnatural-looking shadows. The eye sockets, which are usually dark, appear particularly bright, while the bridge of the nose is unusually dark. The resulting image is unsatisfactory from the customer's perspective.

[0005] Another shortcoming is that the corresponding video camera only provides grayscale images, while customers expect color images for video calls. In addition, due to the installation position of the camera, the face is captured from below and there is no eye contact between the driver and the camera.

[0006] The aforementioned problem can currently only be remedied by installing an additional color camera that is pointed directly at the driver. However, this results in considerable additional costs. Providing additional installation space is also a problem in terms of vehicle interior design.

[0007] For special applications such as E-Call (electronic telephone), unaltered infrared images may be a luxury for customers. It is therefore desirable to process the available infrared images and make them available for other applications, in particular for video telephony.

[0008] In this context, CN 105023269 A describes a method for colorizing infrared images from an infrared camera of a vehicle. First, infrared images are acquired in the vehicle in order to train a classifier using these infrared images. After the training is completed, the infrared image to be colored is handed over to the classifier, which classifies each pixel and generates a result map. The result map is divided into superpixels, and a result histogram is created for each superpixel. The main classification attributes are assigned to the entire superpixel to generate an optimized result map. Finally, an RGB image of the same size as the infrared image to be colored is first generated, and then the color space of the RGB image is converted into the HSV color space. Then, based on the optimized result map and the grayscale values ​​of the infrared image, the pixel values ​​of the image converted into the HSV color space are determined.

[0009] DE 102006044864 A1 describes a method for computer-assisted image processing in a night vision system for a motor vehicle. A detection device detects a night vision image of the motor vehicle's surroundings. Further detection devices detect parameters of the motor vehicle's surroundings. Based on the parameters of the surroundings, image regions of the night vision image are determined, each associated with an image content category. For each image content category, one or more image processing criteria are defined. The detected night vision image is then processed for display on a display device. The night vision image is processed based on the one or more image processing criteria for the image content category of the corresponding image region. Finally, the processed night vision image is reproduced on the display device.

[0010] US 6792136 B1 describes a method for generating a high-resolution color image from an infrared image of a monitored scene. The captured infrared image is analyzed to determine whether an object, such as a face, can be identified in the image. If an object can be identified, the object's characteristics are compared with a plurality of stored object characteristics. If there is a match, the object's color characteristics are determined and the object is colored based on the characteristic information stored in a database. If there is no match or no identifiable object exists and the object's color cannot be identified, the image is analyzed to determine whether a pattern, such as clothing, can be identified within the image. If a pattern can be identified, the color characteristics of the pattern are determined and colored based on the infrared reflection characteristics in combination with the stored pattern information. If a pattern cannot be identified, unpatterned and unfeatured portions of the image are colored based on the infrared reflection characteristics. Summary of the Invention

[0011] The object of the present invention is to provide an improved concept for processing infrared images.

[0012] This object is achieved by a method for processing infrared images for video telephony in a motor vehicle, by a computer-readable storage medium with instructions, and by a device for processing infrared images for video telephony in a motor vehicle. Preferred embodiments of the invention are the subject of the following description.

[0013] According to a first aspect of the invention, a method for processing infrared images for use in videophone in a motor vehicle comprises the following steps:

[0014] - reading in an infrared image of the face of the driver of the motor vehicle captured by an infrared camera installed in the motor vehicle for driver observation;

[0015] - generating an output image by applying a generative adversarial network to the infrared image, wherein the generative adversarial network has been trained to filter shadows in the infrared image caused by active lighting of the driver's face or to reduce lighting effects in the infrared image caused by active lighting of the driver's face; and

[0016] - Export the output image for videophone use.

[0017] Accordingly, a computer-readable storage medium includes the following instructions, which, when executed by a computer, cause the computer to perform the following steps to process infrared images for videophone use in a motor vehicle:

[0018] - reading in an infrared image of the face of the driver of the motor vehicle captured by an infrared camera installed in the motor vehicle for driver observation;

[0019] - generating an output image by applying a generative adversarial network to the infrared image, wherein the generative adversarial network has been trained to filter shadows in the infrared image caused by active lighting of the driver's face or to reduce lighting effects in the infrared image caused by active lighting of the driver's face; and

[0020] - Export the output image for videophone use.

[0021] The term computer is to be understood broadly herein and includes, in particular, workstations, control devices and other processor-based data processing devices.

[0022] Similarly, an apparatus for processing infrared images for use in video telephony in a motor vehicle has:

[0023] an input module for reading in an infrared image of the face of a driver of a motor vehicle captured by an infrared camera installed in the motor vehicle for driver observation;

[0024] - an image processing unit for generating an output image by applying a generative adversarial network to the infrared image, wherein the generative adversarial network has been trained to filter shadows caused by active lighting of the driver's face in the infrared image or to reduce lighting effects caused by active lighting of the driver's face in the infrared image; and - an output terminal for outputting the output image for videophone use.

[0025] The solution according to the invention is based on the idea of ​​using an infrared image of a user, in particular a driver of a motor vehicle, as input data and calculating improved output images with the aid of a generative adversarial network, which can then be used for other applications, in particular for video telephony. The improvement can consist in reducing the effects of unnatural lighting or filtering out shadows. In recent years, considerable progress has been made in the field of machine learning. As computers become increasingly powerful, complex neural networks can also be taught and used. With these complex neural networks, even difficult image processing tasks can be achieved. For example, an example of complex image processing using neural networks is described in [1]. In this example, two different images are used in order to generate a new image based on them. One of the input images specifies the content of the image to be generated. The second input image specifies the structure or texture and the color scheme. According to the invention, the neural network is a generative adversarial network (GAN, for example in German: "erzeugende gegnerische Netzwerke") [2]. The solution according to the invention makes it possible to dispense with additional cameras, which on the one hand saves costs and on the other hand enables efficient use of the available structural space.

[0026] According to one aspect of the present invention, a generative adversarial network is additionally trained to change the perspective of infrared images. By shifting the perspective, it is possible to achieve an output image that appears as if the person in the image is looking directly at the camera. This significantly improves the perceived naturalness of the rendering.

[0027] According to one aspect of the present invention, the output image is a grayscale image. Alternatively, the output image is a color image. In this case, the generative adversarial network is additionally trained to colorize infrared images. While the output image in grayscale is still not a color image, it achieves a more natural appearance. If the generative adversarial network also adds color to the image, this further significantly improves the appearance.

[0028] According to one aspect of the present invention, a generative adversarial network (GAN) can access a color image of the user to colorize infrared images. Alternatively, the user can choose between two or more GANs to colorize infrared images. The result of the processing does not directly correspond to reality; rather, it merely mimics the style of the presentation that was taught. As a result, the colors in the output image may deviate from the true colors. This is not critical for the user's clothing, but these deviations can be quite annoying, for example, in hair color. This problem can be addressed in two ways. First, the GAN can also be taught so that the system is also provided with a color photo of the current user. Based on this image, the system can make a color decision. Alternatively, it is possible to teach different GANs that take into account different personal characteristics. In this case, the user can select the GAN that best represents them.

[0029] According to one aspect of the present invention, the infrared image is part of the infrared video, and the output image is part of the output video. The described solution is not limited to still images in its application. The solution is equally applicable to moving images or image sequences. In this way, the infrared image of an infrared camera can also be used in applications based on video data.

[0030] The method according to the invention or the device according to the invention is particularly advantageously used in a vehicle, in particular a motor vehicle.

[0031] The challenge in machine learning is providing sufficient training data. For this application, a second camera can be used to generate training data. This second camera is only needed during the training phase and can be omitted later. The second camera captures the user in good lighting conditions within the visible light range. This structure is now used to capture a variety of test subjects. The image data from both cameras is then used to teach the generative adversarial network. In this way, the required training data can be easily generated.

[0032] The image from the additional camera preferably has a similar viewing angle to the infrared image from the infrared camera or has a preferred target viewing angle. In particular, the situation where a user in a motor vehicle is photographed from below and is not looking at the camera can be considered in two ways. The first possibility is to train the generative adversarial network so that it also adapts to this viewing angle. To this end, the second camera is not positioned next to the infrared camera to generate training data, but is instead aimed at the user from the front. The generative adversarial network thus implicitly learns the preferred viewing angle. In the second possibility, the viewing angle is transformed in a conventional manner (such as is common in photo manipulation) in a second step by vertically or horizontally deforming the image. To this end, the infrared camera provides the 3D position and orientation of the head. In this case, it is also possible to correct unnatural lighting conditions from below by changing the brightness gradient in the image.

[0033] The images from the additional camera are preferably grayscale or color images. Depending on the application, that is, whether a grayscale or color image is to be output as the output image, the additional camera can record grayscale or color images. This allows the training to be easily adapted to the requirements of the future application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Other features of the present invention will become apparent from the ensuing description and appended claims taken in conjunction with the accompanying drawings.

[0035] Figure 1 A method for processing infrared images is schematically shown;

[0036] Figure 2 A first embodiment of an apparatus for processing infrared images is shown;

[0037] Figure 3 A second embodiment of an apparatus for processing infrared images is shown;

[0038] Figure 4 A motor vehicle is schematically shown in which the solution according to the invention is implemented;

[0039] Figure 5 A method for training a neural network is schematically shown;

[0040] Figure 6 Schematically shows an infrared camera arranged in a motor vehicle;

[0041] Figure 7 Schematically illustrates the architecture of a structure for applying a neural network in a motor vehicle; and

[0042] Figure 8 Schematic diagram showing the architecture used to teach the structure of a neural network. DETAILED DESCRIPTION

[0043] In order to better understand the principles of the present invention, embodiments of the present invention are described in more detail below with reference to the accompanying drawings. It is readily understood that the present invention is not limited to these embodiments and that the features described may be combined or modified without departing from the scope of protection of the present invention, as defined in the appended claims.

[0044] Figure 1 A method for processing infrared images is schematically illustrated. In a first step 10, an infrared image from an infrared camera is read in. The infrared image can be part of an infrared video. Subsequently, an output image is generated 11 by applying a neural network to the infrared image. For this application, the neural network has been trained to filter out shadows or to reduce lighting effects. Additionally, the neural network can be trained to change the perspective of the infrared image or to colorize the infrared image. To colorize the infrared image, the neural network preferably has access to a color image of the user. Alternatively, the user can choose between two or more neural networks to colorize the infrared image. Finally, the output image is output 12 for further use, for example, for video telephony. In particular, the output image can be a grayscale image or a color image. It can also be part of an output video.

[0045] Figure 2A simplified schematic diagram of a first embodiment of a device 20 for processing infrared images is shown. Device 20 has an input 21, via which an infrared image from an infrared camera 41 can be read in, in particular, via an input module 22. The infrared image can be part of an infrared video. An image processing unit 23 generates an output image by applying a neural network to the infrared image. For this application, the neural network has been trained to filter shadows or reduce lighting effects. Alternatively, the neural network may have been trained to change the perspective of the infrared image or to colorize it. To colorize the infrared image, the neural network preferably has access to a color image of the user, which can be stored in a memory 25 of device 20 or received via input 21. Alternatively, the user can select between two or more neural networks to colorize the infrared image. The output image is provided via an output 26 of device 20 for further use, such as for video telephony. The output image can, in particular, be a grayscale image or a color image. It can also be part of an output video.

[0046] The input module 22 and the image processing unit 23 can be controlled by the control unit 24. If necessary, the settings of the input module 22, the image processing unit 23 or the control unit 24 can be changed via the user interface 27. The data accumulated in the device 20 can be stored in the memory 25 when necessary, for example, for later analysis or for use by the components of the device 20. The input module 22, the image processing unit 23 and the control unit 24 can be implemented as dedicated hardware, for example, as an integrated circuit. However, they can also be partially or completely combined or implemented as software running on an appropriate processor, for example, on a GPU or CPU. The input end 21 and the output end 26 can be implemented as separate interfaces or can be implemented as a combined two-way interface.

[0047] Figure 3 A simplified schematic diagram of a second embodiment of a device 30 for processing infrared images is shown. Device 30 includes a processor 32 and a memory 31. Device 30 can be, for example, a computer or a control device. Memory 31 stores instructions that, when executed by processor 32, cause device 30 to carry out the steps of one of the described methods. Thus, the instructions stored in memory 31 represent a program executable by processor 32 that implements the method according to the present invention. Device 30 includes an input 33 for receiving information, particularly images from an infrared camera. Data generated by processor 32 is provided via output 34. This data can also be stored in memory 31. Input 33 and output 34 can be combined to form a bidirectional interface.

[0048] Processor 32 may include one or more processor units, such as a microprocessor, a digital signal processor, or a combination thereof.

[0049] The memories 25 , 31 of the described embodiments may have both volatile and nonvolatile storage areas and may include a wide variety of storage devices and storage media, such as a hard disk, an optical storage medium, or a semiconductor memory.

[0050] Figure 4 A motor vehicle 40 is schematically shown in which the solution according to the invention is implemented. An infrared camera 41 is installed in the motor vehicle 40, for example in the dashboard in the area of ​​the steering wheel 42. The device 20 for processing infrared images processes the infrared images captured by the infrared camera 41 and provides the resulting output image for further use. Figure 4 In the example, device 20 is a separate component, but it could also be installed in control unit 43 of infrared camera 41. Another component of vehicle 40 is a data transmission unit 44, via which a connection to a service provider, for example for video telephony, can be established. A memory 45 is provided for data storage. Data exchange between the various components of vehicle 40 occurs via a network 46.

[0051] Figure 5 A method for training a neural network is schematically illustrated. In a first step 70, an infrared image from an infrared camera is read in. Also read in 71 is an image from an additional camera. The image from the additional camera preferably has a viewing angle similar to that of the infrared image from the infrared camera or has a preferred viewing angle. The image from the additional camera can, in particular, be a grayscale image or a color image. A neural network is trained 72 based on these images and the infrared images. The infrared image from the infrared camera serves as the input image, and the image from the additional camera serves as the output image.

[0052] Then, it should be based on Figures 6 to 8 The preferred embodiment of the present invention is described with the example of use in a motor vehicle. Of course, other uses are also possible.

[0053] Figure 6 An infrared camera 41 is schematically shown positioned in a motor vehicle. In this example, infrared camera 41 is mounted on the dashboard in the area of ​​steering wheel 42. To provide adequate illumination of infrared camera 41's field of view, infrared light sources 47 are mounted to the right and left of infrared camera 41. Due to the positioning of infrared camera 41 and the active illumination of the face from below by infrared light sources 47, the captured images have unusual perspectives and unusual, seemingly unnatural shadows.

[0054] Figure 7The schematic diagram shows the architecture of a structure for applying a neural network in a motor vehicle. The idea behind this structure is to use a trained neural network 51, which generates an improved output image 52 or output video based on an input image or input video. For example, so-called "Generative Adversarial Networks" (GANs, for example "erzeugende gegnerische Netzwerke" in German) can be used here, which generate plausible images [2]. In particular, for video telephony, a color image should be generated as output image 52 based on an infrared image 50 as input image. In the example shown, an infrared camera 41 provides image data or video data captured in the infrared (IR) or near infrared (NIR). The output data here are RGB image data or RGB video data.

[0055] It should be noted that the results do not directly correspond to reality; rather, the taught presentation style is merely simulated. Thus, in this application scenario, the algorithm might generate an image of a driver wearing a blue sweater, even if the driver is wearing a red sweater. Deviations from reality, for example, can have a more pronounced effect on hair color.

[0056] This problem can be solved in two ways. First, different neural networks 51 can be trained that take into account different personal characteristics. In this case, the driver can select the neural network 51 that best represents him. Second, the neural network can also be trained so that the system is also provided with a color photo 53 of the current driver. Based on this image, the system can make a color decision. Color photo 53 can be taken, for example, using a smartphone 54 and transferred to the vehicle using CarNet.

[0057] Figure 8The schematic diagram shows the architecture of the structure for teaching the neural network 51. The challenge in machine learning is to provide sufficient training data. For this application, an additional camera 61 can be embedded in the vehicle to generate the training data. Of course, this second camera 61 is only necessary during the development phase and can be omitted in production vehicles. The additional camera 61 films the driver from a similar perspective to the embedded infrared camera 40. In this case, the image 60 of the additional camera 61 is captured in the visible light range under good lighting conditions. Depending on the application, grayscale images or color images can be captured. In the example shown, the additional camera 61 provides RGB image data or RGB video data. Now, various test persons are filmed using this structure. The image data of the two cameras 41, 61 are then used in a training stage 62 to teach the neural network 51.

[0058] The situation where the driver is filmed from below and is not looking at cameras 41, 61 can be addressed in two ways. Firstly, the neural network 51 can be trained so that it also adapts to this perspective. To this end, the additional camera 61 is not placed next to the infrared camera 41 to generate training data, but is instead directed at the driver from the front. The neural network 51 thus implicitly learns the preferred perspective. Secondly, the perspective conversion can be performed in a conventional second step by vertically or horizontally warping the image, as is common in photo processing. For this purpose, the infrared camera 41 provides the 3D position and orientation of the head. In this case, it is also possible to correct unnatural lighting conditions from below by changing the brightness gradient in the image.

[0059] References

[0060] [1] Gatys et al.: “Image Style Transfer Using Convolutional Neural Networks”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2414–2423.

[0061] [2] Zhu et al.: “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks”, 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2242–2251.

[0062] Reference Signs List

[0063] 10 Reading infrared images

[0064] 11 Generate output image by applying neural network to infrared image

[0065] 12 Output the output image

[0066] 20 devices

[0067] 21 Input

[0068] 22 Input Module

[0069] 23 Image processing unit

[0070] 24 control unit

[0071] 25 Memory

[0072] 26 output terminals

[0073] 27 User Interface

[0074] 30 devices

[0075] 31 Memory

[0076] 32 processors

[0077] 33 Input

[0078] 34 output terminals

[0079] 40 Motor Vehicles

[0080] 41 Infrared Camera

[0081] 42 Steering Wheel

[0082] 43 Control Equipment

[0083] 44 Data Transmission Unit

[0084] 45 Memory

[0085] 46 Network

[0086] 47 Infrared light source

[0087] 50 infrared images

[0088] 51 Neural Networks

[0089] 52 Output Image

[0090] 53 color photos

[0091] 54 smartphones

[0092] 60 images

[0093] 61 Additional Camera

[0094] 62 Training Level

[0095] 70 Read infrared image from infrared camera

[0096] 71 Read in images from additional cameras

[0097] 72 Use the read image and infrared image to train the neural network

Claims

1. A method for processing an infrared image (50) for use in a videophone in a motor vehicle (40), the method comprising the following steps: - reading (10) an infrared image (50) of the face of the driver of the motor vehicle (40) captured by an infrared camera (41) installed in the motor vehicle (40) for driver observation; - generating (11) an output image (52) by applying a generative adversarial network (51) to the infrared image (50), wherein the generative adversarial network (51) has been trained to filter shadows in the infrared image (50) caused by active lighting of the driver's face or to reduce lighting effects in the infrared image (50) caused by active lighting of the driver's face, wherein, The generative adversarial network (51) has additionally been trained for colorizing the infrared image (50) and the output image (52) is a color image, wherein the generative adversarial network (51) has access to a user's color image (53) for colorizing the infrared image (50), or the user can choose between two or more neural networks (51) for colorizing the infrared image (50); and - outputting (12) said output image (52) for use in said videophone.

2. The method according to claim 1, wherein the generative adversarial network (51) has additionally been trained for changing the viewing angle of the infrared image (50).

3. The method of claim 1 or 2, wherein the infrared image (50) is part of an infrared video and the output image (52) is part of an output video.

4. A computer-readable storage medium having instructions which, when executed by a computer, cause the computer to carry out the steps of the method for processing an infrared image (50) according to any one of claims 1 to 3.

5. A device for processing infrared images (50) for use in video telephony in a motor vehicle (40), the device comprising: - an input module (22) for reading (10) an infrared image (50) of the face of the driver of the motor vehicle (40) captured by an infrared camera (41) installed in the motor vehicle (40) for driver observation; - an image processing unit (23) for generating (11) an output image (52) by applying a generative adversarial network (51) to the infrared image (50), wherein the generative adversarial network (51) has been trained to filter shadows in the infrared image (50) caused by active lighting of the driver's face or to reduce lighting effects in the infrared image (50) caused by active lighting of the driver's face, wherein The generative adversarial network (51) has additionally been trained for colorizing the infrared image (50) and the output image (52) is a color image, wherein the generative adversarial network (51) has access to a user's color image (53) for colorizing the infrared image (50), or the user can choose between two or more neural networks (51) for colorizing the infrared image (50); and - an output terminal (26) for outputting (12) the output image (52) for use in the videophone.

6. A motor vehicle (40), characterized in that The motor vehicle (40) has a device (20) according to claim 5 or is configured to carry out a method for processing infrared images (50) according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Vehicle-mounted infrared image colorization method

    CN105023269A

  • Computerized image processing method for night vision system of e.g. passenger car, involves processing detected night vision image for displaying on display unit, where processing takes place depending on image processing criteria

    DE102006044864A1

  • True color infrared photography and video

    US6792136B1

  • Human face attribute identification method and apparatus

    CN108133201A

  • Method and device for generative model

    CN108364029A