Image processing device, imaging system, and image processing method
The image processing device addresses the challenge of converting thermal images to visible light images by selecting suitable training data based on environmental factors, improving image clarity and resolution using a combined visible light and thermal camera system.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-05
- Publication Date
- 2026-05-19
AI Technical Summary
Existing image processing systems using both visible light and thermal cameras face challenges in accurately converting thermal images to visible light images due to the use of unsuitable training data, particularly in harsh environments like fog or haze, leading to poor image capture and resolution.
An image processing device that utilizes a machine learning unit to select effective training data based on environmental information, combining a visible light camera and a thermal camera, with an effectiveness determination unit to assess the suitability of visible light images for training, and a training data selection unit to choose appropriate images for correction and resolution improvement.
The system effectively performs noise reduction and resolution improvement on thermal images using appropriate training data, enhancing image clarity and accuracy through machine learning, even in varying environmental conditions.
Smart Images

Figure 2026081460000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, and particularly to an image processing apparatus that uses both a visible light camera for acquiring a visible light image and a thermal camera for acquiring a thermal image, and performs appropriate correction of an image such as noise reduction and resolution improvement.
Background Art
[0002] In recent years, systems that use both a visible light camera for imaging with visible light and a thermal camera for sensing and visualizing heat with infrared rays have come to be used in various fields. Since each camera has different characteristics, such a system enables precise observation and monitoring in a wide range of environments by using them together. Representative examples include a surveillance camera system, an automatic driving system, an imaging system, and a medical system.
[0003] By the way, since visible light cameras and thermal cameras image different wavelengths, the appearance of the acquired images changes depending on the surrounding environment. For example, in the presence of fog or haze (severe environment), a visible light camera cannot capture distant subjects due to the fog or haze. On the other hand, a thermal camera is less affected by fog or haze and can capture distant subjects. Conversely, when there is sufficient illumination and no fog or haze, a visible light camera can capture images with higher resolution than a thermal camera. This is because the visible light camera has a shorter wavelength for imaging than the thermal camera, and the pixel pitch can be made narrower.
[0004] As a technology used in such a combined system, machine learning using visible light images and infrared images as teacher data, and technologies for reducing noise and improving resolution are known. For example, Patent Document 1 discloses a technology for converting a captured far-infrared image (monochrome image) into a visible light image (color image) according to a generation model by machine learning based on visible light images and non-visible light images captured at different times.
Prior Art Documents
Patent Documents
[0005] [Patent Document 1] Japanese Patent Publication No. 2022-38287 [Overview of the project] [Problems that the invention aims to solve]
[0006] The image processing device described in Patent Document 1 above converts infrared images to visible light images with high accuracy in color values using a generative model that uses visible light images as one of the training data. However, with the technology shown in Patent Document 1, if visible light images taken in harsh environments such as fog or haze are used as training data, the subject may not be properly captured, resulting in unsuitable training data for correct image data.
[0007] The object of the present invention is to provide an image processing device that can perform image correction, such as noise reduction and resolution improvement, through machine learning using appropriate training data, in a combined system of a visible light camera that acquires visible light images and a thermal camera that acquires thermal images. [Means for solving the problem]
[0008] The image processing apparatus of the present invention is configured to perform processing on invisible light images by machine learning using visible light images as training data, and comprises: a machine learning unit that performs learning and inference for processing using training data; an environmental information acquisition unit that acquires information about the surrounding environment; an effectiveness determination unit that determines the effectiveness of visible light images as training data based on the environmental information acquired by the environmental information acquisition unit; and a training data selection unit that determines whether a visible light image is effective as training data for an invisible light image that is temporally corresponding to that visible light image based on the effectiveness determined by the effectiveness determination unit, and selects training data that has been determined to be effective. [Effects of the Invention]
[0009] According to the present invention, an image processing device is provided that combines a visible light camera for acquiring visible light images and a thermal camera for acquiring thermal images, and can perform image correction such as noise reduction and resolution improvement through machine learning using appropriate training data. [Brief explanation of the drawing]
[0010] [Figure 1] This is an image configuration diagram of the imaging system according to Embodiment 1. [Figure 2] This is an overall configuration diagram of the imaging system according to Embodiment 1. [Figure 3] This is a hardware configuration diagram of an image processing device. [Figure 4] This is a hardware configuration diagram of the client device. [Figure 5A] This figure shows an example of a visible light image taken under favorable environmental conditions. [Figure 5B] This figure shows an example of a non-visible light image taken under favorable environmental conditions. [Figure 6A] This figure shows an example of a visible light image taken under poor environmental conditions. [Figure 6B] This figure shows an example of a non-visible light image taken under poor environmental conditions during shooting. [Figure 7] This figure shows an example of a non-visible light image whose contours have been corrected using machine learning. [Figure 8] This flowchart shows the process of storing visible light images as training data and performing machine learning. [Figure 9] This is a block diagram showing the processing during machine learning training according to Embodiment 1. [Figure 10] This is a schematic diagram of a neural network as a machine learning model according to Embodiment 1. [Figure 11] This is a block diagram showing the processing during machine learning inference according to Embodiment 1. [Figure 12] This is an overall configuration diagram of the imaging system according to Embodiment 2. [Figure 13]It is a flowchart showing the output processing for the thermal image according to Embodiment 2. [Figure 14] It is a diagram showing an example of a visible light image with the contour corrected by machine learning. [Figure 15] It is a flowchart showing the process of storing the thermal image as teacher data and performing machine learning. [Figure 16] It is a flowchart showing the output processing for the visible light image according to Embodiment 4. [Figure 17] It is an overall configuration diagram of the imaging system according to Embodiment 5.
Mode for Carrying Out the Invention
[0011] Hereinafter, each embodiment according to the present invention will be described with reference to FIGS. 1 to 17.
[0012] 〔Embodiment 1〕 Hereinafter, Embodiment 1 according to the present invention will be described with reference to FIGS. 1 to 11. First, the image of the imaging system according to Embodiment 1 will be described with reference to FIG. 1. FIG. 1 is an image configuration diagram of the imaging system according to Embodiment 1.
[0013] As shown in FIG. 1, the imaging system 10 has the image processing apparatus 100 and the client apparatus 200 connected by the network 20 to transfer image data.
[0014] The image processing apparatus 100 is a so-called dual-spectrum camera having both functions of visible light imaging and thermal image imaging by infrared rays, and includes a visible light imaging apparatus 310 and a non-visible light imaging apparatus 320. A display apparatus 410 is connected to the client apparatus 200, and the user can thereby confirm the visible light image and the thermal image photographed by the image processing apparatus 100.
[0015] Next, the configuration of the imaging system according to Embodiment 1 will be described with reference to FIGS. 2 to 4. Figure 2 is an overall configuration diagram of the imaging system according to Embodiment 1. Figure 3 is a hardware configuration diagram of the image processing device. Figure 4 is a hardware configuration diagram of the client device.
[0016] As shown in Figure 2, the imaging system 10 has a configuration in which an image processing device 100 and a client device 200 are connected by a network 20.
[0017] Network 20 can be wired or wireless, and it can be a local area network (LAN) or a global network (Internet).
[0018] The client device 200 is a device that displays images processed by the image processing device 100, collects the status of the image processing device 100, and allows the user to instruct the image processing device 100 on what to do.
[0019] The image processing device 100 consists of a control unit 101, an imaging unit 110, a communication unit 102, a storage unit 103, a training data selection unit 104, a machine learning unit 105, an effectiveness calculation unit (effectiveness determination unit) 106, and an environmental information acquisition unit 120.
[0020] The control unit 101 is a functional unit that controls each part of the image processing device 100 and performs image processing.
[0021] The communication unit 102 is a functional unit that performs communication between the image processing device 100 and the client device 200. The communication unit 102 can transmit image data output by the image processing unit 113 of the imaging unit 110 to the client device 200 via the network 20. It can also receive operation information input from the operation input unit 205 of the client device 200.
[0022] The memory unit 103 is a functional unit that stores necessary data and programs for the image processing device 100. The memory unit 103 can store and read image data output by the image processing unit 113 of the imaging unit 110. Furthermore, it is also used as a storage area for programs executed by the control unit 101, a storage area for various parameters, and a work area during program execution.
[0023] The imaging unit 110 has a visible light imaging unit 111 and a non-visible light imaging unit 112. These units each image electromagnetic waves at different wavelengths; the visible light imaging unit 111 is in the visible light wavelength range (approximately 360 nm to 830 nm), while the non-visible light imaging unit 112 is, for example, in the infrared wavelength range (approximately 830 nm to 15000 nm).
[0024] The image processing unit 113 is a functional unit that converts the signals photoelectrically converted by the visible light imaging unit 111 and the non-visible light imaging unit 112 into image data (digital data).
[0025] In the image processing unit 113, pixel data is converted into a digital signal by A / D conversion. The converted digital signal is then converted back into image data through correction and development processes such as black level correction, gamma curve adjustment, temperature correction, scratch correction, noise reduction, and white balance correction. Data compression processes such as MP4 and JPEG are also performed. This image processing unit performs corrections tailored to the wavelength characteristics of each image sensor.
[0026] The visible light imaging unit 111 and the non-visible light imaging unit 112 can capture images at their respective wavelengths in the same time period.
[0027] In this embodiment, as shown in Figure 2, an example is shown in which both a visible light imaging unit 111 and a non-visible light imaging unit 112 are located in a single image processing device 100. However, the visible light imaging unit 111 and the non-visible light imaging unit 112 may be located in separate devices from the image processing device 100. Furthermore, the visible light imaging unit 111 and the non-visible light imaging unit 112 may be located in different devices. However, the imaging areas of the visible light imaging unit 111 and the non-visible light imaging unit 112 must overlap in some respects in order to photograph the same subject.
[0028] The training data selection unit 104 is a functional unit that determines whether the acquired visible light image is valid as training data for a thermal image and selects the visible light image. If the validity level (details described later) is high, it is stored in the storage unit 103 as image data to be used as training data.
[0029] The effectiveness calculation unit 106 is a functional unit that calculates (determines) the effectiveness based on the environmental information acquired by the environmental information acquisition unit 120. The determination for selecting training data is made by comparing the effectiveness calculated from the environmental data acquired by the environmental information acquisition unit 120 with a predetermined threshold (details will be described later). Furthermore, the calculation of effectiveness may be performed not only based on the environmental data acquired by the environmental information acquisition unit 120, but also based on the results of the analysis of the acquired images.
[0030] The machine learning unit 105 consists of a neural network (described later) that performs machine learning based on training data and student data. The image parameters calculated by the neural network enable image processing of the captured image (adjusting color tones such as white balance and contrast, and outputting it in a viewable format such as JPEG). The machine learning unit 105 may also be integrated with the image processing unit 113.
[0031] In this specification, "training data" refers to a dataset containing input data for a model and correct labels (target values) for that data, according to the general definition in machine learning. "Student data" refers to the data that is to be learned from the training data.
[0032] As will be explained in detail later, the training data in this embodiment is a high-resolution visible image, and the student data is a low-resolution thermal image.
[0033] The environmental information acquisition unit 120 is a functional unit that acquires environmental information of the image processing device 100 or the area surrounding the image processing device 100. The environmental information acquisition unit 120 includes an illuminance acquisition unit 121, a weather information acquisition unit 122, a distance acquisition unit 123, an attitude acquisition unit 124, a position acquisition unit 125, and a speed acquisition unit 126.
[0034] The illuminance acquisition unit 121 is implemented as hardware by an illuminance sensor and is a functional unit that processes the acquired digital data to obtain environmental information related to brightness. Alternatively, brightness may be calculated from image data acquired from the imaging unit 110 (shooting conditions such as exposure, gain, aperture, shutter speed, and brightness information). Alternatively, time information may be acquired by an RTC (Real Time Clock) (a clock function built into the device) and brightness may be predicted from the time of day (morning, noon, night).
[0035] The weather information acquisition unit 122 is implemented as hardware using temperature sensors, humidity sensors, rainfall sensors, wind speed sensors, and barometric pressure sensors, and processes the acquired digital data to obtain environmental information related to weather information. Alternatively, the weather may be estimated by analyzing the image data acquired from the imaging unit 110.
[0036] The distance acquisition unit 123 processes data output from distance sensors such as Lidar (Light Detection and Ranging), millimeter-wave radar, and ultrasonic sensors to calculate distance information (environmental information) to the surrounding environment and obstacles of the image processing device 100. Alternatively, distance information may be calculated from the focus evaluation value obtained by phase-detection AF (AUTO FOCUS).
[0037] The attitude acquisition unit 124 processes digital data output from a gyro sensor (angular velocity sensor) and calculates the rotation and orientation changes of the image processing device 100.
[0038] The position acquisition unit 125 receives and processes GPS (Global Positioning System) signals to determine the position of the image processing device 100, and processes electronic compass signals to determine the direction. It can also calculate environmental information around the image processing device 100 from information obtained from the illuminance acquisition unit 121 and the weather information acquisition unit 122. Furthermore, it can calculate environmental information of the area being photographed from the shooting direction and field of view of the imaging unit 110 and distance information obtained by the distance acquisition unit 123.
[0039] The speed acquisition unit 126 processes the digital data output from the acceleration sensor and other sources to calculate the acceleration and motion speed applied to the image processing device 100.
[0040] Some or all of the functions of the environmental information acquisition unit 120 do not have to be functions of the image processing device. For example, it is sufficient if an IoT (Internet of Things) device can communicate with the image processing device 100 or the client device 200. Alternatively, the functions may be located in the client device 200.
[0041] The client device 200 is, for example, an information processing device such as a personal computer, and is composed of a control unit 201, a communication unit 202, a display unit 203, an operation input unit 205, a storage unit 206, and an external device I / F unit 207.
[0042] The control unit 201 is a functional unit that controls each component of the client device 200 and performs system control such as setting various parameters, display control, and data transmission / reception instructions.
[0043] The communication unit 202 is a functional unit that communicates with the image processing device 100 via the network 20.
[0044] The display unit 203 is controlled by the display control unit 204 and is a functional unit that displays images captured by each imaging device and various information, and is implemented using an LCD (Liquid Crystal Display).
[0045] The operation input unit 205 is a functional unit that provides user instructions and inputs data to the client device 200, and is implemented using a keyboard, mouse, or touch panel. The user can control the image processing device 100 and the client device 200 by operating the keyboard, mouse, or touch panel.
[0046] The memory unit 206 is a functional unit that stores the program executed by the client device 200, various parameters, and work data during program execution.
[0047] The external device interface (I / F) section 207 is an interface for connecting the client device 200 to a display device such as a PC (Personal Computer) or a display. The external device interface (I / F) section 207 is also an interface for connecting to an external storage medium (e.g., a hard disk, memory card, SD card, USB memory, etc.).
[0048] Furthermore, the client device 200 may be able to connect to an external server and handle cloud data. Also, some of the functions may reside on the server or another personal computer connected to the server.
[0049] Next, we will explain the hardware configuration of the image processing device using Figure 3. As shown in Figure 3, the image processing device 100 is configured such that a visible light imaging device 310, a non-visible light imaging device 320, a CPU (Central Processing Unit) 301, a main memory 302, a non-volatile memory 303, an environmental information acquisition device group 330, and a communication I / F device 350 are connected by a bus.
[0050] The visible light imaging device 310 includes a visible light imaging optical system 311 and a visible light image sensor 313. Similarly, the non-visible light imaging device 320 includes a non-visible light imaging optical system 321 and a non-visible light image sensor 323.
[0051] The visible light imaging optical system 311 includes a visible light lens 312 (zoom lens, focus lens) and an aperture mechanism, and focuses visible light (wavelength: approximately 360nm to 830nm) from the subject onto the light-receiving surface of the visible light image sensor 313. Here, the visible light lens 312 is a lens with high transmittance in the visible light wavelength band, and the visible light image sensor 313 is an element with high sensitivity in the visible light wavelength band.
[0052] Similarly, the invisible light imaging optical system 321 includes an invisible light lens 322 (zoom lens, focus lens) and an aperture mechanism to focus invisible light (wavelengths other than visible light, for example, infrared wavelengths of approximately 830 nm to 15000 nm) from the subject onto the light-receiving surface of the invisible light image sensor 323. Here, the invisible light lens 322 is a lens with high transmittance in the invisible light wavelength band, and the invisible light image sensor 323 is an element with high sensitivity in the invisible light wavelength band.
[0053] Here, a zoom lens is a lens that moves along the optical axis and can change the magnification of the image, and a focus lens is a lens that moves along the optical axis and can adjust the focus. The aperture mechanism is a mechanism that adjusts the amount of light passing through the optical system.
[0054] The visible light image sensor 313 and the non-visible light image sensor 323 are semiconductor elements such as CMOS (Complementary Metal Oxide Semiconductor) sensors and CCD (Charge Coupled Device) sensors. The visible light image sensor 313 and the non-visible light image sensor 323 convert the light incident from their respective imaging optics into an analog video signal by photoelectric conversion. In particular, the non-visible light image sensor 323 is an infrared sensor sensitive to infrared light, including near-infrared, mid-infrared, and far-infrared. Far-infrared light is commonly used, especially in thermal cameras. Depending on the application, semiconductor elements sensitive to wavelengths other than visible light, such as ultraviolet sensors, may also be used.
[0055] The CPU 301 is a processor that controls and processes each function of the image processing device 100. The CPU 301 controls each function, such as the control unit 101, image processing unit 113, training data selection unit 104, and machine learning unit 105, as shown in the functional diagram in Figure 2.
[0056] The main memory 302 is a volatile semiconductor element such as RAM (Random Access Memory) and stores programs and work data executed by the CPU 301. The non-volatile memory 303 is a non-volatile semiconductor element such as flash memory and stores setting data and installed programs for the image processing device 100. In this embodiment, the non-volatile memory 303 stores visible light image data 30, non-visible light image data 31, machine learning control data 40, image parameters 41, and environmental data 50.
[0057] Visible light image data 30 is data of a visible light image captured by the visible light imaging device 310. Non-visible light image data 31 is data of a non-visible light image captured by the non-visible light imaging device 320. Machine learning control data 40 is data used to control machine learning learning and inference. Image parameters 41 are parameters for processing images. Environmental data 50 is data related to environmental information acquired by the environmental information acquisition device group 330.
[0058] The communication interface device 340 is an interface device for the image processing device 100 to communicate with other devices such as the client device 200.
[0059] In this embodiment, the image processing device 100 has a hardware configuration in which the CPU 301 executes a program installed in the non-volatile memory 303. However, it is not limited to this, and can also be implemented using a single logic circuit element (for example, an ASIC (Application Specific Integrated Circuit)).
[0060] The environmental information acquisition device group 330 consists of various sensor devices for acquiring environmental information. For example, as shown in Figure 3, the environmental information acquisition device group 330 includes an illuminance sensor 331, a wind speed sensor 332, a barometric pressure sensor 333, a temperature sensor 334, a humidity sensor 335, a rainfall sensor 336, a gyro sensor 338, an acceleration sensor 339, and a distance sensor 340. The environmental information acquisition device group 330 also includes a GPS receiver 341 and an electronic compass 342.
[0061] The illuminance sensor 331 is a device that measures the illuminance around the image processing device 110. The wind speed sensor 332, the atmospheric pressure sensor 333, the temperature sensor 334, the humidity sensor 335, and the rainfall sensor 336 are devices that measure wind speed, atmospheric pressure, temperature, humidity, and rainfall at the location where the image processing device 110 is located, respectively.
[0062] The gyro sensor 338 is a device that the image processing device 110 uses to detect the amount of change in angle per unit time (various velocities). The acceleration sensor 339 is a device that measures the acceleration when the image processing device 110 is moved. The distance sensor 340 is a device that calculates the distance to the surrounding environment and obstacles of the image processing device 100, such as Lidar (Light Detection and Ranging), millimeter-wave radar, and ultrasonic sensors.
[0063] The GPS receiver 341 is a device that receives signals from artificial satellites orbiting the Earth and determines the current location of the image processing device 200. The electronic compass 342 is a device that electrically measures the magnitude of the Earth's magnetic field using a magnetic sensor and calculates directional information by calculating the measured value.
[0064] In this embodiment, the image processing device 100 was described as performing a program reading function. This program is supplied to the image processing device 200 via a network or storage medium. In this embodiment, the CPU 301 was described as executing a program installed in the non-volatile memory 303, but this can also be achieved by a single logic integrated circuit (e.g., ASIC: Application Specific Integrated Circuit).
[0065] Next, we will explain the hardware configuration of the client device using Figure 4. The client device 200 is a general information processing device, such as a personal computer. As shown in Figure 4, the client device 200 consists of a CPU 401, main memory 402, non-volatile memory 403, display I / F device 404, external device I / F device 405, communication I / F device 406, and input / output I / F device 407 connected by a bus.
[0066] The CPU 401 controls the various parts of the client device 200 and executes programs. The main memory 402 is a volatile semiconductor element that stores programs executed by the CPU and work data. The non-volatile memory 403 is a non-volatile semiconductor element such as flash memory that stores programs executed by the client device and client device configuration data. The display I / F device 404 is an interface device for connecting a display device 410 such as a display. The external device I / F device 405 is an interface device for connecting the client device and an external device, and performs format conversion according to the standard when the client device 200 is connected to an external device via a wired connection. The communication I / F device 406 is a device for connecting the client device 200 and the image processing device 100 via a wired or wireless connection. The input / output I / F device 407 is an interface device for connecting input / output devices such as a keyboard 420 and a mouse 421.
[0067] Next, we will explain a specific example of correcting thermal images using machine learning with visible light images as training data, using Figures 5A to 7. Figure 5A shows an example of a visible light image taken under favorable environmental conditions. Figure 5B shows an example of a non-visible light image taken under favorable environmental conditions. Figure 6A shows an example of a visible light image taken under poor environmental conditions. Figure 6B shows an example of a non-visible light image taken under poor environmental conditions. Figure 7 shows an example of a non-visible light image whose contours have been corrected using machine learning.
[0068] In this embodiment, the image processing device 100 performs image processing to correct a thermal image (invisible light image) using machine learning with a visible light image as training data. The key idea is to consider the environmental conditions at the time of shooting when selecting the visible light image as training data.
[0069] To perform this type of image processing, it is assumed that the visible light imaging unit 111 and the invisible light imaging unit 112 of the image processing device 100 capture images in a direction where at least a portion of their respective imaging areas overlap. Then, machine learning is performed on the subject (area) captured in this overlapping area. Here, we will explain using the example where the visible light imaging unit 111 and the invisible light imaging unit 112 have an ideal field of view with no parallax, and the size and position of their imaging areas are the same. If there is a difference in the areas captured, the same area should be extracted from each, the pixel positions corrected, and then machine learning should be performed. Furthermore, if there is physical parallax between the visible light imaging unit 111 and the invisible light imaging unit 112, it is desirable to correct the parallax using keystone correction or the like.
[0070] Here, Figure 5A shows a visible light image under favorable environmental conditions during shooting, and Figure 5B shows a thermal image under favorable environmental conditions during shooting. In this embodiment, the concept of "effectiveness" during shooting is introduced. As an indicator of effectiveness, the more suitable the environment is for shooting visible light and invisible light images, the higher the effectiveness is judged to be. For example, when the image processing device 100 is shooting, it may be during the daytime when the sun is out and there is sufficient illumination, there may be no fog, haze, rain, snow, hail, etc., and the shaking of the image processing device 100 may be small. Furthermore, only a single piece of information (for example, brightness information obtained from an illuminance sensor) may be used to determine the effectiveness, or it may be determined by combining different pieces of information. The calculation of effectiveness will be explained in detail later.
[0071] Generally, in environments with high effectiveness, as shown in Figures 5A and 5B, the visible light image 501 has higher resolution (image quality) than the thermal image 601. In this embodiment, an image with clear contours is described as an image with high resolution (image quality). As shown in Figure 5A, the contours (edges) of the subject 701 being photographed are clear in the visible light image 501, while as shown in Figure 5B, the contours of the subject 701 photographed in the thermal image 601 are less clear than in the visible light image 501. In this case, since the thermal image 601 has low resolution, it is possible to generate a thermal image 601 with clear contours by performing image processing using machine learning, with the visible light image 501 as training data and the thermal image 601 as student data. Therefore, when effectiveness is low in this way, the visible light image 501 is saved as training data for the thermal image 601. A specific example of a thermal image corrected by machine learning will be described later.
[0072] On the other hand, when the environmental conditions during shooting are poor, i.e., when the effectiveness is low, the visible light image 501 and thermal image 601 are shown in Figure 6A and Figure 6B, respectively.
[0073] Figure 6A shows how fog degrades the visibility of the visible light image (resulting in low resolution). In such poor shooting conditions (harsh environments), the subject 701 may not be captured in the visible light image 501. If, in such a case, the visible light image 501 were used as training data and the thermal image 601 as student data, the outline of the subject 701 would become even more unclear, resulting in lower resolution. Therefore, when the effectiveness is low in this way, the visible light image 501 should not be selected as training data for the thermal image 601.
[0074] When environmental conditions are favorable, i.e., in an environment with high effectiveness, machine learning is performed using the visible light image 501 shown in Figure 5A as training data and the thermal image shown in Figure 5B as student data, and contour correction is performed, resulting in the image-corrected thermal image 611 shown in Figure 7. As shown in Figure 7, the image-corrected thermal image 611 has a clearer contour of the subject 701 and improved resolution compared to the thermal image 601.
[0075] Next, the relationship between the various environmental information acquired by the environmental information acquisition unit 120 and its effectiveness is shown below. It is desirable to determine the effectiveness comprehensively from various information obtained from the environmental information acquisition unit 120. For the visible light imaging unit 111, if the imaging is performed in a harsh environment, the effectiveness will be judged as low.
[0076] When the illuminance acquired by the illuminance acquisition unit 121 is high, that is, when it is bright, the effectiveness is determined to be high, and when the illuminance is low, that is, when it is dark, the effectiveness is determined to be low.
[0077] In the weather information acquired by the weather information acquisition unit 122, if the weather is clear, the effectiveness is judged to be high. If the weather is bad (fog, haze, rain, sleet, hail, thunderstorm, strong winds, etc.), the effectiveness is judged to be low.
[0078] Regarding the distance information acquired by the distance acquisition unit 123, when the distance between the image processing device 100 and the subject is small, that is, when the subject is photographed at close range, the effectiveness is determined to be high. When the distance between the image processing device 100 and the subject is large, that is, when the subject is photographed at a distance, the effectiveness is set to be low. This is because subjects at close range are less affected by fog, haze, etc., and therefore the effectiveness is considered to be higher.
[0079] Regarding the information acquired by the attitude acquisition unit 124, if the fluctuation (measured vibration) is large, the shaking is large and the effectiveness is determined to be low. If the fluctuation is small, the effectiveness is determined to be high. If the effectiveness is low, the subject will be blurred due to the shaking and will therefore be unsuitable as training data. Furthermore, if the fluctuation determined by the attitude acquisition unit 124 is small, it is desirable not to use either visible light images or thermal images (an example of using thermal images as training data will be described later in Embodiment 2) as training data.
[0080] Regarding the information acquired by the speed acquisition unit 126, similar to the attitude acquisition unit 124, if the fluctuation (movement speed) is large, the shaking is large, and the effectiveness is determined to be low. If the fluctuation is small, the effectiveness is determined to be high. If the effectiveness is low, the subject will be blurred due to the shaking, making it unsuitable for training data. Furthermore, if the movement speed determined by the attitude acquisition unit 124 is large, it is desirable not to use it as training data in either the visible light image or the thermal image (described later in Embodiment 2).
[0081] Next, the image processing performed by the imaging system according to Embodiment 1 will be described using Figures 8 to 11. Figure 8 is a flowchart showing the process of storing visible light images as training data and performing machine learning. Figure 9 is a block diagram showing the processing during machine learning training according to Embodiment 1. Figure 10 is a schematic diagram of a neural network as a machine learning model according to Embodiment 1. Figure 11 is a block diagram showing the processing during machine learning inference according to Embodiment 1.
[0082] First, using Figure 8, we will explain the process of storing visible light images as training data and performing machine learning.
[0083] First, the training data selection unit 104 of the image processing device 100 acquires a visible light image that will serve as a candidate for training data, and a thermal image that is captured simultaneously (S800).
[0084] Next, the training data selection unit 104 of the image processing device 100 determines from the image whether or not the subject 701 is captured in the thermal image 601 (S801). For example, the determination method is based on whether or not there is a heat source. If the subject 701 is captured (S801: YES), the process proceeds to S802. At this time, it is not necessary to obtain sufficient resolution in the thermal image 601. Therefore, it is desirable to make the determination using a method other than contour (edge) detection, such as threshold determination of the temperature difference between the subject and the background, or motion detection. Also, if it is determined that the subject 701 is not captured in the thermal image 601 (S801: NO), the thermal image is unsuitable as student data, and processing is terminated.
[0085] If the subject 701 is captured in the thermal image 601, it is determined from the image whether or not the subject 701 is identified in the visible light image 501 (S802). One method of determination is to determine whether or not the subject 701 has distinctive feature points. That is, image analysis is performed on the visible light image 501 to detect edges and determine whether or not it has feature points based on its shape. For example, if the subject 701 is assumed to be a ship, the result of edge detection is used to determine through image processing whether or not it has distinctive feature points of a ship. If the specific subject is captured (S802: YES), proceed to S803. If the specific subject is not captured (S802: NO), the process ends.
[0086] Here, we have explained an example of identifying the feature points of a subject in the analysis of whether a subject exists in the visible light image 501, but other determination methods such as determining whether or not there is a change in brightness (a change in the image due to a moving object, etc.) may also be used. Furthermore, even if one of S801 or S802 is omitted, the image processing device 100 can, in principle, detect the subject 701. Also, if the user specifies the subject 701 or a specific area to be included in the image, S801 and S802 can be omitted. In addition, the subject 701 may be detected from information obtained from the client device 200, distance sensor 320, GPS 321, etc.
[0087] Next, when it is determined that the subject 701 is captured in the visible light image 501, the effectiveness calculation unit 106 of the image processing device 100 reads the environmental data 50 acquired by the environmental information acquisition unit 120 (S803).
[0088] Next, the effectiveness calculation unit 106 of the image processing device 100 calculates the effectiveness based on the environmental data 50 read in S803 (S804). The specific method for calculating the effectiveness will be described in detail later.
[0089] Next, the training data selection unit 104 of the image processing device 100 determines whether the effectiveness calculated in S804 is equal to or greater than a predetermined threshold (S805). If the effectiveness is equal to or greater than the threshold (S805: YES), the process proceeds to S806; if the effectiveness is less than the threshold (S805: NO), the process ends.
[0090] For example, in the shooting environment of images like those in Figures 5A and 5B, the illumination is sufficient and there is no fog or haze, so the effectiveness is calculated to be a high value. Therefore, if the threshold is set appropriately, the effectiveness will be above the threshold, and the process proceeds to S806. On the other hand, in the shooting environment of images like those in Figures 6A and 6B, the resolution is poor due to the fog, so the effectiveness falls below the threshold, and the visible light image 501 is not stored as training data, and the process ends.
[0091] When the effectiveness is determined to be above a threshold, the training data selection unit 104 of the image processing device 100 stores the visible light image 501 as training data and the thermal image 601 as student data (S806).
[0092] Next, the machine learning unit of the image processing device 100 performs machine learning using the stored training data and student data (S807). Details of the machine learning method will be explained later with reference to Figure 9.
[0093] As described above, the procedure shown in this embodiment makes it possible to obtain appropriate visible light images as training data.
[0094] Furthermore, when inputting image data into machine learning, you may process the images to make them easier to learn from before inputting them. Below, we will explain four examples of image processing. (1) Processing to emphasize the outline of the subject
[0095] The image is enhanced to highlight the contours of the subject. Alternatively, the output may be binarized. (2) Image resizing.
[0096] The area in which the subject is captured is extracted. At this time, the same area (where the areas overlap) is extracted from both the visible light image and the thermal image for training purposes. If the shooting range differs between the visible light image and the thermal image, the size and position of the images are made consistent by extracting the area. If the number of pixels in the visible light image 501 and the thermal image 601 differ, the number of pixels may be made consistent before machine learning is performed. For example, if the visible light image 501 has a resolution of 3840x2160 and the thermal image 601 has a resolution of 1280x720, the thermal image 601 is upscaled to a resolution of 3840x2160 before being input to the machine learning unit 105. (3) Processing to reduce unnecessary areas
[0097] The data in the background area is filled with white or black. Also, since color information cannot be reproduced in the thermal image 601, the color information in the visible light image 501 is deleted. Since the letters of ship names and other text cannot be reproduced in the thermal image 601, it is desirable to process the areas of the letters captured in the visible light image 501, such as blurring them with a low-pass filter, in order to improve the accuracy of the learning process. (4) Parallax correction processing
[0098] Parallax (trapezoidal) correction is performed when the distance to the subject is short. For example, a threshold can be set, such as a distance of 500m or less between the image processing device 100 and the subject, and parallax correction is performed when the distance is shorter than that. It is also desirable to change the image parameters for parallax correction depending on the distance.
[0099] Furthermore, when storing visible light images as training data for machine learning (S806), labeling (classifying data) with environmental information is also permitted. By performing machine learning on data with the same label, the accuracy of image generation can be improved. Examples include nighttime, fog or haze, and distance (numerical values such as xx [km]).
[0100] You may also label subjects based on their classification. For example, ships, people, vehicles, aircraft, drones, etc. Furthermore, if the use case is limited, you may use more detailed classifications. For example, if it is a ship, you may use large, small, sailing ships, etc.
[0101] Furthermore, the weighting of machine learning can be changed based on effectiveness. For training data with effectiveness near a threshold, it is desirable to set a smaller weight for the data that can be treated as training data. This makes it possible to train the machine learning model more appropriately.
[0102] Next, we will explain the details of machine learning using Figure 9.
[0103] This corresponds to process S807 in Figure 8.
[0104] First, the machine learning unit 105 of the image processing device 100 receives the visible light image 501, which is the training data (S901).
[0105] Next, the thermal image 601, which is student data, is input to the machine learning unit 105 of the image processing device 100 (S902). The images input by the teacher data input process and the student data input process are time-corresponding, that is, images taken at the same time.
[0106] Next, the machine learning unit 105 of the image processing device 100 performs image processing using a neural network (S903). In the image processing in S903, the thermal image 601, which is a student image, is enhanced in quality based on the image parameters calculated in S904 (described later).
[0107] Next, the machine learning unit 105 of the image processing device 100 compares the training data input in S901 with the student data after image processing in S903 and calculates the error (S905).
[0108] Next, the machine learning unit 105 of the image processing device 100 calculates the amount to update the image parameters based on the error calculated in S905 (S906).
[0109] Next, the machine learning unit 105 of the image processing device 100 calculates image parameters based on the amount of image parameter update in S906 (S904). The image parameters here are parameters that represent gamma value, brightness, contrast, sharpness, etc.
[0110] As shown above, by performing machine learning for image processing using highly effective training data, it is possible to obtain images with improved resolution (clearer outlines).
[0111] Furthermore, while we have explained machine learning for image processing, it can also be used for machine learning for classification. For example, if the effectiveness is high and the subject is determined to be a ship from image analysis of a visible light image, then ships are used as the positive (training data). In this case, in machine learning using thermal images as student data, ships are trained as the positive. In this way, it can be applied to different types of machine learning, such as classification. In classification, it is not limited to ships, but it is possible to set any classification and make judgments on reefs, people, vehicles, animals, aircraft, etc.
[0112] Next, we will explain the basic concepts of neural networks as machine learning models according to Embodiment 1.
[0113] A neural network is a computer model that mimics the function of nerve cells (neurons) in the human brain. It has multiple layers of nodes (neurons) that are connected to each other to transmit and process information. A neural network consists of an input layer, a hidden layer, and an output layer. In the example in Figure 10, there are two hidden layers, but it can be configured with more layers.
[0114] In a neural network, each node (neuron) adjusts the importance of the input signal using values called weights when passing a signal to the next node. Each node also adds a value called bias, and finally outputs the signal to the connected nodes through an activation function. During the learning process, weights and biases are adjusted, improving the model to make more accurate predictions.
[0115] In this embodiment, the weighting of image parameters is adjusted based on a comparison of the output layer results with the input training data.
[0116] Next, Figure 11 will be used to explain the processing during machine learning inference according to Embodiment 1.
[0117] First, the machine learning unit 115 of the image processing device 100 receives image data of the thermal image 601 captured by the invisible light imaging unit 112 (S911). Next, the input image data is processed by:
[0118] Next, the machine learning unit 115 of the image processing device 100 processes the image using a neural network and performs correction processing on the thermal image 601 based on the image parameters obtained in S913 (S912). The image parameters calculated in S913 in Figure 11 are the same image parameters calculated from the image parameter calculation process in S904 during training, as shown in Figure 9. The image generated by the image processing (thermal image 611 after image correction) is output as a corrected thermal image (S914) and stored in the non-volatile memory 303 or transmitted to the client device 200.
[0119] Next, we will explain the details of the process for calculating effectiveness.
[0120] The effectiveness V for selecting visible light images as training data can be calculated, for example, as a weighted linear sum of environmental factors using the following equation (Equation 1).
number
[0121] Here, w i (1≦i≦n) is the weighting coefficient, and F(i)(1≦i≦n) is the value that the environmental factor i can take. max (i) is the maximum value that the environmental factor i can take.
[0122] Weight coefficient w i The values of terms deemed important are set to be large. Also, each term has a value of F(i) max By dividing by (i), the values are kept within the range of 0 to 1. This normalization allows for the unified treatment of factors with different scales and units, enabling more accurate effectiveness evaluation. In other words, it prevents situations where a particular factor takes an extremely large value compared to other factors, thus preventing bias in the overall effectiveness calculation. Furthermore, setting the weight coefficients can be done intuitively and easily.
[0123] If the specific environmental factors are, for example, illuminance, weather, time of day, and speed of movement, then the effectiveness V will be shown in (Equation 2) below.
number
[0124] Illuminance is the amount of light received by the image processing device 100, as acquired by the illuminance acquisition unit 121. A higher illuminance is considered more effective for the visible light image capture environment. The maximum illuminance value is set to an appropriate value that is statistically possible depending on the shooting environment.
[0125] Weather refers to the amount of light received by the image processing device 100, which is acquired by the weather information acquisition unit 122. When it is raining or cloudy, the environment is not suitable for capturing visible light images, and conversely, when it is sunny, it is deemed suitable for capturing visible light images. Therefore, for example, it is defined as shown in (Equation 3) and (Equation 4) below.
number
[0126] The time period is information obtained by taking into account the calendar accessed by the image processing device 100 via an internet connection, the built-in clock, and location information (considering sunrise and sunset times). The effectiveness of the time period is defined such that the value increases when sunlight is strong, for example, as shown in (Equation 5) and (Equation 6) below.
number
[0127] Now, assuming we know the sunrise and sunset times at the shooting location, we can define each time period as follows:
[0128] Daytime: From about one hour after sunrise to about one hour before sunset
[0129] Dusk: From one hour before sunset until sunset
[0130] Nighttime: From sunset until one hour before sunrise
[0131] Early morning: From one hour before sunrise to one hour after sunrise
[0132] The values and time periods in (Equation 5) can be appropriately determined based on the latitude and longitude of the shooting location.
[0133] Furthermore, the movement speed factor is set to have a high effectiveness when the image processing device 100 is not moving, and a low effectiveness when the movement speed of the image processing device 100 is high. For this purpose, it is defined, for example, as shown in (Equation 7) and (Equation 8) below.
number
[0134] Here, v is the moving speed of the image processing device 100, which is acquired by the speed acquisition unit 126.
[0135] As described above, the image processing apparatus of this embodiment calculates the effectiveness of visible light image capture based on the environmental conditions at the time of shooting. Then, it selects visible light images with high effectiveness as training data and performs machine learning using that training data to learn how to correct thermal images. As a result, it is possible to perform appropriate image correction such as noise reduction and resolution improvement.
[0136] [Embodiment 2] Hereinafter, Embodiment 2 of the present invention will be described with reference to Figures 12 and 13. Figure 12 is an overall configuration diagram of the imaging system according to Embodiment 2. Figure 13 is a flowchart showing the output processing for thermal images according to Embodiment 2. In the inference shown in Figure 11 of Embodiment 1, a high-resolution thermal image can be generated by processing the thermal image based on the image parameters obtained in the learning process in Figure 9.
[0137] In this embodiment, thermal image correction using machine learning, as shown in Embodiment 1, is performed selectively. In this embodiment, whether or not to correct the thermal image 601 is determined based on the effectiveness level. That is, when the effectiveness level is low, it is considered that the environment is harsh for imaging, and the subject cannot be properly captured in the visible light image 501, so it is desirable to generate a high-resolution thermal image. On the other hand, when the effectiveness level is high, it is considered that the environment is not harsh, and a high-resolution image can be captured in the visible light image 501, so the thermal image 611 is not corrected by inference. This reduces the load on data processing.
[0138] The following description will focus on the differences between this embodiment and Embodiment 1.
[0139] In terms of system configuration, as shown in Figure 11, the difference from the configuration in Figure 2 of Embodiment 1 is that in the image processing device 100, the effectiveness calculated by the effectiveness calculation unit 106 is referenced by the machine learning unit 105 (effectiveness calculation unit 106 → machine learning unit 105).
[0140] In the output processing for thermal images, the process is branched based on effectiveness, following the machine learning inference process, to determine whether or not to correct the thermal image.
[0141] First, the machine learning unit 105 of the image processing device 100 acquires a thermal image that is a candidate for correction (S1000).
[0142] Next, the machine learning unit 105 of the image processing device 100 determines from the image whether or not the subject 701 is captured in the thermal image 601 (S1001). The determination method is, for example, whether or not there is a heat source. If the subject 701 is captured (S1001: YES), the process proceeds to S1002. If it is determined that the subject 701 is not captured in the thermal image 601 (S1001: NO), the process proceeds to S1006.
[0143] Next, when it is determined that the subject 701 is captured in the thermal image 601, the effectiveness calculation unit 106 of the image processing device 100 reads the environmental data 50 acquired by the environmental information acquisition unit 120 (S1002).
[0144] Next, the effectiveness calculation unit 106 of the image processing device 100 calculates the effectiveness based on the environmental data 50 read in S1003 (S1003). The method for calculating the effectiveness is the same as described in Embodiment 1.
[0145] Next, the machine learning unit 105 of the image processing device 100 determines whether the effectiveness calculated in S1003 is less than a predetermined threshold (S1004). If the effectiveness is less than the threshold (S1004: YES), the process proceeds to S1005; if the effectiveness is equal to or greater than the threshold (S1005: NO), the process proceeds to S1006.
[0146] If the effectiveness is below the threshold in S1004, the thermal image is corrected by machine learning inference processing, as shown in Figure 11 of Embodiment 1, and the thermal image is output (S1005). This is because when the effectiveness is below the threshold, the shooting environment is harsh, and correction is considered meaningful.
[0147] If subject 701 is not captured in thermal image 601 at S1001, or if the effectiveness is above the threshold at S1004, machine learning inference processing is not performed, and the thermal image is output (S1006). This is because there is no point in applying correction when subject 701 is not captured in thermal image 601, and when the effectiveness is above the threshold, it is considered that the shooting environment is good and the resolution of the thermal image is good.
[0148] As described above, according to the processing of this embodiment, the effectiveness level calculated based on environmental conditions determines whether or not to perform machine learning using inference. If it is determined to be unnecessary, machine learning using inference is not performed, and the image is not corrected. This eliminates the need for the system to generate extra data, thus reducing the amount of data and eliminating the need for unnecessary processing, thereby reducing the processing load on the system.
[0149] [Embodiment 3] Embodiment 3 will be described below with reference to Figures 14 and 15. Figure 14 shows an example of a visible light image whose contours have been corrected using machine learning. Figure 15 is a flowchart showing the process of storing thermal images as training data and performing machine learning.
[0150] Embodiment 1 showed an example where a visible light image was used as training data for a thermal image in order to improve the resolution of the thermal image using machine learning. In this embodiment, the opposite example will be described, where a thermal image is used as training data for a visible light image in order to improve the resolution of the visible light image.
[0151] The visible light image 501 shown in Figure 6A of Embodiment 1 was taken in a harsh environment and therefore lacked sufficient resolution. This embodiment aims to improve the resolution of the visible light image 501 by using the thermal image 601 as training data and applying machine learning to the visible light image 501.
[0152] Figure 14 shows the visible light image 511 after correcting the visible light image 501 using the thermal image 601 as training data. The corrected visible light image 511 in Figure 14 is an image with improved resolution compared to the visible light image 501 in Figure 6A, obtained through machine learning inference.
[0153] Next, using Figure 15, we will explain the process of storing thermal images as training data and performing machine learning.
[0154] First, the training data selection unit 104 of the image processing device 100 acquires a thermal image that will serve as a candidate for training data, along with a visible light image captured at the same time (S1100).
[0155] Next, the training data selection unit 104 of the image processing device 100 determines from the image whether or not the subject 701 is captured in the thermal image 601 (S1101). The determination method is, for example, whether or not there is a heat source. If the subject 701 is captured (S1101: YES), the process proceeds to S1102. If it is determined that the subject 701 is not captured in the thermal image 601 (S1101: NO), the thermal image is unsuitable as training data, and the process is terminated.
[0156] If the subject 701 is captured in the thermal image 601, it is determined from the image whether or not the subject 701 is identified in the visible light image 501 (S1102). One method of determination is to determine whether or not the subject 701 has distinctive feature points. That is, image analysis is performed on the visible light image 501 to detect edges and determine whether or not it has feature points based on its shape. If the specific subject is captured (S1102: YES), the process proceeds to S1103. If the specific subject is not captured (S1102: NO), the process ends.
[0157] Next, when it is determined that the subject 701 is captured in the visible light image 501, the effectiveness calculation unit 106 of the image processing device 100 reads the environmental data 50 acquired by the environmental information acquisition unit 120 (S1103).
[0158] Next, the effectiveness calculation unit 106 of the image processing device 100 calculates the effectiveness based on the environmental data 50 read in S1103 (S1104). The specific method for calculating the effectiveness is the same as in Embodiment 1.
[0159] Next, the training data selection unit 104 of the image processing device 100 determines whether the effectiveness calculated in S1104 is equal to or greater than a predetermined threshold (S1105). If the effectiveness is equal to or greater than the threshold (S1105: YES), the process proceeds to S1106; if the effectiveness is less than the threshold (S1105: NO), the process ends.
[0160] When the effectiveness is determined to be above a threshold, the training data selection unit 104 of the image processing device 100 stores the thermal image 601 as training data and the visible light image 501 as student data (S1106).
[0161] Next, the machine learning unit of the image processing device 100 performs machine learning using the stored training data and student data (S1107). The details of the machine learning method are the same as in Embodiment 1, except that the training data and student data are swapped.
[0162] In Embodiment 1, we described an example where the image data to be input into machine learning was processed to make it easier to learn from before inputting it. The same applies to the learning process in Embodiment 3.
[0163] As described above, the procedure shown in this embodiment makes it possible to obtain appropriate thermal images as training data.
[0164] In general, the threshold value for the determination in S805 in Figure 8 of Embodiment 1 and the threshold value for S1105 in Figure 15 of this embodiment may be different.
[0165] Furthermore, as a machine learning approach, the process of using the visible light image 501 shown in Figure 8 as training data and the thermal image 601 as student data may be performed simultaneously, as well as the process of using the thermal image 601 as training data and the visible light image 501 as student data, as shown in Figure 12. Additionally, at this time, the system may determine which image—the visible light image or the thermal image—should be used as training data based on its effectiveness, and select the appropriate image as training data.
[0166] Next, we will discuss examples of environmental factors that should be particularly considered when calculating the effectiveness of using thermal images as training data.
[0167] Invisible light cameras that capture thermal images detect subjects based on temperature, so the following environmental factors become important. (1) Temperature difference between the subject and the background
[0168] In thermal imaging, the greater the temperature difference between the subject and the background, the clearer the subject is recognized. Therefore, this temperature difference is a major factor in determining the effectiveness of the thermal image. In other words, the temperature of the subject and the temperature of the background are compared, the difference is calculated, and if the temperature difference is large, the effectiveness is high, and if it is small, the effectiveness is low. (2)Humidity
[0169] In high-humidity environments, the accuracy of thermal images can decrease. This is because water vapor in the air obstructs the transmission of infrared rays, resulting in blurred images; therefore, humidity is a crucial factor. For this reason, relative humidity is measured as a percentage (%), and the effectiveness is reduced as humidity increases and decreased as humidity decreases. (3) Wind speed
[0170] High wind speeds can cause rapid changes in the temperature of subjects and backgrounds, affecting the accuracy of invisible light cameras used to capture thermal images. In other words, strong winds cause subjects to change temperature more easily, reducing image accuracy. Therefore, the effectiveness of thermal images should be lowered as wind speeds increase, and higher at low wind speeds, i.e., gentle winds. (4) Time slot
[0171] Invisible light cameras used to capture thermal images can be used day or night, but during the day, the sun's heat affects the entire environment, which can result in smaller temperature differences compared to nighttime. This can slightly reduce the effectiveness of thermal images during the day. Therefore, based on the time of capture, the effectiveness is set lower during the day and higher at night.
[0172] As described above, in this embodiment, thermal images with high effectiveness are selected as training data, and machine learning is performed using this training data to learn how to correct visible light images. Therefore, appropriate image correction such as noise reduction and resolution improvement can be performed.
[0173] [Embodiment 4] Hereinafter, Embodiment 4 of the present invention will be described with reference to Figure 16. Figure 16 is a flowchart showing the output processing for a visible light image according to Embodiment 4.
[0174] In Embodiment 2, the decision of whether or not to correct the thermal image 601 in the machine learning inference of Embodiment 1 was made based on its effectiveness. Based on a similar concept, in Embodiment 3, the decision of whether or not to correct the visible light image 501 in the machine learning inference of Embodiment 3 is made based on its effectiveness.
[0175] The following description will focus on the differences between this embodiment and Embodiment 2. The system configuration is the same as shown in Figure 11 of Embodiment 2.
[0176] In the output processing for visible light images, the process is branched based on effectiveness, according to the machine learning inference process, to determine whether or not to correct the visible light image.
[0177] First, the machine learning unit 105 of the image processing device 100 acquires a visible light image 501 that is a candidate for correction (S1200).
[0178] Next, the machine learning unit 105 of the image processing device 100 determines from the image whether or not the subject 701 is captured in the visible light image 501 (S1201). For example, the determination method is to determine whether or not the subject is present by edge detection. If the subject 701 is captured (S1201: YES), proceed to S1202. If it is determined that the subject 701 is not captured in the visible light image 501 (S1201: NO), proceed to S1206.
[0179] Next, when it is determined that the subject 701 is captured in the visible light image 501, the effectiveness calculation unit 106 of the image processing device 100 reads the environmental data 50 acquired by the environmental information acquisition unit 120 (S1202).
[0180] Next, the effectiveness calculation unit 106 of the image processing device 100 calculates the effectiveness based on the environmental data 50 read in S1203 (S1203). The method for calculating the effectiveness is the same as described in Embodiment 1.
[0181] Next, the machine learning unit 105 of the image processing device 100 determines whether the effectiveness calculated in S1203 is less than a predetermined threshold (S1204). If the effectiveness is less than the threshold (S1204: YES), the process proceeds to S1205; if the effectiveness is equal to or greater than the threshold (S1205: NO), the process proceeds to S1206.
[0182] If the effectiveness is below the threshold in S1204, the visible light image 501 is corrected by machine learning inference processing, as shown in Figure 11 of Embodiment 1, and a thermal image is output (S1205). This is because when the effectiveness is below the threshold, the shooting environment is harsh, and it is considered meaningful to perform the correction.
[0183] If the subject 701 is not captured in the visible light image 501 in S1201, or if the effectiveness is above the threshold in S1204, machine learning inference processing is not performed, and the visible light image 501 is output (S1206). This is because there is no point in applying correction when the subject 701 is not captured in the visible light image 501, and when the effectiveness is above the threshold, it is considered that the shooting environment is good and the resolution of the thermal image is good.
[0184] As described above, according to the processing of this embodiment, similar to Embodiment 2, it is determined whether or not to perform machine learning using inference based on the effectiveness calculated from the environmental conditions. If it is determined that it is unnecessary, machine learning using inference is not performed, and the image is not corrected. This reduces the amount of data because the system does not need to generate extra data, and it also reduces the processing load on the system because unnecessary processing is not required.
[0185] [Embodiment 5] Hereinafter, Embodiment 5 of the present invention will be described with reference to Figure 17. Figure 17 is an overall configuration diagram of the imaging system according to Embodiment 5.
[0186] In Embodiment 1, the machine learning unit 105 is provided in the image processing device 100, and an example was shown in which learning and inference are performed in the image processing device 100 to improve the resolution of the thermal image. In this embodiment, the machine learning unit 105 is provided in the client device 200, and the same processing is performed on the client device 200.
[0187] As shown in Figure 17, the client device 200 of this embodiment includes an image processing unit 103a, a training data selection unit 104a, a machine learning unit 105a, and an effectiveness calculation unit 106a. These have the same functions as the components of the same name in Embodiment 1.
[0188] Furthermore, the control unit 101 of the image processing device 100 operates to transmit the environmental data acquired by the environmental information acquisition unit 120 to the client device 200 at appropriate intervals or when requested by the client device 200.
[0189] Since it is easier to implement the client device 200 as a high-performance PC or similar device and to make the CPU 301 high-performance than with the image processing device 100, it is expected that the processing speed will be improved in machine learning processing and other applications.
[0190] Furthermore, the client device 200 may be connectable to the server and may handle cloud data. Also, some of the functions may reside on the server or another personal computer connected to the server.
[0191] When transmitting images over a network to another device, compressed images are generally used. However, the machine learning unit 105 prefers to receive uncompressed images, which contain more information, rather than compressed images. Therefore, when the machine learning unit 105 is located in the client device 200, as in this embodiment, it is desirable to transmit uncompressed image data from the image processing device 100 to the client device 200, as long as the network bandwidth allows.
[0192] As described above, in this embodiment, machine learning is performed on a high-performance client device, making it possible to perform image correction processing more efficiently as a system.
[0193] (Composition 1) An image processing device that uses visible light images as training data and performs processing on invisible light images by machine learning, A machine learning unit that performs learning and inference to correct non-visible light images using training data, An environmental information acquisition unit that acquires information about the surrounding environment, Based on the environmental information acquired by the aforementioned environmental information acquisition unit, an effectiveness determination unit determines the effectiveness of the environmental information as training data for a visible light image, An image processing apparatus comprising: a training data selection unit that, based on the effectiveness determined by the effectiveness determination unit, determines whether the visible light image is effective as training data for a non-visible light image that temporally corresponds to the visible light image, and selects the training data that has been determined to be effective.
[0194] (Configuration 2) The image processing apparatus according to Configuration 1, characterized in that the environmental factors for determining the effectiveness of the aforementioned environmental information include at least one of the following: brightness at the time of shooting, weather information at the time of shooting, distance between the image processing apparatus and the subject in the image, orientation of the image processing apparatus, and movement speed of the image processing apparatus.
[0195] (Composition 3) The image processing apparatus according to either configuration 1 or configuration 2, characterized in that the effectiveness is determined as a weighted linear sum of the environmental factors.
[0196] (Composition 4) The image processing apparatus according to any one of configurations 1 to 3, characterized in that the training data selection unit determines, based on the effectiveness, whether to perform correction of the invisible light image by inference as the processing.
[0197] (Composition 5) The image processing apparatus according to any one of configurations 1 to 4, characterized in that the training data selection unit changes the weighting of data that can be treated as training data during learning based on the effectiveness.
[0198] (Composition 6) The image processing apparatus according to any one of configurations 1 to 5, characterized in that the machine learning unit learns image parameters that improve the image quality by the machine learning unit.
[0199] (Composition 7) The image processing apparatus according to configuration 6, characterized in that the machine learning unit generates an image with improved image quality.
[0200] (Composition 8) The image processing apparatus according to any one of configurations 1 to 6, characterized in that the machine learning unit corrects and outputs a non-visible light image.
[0201] (Configuration 9) An image processing device that uses a visible light image or a non-visible light image as training data and performs processing on a visible light image or a non-visible light image by machine learning, A machine learning unit that performs learning and inference for the above processing using training data, An environmental information acquisition unit that acquires information about the surrounding environment, Based on the environmental information acquired by the aforementioned environmental information acquisition unit, an effectiveness determination unit determines the effectiveness of visible light images or non-visible light images as training data, An image processing apparatus comprising: a training data selection unit that, based on the effectiveness determined by the effectiveness determination unit, determines whether the visible light image is directed to the invisible light image that corresponds to the visible light image in time, or determines whether the invisible light image is effective as training data for the visible light image that corresponds to the invisible light image in time, and selects the training data that has been determined to be effective.
[0202] (Composition 10) The image processing apparatus according to configuration 9, characterized in that the training data selection unit determines whether or not to use either a visible light image or a non-visible light image as training data based on its effectiveness.
[0203] (Composition 11) The image processing apparatus according to either configuration 9 or configuration 10, characterized in that the machine learning unit determines whether to perform correction of the visible light image by inference based on the effectiveness level.
[0204] (Composition 12) The image processing apparatus according to any one of configurations 9 to 11, characterized in that the machine learning unit processes the visible light image or non-visible light image selected as training data.
[0205] (Composition 13) The image processing apparatus according to any one of configurations 9 to 12, characterized in that the machine learning unit performs parallax correction when the distance between the image processing apparatus and the subject in the image is short when processing the visible light image or non-visible light image selected as training data.
[0206] (Composition 14) The image processing apparatus according to any one of configurations 9 to 13, characterized in that the training data selection unit labels the training data with the environmental information.
[0207] (Composition 15) The image processing apparatus according to any one of configurations 9 to 14, characterized in that the training data selection unit labels the training data by classifying the subjects in the image.
[0208] (Composition 16) The image processing apparatus according to claims 9 to 15, characterized in that the machine learning unit outputs a classification of a subject in a visible light image or a non-visible light image.
[0209] (Composition 17) The image processing apparatus according to any one of claims 9 to 16, characterized in that the machine learning unit corrects and outputs a visible light image or a non-visible light image.
[0210] (Method 1) An image processing method using an image processing device that performs processing on invisible light images by machine learning with visible light images as training data, The image processing device performs a machine learning step of learning and inference for the processing using training data, The image processing device includes an environmental information acquisition step of acquiring information about the surrounding environment, The image processing apparatus performs an effectiveness determination step, which determines the effectiveness of a visible light image as training data based on the environmental information acquired in the environmental information acquisition step, An image processing method characterized in that the image processing device determines, based on the effectiveness calculated in the effectiveness determination step, whether the visible light image is effective as training data for a non-visible light image that temporally corresponds to the visible light image, and selects the training data that has been determined to be effective.
[0211] (Method 2) An image processing method using an image processing device that performs processing on a visible light image or an invisible light image by machine learning, with the visible light image or invisible light image being used as training data, The image processing device performs a machine learning step of learning and inference for the processing using training data, The image processing device includes an environmental information acquisition step of acquiring information about the surrounding environment, The image processing apparatus performs an effectiveness determination step, which determines the effectiveness of a visible light image or a non-visible light image as training data based on the environmental information acquired in the environmental information acquisition step, An image processing method characterized in that the image processing device determines, based on the effectiveness determined in the effectiveness determination step, whether the visible light image is effective for a non-visible light image that is temporally corresponding to the visible light image, or whether the non-visible light image is effective as training data for a visible light image that is temporally corresponding to the non-visible light image, and selects the training data that has been determined to be effective.
[0212] (Program 1) A program for causing a computer to perform the image processing method described in Method 1 or Method 2. [Explanation of Symbols]
[0213] 10...Imaging system, 20...Network, 100...Image processing device, 200...Client device, 30…Visible light image data, 31…Invisible light image data, 40…Machine learning control data, 41…Image parameters, 50…Environmental data, 101...Control unit, 102...Communication unit, 103...Storage unit, 104,104a...Trainer data selection unit, 105,105a...Machine learning unit, 106,106a...Effectiveness calculation unit, 110...Imaging unit, 111...Visible light imaging unit, 112...Invisible light imaging unit, 113,113a...Image processing unit, 120...Environmental information acquisition unit, 121...Illuminance acquisition unit, 122...Weather information acquisition unit, 123...Distance acquisition unit, 124...Attitude acquisition unit, 125...Position acquisition unit, 126...Speed acquisition unit, 201...Control unit, 202...Communication unit, 203...Display unit, 205...Operation input unit, 206...Storage unit, 207...External device I / F unit, 301...CPU, 302...Main memory, 303...Non-volatile memory, 350...Communication I / F device, 310…Visible light imaging device, 311…Visible light imaging optical system, 312…Visible light lens, 313…Visible light image sensor, 320...Invisible light imaging device, 321...Invisible light imaging optical system, 322...Invisible light lens, 323...Invisible light image sensor, 330...Environmental information acquisition device group, 331...Illuminance sensor, 332...Wind speed sensor, 333...Barometric pressure sensor, 334...Temperature sensor, 335...Humidity sensor, 336...Rain amount sensor, 338...Gyroscope sensor, 339...Accelerometer, 340...Distance sensor, 341...GPS receiver, 342...Electronic compass, 401...CPU, 402...Main memory, 403...Non-volatile memory, 404...Display I / F device, 405...External device I / F device, 406...Communication I / F device, 407...Input / output I / F device, 410...Display device, 420...Keyboard, 421...Mouse
Claims
1. An image processing device that uses visible light images as training data and performs processing on invisible light images by machine learning, A machine learning unit that performs learning and inference for the above processing using training data, An environmental information acquisition unit that acquires information about the surrounding environment, Based on the environmental information acquired by the aforementioned environmental information acquisition unit, an effectiveness determination unit determines the effectiveness of visible light images as training data, An image processing apparatus comprising: a training data selection unit that, based on the effectiveness determined by the effectiveness determination unit, determines whether the visible light image is effective as training data for a non-visible light image that temporally corresponds to the visible light image, and selects the training data that has been determined to be effective.
2. The image processing apparatus according to claim 1, characterized in that the environmental factors for determining the effectiveness of the aforementioned environmental information include at least one of the following: brightness at the time of shooting, weather information at the time of shooting, distance between the image processing apparatus and the subject in the image, orientation of the image processing apparatus, and movement speed of the image processing apparatus.
3. The image processing apparatus according to claim 2, characterized in that the effectiveness is determined as a weighted linear sum of the environmental factors.
4. The image processing apparatus according to claim 1, characterized in that the training data selection unit determines, based on the effectiveness, whether to perform correction of the invisible light image by inference as the processing.
5. The image processing apparatus according to claim 1, characterized in that the training data selection unit changes the weighting of data that can be treated as training data during learning based on the effectiveness.
6. The image processing apparatus according to claim 1, characterized in that the machine learning unit learns image parameters that improve the image quality by the machine learning unit.
7. The image processing apparatus according to claim 6, characterized in that the machine learning unit generates an image with improved image quality.
8. The image processing apparatus according to claim 1, characterized in that the machine learning unit corrects and outputs a non-visible light image.
9. An image processing device that uses visible light images or invisible light images as training data and performs processing on visible light images or invisible light images by machine learning, A machine learning unit that performs learning and inference for the above processing using training data, An environmental information acquisition unit that acquires information about the surrounding environment, Based on the environmental information acquired by the aforementioned environmental information acquisition unit, an effectiveness determination unit determines the effectiveness of visible light images or non-visible light images as training data, An image processing apparatus comprising: a training data selection unit that, based on the effectiveness determined by the effectiveness determination unit, determines whether the visible light image is effective for a non-visible light image that is temporally corresponding to the visible light image, or determines whether the non-visible light image is effective as training data for a visible light image that is temporally corresponding to the non-visible image, and selects the training data that has been determined to be effective.
10. The image processing apparatus according to claim 9, characterized in that the training data selection unit determines whether to use a high-visibility image or a non-visible light image as training data based on its effectiveness.
11. The image processing apparatus according to claim 9, characterized in that the machine learning unit determines whether to perform correction of the visible light image by inference based on the effectiveness level.
12. The image processing apparatus according to claim 9, characterized in that the machine learning unit processes the visible light image or non-visible light image selected as training data.
13. The image processing apparatus according to claim 9, characterized in that the machine learning unit performs parallax correction when the distance between the image processing apparatus and the subject in the image is close when processing the visible light image or non-visible light image selected as training data.
14. The image processing apparatus according to claim 9, characterized in that the training data selection unit labels the training data with the environmental information.
15. The image processing apparatus according to claim 9, characterized in that the training data selection unit labels the training data by classifying the subjects in the image.
16. The image processing apparatus according to claim 9, characterized in that the machine learning unit outputs a classification of a subject in a visible light image or a non-visible light image.
17. The image processing apparatus according to claim 9, characterized in that the machine learning unit corrects and outputs a visible light image or a non-visible light image.
18. An image processing method using an image processing device that performs processing on invisible light images by machine learning with visible light images as training data, The image processing device performs a machine learning step of learning and inference for the processing using training data, The image processing device includes an environmental information acquisition step of acquiring information about the surrounding environment, The image processing apparatus performs an effectiveness determination step, which determines the effectiveness of a visible light image as training data based on the environmental information acquired in the environmental information acquisition step, An image processing method characterized in that the image processing device determines, based on the effectiveness calculated in the effectiveness determination step, whether the visible light image is effective as training data for a non-visible light image that temporally corresponds to the visible light image, and selects the training data that has been determined to be effective.
19. An image processing method using an image processing device that performs processing on a visible light image or an invisible light image by machine learning, with the visible light image or invisible light image being used as training data, The image processing device performs a machine learning step of learning and inference for the processing using training data, The image processing device includes an environmental information acquisition step of acquiring information about the surrounding environment, The image processing apparatus performs an effectiveness determination step, which determines the effectiveness of a visible light image or a non-visible light image as training data based on the environmental information acquired in the environmental information acquisition step, An image processing method characterized in that the image processing device determines, based on the effectiveness determined in the effectiveness determination step, whether the visible light image is effective for a non-visible light image that is temporally corresponding to the visible light image, or whether the non-visible light image is effective as training data for a visible light image that is temporally corresponding to the non-visible light image, and selects the training data that has been determined to be effective.
20. A program for causing a computer to execute the image processing method according to claim 18 or 19.