Recognition model, information processing system, information processing method, and recognition model generation method
Patent Information
- Application Number
- PCT/JP2025/007547
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-03-03
- Publication Date
- 2025-10-02
AI Technical Summary
The accuracy of object recognition is compromised when an image sensor captures both visible and infrared light due to saturation of visible light intensity by infrared light, leading to heterogeneous images that affect recognition models.
A method involving an information processing system that generates a subtraction image by subtracting infrared light intensity from visible light intensity in specific regions, using a recognition model trained on both subtraction and virtual images to improve recognition accuracy.
Enhances object recognition accuracy by addressing image heterogeneity caused by infrared light saturation, even when using sensors capturing both light types.
Smart Images

Figure JP2025007547_02102025_PF_FP_ABST
Abstract
Description
Recognition model, information processing system, information processing method, and recognition model generation method Cross-reference to related applications
[0001] This application claims priority from Japanese Patent Application No. 2024-33280 (filed March 5, 2024), the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to a recognition model, an information processing system, an information processing method, and a recognition model generation method.
[0003] As described in Patent Document 1, a device is known that detects the position of an object using an image of the object captured using visible light and an image of the object captured using infrared light.
[0004] International Publication No. 2020 / 174623
[0005] A recognition model according to one embodiment of the present disclosure includes an input unit for inputting an infrared image capturing the infrared light component of an object, and a subtraction image obtained by subtracting the detected intensity of the infrared light in a subtraction region of a visible image capturing a composite component combining the infrared light component and the visible light component of the object, the subtraction region including pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold, and an output unit for outputting the result of recognizing the object from the infrared image and the subtraction image.
[0006] An information processing system according to one embodiment of the present disclosure includes an acquisition unit that acquires an infrared image capturing the infrared light component of an object and a visible image capturing a composite component combining the infrared light component and the visible light component of the object; a subtraction unit that generates a subtraction image by subtracting the infrared light detection intensity in a subtraction region of the visible image that includes pixels where the infrared light detection intensity is equal to or less than a subtraction threshold; and a recognition unit that recognizes the object using a recognition model that outputs a result of recognizing the object from the infrared image and the subtraction image when the infrared image and the subtraction image are input.
[0007] An information processing method according to one embodiment of the present disclosure includes acquiring an infrared image capturing the infrared light component of an object and a visible image capturing a composite component combining the infrared light component and the visible light component of the object, generating a subtraction image by subtracting the detected intensity of the infrared light in a subtraction region of the visible image that includes pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold, and recognizing the object using a recognition model that outputs a result of recognizing the object from the infrared image and the subtraction image when the infrared image and the subtraction image are input.
[0008] A recognition model generation method according to one embodiment of the present disclosure includes acquiring an infrared image capturing the infrared light component of an object and a visible image capturing a composite component combining the infrared light component and the visible light component of the object, generating a subtraction image in which the detected intensity of the infrared light is subtracted in a subtraction region of the visible image that includes pixels in which the detected intensity of the infrared light is equal to or less than a subtraction threshold, and performing learning using the infrared image and the subtraction image as learning data, thereby generating a recognition model that outputs a result of recognizing the object when the infrared image and the subtraction image are input.
[0009] 5A is a block diagram illustrating an example of a schematic configuration of an information processing system according to an embodiment. FIG. 5B is a diagram illustrating generation of an infrared image and a visible image from a captured image. FIG. 6A is a graph illustrating an example of sensitivity characteristics of an imaging device. FIG. 6B is a graph illustrating an example of detection intensities of infrared light and each of the RGB colors. FIG. 6C is a graph illustrating subtraction intensities obtained by subtracting the detection intensity of infrared light from the detection intensities of each of the RGB colors in FIG. 4A. FIG. 6D is a graph illustrating an example of detection intensities of infrared light and each of the RGB colors when the detection intensities of each of the RGB colors exceed the upper detection limit. FIG. 6E is a graph illustrating subtraction intensities obtained by subtracting the detection intensity of infrared light from the upper detection limit, which is the detection intensity of each of the RGB colors in FIG. 5A. FIG. 6F is a diagram illustrating an example of an infrared image. FIG. 6G is a diagram illustrating a subtraction image. FIG. 6H is a diagram illustrating the relationship between the position of the boundary of a subtraction region and the magnitude of a subtraction threshold. FIG. 6H is an example of a virtual image generated using a virtual threshold that is smaller than the subtraction threshold used to generate the subtraction image in FIG. 6B. FIG. 6I is an example of a virtual image generated by further reducing the virtual threshold used to generate the virtual image in FIG. 8A. FIG. 6I is a block diagram illustrating an example of a learning flow for generating a recognition model. FIG. 6I is a block diagram illustrating an example of a processing flow for recognizing an object using a recognition model. 10 is a flowchart illustrating an example of a procedure of a recognition model generation method according to an embodiment;
[0010] When capturing an image of an object using an image sensor that detects both visible light and infrared light, the detected intensity of visible light at some pixels may be saturated by the infrared light. The accuracy of recognizing an object from a visible image that includes pixels with saturated detected intensity of visible light is lower than the accuracy of recognizing an object from a visible image that does not include pixels with saturated detected intensity of visible light. There is a need to improve the accuracy of recognizing objects. According to a recognition model, an information processing system, an information processing method, and a recognition model generation method according to an embodiment of the present disclosure, the accuracy of recognizing objects is improved.
[0011] An image sensor that captures visible light is also sensitive to infrared light, but by using a filter that attenuates infrared light, it is possible to capture an image of an object using visible light so that the infrared light component is not detected. On the other hand, an image sensor that has both pixels that detect visible light and pixels that detect infrared light cannot use a filter that attenuates infrared light to capture infrared light, and the pixels that detect visible light detect both visible light and infrared light components. As a result, an image captured by the pixels that detect visible light is affected by the infrared light component.
[0012] In order to remove the influence of infrared light components from an image captured by pixels that detect visible light, it is conceivable to subtract the detected intensity of infrared light by pixels that detect infrared light from the detected intensity by pixels that detect visible light. However, when light with an intensity exceeding the detection upper limit is incident on a pixel, the detected intensity of light by that pixel becomes saturated. Even if the intensity of the infrared light component is subtracted from the saturated detected intensity, only an intensity that is lower than the original intensity of the visible light component by the amount of the intensity exceeding the detection upper limit is obtained. In other words, the original intensity of the visible light component is not reproduced.
[0013] When the light detection intensity is saturated in only some pixels of an image, the image becomes a heterogeneous image that includes pixels for which the intensity of the original visible light component is calculated and pixels for which the intensity of the original visible light component is not calculated. When recognizing an object using a model from an image capturing an infrared light component and an image capturing a combined component of infrared light and visible light components from which the infrared light component is subtracted, if the image input to the model becomes heterogeneous, the recognition accuracy of the object will be lower than if the image were not heterogeneous.
[0014] Hereinafter, in this disclosure, an example of an embodiment of an information processing system 1, an information processing device 10 (see FIG. 1), and an information processing method that can improve the recognition accuracy of an object when an image input to a model becomes heterogeneous will be described.
[0015] (Configuration Example of Information Processing System 1) As shown in FIG. 1 , an information processing system 1 according to an embodiment of the present disclosure includes an information processing device 10 and an imaging device 20.
[0016] <Information Processing Device 10 > The information processing device 10 includes an acquisition unit 12 , a subtraction unit 14 , a recognition unit 16 , and an output unit 18 .
[0017] The acquisition unit 12 acquires image data from the image capture device 20. The acquisition unit 12 may acquire various other data or information. The acquisition unit 12 may include a communication interface for wired or wireless communication with the image capture device 20 or other devices. The communication interface may be configured to be capable of communication using a communication method based on various communication standards. The communication interface may be configured based on known communication technology.
[0018] The acquisition unit 12 may include an input device that accepts input from a user. The input device may include, for example, a keyboard or physical keys, or a pointing device such as a touch panel, a touch sensor, or a mouse. The input device is not limited to these examples and may include various other devices. The acquisition unit 12 may be configured to be able to communicate with an external input device.
[0019] The subtraction unit 14 generates a subtraction image, which will be described later. The recognition unit 16 recognizes objects in the image and outputs the recognition result. The subtraction unit 14 and the recognition unit 16 may each include at least one processor to provide control and processing capabilities for performing various functions. The functions of the subtraction unit 14 and the recognition unit 16 may be implemented by one processor or several processors. The subtraction unit 14 and the recognition unit 16 may be integrated. The combined functions of the subtraction unit 14 and the recognition unit 16 may be implemented by one processor. The processor may be implemented as a single integrated circuit (IC). The processor may be implemented as multiple communicatively connected integrated circuits or discrete circuits. The processor may also be implemented based on various other known technologies.
[0020] The processor may include a general-purpose processor that loads a specific program to execute a specific function, or a dedicated processor specialized for a specific process. The general-purpose processor may include, for example, a central processing unit (CPU) or a digital signal processor (DSP). The dedicated processor may include an application-specific integrated circuit (ASIC). The processor may include a programmable logic device (PLD). The PLD may include a field-programmable gate array (FPGA). The subtraction unit 14 and the recognition unit 16 may include either a system-on-a-chip (SoC) or a system-in-a-package (SiP) in which one or more processors work together.
[0021] The information processing device 10 may include a storage unit. The storage unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The storage unit stores various information. The storage unit stores programs executed by a processor or the like that functions as the subtraction unit 14 or the recognition unit 16. The storage unit may be configured as a non-transitory readable medium. The storage unit may function as a work memory for the subtraction unit 14 or the recognition unit 16. At least a portion of the storage unit may be configured integrally with the subtraction unit 14 or the recognition unit 16.
[0022] The output unit 18 outputs the recognition result of the object by the recognition unit 16. The output unit 18 may include a display device such as a display. The display may include various types of displays such as an LCD (Liquid Crystal Display), an organic EL (Electro-Luminescence) display, or an inorganic EL display. The recognition unit 16 may display an image of the recognition result of the object on the display device. The recognition unit 16 may display an infrared image or a visible image, or an image in which the recognition result of the object is superimposed on these images, on the display device.
[0023] The output unit 18 may include an audio output device such as a speaker. The recognition unit 16 may output audio from the audio output device to notify that the object has been recognized. The output unit 18 is not limited to these examples and may include various other devices.
[0024] The information processing device 10 may be mounted on a moving object such as a vehicle, or may be mounted on a device such as a roadside device used in a transportation system. The information processing device 10 may be mounted on a moving object or device together with the image capturing device 20. The information processing device 10 may be installed in a location separate from the image capturing device 20.
[0025] <Photographing Device 20> The photographing device 20 has an imaging element that includes both pixels that capture infrared light and pixels that capture visible light. The photographing device 20 generates a captured image that includes both pixels that represent the detected intensity of infrared light and pixels that represent the detected intensity of visible light. The pixels that capture infrared light detect light in a wavelength range of, for example, 780 nm to 1000 nm. The pixels that capture visible light detect light in a wavelength range of, for example, 380 nm to 780 nm.
[0026] As illustrated in Fig. 2 , the captured image includes pixels that capture infrared light represented by I and pixels that capture visible light in the three colors of red, green, and blue represented by R, G, and B. The image capturing device 20 performs demosaicing on the captured image. Demosaicing is a process that separates pixels that are mixed in the captured image to generate multiple images. By performing demosaicing on the captured image, the image capturing device 20 generates an infrared image made up of pixels that capture infrared light represented by I and a visible image made up of pixels that capture visible light in the three colors of red, green, and blue represented by R, G, and B.
[0027] As shown in Fig. 3, the wavelength characteristics of sensitivity in a pixel capturing infrared light have a peak in the infrared band. On the other hand, the wavelength characteristics of sensitivity in a pixel capturing each of the R, G, and B visible light have peaks in both the R, G, and B wavelengths included in the visible band and in the infrared band. In other words, a pixel capturing visible light detects not only visible light components but also infrared light components. Therefore, the detection intensity corresponding to each pixel of the visible image is the detection intensity of a component that combines visible light components and infrared light components.
[0028] The imaging element may be, for example, a charge coupled device image sensor (CCD) or a complementary metal oxide semiconductor (CMOS) sensor.
[0029] The number of image capturing devices 20 is not limited to one and may be two or more. The image capturing device 20 may be mounted on a moving object such as a vehicle, or on a device such as a roadside unit used in a transportation system. The image capturing device 20 may be mounted on a moving object or device together with the information processing device 10.
[0030] (Example of operation of information processing system 1)
[0031] The acquisition unit 12 of the information processing device 10 acquires an infrared image and a visible image from the imaging device 20. As described above, each pixel of the visible image represents the detection intensity of a component that is a combination of a visible light component and an infrared light component. The combination of a visible light component and an infrared light component is also referred to as a composite component. As illustrated in FIG. 4A , the detection intensity at a pixel corresponding to each of the colors red, green, and blue represented by R, G, and B is a value obtained by adding the detection intensity of the infrared light component represented by IR to the detection intensity of the red, green, and blue components when an image of an object is captured using only visible light.
[0032] The subtraction unit 14 of the information processing device 10 calculates a subtraction intensity by subtracting the detection intensity of each pixel of the corresponding infrared image from the detection intensity of each pixel of the visible image. The subtraction intensity is the intensity obtained by subtracting the detection intensity of the infrared light component from the detection intensity at the pixel corresponding to each color of red, green, and blue. The subtraction unit 14 generates a subtraction image based on the calculated subtraction intensity of each pixel. The subtraction image is an image in which each pixel represents a subtraction intensity.
[0033] In Fig. 4A, the detection intensities of red, green, and blue are equal to or less than the upper detection limit. When the detection intensities of each color are equal to or less than the upper detection limit, i.e., when the detection intensities of each color are not saturated, the subtraction unit 14 can calculate the subtraction intensities as the detection intensities of red, green, and blue when the target object is imaged using only visible light, as shown in Fig. 4B.
[0034] On the other hand, as illustrated in FIG. 5A , the detection intensity of the combined component of the infrared light component and the red, green, and blue color components, i.e., the combined component, may exceed the detection upper limit. In other words, the detection intensity of the combined component may become saturated. A region including pixels where the detection intensity of the combined component becomes saturated is also called a saturated region. When the detection intensity of the combined component becomes saturated in at least some pixels of the visible image, the visible image has a saturated region.
[0035] When the detection intensity of the composite component is saturated, the subtraction intensity of each of the red, green, and blue colors calculated by the subtraction unit 14 differs from the detection intensity of the red, green, and blue components when the object is imaged using only visible light. For example, as shown in FIG. 5B , the detection intensity of each of the red, green, and blue colors when the object is imaged using only visible light is the combined intensity of the solid-line rectangle and the dashed-line rectangle. Here, the intensity represented by the dashed-line rectangle is the intensity of the component that exceeds the upper detection limit in FIG. 5A . Therefore, the subtraction intensity of each of the red, green, and blue colors is the intensity represented by only the solid-line rectangle, excluding the dashed-line rectangle.
[0036] When the infrared light component is strong, the detected intensity of visible light is likely to saturate. In the present disclosure, the subtraction unit 14 calculates subtraction intensities for pixels with weak infrared light components and does not calculate subtraction intensities for pixels with strong infrared light components. For example, the detected intensity of infrared light in the infrared image shown in FIG. 6A is strong in the portions of pixels depicting a human as the target, represented as "Infrared: Strong," and weak in the portions represented as "Infrared: Weak." In this case, the subtraction unit 14 calculates subtraction intensities for portions with weak infrared light detection intensities and does not calculate subtraction intensities for portions with strong infrared light detection intensities, thereby generating the subtraction image shown in FIG. 6B.
[0037] The subtraction unit 14 may compare the detected intensity of infrared light with a subtraction threshold, and determine to calculate a subtraction intensity when the detected intensity of infrared light is equal to or less than the subtraction threshold. That is, the subtraction unit 14 may set a subtraction threshold for determining whether to perform subtraction processing for each pixel. The subtraction unit 14 performs subtraction processing to calculate a subtraction intensity in a subtraction region determined based on the subtraction threshold. The subtraction region is a region including pixels whose detected intensity of infrared light is equal to or less than the subtraction threshold. The subtraction region may include at least a portion of the saturated region.
[0038] In the subtraction image, the intensity of the portion where the subtraction intensity was not calculated, represented by "no subtraction," is significantly different from the intensity of the portion where the subtraction intensity was calculated, represented by "subtraction." In particular, the intensity of the pixels on both sides of the boundary between the "no subtraction" portion and the "subtraction" portion changes abruptly. In an image that depicts a human as the object, the detection intensity rarely changes abruptly in the range in which the human is captured. In other words, the subtraction image in FIG. 6B is a heterogeneous image. In the present disclosure, a heterogeneous image is an image in which the detection intensity changes abruptly within the range in which the object is captured in the subtraction image.
[0039] The recognition unit 16 of the information processing device 10 uses a recognition model to recognize objects appearing in the infrared image and the subtraction image. The recognition model is configured to output a result of recognizing objects appearing in the infrared image and the subtraction image when the infrared image and the subtraction image are input. The recognition model may include an input unit and an output unit. The input unit is a unit for inputting the infrared image and the subtraction image. The output unit is a unit for outputting the recognition result of the object. The recognition model may be a trained model obtained by performing training using the infrared image and the subtraction image as training data.
[0040] Here, as described above, the subtraction image may be an image in which the detection intensity changes suddenly among pixels that represent the object, i.e., a heterogeneous image. The recognition model may be a trained model obtained by performing training using subtraction images including heterogeneous images as training data.
[0041] The area of white pixels in the subtraction image that represent the target, i.e., pixels with a high detection intensity, changes depending on the magnitude relationship between the infrared light detection intensity and the subtraction threshold. As illustrated in Figure 7, the boundary when the subtraction threshold is set to a value smaller than a predetermined value, i.e., the "low threshold" boundary represented by the white dashed line, moves in a direction that increases the area of white pixels in the pixels that represent the target, i.e., pixels with a high detection intensity, relative to the boundary when the subtraction threshold is set to a predetermined value, represented by the gray line. Conversely, the boundary when the subtraction threshold is set to a value larger than the predetermined value, i.e., the "high threshold" boundary represented by the black dashed line, moves in a direction that decreases the area of white pixels in the pixels that represent the target, i.e., pixels with a high detection intensity, relative to the boundary when the subtraction threshold is set to a predetermined value, represented by the gray line.
[0042] As described above, in the subtraction image, the area of white pixels representing the object, i.e., pixels with high detection intensity, changes according to the relationship between the detected intensity of infrared light and the subtraction threshold. In other words, even a small change in the detected intensity of infrared light changes the shape of white pixels representing the object, i.e., pixels with high detection intensity, in the subtraction image.
[0043] The recognition model tends to detect areas of white pixels in the subtraction image as objects. Therefore, even if a recognition model is generated by performing learning using subtraction images generated by setting the subtraction threshold to a predetermined value as learning data, the accuracy of object recognition by the recognition model will decrease if the shape of white pixels among the pixels representing the object in the subtraction image, i.e., pixels with high detection intensity, changes. In other words, even a small change in the detection intensity of infrared light will decrease the accuracy of object recognition by the recognition model.
[0044] Therefore, in addition to the actual subtraction image, a virtual image may be used as learning data for generating a recognition model. The virtual image is an image in which the shape of white pixels, i.e., pixels with a large detection intensity, among pixels capturing an object is virtually changed. The virtual image may be generated by performing subtraction processing in a virtual subtraction region determined based on a virtual threshold. The virtual threshold is a value obtained by virtually changing the subtraction threshold. The virtual subtraction region is a region including pixels in which the detection intensity of infrared light is equal to or less than the virtual threshold.
[0045] For example, as shown in Fig. 8A, by setting a virtual threshold that is smaller than the subtraction threshold, a virtual image may be generated in which the area of white pixels is larger than the boundary in the actual subtraction image represented by the gray line. Also, as shown in Fig. 8B, by setting a virtual threshold that is even smaller than the subtraction threshold, a virtual image may be generated in which the area of white pixels is even larger.
[0046] The relationship between each image and the learning of the recognition model in the generation of the above-mentioned recognition model will be described with reference to the block diagram shown in FIG. 9. The subtraction image is generated by performing a subtraction process based on an infrared image and a visible image. The virtual image is generated by performing a subtraction process based on an infrared image and a visible image, setting a virtual threshold instead of a subtraction threshold. The recognition model is generated by performing learning using the infrared image, the subtraction image, and the virtual image as learning data. The learning data may further include a visible image.
[0047] In the block diagram of Fig. 9, a virtual image does not necessarily have to be generated. If a virtual image is not generated, a recognition model is generated by performing learning using an infrared image and a subtraction image as learning data.
[0048] The information processing device 10 recognizes an object from an infrared image and a visible image using the recognition model generated as described above. The recognition process of the information processing device 10 will be described with reference to the block diagram shown in FIG. 10 . The acquisition unit 12 acquires an infrared image and a visible image. The subtraction unit 14 generates a subtraction image by performing subtraction processing based on the infrared image and the visible image. The recognition unit 16 inputs the infrared image and the subtraction image to the recognition model and obtains the recognition result of the object output from the recognition model. The recognition unit 16 may further input a visible image to the recognition model.
[0049] The output unit 18 of the information processing device 10 may output the recognition result of the object. The output unit 18 may superimpose the recognition result of the object on the infrared image, the visible image, or the subtraction image and display it.
[0050] The information processing device 10 may acquire a recognition model from an external device. The information processing device 10 may generate a recognition model by itself by executing a recognition model generation method including the steps of the flowchart illustrated in FIG. 11. The information processing device 10 may generate a recognition model using the recognition unit 16. The information processing device 10 may further include a generation unit for generating the recognition model. The recognition model generation method may be realized as a recognition model generation program executed by a processor included in the information processing device 10. The recognition model generation program may be stored on a non-transitory computer-readable medium.
[0051] The acquisition unit 12 acquires an infrared image and a visible image from the imaging device 20 (step S1). The subtraction unit 14 generates a subtraction image by performing subtraction processing in a subtraction region based on the infrared image and the visible image (step S2). The subtraction unit 14 generates a virtual image by performing subtraction processing in a virtual subtraction region (step S3). The recognition unit 16 or the generation unit generates a recognition model by performing learning using the infrared image, the visible image, the subtraction image, and the virtual image as training data (step S4). After performing the procedure of step S4, the information processing device 10 ends the execution of the procedure of the flowchart in FIG. 11.
[0052] The information processing device 10 may recognize an object from an infrared image and a visible image using a recognition model by executing an information processing method including the steps of the flowchart illustrated in Fig. 12. The information processing method may be realized as an information processing program executed by a processor included in the information processing device 10. The information processing program may be stored in a non-transitory computer-readable medium.
[0053] The acquisition unit 12 acquires an infrared image and a visible image from the imaging device 20 (step S11). The subtraction unit 14 generates a subtraction image by performing subtraction processing based on the infrared image and the visible image (step S12). The recognition unit 16 inputs the infrared image and the subtraction image to a recognition model (step S13). The recognition unit 16 may further input a visible image to the recognition model. The recognition unit 16 acquires a recognition result of the object output from the recognition model (step S14). After executing the procedure of step S14, the information processing device 10 ends execution of the procedure of the flowchart in FIG. 12. After executing the procedure of acquiring the recognition result in step S14, the information processing device 10 may output the recognition result of the object via the output unit 18.
[0054] (Summary) As described above, according to the information processing system 1, information processing device 10, and information processing method disclosed herein, an object is recognized from a subtraction image generated by subtracting an infrared light component from a visible image captured using a composite component that combines an infrared light component and a visible light component. By recognizing the object from the subtraction image, the accuracy of object recognition can be improved even when an image sensor that captures both infrared light and visible light is used.
[0055] Furthermore, by using a recognition model generated by learning using heterogeneous images as training data, the accuracy of object recognition can be improved even when the subtracted image is a heterogeneous image.Furthermore, by using a recognition model generated by learning using virtual images as training data, the accuracy of object recognition can be improved even when various changes occur, such as changes in the detected intensity of infrared light.
[0056] In the above-described embodiments, the information processing device 10 or the image capturing device 20 may be mounted on a moving object such as a vehicle, or on a device such as a roadside device used in a transportation system. The information processing device 10 or the image capturing device 20 may be mounted on, for example, a robot or a robot controller. The information processing device 10 or the image capturing device 20 may be installed in a space where the robot operates.
[0057] The drawings illustrating the embodiments of the present disclosure are schematic, and the dimensional ratios and the like in the drawings do not necessarily correspond to the actual ones.
[0058] Although the embodiments according to the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art could make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to be logically inconsistent, and multiple components can be combined into one or divided. It should be understood that these modifications are also included in the scope of the present disclosure.
[0059] All of the features described in this disclosure and / or all steps of all of the disclosed methods or processes may be combined in any combination except combinations in which these features are mutually exclusive. Furthermore, each feature described in this disclosure may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly denied. Thus, unless expressly denied, each disclosed feature is only one example of a generic series of identical or equivalent features.
[0060] Furthermore, embodiments of the present disclosure are not limited to the specific configurations of any of the above-described embodiments, but rather extend to any novel feature or combination thereof described herein, or any novel method or process step or combination thereof described herein.
[0061] Vehicles according to the present disclosure may include, for example, automobiles, industrial vehicles, rail vehicles, lifestyle vehicles, or fixed-wing aircraft that travel on runways. Automobiles may include, for example, passenger cars, trucks, buses, motorcycles, or trolleybuses. Industrial vehicles may include, for example, industrial vehicles for agriculture or construction. Industrial vehicles may include, for example, forklifts or golf carts. Industrial vehicles for agriculture may include, for example, tractors, cultivators, transplanters, binders, combines, or lawnmowers. Industrial vehicles for construction may include, for example, bulldozers, scrapers, excavators, crane trucks, dump trucks, or road rollers. Vehicles may include vehicles that are powered by human power. Vehicle classifications are not limited to the above examples. For example, automobiles may include industrial vehicles that can travel on roads. Vehicles of the same type may be included in multiple classifications.
[0062] The above has described an embodiment of an information processing method using the information processing system 1, but embodiments of the present disclosure can also be embodied as a storage medium on which a program is recorded (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.), in addition to a method or program for implementing the device.
[0063] Furthermore, the implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, but may also be in the form of a program module incorporated into an operating system. Furthermore, the program may or may not be configured so that all processing is performed solely by the CPU on the control board. The program may also be configured so that part or all of it is executed by another processing unit mounted on an expansion board or expansion unit added to the board as needed.
[0064] In one embodiment, (1) the recognition model includes an input unit for inputting an infrared image capturing the infrared light component of an object, and a subtraction image obtained by subtracting the detected intensity of the infrared light in a subtraction region including pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold from a visible image capturing a composite component combining the infrared light component and the visible light component of the object, and an output unit for outputting the result of recognizing the object from the infrared image and the subtraction image.
[0065] (2) In the recognition model described in (1), the visible image may have a saturated region including pixels in which the detected intensity of the composite component is greater than the upper detection limit of an image sensor that captured the composite component. The subtracted image may be an image obtained by subtracting the intensity of the infrared light in at least a part of the saturated region.
[0066] In one embodiment, (3) the information processing system includes an acquisition unit that acquires an infrared image capturing the infrared light component of an object and a visible image capturing a composite component combining the infrared light component and the visible light component of the object; a subtraction unit that generates a subtraction image by subtracting the infrared light detection intensity in a subtraction region of the visible image that includes pixels where the infrared light detection intensity is equal to or less than a subtraction threshold; and a recognition unit that recognizes the object using a recognition model that outputs a result of recognizing the object from the infrared image and the subtraction image when the infrared image and the subtraction image are input.
[0067] In one embodiment, (4) the information processing method includes acquiring an infrared image capturing the infrared light component of an object and a visible image capturing a composite component combining the infrared light component and the visible light component of the object, generating a subtraction image by subtracting the detected intensity of the infrared light in a subtraction region of the visible image including pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold, and recognizing the object using a recognition model that outputs a result of recognizing the object from the infrared image and the subtraction image when the infrared image and the subtraction image are input.
[0068] In one embodiment, (5) the recognition model generation method includes acquiring an infrared image capturing the infrared light component of an object and a visible image capturing a composite component combining the infrared light component and the visible light component of the object, generating a subtraction image in which the detected intensity of the infrared light is subtracted in a subtraction region of the visible image that includes pixels in which the detected intensity of the infrared light is equal to or less than a subtraction threshold, and performing learning using the infrared image and the subtraction image as learning data, thereby generating a recognition model that outputs a result of recognizing the object when the infrared image and the subtraction image are input.
[0069] (6) The recognition model generation method described in (5) above may further include: setting a virtual threshold that is a modified subtraction threshold; subtracting the intensity of the infrared light in a virtual subtraction region of the visible image where the intensity of the infrared light is equal to or less than the virtual threshold to generate a virtual image; and generating the recognition model by performing learning using the virtual image as further learning data.
[0070] 1 Information processing system 10 Information processing device (12: Acquisition unit, 14: Subtraction unit, 16: Recognition unit, 18: Output unit) 20 Imaging device
Claims
1. A recognition model comprising: an input unit for inputting an infrared image capturing the infrared light component of an object; and a subtraction image obtained by subtracting the detected intensity of the infrared light in a subtraction region including pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold from a visible image capturing a composite component combining the infrared light component and the visible light component of the object; and an output unit for outputting the result of recognizing the object from the infrared image and the subtraction image.
2. The recognition model according to claim 1, wherein the visible image has a saturated region including pixels in which the detected intensity of the composite component is greater than the upper detection limit of the imaging element that captured the composite component, and the subtracted image is an image in which the intensity of the infrared light has been subtracted in at least a portion of the saturated region.
3. An information processing system comprising: an acquisition unit that acquires an infrared image capturing the infrared light component of an object, and a visible image capturing a composite component combining the infrared light component and the visible light component of the object; a subtraction unit that generates a subtraction image by subtracting the detected intensity of the infrared light in a subtraction region of the visible image that includes pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold; and a recognition unit that recognizes the object using a recognition model that outputs a result of recognizing the object from the infrared image and the subtraction image when the infrared image and the subtraction image are input.
4. An information processing method comprising: acquiring an infrared image capturing the infrared light component of an object; and a visible image capturing a composite component combining the infrared light component and the visible light component of the object; generating a subtraction image by subtracting the detected intensity of the infrared light in a subtraction region of the visible image that includes pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold; and recognizing the object using a recognition model that outputs a result of recognizing the object from the infrared image and the subtraction image when the infrared image and the subtraction image are input.
5. A recognition model generation method comprising: acquiring an infrared image capturing an infrared light component of an object; and a visible image capturing a composite component combining the infrared light component and the visible light component of the object; generating a subtraction image by subtracting the detected intensity of the infrared light in a subtraction region of the visible image that includes pixels where the detected intensity of the infrared light is equal to or less than a subtraction threshold; and performing learning using the infrared image and the subtraction image as learning data, thereby generating a recognition model that outputs a result of recognizing the object when the infrared image and the subtraction image are input.
6. The recognition model generating method according to claim 5, further comprising: setting a virtual threshold that is a modified version of the subtraction threshold; and generating a virtual image by subtracting the intensity of the infrared light in a virtual subtraction region of the visible image where the intensity of the infrared light is equal to or less than the virtual threshold; and generating the recognition model by performing learning using the virtual image as further learning data.