Information processing method, information processing device, and information processing system

The information processing system improves object detection accuracy by generating integrated data through cross-attention between infrared and visible images, addressing positional shifts and reducing computational load, thereby enhancing detection precision.

WO2025173596A1PCT designated stage Publication Date: 2025-08-21KYOCERA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/003604
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-13
Filing Date
2025-02-04
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing systems face challenges in accurately detecting objects in images captured at different wavelengths, such as infrared and visible light, due to positional shifts and misalignments between the images, leading to reduced detection accuracy and increased computational load.

Method used

An information processing system that generates integrated data by performing cross-attention between infrared and visible images, setting patches only in relevant regions of the visible image, and considering positional deviation factors to improve detection accuracy while reducing computational load.

Benefits of technology

Enhances detection accuracy by focusing on relevant image regions and minimizing computational requirements, allowing for precise object detection across different wavelength images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025003604_21082025_PF_FP_ABST
    Figure JP2025003604_21082025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing method comprises: acquiring a first image and a second image obtained by capturing, at different wavelengths, a detection target in a real space; setting a plurality of partial regions obtained by dividing the first image into a plurality of regions; setting, on the basis of a factor that causes a shift in position of an object captured in each of the images, a plurality of reference regions corresponding to the respective partial regions of the first image as destination regions to which the respective partial regions are to be shifted in the second image; generating integrated data on the basis of the respective partial regions of the first image and the respective reference regions of the second image; and detecting the detection target from the first image on the basis of the first image and the integrated data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing device, and information processing system Cross-reference to related applications

[0001] This application claims priority from Japanese Patent Application No. 2024-019668 (filed February 13, 2024), the entire disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to an information processing method, an information processing device, and an information processing system.

[0003] A system is known that improves detection accuracy by integrating the results of detecting an object from multiple images of the object to be detected, such as visible images taken with a visible camera and infrared images taken with an infrared camera (see Patent Document 1).

[0004] International Publication No. 2019 / 220622

[0005] An information processing method according to one embodiment of the present disclosure includes acquiring a first image and a second image of a detection target in real space photographed at different wavelengths, setting a plurality of partial regions by dividing the first image into a plurality of regions, setting a plurality of reference regions in the second image corresponding to each of the plurality of partial regions of the first image as regions to which the partial regions are shifted based on factors that cause the position of an object depicted in each image to be shifted, generating integrated data based on each of the partial regions of the first image and each of the reference regions of the second image, and detecting the detection target from the first image based on the first image and the integrated data.

[0006] According to an embodiment of the present disclosure, an information processing device includes an acquisition unit and a control unit. The acquisition unit acquires a first image and a second image of a detection target captured in real space at different wavelengths. The control unit divides the first image into a plurality of regions to set a plurality of partial regions, and based on a cause of a positional shift of an object captured in each image, sets a plurality of reference regions in the second image corresponding to each of the partial regions in the first image as regions to be shifted from each of the partial regions, generates integrated data based on each of the partial regions in the first image and each of the reference regions in the second image, and detects the detection target from the first image based on the first image and the integrated data.

[0007] According to an embodiment of the present disclosure, an information processing system includes an imaging device and an information processing device. The imaging device captures a first image and a second image of a detection target in real space at different wavelengths. The information processing device acquires the first image and the second image, divides the first image into a plurality of regions to set a plurality of partial regions, and sets a plurality of reference regions in the second image corresponding to each of the partial regions in the first image as regions to be shifted from each of the partial regions based on a cause of a positional shift of an object depicted in each image. The information processing device generates integrated data based on each of the partial regions in the first image and each of the reference regions in the second image, and detects the detection target from the first image based on the first image and the integrated data.

[0008] 1 is a block diagram showing a schematic configuration example of an information processing system according to an embodiment;

[0023] FIG. 1 is a diagram showing an infrared image and a visible image of the same detection target, where only a portion of the detection target is captured in the visible image;

[0024] FIG. 2 is a diagram showing an infrared image and a visible image of the same detection target, where almost no detection target is captured in the visible image;

[0025] FIG. 3 is a block diagram showing an example of the flow of information processing;

[0026] FIG. 4 is a diagram showing a portion that is commonly detected from the infrared image and the visible image;

[0027] FIG. 5 is a schematic diagram showing a misalignment in the inclination of the optical axes of two image capturing devices;

[0028] FIG. 6 is a schematic diagram showing the baseline lengths of the optical axes of the two image capturing devices;

[0029] FIG. 7 is a diagram explaining an example configuration when one image capturing element includes pixels that detect visible light and pixels that detect infrared light;

[0029] FIG. 8 is a diagram explaining an operation of reading out an infrared image and a visible image from one image capturing element in a time-division manner;

[0029] FIG. 9 is a diagram showing, by means of superimposed images, the misalignment of the positions of vehicles captured in the infrared image and the visible image;

[0029] FIG. 10 is a diagram showing an example of the correspondence between a patch set in an infrared image and a reference area set in the visible image so as to correspond to the patch;

[0029] FIG. 11 is a diagram explaining cross-attention between a patch set in an infrared image and a reference area set in the visible image. Fig. 3 is a diagram illustrating cross-attention between a patch set in the infrared image and a reference region set in the visible image in the infrared image and the visible image in Fig. 2. Fig. 4 is a flowchart illustrating an example of a procedure of an information processing method according to an embodiment.

[0009] 1 , an information processing system 1 according to an embodiment of the present disclosure includes an information processing device 10, a first image capturing device 20, and a second image capturing device 30. In the present disclosure, the first image capturing device 20 generates an infrared image by capturing an image of a real space using infrared light. The second image capturing device 30 generates a visible image by capturing an image of the real space using visible light.

[0010] The information processing device 10 detects a detection target 80 (see FIG. 2 or FIG. 3, etc.) that appears in an infrared image or a visible image. The detection target 80 is an object that is the target of detection by the information processing device 10. Here, as illustrated in FIGS. 2 and 3, the detection target 80 may appear differently in the infrared image and the visible image.

[0011] In the infrared image of Fig. 2, the entire body of the human being, who is the detection target 80, is clearly visible. On the other hand, in the visible image of Fig. 2, the amount of visible light received by the second image capture device 30 is insufficient, so only parts of the body, such as the head and feet, are clearly visible.

[0012] In the infrared image of Fig. 3, the entire body of the human being, who is the detection target 80, is clearly visible. On the other hand, in the visible image of Fig. 3, the human being, who is the detection target 80, is barely visible because fog or smoke or the like blocks visible light.

[0013] In this way, when the detection target 80 is clearly visible in the infrared image but is only clearly visible in the visible image, the information processing device 10 detects the detection target 80 primarily from the infrared image, while also referring to the visible image to improve detection accuracy.

[0014] The information processing device 10 generates integrated data by performing cross-attention between the infrared image and the visible image.

[0015] Cross-attention between an infrared image and a visible image may be performed according to the following steps (1) to (3).

[0016] (1) The information processing device 10 sets at least one patch in the infrared image and the visible image. A patch is a partial region of an image and is also called a partial region. The information processing device 10 may divide the image into multiple regions and set each region as a patch. Each patch in the infrared image and the visible image may be set as a query, a key, or a value in the cross-attention terminology. For example, a patch in the infrared image may be set as a query and a value in the cross-attention terminology. A patch in the visible image may be set as a key in the cross-attention terminology.

[0017] (2) The information processing device 10 calculates the inner product of a query role vector corresponding to a patch set in the infrared image and a key role vector corresponding to a patch set in the visible image. The vector corresponding to the patch is generated by vectorizing the data of pixels included in the patch.

[0018] (3) The information processing device 10 calculates a dot-product weight vector sum that aggregates the information of the visible image and the infrared image by combining each vector of the value role corresponding to the patch set in the infrared image with the scalar value of the dot product calculated in the procedure (2) above as a weight. The information processing device 10 uses the dot-product weight vector sum calculated using the calculated patch of the infrared image as a query and value and the patch of the visible image as a key as integrated data. In other words, the information processing device 10 generates integrated data by calculating the dot-product weight vector sum of the infrared image and the visible image.

[0019] The information processing device 10 detects the detection target 80 from the infrared image and the visible image based on the integrated data generated by the above procedure. The integrated data is data that has been weighted based on the relevance between the visible image and the infrared image, and reflects areas where the same object is likely to appear in the visible image and the infrared image. By referencing the integrated data that reflects the relationship between the infrared image and the visible image in addition to the infrared image, the information processing device 10 can improve the accuracy of detecting the detection target 80 from the infrared image compared to detecting the detection target 80 from only the infrared image.

[0020] In the above-described step (1) of cross-attention between an infrared image and a visible image, patches are set across the entire infrared image and the entire visible image. In step (2), inner products are calculated for all combinations of patches set across the entire infrared image and the entire visible image. In this case, the computational load required to calculate the inner products increases depending on the image size.

[0021] Here, the information processing device 10 according to the present disclosure does not set patches over the entire visible image, but sets patches only in a portion of the visible image. Specifically, the information processing device 10 sets a reference region in the visible image and sets patches within the reference region. The reference region is a portion of the visible image. The information processing device 10 calculates the dot product of a combination of the patches set within the reference region and the patches set in the infrared image.

[0022] By setting a patch only in a reference region that is a portion of the visible image, a large weight is assigned to a portion of the visible image that is likely to contain the detection target 80. Conversely, portions of the visible image that are likely to be unrelated to the detection target 80 are ignored. In other words, integrated data is generated so as to emphasize data of portions of the visible image that are highly related to the detection target 80. By detecting the detection target 80 from the infrared image and the visible image based on the integrated data generated in this manner, detection accuracy is improved.

[0023] Furthermore, by limiting the number of patches to a reference region, which is a part of the visible image, the number of combinations for calculating the dot product is limited. As a result, the computational load required to calculate the dot product is reduced. In other words, the computational load required to perform cross-attention is reduced.

[0024] Note that the setting of the query, key, or value for each patch of the visible image and the infrared image is not limited to the above example. For example, when using the visible image as a base, the patch of the visible image may be set as the query and value, and the patch of the infrared image may be set as the key. The setting of the query, key, or value for each patch of the visible image and the infrared image may be configured to use cross-attention with different conditions simultaneously in parallel.

[0025] Hereinafter, an embodiment for realizing the information processing device 10 according to the present disclosure maintaining or improving the detection accuracy will be described. Also, an embodiment for further realizing reduction in the calculation load or suppression of an increase in the calculation load will be described.

[0026] (Configuration Example of Information Processing System 1) As described above, the information processing system 1 includes the information processing device 10, the first image capturing device 20, and the second image capturing device 30. A configuration example of each part of the information processing system 1 will be described below.

[0027] <Information Processing Device 10> As shown in FIG. 1, the information processing device 10 includes an acquisition unit 12, a control unit 14, and an output unit 16.

[0028] The acquisition unit 12 acquires image data from the first image capture device 20 and the second image capture device 30. The acquisition unit 12 may acquire various other data or information. The acquisition unit 12 may include a communication interface for wired or wireless communication with the first image capture device 20 or the second image capture device 30, or with another device. The communication interface may be configured to be capable of communication using a communication method based on various communication standards. The communication interface may be configured based on known communication technology.

[0029] The acquisition unit 12 may include an input device that accepts input from a user. The input device may include, for example, a keyboard or physical keys, or a pointing device such as a touch panel, a touch sensor, or a mouse. The input device is not limited to these examples and may include various other devices. The acquisition unit 12 may be configured to be able to communicate with an external input device.

[0030] The controller 14 may include at least one processor to provide control and processing power for performing various functions. The functions of the controller 14 may be implemented by one processor or several processors. The processor may be implemented as a single integrated circuit (IC). The processor may be implemented as multiple communicatively connected integrated circuits or discrete circuits. The processor may be implemented based on various other known technologies.

[0031] The processor may include a general-purpose processor that loads a specific program to execute a specific function, or a dedicated processor specialized for a specific process. The general-purpose processor may include, for example, a central processing unit (CPU) or a digital signal processor (DSP). The dedicated processor may include an application-specific integrated circuit (ASIC). The processor may include a programmable logic device (PLD). The PLD may include a field-programmable gate array (FPGA). The control unit 14 may include either a system-on-a-chip (SoC) or a system in a package (SiP) in which one or more processors work together.

[0032] The information processing device 10 may include a storage unit. The storage unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The storage unit stores various information. The storage unit stores programs to be executed by a processor or the like that functions as the control unit 14. The storage unit may be configured as a non-transitory readable medium. The storage unit may function as a work memory for the control unit 14. At least a portion of the storage unit may be configured integrally with the control unit 14.

[0033] The output unit 16 outputs the detection result of the detection target 80. The output unit 16 may include a display device such as a display. The display may include various types of displays such as an LCD (Liquid Crystal Display), an organic EL (Electro-Luminescence) display, or an inorganic EL display. The control unit 14 may display an image of the detection result of the detection target 80 on the display device. The control unit 14 may display an infrared image or a visible image, or an image in which the detection result of the detection target 80 is superimposed on these images, on the display device.

[0034] The output unit 16 may include an audio output device such as a speaker. The control unit 14 may output audio from the audio output device to notify that the detection target 80 has been detected. The output unit 16 is not limited to these examples and may include various other devices.

[0035] The information processing device 10 may be mounted on a moving object such as a vehicle, or on a device such as a roadside device used in a transportation system. The information processing device 10 may be mounted on a moving object or a device together with the first image capturing device 20 and the second image capturing device 30. The information processing device 10 may be installed in a location separate from the first image capturing device 20 and the second image capturing device 30.

[0036] <First Image Capturing Device 20 and Second Image Capturing Device 30> The first image capturing device 20 includes an image capturing element configured to capture light having wavelengths within a first wavelength range, and generates a first image. In the present disclosure, the first wavelength range is the wavelength range of infrared light. The wavelength range of infrared light is, for example, from 780 nm to 1000 nm. The first image is an infrared image captured using infrared light. The infrared image may be an image that represents the intensity distribution of infrared light in a gray scale or a color scale.

[0037] The second image capturing device 30 includes an image capturing element configured to capture light having wavelengths within a second wavelength range, and generates a second image. In the present disclosure, the second wavelength range is the wavelength range of visible light. The wavelength range of visible light is, for example, from 380 nm to 780 nm. The second image is a visible image captured using visible light. The visible image may be an RGB image.

[0038] The first and second images are images of the detection target 80 in real space captured at different wavelengths.

[0039] The imaging element may be, for example, a charge coupled device image sensor (CCD) or a complementary metal oxide semiconductor (CMOS) sensor.

[0040] The number of the first imaging device 20 and the second imaging device 30 is not limited to one and may be two or more. The first imaging device 20 and the second imaging device 30 may be integrated into a single imaging device that can generate both infrared images and visible images. The first imaging device 20 and the second imaging device 30 may be mounted on a moving object such as a vehicle, or on an apparatus such as a roadside unit used in a transportation system. The first imaging device 20 and the second imaging device 30 may be mounted on a moving object or apparatus together with the information processing apparatus 10.

[0041] (Example of operation of information processing system 1) As described above, in the information processing system 1 according to the present disclosure, the information processing device 10 is configured to achieve both maintenance of detection accuracy and reduction of calculation load. An example of operation of the information processing system 1 according to the present disclosure will be described below.

[0042] The control unit 14 of the information processing device 10 may process information in accordance with the flow of the block diagram illustrated in FIG.

[0043] The control unit 14 uses an object detection model to detect the detection target 80 from the infrared image. The object detection model is configured to receive the infrared image and the integrated data. The object detection model outputs the result of detecting the detection target 80 from the infrared image based on the infrared image and the integrated data. The object detection model may be configured as a model of various types, such as a trained model generated by machine learning, a model that performs pattern matching, or a database-type model that combines images and detection results.

[0044] As described above, the integrated data represents an area where the same object is likely to appear in both the visible image and the infrared image. By using the integrated data, the control unit 14 can detect the detection target 80 from the infrared image by referring to the visible image. By referring to the visible image, the control unit 14 can improve the accuracy of detecting the detection target 80 from the infrared image.

[0045] An example of why the detection accuracy can be improved is described below. As illustrated in Fig. 5 , when a detection target 80 appears in an infrared image, the control unit 14 can detect the detection target 80 from the infrared image by recognizing feature points 82 and 83 that appear in the infrared image. When the control unit 14 detects the detection target 80, it may display a detection frame 81 that surrounds the detection target 80 as the detection result.

[0046] Here, the control unit 14 refers to the visible image to confirm the validity of detecting the detection target 80 based on the recognition of the feature points 82 and 83. When the control unit 14 determines that the feature points 84 and 85 in the visible image correspond to the feature points 82 and 83 in the infrared image, the control unit 14 may decide to detect the detection target 80 including the feature points 82 and 83 from the infrared image. On the other hand, when the control unit 14 determines that the feature points 84 and 85 in the visible image do not correspond to the feature points 82 and 83 in the infrared image, the control unit 14 may decide not to detect the detection target 80 from the infrared image.

[0047] In other words, the integrated data is data that reflects the results of referring to the visible image when detecting the detection target 80 from the infrared image. The control unit 14 generates the integrated data by performing cross-attention between the infrared image and the visible image. The integrated data generated by cross-attention is data that reflects information from the visible image, i.e., data that reflects the results of referring to the visible image when detecting the detection target 80 from the infrared image.

[0048] The control unit 14 sets at least one patch in the infrared image in preparation for cross-attention for generating integrated data. The control unit 14 sets a reference region in the visible image corresponding to the patch set in the infrared image. The control unit 14 sets one reference region in the visible image corresponding to one patch set in the infrared image. The reference regions corresponding to each of the multiple patches set in the infrared image may be set so as to at least partially overlap in the visible image, or may be set so as not to overlap each other. The reference regions corresponding to each of the multiple patches set in the infrared image may be set in different regions of the visible image, or may be set in the same region of the visible image.

[0049] The reference area set in the visible image is an area in which an object related to an object shown in the patch set in the infrared image is likely to appear in the visible image. If the infrared image and the visible image are images of exactly the same scenery, the control unit 14 may set a reference area in the visible image at the same position and size as the patch set in the infrared image.

[0050] However, the infrared image and the visible image may not be images of exactly the same scenery, that is, there may be a discrepancy in the position where the detection target 80 appears between the infrared image and the visible image.

[0051] The deviation in the position where the detection target 80 is captured may occur due to a deviation in the arrangement of the first image capturing device 20 and the second image capturing device 30 .

[0052] 6A , when the first imaging device 20 and the second imaging device 30 are arranged side by side, the optical axis 21 of the first imaging device 20 and the optical axis 31 of the second imaging device 30 may not be parallel. Specifically, the optical axis 31 of the second imaging device 30 is tilted with respect to the parallel line 21P of the optical axis 21 of the first imaging device 20. In this case, the position of the detection target 80 appearing in the infrared image captured by the first imaging device 20 may differ from the position of the detection target 80 appearing in the visible image captured by the second imaging device 30.

[0053] Furthermore, as illustrated in FIG. 6B , even if the optical axis 21 of the first imaging device 20 and the optical axis 31 of the second imaging device 30 are parallel, the distance between the optical axes 21 and 31 may cause a misalignment between the position of the detection target 80 in the infrared image captured by the first imaging device 20 and the position of the detection target 80 in the visible image captured by the second imaging device 30. The distance between the optical axes 21 and 31 is also referred to as the baseline length. The longer the baseline length, the greater the misalignment between the position of the detection target 80 in the infrared image and the position of the detection target 80 in the visible image. The first imaging device 20 and the second imaging device 30 have housings of finite size. Therefore, it is difficult to shorten the baseline length. As a result, it is difficult to eliminate the positional misalignment caused by the baseline length.

[0054] The positional deviation of the detection target 80 may occur due to differences in the times at which the images were captured. For example, if the first image capturing device 20 and the second image capturing device 30 are integrated into a single image capturing device, the image capturing device may not be able to capture an infrared image and a visible image simultaneously. In this case, the time at which the infrared image was captured differs from the time at which the visible image was captured.

[0055] An imaging device in which the first imaging device 20 and the second imaging device 30 are integrated may include an imaging element as exemplified in Fig. 7A. The imaging element includes pixels represented by R, G, and B, and a pixel represented by I. The pixels represented by R, G, and B are pixels for detecting visible light. The pixel represented by I is a pixel for detecting infrared light.

[0056] As illustrated in Fig. 7B , the image capturing device reads out the infrared light detection results and the visible light detection results in a time-division manner to generate an infrared image and a visible image, respectively. At time T1, the image capturing device reads out the infrared light detection results from the pixel represented by I to generate an infrared image. At the next time T2, the image capturing device reads out the visible light detection results from the pixels represented by R, G, and B to generate a visible image. At the next time T3, the image capturing device reads out the infrared light detection results from the pixel represented by I to generate an infrared image. At the next time T4, the image capturing device reads out the visible light detection results from the pixels represented by R, G, and B to generate a visible image.

[0057] Here, if the detection target 80 is moving, the position at which the detection target 80 appears in each image captured at each time may change. In other words, the position of the detection target 80 appearing in the infrared image captured at time T1 may differ from the position of the detection target 80 appearing in the visible image captured at time T2.

[0058] An example of the positional deviation of the detection target 80 is shown in Figure 8. The deviation between the position of the vehicle in the infrared image and the position of the vehicle in the visible image is shown by a superimposed image. The superimposed image is an image in which the infrared image and the visible image are superimposed after enlarging the area enclosed by the dashed line frame.

[0059] As described above, there may be a discrepancy in the position where the detection target 80 appears between the infrared image and the visible image. Returning to Fig. 4, the control unit 14 sets a reference area in the visible image in consideration of the discrepancy in the position where the detection target 80 appears, i.e., the positional deviation factor. The positional deviation factor is a factor that causes a discrepancy in the position of an object appearing in each image, and specifically, a factor that causes a discrepancy between the position of an object appearing in the infrared image and the position of an object appearing in the visible image.

[0060] The positional deviation factors include deviation in the arrangement of the first image capturing device 20 and the second image capturing device 30. The deviation in the arrangement of the first image capturing device 20 and the second image capturing device 30 is identified based on the state in which the first image capturing device 20 and the second image capturing device 30 are actually installed.

[0061] The positional deviation factor includes a difference between the time when the infrared image is captured and the time when the visible image is captured, i.e., a time difference between the captures. If the time when the infrared image is captured and the time when the visible image is captured differ, the detection target 80 moves within the angle of view of the imaging device, causing a deviation in the position at which the detection target 80 is captured between the infrared image and the visible image. The detection target 80 may move within the angle of view of the imaging device as the detection target 80 itself moves. The detection target 80 may move within the angle of view of the imaging device as the imaging device moves. The detection target 80 may move within the angle of view of the imaging device as the direction in which the imaging device captures images changes. Therefore, the positional deviation factor may include information regarding the speed at which the detection target 80 moves within the angle of view of the imaging device.

[0062] When the imaging device is mounted on a vehicle, the speed at which the detection target 80 moves within the field of view of the imaging device is determined based on the driving state of the vehicle. The position deviation factor may include information related to the driving state of the vehicle. The driving state of the vehicle includes the speed or direction of the vehicle. The speed or direction of the vehicle may be measured by an acceleration sensor, an angular velocity sensor, an inertial measurement unit (IMU), or the like mounted on the vehicle.

[0063] The speed at which the detection target 80 moves within the angle of view of the imaging device is determined based on the movement speed of the detection target 80 itself. The movement speed of the detection target 80 itself can be estimated based on the type and orientation of the detection target 80. For example, if the detection target 80 is a human, the movement speed of the detection target 80 may be estimated to be within a range considered to be a walking or running speed of a human. For example, if the detection target 80 is a vehicle, the movement speed of the detection target 80 may be estimated to be within a range considered to be a running speed of a vehicle. Therefore, the position deviation factor may include information regarding the type and orientation of the detection target 80.

[0064] The control unit 14 estimates, based on the positional deviation factor, how the position at which the detection target 80 appears in each image will be shifted if the detection target 80 appears in both the infrared image and the visible image. The control unit 14 sets, in the visible image, reference areas corresponding to the patches set in the infrared image based on the estimated shift in the position at which the detection target 80 appears. When setting multiple patches in the infrared image, the control unit 14 sets, based on the positional deviation factor, multiple reference areas corresponding to the multiple patches in the infrared image as areas shifted from each patch in the visible image.

[0065] 9 , for example, the control unit 14 sets an area in the visible image where an object appearing at the position of patch 41 set in the infrared image is likely to appear in the visible image as reference area 51. The control unit 14 also sets an area in the visible image where an object appearing at the position of patch 42 set in the infrared image is likely to appear in the visible image as reference area 52. The control unit 14 also sets an area in the visible image where an object appearing at the position of patch 43 set in the infrared image is likely to appear in the visible image as reference area 53. In other words, when multiple patches are set in the infrared image, the control unit 14 sets reference areas corresponding to each patch in the visible image.

[0066] When the detection target 80 is a moving vehicle in real space, the infrared image and the visible image may be images captured of an area including the moving vehicle in real space. In this case, the control unit 14 may acquire the state of the moving vehicle as a cause of the positional deviation. The state of the moving vehicle may include information about the speed or direction of the moving vehicle. The control unit 14 may set a reference area in the visible image based on the state of the moving vehicle.

[0067] The control unit 14 may set the reference area in the visible image by converting the patch set in the infrared image into a reference area to be set in the visible image using a transformation matrix. The transformation matrix is ​​a matrix that specifies the relationship between a vector representing the area of ​​the patch set in the infrared image and a vector representing the reference area to be set in the visible image. The control unit 14 may generate the transformation matrix based on a positional deviation factor. In other words, the transformation matrix is ​​a matrix that reflects the positional deviation factor. The control unit 14 may generate a transformation matrix that reflects at least one of a positional deviation between the first image capturing device 20 and the second image capturing device 30 or a difference between the time when the infrared image is captured and the time when the visible image is captured.

[0068] When the control unit 14 acquires an infrared image from the first imaging device 20 and a visible image from the second imaging device 30, the control unit 14 may acquire information representing a deviation between the first imaging device 20 and the second imaging device 30. The deviation between the first imaging device 20 and the second imaging device 30 may include a deviation between a position where the first imaging device 20 is installed and a position where the second imaging device 30 is installed. The deviation between the first imaging device 20 and the second imaging device 30 may include a deviation between a direction in which the first imaging device 20 captures images and a direction in which the second imaging device 30 captures images. The control unit 14 may consider the information representing the deviation between the first imaging device 20 and the second imaging device 30 as a positional deviation factor, and set a reference area in the visible image based on the deviation between the first imaging device 20 and the second imaging device 30.

[0069] When the infrared image and the visible image are acquired from a single image capturing device that captures both the infrared image and the visible image, the control unit 14 may acquire information on the time difference between capturing the first image and the second image in the image capturing device. The control unit 14 may consider the time difference between capturing the first image and the second image in the image capturing device as a positional deviation factor and set the reference area in the visible image based on the time difference between capturing the first image and the second image in the image capturing device.

[0070] Even when the control unit 14 acquires an infrared image from the first imaging device 20 and a visible image from the second imaging device 30, the control unit 14 may acquire information about the time difference between the time when the first image was captured by the first imaging device 20 and the time when the second image was captured by the second imaging device 30. The control unit 14 may set a reference area in the visible image based on the time difference between the time when the first image was captured by the first imaging device 20 and the time when the second image was captured by the second imaging device 30.

[0071] As described above, the control unit 14 sets the reference region in the visible image. Returning to Fig. 4, the control unit 14 performs cross-attention between the patch set in the infrared image and the reference region set in the visible image to generate integrated data.

[0072] 10 , for example, the control unit 14 sets a patch 44 in the infrared image, and sets a reference region 54 corresponding to the patch 44 in the visible image. The control unit 14 sets patches 61, 62, 63, and 64 within the reference region 54 in the visible image. The patch 44 in the infrared image and the patches 61, 62, 63, and 64 in the visible image are set to the same size so that their dot products can be calculated after vectorization. The control unit 14 calculates the dot products between the patch 44 in the infrared image and each of the patches 61, 62, 63, and 64 in the visible image, weights the data obtained by vectorizing the visible image according to the calculated dot product values, and calculates the dot product weight vector sum as integrated data.

[0073] The setting of the patch and reference region will be described based on the example of the infrared image and visible image shown in Fig. 11. The control unit 14 sets a patch 45 in the infrared image. The control unit 14 sets a reference region 55 as a region where an object appearing in the patch 45 of the infrared image is likely to also appear in the visible image. The control unit 14 sets a patch 65 within the reference region 55 in the visible image. The control unit 14 calculates the dot product of the patch 45 of the infrared image and the patch 65 of the visible image.

[0074] Returning to FIG. 4, the control unit 14 inputs the infrared image and the integrated data into the object detection model, and obtains the detection result of the detection target 80 output by the object detection model.

[0075] <Example of Procedure of Information Processing Method> The control unit 14 of the information processing device 10 may execute an information processing method including the procedure of the flowchart illustrated in Fig. 12. The information processing method may be realized as an information processing program executed by a processor constituting the control unit 14. The information processing program may be stored in a non-transitory computer-readable medium.

[0076] The control unit 14 acquires a first image and a second image (step S1). The first image may be an infrared image. The second image may be a visible image. When the first imaging device 20 and the second imaging device 30 are separate, the control unit 14 acquires the first image from the first imaging device 20 and the second image from the second imaging device 30. When the first imaging device 20 and the second imaging device 30 are integrated, the control unit 14 acquires the first image and the second image from one imaging device.

[0077] The control unit 14 sets at least one patch in the first image (step S2), sets reference regions in the second image corresponding to each of the at least one patch set in the first image (step S3), and generates integrated data by performing cross-attention between the patch set in the first image and the reference region set in the second image (step S4).

[0078] The control unit 14 detects the detection target 80 from the first image based on the first image and the integrated data (step S5). The control unit 14 may output the detection result of the detection target 80 from the output unit 16. After executing the procedure of step S5, the control unit 14 ends the execution of the flowchart in FIG. 12 .

[0079] (Summary) As described above, according to the information processing system 1, the information processing device 10, and the information processing method according to the present disclosure, integrated data is generated by performing cross-attention between the first image and the second image.

[0080] By detecting the target object 80 from the first image based on the first image and the integrated data, the accuracy of detecting the target object 80 from the first image is improved, regardless of whether the target object 80 can be detected from the second image.

[0081] In cross-attention, when calculating the dot product between a patch of a first image and a patch of a second image, the patch of the second image is not set over the entire second image, but is set only within a reference area corresponding to the patch of the first image.

[0082] By limiting the range in which patches are set in the second image, the weight of information that is effective for detecting the detection target 80 from the first image is increased compared to when patches are set over the entire second image, thereby improving the accuracy of detecting the detection target 80 from the first image.

[0083] Furthermore, when setting a reference area in the second image, the positional deviation factor is taken into consideration. By taking the positional deviation factor into consideration, the reference area is appropriately set in the visible image so as not to reduce the accuracy of detecting the detection target 80 from the first image.

[0084] As described above, the information processing system 1, the information processing device 10, and the information processing method according to the present disclosure can maintain or improve detection accuracy.

[0085] Furthermore, by limiting the range in which patches are set in the second image, the amount of calculation of the dot product is reduced compared to when patches are set over the entire second image. Reducing the amount of calculation of the dot product reduces the computational load for detecting the detection target 80. In other words, the information processing system 1, information processing device 10, and information processing method according to the present disclosure not only maintain or improve detection accuracy, but also reduce the computational load or suppress an increase in the computational load.

[0086] Furthermore, by utilizing the reduced computational load resulting from limiting the range in which patches are set in the second image, it is possible to further divide the first image into smaller areas and set patches. By setting patches more finely, detection accuracy is improved. In other words, the information processing system 1, information processing device 10, and information processing method according to the present disclosure achieve improved detection accuracy.

[0087] In the above-described embodiment, an infrared image is used as the first image and a visible image is used as the second image. However, various other images, such as a depth image representing the distance to the subject, may be used as the first image or the second image.

[0088] In the above-described embodiments, at least one of the information processing device 10, the first image capturing device 20, and the second image capturing device 30, or the image capturing devices that capture both the first image and the second image, may be mounted on a moving object such as a vehicle, or on a device such as a roadside unit used in a transportation system. At least one of the information processing device 10, the first image capturing device 20, and the second image capturing device 30, or the image capturing devices that capture both the first image and the second image, may be mounted on, for example, a robot or a robot controller. At least one of the information processing device 10, the first image capturing device 20, and the second image capturing device 30, or the image capturing devices that capture both the first image and the second image, may be installed in a space in which the robot operates.

[0089] The drawings illustrating the embodiments of the present disclosure are schematic, and the dimensional ratios and the like in the drawings do not necessarily correspond to the actual ones.

[0090] Although the embodiments according to the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art could make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to be logically inconsistent, and multiple components can be combined into one or divided. It should be understood that these modifications are also included in the scope of the present disclosure.

[0091] All of the features described in this disclosure and / or all steps of all of the disclosed methods or processes may be combined in any combination except combinations in which these features are mutually exclusive. Furthermore, each feature described in this disclosure may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly denied. Thus, unless expressly denied, each disclosed feature is only one example of a generic series of identical or equivalent features.

[0092] Furthermore, embodiments of the present disclosure are not limited to the specific configurations of any of the above-described embodiments, but rather extend to any novel feature or combination thereof described herein, or any novel method or process step or combination thereof described herein.

[0093] Vehicles according to the present disclosure may include, for example, automobiles, industrial vehicles, rail vehicles, lifestyle vehicles, or fixed-wing aircraft that travel on runways. Automobiles may include, for example, passenger cars, trucks, buses, motorcycles, or trolleybuses. Industrial vehicles may include, for example, industrial vehicles for agriculture or construction. Industrial vehicles may include, for example, forklifts or golf carts. Industrial vehicles for agriculture may include, for example, tractors, cultivators, transplanters, binders, combines, or lawnmowers. Industrial vehicles for construction may include, for example, bulldozers, scrapers, excavators, crane trucks, dump trucks, or road rollers. Vehicles may include vehicles that are powered by human power. Vehicle classifications are not limited to the above examples. For example, automobiles may include industrial vehicles that can travel on roads. Vehicles of the same type may be included in multiple classifications.

[0094] In this disclosure, descriptions such as "first" and "second" are identifiers for distinguishing the configuration. In this disclosure, the configurations distinguished by descriptions such as "first" and "second" can have their numbers swapped. For example, the identifiers "first" and "second" can be swapped between the first image and the second image. The identifier swapping is performed simultaneously. The configurations remain distinguished even after the identifier swapping. The identifiers may be deleted. A configuration from which the identifiers have been deleted is distinguished by a symbol. The identifiers "first" and "second" in this disclosure should not be used solely to interpret the order of the configurations or to justify the existence of an identifier with a smaller number.

[0095] The above has described an embodiment of an information processing method using the information processing system 1, but embodiments of the present disclosure can also be embodied as a storage medium on which a program is recorded (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.), in addition to a method or program for implementing the device.

[0096] Furthermore, the implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, but may also be in the form of a program module incorporated into an operating system. Furthermore, the program may or may not be configured so that all processing is performed solely by the CPU on the control board. The program may also be configured so that part or all of it is executed by another processing unit mounted on an expansion board or expansion unit added to the board as needed.

[0097] In one embodiment, (1) an information processing method includes acquiring a first image and a second image of a detection target in real space photographed at different wavelengths, setting a plurality of partial areas by dividing the first image into a plurality of regions, setting a plurality of reference areas in the second image corresponding to each of the plurality of partial areas of the first image as areas to which the partial areas are shifted based on factors that cause the position of an object depicted in each image to be shifted, generating integrated data based on each of the partial areas of the first image and each of the reference areas of the second image, and detecting the detection target from the first image based on the first image and the integrated data.

[0098] (2) In the information processing method described in (1) above, the method may include performing cross-attention to calculate an inner product between a vector corresponding to each of the plurality of partial regions and a vector corresponding to each of the plurality of reference regions in each of the partial regions of the first image and each of the reference regions of the second image, and generating as the integrated data an inner product weighted vector sum calculated by weighting the inner product and combining the vectors of the images of the plurality of reference regions.

[0099] (3) In the information processing method described in (1) or (2), the first image and the second image may be images captured of an area including a traveling vehicle in the real space, and the information processing method may further include acquiring a state of the traveling vehicle and setting the plurality of reference areas based on the state of the traveling vehicle.

[0100] (4) The information processing method described in (1) or (2) above may further include acquiring the state of a moving vehicle equipped with an imaging device that captures the first image and the second image, and setting the multiple reference areas based on the state of the moving vehicle.

[0101] (5) The information processing method described in any one of (1) to (4) above may further include generating a transformation matrix that transforms each of the partial regions of the first image into each of the reference regions of the second image.

[0102] (6) In the information processing method according to any one of (1) to (5), the first image may be an infrared image, and the second image may be a visible image.

[0103] (7) The information processing method described in any one of (1) to (6) above may further include acquiring the first image from a first photographing device and acquiring the second image from a second photographing device, and setting the plurality of reference areas based on a deviation between the first photographing device and the second photographing device.

[0104] (8) The information processing method described in any one of (1) to (7) above may further include acquiring the first image and the second image from a photographing device that captures both the first image and the second image, and setting the multiple reference areas based on the time difference between capturing the first image and the second image in the photographing device.

[0105] In one embodiment, (9) an information processing device includes an acquisition unit and a control unit. The acquisition unit acquires a first image and a second image of a detection target captured in real space at different wavelengths. The control unit sets a plurality of partial regions by dividing the first image into a plurality of regions, sets a plurality of reference regions in the second image corresponding to each of the partial regions in the first image as regions to be shifted from each of the partial regions based on a cause of a positional shift of an object captured in each of the images, generates integrated data based on each of the partial regions in the first image and each of the reference regions in the second image, and detects the detection target from the first image based on the first image and the integrated data.

[0106] In one embodiment, (10) an information processing system includes an imaging device and an information processing device. The imaging device captures a first image and a second image of a detection target in real space at different wavelengths. The information processing device acquires the first image and the second image, divides the first image into a plurality of regions to set a plurality of partial regions, sets a plurality of reference regions in the second image corresponding to each of the partial regions in the first image as regions to be shifted from each of the partial regions based on a cause of a positional shift of an object depicted in each image, generates integrated data based on each of the partial regions in the first image and each of the reference regions in the second image, and detects the detection target from the first image based on the first image and the integrated data.

[0107] 1 Information processing system 10 Information processing device (12: acquisition unit, 14: control unit, 16: output unit) 20 First image capturing device (21: optical axis, 21P: parallel line of optical axis) 30 Second image capturing device (31: optical axis) 41, 42, 43, 44, 45 Patches set in infrared image 51, 52, 53, 54, 55 Reference area 61, 62, 63, 64, 65 Patches set in visible image 80 Detection target 81 Detection frame 82, 83, 84, 85 Feature points

Claims

1. An information processing method comprising: acquiring a first image and a second image of a detection target in real space photographed at different wavelengths; dividing the first image into a plurality of regions to set a plurality of partial regions; setting a plurality of reference regions in the second image corresponding to each of the plurality of partial regions of the first image as regions to which the object in each image is shifted based on factors that cause the position of the object to shift; generating integrated data based on each of the partial regions of the first image and each of the reference regions of the second image; and detecting the detection target from the first image based on the first image and the integrated data.

2. The information processing method of claim 1, further comprising: performing cross-attention to calculate the dot product of a vector corresponding to each of the plurality of partial regions and a vector corresponding to each of the plurality of reference regions in each of the plurality of partial regions of the first image and each of the plurality of reference regions; and generating as the integrated data a dot product weighted vector sum calculated by weighting the dot product and combining the vectors of the images of the plurality of reference regions.

3. The information processing method described in claim 1 or 2, wherein the first image and the second image are images captured of an area including a moving vehicle in the real space, and further comprising: acquiring the state of the moving vehicle; and setting the multiple reference areas based on the state of the moving vehicle.

4. An information processing method as described in claim 1 or 2, further comprising: acquiring the state of a moving vehicle equipped with an imaging device that captures the first image and the second image; and setting the multiple reference areas based on the state of the moving vehicle.

5. An information processing method according to any one of claims 1 to 4, further comprising generating a transformation matrix that transforms each of the partial regions of the first image into each of the reference regions of the second image.

6. An information processing method according to any one of claims 1 to 5, wherein the first image is an infrared image and the second image is a visible image.

7. An information processing method described in any one of claims 1 to 6, further comprising: acquiring the first image from a first image capturing device and acquiring the second image from a second image capturing device; and setting the plurality of reference areas based on a deviation between the first image capturing device and the second image capturing device.

8. An information processing method described in any one of claims 1 to 7, further comprising: acquiring the first image and the second image from a photographing device that captures both the first image and the second image; and setting the multiple reference areas based on the time difference between capturing the first image and the second image in the photographing device.

9. An information processing device comprising an acquisition unit and a control unit, wherein the acquisition unit acquires a first image and a second image of a detection target in real space photographed at different wavelengths, the control unit sets a plurality of partial areas by dividing the first image into a plurality of regions, sets a plurality of reference areas in the second image corresponding to each of the plurality of partial areas of the first image as areas to which each partial area is shifted based on factors that cause the position of an object depicted in each image to be shifted, generates integrated data based on each partial area of ​​the first image and each reference area of ​​the second image, and detects the detection target from the first image based on the first image and the integrated data.

10. An information processing system comprising an imaging device and an information processing device, wherein the imaging device captures a first image and a second image of a detection target in real space at different wavelengths, and the information processing device acquires the first image and the second image, sets a plurality of partial areas by dividing the first image into a plurality of regions, sets a plurality of reference areas in the second image corresponding to each of the plurality of partial areas of the first image as areas to which each partial area is shifted based on factors that cause the position of an object depicted in each image to be shifted, generates integrated data based on each partial area of ​​the first image and each reference area of ​​the second image, and detects the detection target from the first image based on the first image and the integrated data.

Citation Information

Patent Citations

  • Image processing apparatus, distance measuring apparatus, imaging apparatus, and image processing method

    JP2015036841A

  • Vehicle-side device, server, method, and storage medium

    JP2020038362A

  • Image processing apparatus, image processing method, and program

    JP2024016619A