Electronic devices, information processing methods, and programs
The method addresses computational inefficiencies in sensor-based object recognition by using spatial concatenation and Fourier transforms to achieve accurate and efficient object recognition across multiple sensor inputs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KYOCERA CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
Smart Images

Figure 2026064151000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an electronic device, an information processing method, and a program.
Background Art
[0002] Object recognition for recognizing the characteristics of an object included as an image in an image based on the image is used for autonomous driving and the like. In recent years, performing object recognition using a learned estimation model has been studied.
[0003] An image obtained for object recognition may reduce the recognition accuracy depending on external circumstances such as at night or in the presence of fog. Therefore, using a plurality of different types of sensors has been studied. The plurality of sensors are sensors that can generate position-specific information within a measurement range, such as a visible light camera, an infrared camera, a distance measuring device such as LiDAR, and a millimeter wave radar.
[0004] As an estimation model for performing object recognition using the output results of a plurality of sensors, a sensor fusion model using Attention has been proposed (see Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] The estimation model described in Non-Patent Document 1 is computationally expensive. Therefore, recognition of high-resolution images takes a relatively long time. As a result, it has been difficult to apply it to applications that require instantaneous recognition, such as autonomous driving.
[0007] Therefore, the purpose of this disclosure is to provide an electronic device, an information processing method, and a program that perform object recognition with high estimation accuracy and low computational cost based on the output results of multiple sensors. [Means for solving the problem]
[0008] Electronic devices from the first perspective are, A first concatenated map is generated by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band. By performing a Fourier transform on the aforementioned first concatenated map, a frequency domain map is generated. The frequency domain map after the element-wise product of the aforementioned frequency domain map and the trained filter is calculated. A second concatenated map is generated by performing an inverse Fourier transform on the frequency domain map after the element product. The system includes a control unit for dividing the aforementioned second link map.
[0009] In the second perspective on information processing methods, Computer A first concatenated map is generated by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band. By performing a Fourier transform on the aforementioned first concatenated map, a frequency domain map is generated. The frequency domain map after the element-wise product of the aforementioned frequency domain map and the trained filter is calculated. A second concatenated map is generated by performing an inverse Fourier transform on the frequency domain map after the element product. The aforementioned second linked map is split.
[0010] The program from a third perspective is: to the computer A first concatenated map is generated by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band. By performing a Fourier transform on the first concatenated map described above, a frequency domain map is generated. The frequency domain map is calculated after the element-wise product of the aforementioned frequency domain map and the trained filter. A second connected map is generated by performing an inverse Fourier transform on the frequency domain map after the element product. The second linked map described above is split. [Effects of the Invention]
[0011] According to the electronic device, information processing method, and program described above, object recognition can be performed with high estimation accuracy and low computational cost on the output results of multiple sensors. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 is a block diagram showing the schematic configuration of the electronic device according to this embodiment. [Figure 2] This diagram shows the configuration of the Extended Global Filter Layer in this disclosure. [Figure 3] This is a diagram showing the model configuration of the trained model applied to this embodiment. [Figure 4] This is a flowchart illustrating the object recognition process performed by the control unit shown in Figure 1. [Figure 5] This table shows the evaluation results of the examples and comparative examples of this disclosure. [Modes for carrying out the invention]
[0013] Hereinafter, embodiments of an electronic device to which the present disclosure is applied will be described with reference to the drawings.
[0014] FIG. 1 shows a configuration example of an electronic device 10. The electronic device 10 includes a control unit 11. The electronic device 10 may further include an acquisition unit 12. Each functional unit constituting the electronic device 10 may be connected to each other by a network consisting of wired, wireless, or a combination thereof. The electronic device 10 may be constituted by the connected functional units.
[0015] The acquisition unit 12 may acquire an image as information. The acquisition unit 12 may acquire a group of images substantially simultaneously detected by a plurality of sensors whose at least a part of each detection range overlaps. The acquisition unit 12 may acquire an image from a sensor, an electronic device, a storage medium, or the like. Alternatively, the acquisition unit 12 may acquire an image from an electronic device, a storage medium, or the like that has acquired an image detected by a sensor.
[0016] The plurality of sensors may output a detection result including information for each minute range constituting an arbitrary spatial range, in other words, image-like information. The plurality of sensors may include at least a first sensor and a second sensor. The first sensor may perform imaging in a first frequency band and generate a first image. The second sensor may generate a second image captured in a second frequency band different from the first frequency band. Imaging includes detecting a detectable result that can be imaged.
[0017] The first frequency band may be, for example, a visible light band. Therefore, the first image may be a visible light image such as an RGB image. The second frequency band may be, for example, a near-infrared band. The second image may be a distance image that is an output result of a distance measurement sensor using the near-infrared band. The sensor may be, for example, a visible light camera such as RGB, a NIR camera, a FIR camera, LiDar, a millimeter wave radar, or an infrared radar. A part of the first frequency band and the second frequency band may overlap.
[0018] "Substantially simultaneous detection" does not mean strictly simultaneous detection, but rather detection within a range that can be considered simultaneous in terms of sensor detection. Specifically, substantially simultaneous detection may mean detection within the shortest sampling period among multiple sensors.
[0019] The acquisition unit 12 may include, for example, a physical connector and a wireless communication device. The physical connector includes electrical connectors that support transmission by electrical signals, optical connectors that support transmission by optical signals, and electromagnetic connectors that support transmission by electromagnetic waves. The electrical connector includes connectors that comply with IEC60603, connectors that comply with the USB standard, connectors that support RCA terminals, connectors that support S terminals as defined in EIAJ CP-1211A, connectors that support D terminals as defined in EIAJ RC-5237, connectors that comply with the HDMI® standard, and connectors that support coaxial cables including BNC. The optical connector includes various connectors that comply with IEC 61754. The wireless communication device includes wireless communication devices that comply with various standards, including Bluetooth® and IEEE802.11.
[0020] The control unit 11 is configured to include at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for a specific process. The dedicated circuit may be, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 11 may control the operation of the electronic device 10.
[0021] The control unit 11 may further include a storage unit 11a. The storage unit 11a may include any storage device, such as RAM (Random Access Memory) and ROM (Read Only Memory). The storage unit 11a may store data or programs for controlling the entire or a part of the electronic device 10, as well as various programs for making the control unit 11 function, and various information used by the control unit 11.
[0022] The control unit 11 may function as a trained model. The trained model is, for example, an object recognition model. The object recognition model may estimate information about subjects in a scene within a detection range where at least a portion of the first and second images overlap each other, based on the first and second images. Information about subjects may include, for example, the type of subject (person or object), the position of the subject, the number of subjects, the size of the subjects, the movement of the subjects, etc.
[0023] When the trained model is put into operation, the control unit 11 generates a first connected map, generates a frequency domain map, calculates a frequency domain map after element-wise product, generates a second connected map, and splits the second connected map, as described below. The control unit 11 may generate the first connected map, generate a Fourier transform map, calculate a frequency domain map after element-wise product, generate the second connected map, and split the second connected map during the training of the trained model.
[0024] The trained model includes an extended GFL (Global Filter Layer). The extended GFL 13 may include the following steps, as shown in Figure 2: a concatenation layer 14, a Fourier transform layer 15, an elemental product layer 16, an inverse Fourier transform layer 17, and a splitting layer 18. Note that the Fourier transform layer 15, the elemental product layer 16, and the inverse Fourier transform layer 17 may be conventional GFLs.
[0025] The concatenation layer 14 connects the first feature map 19 and the second feature map 20 in space. Connecting the first feature map 19 and the second feature map 20 in space may mean joining the second feature map 20 at any end in the arrangement direction of each element of the first feature map 19. For example, the concatenation layer 14 may connect the first feature map 19, which has W elements arranged in the row direction and H elements arranged in the column direction, with the second feature map 20, which has W elements arranged in the row direction and H elements arranged in the column direction, at their respective ends in the row direction. The concatenation may generate a first concatenated map 21 with 2W elements arranged in the row direction and H elements arranged in the column direction.
[0026] The first feature map 19 is based on the first image. The first feature map 19 may be an image-like feature map obtained by feature extraction from the first image. The first feature map 19 may be generated by convolution using a trained filter. Alternatively, the first feature map 19 may be the first image itself. The second feature map 20 is based on the second image. The second feature map 30 may be an image-like feature map obtained by feature extraction from the second image. The second feature map 20 may be generated by convolution using a trained filter. Alternatively, the second feature map 20 may be the second image itself. The trained filters used to generate the first feature map 19 and the second feature map 20 may be different.
[0027] The first feature map 19 and the second feature map 20 may be the same size. If the first image and the second image are of different sizes, the first feature map 19 and the second feature map 20 may be resized to be the same size before being input to the extended GFL 13.
[0028] The Fourier transform layer 15 applies a Fourier transform to the first concatenated map 21 to generate a frequency domain map. The element-wise product layer 16 calculates a frequency domain map after the element-wise product of the frequency domain map and the trained filter. The inverse Fourier transform layer 17 applies an inverse Fourier transform to the frequency domain map after the element-wise product to generate a second concatenated map 22.
[0029] The partitioning layer 18 may partition the second connected map 22. The partitioned map may be used in the trained model as a feature map showing long-range dependencies in the first and second images.
[0030] The trained model may include CNNs, Transformers, etc. The extended GFL13 described above may constitute a part of the trained model. For example, the trained model 23 may include a Residual Network as part of it, as shown in Figure 3. The trained model 23 may include a single extended GFL13 or multiple extended GFL13s. A Residual Network may be used as a type of CNN.
[0031] The trained model 23 in the example may have first feature extraction layers 24 from the first to the fourth layer for the first image and the second image, respectively. The first feature extraction layers 24 may perform feature extraction on a single type of image or feature map. Some of the first feature extraction layers 24 may output feature maps to the extended GFL 13 and the adder 25. The feature maps output to the extended GFL 13 may be used as the first feature map 19 or the second feature map 20 described above. The adder 25 may add the feature maps output from the first feature extraction layers 24 and the feature maps output from the extended GFL 13.
[0032] The feature maps output from the adder 25 may be output to the lower-level first feature extraction layer 24 and second feature extraction layer 26. The second feature extraction layer 26 may receive feature maps from the channels of the first and second images. The second feature extraction layer 26 may generate a feature map by performing convolution using a trained filter on two different types of feature maps. The second feature extraction layer 26 may output the generated feature map to the detector 27. The detector 27 may output an estimation result based on the feature maps input from each of the multiple second feature extraction layers 26.
[0033] Next, the estimation process performed by the control unit 11 in this embodiment will be explained using the flowchart in Figure 4. The estimation process starts when the acquisition unit 12 acquires a pair of the first image and the second image.
[0034] In step S100, the control unit 11 performs preprocessing such as resizing on the acquired first image and second image. After the preprocessing is completed, the process proceeds to step S101.
[0035] In step S101, the control unit 11 performs feature extraction on the first image and the second image, which were preprocessed in step S100, and generates feature maps. After generation, the process proceeds to step S102.
[0036] In step S102, the control unit 11 connects the two feature maps generated in step S101 in space to generate a first concatenated map 21. After generation, the process proceeds to step S103.
[0037] In step S103, the control unit 11 applies a Fourier transform to the first concatenated map 21 generated in step S102 to generate a frequency domain map. After generation, the process proceeds to step S104.
[0038] In step S104, the control unit 11 calculates the frequency domain map after the element-wise product of the frequency domain map generated in step S103 and the trained filter. After calculation, the process proceeds to step S105.
[0039] In step S105, the control unit 11 applies an inverse Fourier transform to the frequency domain map after element product calculated in step S104 to generate a second concatenated map 22. After generation, the process proceeds to step S106.
[0040] In step S106, the control unit 11 divides the second linked map 22 generated in step S105. After the division, the process proceeds to step S107.
[0041] In step S107, the control unit 11 uses the second concatenated map 22, which was divided in step S106, to estimate information about the subject within the detection range of both the first and second images. After estimation, the estimation process ends.
[0042] The electronic device 10 of this embodiment, configured as described above, includes a control unit 11 that generates a first concatenated map 21 by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band, generates a frequency domain map by performing a Fourier transform on the first concatenated map 21, calculates a frequency domain map after element-wise product of the frequency domain map and a trained filter, generates a second concatenated map 22 by performing an inverse Fourier transform on the frequency domain map after element-wise product, and divides the second concatenated map 22. The GFL (Global Filter Layer) that performs the above-described generation of the frequency domain map by Fourier transform, calculation of the frequency domain map after element-wise product, and inverse Fourier transform is generally capable of extracting features of long-range dependencies, similar to Attention. Furthermore, GFL generally has a lower computational cost than Attention. With the configuration of this disclosure as described above, utilizing the characteristics of such GFL, the electronic device 10 can use the first image and the second image to extract features while maintaining the features of long-range dependencies between the first feature map and the second feature map. Therefore, the electronic device 10 enables estimation using images with different frequency bands in order to maintain the advantages of GFL. As a result, the electronic device 10 can perform object recognition with high estimation accuracy and low computational cost on the output results of multiple sensors. [Examples]
[0043] The present disclosure will be further described below by examples and comparative examples, but will not be limited thereto.
[0044] As a model for the embodiment of this disclosure, a pre-trained model as shown in Figure 3 was constructed. Specifically, an RGB image was used for the first image and a Depth Map was used for the second image. In the pre-trained model, ResNet18 was applied to both the first and second images. In the pre-trained model of the embodiment, as described above, the outputs of Layers 2, 3, and 4 (first feature extraction layers) of the respective ResNet18 for the first and second images were input to the extended GFL13.
[0045] In Comparative Example 1 of this disclosure, Attention was applied to the model instead of the Extended GFL used in the Example model. In Comparative Example 2, Spatial Reduction Attention was applied to the model instead of the Extended GFL used in the Example model. In Spatial Reduction Attention, Key and Value downsampling was performed before applying Attention. Furthermore, when applying Attention and Spatial Reduction Attention, feature maps extracted from RGB and Depth were concatenated. By performing such concatenation, the features of RGB and Depth are learned in a manner similar to the Example.
[0046] Examples, Comparative Example 1, and Comparative Example 2 were trained using 100,000 arbitrarily selected pairs of RGB and Depth Maps from the SHIFT dataset. During training, the resolution was resized from the original 1280x800 to 320x320.
[0047] AP, AP50, and AP75 were measured to evaluate the accuracy of Example, Comparative Example 1, and Comparative Example 2, respectively. Furthermore, to evaluate the processing speed of Example, Comparative Example 1, and Comparative Example 2, the frame rates of object recognition using first and second images of sizes 320×320, 640×640, 960×960, 1280×1280, and 1600×1600 were measured. An RTX (NVIDIA Corporation, registered trademark) 4090 GPU was used for the experiment. FP32 was used for bit precision settings. The evaluation results are shown in Figure 5.
[0048] While embodiments of the electronic device 10 have been described above, embodiments of the present disclosure may also include methods or programs for implementing the device, as well as embodiments of a storage medium on which a program is recorded (for example, an optical disc, magneto-optical disc, CD-ROM, CD-R, CD-RW, magnetic tape, hard disk, or memory card).
[0049] Furthermore, the implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, but may also be in the form of a program module embedded in an operating system. In addition, the program may or may not be configured so that all processing is performed only on the CPU on the control board. The program may also be configured so that some or all of its processing is performed by another processing unit implemented on an expansion board or expansion unit attached to the board, as needed.
[0050] The diagrams illustrating the embodiments described herein are schematic. Dimensions and proportions shown in the drawings do not necessarily correspond to actual dimensions.
[0051] While embodiments relating to this disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art can make various modifications or alterations based on this disclosure. Therefore, it should be noted that these modifications or alterations are within the scope of this disclosure. For example, the functions included in each component can be rearranged in a logically consistent manner, and multiple components can be combined into one or separated. This disclosure can be applied not only to object detection but also to technologies such as image classification, semantic segmentation, instance segmentation, person detection, and super-resolution to increase the resolution of low-resolution images.
[0052] For example, in the estimation process shown in Figure 4, steps S102 to S106 may be repeated any number of times. Specifically, after step S106, it may be determined whether the processes in steps S102 to S106 have been executed any number of times. If the determination indicates that the processes have been executed, the process may proceed to step S107; otherwise, the process may proceed to step S102. Furthermore, in the estimation process, step S101 may also be repeated any number of times along with steps S102 to S106.
[0053] All of the constituent elements described in this disclosure, and / or all of the disclosed methods or steps of processing, can be combined in any combination except for any combination in which these features are mutually exclusive. Furthermore, each of the features described in this disclosure can be replaced by an alternative feature that works for the same, equivalent, or similar purposes, unless expressly disregarded. Thus, unless expressly disregarded, each of the disclosed features is merely an example of a comprehensive set of identical or equivalent features.
[0054] Furthermore, the embodiments relating to this disclosure are not limited to any specific configuration of the embodiments described above. The embodiments relating to this disclosure can be extended to all novel features or combinations thereof described herein, or all novel methods or processing steps or combinations thereof described herein.
[0055] In this disclosure, the designations "First," "Second," etc., are identifiers used to distinguish the configurations. Configurations distinguished by the designations "First," "Second," etc., in this disclosure may have their numbers swapped. For example, Feature Map 1 may swap its identifiers "First" and "Second" with Feature Map 2. The swapping of identifiers occurs simultaneously. The configurations remain distinguishable even after the swapping of identifiers. Identifiers may be deleted. Configurations from which identifiers have been deleted are distinguished by codes. The designations "First," "Second," etc., in this disclosure should not be used alone to interpret the order of the configurations or to justify the existence of smaller numbered identifiers. [Explanation of symbols]
[0056] 10 Electronic equipment 11 Control Unit 11a Storage section 12 Acquisition Department 13 Extended GFL 14 Connectivity layer 15. Fourier transform layer 16-element integral layer 17 Inverse Fourier Transform Layer 18 split layer 19. Feature Map (First Feature) 20. Second Feature Map 21. First Link Map 22 Second Link Map 23 Pre-trained models 24. First Feature Extraction Layer 25 Adder 26. Second Feature Extraction Layer
Claims
1. A first concatenated map is generated by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band. By performing a Fourier transform on the first concatenated map described above, a frequency domain map is generated. The frequency domain map after the element-wise product of the aforementioned frequency domain map and the trained filter is calculated. A second connected map is generated by performing an inverse Fourier transform on the frequency domain map after the element product. The system includes a control unit that divides the second linked map. electronic equipment.
2. In the electronic device described in claim 1, The control unit performs at least one of the following: generating a first feature map by extracting features from the first image, and generating a second feature map by extracting features from the second image. electronic equipment.
3. In the electronic device according to claim 1 or 2, The generation of the first concatenated map, the generation of the frequency domain map, the calculation of the frequency domain map after the element product, the generation of the second concatenated map, and the division of the second concatenated map are performed in a model that estimates information about subjects in a scene from two images captured of a scene in which at least a portion of the detection range overlaps with each other in the first frequency band and the second frequency band. electronic equipment.
4. Computer A first concatenated map is generated by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band. By performing a Fourier transform on the first concatenated map described above, a frequency domain map is generated. The frequency domain map after the element-wise product of the aforementioned frequency domain map and the trained filter is calculated. A second connected map is generated by performing an inverse Fourier transform on the frequency domain map after the element product. The second linked map described above is divided. Information processing methods.
5. to the computer A first concatenated map is generated by spatially concatenating a first feature map based on a first image captured in a first frequency band and a second feature map based on a second image captured in a second frequency band. By performing a Fourier transform on the first concatenated map, a frequency domain map is generated. The frequency domain map is calculated after the element-wise product of the aforementioned frequency domain map and the trained filter. A second connected map is generated by performing an inverse Fourier transform on the frequency domain map after the element product. The second linked map described above is divided. program.