Prediction device, prediction method, prediction program, and storage medium

By extracting a partial rectangular area from vehicle images, focusing on the center and adjusting with speed, the device efficiently predicts driver gaze, reducing processing load without compromising accuracy.

JP2025156586APending Publication Date: 2025-10-14PIONEER IP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025134038
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Conventional techniques for calculating visual saliency impose a large processing load that increases with the size of the image, leading to inefficiencies in gaze prediction.

Method used

The prediction device extracts a rectangular area from an image, with a height smaller than the image, focusing on the center and adjusting based on vehicle speed, and predicts the driver's gaze using a deep learning model on this reduced area.

Benefits of technology

Reduces processing load while maintaining gaze prediction accuracy by calculating visual saliency on a partial image area, particularly beneficial for in-vehicle devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025156586000001_ABST
    Figure 2025156586000001_ABST
Patent Text Reader

Abstract

To reduce a processing load on calculation of visual saliency without degrading the accuracy in line-of-sight prediction.SOLUTION: An extraction unit 151 of a prediction device 10 extracts a partial area from an image captured in a direction of line-of-sight of a driver of a moving object. A prediction unit 152 predicts a position of driver's line-of-sight in the area extracted by the extraction unit 151, from the image. For example, the prediction unit 152 predicts the position of driver's line-of-sight by calculating visual saliency on the basis of the area extracted by the extraction unit 151.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a prediction device, a prediction method, a prediction program, and a storage medium. [Background technology]

[0002] Conventionally, visual saliency is known that can be obtained by estimating the driver's gaze position from an image of the area ahead of the vehicle. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-009825 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the conventional techniques have a problem in that they impose a large processing load, and the amount of processing required to calculate visual saliency increases with the size of the image that is the source of the calculation.

[0005] The present invention has been made in consideration of the above, and aims to provide a prediction device, a prediction method, a prediction program, and a storage medium that can reduce the processing load required for calculating visual saliency without deteriorating the accuracy of gaze prediction. [Means for solving the problem]

[0006] The prediction device described in claim 1 has an extraction unit that extracts a rectangular area from an image captured in the direction of the line of sight of the driver of a moving body, the rectangular area having a height smaller than the height of the image, and the height becomes smaller as the speed of the moving body increases, and a prediction unit that predicts the position of the driver's line of sight from the image in the area extracted by the extraction unit, wherein the extraction unit extracts an area including the center of the image.

[0007] The prediction method described in claim 4 is a prediction method executed by a computer, and includes an extraction step of extracting a rectangular area from an image captured in the direction of the line of sight of the driver of a moving body, the rectangular area having a height smaller than the height of the image, the height decreasing as the speed of the moving body increases, and a prediction step of predicting the position of the driver's line of sight in the area extracted by the extraction step from the image, wherein the extraction step extracts an area including the center of the image.

[0008] The prediction program described in claim 5 causes a computer to execute an extraction step of extracting, from an image captured in the direction of the line of sight of the driver of a moving body, a rectangular area whose height is smaller than the height of the image, and whose height becomes smaller as the speed of the moving body increases, and a prediction step of predicting, from the image, the position of the driver's line of sight in the area extracted by the extraction step, wherein the extraction step extracts an area including the center of the image.

[0009] The storage medium described in claim 6 causes a computer to execute an extraction step of extracting, from an image captured in the direction of the line of sight of the driver of a moving body, a rectangular area whose height is smaller than the height of the image, and whose height becomes smaller as the speed of the moving body increases, and a prediction step of predicting, from the image, the position of the driver's line of sight in the area extracted by the extraction step, wherein the extraction step stores a prediction program that extracts an area including the center of the image. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration of a prediction device according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating visual saliency. [Figure 3] FIG. 3 is a diagram illustrating a method for extracting an area. [Figure 4] FIG. 4 is a diagram for explaining a method for extracting an area. [Figure 5] FIG. 5 is a flowchart showing the flow of processing performed by the prediction device according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, a mode for carrying out the present invention (hereinafter referred to as an embodiment) will be described with reference to the drawings. Note that the present invention is not limited to the embodiment described below. Furthermore, in the description of the drawings, the same parts are given the same reference numerals.

[0012] [First embodiment] The prediction device according to the first embodiment predicts (estimates) the position of the driver's gaze in an image captured from a vehicle. The prediction device predicts the position of the gaze by calculating visual saliency.

[0013] The prediction device may be an in-vehicle device such as a drive recorder or a car navigation system, or may be an information processing device such as a personal computer or a server device.

[0014] 1 is a diagram illustrating an example of the configuration of a prediction device according to the first embodiment. As illustrated in FIG. 1, the prediction device 10 includes a communication unit 11, an imaging unit 12, a positioning unit 13, a storage unit 14, and a control unit 15.

[0015] The communication unit 11 is a communication module that is capable of data communication with other devices via a communication network such as the Internet.

[0016] The imaging unit 12 is, for example, a camera, and may be a camera of a drive recorder.

[0017] The positioning unit 13 receives a predetermined signal and measures the position of the vehicle 10 V. The positioning unit 13 receives signals from a global navigation satellite system (GNSS) or a global positioning system (GPS).

[0018] The storage unit 14 stores various programs executed by the prediction device 10, data necessary for executing processes, and the like.

[0019] The storage unit 14 stores model information 141. The model information 141 is parameters such as weights for constructing a deep learning model that performs computation of visual saliency.

[0020] The control unit 15 is realized by a controller such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs stored in the storage unit 14, and controls the overall operation of the prediction device 10. Note that the control unit 15 is not limited to being realized by a CPU or an MPU, but may also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0021] The control unit 15 has an extraction unit 151 and a prediction unit 152. The extraction unit 151 and the prediction unit 152 perform processing related to visual saliency.

[0022] Visual saliency will be described with reference to Fig. 2. Fig. 2 is a diagram for explaining visual saliency. As shown in Fig. 2, visual saliency is an index obtained by estimating the position of the driver's line of sight for an image showing the area ahead of the vehicle (see, for example, Patent Document 1).

[0023] Visual saliency may be calculated by inputting images into a deep learning model, for example, trained on a large number of images from a wide range of fields and the gaze information of multiple subjects who have viewed them.

[0024] Visual saliency is, for example, an 8-bit (0 to 255) value assigned to each pixel of an image, and is expressed as a larger value the greater the probability that the pixel is in the driver's line of sight. Therefore, if this value is considered to be a luminance value, the visual saliency can be superimposed on the original image as a heat map, as shown in Figure 2. In the following description, the visual saliency value of each pixel may be referred to as a luminance value.

[0025] The extraction unit 151 extracts a partial area from an image captured in the direction of the line of sight of the driver of the moving object. In the first embodiment, the calculation is performed on not the entire area of ​​the image captured by the imaging unit 12 but only on a partial area.

[0026] 3 is an image captured by the imaging unit 12. At this time, the extraction unit 151 extracts a rectangular region 211 whose width is equal to the width of the image 20 and whose height is smaller than the height of the image 20. FIG. 3 is a diagram illustrating a method for extracting a region.

[0027] Here, the computation of visual saliency can be considered a method for predicting which area in the field of view a driver is particularly paying attention to. It is thought that a driver driving on a road often moves their gaze horizontally (left and right) rather than vertically (up and down) in order to confirm the presence of surrounding traffic participants such as vehicles and pedestrians.

[0028] Therefore, it is thought that excluding the areas near the top and bottom edges of the image from the calculations will not have a significant effect on the prediction results, as shown in Figure 3. In fact, when looking at image 20 in Figure 3, there are no objects that require particular attention near the top and bottom edges, and many moving vehicles and other objects are visible in the horizontally long area near the center.

[0029] Furthermore, in order to focus particularly on the central area, the extraction unit 151 extracts a rectangular area whose width is equal to the width of the image 20 and whose height is smaller than the height of the image 20, and which includes the center of the image 20. In the example of FIG. 3, the area 211 overlaps with the center line 212 of the image 20.

[0030] Furthermore, the extraction unit 151 extracts a region that is a part of the image and whose position and size change depending on the state of the moving object.

[0031] For example, it is considered that the area to which the driver pays attention becomes smaller as the vehicle speed increases. Therefore, as shown in Fig. 4, extraction unit 151 extracts area 221 or area 222, which is a partial area of ​​image 20 and becomes smaller as the speed of the moving object increases. Fig. 4 is a diagram for explaining a method of extracting an area.

[0032] As shown in Figure 4, when the vehicle speed is a first speed, the extraction unit 151 extracts area 221, and when the vehicle speed is a second speed that is faster than the first speed, the extraction unit 151 extracts area 222, which has a smaller area than area 221.

[0033] The prediction unit 152 predicts, from the image, the position of the driver's gaze in the area extracted by the extraction unit 151. For example, the prediction unit 152 predicts the position of the driver's gaze by calculating visual saliency based on the area extracted by the extraction unit 151.

[0034] Specifically, the prediction unit 152 regards the region extracted by the extraction unit 151 as an image, inputs the image to a deep learning model constructed from the model information 141, and performs calculations on visual saliency.

[0035] 5 is a flowchart showing the flow of processing by the prediction device according to the first embodiment. As shown in FIG. 5, first, the prediction device 10 captures an image (step S101).

[0036] Next, the prediction device 10 extracts a partial region of the image (step S102). Then, the prediction device 10 predicts the gaze position in the extracted region (step S103). For example, the prediction device 10 obtains visual saliency by inputting the image of the extracted region into a trained deep learning model.

[0037] [Advantages of the first embodiment] As described above, the extraction unit 151 of the prediction device 10 extracts a partial area from an image capturing the gaze direction of a driver of a moving object. The prediction unit 152 predicts the position of the driver's gaze in the area extracted by the extraction unit 151 from the image. For example, the prediction unit 152 predicts the position of the driver's gaze by calculating visual saliency based on the area extracted by the extraction unit 151.

[0038] In this way, in the first embodiment, by calculating visual saliency not for the entire image but for a portion of the image, the processing load can be reduced without deteriorating the gaze prediction accuracy. The first embodiment is particularly useful when performing calculations in an in-vehicle device with limited computer resources.

[0039] The extraction unit 151 extracts a rectangular region whose width is equal to the width of the image and whose height is smaller than the height of the image. In particular, the extraction unit 151 extracts a rectangular region whose width is equal to the width of the image and whose height is smaller than the height of the image, and which includes the center of the image.

[0040] In this way, by extracting the area where the driver's line of sight is likely to be directed, useful calculation results can be obtained with a small amount of calculation.

[0041] The extraction unit 151 extracts a region that is a part of the image and whose position and size change depending on the state of the moving object. In particular, the extraction unit 151 extracts a region that is a part of the image and whose size becomes smaller as the speed of the moving object increases.

[0042] In this way, by changing the region extraction method depending on the situation, it is possible to efficiently perform calculations of visual saliency in accordance with the situation. [Explanation of symbols]

[0043] 10 Prediction Device 11 Communications Department 12 Imaging unit 13 Positioning unit 14 Storage section 15 Control Unit 141 Model Information 151 Extraction part 152 Prediction Department

Claims

1. an extraction unit that extracts, from an image captured in the direction of the line of sight of a driver of a moving object, a rectangular area whose height is smaller than the height of the image, and whose height decreases as the speed of the moving object increases; a prediction unit that predicts, from the image, a position of the driver's line of sight in the area extracted by the extraction unit; and The prediction device is characterized in that the extraction unit extracts a region including a center of the image.

2. The prediction device according to claim 1 , wherein the prediction unit predicts the driver's gaze position by calculating visual saliency based on the area extracted by the extraction unit.

3. 3. The prediction device according to claim 1, wherein the extraction unit extracts the rectangular region having a width equal to a width of the image.

4. 1. A computer-implemented prediction method comprising: an extraction step of extracting, from an image captured in the direction of the line of sight of a driver of a moving object, a rectangular area whose height is smaller than the height of the image, the height decreasing as the speed of the moving object increases; a prediction step of predicting, from the image, a position of the driver's line of sight in the area extracted by the extraction step; and A prediction method characterized in that the extraction step extracts a region including the center of the image.

5. an extraction step of extracting, from an image captured in the direction of the line of sight of a driver of a moving object, a rectangular area whose height is smaller than the height of the image, the height decreasing as the speed of the moving object increases; a prediction step of predicting, from the image, a position of the driver's line of sight in the area extracted by the extraction step; on the computer, The extraction step is a prediction program that extracts a region including the center of the image.

6. an extraction step of extracting, from an image captured in the direction of the line of sight of a driver of a moving object, a rectangular area whose height is smaller than the height of the image, the height decreasing as the speed of the moving object increases; a prediction step of predicting, from the image, a position of the driver's line of sight in the area extracted by the extraction step; on the computer, A storage medium storing a prediction program for extracting an area including a center of the image in the extraction step.

Citation Information

Patent Citations

  • Visual confirmation load amount estimation device, drive support device and visual confirmation load amount estimation program

    JP2013009825A