Determination device
The determination device uses visual splendor distribution analysis to enhance monotonicity detection by considering human gaze patterns, improving accuracy in identifying monotonic landscapes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PIONEER IP
- Filing Date
- 2026-01-29
- Publication Date
- 2026-04-10
AI Technical Summary
Existing systems fail to accurately detect monotonic trends in landscapes due to reliance on landscape changes or road information, neglecting visual cues, leading to incorrect assessments of monotonicity.
A determination device that estimates visual splendor distribution from captured images using statistical quantities, such as standard deviation and gaze movement, to determine monotonicity based on human gaze patterns.
Accurately determines monotonicity by focusing on human gaze patterns, enhancing detection accuracy and addressing limitations of previous methods.
Smart Images

Figure 2026063419000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a determination device that determines whether an image has a monotonic tendency based on one or more images taken of the outside from a moving object. [Background technology]
[0002] Generally, the phenomenon of feeling drowsy or inducing a hypnotic state when driving for long periods on roads with monotonous scenery, such as highways, is known as highway hypnosis. Additionally, drowsiness can occur on roads with little or no change in scenery, or on roads with regularly spaced streetlights.
[0003] As an invention for detecting such monotonous roads, for example, Patent Document 1 describes a landscape monotonicity calculation device comprising an image acquisition means for acquiring an appearance image and a monotonicity calculation means for calculating the monotonicity of the landscape corresponding to the appearance image based on the appearance image acquired by the image acquisition means.
[0004] Furthermore, Patent Document 2 describes an alertness state determination system that can determine the driver's alertness state with high accuracy. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Patent No. 4550116 [Patent Document 2] Japanese Patent Publication No. 2010-128649 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] In the invention described in Patent Document 1, a feature judgment process is performed to classify the landscape, and monotonicity is determined based on the count of changes in the landscape. Therefore, even if the landscape is complex, such as a "streetscape," if the changes in time or position are small, the monotonicity will be judged as low. In the invention described in Patent Document 2, the level of alertness is determined from head movements in sections determined to be monotonic using road information, vehicle information, and inter-vehicle distance information. Since the detection of monotonic sections is based on road information, visual information is not taken into consideration. Therefore, even if there are many signs and markers, if it is a straight road, it may be judged as monotonic.
[0007] One example of a problem that this invention aims to solve is detecting a monotonic trend from captured images. [Means for solving the problem]
[0008] To solve the above problems, the invention described in claim 1 is characterized by comprising: an acquisition unit that acquires visual splendor distribution information obtained by estimating the level of visual splendor within an image based on an image captured from a moving object; and a determination unit that determines whether the image is monotonic using a statistical quantity calculated based on the visual splendor distribution information.
[0009] The invention described in claim 6 is a determination method performed by a determination device that determines whether an image captured from a moving object is monotonic, and is characterized by comprising: an acquisition step of obtaining visual splendor distribution information obtained by inferring the level of visual splendor within the image based on the image; and a determination step of determining whether the image is monotonic using a statistical quantity calculated based on the visual splendor distribution information.
[0010] The invention described in claim 7 is characterized in that the determination method described in claim 6 is executed by a computer.
[0011] The invention described in claim 8 is characterized by storing the determination program described in claim 7.
Brief Description of the Drawings
[0012] [Figure 1] It is a functional configuration diagram of a determination device according to a first embodiment of the present invention. [Figure 2] It is a block diagram illustrating the configuration of the landscape image processing unit shown in FIG. 1. [Figure 3] (a) is a diagram illustrating an image input to the determination device, and (b) is a diagram illustrating a visually salient map estimated for (a). [Figure 4] It is a flowchart illustrating the processing method of the landscape image processing unit shown in FIG. 1. [Figure 5] It is a diagram illustrating in detail the configuration of the non-linear mapping unit. [Figure 6] It is a diagram illustrating the configuration of the intermediate layer. [Figure 7] (a) and (b) are diagrams each showing an example of a convolution process performed by a filter. [Figure 8] (a) is a diagram for explaining the processing of the first pooling unit, (b) is a diagram for explaining the processing of the second pooling unit, and (c) is a diagram for explaining the processing of the unpooling unit. [Figure 9] It is a flowchart of the operation of the monotonic determination unit shown in FIG. 1. [Figure 10] It is a flowchart of the operation of a determination device according to a second embodiment of the present invention. [Figure 11] It is an example of the calculation result of autocorrelation.
Modes for Carrying Out the Invention
[0013] Hereinafter, a determination device according to an embodiment of the present invention will be described. In the determination device according to an embodiment of the present invention, an acquisition unit acquires visual saliency distribution information obtained by estimating the level of visual saliency in an image captured by a moving body of the outside, and a determination unit determines whether the image has a monotonic tendency using a statistic calculated based on the visual saliency distribution information. By doing so, it becomes possible to determine whether there is a monotonic tendency based on the position where a human is likely to gaze from the captured image. Since it is determined based on the position where a human (driver) is likely to gaze, it is possible to determine with a tendency close to what the driver feels is monotonous, and it is possible to determine more accurately.
[0014] Further, a standard deviation calculation unit that calculates the standard deviation of the luminance of each pixel in the image obtained as the visual saliency distribution information may be provided, and the determination unit may determine whether the image has a monotonic tendency based on the calculated standard deviation. By doing so, in one image, it is possible to determine that there is a monotonic tendency when the positions where it is easy to gaze are concentrated.
[0015] Further, an average value calculation unit that calculates the average value of the luminance of each pixel in the image obtained as the visual saliency distribution information may be provided, and the determination unit may determine whether the image has a monotonic tendency based on the calculated average value. By doing so, in one image, it is possible to determine that there is a monotonic tendency when the positions where it is easy to gaze are concentrated. Also, since the determination is made using the average value, the arithmetic processing can be simplified.
[0016] Further, the acquisition unit acquires the visual saliency distribution information in time series, and the determination unit may include a first determination unit that calculates a statistic from the visual saliency distribution information acquired in time series and determines whether there is a monotonic tendency based on the statistic obtained in time series, and a second determination unit that determines whether there is a monotonic tendency based on autocorrelation. By doing so, it is possible to determine the monotonic tendency caused by periodically appearing objects such as street lights that appear during driving, which is difficult to determine only by the statistic, by autocorrelation.
[0017] Furthermore, the system may include a gaze movement calculation unit that calculates the amount of gaze movement between frames based on images obtained in a time series as visual splendor distribution information, and a determination unit that determines whether the image is monotonic based on the calculated gaze movement amount. In this way, for example, if the amount of gaze movement is small, it can be determined that the image is monotonic.
[0018] Furthermore, the acquisition unit includes an input unit that converts an image into intermediate data that can be mapped, a nonlinear mapping unit that converts the intermediate data into mapping data, and an output unit that generates splendor estimation information showing a splendor distribution based on the mapping data. The nonlinear mapping unit may also include a feature extraction unit that extracts features from the intermediate data and an upsampling unit that upsamples the data generated by the feature extraction unit. In this way, visual splendor can be estimated with low computational cost. Moreover, the visual splendor estimated in this way reflects the contextual attentional state.
[0019] Furthermore, in the determination method according to one embodiment of the present invention, in the acquisition step, visual splendor distribution information is obtained by estimating the level of visual splendor within an image based on an image captured from a moving object, and in the determination step, a statistic calculated based on the visual splendor distribution information is used to determine whether the image has a monotonic tendency. In this way, it becomes possible to determine whether the captured image has a monotonic tendency based on a position that is easy for a human to focus on. Since the determination is based on a position that is easy for a human (driver) to focus on, the determination can be made with a tendency that is close to what the driver perceives as monotonic, and the determination accuracy can be improved.
[0020] Furthermore, the aforementioned determination method is performed by a computer. By doing so, it becomes possible to determine whether there is a monotonic trend based on the position that humans tend to focus on, from the images captured using the computer.
[0021] Furthermore, the aforementioned judgment program may be stored in a computer-readable storage medium. This allows the program to be distributed independently, in addition to being incorporated into a device, and facilitates version upgrades and other modifications. [Examples]
[0022] A determination device according to the first embodiment of the present invention will be described with reference to Figures 1 to 9. The determination device according to this embodiment is not limited to being installed on a mobile device such as an automobile, but may also be composed of a server device installed in a business office or the like. In other words, it is not necessary to perform the analysis in real time, and the analysis may be performed after driving or at some other time.
[0023] As shown in Figure 1, the determination device 1 comprises a landscape image acquisition unit 2, a landscape image processing unit 3, and a monotone determination unit 4.
[0024] The landscape image acquisition unit 2 receives an image (e.g., a video) captured by a camera or the like as input and outputs that image as image data. The input video is output as image data that has been broken down into time series, such as frame by frame. Still images may be input to the landscape image acquisition unit 2, but it is preferable to input a group of images consisting of multiple still images arranged in time series.
[0025] The images input to the landscape image acquisition unit 2 include, for example, images capturing the direction of travel of a vehicle. In other words, images continuously captured from a moving object. These images may include images other than the direction of travel, such as panoramic images or images acquired using multiple cameras, such as 180° or 360° in the horizontal direction. Furthermore, the images input to the landscape image acquisition unit 2 are not limited to images captured by cameras; images read from recording media such as hard disk drives or memory cards may also be input.
[0026] The landscape image processing unit 3 receives image data from the landscape image acquisition unit 2 and outputs a visual splendor map as visual splendor estimation information, which will be described later. In other words, the landscape image processing unit 3 functions as an acquisition unit that acquires a visual splendor map (visual splendor distribution information) obtained by estimating the level of visual splendor within an image based on an image of the outside world captured by a moving object.
[0027] Figure 2 is a block diagram illustrating the configuration of the landscape image processing unit 3. The landscape image processing unit 3 according to this embodiment comprises an input unit 310, a nonlinear mapping unit 320, an output unit 330, and a storage unit 390. The input unit 310 converts an image into intermediate data that can be mapped. The nonlinear mapping unit 320 converts the intermediate data into mapping data. The output unit 330 generates spleniality estimation information showing the spleniality distribution based on the mapping data. The nonlinear mapping unit 320 comprises a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. The storage unit 390 holds image data input from the landscape image acquisition unit 2 and the coefficients of filters described later. A detailed explanation follows below.
[0028] Figure 3(a) is an example of an image input to the landscape image processing unit 3, and Figure 3(b) is an example of an image showing the estimated visual splendor distribution for Figure 3(a). The landscape image processing unit 3 according to this embodiment is a device that estimates the visual splendor of each part of an image. Visual splendor means, for example, how easily something stands out or how easily it attracts the viewer's gaze. Specifically, visual splendor is expressed as a probability, etc. Here, the magnitude of the probability corresponds, for example, to the probability that the viewer's gaze will be directed towards that location when looking at the image.
[0029] Figures 3(a) and 3(b) correspond to each other in terms of position. Furthermore, in Figure 3(a), locations with higher visual spleness are displayed with higher brightness in Figure 3(b). The image showing the visual spleness distribution, as in Figure 3(b), is an example of a visual spleness map output by the output unit 330. In this example, visual spleness is visualized using 256 levels of brightness values. An example of a visual spleness map output by the output unit 330 will be described in more detail later.
[0030] Figure 4 is a flowchart illustrating the operation of the landscape image processing unit 3 according to this embodiment. The flowchart shown in Figure 4 is part of a determination method executed by a computer and includes an input step S110, a nonlinear mapping step S120, and an output step S130. In the input step S110, the image is converted into intermediate data that can be mapped. In the nonlinear mapping step S120, the intermediate data is converted into mapping data. In the output step S130, visual splendor estimation information (visual splendor distribution information) showing a splendor distribution based on the mapping data is generated. Here, the nonlinear mapping step S120 includes a feature extraction step S121 that extracts features from the intermediate data and an upsampling step S122 that upsamples the data generated in the feature extraction step S121.
[0031] Returning to Figure 2, let's explain each component of the landscape image processing unit 3. In input step S110, the input unit 310 acquires an image and converts it into intermediate data. The input unit 310 acquires image data from the landscape image acquisition unit 2. The input unit 310 then converts the acquired image into intermediate data. The intermediate data is not particularly limited as long as it is data that the nonlinear mapping unit 320 can accept, but for example, it is a high-dimensional tensor. The intermediate data is, for example, data in which the brightness of the acquired image has been normalized, or data in which each pixel of the acquired image has been converted into a brightness slope. In input step S110, the input unit 310 may further perform image noise reduction, resolution conversion, etc.
[0032] In the nonlinear mapping step S120, the nonlinear mapping unit 320 acquires intermediate data from the input unit 310. Then, the intermediate data is converted into mapping data in the nonlinear mapping unit 320. Here, the mapping data is, for example, a high-dimensional tensor. The mapping process applied to the intermediate data by the nonlinear mapping unit 320 is a mapping process that can be controlled by parameters, for example, and is preferably a process using a function, a functional, or a neural network.
[0033] Figure 5 is a diagram illustrating the configuration of the nonlinear mapping unit 320 in detail, and Figure 6 is a diagram illustrating the configuration of the intermediate layer 323. As described above, the nonlinear mapping unit 320 includes a feature extraction unit 321 and an upsampling unit 322. The feature extraction step S121 is performed in the feature extraction unit 321, and the upsampling step S122 is performed in the upsampling unit 322. In the example shown in this figure, at least one of the feature extraction unit 321 and the upsampling unit 322 is configured to include a neural network containing multiple intermediate layers 323. In the neural network, multiple intermediate layers 323 are connected.
[0034] In particular, the neural network is preferably a convolutional neural network. Specifically, each of the multiple hidden layers 323 includes one or more convolutional layers 324. In the convolutional layers 324, the input data is convolved by multiple filters 325, and the outputs of the multiple filters 325 are subjected to activation processing.
[0035] In the example shown in Figure 5, the feature extraction unit 321 is configured to include a neural network containing multiple hidden layers 323, with a first pooling unit 326 between the multiple hidden layers 323. The upsampling unit 322 is also configured to include a neural network containing multiple hidden layers 323, with an unpooling unit 328 between the multiple hidden layers 323. Furthermore, the feature extraction unit 321 and the upsampling unit 322 are connected to each other via a second pooling unit 327 that performs overlap pooling.
[0036] In the example shown in this figure, each intermediate layer 323 consists of two or more convolutional layers 324. However, at least some of the intermediate layers 323 may consist of only one convolutional layer 324. Adjacent intermediate layers 323 are separated by one of the first pooling section 326, the second pooling section 327, and the unpooling section 328. Here, if an intermediate layer 323 contains two or more convolutional layers 324, it is preferable that the number of filters 325 in those convolutional layers 324 are equal to each other.
[0037] In this diagram, the intermediate layer 323 labeled "A×B" consists of B convolutional layers 324, and each convolutional layer 324 contains A convolutional filters for each channel. Such an intermediate layer 323 will also be referred to as an "A×B intermediate layer" below. For example, a 64×2 intermediate layer 323 consists of two convolutional layers 324, and each convolutional layer 324 contains 64 convolutional filters for each channel.
[0038] In the example shown in this figure, the feature extraction unit 321 includes a 64×2 intermediate layer 323, a 128×2 intermediate layer 323, a 256×3 intermediate layer 323, and a 512×3 intermediate layer 323 in that order. The upsampling unit 322 also includes a 512×3 intermediate layer 323, a 256×3 intermediate layer 323, a 128×2 intermediate layer 323, and a 64×2 intermediate layer 323 in that order. The second pooling unit 327 connects two 512×3 intermediate layers 323 to each other. The number of intermediate layers 323 constituting the nonlinear mapping unit 320 is not particularly limited and can be determined, for example, according to the number of pixels in the image data.
[0039] Note that this figure shows only one example of the configuration of the nonlinear mapping unit 320, and the nonlinear mapping unit 320 may have other configurations. For example, a 64×1 hidden layer 323 may be included instead of a 64×2 hidden layer 323. Reducing the number of convolutional layers 324 included in the hidden layer 323 may further reduce the computation cost. Also, for example, a 32×2 hidden layer 323 may be included instead of a 64×2 hidden layer 323. Reducing the number of channels in the hidden layer 323 may further reduce the computation cost. Furthermore, both the number of convolutional layers 324 and the number of channels in the hidden layer 323 may be reduced.
[0040] Here, in the plurality of intermediate layers 323 included in the feature extraction unit 321, it is preferable that the number of filters 325 increases each time passing through the first pooling unit 326. Specifically, the first intermediate layer 323a and the second intermediate layer 323b are continuous with each other via the first pooling unit 326, and the second intermediate layer 323b is located at the subsequent stage of the first intermediate layer 323a. The first intermediate layer 323a is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N1, and the second intermediate layer 323b is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N2. At this time, it is preferable that N2 > N1 holds. More preferably, N2 = N1 × 2 holds.
[0041] Also, in the plurality of intermediate layers 323 included in the upsampling unit 322, it is preferable that the number of filters 325 decreases each time passing through the unpooling unit 328. Specifically, the third intermediate layer 323c and the fourth intermediate layer 323d are continuous with each other via the unpooling unit 328, and the fourth intermediate layer 323d is located at the subsequent stage of the third intermediate layer 323c. The third intermediate layer 323c is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N3, and the fourth intermediate layer 323d is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N4. At this time, it is preferable that N4 < N3 holds. More preferably, N3 = N4 × 2 holds.
[0042] The feature extraction unit 321 extracts image features with multiple levels of abstraction, such as gradients and shapes, from the intermediate data acquired from the input unit 310 as channels for the intermediate layer 323. Figure 6 illustrates the configuration of a 64x2 intermediate layer 323. The processing in the intermediate layer 323 will be explained with reference to this figure. In the example shown in this figure, the intermediate layer 323 consists of a first convolutional layer 324a and a second convolutional layer 324b, and each convolutional layer 324 is equipped with 64 filters 325. In the first convolutional layer 324a, a convolution process using the filters 325 is applied to each channel of the data input to the intermediate layer 323. For example, if the image input to the input unit 310 is an RGB image, three channels h 0 i The process is applied to each of (i=1..3). Also, in the example shown in this figure, filter 325 consists of 64 types of 3x3 filters, meaning there are a total of 64 x 3 types of filters. As a result of the convolution process, for each channel i, there are 64 result h 0 i,j (i=1..3, j=1..64) is obtained.
[0043] Next, the output of the multiple filters 325 is activated in the activation unit 329. Specifically, for the corresponding result j of all channels, the activation process is applied to the sum of each corresponding element. This activation process results in the 64 channels of result h. 1 i (i=1..64), that is, the output of the first convolutional layer 324a is obtained as an image feature. The activation process is not particularly limited, but a process using at least one of a hyperbolic function, a sigmoid function, and a normalized linear function is preferred.
[0044] Furthermore, the output data of the first convolutional layer 324a is used as the input data for the second convolutional layer 324b, and the same processing as in the first convolutional layer 324a is performed in the second convolutional layer 324b to obtain the 64-channel result h 2 i (i=1..64), that is, the output of the second convolutional layer 324b is obtained as an image feature. The output of the second convolutional layer 324b becomes the output data of this 64×2 hidden layer 323.
[0045] Here, the structure of the filter 325 is not particularly limited, but it is preferably a 3x3 two-dimensional filter. Also, the coefficients of each filter 325 can be set independently. In this embodiment, the coefficients of each filter 325 are stored in the storage unit 390, and the nonlinear mapping unit 320 can read them and use them for processing. Here, the coefficients of the multiple filters 325 may be determined based on correction information generated and modified using machine learning. For example, the correction information includes the coefficients of the multiple filters 325 as multiple correction parameters. The nonlinear mapping unit 320 can further use this correction information to convert the intermediate data into mapping data. The storage unit 390 may be provided in the landscape image processing unit 3 or may be provided outside the landscape image processing unit 3. Also, the nonlinear mapping unit 320 may acquire the correction information from an external source via a communication network.
[0046] Figures 7(a) and 7(b) show examples of convolution processing performed by filter 325, respectively. Both Figures 7(a) and 7(b) show examples of 3x3 convolution. The example in Figure 7(a) is a convolution using nearest neighbor elements. The example in Figure 7(b) is a convolution using neighbor elements with a distance of two or more. It is also possible to perform convolution using neighbor elements with a distance of three or more. Filter 325 preferably performs convolution processing using neighbor elements with a distance of two or more, because it can extract a wider range of features and further improve the accuracy of visual splendor estimation.
[0047] The operation of the 64x2 hidden layer 323 has been explained above. The operation of other hidden layers 323 (128x2 hidden layer 323, 256x3 hidden layer 323, and 512x3 hidden layer 323, etc.) is the same as that of the 64x2 hidden layer 323, except for the number of convolutional layers 324 and the number of channels. Furthermore, the operation of the hidden layer 323 in the feature extraction unit 321 and the operation of the hidden layer 323 in the upsampling unit 322 are the same as described above.
[0048] Figure 8(a) is a diagram illustrating the process of the first pooling section 326, Figure 8(b) is a diagram illustrating the process of the second pooling section 327, and Figure 8(c) is a diagram illustrating the process of the unpooling section 328.
[0049] In the feature extraction unit 321, the data output from the intermediate layer 323 is subjected to pooling processing for each channel in the first pooling unit 326 before being input to the next intermediate layer 323. In the first pooling unit 326, for example, non-overlapping pooling processing is performed. Figure 8(a) shows the process of associating a 2x2 array of four elements 30 with one element 30 for each group of elements contained in each channel. In the first pooling unit 326, this association is performed for all elements 30. Here, the 2x2 array of four elements 30 are selected so as not to overlap with each other. In this example, the number of elements in each channel is reduced to one-quarter. Note that as long as the number of elements is reduced in the first pooling unit 326, the number of elements 30 before and after the association is not particularly limited.
[0050] The data output from the feature extraction unit 321 is input to the upsampling unit 322 via the second pooling unit 327. In the second pooling unit 327, overlap pooling is applied to the output data from the feature extraction unit 321. Figure 8(b) shows the process of associating 2x2 four elements 30 with one element 30 while overlapping some of the elements 30. That is, in repeated associations, some of the 2x2 four elements 30 in one association are also included in the 2x2 four elements 30 in the next association. In the second pooling unit 327 as shown in this figure, the number of elements is not reduced. Note that the number of elements 30 before and after the association in the second pooling unit 327 is not particularly limited.
[0051] The methods of each process performed in the first pooling unit 326 and the second pooling unit 327 are not particularly limited, but examples include matching the maximum value of four elements 30 to one element 30 (max pooling) and matching the average value of four elements 30 to one element 30 (average pooling).
[0052] The data output from the second pooling unit 327 is input to the intermediate layer 323 in the upsampling unit 322. The output data from the intermediate layer 323 of the upsampling unit 322 is then subjected to amplification processing for each channel in the amplification unit 328 before being input to the next intermediate layer 323. Figure 8(c) shows the process of expanding one element 30 into multiple elements 30. The method of amplification is not particularly limited, but one example is the method of duplicating one element 30 into four elements 30 in a 2x2 arrangement.
[0053] The output data from the last intermediate layer 323 of the upsampling unit 322 is output as mapping data from the nonlinear mapping unit 320 and input to the output unit 330. In the output step S130, the output unit 330 generates a visual splendor map by performing operations such as normalization and resolution conversion on the data acquired from the nonlinear mapping unit 320, and outputs it. The visual splendor map is, for example, an image (image data) that visualizes visual splendor using luminance values, as illustrated in Figure 3(b). The visual splendor map may also be an image color-coded according to visual splendor, such as a heat map, or an image in which visual splendor regions with visual splendor higher than a predetermined standard are marked in a way that makes them distinguishable from other locations. Furthermore, the visual splendor estimation information is not limited to map information shown as an image, etc., but may also be a table or the like that lists information indicating visual splendor regions.
[0054] The monotonicity determination unit 4 determines whether the image input to the landscape image acquisition unit 2 has a monotonic tendency based on the visual saliency map acquired by the landscape image processing unit 3. In this embodiment, various statistical quantities are calculated from the visual saliency map, and it is determined whether there is a monotonic tendency based on these statistical quantities. That is, the monotonicity determination unit 4 functions as a determination unit that determines whether the image has a monotonic tendency using the statistical quantity calculated based on the visual saliency map (visual saliency distribution information).
[0055] Fig. 9 shows a flowchart of the operation of the monotonicity determination unit 4. First, the standard deviation of the luminance of each pixel in the image (for example, Fig. 3(b)) constituting the visual saliency map is calculated (step S11). In this step, first, the average value of the luminance of each image in the image constituting the visual saliency map is calculated. If the image constituting the visual saliency map is H pixels × V pixels and the luminance value at an arbitrary coordinate (k, m) is V VC(k,m) then the average value is calculated by the following equation (1).
Equation
[0056] The standard deviation of the luminance of each image in the image constituting the visual saliency map is calculated from the average value calculated by equation (1). The standard deviation SDEV is calculated by the following equation (2).
Equation
[0057] It is determined whether there are multiple output results for the standard deviation calculated in step S11 (step S12). In this step, since the image input from the landscape image acquisition unit 2 is a moving image and the visual saliency map is acquired in frame units, it is determined whether the standard deviations for multiple frames have been calculated in step S11.
[0058] If there are multiple output results (step S12; Yes), the gaze shift amount is calculated (step S13). In this embodiment, the gaze shift amount is determined by the coordinate distance between the maximum (highest) luminance values in the visual spleness maps of the preceding and succeeding frames in time. The gaze shift amount VSA is calculated by the following equation (3), where (x1, y1) is the coordinate of the highest luminance value in the previous frame and (x2, y2) is the coordinate of the highest luminance value in the subsequent frame.
number
[0059] Then, based on the standard deviation calculated in step S11 and the amount of eye movement calculated in S13, it is determined whether there is a monotonic trend (step S14). In this step, if step S12 is No, a threshold is set for the standard deviation calculated in step S11, and the monotonic trend can be determined by comparing it to that threshold. On the other hand, if step S12 is Yes, a threshold is set for the amount of eye movement calculated in step S13, and the monotonic trend can be determined by comparing it to that threshold.
[0060] Specifically, the monotonicity determination unit 4 functions as a standard deviation calculation unit that calculates the standard deviation of the brightness of each pixel in the image obtained as a visual spleness map (visual spleness distribution information), and as a gaze movement amount calculation unit that calculates the amount of gaze movement between frames based on the images obtained in time series as a visual spleness map (visual spleness distribution information).
[0061] The processing result (judgment result) of the monotony determination unit 4 is output to the outside of the determination device 1. For example, based on this processing result, a visual notification that the vehicle is driving on a monotonous road may be displayed on an in-vehicle display device, or a notification may be given audibly through speakers, or through vibrations of seats, etc. Furthermore, based on the processing result of the monotony determination unit 4, the vehicle may be guided to a place where it can rest. The judgment result may also be used for analyzing the factors that led to near misses during driving.
[0062] Furthermore, while the above explanation used the standard deviation to determine whether the trend was monotonic, it is also possible to determine whether the trend was monotonic based on the average brightness of the visual spleness map, that is, the result of equation (1). In the case of the average brightness, a threshold can be set, similar to the standard deviation, and the trend can be determined by comparing it with that threshold. In this case, the monotonicity determination unit 4 functions as the average value calculation unit.
[0063] Furthermore, by making the configuration shown in Figure 1 a program executed by a computer, for example, it can be made into a judgment program that executes the judgment method. In addition, this judgment program is not limited to being stored in the memory of the judgment device 1, but may also be stored in a storage medium such as a memory card or optical disc.
[0064] In this embodiment, the determination device 1 has a landscape image processing unit 3 that estimates the level of visual splendor within an image based on an image captured from a moving object, and acquires a visual splendor map in a time series. The monotonicity determination unit 4 then determines whether the image has a monotonic tendency based on the standard deviation and gaze movement amount calculated based on the visual splendor map. In this way, it becomes possible to determine whether an image has a monotonic tendency based on the position that a human is likely to focus on. Since the determination is based on the position that a human (driver) is likely to focus on, it is possible to determine the tendency to be close to what the driver perceives as monotonic, and thus the determination can be made with greater accuracy.
[0065] Furthermore, the monotonicity determination unit 4 calculates the standard deviation of the brightness of each pixel in the image obtained as a visual splendor map, and then determines whether the image is monotonic based on the calculated standard deviation. In this way, if an image has a concentration of areas that are easy to focus on, it can be determined to be monotonic.
[0066] Furthermore, the monotonicity determination unit 4 may calculate the average brightness of each pixel in the image obtained as a visual spleness map, and then determine whether the image is monotonic based on the calculated average value. In this way, a monotonicity can be determined when there is a concentration of easily focused areas in an image. In addition, since the determination is made using an average value, the calculation process can be simplified.
[0067] Furthermore, the monotonicity determination unit 4 calculates the amount of eye movement between frames based on the images obtained in time series as a visual splendor map, and then determines whether there is a monotonic tendency based on the calculated amount of eye movement. In this way, when making a determination on a moving image, for example, if the amount of eye movement is small, it can be determined that there is a monotonic tendency.
[0068] Furthermore, the landscape image processing unit 3 includes an input unit 310 that converts an image into intermediate data that can be mapped, a nonlinear mapping unit 320 that converts the intermediate data into mapping data, and an output unit 330 that generates splendor estimation information showing a splendor distribution based on the mapping data. The nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. In this way, visual splendor can be estimated with low computational cost. Moreover, the visual splendor estimated in this way reflects the contextual attention state. [Examples]
[0069] Next, a risk information output device according to a second embodiment of the present invention will be described with reference to Figures 10 and 11. Note that parts identical to those in the first embodiment described above are denoted by the same reference numerals and their descriptions are omitted.
[0070] This embodiment allows for the determination of a monotonic trend even in cases where detection is missed in the method of the first embodiment, particularly when there are multiple output results (video). The block configuration and other aspects are the same as in the first embodiment. Figure 10 shows a flowchart of the operation of the monotonicity determination unit 4 according to this embodiment.
[0071] In the flowchart of Figure 10, steps S11 and S13 are the same as in Figure 9. In this embodiment, since autocorrelation is used as described later, the target image is a moving image, so step S12 is omitted. The judgment content of step S14A is the same as that of step S14. In this embodiment, step S14A is performed as a primary judgment on monotonicity.
[0072] Next, if the result of the determination in step S14A is determined to be monotonic (step S15; Yes), the determination result is output to the outside of the determination device 1, as in Figure 9. On the other hand, if the result of the determination in step S14A is determined to be not monotonic (step S15; No), an autocorrelation calculation is performed (step S16).
[0073] In this embodiment, the autocorrelation is calculated using the standard deviation (average brightness) and line-of-sight movement amount calculated in steps S11 and S13. Autocorrelation R (k) Here, E is the expected value, μ is the mean of X, and σ is the variance of X. 2 It is known that the lag can be calculated using the following equation (4), where k is the lag. In this embodiment, the calculation of equation (4) is performed by varying k within a predetermined range, and the largest calculated value is taken as the autocorrelation value.
number
[0074] Then, based on the calculated autocorrelation value, it is determined whether the trend is monotonic (step S17). The determination can be made by setting a threshold for the autocorrelation value, as in the first embodiment, and comparing it with the threshold to determine whether the trend is monotonic. For example, if the autocorrelation value at k=k1 is greater than the threshold, it means that a similar landscape is repeated every k1. If it is determined that the trend is monotonic, the landscape image is classified as an image with a monotonic trend. By calculating such autocorrelation values, it becomes possible to determine roads with a monotonic trend due to periodically arranged objects such as streetlights that are regularly placed at equal intervals.
[0075] Figure 11 shows an example of the autocorrelation calculation results. Figure 11 is a correlogram of the average luminance of the visual spleness map for a driving video. In Figure 11, the vertical axis represents the correlation function (autocorrelation value), and the horizontal axis represents the lag. In Figure 11, the shaded area represents a 95% confidence interval (significance level αs = 0.05). If the null hypothesis is "there is no periodicity when the lag is k" and the alternative hypothesis is "there is periodicity when the lag is k", then the data within this shaded area is judged to have no periodicity because the null hypothesis cannot be rejected, and data beyond the shaded area is judged to have periodicity regardless of whether it is positive or negative.
[0076] Figure 11(a) is a video of driving through a tunnel, an example of periodicity. Figure 11(a) shows that periodicity can be observed at the 10th and 17th points. In the case of tunnels, since tunnel lighting is placed at regular intervals, it is possible to determine a monotonous trend due to the lighting. On the other hand, Figure 11(b) is a video of driving on a public road, an example of no periodicity. Figure 11(b) shows that most of the data falls within the confidence interval.
[0077] By operating as shown in the flowchart in Figure 10, it becomes possible to first determine if the data is monotonic based on the mean, standard deviation, and eye movement, and then perform a secondary determination from the remaining data based on periodicity.
[0078] In this embodiment, the landscape image processing unit 3 acquires a visual splendor map in a time series, and the monotonicity determination unit 4 calculates statistical quantities from the visual splendor map acquired in a time series. It functions as a first determination unit that determines whether there is a monotonicity based on the statistical quantities obtained in the time series, and a second determination unit that determines whether there is a monotonicity based on autocorrelation. In this way, monotonicity caused by periodically appearing objects such as streetlights that appear while driving, which is difficult to determine using statistical quantities alone, can be determined by autocorrelation.
[0079] Furthermore, the present invention is not limited to the embodiments described above. That is, those skilled in the art can implement the invention in various ways, without departing from the core principles, in accordance with conventionally known knowledge. Such modifications, as long as they still incorporate the determination device of the present invention, are of course included within the scope of the present invention. [Explanation of symbols]
[0080] 1 Judgment device 2. Landscape image acquisition unit 3. Landscape Image Processing Unit (Acquisition Unit) 4. Monotonicity determination unit (determination unit, standard deviation calculation unit, mean calculation unit, eye movement amount calculation unit)
Claims
[Claim 1] An acquisition unit that acquires visual splendor distribution information obtained by estimating the level of visual splendor within an image based on an image captured from a moving object, A determination device comprising a determination unit that determines whether the image has a monotonic tendency using a statistical quantity calculated based on the aforementioned visual spleness distribution information.
Citation Information
Patent Citations
Awakening state determining device and awakening state determining method
JP2010128649A
Landscape Monotonicity Calculation Device and Method
JP4550116B2