State output device
The situation output device automates the analysis of driving environments by using visual saliency distribution information to estimate trends, enhancing the efficiency and accuracy of identifying near misses in vehicles.
Patent Information
- Application Number
- JP2025107841
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-28
AI Technical Summary
Existing systems for analyzing the causes of near misses in vehicles require complex sensor communication and manual review of images to identify specific factors, which is time-consuming and inefficient.
A situation output device that utilizes a visual saliency distribution information acquisition unit to estimate visual trends from captured images, allowing for automated analysis of driving environments and outputting situations based on visual features without manual image review.
Enables efficient and automated analysis of driving situations, reducing the need for manual image review and improving the speed and accuracy of identifying near misses by analyzing visual saliency distribution information.
Smart Images

Figure 2025126282000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a situation output device that outputs a situation based on an image of the outside taken from a moving object. [Background technology]
[0002] For example, a drive recorder detects the occurrence of an accident or near miss based on the acceleration of the vehicle, and records images before and after the occurrence.
[0003] Patent Document 1 describes a factor analysis device that includes a common point identification unit 220 that compares vehicle information transmitted from a vehicle with vehicle information accumulated about accidents and near misses that have occurred in the past to identify common points, an environmental factor estimation unit 230 that estimates whether or not an environmental factor is the cause of an accident or near miss that has occurred in the vehicle based on the common points, a driver factor estimation unit 240 that estimates whether or not a driver factor is the cause of an accident or near miss that has occurred in the vehicle based on the common points, and a factor determination unit 250 that determines the main cause of the vehicle accident or near miss based on the estimation results by the environmental factor estimation unit 230 and the driver factor estimation unit 240. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-71492 Summary of the Invention [Problem to be solved by the invention]
[0005] The invention described in Patent Document 1 collects various information such as area, date and time, weather, vehicle information, external sensor data, and driver data in order to estimate the causes of near misses, etc. Therefore, in order to obtain information inside and outside the vehicle, it is necessary to provide a means for communicating with sensors, etc., which is not easy to implement.
[0006] Furthermore, in the invention described in Patent Document 1, only specific factors such as a curve with poor visibility are recorded as estimated factors, which means that it is necessary to check the detailed circumstances of why the factor occurred, which requires time and effort, such as re-examining the image of the near miss or interviewing the driver.
[0007] One example of a problem that the present invention aims to solve is how to estimate the circumstances under which a near miss or the like occurs. [Means for solving the problem]
[0008] In order to solve the above problem, the invention described in claim 1 is characterized by comprising a visual saliency distribution information acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency within an image captured of the outside from a moving body, a visual feature extraction unit that acquires visual trends in the movement environment of the moving body based on the visual saliency distribution information, and a situation output unit that outputs the situation of the image based on the visual trends.
[0009] The invention described in claim 13 is a situation output method executed by a situation output device that outputs a situation based on an image of the outside taken from a moving body, characterized by including a visual saliency distribution information acquisition step of acquiring visual saliency distribution information obtained by estimating the level of visual saliency within the image based on the image, a visual feature extraction step of acquiring a visual tendency of the movement environment of the moving body based on the visual saliency distribution information, and a situation output step of outputting the situation of the image acquired based on the visual tendency.
[0010] The invention as set forth in claim 14 is characterized in that the situation output method as set forth in claim 13 is executed by a computer.
[0011] The invention as set forth in claim 15 is characterized in that the status output program as set forth in claim 14 is stored. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a schematic configuration diagram of a system including a determination device according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a functional configuration diagram of the situation output device shown in FIG. [Figure 3] 2 is a block diagram illustrating the configuration of a visual saliency calculation unit shown in FIG. 1. FIG. [Figure 4] 1A is a diagram illustrating an example of an image input to a determination device, and FIG. 1B is a diagram illustrating an example of a visual saliency map estimated for FIG. 1A. [Figure 5] 2 is a flowchart illustrating a processing method of the visual saliency calculation unit shown in FIG. 1; [Figure 6] FIG. 2 is a diagram illustrating in detail an example of the configuration of a nonlinear mapping unit. [Figure 7] FIG. 2 is a diagram illustrating the configuration of an intermediate layer. [Figure 8] 10(a) and 10(b) are diagrams illustrating examples of convolution processing performed by a filter. [Figure 9] (a) is a diagram for explaining the processing of the first pooling unit, (b) is a diagram for explaining the processing of the second pooling unit, and (c) is a diagram for explaining the processing of the unpooling unit. [Figure 10] 2 is a block diagram illustrating the configuration of a visual attention concentration level calculation unit shown in FIG. 1. FIG. [Figure 11] FIG. 10 is an explanatory diagram of a vector error. [Figure 12] 2 is an example of an image input to the image input unit shown in FIG. 1 and a visual saliency map obtained from the image. [Figure 13] 10 is a graph showing an example of temporal changes in visual attention concentration level. [Figure 14] 2 is a block diagram illustrating the configuration of an inattentive driving tendency calculation unit shown in FIG. 1. FIG. [Figure 15] FIG. 10 is an explanatory diagram of a method for setting a gaze area. [Figure 16] FIG. 10 is an explanatory diagram of an aside-looking detection area. [Figure 17]FIG. 10 is an explanatory diagram of another inattentive looking detection area. [Figure 18] 2 is a flowchart of the operation of an inattentiveness tendency calculation unit shown in FIG. 1. [Figure 19] 2 is a flowchart of the operation of the monotonic tendency calculation unit shown in FIG. 1. [Figure 20] 10 is a flowchart of an operation of a modified example of the monotonic tendency calculation unit. [Figure 21] 10 is an example of a calculation result of autocorrelation. [Figure 22] FIG. 2 is a functional configuration diagram of a visual load tendency calculation unit shown in FIG. [Figure 23] 23 is a waveform diagram showing the operation of each function in the visual load tendency calculation unit shown in FIG. 22. FIG. [Figure 24] FIG. 2 is an explanatory diagram (part 1) showing an example of text generation in the situation output unit shown in FIG. 1. [Figure 25] FIG. 2 is an explanatory diagram (part 2) showing an example of text generation in the situation output unit shown in FIG. [Figure 26] 2 is a flowchart of the operation of the situation output device shown in FIG. 1. [Figure 27] 2 is an example of a screen output by the situation output unit shown in FIG. 1. [Figure 28] 2 is a modified example of the situation output device shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0013] A situation output device according to one embodiment of the present invention will be described below. In the situation output device according to one embodiment of the present invention, a visual saliency distribution information acquisition unit acquires visual saliency distribution information obtained by estimating the level of visual saliency in an image based on an image of the outside captured from a mobile body, and a visual feature extraction unit acquires a visual trend of the moving environment of the mobile body based on the visual saliency distribution information. Then, a situation output unit outputs the situation of the acquired image based on the visual trend. In this way, the situation of the image can be output from the visual trend based on the visual saliency distribution information. Therefore, it is possible to analyze the situation using only the image, and the situation in which a near miss or the like occurred can be estimated without checking the image.
[0014] The status output unit may also output the status of the image as text information, which can be useful for creating reports and the like.
[0015] The situation output unit may also output text information that combines multiple keywords based on each of the multiple pieces of information acquired based on the visual tendency. This allows the situation to be constructed as a sentence that relates not only to a specific factor but also to multiple factors. This makes it easier to understand the detailed situation.
[0016] The visual feature extraction unit may also include a visual attention concentration level acquisition unit that acquires information indicating the level of visual attention concentration in the image based on the visual saliency distribution information and a reference gaze position that is predetermined for the image. In this way, the situation of the image can be estimated based on the level of visual attention concentration.
[0017] Furthermore, the visual attention concentration level acquisition unit may acquire the visual attention concentration level based on the value of each pixel constituting the visual saliency distribution information and the vector error between the position of each pixel and the coordinate position of the reference gaze position. In this way, a value corresponding to the difference between a position of high visual saliency and the reference gaze position is calculated as the visual attention concentration level. Therefore, for example, the value of the visual attention concentration level can be changed depending on the distance between the position of high visual saliency and the reference gaze position.
[0018] The visual feature extraction unit may also include an inattentiveness acquisition unit that acquires information indicating a tendency to look aside based on the visual saliency distribution information, thereby making it possible to estimate the situation of the image based on the information indicating the tendency to look aside.
[0019] The inattentiveness acquisition unit may also include a peak position detection unit that detects at least one peak position in the visual saliency distribution information in a time series, and a gaze range setting unit that sets a range in the image where the driver of the moving object should gaze, and may output information indicating a tendency for inattentiveness if the peak position is out of the gaze range for a predetermined period of time or more. This visual saliency distribution information indicates the statistical likelihood of human gaze being focused. Therefore, the peak of the visual saliency distribution information indicates the position where human gaze is most likely to be focused statistically. Therefore, by using the visual saliency distribution information, it is possible to detect a tendency for inattentiveness with a simple configuration without measuring the actual driver's gaze.
[0020] The visual feature extraction unit may further include a monotony acquisition unit that acquires information indicating a monotony tendency of the image based on the visual saliency distribution information, thereby enabling estimation of the image situation based on the information indicating the monotony tendency.
[0021] The monotony acquisition unit may also acquire information indicating a monotony tendency using statistics calculated based on the visual saliency distribution information. In this way, it becomes possible to determine whether a captured image has a monotony tendency based on positions that are likely to be gazed upon by humans. Since the determination is based on positions that are likely to be gazed upon by humans (drivers), it is possible to determine a tendency that is close to what the driver perceives as monotony, thereby enabling more accurate determination.
[0022] The visual feature extraction unit may further include a visual load acquisition unit that acquires information indicating a visual load tendency of the image based on the visual saliency distribution information. In this way, the situation of the image can be estimated based on the information indicating the visual load tendency.
[0023] The system may also include a storage unit that stores the situations output by the situation output unit for each driver, and an analysis unit that analyzes the driving tendencies of the driver based on the situations stored in the storage unit and outputs the analysis results. This makes it possible to analyze the driving tendencies of the driver and provide guidance, etc. using the analysis results.
[0024] The visual saliency distribution information acquisition unit may include an input unit that converts the image into intermediate data that can be mapped, a nonlinear mapping unit that converts the intermediate data into mapped data, and an output unit that generates saliency estimation information indicating a saliency distribution based on the mapped data, and the nonlinear mapping unit may include a feature extraction unit that extracts features from the intermediate data and an upsampling unit that upsamples the data generated by the feature extraction unit. This allows visual saliency to be estimated with low computational cost. Furthermore, the visual saliency estimated in this manner reflects a contextual attention state.
[0025] Furthermore, in a situation output method according to one embodiment of the present invention, a visual saliency distribution information acquisition step acquires visual saliency distribution information obtained by estimating the level of visual saliency within an image captured of the outside from a mobile body, and a visual feature extraction step acquires a visual trend of the moving environment of the mobile body based on the visual saliency distribution information. Then, a situation output step outputs the situation of the acquired image based on the visual trend. In this way, the situation of the image can be output from the visual trend based on the visual saliency distribution information. Therefore, it is possible to analyze the situation using only the image, and to estimate the situation in which a near miss or the like occurred.
[0026] Furthermore, the above-described situation output method is executed by a computer. In this way, the situation of the image can be output from the visual tendency based on the visual saliency distribution information using a computer. Therefore, it is possible to analyze the situation using only the image, and to estimate the situation in which a near miss or the like occurred.
[0027] The status output program may be stored in a computer-readable storage medium. This allows the program to be distributed as a standalone program rather than being incorporated into a device, and allows for easy version upgrades. [Example]
[0028] A situation output device according to one embodiment of the present invention will be described with reference to Figs. 1 to 28. The situation output device according to this embodiment may be configured as a server device or the like installed in a business establishment or the like, or may be installed in a mobile body such as an automobile (see Fig. 1). The situation output device mainly performs analysis after driving, but may also perform analysis in real time.
[0029] FIG. 1 shows an example in which a situation output device is configured as a server device. In FIG. 1, when a vehicle behavior detection unit 12, such as an acceleration sensor, detects a large acceleration due to sudden braking, sudden acceleration, or other impact in a drive recorder 11 mounted on a vehicle V, images (video images) from a predetermined period before and after the event are transmitted to a determination device 1 via a network N, such as the Internet. The vehicle behavior detection unit 12 shown in FIG. 2 is not limited to an acceleration sensor, but may also be an anti-lock braking system (ABS) or a skid prevention device mounted on the vehicle V. The activation of such a device (sensor) may trigger the transmission or storage of images from a predetermined period before and after the event. By processing images in which sudden braking or other events are detected, as described below, it is possible to narrow down the number of images to be processed, thereby reducing the processing time (processing volume). The transmission of images is not limited to the form of communication shown in FIG. 1 , and images may also be read from a recording medium, such as a hard disk drive or memory card, connected to the server device 1. Furthermore, the status output device is not limited to a server device, but may be an in-vehicle device incorporating each block described below, or a PC terminal at home or at work, or may be configured to distribute processing between these devices and the server.
[0030] As shown in Figure 2, the situation output device 1 includes an image input unit 2, a visual saliency calculation unit 3, a visual attention concentration calculation unit 4, a distraction tendency calculation unit 5, a monotony tendency calculation unit 6, a situation output unit 7, and a visual load tendency calculation unit 8.
[0031] The image input unit 2 receives an image (e.g., a moving image) captured by a camera or the like, and outputs the image as image data. The input moving image is output as image data broken down into a time series, such as for each frame. Although still images may be input as images to the image input unit 2, it is preferable to input them as an image group consisting of a plurality of still images in time series.
[0032] The images input to the image input unit 2 include, for example, images captured in the direction of travel of the vehicle. In other words, images are taken of the outside world continuously from a moving body. These images may be so-called panoramic images or images acquired using multiple cameras, and may include images that include angles other than the direction of travel, such as 180° or 360° in the horizontal direction. Furthermore, the images input to the image input unit 2 are not limited to images captured by a camera, and may also be images read from a recording medium such as a hard disk drive or a memory card.
[0033] The visual saliency calculation unit 3 receives image data from the image input unit 2 and outputs a visual saliency map as visual saliency estimation information (to be described later). That is, the visual saliency calculation unit 3 functions as a visual saliency distribution information acquisition unit that acquires a visual saliency map (visual saliency distribution information) obtained by estimating the level of visual saliency based on an image of the outside captured from a moving object.
[0034] FIG. 3 is a block diagram illustrating the configuration of the visual saliency calculation unit 3. The visual saliency calculation unit 3 according to this embodiment includes an input unit 310, a nonlinear mapping unit 320, an output unit 330, and a storage unit 390. The input unit 310 converts an image into intermediate data that can be subjected to mapping processing. The nonlinear mapping unit 320 converts the intermediate data into mapped data. The output unit 330 generates saliency estimation information indicating a saliency distribution based on the mapped data. The nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. The storage unit 390 stores image data input from the image input unit 2, filter coefficients (described later), and the like. These are described in detail below.
[0035] FIG. 4(a) is a diagram illustrating an example of an image input to the visual saliency calculation unit 3, and FIG. 4(b) is a diagram illustrating an example of an image showing a visual saliency distribution estimated for FIG. 4(a). The visual saliency calculation unit 3 according to this embodiment is a device that estimates the visual saliency of each part in an image. Visual saliency means, for example, how easily something stands out or how easily it attracts attention. Specifically, visual saliency is expressed as a probability or the like. Here, the magnitude of the probability corresponds to, for example, the probability that a person viewing the image will direct their gaze to that position.
[0036] 4(a) and 4(b) correspond to each other in position. In FIG. 4(a), the higher the visual saliency, the higher the brightness displayed in FIG. 4(b). The image showing the visual saliency distribution as shown in FIG. 4(b) is an example of a visual saliency map output by the output unit 330. In this example, visual saliency is visualized using brightness values of 256 levels. An example of the visual saliency map output by the output unit 330 will be described in detail later.
[0037] FIG. 5 is a flowchart illustrating the operation of the visual saliency calculation unit 3 according to this embodiment. The flowchart shown in FIG. 5 is part of a situation output method executed by a computer, and includes an input step S115, a nonlinear mapping step S120, and an output step S130. In the input step S115, an image is converted into intermediate data that can be mapped. In the nonlinear mapping step S120, the intermediate data is converted into mapped data. In the output step S130, visual saliency estimation information (visual saliency distribution information) indicating a saliency distribution is generated based on the mapped data. Here, the nonlinear mapping step S120 includes a feature extraction step S121 that extracts features from the intermediate data, and an upsampling step S122 that upsamples the data generated in the feature extraction step S121.
[0038] Returning to FIG. 3, each component of the visual saliency calculation unit 3 will be described. In input step S115, the input unit 310 acquires an image and converts it into intermediate data. The input unit 310 acquires image data from the image input unit 2. The input unit 310 then converts the acquired image into intermediate data. The intermediate data is not particularly limited as long as it is data that can be accepted by the nonlinear mapping unit 320, and is, for example, a high-dimensional tensor. Furthermore, the intermediate data is, for example, data in which the brightness of the acquired image is normalized, or data in which each pixel of the acquired image is converted into a brightness gradient. In input step S115, the input unit 310 may further perform noise removal, resolution conversion, etc. on the image.
[0039] In the nonlinear mapping step S120, the nonlinear mapping unit 320 acquires intermediate data from the input unit 310. The nonlinear mapping unit 320 then converts the intermediate data into mapping data. Here, the mapping data is, for example, a high-dimensional tensor. The mapping process performed on the intermediate data by the nonlinear mapping unit 320 is, for example, a mapping process that can be controlled by parameters, and is preferably a process using a function, a functional, or a neural network.
[0040] Fig. 6 is a diagram illustrating a detailed configuration of the nonlinear mapping unit 320, and Fig. 7 is a diagram illustrating a configuration of the hidden layer 323. As described above, the nonlinear mapping unit 320 includes a feature extraction unit 321 and an upsampling unit 322. The feature extraction unit 321 performs the feature extraction step S121, and the upsampling unit 322 performs the upsampling step S122. In the example shown in this figure, at least one of the feature extraction unit 321 and the upsampling unit 322 is configured to include a neural network including a plurality of hidden layers 323. In the neural network, a plurality of hidden layers 323 are connected.
[0041] In particular, the neural network is preferably a convolutional neural network. Specifically, each of the multiple hidden layers 323 includes one or more convolutional layers 324. In the convolutional layers 324, input data is convolved by multiple filters 325, and activation processing is performed on the outputs of the multiple filters 325.
[0042] 6, feature extraction unit 321 is configured to include a neural network including a plurality of hidden layers 323, and a first pooling unit 326 is provided between the plurality of hidden layers 323. Furthermore, upsampling unit 322 is configured to include a neural network including a plurality of hidden layers 323, and an unpooling unit 328 is provided between the plurality of hidden layers 323. Furthermore, feature extraction unit 321 and upsampling unit 322 are connected to each other via a second pooling unit 327 that performs overlap pooling.
[0043] In the example shown in this figure, each intermediate layer 323 is made up of two or more convolutional layers 324. However, at least some of the intermediate layers 323 may be made up of only one convolutional layer 324. Adjacent intermediate layers 323 are separated by any of a first pooling unit 326, a second pooling unit 327, and an unpooling unit 328. Here, when an intermediate layer 323 includes two or more convolutional layers 324, it is preferable that the number of filters 325 in those convolutional layers 324 be equal to each other.
[0044] In this figure, an "A×B" hidden layer 323 is composed of B convolution layers 324, and each convolution layer 324 includes A convolution filters for each channel. Such a hidden layer 323 is also referred to as an "A×B hidden layer" below. For example, a 64×2 hidden layer 323 is composed of two convolution layers 324, and each convolution layer 324 includes 64 convolution filters for each channel.
[0045] In the example shown in the figure, the feature extraction unit 321 includes a 64×2 hidden layer 323, a 128×2 hidden layer 323, a 256×3 hidden layer 323, and a 512×3 hidden layer 323, in this order. The upsampling unit 322 includes a 512×3 hidden layer 323, a 256×3 hidden layer 323, a 128×2 hidden layer 323, and a 64×2 hidden layer 323, in this order. The second pooling unit 327 connects the two 512×3 hidden layers 323 to each other. The number of hidden layers 323 constituting the nonlinear mapping unit 320 is not particularly limited and can be determined, for example, according to the number of pixels in the image data.
[0046] Note that this diagram shows an example of the configuration of the nonlinear mapping unit 320, and the nonlinear mapping unit 320 may have other configurations. For example, a 64×1 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. Reducing the number of convolutional layers 324 included in the intermediate layer 323 may further reduce the computational cost. Also, for example, a 32×2 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. Reducing the number of channels in the intermediate layer 323 may further reduce the computational cost. Furthermore, both the number of convolutional layers 324 and the number of channels in the intermediate layer 323 may be reduced.
[0047] Here, in the multiple intermediate layers 323 included in the feature extraction unit 321, it is preferable that the number of filters 325 increases each time the data passes through the first pooling unit 326. Specifically, the first intermediate layer 323a and the second intermediate layer 323b are connected to each other via the first pooling unit 326, and the second intermediate layer 323b is located after the first intermediate layer 323a. The first intermediate layer 323a is configured with convolutional layers 324 in which the number of filters 325 for each channel is N1, and the second intermediate layer 323b is configured with convolutional layers 324 in which the number of filters 325 for each channel is N2. In this case, it is preferable that N2 > N1. It is more preferable that N2 = N1 × 2.
[0048] In addition, in the plurality of intermediate layers 323 included in the upsampling unit 322, it is preferable that the number of filters 325 decreases every time it passes through the unpooling unit 328. Specifically, the third intermediate layer 323c and the fourth intermediate layer 323d are continuous with each other via the unpooling unit 328, and the fourth intermediate layer 323d is located after the third intermediate layer 323c. The third intermediate layer 323c is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N3, and the fourth intermediate layer 323d is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N4. At this time, it is preferable that N4 < N3 holds. More preferably, N3 = N4 × 2 holds.
[0049] In the feature extraction unit 321, image features having a plurality of levels of abstraction, such as gradients and shapes, are extracted from the intermediate data acquired from the input unit 310 as channels of the intermediate layer 323. FIG. 7 illustrates the configuration of the 64×2 intermediate layer 323. Referring to this figure, the processing in the intermediate layer 323 will be described. In the example of this figure, the intermediate layer 323 is composed of a first convolutional layer 324a and a second convolutional layer 324b, and each convolutional layer 324 includes 64 filters 325. In the first convolutional layer 324a, convolution processing using the filter 325 is performed on each channel of the data input to the intermediate layer 323. For example, when the image input to the input unit 310 is an RGB image, processing is performed on each of the three channels h 0 i (i = 1..3). Also, in the example of this figure, the filter 325 is a 64 types of 3×3 filters, that is, a total of 64×3 types of filters. As a result of the convolution processing, for each channel i, 64 results h 0 i,j (i = 1..3, j = 1..64) are obtained.
[0050] Next, activation processing is performed on the outputs of the plurality of filters 325 in the activation unit 329. Specifically, activation processing is performed on the sum of corresponding elements for the corresponding results j of all channels. By this activation processing, 64-channel results h 1i (i=1..64), i.e., the output of the first convolution layer 324a, is obtained as the image feature. The activation process is not particularly limited, but a process using at least one of a hyperbolic function, a sigmoid function, and a rectified linear function is preferable.
[0051] Furthermore, the output data of the first convolution layer 324a is used as input data for the second convolution layer 324b, and the same processing as that of the first convolution layer 324a is performed in the second convolution layer 324b to obtain the result h of 64 channels. 2 i (i=1..64), that is, the output of the second convolutional layer 324b, is obtained as the image features. The output of the second convolutional layer 324b becomes the output data of this 64×2 hidden layer 323.
[0052] Here, the structure of the filter 325 is not particularly limited, but a 3x3 two-dimensional filter is preferable. Furthermore, the coefficients of each filter 325 can be set independently. In this embodiment, the coefficients of each filter 325 are stored in the memory unit 390, and the nonlinear mapping unit 320 can read and use them for processing. Here, the coefficients of the multiple filters 325 may be determined based on correction information generated and corrected using machine learning. For example, the correction information includes the coefficients of the multiple filters 325 as multiple correction parameters. The nonlinear mapping unit 320 can further use this correction information to convert the intermediate data into mapped data. The memory unit 390 may be provided in the visual saliency calculation unit 3 or external to the visual saliency calculation unit 3. Furthermore, the nonlinear mapping unit 320 may obtain the correction information from an external source via a communication network.
[0053] 8(a) and 8(b) are diagrams illustrating examples of convolution processing performed by the filter 325. Both of FIGS. 8(a) and 8(b) illustrate examples of 3×3 convolution. The example in FIG. 8(a) illustrates convolution processing using nearest neighbor elements. The example in FIG. 8(b) illustrates convolution processing using neighbor elements with a distance of two or more. Note that convolution processing using neighbor elements with a distance of three or more is also possible. It is preferable that the filter 325 performs convolution processing using neighbor elements with a distance of two or more. This is because it allows for the extraction of a wider range of features, thereby further improving the accuracy of visual saliency estimation.
[0054] The above has described the operation of the 64×2 hidden layer 323. The operations of the other hidden layers 323 (such as the 128×2 hidden layer 323, the 256×3 hidden layer 323, and the 512×3 hidden layer 323) are the same as the operation of the 64×2 hidden layer 323, except for the number of convolutional layers 324 and the number of channels. Furthermore, the operations of the hidden layer 323 in the feature extraction unit 321 and the hidden layer 323 in the upsampling unit 322 are also the same as those described above.
[0055] 9(a) is a diagram for explaining the processing of the first pooling unit 326, FIG. 9(b) is a diagram for explaining the processing of the second pooling unit 327, and FIG. 9(c) is a diagram for explaining the processing of the unpooling unit 328.
[0056] In the feature extraction unit 321, data output from the intermediate layer 323 is subjected to pooling processing for each channel in the first pooling unit 326, and then input to the next intermediate layer 323. The first pooling unit 326 performs, for example, non-overlapping pooling processing. FIG. 9(a) shows processing for associating four 2×2 elements 30 with one element 30 for a group of elements included in each channel. The first pooling unit 326 performs such association for all elements 30. Here, the four 2×2 elements 30 are selected so that they do not overlap with each other. In this example, the number of elements in each channel is reduced to one-fourth. Note that, as long as the number of elements is reduced in the first pooling unit 326, the number of elements 30 before and after the association is not particularly limited.
[0057] The data output from the feature extraction unit 321 is input to the upsampling unit 322 via the second pooling unit 327. The second pooling unit 327 performs overlap pooling on the output data from the feature extraction unit 321. FIG. 9(b) shows a process of associating four 2×2 elements 30 with one element 30 while overlapping some of the elements 30. That is, in repeated associations, some of the four 2×2 elements 30 in a certain association are also included in the four 2×2 elements 30 in the next association. The second pooling unit 327 shown in this figure does not reduce the number of elements. Note that the number of elements 30 before and after association in the second pooling unit 327 is not particularly limited.
[0058] The methods of processing performed by the first pooling unit 326 and the second pooling unit 327 are not particularly limited, but examples include matching in which the maximum value of four elements 30 is matched to one element 30 (max pooling) and matching in which the average value of four elements 30 is matched to one element 30 (average pooling).
[0059] The data output from the second pooling unit 327 is input to the hidden layer 323 in the upsampling unit 322. Then, the output data from the hidden layer 323 of the upsampling unit 322 undergoes unpooling processing for each channel in the unpooling unit 328, and is then input to the next hidden layer 323. Figure 9(c) shows processing for expanding one element 30 into multiple elements 30. The method of expansion is not particularly limited, but an example is a method of duplicating one element 30 into four elements 30 (2 x 2).
[0060] The output data of the last hidden layer 323 of the upsampling unit 322 is output as mapping data from the nonlinear mapping unit 320 and input to the output unit 330. In the output step S130, the output unit 330 generates and outputs a visual saliency map by performing, for example, normalization or resolution conversion on the data acquired from the nonlinear mapping unit 320. The visual saliency map is, for example, an image (image data) that visualizes visual saliency using brightness values, as illustrated in FIG. 4(b). The visual saliency map may also be, for example, an image that is color-coded according to visual saliency, such as a heat map, or an image in which visual saliency regions with visual saliency higher than a predetermined standard are marked so as to be distinguishable from other positions. Furthermore, the visual saliency estimation information is not limited to map information displayed as an image or the like, but may also be a table listing information indicating visual saliency regions.
[0061] As shown in FIG. 10, the visual attention concentration level calculation unit 4 includes a line-of-sight coordinate setting unit 41 and a vector error calculation unit .
[0062] The gaze coordinate setting unit 41 sets an ideal gaze, described later, on a visual saliency map. The ideal gaze refers to the gaze of a vehicle driver along the direction of travel in an ideal traffic environment with no obstacles or other traffic participants. It is handled as an (x, y) coordinate in image data and on the visual saliency map. In this embodiment, the ideal gaze is a fixed value. However, it may be treated as a function of the speed or road friction coefficient, which affect the stopping distance of a moving object, or may be determined using set route information. A vanishing point corresponding to the current road may also be used to calculate the ideal viewpoint. In this case, the vehicle speed may be detected and the ideal viewpoint may be set two or three seconds after the vanishing point and the vehicle's position. In other words, the gaze coordinate setting unit 41 functions as a gaze position setting unit that sets the ideal gaze (reference gaze position) in the image according to a predetermined rule.
[0063] The vector error calculation unit 42 calculates a vector error based on the visual saliency map output by the visual saliency calculation unit 3 and the ideal gaze set by the gaze coordinate setting unit 41 for the visual saliency map and the image, and calculates a visual attention concentration level Ps (described later) indicating the degree of visual attention concentration based on the vector error. That is, the vector error calculation unit 42 functions as a visual attention concentration level calculation unit that calculates the degree of visual attention concentration in an image based on the visual saliency distribution information and the gaze position.
[0064] Here, the vector error in this embodiment will be described with reference to FIG. 11. FIG. 11 shows an example of a visual saliency map. This visual saliency map is shown with brightness values of 256 gradations of H pixels x V pixels, and similarly to FIG. 4, pixels with higher visual saliency are displayed with higher brightness. In FIG. 11, the coordinates (x, y) of the ideal line of sight are (x, y)=(x im ,y im), the vector error with the pixel at any coordinate (k, m) in the visual saliency map is calculated. When the coordinate with high brightness in the visual saliency map is far from the coordinate of the ideal gaze, it means that the position to be gazed at and the position that is actually easy to gaze at are far apart, and it can be said that the image is likely to distract visual attention. On the other hand, when the coordinate with high brightness is close to the coordinate of the ideal gaze, it means that the position to be gazed at and the position that is actually easy to gaze at are close, and it can be said that the image is likely to focus visual attention on the position to be gazed at.
[0065] Next, we will explain how to calculate the visual attention concentration level Ps in the vector error calculation unit 4. In this embodiment, the visual attention concentration level Ps is calculated by the following equation (1).
number
[0066] In equation (1), V vc is the pixel depth (brightness value), f w is the weighting function, d err indicates the vector error. This weighting function is, for example, V vc is a function that sets a weight based on the distance from the pixel showing the value of α to the coordinate of the ideal gaze. α is a coefficient that makes the visual attention concentration level Ps equal to 1 when the coordinate of the bright spot and the coordinate of the ideal gaze match in a visual saliency map (reference heat map) of one bright spot.
[0067] That is, the vector error calculation unit 42 calculates the degree of visual attention concentration based on the value of each pixel that constitutes the visual saliency map (visual saliency distribution information) and the vector error between the position of each pixel and the coordinate position of the ideal gaze (reference gaze position).
[0068] The visual attention concentration level Ps obtained in this way is the reciprocal of the weighted sum of the vector error and luminance value of all pixel coordinates from the ideal gaze coordinates set on the visual saliency map. The visual attention concentration level Ps is calculated to be a low value when the high-luminance distribution of the visual saliency map is far from the ideal gaze coordinates. In other words, the visual attention concentration level Ps can also be considered as the concentration level relative to the ideal gaze. That is, the visual attention concentration level calculation unit 4 functions as a visual attention concentration level acquisition unit that acquires the visual attention concentration level Ps (information indicating the visual attention concentration level) for an image based on the visual saliency map (visual saliency distribution information) and the ideal gaze (a reference gaze position previously determined for the image).
[0069] Fig. 12 shows an example of an image input to the image input unit 2 and a visual saliency map obtained from that image. Fig. 12(a) is the input image, and (b) is the visual saliency map. In Fig. 12, if the coordinates of the ideal gaze are set on the road, for example, on a truck traveling ahead, the visual attention concentration level Ps in that case can be calculated.
[0070] The visual attention concentration level Ps calculated in this way can be used to determine, based on its change over time, whether the image input from the image input unit 2 is suspected of causing a safety issue such as an accident or near miss while the moving object was in motion.
[0071] Fig. 13 shows an example of how the visual concentration level Ps changes over time. Fig. 13 shows how the visual concentration level Ps changes in a 12-second video. In Fig. 13, the visual concentration level Ps changes suddenly between about 6.5 seconds and about 7 seconds. This occurs, for example, when another vehicle cuts in front of the vehicle, and by detecting such a change, it is possible to detect an incident that could be a near miss.
[0072] As shown in FIG. 14, the inattentive driving tendency calculation unit 5 includes a visual saliency peak detection unit 53 and an inattentive driving tendency determination unit 54.
[0073] The visual saliency peak detection unit 53 detects peak positions (pixels) in the visual saliency map acquired by the visual saliency calculation unit 3. Here, in this embodiment, a peak is a pixel with high visual saliency where the pixel value is maximum (maximum brightness), and the position is expressed by coordinates. That is, the visual saliency peak detection unit 53 functions as a peak position detection unit that detects at least one peak position in the visual saliency map (visual saliency distribution information).
[0074] The inattentive driving tendency determination unit 54 determines whether the image input from the image input unit 2 has an inattentive driving tendency based on the peak position detected by the visual saliency peak detection unit 53. The inattentive driving tendency determination unit 54 first sets a gaze area (a range to be gazed at) for the image input from the image input unit 2. A method for setting the gaze area will be described with reference to Fig. 15. That is, the inattentive driving tendency determination unit 54 functions as a gaze range setting unit that sets a range in the image to be gazed at by the driver of the moving object.
[0075] In the image P shown in FIG. 15, the gaze area G is set around the vanishing point V. That is, the gaze area G (the range to be gazed at) is set based on the vanishing point of the image. The size of this gaze area G (for example, 3 m wide and 2 m high) is set in advance, and the number of pixels of the set size can be calculated from the number of horizontal pixels, the number of vertical pixels, the horizontal angle of view, the vertical angle of view, the distance to the leading vehicle, the mounting height of the camera such as a drive recorder capturing the image, and the like. The vanishing point may be estimated from a white line or the like, or may be estimated using optical flow or the like. Furthermore, the distance to the leading vehicle does not need to be detected actually, and may be set virtually.
[0076] Next, an inattentiveness detection area is set in image P based on the set gaze area G (shaded area in FIG. 16 ). The inattentiveness detection area is divided into an upper area Iu, a lower area Id, a left side area Il, and a right side area Ir. These areas are separated by line segments connecting the vanishing point V and the vertices of gaze area G. That is, the upper area Iu and the left side area Il are separated by line segment L1 connecting the vanishing point V and the vertex Ga of gaze area G. The upper area Iu and the right side area Ir are separated by line segment L2 connecting the vanishing point V and the vertex Gd of gaze area G. The lower area Id and the left side area Il are separated by line segment L3 connecting the vanishing point V and the vertex Gb of gaze area G. The lower area Id and the right side area Ir are separated by line segment L4 connecting the vanishing point V and the vertex Gc of gaze area G.
[0077] Note that the inattentive driving detection area is not limited to the division shown in Fig. 16. For example, it may be as shown in Fig. 17. In Fig. 17, the inattentive driving detection area is divided by line segments extending each side of the gaze area G. The method of Fig. 17 simplifies the shape, thereby reducing the amount of processing required to divide the inattentive driving detection area.
[0078] Next, the determination of the inattentive tendency by the inattentive tendency determination unit 54 will be described. If the peak position detected by the visual saliency peak detection unit 53 is continuously outside the gaze area G for a predetermined period of time or more, it is determined that there is an inattentive tendency. Here, the predetermined period of time can be, for example, 2 seconds, but may be changed as appropriate. In other words, the inattentive tendency determination unit 54 determines whether the peak position has continuously been outside the range to be gazed at for a predetermined period of time or more. Therefore, the inattentive tendency calculation unit 5 functions as an inattentive tendency acquisition unit that acquires information indicating the inattentive tendency based on the visual saliency map (visual saliency distribution information).
[0079] Furthermore, the inattentiveness tendency determination unit 54 may determine that there is a tendency for the driver to look inattentively due to fixed objects when the inattentiveness detection area is the upper area Iu or the lower area Id. This is because, in an image captured of the area ahead from the vehicle, fixed objects such as buildings, traffic signals, signs, and streetlights are generally reflected in the upper area Iu, and road markings such as road signs are generally reflected in the lower area Id. On the other hand, moving objects other than the vehicle itself, such as vehicles traveling in other driving lanes, may be reflected in the left side area Il and the right side area Ir, making it difficult to determine the object of inattentiveness (whether it is a fixed object or a moving object) depending on the area.
[0080] If the peak position is in the left side area Il or the right side area Ir, it is not possible to determine whether the object being looked at is a fixed object or a moving object based on the area alone, so object recognition is used to make the determination. Object recognition (also called object detection) can be performed using any known algorithm, and the specific method is not particularly limited.
[0081] Furthermore, in addition to object recognition, relative speed may also be used to determine whether an object is fixed or moving. This involves calculating the relative speed between the vehicle speed and the frame-to-frame movement speed of the inattentive object, and then using this relative speed to determine whether the inattentive object is fixed or moving. Here, the frame-to-frame movement speed of the inattentive object can be calculated by calculating the frame-to-frame movement speed of the peak position. If the calculated relative speed is equal to or greater than a predetermined threshold, the object can be determined to be fixed in a certain position.
[0082] Next, the operation of the inattentive driving tendency calculation unit 5 will be described with reference to the flowchart of FIG.
[0083] First, the image input unit 2 acquires a driving image (step S104), and the visual saliency calculation unit 3 performs visual saliency image processing (acquisition of a visual saliency map) (step S105). Then, the visual saliency peak detection unit 53 acquires (detects) a peak position based on the visual saliency map acquired by the visual saliency calculation unit 3 in step S105 (step S106).
[0084] Next, the inattentive driving tendency determination unit 54 sets a gaze area G and compares the gaze area G with the peak position acquired by the visual saliency peak detection unit 53 (step S107). If the comparison result shows that the peak position is outside the gaze area G (step S107; outside the gaze area), the inattentive driving tendency determination unit 54 determines whether the dwell timer has started or is stopped (step S108). The dwell timer is a timer that measures the time that the peak position dwells outside the gaze area G. Note that the gaze area G may be set when the image is acquired in step S104.
[0085] If the dwell timer is stopped (step S108; stopped), the inattentive driving tendency determination unit 54 starts the dwell timer (step S109). On the other hand, if the dwell timer has already started (step S108; after starting), the inattentive driving tendency determination unit 54 compares the dwell timer threshold value (step S110). The dwell timer threshold value is a threshold value for the time that the peak position remains outside the gaze area G, and is set to 2 seconds, as described above.
[0086] If the dwell timer exceeds the threshold value (step S110; exceeded threshold value), the inattentive driving tendency determination unit 54 determines that there is an inattentive driving tendency (step S111).
[0087] On the other hand, if the residence timer does not exceed the threshold value (step S110; threshold not exceeded), the inattentive driving tendency determination unit 54 does nothing and returns to step S104.
[0088] Furthermore, if the comparison result in step S107 shows that the peak position is within the gaze area G (step S107; within the gaze area), the inattentive driving tendency determination unit 54 stops the dwell timer (step S112).
[0089] The monotonic trend calculation unit 6 determines whether the image input to the image input unit 2 has a monotonic trend based on the visual saliency map acquired by the visual saliency calculation unit 3. In this embodiment, various statistics are calculated from the visual saliency map, and whether the image has a monotonic trend is determined based on the statistics. That is, the monotonic trend calculation unit 6 determines whether the image has a monotonic trend using the statistics calculated based on the visual saliency map (visual saliency distribution information). Thus, the monotonic trend calculation unit 6 functions as a monotonicity acquisition unit that acquires information indicating that the image has a monotonic trend based on the visual saliency map (visual saliency distribution information).
[0090] FIG. 19 shows a flowchart of the operation of the monotonic tendency calculation unit 6. First, the standard deviation of the luminance of each pixel in the image (for example, FIG. 4(b)) that constitutes the visual saliency map is calculated (step S51). In this step, the average luminance value of each image in the image that constitutes the visual saliency map is calculated. If the image that constitutes the visual saliency map is H pixels by V pixels, and the luminance value at any coordinate (k, m) is V, then VC(k,m) Then, the average value is calculated using the following formula (2).
number
[0091] The standard deviation of the brightness of each image in the image that constitutes the visual saliency map is calculated from the average value calculated by equation (2). The standard deviation SDEV is calculated by the following equation (3).
number
[0092] It is determined whether there are multiple output results for the standard deviation calculated in step S51 (step S52). In this step, the image input from the image input unit 2 is a moving image, a visual saliency map is acquired for each frame, and it is determined in step S51 whether standard deviations for multiple frames have been calculated.
[0093] If there are multiple output results (step S52; Yes), the gaze movement amount is calculated (step S53). In this embodiment, the gaze movement amount is calculated from the coordinate distance of the maximum (highest) luminance value in the visual saliency map of each of the previous and next frames in terms of time. The gaze movement amount VSA is calculated using the following equation (4), where the coordinate of the maximum luminance value in the previous frame is (x1, y1) and the coordinate of the maximum luminance value in the next frame is (x2, y2).
number
[0094] Then, it is determined whether there is a monotonic tendency based on the standard deviation calculated in step S51 and the gaze movement amount calculated in S53 (step S54). In this step, if step S52 is No, a threshold value is set for the standard deviation calculated in step S51, and the standard deviation is compared with the threshold value to determine whether there is a monotonic tendency. On the other hand, if step S52 is Yes, a threshold value is set for the gaze movement amount calculated in step S53, and the standard deviation is compared with the threshold value to determine whether there is a monotonic tendency.
[0095] That is, the monotonic trend calculation unit 6 functions as a standard deviation calculation unit that calculates the standard deviation of the luminance of each pixel in the image obtained as the visual saliency map (visual saliency distribution information), and functions as a gaze movement amount calculation unit that calculates the gaze movement amount between frames based on the images obtained in time series as the visual saliency map (visual saliency distribution information).
[0096] It should be noted that the above-described method may fail to detect monotonous trends, especially when there are multiple output results (moving images). Therefore, a mode that can also handle such cases will be described with reference to Figures 20 and 21. A flowchart of the operation is shown in Figure 20.
[0097] In the flowchart of Fig. 20, steps S51 and S53 are the same as those in Fig. 19. In this embodiment, since autocorrelation is used as will be described later, the target image is a moving image, and step S52 is therefore omitted. In step S54A, the determination content is the same as step S54. In this embodiment, step S54A is performed as a primary determination of monotonic tendency.
[0098] Next, if the result of the determination in step S54A is that there is a monotonic trend (step S55; Yes), the result of the determination is output to the outside, as in Fig. 19. On the other hand, if the result of the determination in step S54A is that there is not a monotonic trend (step S55; No), an autocorrelation calculation is performed (step S56).
[0099] In this embodiment, the autocorrelation is calculated using the standard deviation (average luminance value) and the amount of line-of-sight movement calculated in steps S51 and S53. (k) Let E be the expected value, μ be the mean of X, and σ be the variance of X. 2 It is known that the autocorrelation value can be calculated using the following equation (5), where k is the lag. In this embodiment, the calculation of equation (5) is performed while changing k within a predetermined range, and the largest calculated value is taken as the autocorrelation value.
number
[0100] Then, whether there is a monotonic trend is determined based on the calculated autocorrelation value (step S57). To determine whether there is a monotonic trend, a threshold value is set for the autocorrelation value as in FIG. 19, and the determination can be made by comparing the threshold value. For example, if the autocorrelation value at k=k1 is greater than the threshold value, it means that a similar scene is repeated every k1. If a monotonic trend is determined, the image is classified as an image with a monotonic trend. By calculating such autocorrelation values, it becomes possible to determine whether a road has a monotonic trend due to periodically arranged objects such as streetlights installed regularly at equal intervals.
[0101] Figure 21 shows an example of the results of autocorrelation calculations. Figure 21 is a correlogram of the average luminance value of a visual saliency map for a driving video. The vertical axis of Figure 21 represents the correlation function (autocorrelation value), and the horizontal axis represents the lag. In Figure 21, the shaded area represents a 95% confidence interval (significance level αs = 0.05). If the null hypothesis is "there is no periodicity at lag k" and the alternative hypothesis is "there is periodicity at lag k," the data within this shaded area is determined to have no periodicity because the null hypothesis cannot be rejected, and anything beyond the shaded area is determined to have periodicity, regardless of whether it is positive or negative.
[0102] Figure 21(a) is a video of driving through a tunnel, and is an example of periodicity. Figure 21(a) shows that periodicity can be seen at the 10th and 17th points. In the case of a tunnel, lights inside the tunnel are placed at regular intervals, so it is possible to determine a monotonous trend caused by the lights, etc. On the other hand, Figure 21(b) is a video of driving on an ordinary road, and is an example of no periodicity. Figure 21(b) shows that most of the data falls within the confidence interval.
[0103] By operating it as shown in the flowchart in Figure 20, it is possible to first determine whether something is monotonic based on the mean, standard deviation, and amount of eye movement, and then make a secondary determination from the perspective of periodicity from among the results that are not determined.
[0104] The visual load tendency calculation unit 8 estimates the visual load based on the visual saliency map output by the visual saliency calculation unit 3. The visual load estimated by the visual load tendency calculation unit 8 may be, for example, a scalar quantity or a vector quantity. Alternatively, it may be a single piece of data or multiple pieces of time-series data. As shown in FIG. 22 , the visual load tendency calculation unit 8 includes a gaze point estimation means 81, a gaze point movement amount calculation means 82, a basis component decomposition means 83, a basis component selection means 84, a power calculation means 85, and a power combination means 86.
[0105] The gaze point estimation means 81 estimates gaze point information from the time-series visual saliency map output by the visual saliency calculation unit 3. The definition of the gaze point information is not particularly limited, but it can be, for example, the position (coordinates) where the saliency value is maximum. That is, the gaze point estimation means 81 estimates the estimated gaze point as the position on the image where the visual saliency value is maximum in the visual saliency map (visual saliency distribution information).
[0106] The gaze point movement amount calculation means 82 calculates the time-series gaze point movement amount from the time-series gaze point information estimated by the gaze point estimation means 81. The gaze point movement amount calculated by the gaze point movement amount calculation means 82 is also time-series data. There are no particular limitations on the calculation method, but it can be, for example, the Euclidean distance between gaze point coordinates that are in a previous or next relationship in the time series. That is, the gaze point movement amount calculation means 82 calculates the movement amount of the gaze point (estimated gaze point) based on the generated visual saliency map (visual saliency distribution information).
[0107] The basis component decomposition means 83 decomposes the time series change in the amount of gaze point movement calculated by the gaze point movement amount calculation means 82 into one or more basis component groups. Each basis component decomposed by the basis component decomposition means 83 also becomes time series data corresponding to the amount of movement. There are no particular limitations on the decomposition method, but for example, empirical mode decomposition (EMD) is desirable, and in this case the basis components become intrinsic mode functions (IMF). Other methods include The Rober transform family (basis components are sinusoidal) may also be used.
[0108] The basis component selection means 84 selects one or more basis components from the group of basis components decomposed by the basis component decomposition means 83. Note that the selection method is not limited, and various methods can be used.
[0109] The power calculation means 85 calculates the power for each of the basis components selected by the basis component selection means 84. Here, the power also becomes time-series data corresponding to the basis components. There are no particular limitations on the calculation method, but when empirical mode decomposition is used as the basis component decomposition means, it is desirable to calculate the amplitude component and phase component using a Hilbert transform on the basis components (intrinsic mode functions), and use the amplitude component as the power.
[0110] The power combining means 86 combines the multiple power components calculated by the power calculation means 85 to calculate the visual load. Here, the visual load also becomes time-series data corresponding to the power. There are no particular limitations on the combining means, but it may be, for example, a simple addition of multiple power components. That is, the visual load is estimated based on the time transition of the movement amount of the gaze point (estimated gaze point) calculated by the means from the basis component decomposition means 83 to the power combining means 86.
[0111] Fig. 23 shows example waveforms of the gaze point movement amount output by the gaze point movement amount calculation means 82, a group of basis components obtained by decomposing the gaze point movement amount by the basis component decomposition means 83, and a visual load amount output by combining the power calculated from the group of basis components by the power calculation means 85 and the power combining means 86. The top row in Fig. 23 shows the gaze point movement amount. Rows 2 to 16 in Fig. 23 show the group of basis components. And row 17 (bottom row) in Fig. 23 shows the visual load amount.
[0112] In FIG. 23, it is estimated (determined) that the visual load is large in the portion where the amplitude of the visual load shown in the bottom row is large.
[0113] The visual attention concentration level calculation unit 4, the inattentiveness tendency calculation unit 5, the monotonic tendency calculation unit 6, and the visual load tendency calculation unit 8 described above function as a visual feature extraction unit that acquires the visual tendency of the moving environment of a moving object based on the visual saliency map (visual saliency distribution information).
[0114] The situation output unit 7 estimates and outputs the situation of the image input from the image input unit 2 based on the calculation results (determination results) of the visual attention concentration calculation unit 4, the inattentiveness tendency calculation unit 5, the monotony tendency calculation unit 6, and the visual load tendency calculation unit 8. In this embodiment, the estimation result is output as text information (text), but it may also be classified by labels or given a score by scoring.
[0115] An example of sentence generation based on the calculation results (determination results) of the visual attention concentration level calculation unit 4, the inattentiveness tendency calculation unit 5, the monotonic tendency calculation unit 6, and the visual load tendency calculation unit 8 in the situation output unit 7 will be described with reference to Fig. 24. Fig. 24 shows an example of sentence generation based on the calculation results of the visual attention concentration level calculation unit 4, the inattentiveness tendency calculation unit 5, the monotonic tendency calculation unit 6, and the visual load tendency calculation unit 8.
[0116] The block indicated by A in Fig. 24 is an example of a sentence generated based on the calculation result of the monotonic tendency calculation unit 6. For example, if the calculation result of the monotonic tendency calculation unit 6 indicates a monotonic tendency, a sentence such as "driving on a monotonic road" is generated in block A. On the other hand, if the calculation result of the monotonic tendency calculation unit 6 does not indicate a monotonic tendency, a sentence such as "driving on a complex road" is generated in block A.
[0117] For block A, a sentence can be generated based on the visual load calculated by the visual load tendency calculation unit 8. For example, if the visual load is large, the situation can be said to be complex, so a sentence such as "driving on a complex road" can be generated, and if the visual load is small, the situation can be said to be monotonous, so a sentence such as "driving on a monotonous road" can be generated.
[0118] The block indicated by B in FIG. 24 is an example of a sentence generated based on the calculation results of the inattentive tendency calculation unit 5. For example, if the calculation results of the inattentive tendency calculation unit 5 indicate that the moving object on the right is an inattentive object and that there is an inattentive tendency, then in block B a sentence such as "As a result of being distracted by a traffic participant on the right" is generated. In block B, classifications such as forward, right, left, and above can be determined using the areas shown in FIG. 16. Furthermore, whether it is a traffic participant or scenery outside the road can be determined using area and object recognition, as explained in the operation of the inattentive tendency calculation unit 5. In this way, information indicating the inattentive tendency includes not only the determination result of the inattentive tendency, but also information such as the inattentive object and its position.
[0119] The block indicated by C in FIG. 24 is an example of a sentence generated based on the calculation result of the visual attention concentration level calculation unit 4. For example, if the visual attention concentration level Ps changes significantly over time, a sentence such as "I noticed a sudden change in the road environment" is generated. In other words, if the visual attention concentration level Ps changes suddenly as shown in FIG. 13, it can be said that the concentration level increased due to some kind of change in the environment. In this way, information indicating the visual attention concentration level includes not only the visual attention concentration level Ps but also information such as its change over time.
[0120] Block D in Figure 24 is an example of a sentence that shows what happened as a result of the situation shown in blocks A to C. For example, if the driver applied the brakes too late, resulting in sudden braking, and although an accident did not occur, a safety issue may have occurred. A sentence such as "The danger was avoided, but the delay in avoiding the risk is presumed to have resulted in a near miss" is generated. In block D, whether the danger was avoided can be determined by whether the driver applied the brakes.
[0121] Whether or not sudden braking has occurred can be determined by including in the corresponding image information a detection signal (acceleration, etc.) from the vehicle behavior detection unit 12 shown in Fig. 1. For example, if the acceleration value included in the image information is equal to or greater than a predetermined value, it can be determined that sudden braking has occurred. When determining based on images alone, the determination may also be made based on changes in the moving speed of an object included in the images between frames.
[0122] Whether risk avoidance was delayed can be determined from the amount of change over time in the visual attention concentration level Ps. Also, whether it is a near miss or an accident can be determined from the magnitude of acceleration, or an image may be previously labeled as an accident or not. In this way, the detection signals of the vehicle behavior detection unit 12 can be used as an auxiliary means for documenting the above-mentioned information.
[0123] The sentences generated in blocks A to D are then combined and output. For example, they may be combined as "While driving on a monotonous road, the driver noticed a sudden change in the road environment and avoided the danger, but the delay in avoiding the risk resulted in a near miss." If both a monotonous tendency and a tendency to look away are determined, the sentences in blocks A and B are combined, but if a tendency to look away is not determined, the sentence in block B is not generated. As is clear from the above example, this embodiment is capable of generating sentences that follow the time series of moving images, rather than just describing the state of a specific still image. Therefore, factors such as why the driver suddenly braked can also be automatically analyzed and generated into sentences.
[0124] In the example of Figure 24, blocks C and part of block D are the first keyword based on information indicating the degree of visual attention concentration, block B is the second keyword based on information indicating a tendency to look away, and block A is the third keyword based on information indicating a tendency to be monotonous or a tendency to have visual load.
[0125] Furthermore, the text generation is not limited to that shown in Fig. 24, and further detail may be achieved, for example, by using the results of object recognition. An example of text generation using the results of object recognition is shown in Fig. 25. Object recognition may be performed by the inattentiveness tendency calculation unit 5 described above.
[0126] The block E in Fig. 25 is an example of a sentence generated based on traffic regulations such as traffic signs. For example, if a traffic sign ahead is not noticed, a sentence such as "I did not notice the traffic sign ahead" is generated. In block E, the traffic sign is detected by object recognition, and whether or not the traffic sign has been noticed can be determined by whether or not a high-brightness part of the visual saliency map overlaps.
[0127] The block indicated by F in Figure 25 is an example of a sentence generated based on traffic participants such as other vehicles. For example, if the driver does not notice a vehicle ahead, a sentence such as "I did not notice the traffic participants ahead" is generated. In block F, the vehicle is detected by object recognition, and whether or not the driver has noticed the vehicle can be determined by whether or not the high-brightness parts of the visual saliency map overlap.
[0128] Next, the operation (situation output method) of the situation output device 1 configured as described above will be described with reference to the flowchart in Fig. 26. This flowchart can be configured as a program executed by a computer functioning as the situation output device 1 to form a situation output program. This situation output program is not limited to being stored in a memory or the like possessed by the situation output device 1, but may also be stored in a storage medium such as a memory card or optical disk.
[0129] First, the image input unit 2 outputs the input image as image data to the visual saliency calculation unit 3 (step S11). In this step, the image data input to the image input unit 2 is decomposed into a time series of image frames, etc., and input to the visual saliency calculation unit 3. In this step, image processing such as noise removal and geometric transformation may also be performed.
[0130] Next, the visual saliency calculation unit 3 acquires a visual saliency map (step S12). The visual saliency calculation unit 3 outputs the visual saliency map in time series as shown in FIG. 4(b) by the above-described method.
[0131] Next, the visual attention concentration level calculation unit 4 acquires the visual attention concentration level Ps (step S13). The visual attention concentration level Ps is calculated and acquired by the visual attention concentration level calculation unit 4 using the method described above. Then, in parallel with step S13, the inattentive driving tendency calculation unit 5 acquires information indicating the inattentive driving tendency (step S14). The inattentive driving tendency calculation unit 5 acquires the information indicating the inattentive driving tendency by making a determination using the method described above. Then, in parallel with steps S13 and S14, the monotonic tendency calculation unit 6 acquires information indicating the monotonic tendency (step S15). The monotonic tendency calculation unit 6 acquires the information indicating the monotonic tendency by making a determination using the method described above. Furthermore, in parallel with steps S13 to S15, the visual load tendency calculation unit 8 acquires information indicating the visual load tendency (step S16). The visual load tendency calculation unit 8 acquires the information indicating the visual load tendency by making a determination using the method described above.
[0132] Next, the situation output unit 7 puts into words the situation of the image input from the image input unit 2 based on the results of steps S13 to S16 (step S17). The putting into words is performed by the method shown in FIGS.
[0133] Next, the situation output unit 7 combines and outputs the sentences generated in step S17 (step S18). The output destination may be a display device such as a display, or may be stored in a storage device as text data or the like in association with the target image, or may be transmitted to an external device.
[0134] As is clear from the above description, step S12 functions as a visual saliency distribution information acquisition step, steps S13 to S16 function as a visual feature extraction step, and steps S17 and S18 function as a situation output step.
[0135] An example of a display image constituting part of the output of the above-mentioned situation output device 1 will now be described with reference to Fig. 27. The image shown in Fig. 27 can be output as an analysis result of an image input from the image input unit 2 together with the above-mentioned text. The image shown in Fig. 27 is generated by the situation output unit 7 and displayed on a predetermined display device. The image 50 shown in Fig. 27 includes a driving image display area 51 and a visual attention concentration level display area 52.
[0136] The driving image display area 51 displays a driving image of a vehicle captured by a drive recorder, etc. The driving image display area 51 can display a vanishing point VP, a visual saliency map VM, an estimated gaze position GE, a detected object frame OF, a horizon line and depth distance estimation line HD, an estimated white line WL, and a near-miss judgment frame HF.
[0137] The vanishing point VP may be estimated from a white line estimation WL, which will be described later, or may be estimated using optical flow, etc. The visual saliency map VM displays a visual saliency map (heat map) for the image displayed in the driving image display area 51, superimposed on the image. Note that in FIG. 27, only high-brightness areas on the heat map are visible, but in reality, low-brightness areas are also superimposed on the driving image. In other words, it displays areas in the driving image that are likely to be drawn to.
[0138] In this embodiment, the gaze estimation position GE is estimated to be the position with the highest brightness on the heat map. The detected object frame OF is displayed as a frame surrounding an object detected in the driving image as a result of object detection using a well-known algorithm. In this embodiment, the object detection is performed by specifying the type of object to be detected (vehicle, person, etc.), and only objects belonging to the specified type are detected. The object detection process may be performed by the inattentiveness tendency calculation unit 5 or by another functional block not shown.
[0139] The horizontal line and depth distance estimation line HD indicate the horizontal line and depth distance in the driving image. The white line estimation WL indicates a recognized lane marking such as a white line in the driving image. The near-miss judgment frame HF is formed in a frame shape that follows the four sides of the driving image display area 51, and is displayed when the situation output unit 7 determines that a safety issue such as a near-miss has occurred. Alternatively, the near-miss judgment frame HF may be displayed as a blue frame at all times, and when it is determined that a safety issue is suspected, the display color may be changed to red or the like, or the frame may flash.
[0140] The visual attention concentration level display area 52 is provided to the right of the driving image display area 51. The visual attention concentration level display area 52 displays the visual attention concentration level Ps calculated by the vector error calculation unit 4 in the form of a bar graph. In FIG. 27, reference numeral 52a denotes a bar indicating the visual attention concentration level Ps. When the visual attention concentration level Ps indicates a large value, the bar 52a becomes taller, and when the visual attention concentration level Ps indicates a small value, the bar 52a becomes shorter. In other words, when the bar 52a is in a high position, the concentration level tends to be high (concentrated), and when the bar 52a is in a low position, the concentration level tends to be low (dispersed).
[0141] The height of the bar 52a in the visual attention concentration level display area 52 changes with the progression of image playback time because the visual saliency map is acquired frame by frame. Therefore, for example, a scene in which the height of the bar 52a increases suddenly indicates that the visual attention concentration level Ps has changed rapidly from dispersed to concentrated, and this bar 52a can also visually display the temporal change in the visual attention concentration level Ps.
[0142] According to this embodiment, the situation output device 1 acquires a visual saliency map obtained by estimating the level of visual saliency within an image captured from a moving object. Next, as a visual tendency of the moving object's environment, the visual attention concentration level calculation unit 4 acquires a visual attention concentration level Ps for the image based on the visual saliency map and an ideal line of sight previously determined for the image. The inattentive tendency calculation unit 5 determines the inattentive tendency based on the visual saliency map. The monotonic tendency calculation unit 6 determines whether the image has a monotonic tendency based on the visual saliency map. The visual load tendency calculation unit 8 determines whether the image has a high visual load based on the visual saliency map. The situation output unit 7 then outputs the situation of the acquired image based on the acquired visual attention concentration level Ps, the inattentive tendency determination result, the monotonic tendency determination result, and the visual load tendency determination result. In this manner, the situation of the image can be output based on the visual tendencies (four pieces of information) based on the visual saliency map. Therefore, it is possible to analyze the situation using only the image, and the circumstances under which a near miss or other incident occurred can be estimated without checking the image.
[0143] Furthermore, the status output unit 7 outputs the status of the image as text information. In this way, the status of the image can be output as text, etc., which can be useful for creating reports, daily reports, etc.
[0144] Furthermore, the situation output unit 7 outputs text information that combines multiple keywords based on each of the multiple pieces of information acquired based on the visual tendency. Specifically, the situation output unit 7 outputs text information that combines multiple keywords from among a first keyword based on the visual attention concentration level Ps, a second keyword based on the determination result of the tendency to look aside, and a third keyword based on the determination result of the tendency to monotony or the tendency to visual load. In this way, the situation can be constructed as a sentence that relates not only to a specific factor but also to multiple factors. This makes it easier to grasp the detailed situation.
[0145] Furthermore, the visual attention concentration level calculation unit 4 calculates the visual attention concentration level Ps based on the value of each pixel constituting the visual saliency map and the vector error between the position of each pixel and the coordinate position of the ideal gaze. In this way, a value according to the difference between a position of high visual saliency and the ideal gaze is calculated as the visual attention concentration level Ps. Therefore, for example, the value of the visual attention concentration level Ps can be changed according to the distance between the position of high visual saliency and the ideal gaze.
[0146] The inattentive tendency calculation unit 5 also includes a visual saliency peak detection unit 53 that detects at least one peak position in the visual saliency map in a time series, and a inattentive tendency determination unit 54 that sets a range in the image that the driver of the moving object should focus their attention on. If the peak position is out of the range that the driver should focus their attention on for a predetermined period of time or more, the inattentive tendency calculation unit 5 determines that the driver is inattentive. This visual saliency map indicates the statistical likelihood of human gaze focusing. Therefore, the peak of the visual saliency map indicates the position that is statistically most likely to attract human gaze. Therefore, by using the visual saliency map, it is possible to detect the inattentive tendency with a simple configuration without measuring the actual driver's gaze.
[0147] Furthermore, the monotonicity calculation unit 6 determines whether there is a monotonicity trend using statistics calculated based on the visual saliency map. In this way, it is possible to determine whether there is a monotonicity trend based on the positions that are likely to be gazed at by humans from the captured image. Since the determination is based on the positions that are likely to be gazed at by humans (drivers), it is possible to determine a trend that is close to what the driver perceives as monotonic, and thus it is possible to perform the determination with higher accuracy.
[0148] Furthermore, the visual load tendency calculation unit 8 calculates the amount of movement of the gaze point based on the generated visual saliency map. Then, the visual load is estimated based on the temporal change in the amount of movement of the calculated estimated gaze point. In this way, it is possible to estimate a position where the gaze shift will be large based on the temporal change in the amount of movement of the gaze point. Therefore, it is possible to automatically extract parts that cause visual load.
[0149] The visual saliency calculation unit 3 includes an input unit 310 that converts an image into intermediate data that can be mapped, a nonlinear mapping unit 320 that converts the intermediate data into mapped data, and an output unit 330 that generates saliency estimation information indicating a saliency distribution based on the mapped data. The nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. This allows visual saliency to be estimated with low computational cost. Furthermore, the visual saliency estimated in this manner reflects the contextual attention state.
[0150] Although the above-described embodiment includes all of the visual attention concentration calculation unit 4, the inattentiveness tendency calculation unit 5, the monotony tendency calculation unit 6, and the visual load calculation unit 8, it is sufficient to include at least one of these. In other words, it is not necessary to generate related sentences for all of them.
[0151] Here, a modified example of the situation output device shown in Fig. 1 will be described with reference to Fig. 28. The situation output device 1A shown in Fig. 28 has a situation accumulation unit 9 and a situation analysis unit 10 added to the configuration of Fig. 1.
[0152] The situation accumulation unit 9 accumulates the situation of the image output by the situation output unit 7. It is desirable that the situation accumulation unit 9 accumulates the situation for each driver who was driving the moving object when the image input from the image input unit 2 was captured. Therefore, it is desirable that information that can identify the driver is also added to the image input from the image input unit 2.
[0153] The situation analysis unit 10 analyzes the situations accumulated in the situation accumulation unit 9, for example, analyzes the driving tendency of the driver, and outputs the analysis result. For example, if a large number of images and texts related to near misses related to inattentive driving are accumulated, it can be analyzed that the driver has a tendency to inattentive driving.
[0154] Furthermore, the situation analysis unit 10 may compare drivers based on the analysis results. For example, by comparing a driver who tends to drive relatively safely with a driver who has many near misses, it is possible to provide driving guidance to prevent near misses. It is also possible to analyze trends in organizational units, such as an entire company, from the trends of multiple drivers.
[0155] Furthermore, the present invention is not limited to the above-described embodiments. In other words, a person skilled in the art can implement various modifications in accordance with conventional knowledge without departing from the gist of the present invention. As long as such modifications still include the situation output device of the present invention, they are of course included in the scope of the present invention. [Explanation of symbols]
[0156] 1. Situation output device 2 Image input section 3. Visual saliency calculation unit (visual saliency distribution information acquisition unit) 4. Visual attention concentration calculation unit (visual feature extraction unit, visual attention concentration acquisition unit) 5. Inattentiveness tendency calculation unit (visual feature extraction unit, inattentiveness acquisition unit) 6. Monotonicity trend calculation unit (visual feature extraction unit, monotonicity acquisition unit) 7 Status output section 8. Visual load tendency calculation unit (visual feature extraction unit, visual load acquisition unit)
Claims
[Claim 1] a visual saliency distribution information acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency in an image captured from a moving object; a visual feature extraction unit that acquires a visual tendency of a moving environment of the moving object based on the visual saliency distribution information; a situation output unit that outputs a situation of the image based on the visual tendency; A situation output device comprising:
Citation Information
Patent Citations
Weighting matrix learning device, line-of-sight direction prediction system, warning system and weighting matrix learning method
JP2016130959A
Cause analysis device and cause analysis method
JP2016071492A