Determination device

The determination device uses visual saliency and gaze position analysis to detect near misses and accidents by evaluating temporal changes in attention concentration, addressing the limitations of existing systems that rely on acceleration and braking.

JP2025131914AInactive Publication Date: 2025-09-09PIONEER IP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025107840
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing systems fail to accurately detect near misses that do not involve sudden braking, such as vehicle intrusions from the side or distractions, as they rely solely on acceleration and accelerator position changes.

Method used

A determination device that utilizes visual saliency distribution information and gaze position analysis to estimate visual attention concentration, determining safety issues like accidents or near misses based on temporal changes in attention levels.

Benefits of technology

Accurately identifies safety problems caused by psychological stress or distractions through image analysis, reducing manual work and enhancing detection of near misses beyond sudden braking scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131914000001_ABST
    Figure 2025131914000001_ABST
Patent Text Reader

Abstract

To determine the possibility of occurrence of a safety problem such as an accident or a near-miss caused by a psychological burden.SOLUTION: In a determination device 1, a visual saliency calculation part 3, on the basis of an image of the outside photographed from a moving body, acquires a visual saliency map obtained by estimating the high / low level of visual saliency in the image, and a visual line coordinate setting part 4 sets a coordinate of an ideal visual line. Then, a vector error calculation part 5 calculates a visual attention concentration degree Ps in the image on the basis of the visual saliency map and the ideal visual line. A determination part 6 determines the possibility of occurrence of a safety problem such as a near-miss during running of the moving body on the basis of a temporal change amount in the visual attention concentration degree Ps.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a determination device that performs a predetermined determination process based on an image of the outside taken from a moving body. [Background technology]

[0002] For example, a drive recorder detects the occurrence of an accident or near miss based on the acceleration of the vehicle, and records images before and after the occurrence.

[0003] Patent Document 1 describes a drive recorder device that can accurately detect events such as accidents and near misses and identify the cause of the events. In the invention described in Patent Document 1, when a change in acceleration occurs on the acceleration side, the control unit 21 determines whether the change in accelerator pedal position is equal to or less than a certain value that is determined to be due to the driver's decision. If the change in accelerator pedal position is equal to or less than the certain value, the control unit 21 determines that the event is due to external energy. On the other hand, if the change in accelerator pedal position exceeds the certain value, the control unit 21 determines that the event is due to the driver's decision. In this case, the control unit 21 records the event information, data collected by various sensors, and video information in the recording unit 26 as a near miss case. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-164131 Summary of the Invention [Problem to be solved by the invention]

[0005] If an accident or near miss is determined based solely on acceleration, it will not be possible to detect near misses that do not involve sudden braking, such as a sudden vehicle intrusion from the side, or near misses that occur while the driver is distracted.

[0006] In the invention described in Patent Document 1, the degree of accelerator opening is also used as a criterion for judgment, so it is not possible to detect near misses that do not involve sudden braking, such as near misses caused by a vehicle suddenly entering from the side, or near misses that occur when the driver is not paying attention.

[0007] One example of a problem that the present invention aims to solve is determining whether a safety problem, such as an accident or near miss, has occurred due to psychological stress. [Means for solving the problem]

[0008] In order to solve the above problem, the invention described in claim 1 is characterized by comprising an acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency in an image based on an image of the outside taken from a moving body; a gaze position setting unit that sets a reference gaze position in the image according to predetermined rules; a visual attention concentration calculation unit that calculates the degree of visual attention concentration in the image based on the visual saliency distribution information and the gaze position; and a judgment unit that judges that a safety issue is suspected to have occurred while the moving body is traveling based on the amount of change over time in the degree of visual attention concentration.

[0009] The invention described in claim 5 is a judgment method executed by a judgment device that performs a predetermined judgment process based on an image of the outside taken from a moving body, and is characterized by including: an acquisition step of acquiring visual saliency distribution information obtained by estimating the level of visual saliency in the image based on the image; a gaze position setting step of setting a reference gaze position in the image in accordance with a predetermined rule; a visual attention concentration calculation step of calculating a visual attention concentration level in the image based on the visual saliency distribution information and the gaze position; and a judgment step of determining that a safety problem is suspected to have occurred while the moving body is traveling based on a temporal change in the visual attention concentration level.

[0010] The invention as set forth in claim 6 is characterized in that the determination method as set forth in claim 5 is executed by a computer.

[0011] The invention as set forth in claim 7 is characterized in that the determination program as set forth in claim 6 is stored. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a schematic configuration diagram of a system including a determination device according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a functional configuration diagram of the determination device shown in FIG. [Figure 3] 2 is a block diagram illustrating the configuration of a visual saliency calculation unit shown in FIG. 1. FIG. [Figure 4] 1A is a diagram illustrating an example of an image input to a determination device, and FIG. 1B is a diagram illustrating an example of a visual saliency map estimated for FIG. 1A. [Figure 5] 2 is a flowchart illustrating a processing method of the visual saliency calculation unit shown in FIG. 1; [Figure 6] FIG. 2 is a diagram illustrating in detail an example of the configuration of a nonlinear mapping unit. [Figure 7] FIG. 2 is a diagram illustrating the configuration of an intermediate layer. [Figure 8] 10(a) and 10(b) are diagrams illustrating examples of convolution processing performed by a filter. [Figure 9] (a) is a diagram for explaining the processing of the first pooling unit, (b) is a diagram for explaining the processing of the second pooling unit, and (c) is a diagram for explaining the processing of the unpooling unit. [Figure 10] FIG. 10 is an explanatory diagram of a vector error. [Figure 11] 2 is an example of an image input to the image input unit shown in FIG. 1 and a visual saliency map obtained from the image. [Figure 12] 10 is a graph showing an example of temporal changes in visual attention concentration level. [Figure 13] 2 is a flowchart of the operation of the determination device shown in FIG. 1. [Figure 14] 10 is an example of an output screen of a determination device. DETAILED DESCRIPTION OF THE INVENTION

[0013] A determination device according to one embodiment of the present invention will be described below. In the determination device according to one embodiment of the present invention, an acquisition unit acquires visual saliency distribution information obtained by estimating the level of visual saliency in an image captured from a moving object, and a gaze position setting unit sets a reference gaze position in the image according to a predetermined rule. A visual attention concentration calculation unit calculates the visual attention concentration level in the image based on the visual saliency distribution information and the gaze position. A determination unit determines that a safety issue is suspected to have occurred while the moving object is traveling based on a temporal change in the visual attention concentration level. By using the visual saliency distribution information, it is possible to determine that a safety issue is suspected to have occurred based on a temporal change in a contextual attention state in which the gaze tends to unconsciously focus on objects such as signs and pedestrians included in the image. Therefore, it is possible to determine the suspicion of a safety issue, such as an accident or near miss, caused by psychological stress, based solely on the image.

[0014] The visual attention concentration level calculation unit may also calculate the visual attention concentration level based on the value of each pixel constituting the visual saliency distribution information and the vector error between the position of each pixel and the coordinate position of the reference gaze position. In this way, a value corresponding to the difference between a position of high visual saliency and the reference gaze position is calculated as the visual attention concentration level. Therefore, for example, the value of the visual attention concentration level can be changed depending on the distance between the position of high visual saliency and the reference gaze position.

[0015] The device may also include an output unit that outputs information related to the determination result of the determination unit, thereby enabling the determination result and information based on the determination result to be displayed or communicated to the outside.

[0016] The image may also be obtained by detecting sudden braking acceleration of the moving body using a sensor equipped in the moving body. In this way, for example, in a drive recorder or the like, an image extracted based on the detection of sudden braking can be further judged as a near miss, etc., thereby reducing the amount of manual work required.

[0017] The acquisition unit may also include an input unit that converts the image into intermediate data that can be mapped, a nonlinear mapping unit that converts the intermediate data into mapped data, and an output unit that generates saliency estimation information indicating a saliency distribution based on the mapped data, and the nonlinear mapping unit may include a feature extraction unit that extracts features from the intermediate data and an upsampling unit that upsamples the data generated by the feature extraction unit. This allows visual saliency to be estimated with low computational cost. Furthermore, the visual saliency estimated in this manner reflects a contextual attention state.

[0018] In addition, an information processing method according to one embodiment of the present invention includes an acquisition step of acquiring visual saliency distribution information obtained by estimating the level of visual saliency in an image captured from a moving object, and a gaze position setting step of setting a reference gaze position in the image according to a predetermined rule. Then, a visual attention concentration calculation step of calculating the visual attention concentration in the image based on the visual saliency distribution information and the gaze position. Then, a determination step of determining whether a safety issue has occurred while the moving object is traveling based on a temporal change in the visual attention concentration. By using the visual saliency distribution information, it is possible to determine whether a safety issue has occurred based on a temporal change in the contextual attention state, in which the gaze tends to unconsciously focus on objects such as signs and pedestrians included in the image. Therefore, it is possible to determine whether a safety issue, such as an accident or near miss caused by psychological stress, has occurred based on the image alone.

[0019] Furthermore, the above-described information processing method is executed by a computer. In this way, it is possible to determine whether a safety problem has occurred using the computer. Therefore, it is possible to accurately determine whether a safety problem such as an accident or a near miss has occurred using only an image.

[0020] The information processing program may be stored in a computer-readable storage medium, which allows the program to be distributed as a standalone program rather than being incorporated into a device, and allows for easy version upgrades. [Example]

[0021] A determination device according to an embodiment of the present invention will be described with reference to Figures 1 to 14. The determination device according to this embodiment is not limited to being installed in a mobile object such as an automobile, but may also be configured as a server device or the like installed in a business establishment or the like (see Figure 1). In other words, analysis does not need to be performed in real time, and analysis may be performed after driving, etc.

[0022] FIG. 1 illustrates an example in which a determination device is configured as a server device. In FIG. 1, when a vehicle behavior detection unit 11, such as an acceleration sensor, detects a large acceleration due to sudden braking, sudden acceleration, or other impact in a drive recorder 10 mounted on a vehicle V, images (video images) of a predetermined period before and after the event are transmitted to the determination device 1 via a network N, such as the Internet. The vehicle behavior detection unit 11 shown in FIG. 2 is not limited to an acceleration sensor, but may also be an anti-lock braking system (ABS) or a skid prevention device mounted on the vehicle V. The activation of such a device (sensor) may trigger the transmission or storage of images of a predetermined period before and after the event. By processing images in which such sudden braking or other events are detected, as will be described later, it is possible to narrow down the images to a certain extent, thereby reducing the processing time (processing volume). The transmission of images is not limited to the form of communication as shown in FIG. 1, and images may be read from a recording medium, such as a hard disk drive or memory card, connected to the server device 1. The determination device is not limited to a server device, but may be an in-vehicle device incorporating a determination unit or a PC terminal at home or at work, or may be configured to distribute processing between these devices and a server.

[0023] As shown in FIG. 2, the determination device 1 includes an image input unit 2, a visual saliency calculation unit 3, a gaze coordinate setting unit 4, a vector error calculation unit 5, and a determination unit 6.

[0024] The image input unit 2 receives images (e.g., moving images) captured by a camera such as the drive recorder described above, and outputs the images as image data. The input moving images are output as image data broken down into time series, such as frames. Although still images may be input as images to the image input unit 2, it is preferable to input them as an image group consisting of a plurality of still images in time series.

[0025] The images input to the image input unit 2 include, for example, images captured in the direction of travel of the vehicle. In other words, images are taken of the outside world continuously from a moving body. These images may be so-called panoramic images or images acquired using multiple cameras, and may include images that include angles other than the direction of travel, such as 180° or 360° in the horizontal direction. Furthermore, the images input to the image input unit 2 are not limited to images captured by a camera, and may also be images read from a recording medium such as a hard disk drive or memory card, as described above.

[0026] The visual saliency calculation unit 3 receives image data from the image input unit 2 and outputs a visual saliency map as visual saliency estimation information (described later). That is, the visual saliency calculation unit 3 functions as an acquisition unit that acquires a visual saliency map (visual saliency distribution information) obtained by estimating the level of visual saliency based on an image of the outside captured from a moving object.

[0027] FIG. 3 is a block diagram illustrating the configuration of the visual saliency calculation unit 3. The visual saliency calculation unit 3 according to this embodiment includes an input unit 310, a nonlinear mapping unit 320, an output unit 330, and a storage unit 390. The input unit 310 converts an image into intermediate data that can be subjected to mapping processing. The nonlinear mapping unit 320 converts the intermediate data into mapped data. The output unit 330 generates saliency estimation information indicating a saliency distribution based on the mapped data. The nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. The storage unit 390 stores image data input from the image input unit 2, filter coefficients (described later), and the like. These are described in detail below.

[0028] FIG. 4(a) is a diagram illustrating an example of an image input to the visual saliency calculation unit 3, and FIG. 4(b) is a diagram illustrating an example of an image showing a visual saliency distribution estimated for FIG. 4(a). The visual saliency calculation unit 3 according to this embodiment is a device that estimates the visual saliency of each part in an image. Visual saliency means, for example, how easily something stands out or how easily it attracts attention. Specifically, visual saliency is expressed as a probability or the like. Here, the magnitude of the probability corresponds to, for example, the probability that a person viewing the image will direct their gaze to that position.

[0029] 4(a) and 4(b) correspond to each other in position. In FIG. 4(a), the higher the visual saliency, the higher the brightness displayed in FIG. 4(b). The image showing the visual saliency distribution as shown in FIG. 4(b) is an example of a visual saliency map output by the output unit 330. In this example, visual saliency is visualized using brightness values ​​of 256 levels. An example of the visual saliency map output by the output unit 330 will be described in detail later.

[0030] FIG. 5 is a flowchart illustrating the operation of the visual saliency calculation unit 3 according to this embodiment. The flowchart shown in FIG. 5 is a part of a determination method executed by a computer, and includes an input step S110, a nonlinear mapping step S120, and an output step S130. In the input step S110, an image is converted into intermediate data that can be mapped. In the nonlinear mapping step S120, the intermediate data is converted into mapped data. In the output step S130, visual saliency estimation information (visual saliency distribution information) indicating a saliency distribution is generated based on the mapped data. Here, the nonlinear mapping step S120 includes a feature extraction step S121 that extracts features from the intermediate data, and an upsampling step S122 that upsamples the data generated in the feature extraction step S121.

[0031] Returning to FIG. 3, each component of the visual saliency calculation unit 3 will be described. In input step S110, the input unit 310 acquires an image and converts it into intermediate data. The input unit 310 acquires image data from the image input unit 2. The input unit 310 then converts the acquired image into intermediate data. The intermediate data is not particularly limited as long as it is data that can be accepted by the nonlinear mapping unit 320, and is, for example, a high-dimensional tensor. Furthermore, the intermediate data is, for example, data in which the brightness of the acquired image is normalized, or data in which each pixel of the acquired image is converted into a brightness gradient. In input step S110, the input unit 310 may further perform noise removal, resolution conversion, etc. on the image.

[0032] In the nonlinear mapping step S120, the nonlinear mapping unit 320 acquires intermediate data from the input unit 310. The nonlinear mapping unit 320 then converts the intermediate data into mapping data. Here, the mapping data is, for example, a high-dimensional tensor. The mapping process performed on the intermediate data by the nonlinear mapping unit 320 is, for example, a mapping process that can be controlled by parameters, and is preferably a process using a function, a functional, or a neural network.

[0033] Fig. 6 is a diagram illustrating a detailed configuration of the nonlinear mapping unit 320, and Fig. 7 is a diagram illustrating a configuration of the hidden layer 323. As described above, the nonlinear mapping unit 320 includes a feature extraction unit 321 and an upsampling unit 322. The feature extraction unit 321 performs the feature extraction step S121, and the upsampling unit 322 performs the upsampling step S122. In the example shown in this figure, at least one of the feature extraction unit 321 and the upsampling unit 322 is configured to include a neural network including a plurality of hidden layers 323. In the neural network, a plurality of hidden layers 323 are connected.

[0034] In particular, the neural network is preferably a convolutional neural network. Specifically, each of the multiple hidden layers 323 includes one or more convolutional layers 324. In the convolutional layers 324, input data is convolved by multiple filters 325, and activation processing is performed on the outputs of the multiple filters 325.

[0035] 6, feature extraction unit 321 is configured to include a neural network including a plurality of hidden layers 323, and a first pooling unit 326 is provided between the plurality of hidden layers 323. Furthermore, upsampling unit 322 is configured to include a neural network including a plurality of hidden layers 323, and an unpooling unit 328 is provided between the plurality of hidden layers 323. Furthermore, feature extraction unit 321 and upsampling unit 322 are connected to each other via a second pooling unit 327 that performs overlap pooling.

[0036] In the example shown in this figure, each intermediate layer 323 is made up of two or more convolutional layers 324. However, at least some of the intermediate layers 323 may be made up of only one convolutional layer 324. Adjacent intermediate layers 323 are separated by any of a first pooling unit 326, a second pooling unit 327, and an unpooling unit 328. Here, when an intermediate layer 323 includes two or more convolutional layers 324, it is preferable that the number of filters 325 in those convolutional layers 324 be equal to each other.

[0037] In this figure, an "A×B" hidden layer 323 is composed of B convolution layers 324, and each convolution layer 324 includes A convolution filters for each channel. Such a hidden layer 323 is also referred to as an "A×B hidden layer" below. For example, a 64×2 hidden layer 323 is composed of two convolution layers 324, and each convolution layer 324 includes 64 convolution filters for each channel.

[0038] In the example shown in the figure, the feature extraction unit 321 includes a 64×2 hidden layer 323, a 128×2 hidden layer 323, a 256×3 hidden layer 323, and a 512×3 hidden layer 323, in this order. The upsampling unit 322 includes a 512×3 hidden layer 323, a 256×3 hidden layer 323, a 128×2 hidden layer 323, and a 64×2 hidden layer 323, in this order. The second pooling unit 327 connects the two 512×3 hidden layers 323 to each other. The number of hidden layers 323 constituting the nonlinear mapping unit 320 is not particularly limited and can be determined, for example, according to the number of pixels in the image data.

[0039] Note that this diagram shows an example of the configuration of the nonlinear mapping unit 320, and the nonlinear mapping unit 320 may have other configurations. For example, a 64×1 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. Reducing the number of convolutional layers 324 included in the intermediate layer 323 may further reduce the computational cost. Also, for example, a 32×2 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. Reducing the number of channels in the intermediate layer 323 may further reduce the computational cost. Furthermore, both the number of convolutional layers 324 and the number of channels in the intermediate layer 323 may be reduced.

[0040] Here, in the multiple intermediate layers 323 included in the feature extraction unit 321, it is preferable that the number of filters 325 increases each time the data passes through the first pooling unit 326. Specifically, the first intermediate layer 323a and the second intermediate layer 323b are connected to each other via the first pooling unit 326, and the second intermediate layer 323b is located after the first intermediate layer 323a. The first intermediate layer 323a is configured with convolutional layers 324 in which the number of filters 325 for each channel is N1, and the second intermediate layer 323b is configured with convolutional layers 324 in which the number of filters 325 for each channel is N2. In this case, it is preferable that N2 > N1. It is more preferable that N2 = N1 × 2.

[0041] Further, in the plurality of intermediate layers 323 included in the upsampling unit 322, it is preferable that the number of filters 325 decreases every time passing through the unpooling unit 328. Specifically, the third intermediate layer 323c and the fourth intermediate layer 323d are continuous with each other via the unpooling unit 328, and the fourth intermediate layer 323d is located at the subsequent stage of the third intermediate layer 323c. The third intermediate layer 323c is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N3, and the fourth intermediate layer 323d is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N4. At this time, it is preferable that N4 < N3 holds. More preferably, N3 = N4 × 2 holds.

[0042] In the feature extraction unit 321, image features having a plurality of levels of abstraction such as gradients and shapes are extracted from the intermediate data acquired from the input unit 310 as channels of the intermediate layer 323. FIG. 7 illustrates the configuration of the 64×2 intermediate layer 323. Referring to this figure, the processing in the intermediate layer 323 will be described. In the example of this figure, the intermediate layer 323 is composed of a first convolutional layer 324a and a second convolutional layer 324b, and each convolutional layer 324 includes 64 filters 325. In the first convolutional layer 324a, convolutional processing using the filter 325 is performed on each channel of the data input to the intermediate layer 323. For example, when the image input to the input unit 310 is an RGB image, for each of the three channels h

[0043] , , , 1 , 0 , i,j , i i (i = 1..3), processing is performed. Also, in the example of this figure, the filter 325 is a 64 types of 3×3 filters, that is, a total of 64×3 types of filters. As a result of the convolutional processing, for each channel i, 64 results h 0 i,j (i = 1..3, j = 1..64) are obtained.

[0043] Next, activation processing is performed on the outputs of the plurality of filters 325 in the activation unit 329. Specifically, for the corresponding results j of all channels, activation processing is performed on the sum of the corresponding elements. By this activation processing, the results h of the sixty-four channels 1i (i=1..64), i.e., the output of the first convolution layer 324a, is obtained as the image feature. The activation process is not particularly limited, but a process using at least one of a hyperbolic function, a sigmoid function, and a rectified linear function is preferable.

[0044] Furthermore, the output data of the first convolution layer 324a is used as input data for the second convolution layer 324b, and the same processing as that of the first convolution layer 324a is performed in the second convolution layer 324b to obtain the result h of 64 channels. 2 i (i=1..64), that is, the output of the second convolutional layer 324b, is obtained as the image features. The output of the second convolutional layer 324b becomes the output data of this 64×2 hidden layer 323.

[0045] Here, the structure of the filter 325 is not particularly limited, but a 3x3 two-dimensional filter is preferable. Furthermore, the coefficients of each filter 325 can be set independently. In this embodiment, the coefficients of each filter 325 are stored in the memory unit 390, and the nonlinear mapping unit 320 can read and use them for processing. Here, the coefficients of the multiple filters 325 may be determined based on correction information generated and corrected using machine learning. For example, the correction information includes the coefficients of the multiple filters 325 as multiple correction parameters. The nonlinear mapping unit 320 can further use this correction information to convert the intermediate data into mapped data. The memory unit 390 may be provided in the visual saliency calculation unit 3 or external to the visual saliency calculation unit 3. Furthermore, the nonlinear mapping unit 320 may obtain the correction information from an external source via a communication network.

[0046] 8(a) and 8(b) are diagrams illustrating examples of convolution processing performed by the filter 325. Both of FIGS. 8(a) and 8(b) illustrate examples of 3×3 convolution. The example in FIG. 8(a) illustrates convolution processing using nearest neighbor elements. The example in FIG. 8(b) illustrates convolution processing using neighbor elements with a distance of two or more. Note that convolution processing using neighbor elements with a distance of three or more is also possible. It is preferable that the filter 325 performs convolution processing using neighbor elements with a distance of two or more. This is because it allows for the extraction of a wider range of features, thereby further improving the accuracy of visual saliency estimation.

[0047] The above has described the operation of the 64×2 hidden layer 323. The operations of the other hidden layers 323 (such as the 128×2 hidden layer 323, the 256×3 hidden layer 323, and the 512×3 hidden layer 323) are the same as the operation of the 64×2 hidden layer 323, except for the number of convolutional layers 324 and the number of channels. Furthermore, the operations of the hidden layer 323 in the feature extraction unit 321 and the hidden layer 323 in the upsampling unit 322 are also the same as those described above.

[0048] 9(a) is a diagram for explaining the processing of the first pooling unit 326, FIG. 9(b) is a diagram for explaining the processing of the second pooling unit 327, and FIG. 9(c) is a diagram for explaining the processing of the unpooling unit 328.

[0049] In the feature extraction unit 321, data output from the intermediate layer 323 is subjected to pooling processing for each channel in the first pooling unit 326, and then input to the next intermediate layer 323. The first pooling unit 326 performs, for example, non-overlapping pooling processing. FIG. 9(a) shows processing for associating four 2×2 elements 30 with one element 30 for a group of elements included in each channel. The first pooling unit 326 performs such association for all elements 30. Here, the four 2×2 elements 30 are selected so that they do not overlap with each other. In this example, the number of elements in each channel is reduced to one-fourth. Note that, as long as the number of elements is reduced in the first pooling unit 326, the number of elements 30 before and after the association is not particularly limited.

[0050] The data output from the feature extraction unit 321 is input to the upsampling unit 322 via the second pooling unit 327. The second pooling unit 327 performs overlap pooling on the output data from the feature extraction unit 321. FIG. 9(b) shows a process of associating four 2×2 elements 30 with one element 30 while overlapping some of the elements 30. That is, in repeated associations, some of the four 2×2 elements 30 in a certain association are also included in the four 2×2 elements 30 in the next association. The second pooling unit 327 shown in this figure does not reduce the number of elements. Note that the number of elements 30 before and after association in the second pooling unit 327 is not particularly limited.

[0051] The methods of processing performed by the first pooling unit 326 and the second pooling unit 327 are not particularly limited, but examples include matching in which the maximum value of four elements 30 is matched to one element 30 (max pooling) and matching in which the average value of four elements 30 is matched to one element 30 (average pooling).

[0052] The data output from the second pooling unit 327 is input to the hidden layer 323 in the upsampling unit 322. Then, the output data from the hidden layer 323 of the upsampling unit 322 undergoes unpooling processing for each channel in the unpooling unit 328, and is then input to the next hidden layer 323. Figure 9(c) shows processing for expanding one element 30 into multiple elements 30. The method of expansion is not particularly limited, but an example is a method of duplicating one element 30 into four elements 30 (2 x 2).

[0053] The output data of the last hidden layer 323 of the upsampling unit 322 is output as mapping data from the nonlinear mapping unit 320 and input to the output unit 330. In the output step S130, the output unit 330 generates and outputs a visual saliency map by performing, for example, normalization or resolution conversion on the data acquired from the nonlinear mapping unit 320. The visual saliency map is, for example, an image (image data) that visualizes visual saliency using brightness values, as illustrated in FIG. 4(b). The visual saliency map may also be, for example, an image that is color-coded according to visual saliency, such as a heat map, or an image in which visual saliency regions with visual saliency higher than a predetermined standard are marked so as to be distinguishable from other positions. Furthermore, the visual saliency estimation information is not limited to map information displayed as an image or the like, but may also be a table listing information indicating visual saliency regions.

[0054] The gaze coordinate setting unit 4 sets an ideal gaze, described later, on a visual saliency map. The ideal gaze refers to the gaze of a vehicle driver along the direction of travel in an ideal traffic environment with no obstacles or other traffic participants. It is handled as an (x, y) coordinate in the image data and on the visual saliency map. In this embodiment, the ideal gaze is a fixed value. However, it may be treated as a function of the speed or road friction coefficient, which affect the stopping distance of a moving object, or may be determined using set route information. A vanishing point corresponding to the current road may also be used to calculate the ideal viewpoint. In this case, the vehicle speed may be detected and the ideal viewpoint may be set two or three seconds after the vanishing point and the vehicle position. In other words, the gaze coordinate setting unit 4 functions as a gaze position setting unit that sets the ideal gaze (reference gaze position) in the image according to a predetermined rule.

[0055] The vector error calculation unit 5 calculates a vector error based on the visual saliency map output by the visual saliency calculation unit 3 and the ideal gaze set by the gaze coordinate setting unit 4 for the visual saliency map and the image, and calculates a visual attention concentration level Ps (described later) indicating the degree of visual attention concentration based on the vector error. That is, the vector error calculation unit 5 functions as a visual attention concentration level calculation unit that calculates the degree of visual attention concentration in an image based on the visual saliency distribution information and the gaze position.

[0056] Here, the vector error in this embodiment will be described with reference to FIG. 10. FIG. 10 shows an example of a visual saliency map. This visual saliency map is shown with 256 gradation brightness values ​​of H pixels x V pixels, and similarly to FIG. 4, pixels with higher visual saliency are displayed with higher brightness. In FIG. 10, the coordinates (x, y) of the ideal line of sight are (x im ,y im), the vector error with the pixel at any coordinate (k, m) in the visual saliency map is calculated. When the coordinate with high brightness in the visual saliency map is far from the coordinate of the ideal gaze, it means that the position to be gazed at and the position that is actually easy to gaze at are far apart, and it can be said that the image is likely to distract visual attention. On the other hand, when the coordinate with high brightness is close to the coordinate of the ideal gaze, it means that the position to be gazed at and the position that is actually easy to gaze at are close, and it can be said that the image is likely to focus visual attention on the position to be gazed at.

[0057] Next, we will explain how to calculate the visual attention concentration level Ps in the vector error calculation unit 5. In this embodiment, the visual attention concentration level Ps is calculated by the following equation (1).

number

[0058] In equation (1), V vc is the pixel depth (brightness value), f w is the weighting function, d err indicates the vector error. This weighting function is, for example, V vc is a function that sets a weight based on the distance from the pixel showing the value of α to the coordinate of the ideal gaze. α is a coefficient that makes the visual attention concentration level Ps equal to 1 when the coordinate of the bright spot and the coordinate of the ideal gaze match in a visual saliency map (reference heat map) of one bright spot.

[0059] That is, the vector error calculation unit 5 (visual attention concentration calculation unit) calculates the degree of visual attention concentration based on the value of each pixel that constitutes the visual saliency map (visual saliency distribution information) and the vector error between the position of each pixel and the coordinate position of the ideal gaze (reference gaze position).

[0060] The visual attention concentration level Ps obtained in this way is the reciprocal of the weighted sum of the vector error of the coordinates of all pixels from the coordinates of the ideal gaze set on the visual saliency map and the relationship between the luminance value. This visual attention concentration level Ps is calculated to be a low value when the distribution of high luminance on the visual saliency map is far from the coordinates of the ideal gaze. In other words, the visual attention concentration level Ps can be said to be the concentration level relative to the ideal gaze.

[0061] Fig. 11 shows an example of an image input to the image input unit 2 and a visual saliency map obtained from that image. Fig. 11(a) is the input image, and (b) is the visual saliency map. In Fig. 11, if the coordinates of the ideal gaze are set on the road, for example, on a truck traveling ahead, the visual attention concentration level Ps in that case can be calculated.

[0062] The determination unit 6 determines whether the image input from the image input unit 2 is suspected of causing a safety problem such as an accident or near miss while the moving object was traveling, based on the temporal change in the visual attention concentration level Ps calculated by the vector error calculation unit 5. After the determination, the determination result is output to the outside.

[0063] Fig. 12 shows an example of how the visual concentration level Ps changes over time. Fig. 12 shows how the visual concentration level Ps changes in a 12-second video. In Fig. 12, the visual concentration level Ps changes suddenly between about 6.5 seconds and about 7 seconds. This occurs, for example, when another vehicle cuts in front of the vehicle, and by detecting such a change, it is possible to detect an incident that could be a near miss.

[0064] As shown in FIG. 12, images that may be suspected of being near misses or the like can be extracted by comparing the rate of change or value of change per short period of time of the visual attention concentration level Ps with a predetermined threshold value or the like.

[0065] Next, the operation (determination method) of the determination device 1 configured as described above will be described with reference to the flowchart in Fig. 13. This flowchart can be configured as a program executed by a computer that functions as the determination device 1, thereby forming a determination program. This determination program is not limited to being stored in a memory or the like included in the determination device 1, but may also be stored in a storage medium such as a memory card or an optical disk.

[0066] First, the image input unit 2 outputs the input image as image data to the visual saliency calculation unit 3 (step S11). In this step, the image data input to the image input unit 2 is decomposed into a time series of image frames, etc., and input to the visual saliency calculation unit 3. In this step, image processing such as noise removal and geometric transformation may also be performed.

[0067] Next, the visual saliency calculation unit 3 acquires a visual saliency map (step S12). The visual saliency calculation unit 3 outputs the visual saliency map in time series as shown in FIG. 4(b) by the above-described method.

[0068] Meanwhile, in parallel with step S12, the line-of-sight coordinate setting unit 4 sets the coordinates of the ideal line of sight (step S13). As described above, in this embodiment, these coordinates are fixed positions such as forward gaze.

[0069] Next, the vector error calculation unit 5 calculates the visual attention concentration level Ps from the visual saliency map and the ideal gaze (step S14). That is, as described above, the vector error between the coordinates of the ideal gaze and the coordinates of the visual saliency map is calculated, and the visual attention concentration level Ps is calculated using equation (1) based on the vector error and the value of each pixel.

[0070] Next, the judgment unit 6 judges whether the image input from the image input unit 2 is suspected of causing a safety problem while the moving object is traveling, based on the temporal change in the visual attention concentration level Ps calculated by the vector error calculation unit 5 (step S15).

[0071] Next, the determination unit 6 outputs the determination result of step S16 (step S16). In this step, the determination result is not limited to being simply output, but may be displayed on a display device or the like, or a label according to the determination result may be added to the image input from the image input unit 2. Alternatively, only images that are determined to be suspected of having a safety issue among the images input from the image input unit 2 may be stored in a specific storage device (specific storage area). In other words, the determination unit 6 functions as an output unit that outputs information related to the determination result.

[0072] As is clear from the above description, step S12 functions as an acquisition step, step S13 functions as a gaze position setting step, step S14 functions as a visual attention concentration level calculation step, and step S15 functions as a determination step.

[0073] An example of an image displaying the determination result in the above-mentioned determination device 1 will now be described with reference to Fig. 14. The image shown in Fig. 14 is generated by the determination unit 6 and displayed on a predetermined display device. The image 50 shown in Fig. 14 includes a driving image display area 51 and a visual attention concentration level display area 52.

[0074] The driving image display area 51 displays a driving image of a vehicle captured by a drive recorder, etc. The driving image display area 51 can display a vanishing point VP, a visual saliency map VM, an estimated gaze position GE, a detected object frame OF, a horizon line and depth distance estimation line HD, an estimated white line WL, and a near-miss judgment frame HF.

[0075] The vanishing point VP may be estimated from a white line estimation WL, which will be described later, or may be estimated using optical flow, etc. The visual saliency map VM displays a visual saliency map (heat map) for the image displayed in the driving image display area 51, superimposed on the image. Note that in FIG. 14, only high-brightness areas on the heat map are visible, but in reality, low-brightness areas are also superimposed on the driving image. In other words, it displays areas in the driving image that are likely to be drawn to.

[0076] In this embodiment, the gaze estimation position GE is estimated as the gaze position at the position with the highest brightness on the heat map. The detected object frame OF is displayed as a frame surrounding an object detected in the driving image as a result of object detection using a well-known algorithm. In this embodiment, the object detection specifies the type of object to be detected (vehicle, person, etc.), and only objects belonging to the specified type are detected. The object detection process may be performed by the determination unit 6 or by another block not shown in FIG. 1.

[0077] The horizontal line and depth distance estimation line HD indicate the horizontal line and depth distance in the driving image. The white line estimation WL indicates a recognized lane marking such as a white line in the driving image. The near-miss judgment frame HF is formed in a frame shape that follows the four sides of the driving image display area 51, and is displayed when the judgment unit 6 judges that a safety issue such as a near-miss is suspected to have occurred. Alternatively, the near-miss judgment frame HF may be displayed as a frame that is always blue, and when it is judged that a safety issue is suspected to have occurred, the display color may be changed to red, or the frame may flash.

[0078] The visual attention concentration level display area 52 is provided to the right of the driving image display area 51. The visual attention concentration level display area 52 displays the visual attention concentration level Ps calculated by the vector error calculation unit 5 in the form of a bar graph. In FIG. 14, reference numeral 52a denotes a bar indicating the visual attention concentration level Ps. When the visual attention concentration level Ps indicates a large value, the bar 52a becomes taller, and when the visual attention concentration level Ps indicates a small value, the bar 52a becomes shorter. In other words, when the bar 52a is in a high position, the concentration level tends to be high (concentrated), and when the bar 52a is in a low position, the concentration level tends to be low (dispersed).

[0079] The height of the bar 52a in the visual attention concentration display area 52 changes with the progression of image playback time because the visual saliency map is acquired frame by frame. Therefore, for example, a scene in which the height of the bar 52a increases suddenly indicates that the visual attention concentration has changed rapidly from dispersed to concentrated, and this bar 52a can also visually indicate that a safety issue is suspected by the determination unit 6.

[0080] According to this embodiment, the determination device 1 acquires a visual saliency map obtained by estimating the level of visual saliency in an image captured from a moving object, and a gaze coordinate setting unit 4 sets the coordinates of an ideal gaze. The vector error calculation unit 5 then calculates the visual attention concentration level Ps for the image based on the visual saliency map and the ideal gaze. The determination unit 6 determines whether a safety issue, such as a near miss, has occurred while the moving object is traveling, based on the temporal change in the visual attention concentration level Ps. By using the visual saliency map in this manner, it is possible to determine whether a safety issue, such as a near miss, has occurred based on the temporal change in the contextual attention state, in which the gaze tends to unconsciously focus on objects such as signs and pedestrians included in the image. Therefore, it is possible to determine whether a safety issue, such as an accident or near miss, has occurred due to psychological stress, based solely on the image.

[0081] Furthermore, the vector error calculation unit 5 calculates the visual attention concentration level Ps based on the value of each pixel constituting the visual saliency map and the vector error between the position of each pixel and the coordinate position of the ideal gaze. In this way, a value according to the difference between the position of high visual saliency and the ideal gaze is calculated as the visual attention concentration level Ps. Therefore, for example, the value of the visual attention concentration level Ps can be changed according to the distance between the position of high visual saliency and the ideal gaze.

[0082] Furthermore, the determination unit 6 outputs information relating to the determination result, which allows the determination result and information based on the determination result to be displayed or communicated to the outside.

[0083] Furthermore, the image input to the image input unit 2 may be obtained when a sensor equipped in the moving object detects sudden braking, e.g., when the acceleration of the moving object is equal to or greater than a reference value. This allows, for example, a drive recorder or the like to further determine whether the image extracted based on acceleration is a near miss or other incident, thereby eliminating the need for manual determination of whether the sudden braking is related to an accident or near miss or to factors other than an accident or near miss (e.g., the moving object going over a bump or reckless driving). In addition to sudden braking, sudden driving maneuvers (in other words, dangerous behavior) may also be included. For example, an image may be extracted based on the possibility of a near miss occurring, such as an operation to avoid something due to lateral acceleration or an operation to return to normal after noticing that the vehicle has run off a white line. Furthermore, sudden acceleration may also be included. For example, an image may be extracted based on the possibility that the driver suddenly accelerated after noticing a decrease in speed due to distracted driving on a highway.

[0084] The visual saliency calculation unit 3 includes an input unit 310 that converts an image into intermediate data that can be mapped, a nonlinear mapping unit 320 that converts the intermediate data into mapped data, and an output unit 330 that generates saliency estimation information indicating a saliency distribution based on the mapped data. The nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. This allows visual saliency to be estimated with low computational cost. Furthermore, the visual saliency estimated in this manner reflects the contextual attention state.

[0085] Furthermore, the present invention is not limited to the above-described embodiments. That is, a person skilled in the art can implement various modifications in accordance with conventionally known knowledge without departing from the gist of the present invention. As long as such modifications still include the determination device of the present invention, they are of course included in the scope of the present invention. [Explanation of symbols]

[0086] 1 Judgment device 2 Image input section 3. Visual saliency calculation unit (acquisition unit) 4 Line-of-sight coordinate setting section (line-of-sight position setting section) 5. Vector error calculation unit (visual attention concentration calculation unit) 6 Judgment section

Claims

[Claim 1] an acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency in an image captured from a moving object; a gaze position setting unit that sets a reference gaze position in the image according to a predetermined rule; a visual attention concentration calculation unit that calculates a visual attention concentration level in the image based on the visual saliency distribution information and the gaze position; a determination unit that determines that a safety problem is suspected to have occurred while the moving object is traveling, based on a temporal change in the degree of visual attention concentration; A determination device comprising:

Citation Information

Patent Citations

  • Device for operating degree of risk of driving behavior

    JP2003099899A

  • Weighting matrix learning device, line-of-sight direction prediction system, warning system and weighting matrix learning method

    JP2016130959A

  • Drive recorder device and event identification of the same

    JP2012164131A