Map data generation device

The map data generation device addresses the limitation of traditional map systems by using visual saliency analysis to identify and mark visually demanding areas, enhancing safety by reducing driver distraction and fatigue.

JP2025107365AInactive Publication Date: 2025-07-17PIONEER IP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025077791
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-10
Filing Date
2025-05-08
Publication Date
2025-07-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing map data systems fail to adequately identify and highlight areas requiring visual attention beyond traditional hazards like intersections and sharp curves, such as roads with high visual distraction or monotony, which can lead to driver fatigue and accidents.

Method used

A map data generation device that acquires images from a moving object, estimates visual saliency, and adds points or sections requiring attention to the map data based on visual saliency distribution information, using methods like convolutional neural networks to analyze and process image data.

Benefits of technology

Enhances map data by identifying visually demanding areas, reducing driver distraction and fatigue by highlighting areas with high visual load, monotony, or potential distractions, thereby improving road safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025107365000001_ABST
    Figure 2025107365000001_ABST
Patent Text Reader

Abstract

To add a point requiring visual attention to map data.SOLUTION: A map data generation device 1 includes: input means 2 for acquiring image data obtained by imaging an outside from a vehicle and point data of the vehicle so as to correlate both of the data with each other; visual saliency extraction means 3 for generating a visual saliency map obtained by estimating the level of visual saliency based on the image data; analysis means 4 for analyzing whether a point / a section indicated by position information corresponding to the visual saliency map is the point / the section requiring visual attention based on the visual saliency map; and addition means 5 for adding the point / the section requiring visual attention to map data based on an analysis result of the analysis means 4.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a map data generation device that adds predetermined information to map data based on an image captured from a moving object to the outside.

Background Art

[0002] When a vehicle, for example, travels as a moving object, it is already known to display on a map points that should be driven with particular caution, such as intersections where accidents are likely to occur, level crossings, and sharp curves (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Points that should be driven with caution are not limited to the above-mentioned intersections where accidents are likely to occur, level crossings, and sharp curves. For example, even if it is not a sharp curve or the like, attention is required on roads where visual load is felt, the risk of distraction is high, or the road is monotonous.

[0005] As an example of the problems to be solved by the present invention, adding points that require visual attention and the like to map data can be cited.

Means for Solving the Problems

[0006] In order to solve the above problems, the invention according to claim 1 includes a first acquisition unit that acquires input information in which an image captured from a moving body of the outside and the position information of the moving body are associated, a second acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency in the image based on the image, an analysis unit that analyzes whether a point or section indicated by the position information corresponding to the visual saliency distribution information is a point or section that requires visual attention based on the visual saliency distribution information, and an addition unit that adds the point or section that requires visual attention to map data based on the analysis result of the analysis unit.

[0007] The invention according to claim 9 is a map data generation method executed by a map data generation device that adds predetermined information to map data based on an image captured from a moving body of the outside, the method including a first acquisition step of acquiring input information in which the image and the position information of the moving body are associated, a second acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency in the image based on the image, an analysis step of analyzing whether a point or section indicated by the position information corresponding to the visual saliency distribution information is a point or section that requires visual attention based on the visual saliency distribution information, and an addition step of adding the point or section that requires visual attention to map data based on the analysis result of the analysis step.

[0008] The invention according to claim 10 is characterized in that the map data generation method according to claim 9 is executed by a computer.

[0009] The invention according to claim 11 is characterized in that the map data generation program according to claim 10 is stored.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Mode for Carrying Out the Invention

[0011] Hereinafter, a map data generation device according to an embodiment of the present invention will be described. In the map data generation device according to an embodiment of the present invention, a first acquisition unit acquires input information in which an image obtained by imaging the outside from a moving body and the position information of the moving body are associated, and a second acquisition unit acquires visual saliency distribution information obtained by estimating the level of visual saliency in the image based on the image. Then, an analysis unit analyzes, based on the visual saliency distribution information, whether the point or section indicated by the position information corresponding to the visual saliency distribution information is a point or section that requires visual attention, and an addition unit adds a point or section that requires visual attention to the map data based on the analysis result of the analysis unit. By doing so, it is possible to estimate visual saliency based on an image obtained by imaging the outside from a moving body and add a point or the like that requires visual attention to the map data based on the estimated characteristics.

[0012] Further, the analysis unit includes a movement amount calculation unit that calculates the movement amount of the estimated fixation point based on the visual saliency distribution information, and a first determination unit that determines whether the point or section indicated by the position information corresponding to the visual saliency distribution information has a high visual recognition load by comparing the calculated movement amount of the estimated fixation point with a first threshold value. The addition unit may add the point or section determined to have a high visual recognition load as the point that requires visual attention to the map data. By doing so, it is possible to easily determine whether the visual recognition load is high by comparing the movement amount of the estimated fixation point with the first threshold value, and add points that require attention, etc. to the map data based on this determination result.

[0013] Further, the movement amount calculation unit may calculate the movement amount by estimating the estimated fixation point as the position on the image where the visual saliency is the maximum value in the visual saliency distribution information. By doing so, the movement amount can be calculated based on the position estimated to be the most visually recognized.

[0014] Further, the analysis unit includes a line-of-sight position setting unit that sets a reference line-of-sight position in the image according to a predetermined rule, a visual attention concentration calculation unit that calculates the concentration of visual attention in the image based on the visual saliency distribution information and the reference line-of-sight position, and a second determination unit that determines whether the point or section indicated by the position information corresponding to the visual saliency distribution information is a point or section that requires visual attention based on the concentration of visual attention. The addition unit may add the point or section determined to be a point or section that requires visual attention to the map data. By doing so, it is possible to determine points that require attention, etc. based on the concentration of visual attention obtained from the visual saliency distribution information and add them to the map data.

[0015] Further, the second acquisition unit acquires visual saliency distribution information for each approach road, which is the road when entering an intersection, from the images for each approach road. The line-of-sight position setting unit sets the reference line-of-sight position in the image for each exit road, which is the road to exit after entering the intersection, with respect to the visual saliency distribution information. The visual attention concentration calculation unit calculates the concentration of the visual attention for each exit road in the image based on the visual saliency distribution information and the reference line-of-sight position. The second determination unit may determine whether the intersection is a visually attention-required location based on the concentration of the visual attention for each exit road. By doing so, for the intersection, it can be determined whether it is a location that requires attention and added to the map data.

[0016] Further, the analysis unit includes a peak position detection unit that detects at least one peak position in the visual saliency distribution information in time series, a fixation range setting unit that sets a range to be fixated by the driver of the moving object in the image, a side-glance output unit that outputs information indicating a tendency of side-glancing when the peak position has continuously deviated from the range to be fixated for a predetermined time or more, and a third determination unit that determines whether the location or section indicated by the position information corresponding to the visual saliency distribution information is a visually attention-required location or section based on the information indicating a tendency of side-glancing. The addition unit may add the location or section determined to be a visually attention-required location or section to the map data. By doing so, a location having a tendency of side-glancing or the like can be determined as a location that requires attention or the like and added to the map data.

[0017] Further, the analysis unit includes a monotonic determination unit that determines whether the image has a monotonic tendency using a statistic calculated based on the visual saliency distribution information, and a fourth determination unit that determines whether the location or section indicated by the position information corresponding to the visual saliency distribution information is a visually attention-required location or section based on the determination result of the monotonic determination unit. The addition unit may add the location or section determined to be a visually attention-required location or section to the map data. By doing so, a location or the like determined to have a monotonic tendency can be determined as a location that requires attention or the like and added to the map data.

[0018] Further, the second acquisition unit includes an input unit that converts an image into intermediate data that can be subjected to mapping processing, a non-linear mapping unit that converts the intermediate data into mapping data, and an output unit that generates saliency estimation information indicating a saliency distribution based on the mapping data. The non-linear mapping unit may include a feature extraction unit that extracts features from the intermediate data, and an upsampling unit that upsamples the data generated by the feature extraction unit. By doing so, visual saliency can be estimated with a small computational cost.

[0019] Also, a map data generation method according to an embodiment of the present invention includes, in a first acquisition step, acquiring input information in which an image captured from the outside by a moving body and position information of the moving body are associated, and in a second acquisition step, acquiring visual saliency distribution information obtained by estimating the level of visual saliency in the image based on the image. Then, in an analysis step, based on the visual saliency distribution information, it is analyzed whether a point or section indicated by the position information corresponding to the visual saliency distribution information is a point or section that requires visual attention, and in an addition step, based on the analysis result of the analysis step, a point or section that requires visual attention is added to the map data. By doing so, visual saliency can be estimated based on an image captured from the outside by a moving body, and a point or the like that requires attention can be added to the map data based on the estimated features.

[0020] Also, the above-described map data generation method is executed by a computer. By doing so, using a computer, visual saliency can be estimated based on an image captured from the outside by a moving body, and a point or the like that requires visual attention can be added to the map data based on the estimated features.

[0021] Also, the above-described map data generation program may be stored in a computer-readable storage medium. By doing so, the program can be distributed not only when incorporated into a device but also as a single entity, and version updates and the like can be easily performed.

Example

[0022] A map data generation device according to an embodiment of the present invention will be described with reference to FIGS. 1 to 11. The map data generation device according to this embodiment can be configured by, for example, a server device installed in an office or the like.

[0023] As shown in FIG. 1, the map data generation device 1 includes an input means 2, a visual saliency extraction means 3, an analysis means 4, and an addition means 5.

[0024] The input means 2 receives, for example, an image (moving image) captured by a camera or the like and position information (point data) output from a GPS (Global Positioning System) receiver or the like, and outputs the image in association with the point data. The input moving image is output as image data decomposed in time series, for example, for each frame. A still image may be input as the image input to the input means 2, but it is preferably input as an image group composed of a plurality of still images along the time series.

[0025] Examples of the image input to the input means 2 include an image in which the traveling direction of a vehicle is imaged. That is, it is an image continuously captured of the outside from a moving body. This image may be an image including 180° or 360° or the like in the horizontal direction other than the traveling direction, such as a so-called panoramic image or an image obtained using a plurality of cameras. Further, the input to the input means 2 is not limited to directly inputting an image captured by a camera, and may be an image read from a recording medium such as a hard disk drive or a memory card. That is, the input means 2 functions as a first acquisition unit that acquires input information in which an image captured of the outside from a moving body is associated with the position information of the moving body.

[0026] The visual saliency extraction means 3 receives image data from the input means 2 and outputs a visual saliency map as visual saliency estimation information described later. That is, the visual saliency extraction means 3 functions as a second acquisition unit that acquires a visual saliency map (visual saliency distribution information) obtained by estimating the level of visual saliency based on an image captured of the outside from a moving body.

[0027] FIG. 2 is a block diagram illustrating the configuration of the visual saliency extraction means 3. The visual saliency extraction means 3 according to the present embodiment includes an input unit 310, a non-linear mapping unit 320, and an output unit 330. The input unit 310 converts an image into intermediate data capable of being mapped. The non-linear mapping unit 320 converts the intermediate data into mapped data. The output unit 330 generates saliency estimation information indicating a saliency distribution based on the mapped data. The non-linear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that performs upsampling of the data generated by the feature extraction unit 321. This will be described in detail below.

[0028] FIG. 3(a) is a diagram illustrating an image input to the visual saliency extraction means 3, and FIG. 3(b) is a diagram illustrating an image showing the visual saliency distribution estimated for FIG. 3(a). The visual saliency extraction means 3 according to the present embodiment is an apparatus that estimates the visual saliency of each part in an image. Visual saliency means, for example, conspicuousness or ease of concentration of the line of sight. Specifically, visual saliency is indicated by probability or the like. Here, the magnitude of the probability corresponds to, for example, the magnitude of the probability that the line of sight of a person who has seen the image will be directed to that position.

[0029] FIG. 3(a) and FIG. 3(b) are in corresponding positions. In FIG. 3(a), the higher the visual saliency, the higher the luminance is displayed in FIG. 3(b). An image showing the visual saliency distribution like FIG. 3(b) is an example of the visual saliency map output by the output unit 330. In the example of this figure, the visual saliency is visualized with a luminance value of 256 gradations. An example of the visual saliency map output by the output unit 330 will be described in detail later.

[0030] FIG. 4 is a flowchart illustrating the operation of the visual saliency extraction means 3 according to the present embodiment. The flowchart shown in FIG. 4 is part of a map data generation method executed by a computer, and includes an input step S115, a non-linear mapping step S120, and an output step S130. In the input step S115, an image is converted into intermediate data that can be subjected to mapping processing. In the non-linear mapping step S120, the intermediate data is converted into mapped data. In the output step S130, visual saliency estimation information indicating a saliency distribution is generated based on the mapped data. Here, the non-linear mapping step S120 includes a feature extraction step S121 for extracting features from the intermediate data, and an upsampling step S122 for upsampling the data generated in the feature extraction step S121.

[0031] Returning to FIG. 2, each component of the visual saliency extraction means 3 will be described. In the input step S115, the input unit 310 acquires an image and converts it into intermediate data. The input unit 310 acquires image data from the input means 2. Then, the input unit 310 converts the acquired image into intermediate data. The intermediate data is not particularly limited as long as it is data that can be received by the non-linear mapping unit 320, and is, for example, a high-dimensional tensor. Further, the intermediate data is, for example, data obtained by normalizing the luminance of the acquired image, or data obtained by converting each pixel of the acquired image into a luminance gradient. In the input step S115, the input unit 310 may further perform noise removal or resolution conversion of the image.

[0032] In the non-linear mapping step S120, the non-linear mapping unit 320 acquires intermediate data from the input unit 310. Then, in the non-linear mapping unit 320, the intermediate data is converted into mapped data. Here, the mapped data is, for example, a high-dimensional tensor. The mapping process applied to the intermediate data by the non-linear mapping unit 320 is, for example, a mapping process controllable by parameters or the like, and is preferably a process by a function, a functional, or a neural network.

[0033] FIG. 5 is a diagram illustrating in detail the configuration of the non-linear mapping unit 320, and FIG. 6 is a diagram illustrating the configuration of the intermediate layer 323. As described above, the non-linear mapping unit 320 includes a feature extraction unit 321 and an upsampling unit 322. In the feature extraction unit 321, a feature extraction step S121 is performed, and in the upsampling unit 322, an upsampling step S122 is performed. Also, in the example of this figure, at least one of the feature extraction unit 321 and the upsampling unit 322 is configured to include a neural network including a plurality of intermediate layers 323. In the neural network, a plurality of intermediate layers 323 are combined.

[0034] In particular, the neural network is preferably a convolutional neural network. Specifically, each of the plurality of intermediate layers 323 includes one or more convolutional layers 324. And in the convolutional layer 324, convolution of the input data is performed by a plurality of filters 325, and activation processing is performed on the outputs of the plurality of filters 325.

[0035] In the example of FIG. 5, the feature extraction unit 321 is configured to include a neural network including a plurality of intermediate layers 323, and includes a first pooling unit 326 between the plurality of intermediate layers 323. Also, the upsampling unit 322 is configured to include a neural network including a plurality of intermediate layers 323, and includes an unpooling unit 328 between the plurality of intermediate layers 323. Further, the feature extraction unit 321 and the upsampling unit 322 are connected to each other via a second pooling unit 327 that performs overlapping pooling.

[0036] Note that in the example of this figure, each intermediate layer 323 is composed of two or more convolutional layers 324. However, at least some of the intermediate layers 323 may be composed of only one convolutional layer 324. Adjacent intermediate layers 323 are separated by any one of the first pooling unit 326, the second pooling unit 327, and the unpooling unit 328. Here, when two or more convolutional layers 324 are included in the intermediate layer 323, it is preferable that the number of filters 325 in those convolutional layers 324 is equal to each other.

[0037] In this figure, the intermediate layer 323 denoted as "A×B" consists of B convolutional layers 324, and each convolutional layer 324 means including A convolutional filters for each channel. Such an intermediate layer 323 is also referred to as an "A×B intermediate layer" hereinafter. For example, the 64×2 intermediate layer 323 consists of 2 convolutional layers 324, and each convolutional layer 324 means including 64 convolutional filters for each channel.

[0038] In the example of this figure, the feature extraction unit 321 includes the 64×2 intermediate layer 323, the 128×2 intermediate layer 323, the 256×3 intermediate layer 323, and the 512×3 intermediate layer 323 in this order. Also, the upsampling unit 322 includes the 512×3 intermediate layer 323, the 256×3 intermediate layer 323, the 128×2 intermediate layer 323, and the 64×2 intermediate layer 323 in this order. Also, the second pooling unit 327 connects two 512×3 intermediate layers 323 to each other. Note that the number of intermediate layers 323 constituting the non-linear mapping unit 320 is not particularly limited, and can be determined according to, for example, the number of pixels of the image data.

[0039] Note that this figure is an example of the configuration of the non-linear mapping unit 320, and the non-linear mapping unit 320 may have other configurations. For example, a 64×1 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. By reducing the number of convolutional layers 324 included in the intermediate layer 323, the calculation cost may be further reduced. Also, for example, a 32×2 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. By reducing the number of channels of the intermediate layer 323, the calculation cost may be further reduced. Furthermore, both the number of convolutional layers 324 and the number of channels in the intermediate layer 323 may be reduced.

[0040] Here, in the plurality of intermediate layers 323 included in the feature extraction unit 321, it is preferable that the number of filters 325 increases every time it passes through the first pooling unit 326. Specifically, the first intermediate layer 323a and the second intermediate layer 323b are continuous with each other via the first pooling unit 326, and the second intermediate layer 323b is located at the subsequent stage of the first intermediate layer 323a. The first intermediate layer 323a is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N1, and the second intermediate layer 323b is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N2. At this time, it is preferable that N2 > N1 holds. More preferably, N2 = N1 × 2 holds.

[0041] Also, in the plurality of intermediate layers 323 included in the upsampling unit 322, it is preferable that the number of filters 325 decreases every time it passes through the unpooling unit 328. Specifically, the third intermediate layer 323c and the fourth intermediate layer 323d are continuous with each other via the unpooling unit 328, and the fourth intermediate layer 323d is located at the subsequent stage of the third intermediate layer 323c. The third intermediate layer 323c is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N3, and the fourth intermediate layer 323d is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N4. At this time, it is preferable that N4 < N3 holds. More preferably, N3 = N4 × 2 holds.

[0042] In the feature extraction unit 321, image features with multiple levels of abstraction, such as gradients and shapes, are extracted from the intermediate data acquired from the input unit 310 as channels of the intermediate layer 323. FIG. 6 illustrates the configuration of the 64×2 intermediate layer 323. With reference to this figure, the processing in the intermediate layer 323 will be described. In the example of this figure, the intermediate layer 323 is composed of a first convolutional layer 324a and a second convolutional layer 324b, and each convolutional layer 324 is provided with 64 filters 325. In the first convolutional layer 324a, a convolutional process using the filter 325 is performed on each channel of the data input to the intermediate layer 323. For example, when the image input to the input unit 310 is an RGB image, processing is performed on each of the three channels h 0 i (i = 1..3). Also, in the example of this figure, the filter 325 is a 3×3 filter of 64 types, that is, a total of 64×3 types of filters. As a result of the convolutional process, for each channel i, 64 results h 0 i,j (i = 1..3, j = 1..64) are obtained.

[0043] Next, activation processing is performed on the outputs of the plurality of filters 325 in the activation unit 329. Specifically, for the corresponding results j of all channels, activation processing is performed on the sum of the corresponding elements. By this activation processing, 64-channel results h 1 i (i = 1..64), that is, the output of the first convolutional layer 324a, are obtained as image features. The activation processing is not particularly limited, but processing using at least any one of a hyperbolic function, a sigmoid function, and a rectified linear function is preferable.

[0044] Furthermore, the output data of the first convolutional layer 324a is used as the input data of the second convolutional layer 324b, and the same processing as that of the first convolutional layer 324a is performed in the second convolutional layer 324b to obtain 64-channel results h 2 i (i = 1..64), that is, the output of the second convolutional layer 324b, is obtained as image features. The output of the second convolutional layer 324b becomes the output data of this 64×2 intermediate layer 323.

[0045] Here, the structure of the filter 325 is not particularly limited, but it is preferably a 3×3 two-dimensional filter. Also, the coefficients of each filter 325 can be set independently. In this embodiment, the coefficients of each filter 325 are held in the storage unit 390, and the non-linear mapping unit 320 can read them out and use them for processing. Here, the coefficients of the plurality of filters 325 may be determined based on correction information generated and corrected using machine learning. For example, the correction information includes the coefficients of the plurality of filters 325 as a plurality of correction parameters. The non-linear mapping unit 320 can further use this correction information to convert the intermediate data into mapped data. The storage unit 390 may be provided in the visual saliency extraction means 3, or may be provided outside the visual saliency extraction means 3. Also, the non-linear mapping unit 320 may acquire the correction information from the outside via a communication network.

[0046] FIG. 7(a) and FIG. 7(b) are diagrams showing examples of convolution processing performed by the filter 325. In FIG. 7(a) and FIG. 7(b), examples of 3×3 convolution are shown. The example in FIG. 7(a) is convolution processing using the nearest neighbor elements. The example in FIG. 7(b) is convolution processing using neighboring elements with a distance of two or more. Note that convolution processing using neighboring elements with a distance of three or more is also possible. The filter 325 preferably performs convolution processing using neighboring elements with a distance of two or more. This is because more extensive features can be extracted, and the estimation accuracy of visual saliency can be further improved.

[0047] The operation of the 64×2 intermediate layer 323 has been described above. The operations of the other intermediate layers 323 (such as the 128×2 intermediate layer 323, the 256×3 intermediate layer 323, and the 512×3 intermediate layer 323) are the same as the operation of the 64×2 intermediate layer 323, except for the number of convolution layers 324 and the number of channels. Also, the operation of the intermediate layer 323 in the feature extraction unit 321 and the operation of the intermediate layer 323 in the upsampling unit 322 are the same as described above.

[0048] FIG. 8(a) is a diagram for explaining the processing of the first pooling unit 326, FIG. 8(b) is a diagram for explaining the processing of the second pooling unit 327, and FIG. 8(c) is a diagram for explaining the processing of the unpooling unit 328.

[0049] In the feature extraction unit 321, the data output from the intermediate layer 323 is input to the next intermediate layer 323 after being subjected to pooling processing for each channel in the first pooling unit 326. In the first pooling unit 326, for example, non-overlapping pooling processing is performed. FIG. 8(a) shows the processing of associating four 2×2 elements 30 included in each channel with one element 30. In the first pooling unit 326, such association is performed for all the elements 30. Here, the four 2×2 elements 30 are selected so as not to overlap with each other. In this example, the number of elements in each channel is reduced to one-fourth. Note that as long as the number of elements is reduced in the first pooling unit 326, the number of elements 30 before and after the association is not particularly limited.

[0050] The data output from the feature extraction unit 321 is input to the upsampling unit 322 via the second pooling unit 327. In the second pooling unit 327, overlapping pooling is performed on the output data from the feature extraction unit 321. FIG. 8(b) shows the processing of associating four 2×2 elements 30 with one element 30 while overlapping some of the elements 30. That is, in the repeated association, some of the four 2×2 elements 30 in a certain association are also included in the four 2×2 elements 30 in the next association. In the second pooling unit 327 as shown in this figure, the number of elements is not reduced. Note that the number of elements 30 before and after the association in the second pooling unit 327 is not particularly limited.

[0051] The method of each process performed by the first pooling unit 326 and the second pooling unit 327 is not particularly limited. For example, there are max pooling that associates the maximum value of four elements 30 with one element 30 and average pooling that associates the average value of four elements 30 with one element 30.

[0052] The data output from the second pooling unit 327 is input to the intermediate layer 323 in the upsampling unit 322. Then, the output data from the intermediate layer 323 of the upsampling unit 322 is input to the next intermediate layer 323 after being subjected to unpooling processing for each channel in the unpooling unit 328. FIG. 8(c) shows a process of expanding one element 30 into a plurality of elements 30. The expansion method is not particularly limited, and an example is a method of replicating one element 30 into four 2×2 elements 30.

[0053] The output data of the last intermediate layer 323 of the upsampling unit 322 is output from the non-linear mapping unit 320 as mapping data and input to the output unit 330. In the output step S130, the output unit 330 generates and outputs a visual saliency map by performing, for example, normalization or resolution conversion on the data acquired from the non-linear mapping unit 320. The visual saliency map is, for example, an image (image data) in which visual saliency is visualized as luminance values as illustrated in FIG. 3(b). Further, the visual saliency map may be an image color-coded according to visual saliency like a heat map, or an image in which a visually salient region having a visual saliency higher than a predetermined standard is marked distinguishable from other positions. Furthermore, the visual saliency estimation information is not limited to map information shown as an image or the like, and may be a table or the like listing information indicating visually salient regions.

[0054] The analysis means 4 analyzes whether a point corresponding to the visual saliency map output by the visual saliency extraction means 3 tends to have a high visual recognition load based on the visual saliency map. As shown in FIG. 9, the analysis means 4 includes a visual recognition load amount calculation means 41 and a visual recognition load determination means 42.

[0055] The visual load amount calculation means 41 calculates the visual load amount based on the visual saliency map output by the visual saliency extraction means 3. The visual load amount, which is the result calculated by the visual load amount calculation means 41, may be, for example, a scalar quantity or a vector quantity. Alternatively, it may be single data or a plurality of time-series data. The visual load amount calculation means 41 estimates fixation point information and calculates the fixation point movement amount as the visual load amount.

[0056] The details of the visual load amount calculation means 41 will be described. First, fixation point information is estimated from the time-series visual saliency maps output by the visual saliency extraction means 3. Although the definition of the fixation point information is not particularly limited, for example, it can be the position (coordinates) where the saliency value is the maximum value. That is, the visual load amount calculation means 41 estimates the fixation point information as the position on the image where the visual saliency is the maximum value in the visual saliency map (visual saliency distribution information).

[0057] Then, the time-series fixation point movement amount is calculated from the estimated time-series fixation point information. The calculated fixation point movement amount also becomes time-series data. The calculation method is not particularly limited, but for example, it can be the Euclidean distance between fixation point coordinates that are in a sequential relationship in time series. That is, in this embodiment, the fixation point movement amount is calculated as the visual load amount. That is, the visual load amount calculation means 41 functions as a movement amount calculation unit that calculates the movement amount of the fixation point (estimated fixation point) based on the generated visual saliency map (visual saliency distribution information).

[0058] The visual load determination means 42 determines whether the target point or section has a large visual load based on the movement amount calculated by the visual load amount calculation means 41. The determination method in the visual load determination means 42 will be described later.

[0059] The addition means 5 adds attention point information to the acquired map data based on the analysis result in the analysis means 4. That is, the addition means 5 adds the point determined to have a large visual load by the visual load determination means 42 to the map data as a point that requires attention.

[0060] Next, the operation (map data generation method) in the map data generation device 1 having the above-described configuration will be described with reference to the flowchart of FIG. 10. Further, this flowchart can be configured as a program executed by a computer functioning as the map data generation device 1 to obtain a map data generation program. Further, this map data generation program is not limited to being stored in a memory or the like included in the map data generation device 1, and may be stored in a storage medium such as a memory card or an optical disk.

[0061] First, the input means 2 acquires point data (step S210). The point data may be acquired from a GPS receiver or the like as described above.

[0062] Next, the input means 2 acquires a driving video (image data) (step S220). In this step, the image data input to the input means 2 is decomposed into a time series such as image frames and input to the visual saliency extraction means 3 in association with the point data acquired in step S210. Further, image processing such as noise removal and geometric transformation may be performed in this step. Note that the order of steps S210 and S220 may be reversed.

[0063] Next, the visual saliency extraction means 3 extracts a visual saliency map (step S230). The visual saliency map is output in a time series as a visual saliency map as shown in FIG. 3(b) by the above-described method in the visual saliency extraction means 3.

[0064] Next, the fixation point movement amount calculation means 41 calculates the fixation point movement amount by the above-described method (step S240).

[0065] Next, the visual load determination means 42 determines whether or not the amount of fixation point movement calculated in step S240 is equal to or greater than a predetermined threshold (step S250). This threshold is a threshold related to the amount of fixation point movement. That is, the visual load determination means 42 functions as a first determination unit that determines whether the point or section indicated by the point data (position information) corresponding to the visual saliency map (visual saliency distribution information) has a high visual load tendency by comparing the calculated amount of fixation point movement with the first threshold. As a result of the determination in step S250, if the amount of fixation point movement is equal to or greater than the predetermined threshold (step S250: YES), the addition means 5 registers (adds) the target point as a caution point with a large visual load amount in the map data (step S260).

[0066] Also, as a result of the determination in step S250, if the amount of fixation point movement is less than the predetermined threshold (step S250: NO), since the target point does not have a large visual load amount, it is not registered as a caution point.

[0067] Here, an example of a map in which caution points are registered is shown in FIG. 11. The round marks indicated by the symbol W in FIG. 11 indicate caution points. FIG. 11 is an example showing points with a large visual load amount. Here, the color or density of the round marks may be changed according to the magnitude of the visual load amount, or the size of the round marks may be changed.

[0068] According to this embodiment, the map data generation device 1 acquires the image data captured from the outside by the vehicle by the input means 2 and the point data of the vehicle, associates both data, and generates a visual saliency map obtained by estimating the level of visual saliency based on the image data by the visual saliency extraction means 3. Then, based on the visual saliency map, the analysis means 4 analyzes whether the point or section indicated by the position information corresponding to the visual saliency map has a high visual load tendency, and the addition means 5 adds the point or section showing a high visual load tendency to the map data based on the analysis result of the analysis means 4. By doing so, it is possible to estimate the visual saliency based on the image captured from the outside of the vehicle and add the points that visually feel a load based on the estimated characteristics to the map data.

[0069] Further, the analysis means 4 includes a visual load amount calculation means 41 that calculates the fixation point movement amount based on the visual saliency map, and a visual load determination means 42 that determines whether the point or section indicated by the point data corresponding to the visual saliency map has a high visual load tendency by comparing the calculated fixation point movement amount with a first threshold value. By doing so, it is possible to easily determine whether the visual load amount tends to be high by comparing the fixation point movement amount with the first threshold value.

[0070] Further, the visual load amount calculation means 41 calculates the movement amount by estimating the fixation point as the position on the image where the visual saliency is the maximum value in the visual saliency map. By doing so, the movement amount can be calculated based on the position estimated to be the most visually recognized.

[0071] Further, the visual saliency extraction means 3 includes an input unit 310 that converts an image into intermediate data capable of mapping processing, a non-linear mapping unit 320 that converts the intermediate data into mapping data, and an output unit 330 that generates saliency estimation information indicating a saliency distribution based on the mapping data. The non-linear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. By doing so, visual saliency can be estimated with a small computational cost.

Embodiment

[0072] Next, a map data generation device according to a second embodiment of the present invention will be described with reference to FIGS. 12 to 15. The same parts as those in the first embodiment described above are denoted by the same reference numerals and the description thereof will be omitted.

[0073] In this embodiment, instead of the visual load described in the first embodiment, the visual attention concentration degree is calculated, and a point that requires visual attention or the like is added to the map data based on the visual attention concentration degree. The visual attention concentration degree will be described later.

[0074] As shown in FIG. 12, the analysis means 4 according to this embodiment includes a line-of-sight coordinate setting means 43, a vector error calculation means 44, and an output means 45.

[0075] The line-of-sight coordinate setting means 43 sets an ideal line of sight, which will be described later, on the visual saliency map. The ideal line of sight refers to the line of sight that a driver of an automobile directs along the traveling direction in an ideal traffic environment where there are no obstacles or traffic participants other than oneself. On the image data and the visual saliency map, it is handled as (x, y) coordinates. In this embodiment, the ideal line of sight is a fixed value, but it may be handled as a function of the speed that affects the stopping distance of the moving body and the road friction coefficient, or may be determined using the set route information. That is, the line-of-sight coordinate setting means 43 functions as a line-of-sight position setting unit that sets an ideal line of sight (reference line-of-sight position) in the image according to a predetermined rule.

[0076] The vector error calculation means 44 calculates a vector error based on the visual saliency map output by the visual saliency extraction means 3 and the ideal line of sight set by the line-of-sight coordinate setting means 43 for the visual saliency map and the image, and calculates a visual attention concentration degree Ps, which will be described later, indicating the concentration degree of visual attention based on the vector error. That is, the vector error calculation means 44 functions as a visual attention concentration degree calculation unit that calculates the concentration degree of visual attention in the image based on the visual saliency distribution information and the line-of-sight position.

[0077] Here, the vector error in this embodiment will be described with reference to FIG. 13. FIG. 13 shows an example of a visual saliency map. This visual saliency map is represented by luminance values of 256 gradations of H pixels × V pixels, and the higher the visual saliency, the higher the luminance is displayed, similar to FIG. 3. In FIG. 13, the coordinates of the ideal line of sight are (x, y) = (x im , y imWhen [the above is done], the vector error with the pixel at any coordinate (k, m) in the visual saliency map is calculated. When the coordinates with high luminance in the visual saliency map are far from the coordinates of the ideal line of sight, it means that the position to be fixated and the position where fixation is actually likely to occur are far apart, and it can be said that the image is one where visual attention is likely to be scattered. On the other hand, when the coordinates with high luminance and the coordinates of the ideal line of sight are close, it means that the position to be fixated and the position where fixation is actually likely to occur are close, and it can be said that the image is one where visual attention is likely to be concentrated on the position to be fixated.

[0078] Next, a method for calculating the visual attention concentration degree Ps in the vector error calculation means 44 will be described. In this embodiment, the visual attention concentration degree Ps is calculated by the following formula (1).

Equation

[0079] In formula (1), V vc is the pixel depth (luminance value), f w is the weighting function, and d err indicates the vector error. This weighting function is a function whose weights are set based on, for example, the distance from the pixel indicating the value of V vc to the coordinates of the ideal line of sight. α is a coefficient such that the visual attention concentration degree Ps becomes 1 when the coordinates of the highlight and the coordinates of the ideal line of sight coincide in the visual saliency map (reference heat map) of one highlight point.

[0080] That is, the vector error calculation means 44 (visual attention concentration calculation unit) calculates the concentration degree of visual attention based on the value of each pixel constituting the visual saliency map (visual saliency distribution information) and the vector error between the position of each pixel and the coordinate position of the ideal line of sight (reference line of sight position).

[0081] The visually focused attention degree Ps obtained in this way is the reciprocal of the sum obtained by weighting the relationship between the vector error of the coordinates of all pixels from the coordinates of the ideal line of sight set on the visual saliency map and the luminance value. This visually focused attention degree Ps is calculated to have a low value when the distribution of high luminance on the visual saliency map is far from the coordinates of the ideal line of sight. That is to say, the visually focused attention degree Ps can also be said to be the degree of concentration on the ideal line of sight.

[0082] FIG. 14 shows an example of an image input to the input means 2 and a visual saliency map obtained from the image. FIG. 14(a) is the input image, and (b) is the visual saliency map. In such FIG. 14, when the coordinates of the ideal line of sight are set, for example, on a road such as a track traveling forward, the visually focused attention degree Ps in that case is calculated.

[0083] The output means 45 outputs information regarding the risk of the scene shown in the image for which the visually focused attention degree Ps was calculated based on the visually focused attention degree Ps calculated by the vector error calculation means 44. As information regarding the risk, for example, a predetermined threshold is set for the visually focused attention degree Ps, and when the calculated visually focused attention degree Ps is equal to or lower than the threshold, information indicating that it is a scene with a high risk is output. For example, when the visually focused attention degree Ps calculated in FIG. 10 is equal to or lower than the threshold, it can be determined that it is a scene with a high risk, and information such as "risk present (or high risk)" can be output.

[0084] Also, the output means 45 may output information regarding the risk based on the temporal change of the visually focused attention degree Ps calculated by the vector error calculation means 44. FIG. 15 shows an example of the temporal change of the visually focused attention degree Ps. FIG. 15 shows the change of the visually focused attention degree Ps in a moving image of 12 seconds. In FIG. 15, the visually focused attention degree Ps changes rapidly between about 6.5 seconds and about 7 seconds. This is, for example, the case when another vehicle cuts in front of the host vehicle.

[0085] As shown in FIG. 15, by comparing the change rate and change value of the visual attention concentration Ps per unit time with a predetermined threshold value, it can be determined that the scene has a high risk, and information indicating the presence of risk (or high risk) may be output. Further, for example, the presence or absence (high or low) of risk may be determined based on a pattern of change such as the visual attention concentration Ps once decreased increasing.

[0086] And in this embodiment, the adding means 5 registers (adds) the point or section shown in the processed image as an attention point (a point that requires attention) to the map data when the information regarding the risk output from the output means 45 includes information indicating the presence of risk, for example. Note that the example of the map is the same as that in FIG. 11.

[0087] According to this embodiment, the line-of-sight coordinate setting means 43 sets the coordinates of the ideal line of sight at a predetermined fixed position. Then, the vector error calculation means 44 calculates the visual attention concentration Ps in the image based on the visual saliency map and the ideal line of sight. By doing so, since the visual saliency map is used, it is possible to reflect the contextual attention state of what objects such as signs and pedestrians included in the image are. Therefore, it is possible to accurately calculate the visual attention concentration Ps. And the risk points based on the visual attention concentration Ps calculated in this way can be added to the map data.

[0088] Further, the vector error calculation means 44 calculates the visual attention concentration Ps based on the value of each pixel constituting the visual saliency map and the vector error between the position of each pixel and the coordinate position of the ideal line of sight. By doing so, a value corresponding to the difference between the position with high visual saliency and the ideal line of sight is calculated as the visual attention concentration Ps. Therefore, for example, the value of the visual attention concentration Ps can be changed according to the distance between the position with high visual saliency and the ideal line of sight.

[0089] In addition, an output means 45 is provided which outputs risk information at the location shown in the image based on the temporal change of the visual attention concentration Ps. By doing so, for example, it becomes possible to output a location where the temporal change of the visual attention concentration Ps is large as an accident risk location or the like.

Embodiment

[0090] Next, a map data generation device according to the third embodiment of the present invention will be described with reference to FIGS. 16 to 20. Note that the same parts as those in the above-described first and second embodiments are denoted by the same reference numerals and the description thereof is omitted.

[0091] This embodiment is a modification of the second embodiment, and the method of calculating the visual attention concentration is the same. The image input from the input means 2 in this embodiment is an image entering an intersection, and the method of determining risk in the output means 45 and the like are different.

[0092] An example of an intersection for which risk information is output in this embodiment is shown in FIG. 16. FIG. 16 is an intersection constituting a four-way intersection (crossroads). In this intersection, images of the intersection directions (travel directions) when entering from the A direction, B direction, and C direction are respectively shown. That is, the images shown in FIG. 16 are images when the A direction, B direction, and C direction are taken as approach roads which are the roads when entering the intersection.

[0093] For the images shown in FIG. 16, visual saliency maps are respectively obtained as described in the previous embodiments. Then, for each image, ideal lines of sight are set for each of the straight-ahead, right-turn, and left-turn travel directions, and the visual attention concentration Ps is calculated for each of the ideal lines of sight (FIG. 17). That is, for each exit road which is the road to exit after entering the intersection, the ideal line of sight (reference line-of-sight position) in the image is set respectively, and the vector error calculation means 44 calculates the visual attention concentration Ps with respect to each ideal line of sight.

[0094] Here, the temporal change in the visual attention concentration Ps when entering the intersection from each approach road is shown in FIG. 18. Such a temporal change is based on the calculation result of the vector error calculation means 44. The graph in FIG. 18 shows the visual attention concentration Ps on the vertical axis and time on the horizontal axis. The thick line shows the case where the ideal line of sight is set in the straight-ahead direction, the thin line shows the left-turn direction, and the broken line shows the right-turn direction. And FIG. 18(a) shows the case of entering from direction A, FIG. 18(b) shows the case of entering from direction B, and FIG. 18(c) shows the case of entering from direction C.

[0095] According to FIG. 18, when approaching the intersection, the visual attention concentration Ps tends to decrease, but there are cases where it decreases rapidly near the intersection as shown in FIG. 18(b). Also, according to FIG. 18, the visual attention concentration Ps when assuming that the line of sight is directed to either the left or right for a right turn or a left turn is lower than the visual attention concentration Ps when assuming that the front is looked straight ahead for going straight.

[0096] Next, the output means 45 calculates the ratio of the visual attention concentration Ps in the right or left direction to the visual attention concentration Ps in the straight-ahead direction using the temporal change in the visual attention concentration Ps calculated in FIG. 18. The change in the calculated ratio is shown in FIG. 19. The graph in FIG. 19 shows the ratio on the vertical axis and time on the horizontal axis. The thick line shows the left-turn / straight-ahead ratio (L / C), and the thin line shows the right-turn / straight-ahead ratio (R / C). And FIG. 19(a) shows the case of entering from direction A, FIG. 19(b) shows the case of entering from direction B, and FIG. 19(c) shows the case of entering from direction C. For example, I in FIG. 19(a) AL is PS LA (visual attention concentration in the left direction of direction A) / PS CA (visual attention concentration in the straight-ahead direction of direction A), and I AR is PS RA (visual attention concentration in the right direction of direction A) / PS CA (visual attention concentration in the straight-ahead direction of direction A). I in FIG. 19(b) BL , I BR , I in FIG. 19(c) CL , I CR have the same meaning except that the approach directions are different.

[0097] According to FIG. 19, when the ratio of visual attention concentration such as I AL or I AR is less than 1, it can be said that at intersections where the driver's concentration (= visual attention concentration Ps) drops more when turning right or left than when moving the line of sight to go straight. Conversely, when the ratio is greater than 1, it can be said that it represents an intersection where the driver's concentration drops when going straight.

[0098] Therefore, the output means 45 can determine the risk state of the target intersection based on the temporal change of the visual attention concentration Ps and the ratio of the visual attention concentration Ps as described above, and output the determination result as information regarding the risk. Then, based on the information regarding the risk, the attention point can be added to the map data.

[0099] According to the present embodiment, the line-of-sight coordinate setting means 43 sets the coordinates of the ideal line of sight in the image for each exit road that becomes the road to exit after entering the intersection for the visual saliency map. Then, the visual attention concentration Ps for each exit road in the image is calculated based on the vector error calculation means 44, the visual saliency map, and the ideal line of sight, and the output means 45 outputs risk information at the intersection based on the visual attention concentration Ps calculated for each exit road. By doing so, the risk of the target intersection can be evaluated, the risk information can be output, and added to the map data.

[0100] Further, the output means 45 outputs risk information based on the ratio between the visual attention concentration Ps of the straight-ahead exit road and the visual attention concentration Ps of the right-turn or left-turn exit road among the exit roads. By doing so, it is possible to evaluate which of going straight and turning right or left is more likely to attract attention and output the evaluation result.

[0101] Further, the output means 45 may output risk information based on the temporal change of the visual attention concentration Ps. By doing so, for example, when the visual attention concentration Ps changes rapidly, etc., it is possible to detect and output risk information.

[0102] In the third embodiment, for example, in FIG. 17, when entering the intersection from the A direction, the visual attention concentration Ps for a right turn (heading in the B direction) decreases. When entering the intersection from the B direction, the visual attention concentration Ps for a right or left turn (heading in the A or C direction) decreases. When entering the intersection from the C direction, the visual attention concentration Ps for a right turn decreases.

[0103] In this case, for example, the route of turning right from the A direction and the route of turning left from the B direction are both routes where the visual attention concentration Ps decreases more than other routes, and when the entry route and the exit route are swapped, they become the same route. Therefore, at this intersection, information regarding risks such as "risky (or high risk)" may be output for this route.

[0104] Also, although the third embodiment has been described for an intersection, this concept can also be applied to a curve on a road. This will be described with reference to FIG. 20.

[0105] FIG. 20 is an example of a curved road. This road may be traveled as a left curve from the D direction (lower side of the figure) or as a right curve from the E direction (left side of the figure). Here, for example, when entering the curve from the D direction, not only an ideal line of sight is set in the left direction, which is the bending direction of the road, but also an ideal line of sight is set in the direction (D' direction) assuming the road was going straight, and the visual attention concentration Ps is calculated for each. Similarly, when entering the curve from the E direction, not only an ideal line of sight is set in the right direction, which is the bending direction of the road, but also an ideal line of sight is set in the direction (E' direction) assuming the road was going straight, and the visual attention concentration Ps is calculated for each.

[0106] Then, based on the calculated visual attention concentration Ps, the risk may be determined based on time-series changes, ratios, etc., similar to intersections.

[0107] In addition, when the curvature of the curve is large as shown in Fig. 20, not only in the straight-ahead direction, but also a virtual ideal line of sight may be set in the direction opposite to the bending direction of the curve. In the case of Fig. 20, when entering from the D direction, the ideal line of sight may be set not only in the D' direction but also in the E' direction to calculate the visual attention concentration Ps. That is, the ideal line of sight may be set in a direction different from the bending direction of the curve.

[0108] That is, the visual saliency extraction means 3 acquires a visual saliency map obtained by estimating the level of visual saliency in the image from the image when entering a curve on the road. The line-of-sight coordinate setting means 43 sets the coordinates of the ideal line of sight in the image in the bending direction of the curve and in a direction different from the bending direction for the visual saliency map. Then, the vector error calculation means 44 calculates the visual attention concentration Ps in the direction different from the bending direction in the image based on the visual saliency map and the ideal line of sight, and the output means 45 outputs the risk information in the curve based on the visual attention concentration Ps calculated for each exit route.

[0109] By doing so, it is possible to evaluate the risk for the target curve and output information regarding the risk.

Example

[0110] Next, a map data generation device according to the fourth embodiment of the present invention will be described with reference to Figs. 21 to 25. Note that the same parts as those in the first to third embodiments described above are denoted by the same reference numerals and the description thereof is omitted.

[0111] In this embodiment, it detects the tendency of distraction and adds visually attention-required points, etc. to the map data based on the tendency of distraction. As shown in Fig. 21, the analysis means 4 according to this embodiment includes a visual saliency peak detection means 46 and a distraction tendency determination means 47.

[0112] The visual saliency peak detection means 46 detects the position (pixel) that becomes a peak in the visual saliency map acquired by the visual saliency extraction means 3. Here, in this embodiment, a peak is a pixel with high visual saliency where the pixel value is the maximum value (the luminance is the maximum), and the position is represented by coordinates. That is, the visual saliency peak detection means 46 functions as a peak position detection unit that detects at least one peak position in the visual saliency map (visual saliency distribution information).

[0113] The peripheral view tendency determination means 47 determines whether the image input from the input means 2 has a tendency of peripheral view based on the position that becomes a peak detected by the visual saliency peak detection means 46. The peripheral view tendency determination means 47 first sets a fixation area (area to be fixated) for the image input from the input means 2. The method of setting the fixation area will be described with reference to FIG. 22. That is, the peripheral view tendency determination means 47 functions as a fixation range setting unit that sets the range to be fixated by the driver of the moving object in the image.

[0114] In the image P shown in FIG. 22, the fixation area G is set around the vanishing point V. That is, the fixation area G (area to be fixated) is set based on the vanishing point of the image. This fixation area G has a size (for example, width 3 m, height 2 m) set in advance, and it is possible to calculate the number of pixels of the set size from the number of horizontal pixels, the number of vertical pixels, the horizontal field of view, the vertical field of view, the inter-vehicle distance to the preceding vehicle, the mounting height of a camera such as a drive recorder that captures the image, etc. of the image P. Note that the vanishing point may be estimated from a white line or the like, or may be estimated using optical flow or the like. Also, the inter-vehicle distance to the preceding vehicle does not need to detect an actual preceding vehicle and may be set virtually.

[0115] Next, a peripheral vision detection area in the image P is set based on the set fixation area G (the shaded area in Fig. 23). In this peripheral vision detection area, an upper area Iu, a lower area Id, a left lateral area Il, and a right lateral area Ir are set respectively. These areas are divided by line segments connecting the vanishing point V and the vertices of the fixation area G. That is, the upper area Iu and the left lateral area Il are separated by a line segment L1 connecting the vanishing point V and the vertex Ga of the fixation area G. The upper area Iu and the right lateral area Ir are separated by a line segment L2 connecting the vanishing point V and the vertex Gd of the fixation area G. The lower area Id and the left lateral area Il are separated by a line segment L3 connecting the vanishing point V and the vertex Gb of the fixation area G. The lower area Id and the right lateral area Ir are separated by a line segment L4 connecting the vanishing point V and the vertex Gc of the fixation area G.

[0116] Note that the division of the peripheral vision detection area is not limited to the division as shown in Fig. 23. For example, it may be as shown in Fig. 24. Fig. 24 divides the peripheral vision detection area by line segments obtained by extending each side of the fixation area G. Since the method in Fig. 24 has a simpler shape, the processing related to the division of the peripheral vision detection area can be reduced.

[0117] Next, the determination of the peripheral vision tendency in the peripheral vision tendency determination means 47 will be described. When the peak position detected by the visual saliency peak detection means 46 has continuously deviated from the fixation area G for a predetermined time or more, it is determined that there is a peripheral vision tendency. Here, the predetermined time can be, for example, 2 seconds, but it may be changed as appropriate. That is, the peripheral vision tendency determination means 47 determines whether the peak position has continuously deviated from the range to be fixated for a predetermined time or more.

[0118] Further, when the peeping detection area is the upper area Iu or the lower area Id, the peeping tendency determination means 47 may determine that there is a tendency of peeping by a stationary object. In the case of an image obtained by imaging the front from the vehicle, generally, stationary objects such as buildings, traffic signals, signs, and streetlights are reflected in the upper area Iu, and road paint such as road signs is generally reflected in the lower area Id. On the other hand, in the left side area Il and the right side area Ir, moving objects other than the host vehicle, such as vehicles traveling in other driving lanes, may be reflected, and it is difficult to determine whether the peeping target object is a stationary object or a moving object only based on the area.

[0119] When the peak position is in the left side area Il or the right side area Ir, since it is impossible to determine whether the peeping target object is a stationary object or a moving object only based on the area, the determination is made using object recognition. Object recognition (also referred to as object detection) may use a well-known algorithm, and the specific method is not particularly limited.

[0120] Further, not limited to object recognition, the determination of whether it is a stationary object or a moving object may be made using the relative speed. This is to obtain the relative speed from the vehicle speed and the moving speed between the frames of the peeping target object, and to determine whether the peeping target object is a stationary object or a moving object from the relative speed. Here, the moving speed between the frames of the peeping target object may be obtained by obtaining the moving speed between the frames of the peak position. When the obtained relative speed is equal to or greater than a predetermined threshold value, it can be determined that the object is fixed at a certain position.

[0121] Next, the operation of the map data generation device 1 of the present embodiment will be described with reference to the flowchart of FIG. 25.

[0122] First, the input means 2 acquires a driving image (step S104), and the visual saliency extraction means 3 performs visual saliency image processing (acquisition of a visual saliency map) (step S105). Then, the visual saliency peak detection means 46 acquires (detects) the peak position based on the visual saliency map acquired by the visual saliency extraction means 3 in step S105 (step S106).

[0123] Next, the peripheral view tendency determination means 47 sets the fixation area G and compares the fixation area G with the peak position acquired by the visual saliency peak detection means 46 (step S107). If as a result of the comparison the peak position is outside the fixation area G (step S107; outside the fixation area), the peripheral view tendency determination means 47 determines whether the residence timer has started or is stopped (step S108). The residence timer is a timer that measures the time during which the peak position stays outside the fixation area G. Note that the setting of the fixation area G may be performed when the image is acquired in step S104.

[0124] If the residence timer is stopped (step S108; stopped), the peripheral view tendency determination means 47 starts the residence timer (step S109). On the other hand, if the residence timer has started (step S108; after start), the peripheral view tendency determination means 47 compares with the residence timer threshold value (step S110). The residence timer threshold value is the threshold value of the time during which the peak position stays outside the fixation area G, and is set to 2 seconds or the like as described above.

[0125] If the residence timer has exceeded the threshold value (step S110; exceeded the threshold value), the peripheral view tendency determination means 47 determines the point corresponding to the image to be determined as a caution point (a point that requires visual attention) corresponding to a point or section having a peripheral view tendency, and the addition means 5 registers (adds) it to the map data according to the determination (step S111). Note that an example of the map data is the same as that in FIG. 11.

[0126] On the other hand, if the residence timer does not exceed the threshold value (step S110; does not exceed the threshold value), the peripheral view tendency determination means 47 returns to step S101 without doing anything.

[0127] Also, if as a result of the comparison in step S107 the peak position is within the fixation area G (step S107; within the fixation area), the peripheral view tendency determination means 47 stops the residence timer (step S112).

[0128] According to this embodiment, the visual saliency peak detection means 46 detects at least one peak position in the visual saliency map in time series. Then, the peripheral view tendency determination means 47 sets the fixation area G in the image, and when the peak position has been continuously outside the fixation area G for a predetermined time or more, outputs information indicating that there is a tendency of peripheral view, and based on this information, the addition means 5 adds it to the map data as a point that requires attention, etc. This visual saliency map shows the statistical ease of human gaze aggregation. Therefore, by using the visual saliency map, it is possible to detect the tendency of peripheral view with a simple configuration and add it to the map data without measuring the actual driver's gaze.

[0129] Further, the peripheral view tendency determination means 47 sets the fixation area G based on the vanishing point V of the image. By doing so, it becomes possible to easily set the fixation area G without detecting, for example, a preceding vehicle or the like.

[0130] Further, when the peak position has been continuously located above or below the fixation area G for a predetermined time or more, the peripheral view warning unit 6 may output information indicating that there is a tendency of peripheral view due to a fixed object. Above the fixation area G is generally an area where fixed objects such as buildings, traffic signals, signs, and streetlights are reflected, and below the fixation area G is generally an area where road paint such as road signs is reflected. That is, when the peak position is included in the range, it is possible to specify that the object of peripheral view due to peripheral view is a fixed object.

[0131] Note that the attention area G is not necessarily set to a fixed range. For example, it may be changed according to the moving speed of the moving body. For example, it is known that the driver's field of vision becomes narrower during high-speed driving. Therefore, for example, the glance tendency determination means 47 may acquire the vehicle speed from a speed sensor or the like mounted on the vehicle, and narrow the range of the attention area G as the speed increases. Further, since an appropriate inter-vehicle distance also changes according to the moving speed, the range of the attention area G by the calculation method described with reference to FIG. 22 may also be changed. The speed of the vehicle is not limited to the speed sensor, and may be obtained from an acceleration sensor or a captured image.

[0132] Further, the attention area G may be changed according to the driving position and situation of a vehicle or the like. In a situation where attention to the surroundings is required, it is necessary to widen the attention area G. For example, the range to be noted changes depending on the driving position such as a residential area, an arterial road, or a busy street. In a residential area, there are few pedestrians, but it is necessary to pay attention to sudden jumps out, and the attention area G cannot be narrowed. On the other hand, on an arterial road, the driving speed becomes high and the field of vision becomes narrow as described above.

[0133] Specific examples are as follows. There is a risk of children jumping out on the way to school, in parks, and near schools. There are many pedestrians near stations, schools, event locations, and tourist spots. There are many bicycles near bicycle parking lots and schools. There are many drunk people near entertainment districts. Locations such as the above are situations where attention to the surroundings is required, and the attention area G may be widened and the area determined as the glance tendency may be narrowed. On the other hand, when driving on a highway or in an area with low traffic volume and population density, the driving speed tends to be high, and the attention area G may be narrowed and the area determined as the glance tendency may be widened.

[0134] Also, the attention area G may be changed according to time zones, events, etc. For example, during the commuting hours, it is a situation where attention to the surroundings is necessary, and the attention area G may be widened compared to normal hours, and the area determined to have a side glance tendency may be narrowed. Alternatively, similarly, the attention area G may be widened and the area determined to have a side glance tendency may be narrowed from dusk to night. On the other hand, at midnight, the attention area G may be narrowed and the area determined to have a side glance tendency may be widened.

[0135] Furthermore, the attention area G may be changed according to event information. For example, since events and the like are places and time zones with a lot of human traffic, the attention area G may be widened more than usual to relax the determination of the side glance tendency.

[0136] Such location information can be obtained by the side glance tendency determination means 47 from means capable of discriminating the current position such as a GPS receiver and map data and the area where the vehicle is traveling, and associating it with the image data, so that the range of the attention area G can be changed. The time information may be obtained by the information output device 1 from inside or outside. The event information may be obtained from an external site or the like. Also, the determination of the change may be made by combining the position, time, and date, or the determination of the change may be made using any one of them.

[0137] Furthermore, when traveling at high speed, the dwell timer threshold value may be shortened. This is because when traveling at high speed, even a short side glance can put the vehicle in a dangerous state.

Example

[0138] Next, the map data generation device according to the fifth embodiment of the present invention will be described with reference to FIG. 26. Note that the same parts as those in the first to fourth embodiments described above are denoted by the same reference numerals and the description thereof is omitted.

[0139] In this embodiment, a monotonous road is detected, and based on the detection result, visually attention-requiring points and the like are added to the map data. As described above, the analysis means 4 according to this embodiment determines a monotonous road (monotony tendency). A monotonous road generally refers to a road with no change in scenery or scarce scenery changes, a road with street lights installed regularly at equal intervals, or a road with a monotonous landscape such as a highway where a monotonous scenery continues.

[0140] Based on the visual saliency map acquired by the visual saliency extraction means 3, the analysis means 4 determines whether the image input to the input means 2 has a monotony tendency. In this embodiment, various statistical quantities are calculated from the visual saliency map, and based on these statistical quantities, it is determined whether there is a monotony tendency. That is, the analysis means 4 functions as a monotony determination unit that determines whether the image has a monotony tendency using the statistical quantities calculated based on the visual saliency map (visual saliency distribution information).

[0141] Fig. 26 shows a flowchart of the operation of the analysis means 4. First, the standard deviation of the luminance of each pixel in the image (for example, Fig. 3(b)) constituting the visual saliency map is calculated (step S51). In this step, first, the average value of the luminance of each image in the image constituting the visual saliency map is calculated. If the image constituting the visual saliency map is H pixels × V pixels and the luminance value at an arbitrary coordinate (k, m) is V VC(k,m) then the average value is calculated by the following formula (2).

Equation

[0142] From the average value calculated by formula (2), the standard deviation of the luminance of each image in the image constituting the visual saliency map is calculated. The standard deviation SDEV is calculated by the following formula (3).

Equation

[0143] Determine whether there are multiple output results for the standard deviation calculated in step S51 (step S52). In this step, the image input from the input means 2 is a moving image, and the visual saliency map is acquired for each frame, and in step S51, it is determined whether the standard deviation for a plurality of frames has been calculated.

[0144] If there are multiple output results (step S52; Yes), calculate the amount of eye movement (step S53). In this embodiment, the amount of eye movement is obtained based on the coordinate distance of the maximum (highest) luminance value in the visual saliency maps of the frames before and after in time. The amount of eye movement VSA is calculated by the following formula (4) where the coordinates of the maximum luminance value in the previous frame are (x1, y1) and the coordinates of the maximum luminance value in the subsequent frame are (x2, y2).

Equation

[0145] Then, determine whether there is a monotonic trend based on the standard deviation calculated in step S11 and the amount of eye movement calculated in step S13 (step S14). In this step, if step S12 is No, a threshold is set for the standard deviation calculated in step S11, and by comparing it with the threshold, it can be determined whether there is a monotonic trend. On the other hand, if step S12 is Yes, a threshold is set for the amount of eye movement calculated in step S13, and by comparing it with the threshold, it can be determined whether there is a monotonic trend.

[0146] That is, the analysis means 4 functions as a standard deviation calculation unit that calculates the standard deviation of the luminance of each pixel in the image obtained as the visual saliency map (visual saliency distribution information), and functions as an eye movement amount calculation unit that calculates the amount of eye movement between frames based on the images obtained in time series as the visual saliency map (visual saliency distribution information).

[0147] In addition, when the result (judgment result) of the processing by the analysis means 4 is judged to have a monotonic tendency, the judgment result is output to the addition means 5. Then, the addition means 5 registers (adds) the point or section indicated by the image judged to have a monotonic tendency as a point of attention (a point that requires attention) in the map data. Note that the example of the map is the same as that in FIG. 11.

[0148] According to this embodiment, the analysis means 4 determines whether the image has a monotonic tendency based on the standard deviation, the amount of line-of-sight movement, etc. calculated based on the visual saliency map. By doing so, it becomes possible to determine whether there is a monotonic tendency from the captured image based on the position where a human is likely to fixate. Since the determination is made based on the position where a human (driver) is likely to fixate, it is possible to make a determination in a tendency close to what the driver feels is monotonous, make a more accurate determination, and add it to the map data based on the determination result.

[0149] Further, the analysis means 4 may calculate the average value of the luminance of each pixel in the image obtained as the visual saliency map, and determine whether the image has a monotonic tendency based on the calculated average value. By doing so, in one image, it is possible to determine that there is a monotonic tendency when the positions where fixation is likely are concentrated. In addition, since the determination is made using the average value, the arithmetic processing can be simplified.

[0150] Further, the analysis means 4 calculates the amount of line-of-sight movement between frames based on the images obtained in time series as the visual saliency map, and determines whether there is a monotonic tendency based on the calculated amount of line-of-sight movement. By doing so, when determining for a moving image, for example, when the amount of line-of-sight movement is small, it can be determined that there is a monotonic tendency.

Example

[0151] Next, a map data generation device according to a sixth embodiment of the present invention will be described with reference to FIGS. 27 and 28. Note that the same parts as those in the above-described first to fifth embodiments are denoted by the same reference numerals and the description thereof is omitted.

[0152] This embodiment is to enable the determination of a monotonic trend even in the case of detection omission in the method of the fifth embodiment, particularly when there are multiple output results (video). The block configuration and the like are the same as those of the fifth embodiment. A flowchart of the operation of the analysis means 4 according to this embodiment is shown in FIG. 27.

[0153] In the flowchart of FIG. 27, steps S51 and S53 are the same as those in FIG. 26. In this embodiment, since autocorrelation is used as described later, and the target image is a moving image, step S52 is omitted. The determination content of step S54A is the same as that of step S54. In this embodiment, step S54A is performed as a primary determination of the monotonic trend.

[0154] Next, if it is determined that there is a monotonic trend as a result of the determination in step S54A (step S55; Yes), the determination result is output to the outside of the determination device 1 in the same manner as in FIG. 26. On the other hand, if it is determined that there is no monotonic trend as a result of the determination in step S54A (step S55; No), autocorrelation calculation is performed (step S56).

[0155] In this embodiment, autocorrelation is calculated using the standard deviation (average luminance value) and the amount of gaze movement calculated in steps S51 and S53. The autocorrelation R (k) is known to be calculated by the following equation (5) with the expected value being E, the average of X being μ, the variance of X being σ2, and the lag being k. In this embodiment, k is changed within a predetermined range to perform the calculation of equation (5), and the largest calculated value is used as the autocorrelation value.

Equation

[0156] Then, it is determined whether there is a monotonic trend based on the calculated autocorrelation value (step S57). The determination may be made by setting a threshold value for the autocorrelation value and comparing it with the threshold value in the same manner as in the fifth embodiment to determine whether there is a monotonic trend. For example, if the autocorrelation value at k = k1 is greater than the threshold value, it means that the same scenery is repeated every k1. When it is determined that there is a monotonic trend, the scenery image is classified as an image with a monotonic trend. By calculating such an autocorrelation value, it becomes possible to determine a road with a monotonic trend caused by periodically arranged objects such as streetlights regularly installed at equal intervals.

[0157] Fig. 28 shows an example of the calculation result of autocorrelation. Fig. 28 is a correlogram of the luminance average value of the visual saliency map for the driving video. The vertical axis in Fig. 28 represents the correlation function (autocorrelation value), and the horizontal axis represents the lag. Also, in Fig. 28, the shaded part is the confidence interval 95% (significance level αs = 0.05). Assuming the null hypothesis is "there is no periodicity at lag k" and the alternative hypothesis is "there is periodicity at lag k", the data within this shaded part cannot reject the null hypothesis, so it is determined that there is no periodicity, and those exceeding the shaded part are determined to have periodicity regardless of positive or negative.

[0158] Fig. 28(a) is a video of driving in a tunnel, which is an example of having periodicity. According to Fig. 28(a), it can be seen that periodicity is observed at the 10th and 17th. In the case of a tunnel, since the tunnel lighting is arranged at regular intervals, it is possible to determine the monotonic trend caused by the lighting and the like. On the other hand, Fig. 28(b) is a video of driving on a general road, which is an example of having no periodicity. According to Fig. 28(b), it can be seen that most of the data is within the confidence interval.

[0159] By operating as in the flowchart of Fig. 27, first, it is determined whether it is monotonic based on the average, standard deviation, and line-of-sight movement amount, and then a secondary determination can be made from the perspective of periodicity among those that are missed.

[0160] According to this embodiment, a visual saliency map is acquired in a time series. The analysis means 4 functions as a primary determination unit that calculates a statistic from the visual saliency map acquired in the time series and determines whether there is a monotonic trend based on the statistic obtained in the time series, and a secondary determination unit that determines whether there is a monotonic trend based on autocorrelation. By doing so, it is possible to determine a monotonic trend caused by an object that appears periodically, such as a street lamp that appears during driving and is difficult to determine only by the statistic, by autocorrelation.

[0161] In addition, the above-described first to sixth embodiments may be combined. That is, information of a plurality of embodiments may be simultaneously displayed on the map shown in FIG. 11. When displaying them simultaneously, it is preferable to change the color, shape, etc. so that it can be determined which attention point it is.

[0162] Further, the present invention is not limited to the above embodiments. That is, those skilled in the art can variously modify and implement it according to the conventionally known knowledge without departing from the gist of the present invention. As long as the map data generation device of the present invention is still provided even by such a modification, of course, it is included in the scope of the present invention.

Explanation of Reference Numerals

[0163] 1 Map data generation device 2 Input means (acquisition unit) 3 Visual saliency extraction means (generation unit) 4 Analysis means (analysis unit, monotonic determination unit, fourth determination unit) 5 Addition means (addition unit) 41 Visual load amount calculation means (movement amount calculation unit) 42 Visual load determination means (first determination unit) 43 Line-of-sight coordinate setting means (line-of-sight position setting unit) 44 Visual attention concentration calculation means (visual attention concentration calculation unit) 45 Output means (second determination unit) 46 Visual saliency peak detection means (peak position detection unit, fixation range setting unit) 47 Glancing tendency determination unit (glancing output unit, third determination unit)

Claims

【Claim 1】 A map data generation method executed by a map data generation device that adds predetermined information to map data based on an image captured from the outside by a moving body, comprising: a first acquisition step of acquiring input information in which the image and the position information of the moving body are associated; a second acquisition unit that acquires visual saliency distribution information obtained by estimating the level of visual saliency within the image based on the image; an analysis step of analyzing, based on the visual saliency distribution information, whether the point or section indicated by the position information corresponding to the visual saliency distribution information is a point or section that requires visual attention; an addition step of adding the visually attention-requiring point or section to the map data based on the analysis result of the analysis step; The map data generation method characterized by including the above.

Citation Information

Patent Citations

  • Device and method for generating an indication signal to the driver of a vehicle

    EP2511121A1

  • Navigation device for vehicle

    JP2006258656A