Saliency estimation device
The saliency estimation device corrects image brightness based on movement speed and relative positions to accurately detect salient regions in images viewed by a moving person, addressing the limitations of existing technologies.
Patent Information
- Application Number
- JP2025025674
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to accurately detect salient regions in images viewed by a moving person, as they do not effectively account for the changing field of view and brightness with movement speed.
A saliency estimation device that corrects image brightness using speed information and relative positions, generating salience estimation information by processing the corrected image to indicate salience distribution.
The device accurately detects salient regions in images viewed by a moving person, improving the accuracy of salience estimation by accounting for the effects of movement speed on the field of view and image brightness.
Smart Images

Figure 2025071235000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a saliency estimation device, a saliency estimation method, and a program. [Background technology]
[0002] A technology has been proposed for automatically detecting salient regions in an image. On the other hand, when a person is moving, the effective visual field of the person narrows as the person's moving speed increases. Non-Patent Document 1 describes automatic detection of salient regions taking the effective visual field into consideration. Specifically, Non-Patent Document 1 describes a method of reducing the resolution and saturation of a target image according to the distance from the gaze point, and then estimating saliency. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Morimoto et al., "Saliency Estimation Model for Moving Images Considering Response Characteristics of the MST Area," Institute of Image Television Engineers Technical Report, Vol. 38, No. 10, pp. 57-60, 2014. Summary of the Invention [Problem to be solved by the invention]
[0004] The present inventor has studied a method for detecting, with high accuracy, an area from an image that is perceived as highly salient when seen by a person on the move. One example of a problem to be solved by the present invention is to detect, with high accuracy, an area from an image that is perceived as highly salient when seen by a person on the move. [Means for solving the problem]
[0005] One example of the invention according to the present disclosure includes a correction unit that generates a corrected image by correcting an image of a scene seen from a first viewpoint; a saliency estimation unit that processes the corrected image to generate saliency estimation information indicative of a saliency distribution in the corrected image or in the image; Equipped with The correction unit is generating brightness information indicative of a change in brightness of at least a portion of the image using speed information relating to a moving speed of the first viewpoint and a relative position of the at least a portion from a reference point in the image; The saliency estimation device corrects the at least part of the brightness using the brightness information.
[0006] In one embodiment of the present disclosure, a computer generating a corrected image by correcting an image of the scene from a first viewpoint; processing the corrected image to generate saliency estimation information indicative of a saliency distribution in or within the corrected image; The computer further comprises: generating brightness information indicative of a change in brightness of at least a portion of the image using speed information relating to a moving speed of the first viewpoint and a relative position of the at least a portion from a reference point in the image; The saliency estimation method corrects the luminance of at least a portion of the image using the luminance information.
[0007] In one embodiment of the present disclosure, a computer includes: a correction function for generating a corrected image by correcting an image of the scene seen from the first viewpoint; an estimation function for generating saliency estimation information indicative of a saliency distribution in the corrected image or in the image by processing the corrected image; Let them have Furthermore, as at least a part of the correction function, generating brightness information indicating a change in brightness of at least a portion of the image using speed information relating to a moving speed of the first viewpoint and a relative position from a reference point in the image to the at least a portion of the image; a function of correcting the at least a portion of the brightness using the brightness information; It is a program that allows people to have the following: [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating a functional configuration of a saliency estimation device according to a first embodiment. [Diagram 2] FIG. 11 is a diagram for explaining a method for setting viewing angle information. [Diagram 3] FIG. 11 is a diagram for explaining an example of viewing angle information. [Figure 4] FIG. 11 is a diagram for explaining brightness information. [Diagram 5] FIG. 11 is a diagram for explaining resolution information. [Figure 6] FIG. 11 is a diagram for explaining saturation information. [Figure 7] FIG. 4 illustrates an example of a functional configuration of a correction processing unit. [Figure 8] 10 is a block diagram illustrating an example of the configuration of a saliency estimation unit; FIG. [Figure 9] FIG. 2A is a diagram illustrating an example of an image to be input to a saliency estimation unit, and FIG. 2B is a diagram illustrating an example of an image showing a saliency distribution estimated for FIG. [Figure 10] 4 is a flowchart illustrating a processing method according to the first configuration example. [Figure 11] FIG. 2 is a diagram illustrating in detail an example of the configuration of a nonlinear mapping unit. [Figure 12] FIG. 2 is a diagram illustrating a configuration of an intermediate layer. [Figure 13] 13A and 13B are diagrams illustrating an example of convolution processing performed in a filter. [Figure 14] FIG. 1A is a diagram for explaining the processing of a first pooling unit, FIG. 1B is a diagram for explaining the processing of a second pooling unit, and FIG. 1C is a diagram for explaining the processing of an unpooling unit. [Figure 15] FIG. 2 is a block diagram illustrating a hardware configuration of a saliency estimation device. [Figure 16] FIG. 11 is a diagram illustrating a functional configuration of a saliency estimation device according to a second embodiment. [Figure 17] FIG. 11 is a diagram illustrating a functional configuration of a saliency estimation device according to a third embodiment. [Figure 18] 11 is a diagram for explaining an example of the operation of a reference point setting unit; FIG. [Figure 19] FIG. 13 is a diagram illustrating a functional configuration of a saliency estimation device according to a fourth embodiment. [Figure 20] FIG. 13 is a diagram illustrating a configuration of a saliency estimation unit according to the fifth embodiment. [Figure 21] 13 is a flowchart illustrating a learning operation according to the fifth embodiment. [Figure 22] FIG. 13 is a diagram illustrating the configuration and usage environment of a computing device according to a sixth embodiment. [Diagram 23] FIG. 13 is a diagram illustrating a configuration of a saliency estimation unit according to the seventh embodiment. [Figure 24] 11 is a diagram illustrating an example of an image indicated by synthesis information generated by a synthesis unit; FIG. [Diagram 25] FIG. 23 is a diagram illustrating a configuration of a saliency estimation unit according to the eighth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are given the same reference numerals and the description will be omitted as appropriate.
[0010] (First embodiment) FIG. 1 is a diagram showing a functional configuration of a saliency estimation device 10 according to a first embodiment. The saliency estimation device 10 shown in this figure includes an input unit 110, a correction unit 120, and a saliency estimation unit 130. The input unit 110 acquires video data and outputs each of the frame images constituting the acquired video data to the correction unit 120. These frame images are images of a scene seen from a first viewpoint. The first viewpoint is, for example, a position where a camera that generated the video data was located. The correction unit 120 generates a corrected image by correcting the frame image output by the input unit 110. The saliency estimation unit 130 processes the corrected image to generate saliency estimation information. The saliency estimation information indicates a saliency distribution in the corrected image or in the frame image. Here, the correction unit 120 generates luminance information indicating a change in luminance of at least a part of the frame image using speed information related to the movement speed of the first viewpoint and a relative position from a reference point in the frame image to the at least part described above. Then, the correction unit 120 corrects at least a part of the brightness using this brightness information. The saliency estimation device 10 will be described in detail below.
[0011] As described above, moving image data is input to the input unit 110. The input unit 110 outputs to the correction unit 120 each of a plurality of frame images included in the moving image data.
[0012] The correction unit 120 corrects the frame images input from the input unit 110. Specifically, the correction unit 120 has a field of view setting unit 122, a correction information generating unit 124, and a correction processing unit 126.
[0013] The field of view setting unit 122 receives speed information and reference point information from the outside. The speed information indicates a moving speed. This moving speed is, for example, but is not limited to, the speed of the camera when the video data acquired by the input unit 110 is shot. The reference point information specifies a position that should be the gaze point in the frame image acquired by the correction unit 120. The field of view setting unit 122 uses the speed information to generate field of view angle information that indicates the field of view angle of a person when the person moves at that speed. Then, the field of view setting unit 122 outputs the speed information, the reference point information, and the field of view angle information to the correction information generation unit 124.
[0014] Fig. 2 is a diagram for explaining a method for setting the viewing angle information. The viewing angle of a moving person becomes narrower as the moving speed increases. The viewing angle setting unit 122 stores data showing the relationship between speed and viewing angle, for example, as shown in Fig. 2, and uses this data to specify the viewing angle corresponding to the input moving speed and generate viewing angle information showing the specified viewing angle.
[0015] FIG. 3 is a diagram for explaining an example of viewing angle information. In the example shown in this figure, the viewing angle information is information for dividing an image into a plurality of regions based on the distance from a reference point. Specifically, the 0th region including the reference point is a region where the image should be clear. Then, the correction information generating unit 124 and the correction processing unit 126 correct the image so that the region becomes gradually unclear and darker as it becomes a first region located outside the 0th region, a second region located outside the first region, and so on.
[0016] The field of view setting unit 122 determines the size of at least the 0th region using the speed information. For example, when the speed indicated by the speed information is low, the field of view setting unit 122 enlarges the 0th region and narrows the other regions. Note that the field of view setting unit 122 may set the number of regions to be set in addition to the size of each region using the speed information. In this case, the field of view setting unit 122 increases the number of regions to be set as the speed increases.
[0017] In the example shown in Figure 3, the outline of each region is rectangular, but the outline may be another shape (for example, circular or elliptical).
[0018] Returning to FIG. 1, the correction information generation unit 124 generates correction information using the reference point information acquired from the field of view setting unit 122. The correction information is information that specifies the amount of correction for the value of each pixel of the frame image. The field of view angle information is generated using the speed information and the distance from the reference point as described above. Therefore, the field of view setting unit 122 essentially generates correction information using the relative position from the reference point and the speed information.
[0019] In detail, the correction information includes brightness information, resolution information, and saturation information. The brightness information indicates a change in brightness of at least a part of the image, the resolution information indicates a change in resolution of at least a part of the image, and the saturation information indicates a change in saturation of at least a part of the image. Then, the correction information generation unit 124 outputs the viewing angle information, the reference point information, and the correction information to the correction processing unit 126.
[0020] FIG. 4 is a diagram for explaining lightness information, FIG. 5 is a diagram for explaining resolution information, and FIG. 6 is a diagram for explaining saturation information. As shown in these diagrams, the correction information generated by the correction information generating unit 124 indicates that the lightness, resolution, and saturation are all decreased as the distance from the reference point increases. Specifically, for each of the lightness, resolution, and saturation, a correction amount is set for each value of k in the "kth region" shown in FIG. 3. And, as the value of k increases, the lightness, resolution, and saturation all decrease.
[0021] Returning to Fig. 1, the correction processing unit 126 acquires a frame image from the input unit 110, and corrects this frame image using the viewing angle information, the reference point information, and the correction information. Specifically, the correction processing unit 126 defines a reference point in the frame image using the reference point information. Then, the correction processing unit 126 divides the frame image into each region shown in Fig. 3 using the reference point and the viewing angle information. Then, the correction processing unit 126 performs correction on each region according to the correction information.
[0022] 7 is a diagram showing an example of the functional configuration of the correction processing unit 126. The correction processing unit 126 has a resolution correction unit 202, a saturation correction unit 204, and a brightness correction unit 206. The resolution correction unit 202 corrects the resolution of the frame image for each region using the resolution information included in the correction information. The saturation correction unit 204 corrects the saturation of the frame image for each region using the color information included in the correction information. The brightness correction unit 206 corrects the brightness of the frame image for each region using the brightness information included in the correction information. In the example shown in this figure, the resolution correction unit 202, the saturation correction unit 204, and the brightness correction unit 206 are arranged in series in this order, but the order of these is not limited to the example shown in FIG. 7.
[0023] Then, the correction section 120 outputs the frame image after the correction (corrected image) to the saliency estimation section .
[0024] <Configuration example of the saliency estimation unit 130> FIG. 8 is a block diagram illustrating an example of the configuration of the saliency estimation unit 130. The saliency estimation unit 130 generates saliency estimation information by inputting a corrected frame image to a model generated by machine learning. In detail, the saliency estimation unit 130 includes an input unit 310, a nonlinear mapping unit 320, and an output unit 330. The input unit 310 converts the input frame image (hereinafter, described as an image in the description of the saliency estimation unit 130) into intermediate data that can be subjected to mapping processing. The nonlinear mapping unit 320 converts the intermediate data into mapping data. The output unit 330 generates saliency estimation information based on the mapping data. The nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. This will be described in detail below.
[0025] FIG. 9(a) is a diagram illustrating an example of an image input to the saliency estimation unit 130, and FIG. 9(b) is a diagram illustrating an example of an image showing a saliency distribution estimated for FIG. 9(a). For the sake of explanation, these diagrams show frame images before correction by the correction unit 120. The saliency estimation unit 130 according to this configuration example estimates the saliency of each part in an image. Saliency means, for example, how easily it stands out or how easily it attracts attention. Specifically, saliency is expressed by a probability or the like. Here, the magnitude of the probability corresponds to, for example, the probability that the gaze of a person who sees the image will be directed to that position.
[0026] 9(a) and 9(b) correspond to each other in terms of position. In FIG. 9(a), the higher the saliency of a position, the higher the luminance displayed in FIG. 9(b). An image showing a saliency distribution as in FIG. 9(b) is an example of saliency estimation information output by the output unit 330. In the example of this figure, saliency is visualized by 256 gradation luminance values. An example of saliency estimation information output by the output unit 330 will be described in detail later.
[0027] The estimation result of the saliency distribution can be used in various fields, such as, for example, predicting the gaze of traffic participants such as drivers and pedestrians, preventing traffic participants from being overlooked, evaluating the appearance of content such as advertising media, gaze guidance, digitizing the know-how of athletes and skilled craftsmen, understanding the visual cognition of living organisms, etc. Furthermore, the saliency estimation unit 130 and the processing method according to this configuration example can be applied to the mobility field such as autonomous driving, advanced driver assistance systems (ADAS), and road traffic systems, the entertainment field such as virtual reality (VR), augmented reality (AR), and games, the content field such as documents, video content, and signage, and the medical field such as image diagnosis, surgery support, and nursing care services.
[0028] FIG. 10 is a flowchart illustrating a processing method according to the first configuration example. The processing method according to this configuration example is a processing method executed by a computer, and includes an input step S110, a nonlinear mapping step S120, and an output step S130. In the input step S110, an image is converted into intermediate data that can be subjected to mapping processing. In the nonlinear mapping step S120, the intermediate data is converted into mapping data. In the output step S130, saliency estimation information indicating a saliency distribution is generated based on the mapping data. Here, the nonlinear mapping step S120 includes a feature extraction step S121 that extracts features from the intermediate data, and an upsampling step S122 that upsamples the data generated in the feature extraction step S121. The processing method according to this configuration example is realized by the saliency estimation unit 130 according to the configuration example.
[0029] Returning to FIG. 8, each component of the saliency estimation unit 130 will be described. In an input step S110, the input unit 310 acquires an image and converts it into intermediate data. The input unit 310 acquires an image from the correction unit 120. The input unit 310 then converts the acquired image into intermediate data. The intermediate data is not particularly limited as long as it is data that the nonlinear mapping unit 320 can accept, and is, for example, a high-dimensional tensor. In addition, the intermediate data is, for example, data in which the luminance of the acquired image is normalized, or data in which each pixel of the acquired image is converted into a luminance gradient. In the input step S110, the input unit 310 may further perform noise removal, resolution conversion, and the like on the image.
[0030] In the nonlinear mapping step S120, the nonlinear mapping unit 320 acquires intermediate data from the input unit 310. Then, the nonlinear mapping unit 320 converts the intermediate data into mapping data. Here, the mapping data is, for example, a high-dimensional tensor. The mapping process applied to the intermediate data by the nonlinear mapping unit 320 is, for example, a mapping process controllable by parameters or the like, and is preferably a process using a function, a functional, or a neural network.
[0031] Fig. 11 is a diagram illustrating a detailed configuration of the nonlinear mapping unit 320, and Fig. 12 is a diagram illustrating a configuration of the intermediate layer 323. As described above, the nonlinear mapping unit 320 includes a feature extraction unit 321 and an upsampling unit 322. The feature extraction unit 321 performs a feature extraction step S121, and the upsampling unit 322 performs an upsampling step S122. In the example of this figure, at least one of the feature extraction unit 321 and the upsampling unit 322 is configured to include a neural network including a plurality of intermediate layers 323. In the neural network, a plurality of intermediate layers 323 are connected.
[0032] In particular, the neural network is preferably a convolutional neural network. Specifically, each of the multiple intermediate layers 323 includes one or more convolutional layers 324. In the convolutional layer 324, input data is convolved by multiple filters 325, and activation processing is performed on the outputs of the multiple filters 325.
[0033] 11, feature extraction unit 321 is configured to include a neural network including a plurality of hidden layers 323, and includes a first pooling unit 326 between the plurality of hidden layers 323. Also, upsampling unit 322 is configured to include a neural network including a plurality of hidden layers 323, and includes an unpooling unit 328 between the plurality of hidden layers 323. Furthermore, feature extraction unit 321 and upsampling unit 322 are connected to each other via a second pooling unit 327 that performs overlap pooling.
[0034] In the example shown in the figure, each intermediate layer 323 is composed of two or more convolution layers 324. However, at least some of the intermediate layers 323 may be composed of only one convolution layer 324. Adjacent intermediate layers 323 are separated by any of a first pooling unit 326, a second pooling unit 327, and an unpooling unit 328. Here, when an intermediate layer 323 includes two or more convolution layers 324, it is preferable that the numbers of filters 325 in those convolution layers 324 are equal to each other.
[0035] In this figure, the intermediate layer 323 marked with "A×B" means that it is composed of B convolution layers 324, and each convolution layer 324 includes A convolution filters for each channel. Such an intermediate layer 323 is also referred to as an "A×B intermediate layer" below. For example, a 64×2 intermediate layer 323 means that it is composed of two convolution layers 324, and each convolution layer 324 includes 64 convolution filters for each channel.
[0036] In the example of this figure, the feature extraction unit 321 includes a 64×2 intermediate layer 323, a 128×2 intermediate layer 323, a 256×3 intermediate layer 323, and a 512×3 intermediate layer 323 in this order. The upsampling unit 322 includes a 512×3 intermediate layer 323, a 256×3 intermediate layer 323, a 128×2 intermediate layer 323, and a 64×2 intermediate layer 323 in this order. The second pooling unit 327 connects the two 512×3 intermediate layers 323 to each other. The number of intermediate layers 323 constituting the nonlinear mapping unit 320 is not particularly limited, and can be determined according to the number of pixels of the image data, for example.
[0037] Note that this figure is an example of the configuration of the nonlinear mapping unit 320, and the nonlinear mapping unit 320 may have another configuration. For example, a 64×1 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. By reducing the number of convolution layers 324 included in the intermediate layer 323, the calculation cost may be further reduced. Also, for example, a 32×2 intermediate layer 323 may be included instead of the 64×2 intermediate layer 323. By reducing the number of channels of the intermediate layer 323, the calculation cost may be further reduced. Furthermore, both the number of convolution layers 324 and the number of channels in the intermediate layer 323 may be reduced.
[0038] Here, in the multiple intermediate layers 323 included in the feature extraction unit 321, it is preferable that the number of filters 325 increases each time the first pooling unit 326 is passed through. Specifically, the first intermediate layer 323a and the second intermediate layer 323b are continuous with each other via the first pooling unit 326, and the second intermediate layer 323b is located after the first intermediate layer 323a. The first intermediate layer 323a is composed of a convolution layer 324 in which the number of filters 325 for each channel is N1, and the second intermediate layer 323b is composed of a convolution layer 324 in which the number of filters 325 for each channel is N2. At this time, it is preferable that N2>N1 holds. It is more preferable that N2=N1×2 holds.
[0039] Also, in the plurality of intermediate layers 323 included in the upsampling unit 322, it is preferable that the number of filters 325 decreases every time it passes through the upsampling unit 328. Specifically, the third intermediate layer 323c and the fourth intermediate layer 323d are continuous with each other via the upsampling unit 328, and the fourth intermediate layer 323d is located after the third intermediate layer 323c. The third intermediate layer 323c is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N3, and the fourth intermediate layer 323d is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N4. At this time, it is preferable that N4 < N3 holds. More preferably, N3 = N4 × 2 holds.
[0040] In the feature extraction unit 321, image features having a plurality of levels of abstraction, such as gradients and shapes, are extracted from the intermediate data acquired from the input unit 310 as channels of the intermediate layer 323. FIG. 12 illustrates the configuration of the 64×2 intermediate layer 323. With reference to this figure, the processing in the intermediate layer 323 will be described. In the example of this figure, the intermediate layer 323 is composed of a first convolutional layer 324a and a second convolutional layer 324b, and each convolutional layer 324 includes 64 filters 325. In the first convolutional layer 324a, convolution processing using the filter 325 is performed on each channel of the data input to the intermediate layer 323. For example, when the image input to the input unit 310 is an RGB image, processing is performed on each of the three channels h0i (i = 1..3). Also, in the example of this figure, the filter 325 is a 64 types of 3×3 filters, that is, a total of 64×3 types of filters. As a result of the convolution processing, 64 results h0i,j (i = 1..3, j = 1..64) are obtained for each channel i.
[0041] Next, the activation unit 329 performs activation processing on the outputs of the multiple filters 325. Specifically, activation processing is performed on the sum of the corresponding elements for the corresponding results j of all channels. By this activation processing, the results h1i (i=1..64) of 64 channels, that is, the output of the first convolution layer 324a, are obtained as image features. The activation processing is not particularly limited, but it is preferable to use at least one of a hyperbolic function, a sigmoid function, and a normalized linear function.
[0042] Furthermore, the output data of the first convolution layer 324a is used as input data for the second convolution layer 324b, and the same processing as the first convolution layer 324a is performed in the second convolution layer 324b, and the result of 64 channels h2i (i=1..64), that is, the output of the second convolution layer 324b, is obtained as an image feature. The output of the second convolution layer 324b becomes the output data of this 64×2 intermediate layer 323.
[0043] Here, the structure of the filter 325 is not particularly limited, but is preferably a 3×3 two-dimensional filter. Moreover, the coefficients of each filter 325 can be set independently. In this configuration example, the coefficients of each filter 325 are held in the storage unit 390, and the nonlinear mapping unit 320 can read them and use them for processing. Here, the coefficients of the multiple filters 325 may be determined based on correction information generated and corrected using machine learning. For example, the correction information includes the coefficients of the multiple filters 325 as multiple correction parameters. The nonlinear mapping unit 320 can further use this correction information to convert the intermediate data into mapping data. The storage unit 390 may be provided in the saliency estimation unit 130, or may be provided outside the saliency estimation unit 130. Moreover, the nonlinear mapping unit 320 may acquire the correction information from the outside via a communication network.
[0044] 13(a) and 13(b) are diagrams showing examples of convolution processing performed by the filter 325. In both of FIG. 13(a) and FIG. 13(b), an example of 3×3 convolution is shown. The example in FIG. 13(a) is convolution processing using nearest neighbor elements. The example in FIG. 13(b) is convolution processing using neighbor elements with a distance of two or more. Note that convolution processing using neighbor elements with a distance of three or more is also possible. It is preferable that the filter 325 performs convolution processing using neighbor elements with a distance of two or more. This is because a wider range of features can be extracted, and the estimation accuracy of saliency can be further improved.
[0045] The above describes the operation of the 64×2 intermediate layer 323. The operations of the other intermediate layers 323 (such as the 128×2 intermediate layer 323, the 256×3 intermediate layer 323, and the 512×3 intermediate layer 323) are the same as the operation of the 64×2 intermediate layer 323, except for the number of convolution layers 324 and the number of channels. In addition, the operations of the intermediate layer 323 in the feature extraction unit 321 and the operation of the intermediate layer 323 in the upsampling unit 322 are also the same as those described above.
[0046] FIG. 14(a) is a diagram for explaining the processing of the first pooling unit 326, FIG. 14(b) is a diagram for explaining the processing of the second pooling unit 327, and FIG. 14(c) is a diagram for explaining the processing of the unpooling unit 328.
[0047] In the feature extraction unit 321, the data output from the intermediate layer 323 is subjected to pooling processing for each channel in the first pooling unit 326, and then input to the next intermediate layer 323. In the first pooling unit 326, for example, non-overlapping pooling processing is performed. FIG. 14(a) shows a process of associating four 2×2 elements 30 with one element 30 for an element group included in each channel. In the first pooling unit 326, such association is performed for all elements 30. Here, the four 2×2 elements 30 are selected so that they do not overlap with each other. In this example, the number of elements in each channel is reduced to one-fourth. Note that, as long as the number of elements is reduced in the first pooling unit 326, the number of elements 30 before and after the association is not particularly limited.
[0048] The data output from the feature extraction unit 321 is input to the upsampling unit 322 via the second pooling unit 327. In the second pooling unit 327, overlap pooling is performed on the output data from the feature extraction unit 321. FIG. 14(b) shows a process of associating four 2×2 elements 30 with one element 30 while overlapping some of the elements 30. That is, in repeated association, some of the four 2×2 elements 30 in a certain association are also included in the four 2×2 elements 30 in the next association. The number of elements is not reduced in the second pooling unit 327 as shown in this figure. Note that the number of elements 30 before and after the association in the second pooling unit 327 is not particularly limited.
[0049] The methods of processing performed by the first pooling unit 326 and the second pooling unit 327 are not particularly limited, but examples include matching in which the maximum value of four elements 30 is matched to one element 30 (max pooling) and matching in which the average value of four elements 30 is matched to one element 30 (average pooling).
[0050] The data output from the second pooling unit 327 is input to the intermediate layer 323 in the upsampling unit 322. Then, the output data from the intermediate layer 323 of the upsampling unit 322 is subjected to an unpooling process for each channel in the unpooling unit 328, and then input to the next intermediate layer 323. Fig. 14(c) shows a process of expanding one element 30 to multiple elements 30. The method of expansion is not particularly limited, but a method of duplicating one element 30 to four elements 30 (2 x 2) can be given as an example.
[0051] The output data of the last intermediate layer 323 of the upsampling unit 322 is output from the nonlinear mapping unit 320 as mapping data and input to the output unit 330. In the output step S130, the output unit 330 performs, for example, normalization, resolution conversion, etc. on the data acquired from the nonlinear mapping unit 320 to generate and output saliency estimation information. The saliency estimation information is, for example, an image (image data) in which saliency is visualized by brightness values as illustrated in FIG. 9(b). The saliency estimation information may be, for example, an image colored according to saliency like a heat map, or an image in which salient regions with saliency higher than a predetermined standard are marked so as to be distinguishable from other positions. Furthermore, the saliency estimation information is not limited to an image, and may be a table or the like that lists information indicating salient regions.
[0052] The saliency estimation information output from the output unit 330 may be subjected to various computer vision processes such as image segmentation, object recognition, and image classification within the saliency estimation unit 130 or outside the saliency estimation unit 130.
[0053] <Hardware configuration example> Fig. 15 is a block diagram illustrating a hardware configuration of the saliency estimation device 10 shown in Fig. 1. The saliency estimation device 10 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, an input / output interface 1050, and a network interface 1060.
[0054] The bus 1010 is a data transmission path for transmitting and receiving data among the processor 1020, memory 1030, storage device 1040, input / output interface 1050, and network interface 1060. However, the method of connecting the processor 1020 and the like to each other is not limited to bus connection.
[0055] The processor 1020 is implemented by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0056] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.
[0057] The storage device 1040 is an auxiliary storage device realized by a hard disk drive (HDD), a solid state drive (SSD), a memory card, a read only memory (ROM), etc. The storage device 1040 stores program modules that realize each function of the saliency estimation device 10. The processor 1020 loads each of these program modules onto the memory 1030 and executes them, thereby realizing each function corresponding to the program module.
[0058] The input / output interface 1050 is an interface for connecting the saliency estimation device 10 to various input / output devices.
[0059] The network interface 1060 is an interface for connecting the saliency estimation device 10 to a network. This network is, for example, a local area network (LAN) or a wide area network (WAN). The method for connecting the network interface 1060 to the network may be a wireless connection or a wired connection.
[0060] As described above, according to this embodiment, the brightness of the frame images is changed using the speed information before the saliency estimation process, so that the saliency estimation unit 130 can detect with high accuracy an area that would be perceived as highly salient by a moving person from each frame image constituting a video.
[0061] Moreover, the saliency estimation unit 130 includes a feature extraction unit 321 that extracts features from the intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, saliency can be estimated with low calculation cost.
[0062] The saliency estimation device 10 can also process still images. In this case, the above-mentioned effects can also be obtained.
[0063] Second Embodiment 16 is a diagram showing a functional configuration of the saliency estimation device 10 according to the second embodiment. The saliency estimation device 10 according to the present embodiment has the same configuration as the saliency estimation device 10 according to the first embodiment, except that it includes a speed estimation unit 140.
[0064] The speed estimation unit 140 estimates the moving speed by processing the video data input to the input unit 110. An existing algorithm such as optical flow estimation can be used as an estimation algorithm for this moving speed. The speed estimation unit 140 then outputs the estimated speed to the field of view setting unit 122 as speed information.
[0065] This embodiment also provides the same effects as those of the first embodiment. Furthermore, since the speed estimation unit 140 generates the speed information, there is no need to input the speed information from outside.
[0066] (Third embodiment) 17 is a diagram showing a functional configuration of the saliency estimation device 10 according to the third embodiment. The saliency estimation device 10 according to the present embodiment has the same configuration as the saliency estimation device 10 according to the second embodiment, except that it includes a reference point setting unit 150.
[0067] The reference point setting unit 150 processes at least one frame image of the video data input to the input unit 110 to generate reference point information.
[0068] For example, when a frame image includes a predetermined object, such as a specific traffic sign or other object that is easily noticeable to people, the reference point setting unit 150 sets the position of the object as the reference point. This object detection is performed, for example, by a matching process of feature amounts. The feature amounts used here are stored in advance in the saliency estimation device 10.
[0069] Furthermore, when a road is included in the frame image and the length of the straight portion of the road is equal to or longer than a reference, the reference point setting unit 150 sets the vanishing point as the reference point. Existing algorithms such as Hough transform can be used to detect the vanishing point.
[0070] Furthermore, when the scenery shown in the frame image satisfies a specific condition, the reference point setting unit 150 may set a reference point by performing processing according to the condition. For example, as shown in Figures 18(a) and 18(b), when a road is included in the frame image and the road is curved, a part of the road located in the direction of the curve from the center (e.g., the median line) is set as the reference point.
[0071] This embodiment also provides the same effects as those of the second embodiment. Furthermore, since the reference point setting unit 150 is provided, there is no need to input reference point information from the outside.
[0072] (Fourth embodiment) 19 is a diagram showing a functional configuration of a saliency estimation device 10 according to the fourth embodiment. The saliency estimation device 10 according to the present embodiment has the same configuration as the saliency estimation device 10 according to the first embodiment, except that it includes a trimming unit 160.
[0073] The trimming unit 160 acquires the generation conditions of the video data acquired by the input unit 110, and trims the frame images using the generation conditions. The correction unit 120 processes the frame images trimmed by the trimming unit 160. The generation conditions input to the trimming unit 160 are, for example, the type of lens (wide-angle lens or fisheye lens) of the camera that generated the video data. Depending on the generation conditions of the video data, the range of the scenery captured in the frame images may be wider than the field of view of a stationary person. The trimming unit 160 trims the frame images to match the range of the scenery captured in the frame images to the field of view of a stationary person. The range to be trimmed from the frame images is, for example, stored in advance by the trimming unit 160 for each generation condition.
[0074] This embodiment also provides the same effect as the first embodiment. The trimming unit 160 trims the frame image so that the range of the scenery shown in the frame image fits the field of view of a stationary person. This makes it possible to detect with even higher accuracy an area from the image that would be perceived as highly salient by a moving person.
[0075] The trimming unit 160 according to this embodiment may be provided in the saliency estimation device 10 shown in the second or third embodiment.
[0076] Fifth embodiment The saliency estimation device 10 according to this embodiment has the same configuration as the saliency estimation device 10 according to any of the above-described embodiments, except for the functional configuration of the saliency estimation unit 130.
[0077] 20 is a diagram illustrating a configuration of the saliency estimation unit 130 according to the present embodiment. The saliency estimation unit 130 according to the present embodiment is the same as the saliency estimation unit 130 according to the first embodiment, except that it further includes an error calculation unit 340 and a correction unit 350. The error calculation unit 340 uses the saliency estimation information generated for an image and the saliency actual measurement information indicating the saliency distribution actually measured for the image to calculate an error between the saliency distribution indicated by the saliency estimation information and the saliency distribution indicated by the saliency actual measurement information. Then, the correction unit 350 corrects the correction information based on the calculated error.
[0078] The saliency estimation unit 130 according to this embodiment performs an estimation operation and a learning operation. In the estimation operation, saliency estimation information for an input image is generated and output. The estimation operation is the same as that described in the first embodiment. In particular, in this embodiment, the nonlinear mapping unit 320 converts intermediate data into mapping data using correction information. Meanwhile, in the learning operation, machine learning is performed using a teacher image and saliency measurement information for the teacher image, and correction information is generated or modified (updated). The correction information is information used by the nonlinear mapping unit 320, and includes, for example, a plurality of correction parameters.
[0079] In this embodiment, the nonlinear mapping unit 320 converts the intermediate data into mapping data using the correction information. The correction information is information that is at least one of generated and corrected using machine learning. Specifically, the nonlinear mapping unit 320 includes a plurality of filters 325 as described in the first embodiment, and the coefficients of the plurality of filters 325 are determined based on the correction information. For example, the correction information includes the coefficients of the plurality of filters 325 as a plurality of correction parameters.
[0080] FIG. 21 is a flow chart illustrating a learning operation according to this embodiment. The learning operation will be described in detail below. For the learning operation, a teacher image and saliency measurement information for the teacher image are prepared. For example, the teacher image and the saliency measurement information are associated with each other and stored in the storage unit 390. The input unit 310 and the error calculation unit 340 can read out this information from the storage unit 390 and use it.
[0081] The teacher image is any image such as a photograph. The saliency measurement information is generated based on the results of actually measuring the line of sight of a person when looking at the teacher image using an eye tracker. The saliency measurement information can have the same form as the saliency estimation information. That is, the saliency measurement information may be an image in which saliency is visualized by a brightness value, or the saliency measurement information may be an image that is color-coded according to saliency, such as a heat map.
[0082] In the learning operation, the input step S110, the nonlinear mapping step S120, and the output step S130 are performed in the same manner as the input step S110, the nonlinear mapping step S120, and the output step S130 according to the first embodiment. However, the image acquired by the input unit 310 in the input step S110 is a teacher image. In addition, in the nonlinear mapping step S120, the nonlinear mapping unit 320 reads out correction information from the storage unit 390. Then, the intermediate data is converted into mapping data using the correction information. Note that the nonlinear mapping unit 320 may directly acquire the correction information from the correction unit 350 instead of reading it from the storage unit 390. Also, in the initial state, the correction parameters included in the correction information can be set to any value.
[0083] Next, in an error calculation step S140, the error calculation unit 340 acquires saliency estimation information from the output unit 330. The error calculation unit 340 also acquires saliency measurement information associated with the teacher image that is the source of the saliency estimation information. The error calculation unit 340 then calculates the error between the acquired saliency estimation information and the saliency measurement information. There are no particular limitations on the method of calculating the error, but it is preferable to calculate at least one of the L1 distance, the L2 distance (Euclidean distance, mean square error), the Kullback-Leibler distance, the Jensen-Shannon distance, and the Pearson correlation coefficient, for example.
[0084] Specifically, the Euclidean distance is calculated by the following formula (1), the Kullback-Leibler distance is calculated by the following formula (2), and the Jensen-Shannon distance is calculated by the following formula (3): where pi indicates the estimation result (value based on saliency estimation information), and qi indicates the true value (value based on saliency actual measurement information).
[0085]
number
number
number
[0086] Next, in a correction step S150, the correction unit 350 acquires the error from the error calculation unit 340 and corrects the correction parameters so as to reduce the error. Then, the correction parameters held in the storage unit 390 are replaced with the corrected correction parameters. Here, the method of correcting the correction parameters is not particularly limited, but it is preferable to use at least one of the least squares method, quadratic programming, stochastic gradient descent (SGD), adaptive moment estimation (ADAM), and the calculus of variations, for example.
[0087] Here, there are many correction parameters to be corrected, and in order to efficiently determine the values of those parameters and estimate the saliency with high accuracy, it is preferable to use statistical learning (machine learning) using a large amount of teacher data. Therefore, in the learning operation, it is preferable that machine learning is performed by cooperation between the nonlinear mapping unit 320, the error calculation unit 340, and the correction unit 350.
[0088] Note that instead of replacing the correction parameters stored in the storage unit 390 with the corrected correction parameters, the correction unit 350 may output the corrected correction parameters directly to the nonlinear mapping unit 320. In the next nonlinear mapping step S120, the nonlinear mapping unit 320 performs processing using the corrected correction parameters.
[0089] Note that one or more pieces of saliency measurement information may be associated with one teacher image. When multiple pieces of saliency measurement information are associated with one teacher image, the multiple pieces of saliency measurement information are based on different measurement results. Then, the error calculation unit 340 calculates the error between the saliency estimation information and each piece of saliency measurement information. Furthermore, the modification unit 350 modifies the correction parameters, for example, so that the sum of all the errors becomes smaller.
[0090] The learning operation may be performed on multiple pairs of teacher images and saliency measurement information. By repeating the learning operation, the accuracy of saliency estimation is further improved.
[0091] The timing at which the learning operation is performed is not particularly limited. For example, the saliency estimation unit 130 can accept an operation by a user to start the learning operation. Then, based on the operation to start the learning operation, the saliency estimation unit 130 can start the learning operation. Also, the saliency estimation unit 130 can end the learning operation based on a termination operation by the user or a predetermined termination condition. Examples of the termination condition include satisfying a predetermined number of repetitions of the learning operation or an error being equal to or less than a predetermined reference value.
[0092] As described above, according to this embodiment, similarly to the first embodiment, the nonlinear mapping unit 320 includes the feature extraction unit 321 that extracts features from the intermediate data, and the upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, saliency can be estimated with small calculation cost.
[0093] In addition, according to this embodiment, the saliency estimation unit 130 includes an error calculation unit 340 and a correction unit 350. Therefore, using the correction information corrected by the learning operation, more accurate saliency estimation is realized.
[0094] Sixth embodiment The saliency estimation device 10 according to this embodiment has the same configuration as the saliency estimation device 10 according to any of the above-described embodiments, except for the functional configuration of the saliency estimation unit 130.
[0095] FIG. 22 is a diagram illustrating a configuration and a usage environment of a calculation device 40 according to this embodiment. The calculation device 40 according to this embodiment is a device that generates correction information used in the saliency estimation unit 130. The calculation device 40 includes an error calculation unit 440 and a correction unit 450. The error calculation unit 440 uses the saliency estimation information generated for the teacher image and the saliency actual measurement information indicating the saliency distribution actually measured for the teacher image to calculate an error between the saliency distribution indicated by the saliency estimation information and the saliency distribution indicated by the saliency actual measurement information. The correction unit 450 calculates the correction information based on the error.
[0096] The saliency estimation unit 130 according to this embodiment is the same as the saliency estimation unit 130 according to the first embodiment. The saliency estimation unit 130 according to this embodiment includes an input unit 310, a nonlinear mapping unit 320, and an output unit 330. The saliency estimation unit 130 according to this embodiment does not need to include the error calculation unit 340 and the correction unit 350 described in the fifth embodiment. The input unit 310 according to this embodiment is the same as the input unit 310 according to at least one of the first and fifth embodiments, the nonlinear mapping unit 320 according to this embodiment is the same as the nonlinear mapping unit 320 according to at least one of the first and fifth embodiments, and the output unit 330 according to this embodiment is the same as the output unit 330 according to at least one of the first and fifth embodiments. The operation of the error calculation unit 440 according to this embodiment is the same as the operation of the error calculation unit 340 according to the fifth embodiment, and the operation of the correction unit 450 according to this embodiment is the same as the operation of the correction unit 350 according to the fifth embodiment. The saliency estimation unit 130 and the calculation device 40 cooperate to perform the learning operation and the estimation operation described in the fifth embodiment. In addition, the saliency estimation unit 130 and the calculation device 40 may be physically separated from each other, and may be connected to each other via, for example, a communication network.
[0097] In the learning operation according to this embodiment, it is preferable that the nonlinear mapping section 320, the error calculation section 440, and the correction section 450 cooperate to perform machine learning.
[0098] The output unit 330 may temporarily store the generated saliency estimation information in the storage unit 390, and the error calculation unit 440 may read out the saliency estimation information stored in the storage unit 390 and use it.
[0099] In the example of the figure, the storage unit 390 is provided separately from the saliency estimation unit 130 and the calculation device 40, but this is not limited to this example, and the storage unit 390 may be provided in the saliency estimation unit 130 or the calculation device 40. When the storage unit 390 is provided inside the calculation device 40, for example, the storage unit 390 is realized using the storage device 1080 of the computer 1000 that realizes the calculation device 40. Furthermore, the storage unit 390 may be realized by cooperation between the storage device 1080 of the computer 1000 that realizes the saliency estimation unit 130 and the storage device 1080 of the computer 1000 that realizes the calculation device 40.
[0100] As described above, according to this embodiment, similarly to the first embodiment, the nonlinear mapping unit 320 includes the feature extraction unit 321 that extracts features from the intermediate data, and the upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, saliency can be estimated with small calculation cost.
[0101] In addition, according to this embodiment, the arithmetic device 40 includes an error calculation section 440 and a correction section 450. Therefore, using the correction information corrected by the learning operation, more accurate saliency estimation is realized.
[0102] Seventh embodiment The saliency estimation device 10 according to this embodiment has the same configuration as the saliency estimation device 10 according to any of the above-described embodiments, except for the functional configuration of the saliency estimation unit 130.
[0103] 23 is a diagram illustrating a configuration of the saliency estimation unit 130 according to the present embodiment. The saliency estimation unit 130 according to the present embodiment is the same as the saliency estimation unit 130 according to at least one of the first and fifth embodiments, except that the saliency estimation unit 130 further includes a synthesis unit 360 and a display unit 380.
[0104] The synthesis unit 360 generates synthesis information by synthesizing the saliency distribution indicated by the saliency estimation information and the image (input image) input to the input unit 310. Specifically, the synthesis unit 360 acquires the saliency estimation information from the output unit 330, and acquires the input image from, for example, the storage unit 390. Then, the synthesis unit 360 outputs synthesis information indicating the input image and the saliency distribution together. The synthesis information is output to, for example, a display unit 380 provided in the saliency estimation unit 130. Furthermore, the synthesis information output from the synthesis unit 360 may be held in the storage unit 390 or acquired by an external device.
[0105] FIG. 24 is a diagram illustrating an example of an image indicated by the synthesis information generated by the synthesis unit 360. In the example of this figure, the synthesis information is an image in which the input image and a heat map indicating saliency are superimposed. The format of the synthesis information is not particularly limited. For example, the synthesis information may be an image in which a salient region is surrounded by a circle or a square in the input image. The synthesis method is also not particularly limited, and examples thereof include alpha blending.
[0106] The saliency estimation unit 130 according to this embodiment can be implemented in, for example, a mobile terminal (smartphone, tablet, etc.) equipped with an imaging device such as a camera. In this way, while taking an image with the mobile terminal, important objects with high saliency can be extracted on the spot and visualized with good visibility.
[0107] As described above, according to this embodiment, similarly to the first embodiment, the nonlinear mapping unit 320 includes the feature extraction unit 321 that extracts features from the intermediate data, and the upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, saliency can be estimated with small calculation cost.
[0108] In addition, according to this embodiment, the saliency estimation unit 130 further includes a synthesis unit 360. Therefore, the saliency at each position of the image can be visualized with good visibility.
[0109] Eighth embodiment The saliency estimation device 10 according to this embodiment has the same configuration as the saliency estimation device 10 according to any of the above-described embodiments, except for the functional configuration of the saliency estimation unit 130.
[0110] 25 is a diagram illustrating a configuration of the saliency estimation unit 130 according to the present embodiment. The saliency estimation unit 130 according to the present embodiment is the same as the saliency estimation unit 130 according to at least any one of the first, fifth, and seventh embodiments, except that the saliency estimation unit 130 further includes a mask image generation unit 370, a region extraction unit 372, and an object detection unit 374.
[0111] The mask image generating unit 370 acquires the saliency estimation information from the output unit 330 and generates a mask image. Specifically, the mask image generating unit 370 generates a mask image in which a region in the saliency distribution indicated by the saliency estimation information, whose saliency is lower than a predetermined criterion, is a masked region, and a region in which the saliency is equal to or higher than the predetermined criterion, is a non-masked region. That is, the mask image generating unit 370 binarizes the saliency distribution. Here, the criterion is set in advance and stored in the storage unit 390, and the mask image generating unit 370 can read and use it.
[0112] The region extraction unit 372 acquires an input image and a mask image. Then, the mask image is applied to the input image to extract a highly salient region from the input image. For example, the region extraction unit 372 can extract a highly salient region from the input image by performing a logical operation on the input image and the mask image.
[0113] Then, the object detection unit 374 detects an object from the region extracted by the region extraction unit 372. The method of detecting an object is not particularly limited, but may be, for example, a method using a Single Shot Multibox Detector (SSD). The saliency estimation unit 130 according to this embodiment extracts a region with high saliency in advance, and performs object detection only in the extracted region, thereby suppressing erroneous detection.
[0114] The saliency estimation unit 130 according to this embodiment is mounted on a moving body such as an automobile, etc. The object detection result by the object detection unit 374 can be used for automatic driving and driving assistance.
[0115] As described above, according to this embodiment, similarly to the first embodiment, the nonlinear mapping unit 320 includes the feature extraction unit 321 that extracts features from the intermediate data, and the upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, saliency can be estimated with small calculation cost.
[0116] In addition, according to this embodiment, the saliency estimation unit 130 further includes a mask image generation unit 370, a region extraction unit 372, and an object detection unit 374. Therefore, highly accurate object detection can be performed in the input image.
[0117] Although the embodiment and examples have been described above with reference to the drawings, these are merely examples of the present invention, and various configurations other than those described above can also be adopted. [Explanation of symbols]
[0118] 10. Saliency Estimation Device 110 Input section 120 Correction section 122 Field of view setting unit 124 Correction information generation unit 126 Correction processing section 130 Saliency Estimation Unit 140 Speed estimation part 150 Reference point setting section 160 Trimming Section 202 Resolution correction section 204 Saturation correction section 206 Brightness correction section
Claims
[Claim 1] a correction unit for correcting an image to generate a corrected image; a saliency estimation unit that processes the corrected image to generate saliency estimation information indicative of a saliency distribution in the corrected image or in the image; Equipped with The correction unit is generating brightness information indicative of a change in brightness of at least a portion of the image using speed information relating to a moving speed at the time of capturing the image and a distance from a reference point in the image; A saliency estimation device that generates the saliency estimation information based on a corrected image in which at least a portion of the brightness is corrected using the brightness information.
Citation Information
Patent Citations
Image processing apparatus
JP2018019435A
Display device and program
WO2018062538A1