Sampling estimation device, sampling estimation method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-14
Smart Images

Figure 2026131818000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a saliency estimation device, a saliency estimation method, and a program.
Background Art
[0002] Techniques for automatically detecting salient regions in an image have been proposed. On the other hand, when a person is moving, as the person's moving speed increases, the person's effective visual field becomes narrower. Non-Patent Document 1 describes automatically detecting a salient region in consideration of the effective visual field. Specifically, Non-Patent Document 1 describes reducing the resolution and saturation of a target image according to the distance from a fixation point and then estimating saliency.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The inventor has studied a method for highly accurately detecting a region that is felt to have high saliency when viewed by a person in motion from an image. As an example of the problem to be solved by the present invention, highly accurately detecting a region that is felt to have high saliency when viewed by a person in motion from an image can be cited.
Means for Solving the Problems
[0005] An example of the invention according to the present disclosure includes a correction unit that generates a corrected image by correcting an image of a scene viewed from a first viewpoint, A spleness estimation unit that processes the corrected image to generate spleness estimation information showing the spleness distribution within the corrected image or within the image, Equipped with, The correction unit, Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using velocity information relating to the movement speed of the first viewpoint and the relative position from the reference point to at least the portion in the image. This is a spleness estimation device that corrects at least some of the brightness using the aforementioned brightness information. Another example of the invention described herein is a correction unit that generates a corrected image by correcting an image, A spleness estimation unit that processes the corrected image to generate spleness estimation information showing the spleness distribution within the corrected image or within the image, Equipped with, The correction unit, Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using speed information relating to the movement speed of a first viewpoint, which is the position where the camera that took the image was positioned at the time the image was taken, and reference point information that identifies the position in the image that should be the point of focus, based on the distance from a reference point set in the image. This is a spleness estimation device that, when generating the corrected image, corrects at least some of the brightness using the brightness information.
[0006] One example of the invention described herein is a computer that, By correcting the image of the scenery viewed from the first viewpoint, a corrected image is generated. By processing the corrected image, significance estimation information is generated that shows the significance distribution within the corrected image or within the image. Furthermore, the aforementioned computer Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using velocity information relating to the movement speed of the first viewpoint and the relative position from the reference point to at least the portion in the image. This is a method for estimating splendor that corrects at least some of the splendor using the aforementioned splendor information. Another example of the invention relating to this disclosure is a computer, By correcting the image, a corrected image is generated. By processing the corrected image, significance estimation information is generated that shows the significance distribution within the corrected image or within the image. When generating the corrected image, Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using speed information relating to the movement speed of a first viewpoint, which is the position where the camera that took the image was positioned at the time the image was taken, and reference point information that identifies the position in the image that should be the point of focus, based on the distance from a reference point set in the image. This is a spleness estimation method that, when generating the corrected image, corrects at least some of the brightness using the brightness information.
[0007] One example of the invention described herein is a computer, A correction function that generates a corrected image by correcting an image of the scenery viewed from the first viewpoint, An estimation function that processes the corrected image to generate sampling estimation information showing the sampling distribution within the corrected image or within the image, Give it to him Furthermore, as at least a part of the correction function, A function to generate brightness information indicating a change in brightness of at least a portion of the aforementioned image, using velocity information relating to the movement speed of the first viewpoint, and the relative position from a reference point to at least the aforementioned portion in the image, A function to correct at least some of the brightness using the brightness information, This is a program that gives it a function. Another example of the invention relating to this disclosure is a computer, A correction function that generates a corrected image by correcting the image, A spleness estimation function that processes the corrected image to generate spleness estimation information showing the spleness distribution within the corrected image or within the image, Give it to him Furthermore, the correction function is Brightness information indicating the change in brightness of at least a part of the image is generated using the distance from a reference point set in the image based on speed information regarding the moving speed of a first viewpoint, which is the position where the camera that captured the image was placed at the time of capturing the image, and reference point information for specifying a position to be a fixation point in the image, A program that corrects the brightness of at least a part using the brightness information when generating the corrected image.
Brief Description of Drawings
[0008] [Figure 1] It is a diagram showing the functional configuration of a saliency estimation device according to a first embodiment. [Figure 2] It is a diagram for explaining a method of setting viewing angle information. [Figure 3] It is a diagram for explaining an example of viewing angle information. [Figure 4] It is a diagram for explaining brightness information. [Figure 5] It is a diagram for explaining resolution information. [Figure 6] It is a diagram for explaining saturation information. [Figure 7] It is a diagram showing an example of the functional configuration of a correction processing unit. [Figure 8] It is a block diagram exemplifying a configuration example of a saliency estimation unit. [Figure 9] (a) is a diagram exemplifying an image input to a saliency estimation unit, and (b) is a diagram exemplifying an image showing a saliency distribution estimated for (a). [Figure 10] It is a flowchart exemplifying a processing method according to a first configuration example. [Figure 11] It is a diagram showing in detail an example of the configuration of a non-linear mapping unit. [Figure 12] It is a diagram exemplifying the configuration of an intermediate layer. [Figure 13] (a) and (b) are diagrams each showing an example of a convolution process performed by a filter. [Figure 14](a) is a diagram illustrating the process of the first pooling section, (b) is a diagram illustrating the process of the second pooling section, and (c) is a diagram illustrating the process of the unpooling section. [Figure 15] This is a block diagram illustrating the hardware configuration of a sampling estimation device. [Figure 16] This figure shows the functional configuration of the prominentness estimation device according to the second embodiment. [Figure 17] This figure shows the functional configuration of the prominentness estimation device according to the third embodiment. [Figure 18] This is a diagram illustrating an example of the operation of the reference point setting unit. [Figure 19] This figure shows the functional configuration of the prominentness estimation device according to the fourth embodiment. [Figure 20] This figure illustrates the configuration of the prominentness estimation unit according to the fifth embodiment. [Figure 21] This is a flowchart illustrating the learning process according to the fifth embodiment. [Figure 22] This figure illustrates the configuration and operating environment of the computing device according to the sixth embodiment. [Figure 23] This figure illustrates the configuration of the significance estimation unit according to the seventh embodiment. [Figure 24] This figure illustrates an image represented by the composite information generated in the synthesis unit. [Figure 25] This figure illustrates the configuration of the significance estimation unit according to the eighth embodiment. [Modes for carrying out the invention]
[0009] Embodiments of the present invention will be described below with reference to the drawings. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted as appropriate.
[0010] (First embodiment) Figure 1 is a diagram showing the functional configuration of the spleness estimation device 10 according to the first embodiment. The spleness estimation device 10 shown in this figure comprises an input unit 110, a correction unit 120, and a spleness estimation unit 130. The input unit 110 acquires video data and outputs each of the frame images constituting the acquired video data to the correction unit 120. These frame images are images of the scenery as seen from a first viewpoint. The first viewpoint is, for example, the position where the camera that generated the video data was positioned. The correction unit 120 generates a corrected image by correcting the frame images output by the input unit 110. The spleness estimation unit 130 generates spleness estimation information by processing the corrected image. The spleness estimation information indicates the spleness distribution within the corrected image or within the frame images. Here, the correction unit 120 generates brightness information indicating a change in brightness of at least a part of the frame image using speed information related to the movement speed of the first viewpoint and the relative position from a reference point in the frame image to at least the above-mentioned part. The correction unit 120 then uses this brightness information to correct at least some of the brightness values described above. The spleness estimation device 10 will now be described in detail.
[0011] As described above, video data is input to the input unit 110. The input unit 110 outputs each of the multiple frame images contained in the video data to the correction unit 120.
[0012] The correction unit 120 corrects the frame image input from the input unit 110. Specifically, the correction unit 120 includes a field setting unit 122, a correction information generation unit 124, and a correction processing unit 126.
[0013] The field of view setting unit 122 receives speed information and reference point information from an external source. The speed information indicates the speed of movement. This speed is, for example, the speed of the camera when the video data acquired by the input unit 110 was captured, but is not limited to this. The reference point information identifies the position that should be the point of focus in the frame image acquired by the correction unit 120. The field of view setting unit 122 uses the speed information to generate field of view information that indicates the field of view of a person when they move at that speed. The field of view setting unit 122 then outputs the speed information, reference point information, and field of view information to the correction information generation unit 124.
[0014] Figure 2 is a diagram illustrating how to set field of view information. The field of view of a moving person narrows as their speed increases. The field of view setting unit 122 stores data showing the relationship between speed and field of view, such as shown in Figure 2, and uses this data to identify the field of view corresponding to the input speed of movement and generates field of view information indicating the identified field of view.
[0015] Figure 3 is a diagram illustrating an example of field of view information. In the example shown in this figure, the field of view information is information that divides the image into multiple regions based on the distance from a reference point. Specifically, the 0th region, which includes the reference point, is the region where the image should be clear. The correction information generation unit 124 and the correction processing unit 126 then correct the image so that the regions become progressively less clear and darker as they move from the 0th region to the 1st region, the 2nd region, and so on.
[0016] The field of view setting unit 122 determines the size of at least the zeroth region using the velocity information. For example, if the velocity indicated by the velocity information is low, the zeroth region is enlarged and the other regions are narrowed. In addition to the size of each region, the field of view setting unit 122 may also set the number of regions to be set using the velocity information. In this case, the field of view setting unit 122 increases the number of regions to be set as the velocity increases.
[0017] In the example shown in Figure 3, the outlines of each region are rectangles. However, these outlines may be of other shapes (e.g., circles or ellipses).
[0018] Returning to Figure 1, the correction information generation unit 124 generates correction information using the reference point information obtained from the field of view setting unit 122. The correction information is information that specifies the amount of correction for the value of each pixel in the frame image. The field of view angle information is generated using the velocity information and the distance from the reference point, as described above. Therefore, the field of view setting unit 122 essentially generates correction information using the relative position from the reference point and the velocity information.
[0019] In detail, the correction information includes brightness information, resolution information, and saturation information. The brightness information indicates a change in the brightness of at least a portion of the image, the resolution information indicates a change in the resolution of at least a portion of the image, and the saturation information indicates a change in the saturation of at least a portion of the image. The correction information generation unit 124 then outputs the field of view information, reference point information, and correction information to the correction processing unit 126.
[0020] Figure 4 illustrates brightness information, Figure 5 illustrates resolution information, and Figure 6 illustrates saturation information. As shown in these figures, the correction information generated by the correction information generation unit 124 indicates that brightness, resolution, and saturation are all reduced as the distance from the reference point increases. Specifically, for brightness, resolution, and saturation, a correction amount is set for each value of k in the "region k" shown in Figure 3. As the value of k increases, brightness, resolution, and saturation all decrease.
[0021] Returning to Figure 1, the correction processing unit 126 acquires a frame image from the input unit 110 and corrects this frame image using field of view information, reference point information, and correction information. Specifically, the correction processing unit 126 defines reference points within the frame image using the reference point information. Then, using the reference points and field of view information, the correction processing unit 126 divides the frame image into the regions shown in Figure 3. Finally, the correction processing unit 126 performs corrections on each region according to the correction information.
[0022] Figure 7 shows an example of the functional configuration of the correction processing unit 126. The correction processing unit 126 includes a resolution correction unit 202, a saturation correction unit 204, and a brightness correction unit 206. The resolution correction unit 202 corrects the resolution of the frame image region by region using the resolution information contained in the correction information. The saturation correction unit 204 corrects the saturation of the frame image region by region using the saturation information contained in the correction information. The brightness correction unit 206 corrects the brightness of the frame image region by region using the brightness information contained in the correction information. In the example shown in this figure, the resolution correction unit 202, the saturation correction unit 204, and the brightness correction unit 206 are arranged in series in this order, but their arrangement is not limited to the example shown in Figure 7.
[0023] The correction unit 120 then outputs the corrected frame image (corrected image) to the spleness estimation unit 130.
[0024] <Example of configuration of the significance estimation unit 130> Figure 8 is a block diagram illustrating an example configuration of the spleniality estimation unit 130. The spleniality estimation unit 130 generates spleniality estimation information by inputting corrected frame images into a model generated by machine learning. In detail, the spleniality estimation unit 130 comprises an input unit 310, a nonlinear mapping unit 320, and an output unit 330. The input unit 310 converts the input frame image (hereinafter referred to as "image" in the description of the spleniality estimation unit 130) into intermediate data that can be mapped. The nonlinear mapping unit 320 converts the intermediate data into mapping data. The output unit 330 generates spleniality estimation information based on the mapping data. The nonlinear mapping unit 320 further comprises a feature extraction unit 321 that extracts features from the intermediate data and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. These will be explained in detail below.
[0025] Figure 9(a) is an example of an image input to the spleness estimation unit 130, and Figure 9(b) is an example of an image showing the estimated spleness distribution for Figure 9(a). For explanatory purposes, these figures show the frame image before correction by the correction unit 120. The spleness estimation unit 130 in this configuration example estimates the spleness of each part of the image. Spleness refers to, for example, how easily something stands out or how easily it attracts attention. Specifically, spleness is expressed as a probability, etc. Here, the magnitude of the probability corresponds, for example, to the probability that the gaze of a person looking at the image will be directed to that location.
[0026] Figures 9(a) and 9(b) correspond to each other in terms of position. Furthermore, in Figure 9(a), locations with higher sampling are displayed with higher brightness in Figure 9(b). An image showing a sampling distribution like Figure 9(b) is an example of the sampling estimation information output by the output unit 330. In this example, sampling is visualized using 256 levels of brightness values. An example of the sampling estimation information output by the output unit 330 will be described in detail later.
[0027] The estimation results of the sampling distribution can be used in various fields, such as predicting the gaze of traffic participants like drivers and pedestrians, preventing traffic participants from being overlooked, evaluating the visual appeal of content such as advertising media, guiding the viewer's gaze, digitizing the know-how of athletes and skilled workers, and understanding the visual perception of living organisms. Furthermore, the sampling estimation unit 130 and processing method according to this example configuration can be applied to mobility fields such as autonomous driving, advanced driver-assistance systems (ADAS), and road traffic systems; entertainment fields such as virtual reality (VR), augmented reality (AR), and games; content fields such as documents, video content, and signage; and medical fields such as diagnostic imaging, surgical assistance, and nursing care services.
[0028] Figure 10 is a flowchart illustrating a processing method according to the first configuration example. The processing method according to this configuration example is a processing method executed by a computer and includes an input step S110, a nonlinear mapping step S120, and an output step S130. In the input step S110, an image is converted into intermediate data that can be mapped. In the nonlinear mapping step S120, the intermediate data is converted into mapping data. In the output step S130, splendor estimation information showing a splendor distribution is generated based on the mapping data. Here, the nonlinear mapping step S120 includes a feature extraction step S121 that extracts features from the intermediate data, and an upsampling step S122 that upsamples the data generated in the feature extraction step S121. The processing method according to this configuration example is realized by the splendor estimation unit 130 according to the configuration example.
[0029] Returning to Figure 8, let's explain each component of the sampling estimation unit 130. In input step S110, the input unit 310 acquires an image and converts it into intermediate data. The input unit 310 acquires an image from the correction unit 120. The input unit 310 then converts the acquired image into intermediate data. The intermediate data is not particularly limited as long as it is data that the nonlinear mapping unit 320 can accept, but for example, it is a high-dimensional tensor. The intermediate data is, for example, data in which the brightness has been normalized for the acquired image, or data in which each pixel of the acquired image has been converted into a brightness slope. In input step S110, the input unit 310 may further perform image noise reduction, resolution conversion, etc.
[0030] In the nonlinear mapping step S120, the nonlinear mapping unit 320 acquires intermediate data from the input unit 310. Then, the intermediate data is converted into mapping data in the nonlinear mapping unit 320. Here, the mapping data is, for example, a high-dimensional tensor. The mapping process applied to the intermediate data by the nonlinear mapping unit 320 is a mapping process that can be controlled by parameters, for example, and is preferably a process using a function, a functional, or a neural network.
[0031] Figure 11 is a diagram illustrating the configuration of the nonlinear mapping unit 320 in detail, and Figure 12 is a diagram illustrating the configuration of the intermediate layer 323. As described above, the nonlinear mapping unit 320 includes a feature extraction unit 321 and an upsampling unit 322. The feature extraction step S121 is performed in the feature extraction unit 321, and the upsampling step S122 is performed in the upsampling unit 322. In the example shown in this figure, at least one of the feature extraction unit 321 and the upsampling unit 322 is configured to include a neural network containing multiple intermediate layers 323. In the neural network, multiple intermediate layers 323 are connected.
[0032] In particular, the neural network is preferably a convolutional neural network. Specifically, each of the multiple hidden layers 323 includes one or more convolutional layers 324. In the convolutional layers 324, the input data is convolved by multiple filters 325, and the outputs of the multiple filters 325 are subjected to activation processing.
[0033] In the example shown in Figure 11, the feature extraction unit 321 is configured to include a neural network containing multiple hidden layers 323, with a first pooling unit 326 between the hidden layers 323. The upsampling unit 322 is also configured to include a neural network containing multiple hidden layers 323, with an unpooling unit 328 between the hidden layers 323. Furthermore, the feature extraction unit 321 and the upsampling unit 322 are connected to each other via a second pooling unit 327 that performs overlap pooling.
[0034] In the example shown in this figure, each intermediate layer 323 consists of two or more convolutional layers 324. However, at least some of the intermediate layers 323 may consist of only one convolutional layer 324. Adjacent intermediate layers 323 are separated by one of the first pooling section 326, the second pooling section 327, and the unpooling section 328. Here, if an intermediate layer 323 contains two or more convolutional layers 324, it is preferable that the number of filters 325 in those convolutional layers 324 are equal to each other.
[0035] In this diagram, the intermediate layer 323 labeled "A×B" consists of B convolutional layers 324, and each convolutional layer 324 contains A convolutional filters for each channel. Such an intermediate layer 323 will also be referred to as an "A×B intermediate layer" below. For example, a 64×2 intermediate layer 323 consists of two convolutional layers 324, and each convolutional layer 324 contains 64 convolutional filters for each channel.
[0036] In the example shown in this figure, the feature extraction unit 321 includes a 64×2 intermediate layer 323, a 128×2 intermediate layer 323, a 256×3 intermediate layer 323, and a 512×3 intermediate layer 323 in that order. The upsampling unit 322 also includes a 512×3 intermediate layer 323, a 256×3 intermediate layer 323, a 128×2 intermediate layer 323, and a 64×2 intermediate layer 323 in that order. The second pooling unit 327 connects two 512×3 intermediate layers 323 to each other. The number of intermediate layers 323 constituting the nonlinear mapping unit 320 is not particularly limited and can be determined, for example, according to the number of pixels in the image data.
[0037] Note that this figure shows only one example of the configuration of the nonlinear mapping unit 320, and the nonlinear mapping unit 320 may have other configurations. For example, a 64×1 hidden layer 323 may be included instead of a 64×2 hidden layer 323. Reducing the number of convolutional layers 324 included in the hidden layer 323 may further reduce the computation cost. Also, for example, a 32×2 hidden layer 323 may be included instead of a 64×2 hidden layer 323. Reducing the number of channels in the hidden layer 323 may further reduce the computation cost. Furthermore, both the number of convolutional layers 324 and the number of channels in the hidden layer 323 may be reduced.
[0038] Here, in the multiple intermediate layers 323 included in the feature extraction unit 321, it is preferable that the number of filters 325 increases each time the signal passes through the first pooling unit 326. Specifically, the first intermediate layer 323a and the second intermediate layer 323b are continuous with each other via the first pooling unit 326, with the second intermediate layer 323b located after the first intermediate layer 323a. The first intermediate layer 323a is composed of a convolutional layer 324 with N1 filters 325 for each channel, and the second intermediate layer 323b is composed of a convolutional layer 324 with N2 filters 325 for each channel. In this case, it is preferable that N2 > N1 holds. It is even more preferable that N2 = N1 × 2 holds.
[0039] Also, in the plurality of intermediate layers 323 included in the upsampling unit 322, it is preferable that the number of filters 325 decreases every time it passes through the unpooling unit 328. Specifically, the third intermediate layer 323c and the fourth intermediate layer 323d are continuously connected to each other via the unpooling unit 328, and the fourth intermediate layer 323d is located at the subsequent stage of the third intermediate layer 323c. The third intermediate layer 323c is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N3, and the fourth intermediate layer 323d is composed of a convolutional layer 324 in which the number of filters 325 for each channel is N4. At this time, it is preferable that N4 < N3 holds. More preferably, N3 = N4 × 2 holds.
[0040] In the feature extraction unit 321, image features having a plurality of levels of abstraction, such as gradients and shapes, are extracted from the intermediate data acquired from the input unit 310 as channels of the intermediate layer 323. FIG. 12 illustrates the configuration of the 64×2 intermediate layer 323. Referring to this figure, the processing in the intermediate layer 323 will be described. In the example of this figure, the intermediate layer 323 is composed of a first convolutional layer 324a and a second convolutional layer 324b, and each convolutional layer 324 includes 64 filters 325. In the first convolutional layer 324a, convolutional processing using the filter 325 is performed on each channel of the data input to the intermediate layer 323. For example, when the image input to the input unit 310 is an RGB image, processing is performed on each of the three channels h0i (i = 1..3). Also, in the example of this figure, the filter 325 is a 64 types of 3×3 filters, that is, a total of 64×3 types of filters. As a result of the convolutional processing, 64 results h0i,j (i = 1..3, j = 1..64) are obtained for each channel i.
[0041] Next, the output of the multiple filters 325 is activated in the activation unit 329. Specifically, for the corresponding result j of all channels, the activation process is applied to the sum of the corresponding elements. This activation process yields the 64-channel result h1i (i=1..64), i.e., the output of the first convolutional layer 324a, as image features. The activation process is not particularly limited, but it is preferable to use at least one of a hyperbolic function, a sigmoid function, and a normalized linear function.
[0042] Furthermore, the output data of the first convolutional layer 324a is used as input data for the second convolutional layer 324b, and the same processing as in the first convolutional layer 324a is performed in the second convolutional layer 324b to obtain the 64-channel result h2i (i=1..64), that is, the output of the second convolutional layer 324b, as image features. The output of the second convolutional layer 324b becomes the output data of this 64×2 intermediate layer 323.
[0043] Here, the structure of the filter 325 is not particularly limited, but it is preferably a 3x3 two-dimensional filter. Also, the coefficients of each filter 325 can be set independently. In this configuration example, the coefficients of each filter 325 are stored in the storage unit 390, and the nonlinear mapping unit 320 can read them and use them for processing. Here, the coefficients of multiple filters 325 may be determined based on correction information generated and modified using machine learning. For example, the correction information includes the coefficients of multiple filters 325 as multiple correction parameters. The nonlinear mapping unit 320 can further use this correction information to convert intermediate data into mapping data. The storage unit 390 may be provided in the splenishing estimation unit 130, or it may be provided outside the splenishing estimation unit 130. Also, the nonlinear mapping unit 320 may acquire the correction information from an external source via a communication network.
[0044] Figures 13(a) and 13(b) show examples of convolution processing performed by filter 325, respectively. Both Figures 13(a) and 13(b) show examples of 3x3 convolution. The example in Figure 13(a) is a convolution processing using nearest neighbor elements. The example in Figure 13(b) is a convolution processing using neighbor elements with a distance of two or more. It is also possible to perform convolution processing using neighbor elements with a distance of three or more. Filter 325 prefers to perform convolution processing using neighbor elements with a distance of two or more, because it can extract a wider range of features and further improve the accuracy of sampling estimation.
[0045] The operation of the 64x2 hidden layer 323 has been explained above. The operation of other hidden layers 323 (128x2 hidden layer 323, 256x3 hidden layer 323, and 512x3 hidden layer 323, etc.) is the same as that of the 64x2 hidden layer 323, except for the number of convolutional layers 324 and the number of channels. Furthermore, the operation of the hidden layer 323 in the feature extraction unit 321 and the operation of the hidden layer 323 in the upsampling unit 322 are the same as described above.
[0046] Figure 14(a) is a diagram illustrating the process of the first pooling section 326, Figure 14(b) is a diagram illustrating the process of the second pooling section 327, and Figure 14(c) is a diagram illustrating the process of the unpooling section 328.
[0047] In the feature extraction unit 321, the data output from the intermediate layer 323 is subjected to pooling processing for each channel in the first pooling unit 326 before being input to the next intermediate layer 323. In the first pooling unit 326, for example, non-overlapping pooling processing is performed. Figure 14(a) shows the process of associating a 2x2 array of four elements 30 with one element 30 for each group of elements contained in each channel. In the first pooling unit 326, such association is performed for all elements 30. Here, the 2x2 array of four elements 30 are selected so as not to overlap with each other. In this example, the number of elements in each channel is reduced to one-quarter. Note that as long as the number of elements is reduced in the first pooling unit 326, the number of elements 30 before and after the association is not particularly limited.
[0048] The data output from the feature extraction unit 321 is input to the upsampling unit 322 via the second pooling unit 327. In the second pooling unit 327, overlap pooling is applied to the output data from the feature extraction unit 321. Figure 14(b) shows the process of associating 2x2 four elements 30 with one element 30 while overlapping some of the elements 30. That is, in repeated associations, some of the 2x2 four elements 30 in one association are also included in the 2x2 four elements 30 in the next association. In the second pooling unit 327 as shown in this figure, the number of elements is not reduced. Note that the number of elements 30 before and after the association in the second pooling unit 327 is not particularly limited.
[0049] The methods of each process performed in the first pooling unit 326 and the second pooling unit 327 are not particularly limited, but examples include matching the maximum value of four elements 30 to one element 30 (max pooling) and matching the average value of four elements 30 to one element 30 (average pooling).
[0050] The data output from the second pooling unit 327 is input to the intermediate layer 323 in the upsampling unit 322. The output data from the intermediate layer 323 of the upsampling unit 322 is then subjected to amplification processing for each channel in the amplification unit 328 before being input to the next intermediate layer 323. Figure 14(c) shows the process of expanding one element 30 into multiple elements 30. The method of amplification is not particularly limited, but one example is the method of duplicating one element 30 into four elements 30 in a 2x2 arrangement.
[0051] The output data from the last intermediate layer 323 of the upsampling unit 322 is output as mapping data from the nonlinear mapping unit 320 and input to the output unit 330. In the output step S130, the output unit 330 generates splendor estimation information by performing operations such as normalization and resolution conversion on the data acquired from the nonlinear mapping unit 320, and outputs it. The splendor estimation information is, for example, an image (image data) that visualizes splendor using luminance values, as illustrated in Figure 9(b). The splendor estimation information may also be, for example, an image color-coded according to splendor, such as a heat map, or an image in which splendor regions with splendor higher than a predetermined standard are marked in a way that makes them distinguishable from other locations. Furthermore, the splendor estimation information is not limited to images, but may also be a table or the like that lists information indicating splendor regions.
[0052] The sampling estimation information output from the output unit 330 may be subjected to various computer vision processes, such as image segmentation, object recognition, and image classification, either within or outside the sampling estimation unit 130.
[0053] <Example Hardware Configuration> Figure 15 is a block diagram illustrating the hardware configuration of the sampling estimation device 10 shown in Figure 1. The sampling estimation device 10 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, an input / output interface 1050, and a network interface 1060.
[0054] Bus 1010 is a data transmission path for the processor 1020, memory 1030, storage device 1040, input / output interface 1050, and network interface 1060 to send and receive data to and from each other. However, the method of connecting the processor 1020 and the other components to each other is not limited to bus connection.
[0055] Processor 1020 is a processor implemented in components such as the CPU (Central Processing Unit) and GPU (Graphics Processing Unit).
[0056] Memory 1030 is a main memory device implemented using RAM (Random Access Memory), etc.
[0057] The storage device 1040 is an auxiliary storage device implemented as an HDD (Hard Disk Drive), SSD (Solid State Drive), memory card, or ROM (Read Only Memory). The storage device 1040 stores program modules that implement each function of the sampling estimation device 10. The processor 1020 reads these program modules into the memory 1030 and executes them, thereby realizing each function corresponding to that program module.
[0058] The input / output interface 1050 is an interface for connecting the sampling estimation device 10 with various input / output devices.
[0059] The network interface 1060 is an interface for connecting the sampling estimation device 10 to a network. This network may be, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1060 may connect to the network via a wireless connection or a wired connection.
[0060] As described above, according to this embodiment, the brightness of the frame image is changed using speed information before the sampling estimation process. Therefore, the sampling estimation unit 130 can detect with high accuracy the regions that a moving person would perceive as highly sampling from each frame image constituting the video.
[0061] Furthermore, the sampling estimation unit 130 includes a feature extraction unit 321 that extracts features from intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, sampling can be estimated with low computational cost.
[0062] Furthermore, the sampling estimation device 10 can also process still images. In this case as well, the effects described above can be obtained.
[0063] (Second embodiment) Figure 16 shows the functional configuration of the spleness estimation device 10 according to the second embodiment. The spleness estimation device 10 according to this embodiment has the same configuration as the spleness estimation device 10 according to the first embodiment, except that it includes a speed estimation unit 140.
[0064] The velocity estimation unit 140 estimates the speed of movement by processing the video data input to the input unit 110. Existing algorithms, such as optical flow estimation, can be used as the speed estimation algorithm. The velocity estimation unit 140 then outputs the estimated speed as speed information to the field of view setting unit 122.
[0065] This embodiment also provides the same effects as the first embodiment. Furthermore, since the speed estimation unit 140 generates speed information, there is no need to input speed information from an external source.
[0066] (Third embodiment) Figure 17 shows the functional configuration of the splendor estimation device 10 according to the third embodiment. The splendor estimation device 10 according to this embodiment has the same configuration as the splendor estimation device 10 according to the second embodiment, except that it includes a reference point setting unit 150.
[0067] The reference point setting unit 150 generates reference point information by processing at least one frame image of the video data input to the input unit 110.
[0068] For example, if the frame image contains a predetermined object, such as a specific traffic sign or other object that is easily noticeable to the human eye, the reference point setting unit 150 sets the position of that object as a reference point. This object detection is performed, for example, by feature matching processing. The features used here are stored in advance in the spleness estimation device 10.
[0069] Furthermore, the reference point setting unit 150 sets a vanishing point as a reference point if the frame image includes a road and the length of the straight portion of that road is greater than or equal to a standard. Existing algorithms, such as the Hough transform, can be used to detect the vanishing point.
[0070] The reference point setting unit 150 may also set a reference point by performing processing according to specific conditions when the landscape shown in the frame image satisfies those conditions. For example, as shown in Figures 18(a) and (b), if the frame image includes a road and that road is curved, the portion of the road located in the direction of the curve from the center (e.g., the median line) is set as the reference point.
[0071] This embodiment also provides the same effects as the second embodiment. Furthermore, since it has a reference point setting unit 150, there is no need to input reference point information from an external source.
[0072] (Fourth embodiment) Figure 19 shows the functional configuration of the spleness estimation device 10 according to the fourth embodiment. The spleness estimation device 10 according to this embodiment has the same configuration as the spleness estimation device 10 according to the first embodiment, except that it includes a trimming unit 160.
[0073] The trimming unit 160 obtains the video data generation conditions acquired by the input unit 110 and trims the frame image using the generation conditions. The correction unit 120 processes the frame image trimmed by the trimming unit 160. The generation conditions input to the trimming unit 160 are, for example, the type of lens (wide-angle lens or fisheye lens) of the camera that generated the video data. Depending on the video data generation conditions, the range of the scenery captured in the frame image may be wider than the field of view of a stationary person. The trimming unit 160 trims the frame image to match the range of the scenery captured in the frame image to the field of view of a stationary person. The range to be trimmed from the frame image is, for example, stored in advance by the trimming unit 160 for each generation condition.
[0074] This embodiment also provides the same effects as the first embodiment. Furthermore, the trimming unit 160 trims the frame image to match the field of view of a stationary person to the range of the scenery depicted in the frame image. As a result, it is possible to detect areas from the image that would be perceived as highly prominent by a moving person with even higher accuracy.
[0075] Furthermore, the significantness estimation device 10 shown in the second or third embodiment may be provided with the trimming unit 160 according to this embodiment.
[0076] (Fifth embodiment) The prominentness estimation device 10 according to this embodiment has the same configuration as the prominentness estimation device 10 according to any of the above embodiments, except for the functional configuration of the prominentness estimation unit 130.
[0077] Figure 20 is a diagram illustrating the configuration of the spleness estimation unit 130 according to this embodiment. The spleness estimation unit 130 according to this embodiment is the same as the spleness estimation unit 130 according to the first embodiment, except that it further comprises an error calculation unit 340 and a correction unit 350. The error calculation unit 340 uses the spleness estimation information generated for the image and the measured spleness information showing the spleness distribution measured for the image to calculate the error between the spleness distribution shown by the spleness estimation information and the spleness distribution shown by the measured spleness information. The correction unit 350 then corrects the correction information based on the calculated error.
[0078] The spleness estimation unit 130 according to this embodiment performs estimation and learning operations. In the estimation operation, spleness estimation information for the input image is generated and output. The estimation operation is the same as described in the first embodiment. In particular, in this embodiment, the nonlinear mapping unit 320 converts intermediate data into mapping data using correction information. On the other hand, in the learning operation, machine learning is performed using the training image and the measured spleness information for the training image, and correction information is generated or modified (updated). The correction information is information used by the nonlinear mapping unit 320 and includes, for example, a plurality of correction parameters.
[0079] In this embodiment, the nonlinear mapping unit 320 converts intermediate data into mapping data using correction information. The correction information is information that has been generated and modified using machine learning. Specifically, the nonlinear mapping unit 320 includes a plurality of filters 325 as described in the first embodiment, and the coefficients of the plurality of filters 325 are determined based on the correction information. For example, the correction information includes the coefficients of the plurality of filters 325 as a plurality of correction parameters.
[0080] Figure 21 is a flowchart illustrating the learning operation according to this embodiment. The learning operation will be described in detail below. For the learning operation, a training image and splendor measurement information for that training image are prepared. For example, the training image and the splendor measurement information are associated with each other and stored in the storage unit 390. The input unit 310 and the error calculation unit 340 can read and use this information from the storage unit 390.
[0081] The teacher image is any image, such as a photograph. The splendor measurement information is generated based on the results of measuring, for example, the gaze of a person when they look at the teacher image using an eye tracker. The splendor measurement information can take the same form as the splendor estimation information. That is, the splendor measurement information may be an image that visualizes splendor using luminance values, or it may be an image that is color-coded according to splendor, such as a heat map.
[0082] In the learning operation, the input step S110, the nonlinear mapping step S120, and the output step S130 are performed in the same manner as in the input step S110, the nonlinear mapping step S120, and the output step S130 according to the first embodiment. However, the image acquired by the input unit 310 in the input step S110 is a training image. In the nonlinear mapping step S120, the nonlinear mapping unit 320 reads correction information from the storage unit 390. Then, it uses the correction information to convert the intermediate data into mapping data. The nonlinear mapping unit 320 may also directly acquire the correction information from the correction unit 350 instead of reading it from the storage unit 390. In addition, the correction parameters included in the correction information in the initial state can be any value.
[0083] Next, in the error calculation step S140, the error calculation unit 340 acquires splendor estimation information from the output unit 330. The error calculation unit 340 also acquires actual splendor measurement information associated with the training image that was the source of the splendor estimation information. The error calculation unit 340 then calculates the error between the acquired splendor estimation information and the actual splendor measurement information. The method for calculating the error is not particularly limited, but it is preferable to calculate at least one of the following: L1 distance, L2 distance (Euclidean distance, mean squared error), Kullback-Leibler distance, Jensen-Shannon distance, and Pearson correlation coefficient.
[0084] Specifically, the Euclidean distance is calculated using equation (1) below, the Kullback-Leibler distance using equation (2) below, and the Jensen-Shannon distance using equation (3) below. Here, pi represents the estimated result (a value based on sampling estimation information), and qi represents the true value (a value based on observed sampling information).
[0085]
number
number
number
[0086] Next, in the correction step S150, the correction unit 350 obtains the error from the error calculation unit 340 and modifies the correction parameters so that this error is reduced. Then, the correction parameters held in the storage unit 390 are replaced with the modified correction parameters. Here, the method for modifying the correction parameters is not particularly limited, but it is preferable to use at least one of the following methods: least squares method, quadratic programming, stochastic gradient descent (SGD), adaptive moment estimation (ADAM), and variational method.
[0087] In this case, there are many correction parameters that need to be modified, and in order to efficiently determine their values and estimate sampling with high accuracy, it is preferable to use statistical learning (machine learning) with a large amount of training data. Therefore, in the learning operation, it is preferable that machine learning is performed through the cooperation of the nonlinear mapping unit 320, the error calculation unit 340, and the correction unit 350.
[0088] Alternatively, the correction unit 350 may output the corrected correction parameters directly to the nonlinear mapping unit 320 instead of replacing the correction parameters held in the memory unit 390 with the corrected correction parameters. In the next nonlinear mapping step S120, the nonlinear mapping unit 320 performs processing using the corrected correction parameters.
[0089] Note that a single training image may be associated with one or more measured splendor information items. When multiple measured splendor information items are associated with a single training image, these items are based on different measurement results. The error calculation unit 340 then calculates the error between the estimated splendor information and each measured splendor information item. The correction unit 350 then modifies the correction parameters, for example, so that the sum of all errors is reduced.
[0090] The learning process may be performed on multiple sets of training images and measured spleness information. Repeated learning processes further improve the accuracy of spleness estimation.
[0091] The timing of the learning operation is not particularly limited. For example, the sampling estimation unit 130 can accept an operation from the user to start the learning operation. Based on this operation, the sampling estimation unit 130 can start the learning operation. The sampling estimation unit 130 can also terminate the learning operation based on an operation from the user to terminate it or based on predetermined termination conditions. Examples of termination conditions include fulfilling a predetermined number of iterations of the learning operation or the error falling below a predetermined threshold value.
[0092] As described above, according to this embodiment, similar to the first embodiment, the nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, sampling can be estimated with low computational cost.
[0093] Furthermore, according to this embodiment, the splendor estimation unit 130 includes an error calculation unit 340 and a correction unit 350. Therefore, more accurate splendor estimation is achieved using the correction information corrected by the learning operation.
[0094] (Sixth embodiment) The prominentness estimation device 10 according to this embodiment has the same configuration as the prominentness estimation device 10 according to any of the above embodiments, except for the functional configuration of the prominentness estimation unit 130.
[0095] Figure 22 is a diagram illustrating the configuration and operating environment of the arithmetic unit 40 according to this embodiment. The arithmetic unit 40 according to this embodiment is a device that generates correction information used in the spleness estimation unit 130. The arithmetic unit 40 comprises an error calculation unit 440 and a correction unit 450. The error calculation unit 440 uses the spleness estimation information generated for the training image and the measured spleness information showing the spleness distribution measured for the training image to calculate the error between the spleness distribution shown by the spleness estimation information and the spleness distribution shown by the measured spleness information. The correction unit 450 calculates correction information based on the error.
[0096] The sampling estimation unit 130 according to this embodiment is the same as the sampling estimation unit 130 according to the first embodiment. The sampling estimation unit 130 according to this embodiment includes an input unit 310, a nonlinear mapping unit 320, and an output unit 330. Furthermore, the sampling estimation unit 130 according to this embodiment does not need to include the error calculation unit 340 and correction unit 350 described in the fifth embodiment. The input unit 310 according to this embodiment is the same as the input unit 310 according to at least one of the first and fifth embodiments, the nonlinear mapping unit 320 according to this embodiment is the same as the nonlinear mapping unit 320 according to at least one of the first and fifth embodiments, and the output unit 330 according to this embodiment is the same as the output unit 330 according to at least one of the first and fifth embodiments. The operation of the error calculation unit 440 according to this embodiment is the same as the operation of the error calculation unit 340 according to the fifth embodiment, and the operation of the correction unit 450 according to this embodiment is the same as the operation of the correction unit 350 according to the fifth embodiment. The sampling estimation unit 130 and the arithmetic unit 40 cooperate to perform the learning and estimation operations described in the fifth embodiment. Furthermore, the sampling estimation unit 130 and the arithmetic unit 40 may be physically separated, or they may be connected to each other, for example, via a communication network.
[0097] Furthermore, in the learning operation according to this embodiment, it is preferable that machine learning is performed through the cooperation of the nonlinear mapping unit 320, the error calculation unit 440, and the correction unit 450.
[0098] Alternatively, the output unit 330 may temporarily store the generated sampling estimation information in the storage unit 390, and the error calculation unit 440 may read the sampling estimation information stored in the storage unit 390 and use it.
[0099] In the example shown in this figure, the storage unit 390 is provided separately from the spleness estimation unit 130 and the arithmetic unit 40, but this is not limited to this example, and the storage unit 390 may be provided in the spleness estimation unit 130 or in the arithmetic unit 40. When the storage unit 390 is provided inside the arithmetic unit 40, for example, the storage unit 390 is implemented using the storage device 1080 of the computer 1000 that implements the arithmetic unit 40. Alternatively, the storage unit 390 may be implemented through the cooperation of the storage device 1080 of the computer 1000 that implements the spleness estimation unit 130 and the storage device 1080 of the computer 1000 that implements the arithmetic unit 40.
[0100] As described above, according to this embodiment, similar to the first embodiment, the nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, sampling can be estimated with low computational cost.
[0101] Furthermore, according to this embodiment, the arithmetic unit 40 includes an error calculation unit 440 and a correction unit 450. Therefore, more accurate splendor estimation can be achieved using the correction information corrected by the learning operation.
[0102] (Seventh Embodiment) The prominentness estimation device 10 according to this embodiment has the same configuration as the prominentness estimation device 10 according to any of the above embodiments, except for the functional configuration of the prominentness estimation unit 130.
[0103] Figure 23 is a diagram illustrating the configuration of the distinctiveness estimation unit 130 according to this embodiment. The distinctiveness estimation unit 130 according to this embodiment is the same as the distinctiveness estimation unit 130 according to at least one of the first and fifth embodiments, except that it further comprises a synthesis unit 360 and a display unit 380.
[0104] The synthesis unit 360 generates synthesized information by combining the sampling distribution shown by the sampling estimation information with the image input to the input unit 310 (input image). Specifically, the synthesis unit 360 acquires the sampling estimation information from the output unit 330 and, for example, acquires the input image from the storage unit 390. Then, it outputs synthesized information that shows the input image and the sampling distribution together. The synthesized information is output to, for example, the display unit 380 provided in the sampling estimation unit 130. The synthesized information output from the synthesis unit 360 may also be stored in the storage unit 390 or acquired by an external device.
[0105] Figure 24 illustrates an image showing the composite information generated by the compositing unit 360. In this example, the composite information is an image in which the input image and a heatmap indicating saturation are superimposed. The format of the composite information is not particularly limited. For example, the composite information may be an image in which the saturation region is enclosed by a circle or rectangle in the input image. The compositing method is also not particularly limited and can include alpha blending, for example.
[0106] The prominentness estimation unit 130 according to this embodiment can be implemented, for example, in a mobile terminal (smartphone, tablet, etc.) equipped with an imaging device such as a camera. This would allow for the extraction of highly prominent and important objects on the spot while taking pictures with the mobile terminal, and enable visualization with good visibility.
[0107] As described above, according to this embodiment, similar to the first embodiment, the nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, sampling can be estimated with low computational cost.
[0108] In addition, according to this embodiment, the sampling estimation unit 130 further comprises a synthesis unit 360. Therefore, the sampling at each position in the image can be visualized with good visibility.
[0109] (Eighth embodiment) The prominentness estimation device 10 according to this embodiment has the same configuration as the prominentness estimation device 10 according to any of the above embodiments, except for the functional configuration of the prominentness estimation unit 130.
[0110] Figure 25 is a diagram illustrating the configuration of the sampling estimation unit 130 according to this embodiment. The sampling estimation unit 130 according to this embodiment is the same as the sampling estimation unit 130 according to at least one of the first, fifth, and seventh embodiments, except that it further comprises a mask image generation unit 370, a region extraction unit 372, and an object detection unit 374.
[0111] The mask image generation unit 370 acquires spleness estimation information from the output unit 330 and generates a mask image. Specifically, the mask image generation unit 370 generates a mask image in which the region where the spleness is lower than a predetermined standard in the spleness distribution shown by the spleness estimation information is designated as a mask region, and the region where the spleness is equal to or greater than the predetermined standard is designated as an unmasked region. In other words, the mask image generation unit 370 performs binarization of the spleness distribution. Here, the standard is set in advance and stored in the storage unit 390, and the mask image generation unit 370 can read it and use it.
[0112] The region extraction unit 372 acquires an input image and a mask image. Then, by applying the mask image to the input image, it extracts regions with high prominence from the input image. For example, the region extraction unit 372 can extract regions with high prominence from the input image by performing logical operations on the input image and the mask image.
[0113] The object detection unit 374 then detects objects from the regions extracted by the region extraction unit 372. The method of object detection is not particularly limited, but one example is the use of a Single Shot Multibox Detector (SSD). In the prominentness estimation unit 130 of this embodiment, regions with high prominentness are extracted in advance, and object detection is performed only in the extracted regions, thereby suppressing false detections.
[0114] The distinctiveness estimation unit 130 according to this embodiment is mounted on a moving object such as an automobile. The object detection results from the object detection unit 374 can then be used for autonomous driving or driver assistance.
[0115] As described above, according to this embodiment, similar to the first embodiment, the nonlinear mapping unit 320 includes a feature extraction unit 321 that extracts features from intermediate data, and an upsampling unit 322 that upsamples the data generated by the feature extraction unit 321. Therefore, sampling can be estimated with low computational cost.
[0116] In addition, according to this embodiment, the sampling estimation unit 130 further comprises a mask image generation unit 370, a region extraction unit 372, and an object detection unit 374. Therefore, high-precision object detection can be performed in the input image.
[0117] The embodiments and examples described above with reference to the drawings are illustrative examples of the present invention, and various other configurations can also be adopted. [Explanation of symbols]
[0118] 10. Sampling Estimation Device 110 Input Section 120 Correction section 122 Field of View Setting Section 124 Correction information generation unit 126 Correction Processing Unit 130 Sampling Estimation Unit 140 Speed estimation part 150 Reference point setting section 160 Trimming section 202 Resolution Correction Section 204 Saturation correction section 206 Brightness Correction Unit
Claims
1. A correction unit generates a corrected image by correcting an image of the scenery viewed from a first viewpoint, A spleness estimation unit that processes the corrected image to generate spleness estimation information showing the spleness distribution within the corrected image or within the image, Equipped with, The correction unit, Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using velocity information relating to the movement speed of the first viewpoint and the relative position from the reference point to at least the portion in the image. A spleness estimation device that generates spleness estimation information based on a corrected image obtained by correcting at least a portion of the brightness using the brightness information.
2. In the prominentness estimation device according to claim 1, The correction unit, Furthermore, resolution information indicating a change in at least some of the resolution, and saturation information indicating a change in at least some of the saturation are generated using the velocity information and the relative position. A saturation estimation device that corrects the resolution and saturation of at least a portion of the image using the resolution information and saturation information.
3. In the significance estimation device according to claim 1 or 2, The correction unit is a spleness estimation device that generates brightness information such that the brightness decreases as the distance from a reference point in the image to at least a portion of it increases.
4. In the remarkableness estimation device according to any one of claims 1 to 3, The aforementioned image is a frame image included in a video, representing a sampling estimation device.
5. In the remarkableness estimation device according to any one of claims 1 to 4, The correction unit is a spleness estimation device that sets the reference point by processing the image.
6. In the remarkableness estimation device according to any one of claims 1 to 5, The image is further equipped with a trimming unit that trims the image using the image generation conditions, The correction unit is a spleness estimation device that processes the trimmed image to generate the corrected image.
7. In the prominentness estimation device according to any one of claims 1 to 6, The aforementioned spleness estimation unit is a spleness estimation device that generates spleness estimation information by inputting the corrected image into a model generated by machine learning.
8. Computers By correcting the image of the scenery viewed from the first viewpoint, a corrected image is generated. By processing the corrected image, significance estimation information is generated that shows the significance distribution within the corrected image or within the image. Furthermore, the aforementioned computer Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using velocity information relating to the movement speed of the first viewpoint and the relative position from the reference point to at least the portion in the image. A method for estimating splendor, which generates splendor estimation information based on a corrected image obtained by correcting at least a portion of the brightness using the brightness information.
9. On the computer, A correction function that generates a corrected image by correcting an image of the scenery viewed from the first viewpoint, An estimation function that processes the corrected image to generate sampling estimation information showing the sampling distribution within the corrected image or within the image, Give it to him Furthermore, as at least a part of the correction function, A function to generate brightness information indicating a change in brightness of at least a portion of the aforementioned image, using velocity information relating to the movement speed of the first viewpoint, and the relative position from a reference point to at least the aforementioned portion in the image, A function to generate the splendor estimation information based on a corrected image obtained by correcting at least a portion of the brightness using the brightness information, A program to give it a hand.
10. A correction unit that generates a corrected image by correcting the image, A spleness estimation unit that processes the corrected image to generate spleness estimation information showing the spleness distribution within the corrected image or within the image, Equipped with, The correction unit, Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using speed information relating to the movement speed of a first viewpoint, which is the position where the camera that took the image was positioned at the time the image was taken, and reference point information that identifies the position in the image that should be the point of focus, based on the distance from a reference point set in the image. A spleness estimation device that, when generating the corrected image, corrects at least some of the brightness using the brightness information.
11. In the prominentness estimation device according to claim 10, The correction unit, Furthermore, resolution information indicating a change in at least a portion of the resolution, and saturation information indicating a change in at least a portion of the saturation are generated using the velocity information and the relative position from the reference point in the image to at least a portion of the image. A saturation estimation device that corrects the resolution and saturation of at least a portion of the image using the resolution information and saturation information.
12. In the prominentness estimation device according to claim 10 or 11, The correction unit is a spleness estimation device that generates brightness information such that the brightness decreases as the distance from a reference point in the image to at least a portion of it increases.
13. In the remarkableness estimation device according to any one of claims 10 to 12, The aforementioned image is a frame image included in a video, representing a sampling estimation device.
14. In the prominentness estimation device according to claim 11, The correction unit is a spleness estimation device that sets the reference point by processing the image.
15. In the remarkableness estimation device according to any one of claims 10 to 14, The image is further equipped with a trimming unit that trims the image using the image generation conditions, The correction unit is a spleness estimation device that processes the trimmed image to generate the corrected image.
16. In the remarkableness estimation device according to any one of claims 10 to 15, The aforementioned spleness estimation unit is a spleness estimation device that generates spleness estimation information by inputting the corrected image into a model generated by machine learning.
17. Computers By correcting the image, a corrected image is generated. By processing the corrected image, significance estimation information is generated that shows the significance distribution within the corrected image or within the image. When generating the corrected image, Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using speed information relating to the movement speed of a first viewpoint, which is the position where the camera that took the image was positioned at the time the image was taken, and reference point information that identifies the position in the image that should be the point of focus, based on the distance from a reference point set in the image. A method for estimating splendor, which involves correcting at least some of the brightness using the brightness information when generating the corrected image.
18. On the computer, A correction function that generates a corrected image by correcting the image, A spleness estimation function that processes the corrected image to generate spleness estimation information showing the spleness distribution within the corrected image or within the image, Give it to him Furthermore, the correction function is Brightness information indicating a change in brightness of at least a portion of the aforementioned image is generated using speed information relating to the movement speed of a first viewpoint, which is the position where the camera that took the image was positioned at the time the image was taken, and reference point information that identifies the position in the image that should be the point of focus, based on the distance from a reference point set in the image. A program that, when generating the corrected image, corrects at least some of the brightness using the brightness information.