Image-based object estimation device
The object estimation device addresses the challenges of conventional technologies by dividing images into grids, generating diverse learning data, and performing additional learning to update the estimation model, resulting in improved performance and accuracy in object identification.
Patent Information
- Application Number
- JP2021059726
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-03-31
AI Technical Summary
Conventional technologies for object estimation in images using machine learning models face challenges such as lack of diversity in training data, over-learning, and inefficiency in data generation, as well as potential inaccuracies in performance evaluation and data suitability judgment.
An object estimation device that divides images into grids, generates learning data by associating object identification information with grid images, and performs additional learning to update the estimation model based on evaluation data, ensuring accurate and efficient object estimation.
The proposed solution improves the performance of object estimation by generating diverse and accurate training data, preventing over-learning, and ensuring efficient data generation and model updates, leading to more accurate and reliable object identification in images.
Smart Images

Figure 0007681863000001 
Figure 0007681863000002 
Figure 0007681863000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a device capable of estimating the identity of an object appearing in an image. [Background technology]
[0002] Devices that use machine learning models such as deep learning to estimate what objects appear in an image are have been proposed and realized. For example, Patent Literature 1 discloses a device that divides an image into grids and estimates what objects appear in each grid.
[0003] In an estimation device that uses a machine learning model such as deep learning, it is necessary to train the model based on a large amount of training data, and there is a demand for efficiently generating a large amount of appropriate training data.
[0004] In this regard, Patent Document 1 discloses a method of using an image to be estimated (a processed image) and the estimation result as training data for re-learning, thereby enabling efficient generation of training data.
[0005] Furthermore, Patent Document 2 discloses a device that uses training data to determine the accuracy rate of a model that has undergone additional training using additional training data and a model that has not undergone additional training, and that determines, depending on the result, whether to replace an inference means with one that has undergone additional training or to use the one before the additional training.
[0006] Therefore, the inference means can be updated to have higher accuracy through additional learning.
[0007] Patent Document 3 discloses that candidate data is provided to a learning model, and the suitability of adopting the data as learning data is determined based on the magnitude of the estimated probability, and only appropriate learning data is used. This allows more appropriate learning to be performed using only appropriate learning data. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Patent Publication 2019-139497 [Patent Document 2] Patent Publication No. 2020-38538 [Patent Document 3] Patent Publication No. 2020-166397 Summary of the Invention [Problem to be solved by the invention]
[0009] Conventional technologies such as that described in Patent Document 1 have the advantage of being able to automatically generate a large amount of training data, but they have the problem that the obtained training data is not necessarily appropriate in terms of diversity as training data or prevention of over-learning.
[0010] In addition, automatically created learning data has the problem of lacking accuracy because the teaching content is obtained by inference, and on the other hand, if all of this were to be ascribed to teaching content by hand, there is the problem of inefficiency.
[0011] The conventional technology such as that of Patent Document 2 has the advantage that it is possible to perform a performance evaluation using training data for a model that has undergone additional training and a model that has not undergone additional training, and to determine whether or not additional training is advisable. However, since the performance evaluation is performed based on the training data, there is a possibility that the performance evaluation itself will not produce appropriate results.
[0012] In the conventional technology such as that in Patent Document 3, candidate data is given to a learning model, and the suitability of adopting it as learning data is judged based on the magnitude of its estimated probability. However, there is a problem in that the suitability of the data as learning data cannot be appropriately judged based only on the estimated probability by the learning model.
[0013] An object of the present invention is to solve at least any of the above problems and to provide an object estimation device or a learning data generation device with improved performance. [Means for solving the problem]
[0014] Independently applicable features of the present invention are listed below.
[0015] (1)(2)(8)(9) An object estimation device according to the present invention is an image-based object estimation device comprising: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image; and an object estimation means for learning based on learning data in which a learning original image is divided into grids and object identification information appearing in the learning grid image is associated with the learning original image, and for estimating what object appears in the processing grid image obtained by the division means. The object estimation device further comprises a learning data generation means for generating the learning data. The learning data generation means comprises: a learning division means for dividing a learning annotation original image, which has an annotation for distinguishing objects in the captured learning original image, into grids to obtain an annotation grid image; and a learning data acquisition means for obtaining object identification information of the grid based on annotations of the objects included in the annotation grid image, and adding the object identification information to the learning grid image obtained by dividing the learning original image to obtain learning data.
[0016] Therefore, the learning grid image can be generated efficiently.
[0017] (3)(10) The object estimation device according to the present invention is characterized in that it further comprises a learning data generating means for generating, as learning data, a grid image in which the object estimation means has estimated an object with a probability higher than a predetermined value, and object identification information, which is the object estimation result.
[0018] Therefore, the generation of the training grid images can be automated.
[0019] (4)(11) The object estimation device of the present invention further includes an updating means for performing additional learning on the trained object estimation means based on additional learning data and updating the object estimation means in accordance with the results of the additional learning, wherein the updating means uses the object estimation means after learning based on the additional learning data as a provisional estimation means, provides evaluation data to the current object estimation means before the additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the provisional object estimation means, and if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is.
[0020] Therefore, inappropriate additional learning can be eliminated, and additional learning of the object estimation means can be appropriately performed. In addition, since the judgment is performed using evaluation data for appropriate evaluation, it is possible to perform accurate judgment.
[0021] (5)(12) The object estimation device of the present invention further comprises an additional training data generation means, which is characterized in that it comprises a training image feature determination means for dividing each pixel included in each learned image already used in training based on features including its color or density, and determining a representative feature of the learned image based on the number of pixels having each feature, a feature distribution calculation means for calculating the distribution of representative features in the calculated plurality of learned images, a candidate image feature determination means for dividing each pixel included in a candidate image based on features including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature, and a training image selection means for selecting the candidate image as a training image if the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the learned images.
[0022] Therefore, additional learning can be performed so that image features are not biased.
[0023] (6)(13) The object estimation device according to the present invention further includes an additional training data generation means, which includes a trained original image feature determination means for dividing, for each of the trained grid images constituting a trained original image already used for training, each pixel included in the trained grid image based on features including its color or density, and determining a representative feature of the trained grid image based on the number of pixels having each feature; a trained image feature distribution calculation means for calculating a distribution of representative features for each grid position in the calculated plurality of trained original images; a candidate image feature determination means for dividing, for each of the candidate grid images constituting a candidate original image, each pixel included in the candidate grid image based on features including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a training original image selection means for selecting the candidate image as a training original image if the representative feature of each candidate grid image of the candidate images is not contrary to the equalization of the distribution of the representative features of the training images; and a training original image division means for dividing the selected training original image into grids of a predetermined size to obtain a training grid image.
[0024] Therefore, additional learning can be performed so that image features are not biased.
[0025] (7)(14) The object estimation device of the present invention further includes an additional training data generation means, which includes a trained original image frequency calculation means for classifying a plurality of training grid images already used for training by their positions on the training original image and calculating the number of objects appearing in the trained grid image for each position based on the object identification information; a trained image frequency distribution calculation means for calculating a distribution of the number of objects appearing for each grid position in the calculated plurality of trained original images; a candidate original image position identification means for identifying a grid position at which an object appears among candidate grid images constituting a candidate original image; a training image selection means for determining whether or not the candidate original image is to be used as a training image based on the distribution of the number of objects appearing by grid position in the training original image and the grid position at which the object appears in the candidate original image; and a training original image division means for dividing the selected training original image into grids of a predetermined size to obtain a training grid image.
[0026] Therefore, learning can be performed using learning data that is not biased in the positions at which objects appear.
[0027] (15)(16)(20)(21) An object estimation device according to the present invention includes a division means for taking an image captured while driving on a road as a processed image, dividing the processed image into grids of a predetermined size to obtain a processed grid image, an object estimation means which is trained based on a learning grid image obtained by dividing a learning image into grids and learning data in which object identification information appearing in the learning grid image is associated with each other, and which estimates what object is appearing in the processed grid image obtained by the division means, and an object estimation means which performs additional learning on the trained object estimation means based on additional learning data and estimates the object according to the results of the additional learning. In an image-based object estimation device having an updating means for updating the estimation means, the updating means sets the object estimation means after learning based on additional learning data as a provisional estimation means, provides evaluation data to the current object estimation means before additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the provisional object estimation means, and if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is.
[0028] Therefore, inappropriate additional learning can be eliminated, and additional learning of the object estimation means can be performed appropriately. Also, since the judgment is performed using evaluation data for performing an appropriate evaluation, it is possible to perform a highly accurate judgment.
[0029] (17)(22) The object estimation device of the present invention further comprises an additional training data generation means for generating additional training data, the additional training data generation means comprising: a training image representative feature determination means for dividing each pixel included in each training image already used in training based on features including its color or density, and determining a representative feature of the training image based on the number of pixels having each feature; a representative feature distribution calculation means for calculating the distribution of representative features in the calculated plurality of training images; a candidate image representative feature determination means for dividing each pixel included in a candidate image based on features including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; and a training image selection means for selecting the candidate image as a training image if the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the training images.
[0030] Therefore, additional learning can be performed so that image features are not biased.
[0031] (18)(23) The object estimation device according to the present invention further includes an additional training data generation means for generating additional training data, and the additional training data generation means includes a trained original image feature determination means for dividing, for each of the trained grid images constituting a trained original image already used for training, each pixel included in the trained grid image based on features including its color or density, and determining a representative feature of the trained grid image based on the number of pixels having each feature; a trained image feature distribution calculation means for calculating a distribution of representative features for each grid position in the calculated plurality of trained original images; a candidate image feature determination means for dividing, for each of the candidate grid images constituting a candidate original image, each pixel included in the candidate grid image based on features including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a training original image selection means for selecting the candidate image as a training original image if the representative feature of each candidate grid image of the candidate images is not contrary to the equalization of the distribution of the representative features of the training images; and a training original image division means for dividing the selected training original image into grids of a predetermined size to obtain a training grid image.
[0032] Therefore, additional learning can be performed so that image features are not biased.
[0033] (19)(24) The object estimation device of the present invention further includes an additional training data generation means for generating additional training data, and the additional training data generation means includes a trained original image frequency calculation means for classifying a plurality of training grid images already used for training by their positions on the training original image and calculating the number of objects appearing in the trained grid image for each position based on the object identification information, a trained image frequency distribution calculation means for calculating a distribution of the number of objects appearing for each grid position in the calculated plurality of trained original images, a candidate original image position identification means for identifying grid positions at which objects appear among candidate grid images constituting a candidate original image, a training image selection means for determining whether or not the candidate original image is to be a training image based on the distribution of the number of objects appearing by grid position in the training original image and the grid position at which the object appears in the candidate original image, and a training original image division means for dividing the selected training original image into grids of a predetermined size to obtain a training grid image.
[0034] Therefore, learning can be performed using learning data that is not biased in the positions at which objects appear.
[0035] (25)(26)(28(29)The object estimation device according to the present invention is an image-based object estimation device having an object estimation means that is trained based on training data in which a training image and object identification information appearing in the training image are associated with each other, and that estimates what an object appears in a processed image is, and further comprises a training data generation means that generates the training data, and the training data generation means comprises: a training image feature determination means that classifies each pixel included in each learned image that has already been used in training based on a feature including its color or density, and determines a representative feature of the learned image based on the number of pixels having each feature; a feature distribution calculation means that calculates the distribution of the calculated representative features in the plurality of learned images; a candidate image feature determination means that classifies each pixel included in a candidate image based on a feature including its color or density, and determines a representative feature of the candidate image based on the number of pixels having each feature; and a training image selection means that selects the candidate image as a training image if the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the learned images.
[0036] Therefore, additional learning can be performed so that image features are not biased.
[0037] (27)(30) The object estimation device according to the present invention is characterized in that the training images and the candidate images are grid images obtained by dividing an image into grids of a predetermined size.
[0038] Therefore, additional learning can be performed on the grid image so that the image features are not biased.
[0039] (31)(32)(34)(35) An object estimation device according to the present invention is an image-based object estimation device including an original image division means for dividing an image captured while driving on a road into a processing original image and obtaining a processing grid image by dividing the original image into grids of a predetermined size, and an object estimation means for estimating what object is depicted in the processing grid image obtained by dividing the original image into grids and learning data in which object identification information depicted in the learning grid image is associated with each other, and further including a learning data generation means for generating the learning data, wherein the learning data generation means generates the learning data by dividing each pixel included in the learned grid image based on features including its color or density for each of the learned grid images constituting the learned original image already used for learning. the number of pixels having each feature being equal to or less than the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each grid position in the calculated plurality of learned original images; a candidate image feature determination means for dividing each pixel contained in a candidate grid image constituting a candidate original image based on features including its color or density and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning image selection means for selecting a candidate image as a learning image if the representative feature of each candidate grid image of the candidate images is not contrary to the equalization of the distribution of the representative features of the learning images; and a learning image division means for dividing the selected learning image into grids of a predetermined size to obtain a learning grid image.
[0040] Therefore, additional learning can be performed so that image features are not biased.
[0041] (33)(36) The object estimation device of the present invention includes a learned attention grid feature distribution calculation means for selecting a learned attention grid image including an object from among a plurality of learned grid images constituting a plurality of learned original images, and calculating a distribution of representative features of the learned attention grid image, a candidate original image attention grid feature determination means for selecting a candidate attention grid image including an object from among candidate grid images constituting a candidate original image, and calculating representative features of each candidate attention grid image, and a second learning image selection means for determining whether or not to select a candidate original image as a learning original image based on the distribution of representative features of the learned attention grid images and the representative features of the candidate grid images of the candidate original images.
[0042] Therefore, additional learning can be performed without biasing the features of the grid image including the target object.
[0043] (37)(38)(39(40)The object estimation device according to the present invention is an image-based object estimation device comprising: an original image division means for dividing an image captured while driving on a road into a processing original image and obtaining a processing grid image by dividing the original image into grids of a predetermined size; and an object estimation means for learning based on learning grid images obtained by dividing the learning original image into grids and learning data in which object identification information appearing in the learning grid images is associated with each other, and for estimating what object appears in the processing grid image obtained by the division means, further comprising a learning data generation means for generating the learning data, and the learning data generation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and estimating what object appears in the processing grid image obtained by the division means. The system includes a learned original image frequency calculation means for calculating the number of occurrences of objects appearing in a learned grid image for each position based on the specific information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position identification means for identifying grid positions at which objects appear among candidate grid images constituting a candidate original image; a learning image selection means for determining whether or not the candidate original image is to be a learning image based on the distribution of the number of occurrences of the objects by grid position in the learning original image and the grid positions at which the objects appear in the candidate original image; and a learning original image division means for dividing the selected learning original image into grids of a predetermined size to obtain a learning grid image.
[0044] Therefore, learning can be performed using learning data that is not biased in the positions at which objects appear.
[0045] (41)(42) The object estimation device of the present invention is an image-based object estimation device comprising: a division means for dividing an original image captured while driving on a road into grids of a predetermined size to obtain a processing grid image; and an object estimation means for estimating what objects appear in the processing grid image obtained by the division means, which is trained based on learning data that corresponds a learning grid image obtained by dividing the original image into grids and object identification information appearing in the learning grid image, and which estimates what objects appear in the processing grid image obtained by the division means. The object estimation means is configured to comprise an individual object estimation means for estimating whether or not an object is included in the processing grid image for each object, and when multiple objects are included in the processing grid image, each of the corresponding individual object estimation means estimates an object, thereby making it possible to estimate the multiple objects.
[0046] Therefore, even if a grid image contains multiple objects, accurate estimation can be performed.
[0047] In this embodiment, step S2 corresponds to the "division means."
[0048] In this embodiment, step S4 corresponds to the "object estimation means."
[0049] In this embodiment, step S12 corresponds to the "learning division means."
[0050] In the embodiment, steps S13 to S15 correspond to the "learning data acquiring means."
[0051] In this embodiment, step S44 corresponds to the "additional learning means."
[0052] In this embodiment, steps S40 and S46 correspond to the "estimation accuracy determining means."
[0053] In this embodiment, steps S47, S48, and S49 correspond to the "replacement means."
[0054] In the embodiment, steps S52 and S53 correspond to the "learned image feature determining means."
[0055] In this embodiment, step S55 corresponds to the "characteristic distribution calculation means."
[0056] In this embodiment, steps S58 and S59 correspond to the "candidate image feature determining means."
[0057] In the embodiment, the "learning image selection means" corresponds to steps S60, S61, steps S91, S92, and steps S166, S167.
[0058] In the embodiment, steps S73 and S74 correspond to the "learned original image feature determining means."
[0059] In the embodiment, step S76 corresponds to the "trained original image feature distribution calculation means."
[0060] In this embodiment, steps S86 and S87 correspond to the "candidate original image feature determining means."
[0061] In this embodiment, step S152 corresponds to the "learned original image frequency calculation means."
[0062] In the embodiment, step S154 corresponds to the "learned original image frequency distribution calculation means."
[0063] In this embodiment, step S164 corresponds to the "candidate original image frequency calculation means."
[0064] The term "program" is a concept that includes not only programs that can be executed directly by a CPU, but also programs in source format, compressed programs, encrypted programs, and the like. [Brief description of the drawings]
[0065] [Figure 1] 1 illustrates a functional configuration of an object estimation device according to a first embodiment. [Diagram 2] 2 shows a hardware configuration of the object estimation device. [Diagram 3] 13 is a flowchart of an object estimation program. [Figure 4] 1 is an example of a processed image. [Diagram 5] 13 is an example of a processing grid image obtained by dividing a processing image. [Figure 6] FIG. 1 is a diagram showing types of features when features are treated as objects. [Figure 7] 2 is a detailed configuration example of an object estimation means 4. [Figure 8] 13 is an example of associating estimated features with a processed grid image. [Figure 9] 13 is an example in which an image of a feature is associated with feature-specific information. [Figure 10] 13 is a flowchart of learning data generation. [Figure 11] 1 is an example of a learning source image. [Figure 12] This is an example of annotating the original learning image. [Figure 13] This is an example of gridding annotated training source image. [Figure 14] FIG. 13 is a diagram showing feature identification information associated with each learning grid image. [Figure 15] FIG. 13 is a diagram showing feature identification information associated with each learning grid image. [Figure 16] 13 is a flowchart of a learning process. [Figure 17] 13 is a flowchart of learning data generation. [Figure 18] 13 illustrates a functional configuration of an object estimation device according to a second embodiment. [Figure 19] 13 is a flowchart of an additional learning process. [Figure 20] 13 illustrates a functional configuration of an object estimation device according to a third embodiment. [Figure 21] 13 is a flowchart of a learning image generation process. [Figure 22] FIG. 13 is a diagram illustrating an example of color division. [Diagram 23] FIG. 13 is a diagram showing a histogram of a learned image by color division. [Figure 24] 13 illustrates a functional configuration of an object estimation device according to a fourth embodiment. [Diagram 25] 13 is a flowchart of a learning image generation process. [Figure 26] 13 is a flowchart of a learning image generation process. [Figure 27] 13 is a flowchart of a learning image generation process. [Figure 28] FIG. 13 is a diagram showing a histogram of a learned image by color division. [Figure 29] 13 illustrates a functional configuration of an object estimation device according to a fifth embodiment. [Diagram 30] 13 is a flowchart of a learning image generation process. [Diagram 31] 13 is a flowchart of a learning image generation process. [Diagram 32] FIG. 13 is a diagram showing a histogram of object appearance positions in a learned image. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0066] 1. First embodiment 1.1 Functional configuration The functional configuration of an object estimation device according to an embodiment of the present invention is shown in Fig. 1. In this embodiment, an image to be estimated (original image to be processed) is an image captured while driving on a road.
[0067] The division means 2 receives the original image to be processed and divides it into grids of a predetermined size to obtain a processed grid image. The object estimation means 4 estimates what objects (in this embodiment, guardrails, roadsides, slopes, vegetation, etc.) are captured in the processed grid image. The object estimation means 4 can use, for example, an estimation model trained based on a grid image for learning and object identification information for identifying the object.
[0068] In this way, by identifying objects in grid units in images captured while traveling along a road, it is possible to determine which objects are located at which positions near the road.
[0069] As described above, the object estimation means 4 is trained based on training data. This device includes training data generation means 12 for generating the training data.
[0070] In this embodiment, images captured while driving on a road are used as learning source images for learning. A human being checks the learning source images on a screen and generates images (learning annotation source images) in which each object is given a meaning (for example, each object is colored in a different color).
[0071] The learning division means 8 receives the learning annotation original image and divides it into grids of a predetermined size to obtain a learning grid image.
[0072] The learning data acquisition means 10 generates object identification information based on annotations indicating objects included in the learning grid image. The learning data acquisition means 10 generates learning data by attaching the object identification information to the learning grid image.
[0073] The learning means 6 learns or additionally learns the object estimation means 4 based on the generated learning data.
[0074] In this embodiment, the learning data is generated by manually annotating the entire learning source image, dividing it into grids, and adding object identification information based on the annotations. Therefore, accurate learning data can be generated efficiently.
[0075] 1.2 Hardware Configuration The hardware configuration of the object estimation device is shown in Fig. 2. A memory 32, a display 34, an SSD 36, a DVD-ROM drive 38, a keyboard / mouse 40, and a communication circuit 42 are connected to a CPU 30.
[0076] The communication circuit 42 is for connecting to the Internet. An operating system 44 and an object estimation program 46 are recorded in the SSD 36. The object estimation program 46 performs its functions in cooperation with the operating system 44. These programs were recorded on a DVD-ROM 52 and installed in the SSD 36 via the DVD-ROM drive 38. In addition, a learning source image 48 and a processing source image 50 are also recorded in the SSD 36.
[0077] In addition to the CPU 30, a GPU may be used for image processing.
[0078] 1.3 Object Estimation Processing In this embodiment, object estimation is performed using a trained model in which a CNN-based deep learning model is trained using training data.
[0079] 3 shows a flowchart of the object estimation program 46. The CPU 30 acquires an original image recorded in the SSD 36 (step S1). In this embodiment, the original image is a video captured in the forward direction while traveling on a road. In this original image, the position of the vehicle (latitude and longitude information) at the time of capturing the image, acquired by the GPS receiver, is added to each still image constituting the video.
[0080] The original image to be processed may be a road driving video recorded on a portable recording medium that is transferred to the SSD 36. Alternatively, the original image may be transferred via the Internet.
[0081] FIG. 4 shows one still image constituting a moving image, which is the source image for processing. The CPU 30 divides this source image into a predetermined number of grids, and each division is set as a processing grid image (step S2). An example of a processing grid image is shown in FIG. 5. In this embodiment, the size of the processing grid image is determined in advance, and division is performed at grids of that size (division is performed starting from the upper left or center of the image). Therefore, the size of the grids becomes smaller at the edges of the source image for processing. In this case, the grids at the edges are not processed and are discarded (not adopted as processing grid images).
[0082] Next, the CPU 30 estimates which object is shown in the processing grid image by using the trained model (step S4). As described later, the trained model is trained to estimate which feature is shown in the processing grid image, such as a guard rail, road shoulder, slope, or planting. Therefore, it is possible to obtain which feature is shown in the processing grid image. In this embodiment, features such as those shown in FIG. 6 are estimated.
[0083] As shown in Fig. 7, in this embodiment, a trained model is provided for each feature to estimate whether it is the feature or not. That is, a roadway portion estimation means P1, a roadside shoulder estimation means P2, a parking lane estimation means P3,..., a fence estimation means Pn are provided. The roadway portion estimation means P1 outputs the probability that a roadway portion is included in what is shown in the processing grid image. Similarly, the roadside shoulder estimation means P2, a parking lane estimation means P3,..., a fence estimation means Pn output the probability that a roadside is included, the probability that a parking lane is included,..., the probability that a fence is included.
[0084] The integration means 20 receives the output of each estimation means P1 to Pn, selects from among them those that exceed a predetermined probability (e.g., 80% probability), and outputs the feature as an estimation result. Therefore, if there are multiple features that exceed the predetermined probability, multiple features will be output as estimation results.
[0085] In this embodiment, since estimation is performed on an image divided into grids, multiple features may appear in the grid. By performing estimation with the configuration shown in Fig. 7, even if multiple features are included, it is possible to accurately estimate them.
[0086] Next, the CPU 30 records the GPS position information and the estimated feature as object identification information in association with the processing grid image (step S5).
[0087] The CPU 30 performs the above process for all the divided processing grid images (steps S1 to S6).
[0088] Fig. 8 shows the image of Fig. 5 to which estimated features and position information have been added. Although only a portion of the image is shown in Fig. 8, features are identified and position information is added for each grid.
[0089] When the processing for one processing image is completed, the CPU 30 reads the next processing image from the SSD 36, divides it into processing grid images in the same manner as above, and identifies features for each processing grid image (steps S1 to S6).
[0090] This is performed for all the processing images to be processed (step S7). When the processing is completed for all, the CPU 30 records a set of grid images of the estimated features and the positions of the features (step S8). An example is shown in FIG.
[0091] In Fig. 9, the largest image of the estimated feature is selected. The processed images were taken while driving, so the same feature will be shown in multiple images.
[0092] In this embodiment, for the same feature, the grid image in which the feature appears is selected from the largest (oldest) processed image. If the same type of feature appears in the same processed grid image or in adjacent grid images 8 squares above and below in consecutive still images in the processed image, it is determined that they are the same feature.
[0093] The position information is recorded as follows: GPS position information (vehicle position) when the feature determined to be the same as the above appears on the screen, and GPS position information (vehicle position) when the feature disappears from the screen. An intermediate position between the two may also be recorded.
[0094] In this manner, the object can be identified and recorded in association with its position and image.
[0095] 1.4 Learning process for object estimation As described above, in this embodiment, a deep learning model using CNN is used for object estimation. Therefore, this model needs to be trained. In this embodiment, the training is performed using the hardware shown in FIG. 2, but the training may be performed using other computers.
[0096] A processing flowchart for generating learning data is shown in Fig. 10. In this embodiment, this processing is executed as a part of the mode (learning mode) of the object estimation program 46.
[0097] In the learning mode, the CPU 30 acquires the learning source image recorded in the SSD 36 (step S10). In this embodiment, the learning source image is a video captured in the forward direction while traveling on a road. In this learning source image, the vehicle position (latitude and longitude information) at the time of capturing the image, acquired by the GPS receiver, is added to each still image constituting the video.
[0098] The learning source images can be road driving videos recorded on a portable recording medium and transferred to the SSD 36. Alternatively, they may be transferred via the Internet.
[0099] The CPU 30 displays the acquired learning source image on the display 34. An example of the learning source image displayed on the display 34 is shown in FIG. 11. The operator looks at the displayed learning source image and operates the mouse 40 to annotate each of the features appearing in the photograph by encircling each type of feature (step S11). This annotation is specified as a circle or a polygon. FIG. 12 shows an image on which features such as roadways and signboards have been annotated. Annotations are made in distinguishable colors that are predetermined according to the type of feature, such as yellow for roadways (horizontal stripes in the figure), red for signboards (diagonal stripes in the figure), and blue for demarcation lines (stripes in the figure). The operator similarly annotates other features that appear.
[0100] Next, the CPU 30 divides the learning source image with annotations into grids, and each divided image is used as an annotation grid image (step S12). As in step S2, the size of the learning grid image is determined in advance (preferably the same size as in step S2), and the image is divided into grids of that size (division is performed starting from the upper left or center of the image). Therefore, the size of the grid becomes smaller at the edge of the learning source image. In this case, the grids at the edge are not processed and are discarded (not adopted as annotation grid images). In the example of FIG. 13, 66 grids are provided from grid 1,1 at the upper left to grid 6,11 at the lower right.
[0101] Next, the CPU 30 determines which features are included in each of the annotation grid images based on the annotations (step S13), and further records the types of features included in each grid (object identification information) in association with each grid (step S14).
[0102] Fig. 14 shows the types of features recorded in association with the grids. When no features are included, such as grid 1,1, the type of feature is not recorded, and when even a part of a grid includes an annotation of a feature, the type of that feature is recorded. When annotations of multiple features are included, such as grid 6,2 and grid 6,3, the types of multiple features are recorded.
[0103] Next, the CPU 30 records the features of the grids in association with each image obtained by removing the annotations from each annotation grid image (i.e., the learning source image divided into grids. This is called a learning grid image) and sets the data as learning data (step S15). An example of the recorded learning data is shown in Fig. 15. The grid position (grid ID), learning grid image, and type of feature are recorded.
[0104] When multiple learning data are generated based on one learning source image in the above manner, the next learning source image is obtained from the SSD 36, and learning data is generated and recorded in the same manner (steps S10 to S15). This is repeated for all target learning source images.
[0105] In this manner, a large amount of learning data is generated. The learning source images are captured while driving, and adjacent frames are almost identical images. Therefore, when generating the learning data, learning source images may be selected at intervals of a predetermined number of frames, rather than all frames. Also, images captured at locations that are a predetermined distance or more away may be selected based on imaging location information from a GPS.
[0106] Next, the CPU 30 generates learning data for each feature based on the learning data in Fig. 15 (step S17). In this embodiment, an estimation model is provided for each feature. Therefore, learning data is also generated for each estimation model.
[0107] In Fig. 15, the CPU 30 separates the data into data that includes the target feature and data that does not include the target feature (probability of being the target feature is low), and adds information to the former that the target feature is included and to the latter that the target feature is not included. In this way, learning data for the features is generated. This is performed for all features, and learning data for each feature is obtained.
[0108] Next, based on the learning data for each feature generated as described above, an estimation model of the corresponding feature is learned. A flowchart of the learning process is shown in Fig. 16. As described above, in this embodiment, an estimation model is provided for each feature. Therefore, learning is also performed for each estimation model.
[0109] For example, the case of learning an estimation model for estimating a roadway portion will be described below.
[0110] The CPU 30 reads out the learning data of the roadway section from the SSD 36 (step S20). The learning grid image of the learning data is provided to an estimation model (roadway section estimation model) that estimates the roadway section to execute estimation. As a result, the roadway section estimation model outputs the probability that the learning grid image includes the roadway section (step S21).
[0111] At this time, if the given learning grid image includes a roadway portion, the parameters of the estimation model are changed so as to increase the above probability, and if the given learning grid image does not include a roadway portion, the parameters of the estimation model are changed so as to decrease the above probability (step S22).
[0112] The above steps S20 to S22 are repeatedly executed to proceed with learning of the roadway section estimation model. When a predetermined learning end condition (a condition based on the number of data used in learning, the degree of agreement of the estimation probability, etc.) is satisfied, the CPU 30 ends the learning process (step S23). In this manner, a learned roadway section estimation model is generated.
[0113] During learning, it is preferable to provide training data including road sections and training data not including road sections in approximately the same number.
[0114] Similar training is performed on estimation models for other features to generate trained estimation models.
[0115] 1.5 Other (1) In the above embodiment, the size of the grid is determined in advance, the source image for processing and the source image for learning are divided, and grids at the end of the image that do not meet the predetermined size are discarded (deleted, filled with black, etc.). However, the length and width of the source image for processing and the source image for learning may be enlarged or reduced to an integer multiple of the length and width of the grid, and then the grid may be divided.
[0116] Alternatively, overlap with adjacent grids may be permitted at the ends and used as a grid.
[0117] (2) In the above embodiment, the objects to be estimated are features, but people, vehicles, etc. may also be used as the objects to be estimated.
[0118] (3) In the above embodiment, an estimation means is provided for estimating the probability that each object is included. This makes it possible to re-learn an estimation model for each object individually, which is efficient.
[0119] However, it is also possible to provide only one estimation means for estimating the probability that each object is included in the processing grid image, in which case it is preferable to set a low threshold value for the probability that the object is included, taking into consideration the case where the processing grid image includes multiple objects.
[0120] In addition, for similar objects (such as a lane marking and a crosswalk), multiple objects may be estimated using a single estimation model. This is because the estimation accuracy is improved by treating similar objects as a single model.
[0121] Also, an estimation means may be provided that estimates the probability that only the object is included for each object, and an estimation means that estimates the probability that a possible combination of objects (for example, a roadway and a lane marking) is included. This requires a large number of estimation means, but makes it possible to more accurately estimate the objects included in the processing grid image.
[0122] (4) In the above embodiment, the device is constructed as a device for obtaining the position of an object and an image thereof. However, the device may be used to identify the type of feature in the 3D point cloud data obtained by traveling and to associate the image. In other words, the device may be generally used when associating the type of object, such as a feature, with an image.
[0123] (5) In the above embodiment, as shown in FIG. 11, an operator adds annotations to generate learning data. However, as shown in the flowchart of FIG. 17, learning grid images may be estimated using a trained estimation model, and those with a high probability of estimation of the object (for example, 95% or more or 5% or less) may be used as learning data. That is, learning data may be generated by adding object identification information indicating that the learning grid image is an object to a learning grid image that is certain to include an object, and adding object identification information indicating that the learning grid image is not an object (background) to a learning grid image that is certain not to include an object. This can be used when generating additional learning data when a trained model has already been formed.
[0124] In step S30, the CPU 30 acquires a learning source image from the SSD 36. Next, the CPU 30 divides the learning source image into learning grid images (step S31). The learning grid images thus obtained are applied to a trained model for each object to obtain an estimated probability (step S34).
[0125] CPU 30 judges whether this estimation probability is equal to or higher than a predetermined high probability (step S35). If it is equal to or higher than the predetermined high probability (e.g., 100%), the learning grid image, the estimated object, and the grid position are recorded as learning data (step S36). Similarly, if it is equal to or lower than a predetermined low probability (e.g., 0%), the learning grid image, the estimated object (background), and the grid image are recorded as learning data. If the predetermined high probability or low probability is not reached, the learning grid image is not adopted as learning data.
[0126] The above process is performed for the trained models of all objects (steps S33 to S37).
[0127] Furthermore, this process is performed for all of the learning grid images obtained by division (steps S32 to S37).
[0128] The CPU 30 judges whether or not the processing has been completed for all of the target learning source images (step S38), and if there are any unprocessed learning source images, it repeatedly executes step S30 and the following steps. When the processing has been completed for all of the learning source images, the processing ends.
[0129] (6) In the above embodiment, the object estimation device is configured to have a learning function and an additional learning function. However, the object estimation device may be configured as a learning device or an additional learning device that does not have the object estimation function.
[0130] (7) The above-described embodiments and modifications may be combined with other embodiments and modifications as long as this does not go against the essence of the invention.
[0131] 2. Second embodiment 2.1 Functional Configuration 18 shows the functional configuration of an object estimation device according to the second embodiment. In this embodiment, an image to be estimated (original image to be processed) is an image captured while driving on a road.
[0132] The division means 2 receives the original image to be processed and divides it into grids of a predetermined size to obtain a processed grid image. The object estimation means 4 estimates what objects (in this embodiment, guardrails, roadsides, slopes, vegetation, etc.) are captured in the processed grid image. The object estimation means 4 can use, for example, an estimation model trained based on a grid image for learning and object identification information for identifying the object.
[0133] In this way, by identifying objects in grid units in images captured while traveling along a road, it is possible to determine which objects are located at which positions near the road.
[0134] The above points are the same as those in the first embodiment. Also, the object estimation means 4 is formed by learning an estimation model such as deep learning.
[0135] In this embodiment, there is provided an update means 20 for appropriately additionally learning the trained object estimation means 4. The additional learning is learning performed using additional learning data in order to further improve the estimation accuracy of the trained estimation model.
[0136] The additional learning means 22 additionally learns the object estimation means 4 based on the additional learning data, and records the object estimation means 4 after the additional learning as a provisional estimation means .
[0137] The estimation accuracy determination means 24 provides the evaluation data to the provisional estimation means 28 and calculates how accurately the estimation can be performed (estimation accuracy). The same evaluation data is provided to the object estimation means 4 before additional learning and calculates how accurately the estimation can be performed (estimation accuracy). The evaluation data is unbiased data such as image features and object appearance positions so that the data is suitable for calculating the estimation accuracy of the estimation means.
[0138] If the estimation accuracy of the provisional estimation means 28 is higher than that of the object estimation means 4 before learning, the replacement means 26 replaces the object estimation means 4 with the provisional estimation means 28. Therefore, the object estimation means trained using the additional training data will be used thereafter. On the other hand, if the estimation accuracy of the provisional estimation means 28 is lower than that of the object estimation means 4 before learning, the object estimation means 4 before learning will be used as is.
[0139] By doing so, the object estimation means 4 can perform additional learning more preferably.
[0140] 2.2 Hardware Configuration The hardware configuration is the same as that shown in FIG. 2 in the first embodiment.
[0141] 2.3 Additional learning process FIG. 19 shows a flowchart of the additional learning process, which is one function of the object estimation program.
[0142] The CPU 40 provides evaluation data to the estimation model before additional learning, causes it to perform inference, and evaluates the accuracy (precision) of the inference (step S40). In this embodiment, since an estimation model is provided for each object, the accuracy is evaluated for each object. Also, evaluation data is provided for each object.
[0143] Here, the evaluation data is composed of evaluation grid images that include the object to be estimated (recorded in association with the inclusion of the object) and evaluation grid images that do not include the object (recorded in association with the inclusion of the object).
[0144] The evaluation data is data that is suitable for calculating the estimation accuracy of the estimation means and is free of bias in image features, object appearance positions, etc. Such data may be generated by manual selection by an operator, but may also be generated using the same method as that for preventing bias in the additional learning data shown in the third, fourth, and fifth embodiments.
[0145] The CPU 40 performs the evaluation as follows. An evaluation grid image is provided to the estimation model to obtain an estimation result. When an evaluation grid image including an object is provided, if the probability that it is the object exceeds a predetermined value (e.g., 80%), an evaluation value of "1" is provided; if not, an evaluation value of "0" is provided. When an evaluation grid image not including an object is provided, if the probability that it is the object is below a predetermined value (e.g., 20%), an evaluation value of "1" is provided; if not, an evaluation value of "0" is provided.
[0146] The above is performed for all evaluation grid images, and the evaluation values are summed up to become the evaluation value of the estimation model before the additional learning.
[0147] Then, before the additional learning is performed, the parameters of the estimation model are recorded (step S41). This is to enable reproduction of the estimation model before the additional learning. Note that the entire estimation model before the additional learning may be recorded.
[0148] Next, the estimation model is additionally trained using additional training data (steps S42 to S45). Here, the additional training data is similar to the training data shown in Fig. 15. The additional training process is similar to the training process shown in Fig. 16.
[0149] When the additional learning is completed, the CPU 30 evaluates the estimation model after the additional learning using the same evaluation data as before (step S46). The evaluation method is the same as that of step S40.
[0150] The CPU 30 compares the evaluation value of the estimation model before the additional learning with the evaluation value of the estimation model after the additional learning (step S47). If the evaluation value of the estimation model after the additional learning is higher, the estimation model after the additional learning will be used from now on (step S48). In other words, the original estimation model is replaced with the estimation model after the additional learning.
[0151] On the other hand, if the estimation model before the additional learning has a higher evaluation value, the estimation model before the additional learning is restored using the parameters recorded in step S41 and used from now on. In other words, the estimation model before the additional learning continues to be used.
[0152] In this embodiment, an estimation model is constructed for each object to be estimated, and the above process is therefore performed for each estimation model of the object.
[0153] 2.5 Other (1) In the above embodiment, it is determined whether to use an estimation model that has undergone additional learning for each object. However, if the number of estimation models whose estimation accuracy has been improved by performing additional learning is equal to or greater than a predetermined number (e.g., more than half), all estimation models, including estimation models whose accuracy has not been improved, may be replaced with the learned ones.
[0154] (2) In the above embodiment, an estimation model is provided for each object. However, the present invention can be applied to a case where a single estimation model is used to distinguish and judge a plurality of objects or to judge a combination of a plurality of objects.
[0155] (3) In the above embodiment, the final evaluation value is obtained by adding up the evaluation values "1" or "0." However, for evaluation grid images that include an object, the estimated probability may be added, and for evaluation grid images that do not include an object, the estimated probability may be subtracted, and the total value may be used as the evaluation value.
[0156] (4) In the above embodiment, the object estimation device has been described as having a function of performing additional learning. However, the object estimation device may be configured as an additional learning device without the object estimation function.
[0157] (5) The above-described embodiments and modifications may be combined with other embodiments and modifications as long as this does not go against the essence of the invention.
[0158] 3. Third embodiment 3.1 Functional configuration 20 shows a functional configuration of a training data generation device 60 according to the third embodiment. In the figure, in addition to the training data generation device 60, an object estimation device having a learning means 6 and an object estimation means 4 is also shown.
[0159] In addition, the object estimation device may be configured to include the function of the learning data generation device 60.
[0160] In this embodiment, the training data generating device 60 includes a trained image feature determining means 62 , a candidate image feature determining means 64 , a feature distribution calculating means 66 , and a training image selecting means 68 .
[0161] The learned image feature determination means 62 acquires a learned image. Here, the learned image refers to an image that has already been used to train the object estimation means 4 by the learning means 6. The learned image feature determination means 62 classifies each pixel included in the learned image based on features including its color or density. For example, the pixel color is classified into a predetermined number of categories, and it is determined to which category each pixel belongs. Then, it is determined to which category the majority of pixels belong for the entire learned image, and the category with the greatest number of pixels is set as the representative feature of the learned image.
[0162] The learned image feature determining means 62 calculates representative features for all the learning pixels.
[0163] The feature distribution calculation means 66 calculates how the representative features are distributed for all learned images.
[0164] The candidate image feature determination means 64 acquires a candidate image. Here, the candidate image refers to an image that is a candidate for the learning image. The candidate image feature determination means 64 classifies each image included in the candidate image based on features including its color or density, and calculates a representative feature. The method is the same as the calculation of the representative feature of the pre-learned image.
[0165] The learning image selection means 68 determines whether to use the candidate image as a learning image based on the distribution of the representative features of the pre-learned images and the representative features of the candidate image. In this embodiment, it is determined whether the representative features of the pre-learned images are normalized when the representative features of the candidate image are added to the distribution of the representative features of the pre-learned images. If it seems to be normalized, the candidate image is adopted as a learning image, and if it does not seem to be normalized, the candidate image is not adopted as a learning image.
[0166] According to this embodiment, it is possible to obtain a learning image that prevents the bias of the learning data and appropriately learns the object estimation means.
[0167] 3.2 Hardware Configuration The hardware configuration is the same as that in FIG. 2 of the first embodiment.
[0168] 3.3 Learning Data Generation Process In this embodiment, for example, an apparatus for generating learning data for performing additional learning on a learned model that estimates whether an object is shown in a processed image will be described.
[0169] The trained model can be constructed by training a deep learning model such as CNN using training images that include the object (with information indicating that it is an object) and training images that do not include the object (with information indicating that it is not an object). Note that it is preferable that the number of images that include the object (referred to as object images) and images that do not include the object (referred to as background images) used for training are approximately the same.
[0170] In this embodiment, unlike the first and second embodiments, an example is described in which a processed image is not gridded but is given to a trained model to estimate the probability that the image is the target object. As described later, the present invention can also be applied to the first and second embodiments in which the processed image is gridded and estimation processing is performed.
[0171] Additional learning is performed to improve the estimation accuracy of a trained model. However, if the training data for additional learning is not appropriate, the estimation accuracy of the trained model may not improve as expected.
[0172] The following describes the process of generating learning data that is expected to be effective for additional learning.
[0173] A flowchart of the learning data generation process is shown in Fig. 21. In generating learning data, it is necessary to generate both a learning image that includes an object (object learning image) and a learning image that does not include an object (background learning image). The flowchart in Fig. 21 shows the process for generating an object learning image, but the process for generating a background learning image is similar.
[0174] The CPU 30 acquires learned images including an object (learned image of object) from among learned images of learned data (learned data that has already been used for learning) recorded in the SSD 36 (step S51).
[0175] Next, the CPU 30 classifies the color of each pixel of the object learned image into a predetermined type (step S52). In this embodiment, as shown in Fig. 22, the RGB values are divided into four groups, 0 to 63, 64 to 127, 128 to 191, and 192 to 255, and the combinations of these groups are divided into 64 sections. Here, RGB is divided into four groups, but it may be divided into three or less groups, or five or more groups. Also, instead of dividing RGB into the same number, it may be divided into different numbers.
[0176] The CPU 30 determines which color has the most pixels in the learned object image, and sets the most abundant color as the representative color of the learned object image (step S53).
[0177] The above is performed for all of the object learning images to obtain the representative colors of each (steps S50 to S54).
[0178] Next, the CPU 30 generates a histogram of representative colors of the learned object image, as shown in FIG. 23A (step S55).
[0179] Next, the CPU 30 acquires candidate images including the object (object candidate images) from the SSD 36 (step S57). In this embodiment, the candidate images are provided to the object estimation means 4 for estimation, and an image with a predetermined high probability (e.g., 95% or higher) of the object is selected as the object candidate image. Of course, the object candidate images may be selected by a human being.
[0180] Next, in the same manner as in steps S52 and S53, the representative color of the object candidate image is determined (steps S58 and S59).
[0181] Next, CPU 30 judges whether the representative color of the object candidate image is a representative color that has reached the upper limit value of the histogram (step S60). Here, the upper limit value is a value obtained by adding (subtracting) a predetermined percentage (e.g., 20%) to the average value of the histogram, as shown in FIG. 23B.
[0182] If the representative color of the object candidate image does not reach the upper limit value in the histogram, it is adopted as a learning image since it does not violate the histogram equalization (step S61). That is, the object candidate image is recorded as learning data with a mark indicating that it is an object. In this case, the histogram is updated with the adopted new learning image.
[0183] If the representative color of the target candidate image reaches the upper limit value in the histogram, it is not adopted as a learning image because it goes against the equalization of the histogram.
[0184] 23B, the color divisions 3, 4, and 64 have reached the upper limit, and the color divisions 1, 2, 5, and 63 have not reached the upper limit. Therefore, if the representative color of the target candidate image is in the color divisions 3, 4, or 64, it will not be used as a learning image, and if it is in the color divisions 1, 2, 5, or 63, it will be used as a learning image.
[0185] The CPU 30 retrieves the next object candidate image from the SSD 36 and performs the same process as above to determine whether or not to adopt the object candidate image as a learning image. This is performed for all object candidate images (steps S56 to S62). In this manner, learning data including the object can be obtained.
[0186] The CPU 30 also performs the same process as in FIG. 21 on candidate images that do not include a target object (background candidate images) to obtain learning data.
[0187] In addition, it is preferable to use approximately the same amount of object learning data and background learning data in additional learning. However, it is possible that a large amount of one of the data is obtained. In that case, the larger amount of additional learning data is selected so that it matches the smaller amount of additional learning data. In this case, it is preferable to select the larger amount of additional learning data so that the representative color histogram of the additional learning data is close to flat.
[0188] Moreover, the unselected larger amount of additional learning data can be used in the next or subsequent additional learning. In particular, in the next or subsequent additional learning, the larger amount of additional learning data can be used as a selection target for making the representative color histogram of the larger amount of additional learning data closer to flat.
[0189] 3.4 Other (1) In the above embodiment, RGB data is used as the color features. However, saturation or chromaticity may also be used. Also, density features may be used instead of color features. Furthermore, features based on the direction of edges appearing in an image, geometric features of edges, and spatial and temporal waveform features of pixels (which can be calculated by Fourier transform, etc.) may also be used.
[0190] (2) In the above embodiment, a case has been described in which learning data is generated for training an object estimation means that estimates whether an object is included in a processing image. However, the present invention can also be applied to a case in which learning data (learning grid image) is generated for training an object estimation means that performs estimation by dividing an original processing image shown in the first embodiment into grids. In this case, the candidate image is a candidate grid image.
[0191] (3) In the above embodiment, the most common color segment is used as the representative color of the candidate image. However, the average color of all pixels may be used as the representative color.
[0192] (4) The above-described embodiments and modifications may be combined with other embodiments and modifications as long as this does not go against the essence of the invention.
[0193] 4. Fourth embodiment 4.1 Functional configuration 24 shows a functional configuration of a training data generation device according to the fourth embodiment. In the figure, in addition to a training data generation device 70, an object estimation device having a division means 2, an object estimation means 4, and a learning means 6 is also shown.
[0194] It should be noted that the object estimation device may be constructed to include the function of the learning data generation device 70.
[0195] In this embodiment, the training data generation device 70 includes a trained original image feature determination means 72, a candidate image feature determination means 74, a trained original image feature distribution calculation means 76, a first training image selection means 78, a candidate image attention grid feature determination means 82, a trained attention grid feature distribution calculation means 84, and a second training image selection means 86.
[0196] The trained original image feature determining means 72 acquires a training original image that has already been used for training. Here, a trained original image refers to an image in which one or more grid images have been used as training data.
[0197] The learned original image feature determination means 72 classifies each pixel contained in each learned grid image (including those not actually used in learning) that constitutes the learned image based on features including its color or density. For example, the pixel color is classified into a predetermined number of categories, and it is determined which category each pixel belongs to. Then, it is determined which category contains the most pixels of the learned grid image, and the category with the most pixels is set as the representative feature of the learned grid image. Since the learned original image includes multiple learned grid images, a representative feature is obtained for each learned grid.
[0198] The learned original image feature distribution calculation means 76 calculates the distribution of representative features for each grid position of a plurality of learned original images.
[0199] The candidate original image feature determination means 74 calculates the representative feature of each candidate grid image constituting the candidate original image. The calculation method is the same as that of the learned original image feature determination means 72. Since the candidate original image includes multiple candidate grid images, a representative feature is obtained for each candidate grid.
[0200] The first training image selection means 78 judges whether adding the candidate original image to the training images contradicts the equalization of the distribution of the representative features of the already-trained original images. If it contradicts, the candidate original image is not selected as a training image. If it does not contradict, the candidate original image is selected as a training image.
[0201] By making such a selection, it is possible to obtain additional training data that enables training to be performed without bias in the color features of the image.
[0202] Although it is possible to use images selected in this way as learning images, in this embodiment, further narrowing down the selection is performed.
[0203] The learned attention grid feature distribution calculation means 84 calculates the representative color distribution only for grids (learned attention grid images) that include an object among all learned grid images of all learned original images. Here, the distribution is calculated as a whole without considering the position of the grid. Similarly, the representative color distribution is calculated only for grids that do not include an object among all learned attention grids of all learned original images.
[0204] The candidate original image attention grid feature determination means 82 calculates representative colors only for grids that include objects among all the candidate grid images of the candidate original image. Similarly, the candidate original image attention grid calculates a distribution of representative colors only for grids that do not include objects among all the candidate original image attention grids.
[0205] The second learning image selection means 86 judges whether to select the candidate original image as a learning image based on the viewpoint of whether the representative color of the candidate grid including the object violates the equalization of the distribution of the representative color of the learned grid of interest including the object, and based on the viewpoint of whether the representative color of the candidate grid not including the object violates the equalization of the distribution of the representative color of the learned grid of interest not including the object, the distribution of the representative color of the candidate grid not including the object. Learning data is generated by associating the learning image with the object identification information.
[0206] 4.2 Hardware Configuration The hardware configuration is the same as that shown in FIG. 2 in the first embodiment.
[0207] 4.3 Training data generation process 25 to 27 show flowcharts of the learning data generation process. In this embodiment, the learning data generation process is implemented as one function of the object estimation program.
[0208] In this embodiment, a trained model is formed for each object in the object estimation means 4. Therefore, the processes shown in Figs. 25 to 27 are executed for each object.
[0209] The CPU 30 acquires a learned original image from the SSD 36 (step S71). Here, the learned original image refers to an original image in which any of its element grid images has been used for learning. A grid image constituting the learned original image is called a learned grid image. The expression "learned grid image" means a grid image constituting the learned original image. Therefore, the learned original image includes not only a grid image that has actually been used for learning, but also a grid image that has not been used for learning.
[0210] Next, the CPU 30 acquires the trained grid image constituting the trained original image, classifies the color of each pixel, and determines the representative color (steps S73 and S74). This is performed for all trained grid images included in the trained original image, and obtains the representative color for each (steps S72 to S75).
[0211] The above process is executed for all learned original images (steps S70 to S75). Therefore, the representative color for each learned grid image can be obtained for all learned original images.
[0212] As shown in FIG. 28, the CPU 30 calculates a histogram of representative colors of the pre - learned grid images for each grid (step S76).
[0213] Next, the CPU 30 acquires a candidate source image from the SSD 36 (step S78). In this embodiment, an image captured while the vehicle is running is used as the candidate source image. The CPU 30 divides this candidate source image into grids (step S79). Further, the CPU 30 estimates the probability that an object is included in each candidate grid image using the learned object estimation means 4 in the first embodiment, etc. (step S80).
[0214] If there is at least one grid that can be regarded as definitely containing an object (a grid with a probability of containing an object equal to or higher than a predetermined ratio (for example, 95% or more)), the candidate source image is maintained (step S83). If there is not even one grid that can be regarded as definitely containing an object, the candidate source image is not used as an object for subsequent processing for that object (step S83). Note that candidate source images thus excluded from the processing may also be used as learning data for other objects.
[0215] The above processing is performed for all candidate source images (steps S77 to S84), and a candidate source image in which an object is included in at least one of its grids can be selected.
[0216] Next, the CPU 30 classifies the colors of each pixel of the candidate grid images constituting the selected candidate source image and determines the representative color (steps S86, S87). Subsequently, the CPU 30 determines whether the representative color of the candidate grid image is a representative color that reaches the upper limit value of the histogram calculated in step S76 (step S88). Here, as shown in FIG. 23B, the upper limit value is a value obtained by adding (subtracting) a predetermined ratio (such as 20%) to the average value of the histogram.
[0217] If the candidate grid image does not have a representative color that has not reached the upper limit, the adoption evaluation point of the candidate original image (initial value is 0) is incremented by +1 (step S89). If the candidate grid image has a representative color that has not reached the upper limit, the adoption evaluation point of the candidate original image is left unchanged.
[0218] The CPU 30 repeats the above process for all the candidate grid images constituting the candidate original image (steps S85 to S90), thereby calculating the adoption evaluation points for the candidate original images.
[0219] When the CPU 30 calculates the adoption evaluation score for the candidate original image, it determines whether the adoption evaluation score is equal to or greater than a predetermined score (step S91). If the adoption evaluation score is equal to or greater than the predetermined score, the candidate original image is adopted as a learning image (step S92). If the adoption evaluation score is less than the predetermined score, the candidate original image is not adopted as a learning image.
[0220] The CPU 30 repeats the above process to obtain candidate original images to be adopted as learning images (adopted candidate original images) (steps S84 to S93).
[0221] By obtaining training images in the above manner, it is possible to add candidate original images as training images so that they are evenly distributed overall, taking into account the already-trained original images. This allows for unbiased training, including not only the target object but also the background.
[0222] Although this original image candidate for adoption can be used for learning, in this embodiment, as described below, the original image candidates for adoption as learning data are further narrowed down.
[0223] The CPU 30 extracts only candidate grid images including the target object from all the learned original images (step S94).The CPU 30 calculates histograms of the representative colors of all the extracted learned grid images (step S95).
[0224] Next, the CPU 30 selects only the candidate grid images that can be reliably regarded as including the object by the object estimation means 4 from among the candidate grid images that constitute one adopted candidate original image (step S97).
[0225] Next, the representative color of the candidate grid image that can be regarded as definitely including the object is calculated, and it is judged whether the representative color reaches the upper limit of the histogram in step S95 (see FIG. 23B) (step S98). This is performed for all candidate grid images that can be regarded as definitely including the object of the original candidate image.
[0226] The CPU 30 judges whether there is even one candidate grid image having a representative color that does not reach the upper limit (step S99). If there is, the adopted candidate image is adopted as a learning source image (step S100). If there is not, the adopted candidate image is not adopted as a learning source image.
[0227] The CPU 30 executes the above process for all the adoption candidate images, thereby obtaining a group of adoption candidate images to be used as learning source images.
[0228] The learning source image thus obtained is used for learning as follows. First, the CPU 30 divides this learning source image into grids. Furthermore, the object estimation means 4 extracts learning grid images that are deemed to definitely include the object (e.g., estimated probability 95% or more) and learning grid images that are deemed to definitely not include the object (e.g., estimated probability 5% or less). The former are given object identification information indicating that they are objects, and the latter are given object identification information indicating that they are not objects, and these are used as learning data.
[0229] As in the third embodiment, it is preferable to make the number of target images and the number of non-target images approximately equal in the creation of additional training data. In addition, it is also preferable to select the images so that the histogram is as flat as possible.
[0230] 4.4 Other (1) In the above embodiment, the candidate original images that have been narrowed down are further narrowed down by the process shown in Fig. 27. However, the candidate original images selected by the process shown in Fig. 27 may be used as learning original images.
[0231] (2) In the above embodiment, the object estimation means 4 determines whether or not an object is included (step S80, etc.). However, the determination may be made by a human being.
[0232] (3) In the above embodiment, if the estimated probability is 95% or more, it is determined that the target object is definitely included. However, it may be 80% or more or 100%. It is preferable to set the predetermined value to 90% or more.
[0233] (4) In the above embodiment, in step S81, a candidate original image including a grid that is certain to include an object is selected. This is because most candidate original images include grids that are certain to not include an object.
[0234] However, it is also possible to select candidate original images that include not only grids that are certain to include the object, but also grids that are certain to not include the object.
[0235] (5) In the above embodiment, in the process of Fig. 27, a histogram of representative colors of grids including an object in the learned original image is calculated, and whether or not to adopt is determined based on the representative colors of grids in the candidate original image that are sure to include the object. However, a histogram of representative colors of grids not including an object in the learned image may be calculated, and whether or not to adopt may be determined based on the representative colors of grids in the candidate original image that are sure to not include the object. Alternatively, the determination may be made by taking both of these into consideration.
[0236] (6) In the above embodiment, in step S99, if there is even one grid image in the candidate original image that does not reach the upper limit, it is adopted as a learning image. However, it may be determined whether to adopt the image as a learning image based on the ratio of grid images in the candidate original image that do not reach the upper limit (e.g., adopt a predetermined ratio or more) out of all grid images including the target object.
[0237] (7) The above-described embodiments and modifications may be combined with other embodiments and modifications as long as this does not go against the essence of the invention.
[0238] 5. Fifth embodiment 5.1 Functional Configuration 29 shows the functional configuration of a training data generation device according to the fifth embodiment. In the figure, in addition to a training data generation device 90, an object estimation device having a division means 2, an object estimation means 4, and a learning means 6 is also shown.
[0239] It should be noted that the object estimation device may be constructed to include the function of the learning data generation device 90.
[0240] In this embodiment, the training data generating device 90 includes a trained original image frequency calculating means 92 , a candidate image position specifying means 94 , a trained image frequency distribution calculating means 96 , and a training image selecting means 98 .
[0241] The learned original image frequency calculation means 92 acquires a learning original image that has already been used for learning. Here, a learned original image refers to one or more of its grid images that have been used as learning data.
[0242] The learned original image frequency calculation means 92 calculates the frequency with which an object appears at the position of each learned grid image for all learned images.
[0243] The learned image frequency distribution calculation means 96 calculates the distribution of the appearance frequency of the object at each calculated grid position.
[0244] The candidate original image position specifying means 94 specifies the position of a candidate grid image that is sure to include a target object among the candidate grid images that constitute the candidate original image.
[0245] The training image selection means 98 judges whether adding the candidate original image to the training image contradicts the equalization of the appearance frequency distribution of the object in the grid of the trained original image. If it contradicts, the candidate original image is not selected as a training image. If it does not contradict, the candidate original image is selected as a training image.
[0246] By making such a selection, it is possible to obtain additional learning data that enables learning to be performed without bias in the frequency of occurrence of the target object positions.
[0247] 5.2 Hardware Configuration The hardware configuration is the same as that shown in FIG. 2 in the first embodiment.
[0248] 5.3 Training data generation process 30 and 31 show flowcharts of the learning data generation process. In this embodiment, the learning data generation process is implemented as one function of the object estimation program.
[0249] In this embodiment, the object estimation means 4 forms a trained model for each object. Therefore, the processes shown in Fig. 30 and Fig. 31 are executed for each object.
[0250] The CPU 30 acquires a learned original image from the SSD 36 (step S71). Here, the learned original image refers to an original image in which any of its element grid images has been used for learning. A grid image constituting the learned original image is called a learned grid image. The expression "learned grid image" means a grid image constituting the learned original image. Therefore, the learned original image includes not only a grid image that has actually been used for learning, but also a grid image that has not been used for learning.
[0251] Next, the CPU 30 acquires the learned grid images constituting the learned original image, and records the positions of the learned grid images including the object in the learned original image (step S152). Here, the positions in the learned original image are positions 1,1, 1,2, 1,3,..., 6,11 in the example of FIG.
[0252] The CPU 30 executes the above process for all the learned original images (steps S150 to S153).
[0253] Next, the CPU 30 calculates the number of objects appearing for each grid position and generates a histogram (step S154). An example of the generated histogram is shown in FIG.
[0254] Next, the CPU 30 reads out the candidate original image from the SSD 36 (step S156). The CPU 30 divides the candidate original image into candidate grid images (step S157). Furthermore, the CPU 30 uses the object estimation means 4 to estimate the probability that an object is included in each of these candidate grid images (step S158).
[0255] If any of the candidate grid images constituting the candidate original image is found to include a target object, the candidate original image is adopted and is subject to the following learning data generation process (step S161). If there is no candidate grid image that is found to include a target object, the candidate original image is not adopted and is not subject to the following learning data generation process (step S160).
[0256] The determination as to whether or not it is certain that the target object is included can be made by the same method as that described in the fourth embodiment.
[0257] Next, for one selected candidate original image, the CPU 30 specifies a position of a candidate grid image where the object is surely present (step S164). The object may be present at a plurality of positions.
[0258] The CPU 30 judges whether the grid position including the object reaches the upper limit (see FIG. 23B) in the histogram of FIG. 32 (step S165). If there is even one candidate grid image that does not reach the upper limit (step S166), the CPU 30 adopts the candidate original image as a learning original image (step S167).
[0259] If there is no candidate grid image for which the upper limit has not been reached (step S166), the CPU 30 does not adopt the candidate original image as a learning original image.
[0260] The CPU 30 repeats the above process for all selected candidate original images (steps S163 to S168).
[0261] In this manner, a group of candidate original images to be used as learning data can be obtained.
[0262] The learning source image thus obtained is used for learning as follows. First, the CPU 30 divides this learning source image into grids. Furthermore, the object estimation means 4 extracts learning grid images that are deemed to definitely include the object (e.g., estimated probability 95% or more) and learning grid images that are deemed to definitely not include the object (e.g., estimated probability 5% or less). The former are given object identification information indicating that they are objects, and the latter are given object identification information indicating that they are not objects, and these are used as learning data.
[0263] As in the third embodiment, it is preferable to make the number of target images and the number of non-target images approximately equal in the creation of additional training data. In addition, it is also preferable to select the number of target images so that the histogram is as flat as possible.
[0264] 5.4 Other (1) The above-described embodiments and modifications may be combined with other embodiments and modifications as long as such combination does not violate the essence of the embodiments and modifications.
Claims
1. a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning source image into grids and object identification information shown in the learning grid image are associated with each other; An image-based object estimation apparatus comprising: A learning data generating means for generating the learning data is further provided, The learning data generating means A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain an annotation grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, to obtain learning data; An image-based object estimation apparatus comprising: Further comprising additional learning data generating means, The additional learning data generating means includes: a learning image feature determining means for classifying each pixel included in each learned image that has already been used in learning based on a feature including a color or density thereof, and determining a representative feature of the learned image based on the number of pixels having each feature; A feature distribution calculation means for calculating a distribution of representative features in the plurality of learned images; a candidate image feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a learning image selection means for selecting the candidate image as a learning image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the learned images; An object estimation device comprising:
2. a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning source image into grids and object identification information shown in the learning grid image are associated with each other; An image-based object estimation apparatus comprising: A learning data generating means for generating the learning data is further provided, The learning data generating means A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain an annotation grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; An image-based object estimation apparatus comprising: Further comprising additional learning data generating means, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation device comprising:
3. a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning source image into grids and object identification information shown in the learning grid image are associated with each other; An image-based object estimation apparatus comprising: A learning data generating means for generating the learning data is further provided, The learning data generating means A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain an annotation grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, to obtain learning data; An image-based object estimation apparatus comprising: Further comprising additional learning data generating means, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation device comprising:
4. A learning data generation device for generating learning data for training an object estimation means in an image-based object estimation device including: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image; and an object estimation means for estimating what an object is captured in the processing grid image obtained by dividing the original image into grids, the object estimation means being trained based on learning data in which object identification information captured in the learning grid image is associated with the learning grid image, the object estimation means comprising: A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain a learning grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; A training data generating device comprising: Further comprising additional learning data generating means, The additional learning data generating means includes: a learning image feature determining means for classifying each pixel included in each learned image that has already been used in learning based on a feature including a color or density thereof, and determining a representative feature of the learned image based on the number of pixels having each feature; A feature distribution calculation means for calculating a distribution of representative features in the plurality of learned images; a candidate image feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a learning image selection means for selecting the candidate image as a learning image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the learned images; A training data generating device comprising:
5. A learning data generation device for generating learning data for training an object estimation means in an image-based object estimation device including: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image; and an object estimation means for estimating what an object is captured in the processing grid image obtained by dividing the original image into grids, the object estimation means being trained based on learning data in which object identification information captured in the learning grid image is associated with the learning grid image, the object estimation means comprising: A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain a learning grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; A training data generating device comprising: Further comprising additional learning data generating means, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning data generating device comprising:
6. A learning data generation device for generating learning data for training an object estimation means in an image-based object estimation device including: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image; and an object estimation means for estimating what an object is captured in the processing grid image obtained by dividing the original image into grids, the object estimation means being trained based on learning data in which object identification information captured in the learning grid image is associated with the learning grid image, the object estimation means comprising: A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain a learning grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; A training data generating device comprising: Further comprising additional learning data generating means, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a trained image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated trained original images; a candidate original image position specifying means for specifying a grid position where an object appears among candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning data generating device comprising:
7. In the device according to any one of claims 1 to 6, The apparatus further comprises a learning data generating means for generating, as learning data, a grid image in which the object estimation means estimates an object with a probability higher than a predetermined value, and object identification information which is the object estimation result.
8. In the device according to any one of claims 1 to 7, and an update means for performing additional learning on the trained object estimation means based on additional learning data and updating the object estimation means in accordance with a result of the additional learning, The update means sets the object estimation means after learning based on additional learning data as a provisional object estimation means, provides evaluation data to the current object estimation means before additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the provisional object estimation means, and if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is.
9. An object estimation program for implementing an object estimation device by a computer, the program comprising: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning source image into grids and object identification information shown in the learning grid image are associated with each other; a learning data generating unit that generates the learning data, The learning data generating means A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain an annotation grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; In the image-based object estimation program, the computer further comprises: It functions as an additional learning data generating means, The additional learning data generating means includes: a learning image feature determining means for classifying each pixel included in each learned image that has already been used in learning based on a feature including a color or density thereof, and determining a representative feature of the learned image based on the number of pixels having each feature; A feature distribution calculation means for calculating a distribution of representative features in the plurality of learned images; a candidate image feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a learning image selection means for selecting the candidate image as a learning image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the learned images; An object estimation program comprising:
10. An object estimation program for implementing an object estimation device by a computer, the program comprising: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning source image into grids and object identification information shown in the learning grid image are associated with each other; a learning data generating unit that generates the learning data, The learning data generating means A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain an annotation grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; In the image-based object estimation program, the computer further comprises: It functions as an additional learning data generating means, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation program comprising:
11. An object estimation program for implementing an object estimation device by a computer, the program comprising: a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning source image into grids and what object is identified in the learning grid image, the object estimation means being trained based on learning data in which a learning grid image is associated with information on identifying an object shown in the learning grid image, the object estimation means being trained based on learning data in which a learning source image is divided into grids and what object is identified in the processing grid image obtained by the dividing means; a learning data generating unit that generates the learning data, The learning data generating means A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain an annotation grid image; a learning data acquisition means for acquiring object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing the learning source image, thereby obtaining learning data; In the image-based object estimation program, the computer further comprises: It functions as an additional learning data generating means, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation program comprising:
12. A learning data generation program for realizing, by a computer, a learning data generation device for generating learning data for training an object estimation means in an image-based object estimation device including a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image, and an object estimation means for estimating what an object is captured in the processing grid image obtained by dividing the original image into grids and learning data in which object identification information captured in the learning grid image is associated with each other, the learning data generation device comprising: A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain a learning grid image; A learning data generation program for causing a computer to function as learning data acquisition means for obtaining object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing a learning source image to obtain learning data, the computer further comprising: It functions as an additional learning data generating means, The additional learning data generating means includes: a learning image feature determining means for classifying each pixel included in each learned image that has already been used in learning based on a feature including a color or density thereof, and determining a representative feature of the learned image based on the number of pixels having each feature; A feature distribution calculation means for calculating a distribution of representative features in the plurality of learned images; a candidate image feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a training image selection means for selecting the candidate image as a training image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the trained images; A learning data generation program comprising:
13. A learning data generation program for realizing, by a computer, a learning data generation device for generating learning data for training an object estimation means in an image-based object estimation device including a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image, and an object estimation means for estimating what an object is captured in the processing grid image obtained by dividing the original image into grids and learning data in which object identification information captured in the learning grid image is associated with each other, the learning data generation device comprising: A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain a learning grid image; A learning data generation program for causing a computer to function as learning data acquisition means for obtaining object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing a learning source image to obtain learning data, the computer further comprising: It functions as an additional learning data generating means, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning data generation program comprising:
14. A learning data generation program for realizing, by a computer, a learning data generation device for generating learning data for training an object estimation means in an image-based object estimation device including a division means for dividing an original image captured while traveling on a road into grids of a predetermined size to obtain a processing grid image, and an object estimation means for estimating what an object is captured in the processing grid image obtained by dividing the original image into grids and learning data in which object identification information captured in the learning grid image is associated with each other, the learning data generation device comprising: A learning division means for dividing a learning annotation original image, which has been subjected to annotations for classifying objects in a captured learning original image, into grids to obtain a learning grid image; A learning data generation program for causing a computer to function as learning data acquisition means for obtaining object identification information of an annotation grid image based on an annotation of an object included in the annotation grid image, and adding the object identification information to a learning grid image obtained by dividing a learning source image to obtain learning data, the computer further comprising: It functions as an additional learning data generating means, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a trained image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated trained original images; a candidate original image position specifying means for specifying a grid position where an object appears among candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning data generation program comprising:
15. In any one of claims 9 to 14, the computer further comprises: a program that functions as a learning data generating means that generates, as learning data, a grid image in which the object estimation means has estimated an object with a probability higher than a predetermined value, and object identification information that is the object estimation result.
16. In any one of claims 9 to 15, the computer further comprises: The object estimation means functions as an update means for performing additional learning on the trained object estimation means based on additional learning data and updating the object estimation means in accordance with the results of the additional learning, The update means sets the object estimation means after learning based on additional learning data as a provisional object estimation means, provides evaluation data to the current object estimation means before the additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the provisional object estimation means, and if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is, characterized in that
17. a division means for dividing an image captured while traveling on a road into a processed image and dividing the processed image into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning image into grids and object identification information shown in the learning grid image are associated with each other; an updating means for performing additional learning on the trained object estimation means based on additional learning data and updating the object estimation means in accordance with a result of the additional learning; An image-based object estimation apparatus comprising: In the object estimation device, the update means sets the object estimation means after learning based on additional learning data as a tentative object estimation means, provides evaluation data to the current object estimation means before additional learning and the tentative object estimation means to calculate the object estimation accuracy of each, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is, further comprising additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learning image representative feature determining means for classifying each pixel included in each learning image already used in learning based on a feature including a color or density thereof, and determining a representative feature of the learning image based on the number of pixels having each feature; a representative feature distribution calculation means for calculating a distribution of representative features in the plurality of learning images; a candidate image representative feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a training image selection means for selecting the candidate image as a training image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the training images; An object estimation device comprising:
18. a division means for dividing an image captured while traveling on a road into a processed image and dividing the processed image into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning image into grids and object identification information shown in the learning grid image are associated with each other; an updating means for performing additional learning on the trained object estimation means based on additional learning data and updating the object estimation means according to a result of the additional learning; An image-based object estimation apparatus comprising: In the object estimation device, the update means sets the object estimation means after learning based on additional learning data as a tentative object estimation means, provides evaluation data to the current object estimation means before additional learning and the tentative object estimation means to calculate the object estimation accuracy of each, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is, further comprising additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when the representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the learning image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation device comprising:
19. a division means for dividing an image captured while traveling on a road into a processed image and dividing the processed image into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning image into grids and object identification information shown in the learning grid image are associated with each other; an updating means for performing additional learning on the trained object estimation means based on additional learning data and updating the object estimation means in accordance with a result of the additional learning; An image-based object estimation apparatus comprising: In the object estimation device, the update means sets the object estimation means after learning based on additional learning data as a tentative object estimation means, provides evaluation data to the current object estimation means before additional learning and the tentative object estimation means to calculate the object estimation accuracy of each, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as is, further comprising additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation device comprising:
20. An additional learning device for additionally learning an image-based object estimation device, comprising: a division means for taking an image captured while traveling on a road as a processed image, dividing the processed image into grids of a predetermined size to obtain a processed grid image; and an object estimation means for estimating what an object appears in the processed grid image obtained by dividing a learning image into grids and learning data in which object identification information appearing in the learning grid image is associated with each other, the additional learning device comprising: an update means for updating an object estimation means after learning based on additional learning data as a tentative object estimation means, a current object estimation means before additional learning and said tentative object estimation means are provided with evaluation data to calculate the object estimation accuracy of each of the object estimation means, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, the current object estimation means is replaced with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, the current object estimation means is used as is, further comprising additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learning image representative feature determining means for classifying each pixel included in each learning image already used in learning based on a feature including a color or density thereof, and determining a representative feature of the learning image based on the number of pixels having each feature; a representative feature distribution calculation means for calculating a distribution of representative features in the plurality of learning images; a candidate image representative feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a training image selection means for selecting the candidate image as a training image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the training images; An additional learning device comprising:
21. An additional learning device for additionally learning an image-based object estimation device, comprising: a division means for taking an image captured while traveling on a road as a processed image, dividing the processed image into grids of a predetermined size to obtain a processed grid image; and an object estimation means for estimating what an object appears in the processed grid image obtained by dividing a learning image into grids and learning data in which object identification information appearing in the learning grid image is associated with each other, the additional learning device comprising: an update means for updating an object estimation means after learning based on additional learning data as a tentative object estimation means, a current object estimation means before additional learning and said tentative object estimation means are provided with evaluation data to calculate the object estimation accuracy of each of the object estimation means, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, the current object estimation means is replaced with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, the current object estimation means is used as is, further comprising additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when the representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the learning image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; Additional learning device with.
22. An additional learning device for additionally learning an image-based object estimation device, comprising: a division means for taking an image captured while traveling on a road as a processed image, dividing the processed image into grids of a predetermined size to obtain a processed grid image; and an object estimation means for estimating what an object appears in the processed grid image obtained by dividing a learning image into grids and learning data in which object identification information appearing in the learning grid image is associated with each other, the additional learning device comprising: an update means for updating an object estimation means after learning based on additional learning data as a tentative object estimation means, a current object estimation means before additional learning and said tentative object estimation means are provided with evaluation data to calculate the object estimation accuracy of each of the object estimation means, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, the current object estimation means is replaced with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, the current object estimation means is used as is, further comprising additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a trained image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated trained original images; a candidate original image position specifying means for specifying a grid position where an object appears among candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; Additional learning device with.
23. An object estimation program for implementing an object estimation device by a computer, the program comprising: a division means for dividing an image captured while traveling on a road into a processed image and dividing the processed image into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning image into grids and object identification information shown in the learning grid image are associated with each other; An object estimation program for causing a trained object estimation means to function as an update means for performing additional learning based on additional learning data and updating the object estimation means in accordance with a result of the additional learning, comprising: an update means for updating the tentative object estimation means after learning based on additional learning data, and for providing evaluation data to the current object estimation means before the additional learning and the tentative object estimation means to calculate the object estimation accuracy of each of the tentative object estimation means; if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, the current object estimation means is replaced with the tentative object estimation means; and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, the current object estimation means is used as is; The computer is further configured to function as additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learning image representative feature determining means for classifying each pixel included in each learning image already used in learning based on a feature including a color or density thereof, and determining a representative feature of the learning image based on the number of pixels having each feature; a representative feature distribution calculation means for calculating a distribution of representative features in the plurality of learning images; a candidate image representative feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a training image selection means for selecting the candidate image as a training image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the training images; An object estimation program comprising:
24. An object estimation program for implementing an object estimation device by a computer, the program comprising: a division means for dividing an image captured while traveling on a road into a processed image and dividing the processed image into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning image into grids and object identification information shown in the learning grid image are associated with each other; An object estimation program for causing a trained object estimation means to function as an update means for performing additional learning based on additional learning data and updating the object estimation means in accordance with a result of the additional learning, comprising: an update means for updating the tentative object estimation means after learning based on additional learning data, and for providing evaluation data to the current object estimation means before the additional learning and the tentative object estimation means to calculate the object estimation accuracy of each of the tentative object estimation means; if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, the current object estimation means is replaced with the tentative object estimation means; and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, the current object estimation means is used as is; The computer is further configured to function as additional training data generating means for generating the additional training data, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when the representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the learning image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation program comprising:
25. An object estimation program for implementing an object estimation device by a computer, the program comprising: a division means for dividing an image captured while traveling on a road into a processed image and dividing the processed image into grids of a predetermined size to obtain a processed grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the learning image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the learning image into grids and object identification information shown in the learning grid image are associated with each other; An object estimation program for causing a trained object estimation means to function as an update means for performing additional learning based on additional learning data and updating the object estimation means in accordance with a result of the additional learning, comprising: In the object estimation program, the updating means sets the object estimation means after learning based on additional learning data as a tentative object estimation means, provides evaluation data to the current object estimation means before the additional learning and the tentative object estimation means to calculate the object estimation accuracy of each, and if the tentative object estimation means has a higher estimation accuracy than the current object estimation means, replaces the current object estimation means with the tentative object estimation means, and if the tentative object estimation means has a lower estimation accuracy than the current object estimation means, uses the current object estimation means as it is, further comprising: and causing the additional learning data generating means to generate the additional learning data, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation program comprising:
26. An additional learning program for implementing, by a computer, an additional learning device for additional learning of an image-based object estimation device, the image being captured while traveling on a road, the image being a processed image, the processed image being divided into grids of a predetermined size to obtain a processed grid image, and an object estimation means being trained based on learning data in which a learning grid image obtained by dividing a learning image into grids and object identification information appearing in the learning grid image are associated with each other, the additional learning program comprising: an additional learning program for functioning as an update means which sets an object estimation means after learning based on additional learning data as a provisional object estimation means, provides evaluation data to a current object estimation means before additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and replaces the current object estimation means with the provisional object estimation means if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, and uses the current object estimation means as it is if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, The computer is further and causing the additional learning data generating means to generate the additional learning data, The additional learning data generating means includes: a learning image representative feature determining means for classifying each pixel included in each learning image already used in learning based on a feature including a color or density thereof, and determining a representative feature of the learning image based on the number of pixels having each feature; a representative feature distribution calculation means for calculating a distribution of representative features in the plurality of learning images; a candidate image representative feature determining means for classifying each pixel included in the candidate image based on a feature including its color or density, and determining a representative feature of the candidate image based on the number of pixels having each feature; a training image selection means for selecting the candidate image as a training image when the representative feature of the candidate image is not contrary to the equalization of the distribution of the representative features of the training images; An additional learning program characterized by comprising:
27. An additional learning program for implementing, by a computer, an additional learning device for additional learning of an image-based object estimation device, the image being captured while traveling on a road, the image being a processed image, the processed image being divided into grids of a predetermined size to obtain a processed grid image, and an object estimation means being trained based on learning data in which a learning grid image obtained by dividing a learning image into grids and object identification information appearing in the learning grid image are associated with each other, the additional learning program comprising: an additional learning program for functioning as an update means which sets an object estimation means after learning based on additional learning data as a provisional object estimation means, provides evaluation data to a current object estimation means before additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and replaces the current object estimation means with the provisional object estimation means if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, and uses the current object estimation means as it is if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, The computer is further and causing the additional learning data generating means to generate the additional learning data, The additional learning data generating means includes: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when the representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the learning image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning program characterized by comprising:
28. An additional learning program for implementing, by a computer, an additional learning device for additional learning of an image-based object estimation device, the image being captured while traveling on a road, the image being a processed image, the processed image being divided into grids of a predetermined size to obtain a processed grid image, and an object estimation means being trained based on learning data in which a learning grid image obtained by dividing a learning image into grids and object identification information appearing in the learning grid image are associated with each other, the additional learning program comprising: an additional learning program for functioning as an update means which sets an object estimation means after learning based on additional learning data as a provisional object estimation means, provides evaluation data to a current object estimation means before additional learning and the provisional object estimation means to calculate the object estimation accuracy of each, and replaces the current object estimation means with the provisional object estimation means if the provisional object estimation means has a higher estimation accuracy than the current object estimation means, and uses the current object estimation means as it is if the provisional object estimation means has a lower estimation accuracy than the current object estimation means, The computer is further and causing the additional learning data generating means to generate the additional learning data, The additional learning data generating means includes: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a trained image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated trained original images; a candidate original image position specifying means for specifying a grid position where an object appears among candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning program characterized by comprising:
29. an original image division means for dividing an original image captured while driving on a road into grids of a predetermined size to obtain a processing grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the processing source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the processing source image into grids and object identification information shown in the learning grid image are associated with each other; An image-based object estimation apparatus comprising: A learning data generating means for generating the learning data is further provided, The learning data generating means a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation device comprising:
30. A learning data generation device for generating learning data for training an image-based object estimation device, the device comprising: a source image division means for dividing an image captured while driving on a road into a source image and obtaining a processing grid image by dividing the source image into grids of a predetermined size; and an object estimation means for estimating what an object appears in the processing grid image obtained by dividing a learning source image into grids and learning data in which object identification information appearing in the learning grid image is associated with each other, the device comprising: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning data generating device comprising:
31. 31. The apparatus of claim 29 or 30, a learned grid-of-interest feature distribution calculation means for selecting a learned grid-of-interest image including an object from among a plurality of learned grid images constituting a plurality of learned original images, and calculating a distribution of representative features of the learned grid-of-interest image; a candidate original image attention grid feature determining means for selecting a candidate attention grid image including an object from among the candidate grid images constituting the candidate original image, and calculating a representative feature of each candidate attention grid image; a second learning image selection means for determining whether or not a candidate original image is to be a learning original image based on a distribution of representative features of the learned grid image of interest and representative features of the candidate grid image of the candidate original image; An apparatus comprising:
32. An object estimation program for implementing an object estimation device by a computer, the program comprising: an original image division means for dividing an original image captured while driving on a road into grids of a predetermined size to obtain a processing grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the processing source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the processing source image into grids and object identification information shown in the learning grid image are associated with each other; a learning data generating unit that generates the learning data, The learning data generating means a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation program comprising:
33. A learning data generation program for realizing, by a computer, a learning data generation device that generates learning data for training an image-based object estimation device, the image being provided with: an original image division means that divides an image captured while traveling on a road into a processing original image, and obtains a processing grid image by dividing the original image into grids of a predetermined size; and an object estimation means that is trained based on learning data in which a learning grid image obtained by dividing a learning original image into grids and object identification information appearing in the learning grid image are associated with each other, and that estimates what objects appear in the processing grid image obtained by the original image division means, the computer comprising: a learned original image feature determining means for classifying, for each learned grid image constituting a learned original image that has already been used for learning, each pixel included in the learned grid image based on a feature including its color or density, and determining a representative feature of the learned grid image based on the number of pixels having each feature; a learned image feature distribution calculation means for calculating a distribution of representative features for each of the calculated grid positions in the plurality of learned original images; a candidate original image feature determining means for classifying, for each of the candidate grid images constituting the candidate original image, each pixel included in the candidate grid image based on a feature including its color or density, and determining a representative feature of the candidate grid image based on the number of pixels having each feature; a learning source image selection means for selecting the candidate original image as a learning source image when a representative feature of each candidate grid image of the candidate original image is not contrary to the equalization of the distribution of the representative features of the already-learned original image; a learning data generation program for causing the selected learning source image to function as learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image;
34. The program of claim 32 or 33, further comprising the steps of: a learned grid feature distribution calculation means for selecting a learned grid image of interest including an object from among a plurality of learned grid images constituting a plurality of learned original images, and calculating a distribution of representative features of the learned grid image of interest; a candidate original image attention grid feature determining means for selecting a candidate attention grid image including an object from among the candidate grid images constituting the candidate original image, and calculating a representative feature of each candidate attention grid image; A program for functioning as a second learning image selection means for determining whether or not to select a candidate original image as a learning original image based on the distribution of representative features of the already-learned grid image of interest and the representative features of the candidate grid image of the candidate original image.
35. an original image division means for dividing an original image captured while driving on a road into grids of a predetermined size to obtain a processing grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the processing source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the processing source image into grids and object identification information shown in the learning grid image are associated with each other; An image-based object estimation apparatus comprising: A learning data generating means for generating the learning data is further provided, The learning data generating means a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a trained image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated trained original images; a candidate original image position specifying means for specifying a grid position where an object appears among candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation device comprising:
36. A learning data generation device for generating learning data for training an image-based object estimation device, the device comprising: a source image division means for dividing an image captured while driving on a road into a source image and obtaining a processing grid image by dividing the source image into grids of a predetermined size; and an object estimation means for estimating what an object appears in the processing grid image obtained by dividing a learning source image into grids and learning data in which object identification information appearing in the learning grid image is associated with each other, the device comprising: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; A learning data generating device comprising:
37. An object estimation program for implementing an object estimation device by a computer, the program comprising: an original image division means for dividing an original image captured while driving on a road into grids of a predetermined size to obtain a processing grid image; an object estimation means for estimating what object is shown in the processing grid image obtained by dividing the processing source image into grids, the object estimation means being trained based on learning data in which a learning grid image obtained by dividing the processing source image into grids and object identification information shown in the learning grid image are associated with each other; a learning data generating unit that generates the learning data, The learning data generating means a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image; An object estimation program comprising:
38. A learning data generation program for realizing, by a computer, a learning data generation device that generates learning data for training an image-based object estimation device, the image being provided with: an original image division means that divides an image captured while traveling on a road into a processing original image, and obtains a processing grid image by dividing the original image into grids of a predetermined size; and an object estimation means that is trained based on learning data in which a learning grid image obtained by dividing a learning original image into grids and object identification information appearing in the learning grid image are associated with each other, and that estimates what objects appear in the processing grid image obtained by the original image division means, the computer comprising: a learned original image frequency calculation means for classifying a plurality of learning grid images already used in learning by their positions on the learning original image, and calculating the number of objects appearing in the learned grid image for each of the positions based on the object identification information; a learned image frequency distribution calculation means for calculating a distribution of the number of occurrences of objects for each grid position in the calculated plurality of learned original images; a candidate original image position specifying means for specifying a grid position where an object appears among the candidate grid images constituting the candidate original image; a learning image selection means for selecting a candidate original image as a learning original image when it is determined that the candidate original image is appropriate to be used as a learning original image based on a distribution of the number of appearances of the object according to grid positions in the learning original image and the grid positions at which the object appears in the candidate original image; a learning data generation program for causing the selected learning source image to function as learning source image dividing means for dividing the selected learning source image into grids of a predetermined size to obtain a learning grid image;
Citation Information
Patent Citations
Object recognizing method and object recognizing device
JP2000242784A
Methods and systems for wound assessment and management
JP2017504370A
Image processing system and image processing method
JP2019139497A
Image determination system, model update method, and model update program
JP2019152948A
Information processing device, system, control method thereof, and program
JP2020038538A