Learning data selection device, learning data selection method and learning data selection program

The learning data selection device addresses biases in feature amounts by separately extracting object and background features, enhancing the generalization performance of machine learning models through tailored data selection.

JP2025094394APending Publication Date: 2025-06-25KONICA MINOLTA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023209877
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Existing machine learning techniques for selecting learning data fail to account for biases in feature amounts extracted from both objects and backgrounds in images, leading to decreased generalization performance and unintended learning outcomes.

Method used

A learning data selection device that separately extracts object and background feature amounts from images, adjusting their weights based on importance to calculate a final feature amount for selecting appropriate learning data.

Benefits of technology

This approach allows for the selection of unbiased learning data, improving the generalization performance of machine learning models by ensuring that the selected data varies appropriately based on the application's requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094394000001_ABST
    Figure 2025094394000001_ABST
Patent Text Reader

Abstract

To provide a technology for selecting learning data based on feature amounts extracted individually from objects and backgrounds captured in images.SOLUTION: A learning data selection device 100 comprises an object feature amount extraction unit 102 for extracting an object feature amount 120 from each of one or more images 110, a background feature amount extraction unit 104 for extracting a background feature amount 130 from each of the one or more images 110, and a selection unit 106 for selecting one or more learning data for machine learning from the one or more images 110. The selection unit 106 calculates a final feature amount for each of the one or more images 110 based on the object feature amount 120 and the background feature amount 130, and selects one or more learning data based on the final feature amount for each of the one or more images 110.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for selecting images used as learning data for machine learning, and more specifically, to a method for calculating feature amounts of images.

Background Art

[0002] Machine learning requires a large amount of learning data. As an example, when creating a model for analyzing images captured by a surveillance camera, the learning engine requires a large number of images as learning data. However, when learning data is randomly selected, there may be a bias in the learning data or data unnecessary for learning may be selected. As a result, the generalization performance of the model may decrease or the model may learn unintentionally. Therefore, in order to improve the generalization performance of the model, a technique for easily selecting unbiased learning data is required.

[0003] Regarding the technique for selecting learning data, for example, Japanese Unexamined Patent Application Publication No. 2022-112819 (Patent Document 1) discloses an image data set generation device that "includes a generation unit and an exclusion unit. The generation unit generates an image data set by cutting out a region corresponding to the detection range of an object from a group of images and labeling the region with a class name. The exclusion unit excludes at least one of the image data having similar feature amounts from the image data set" (see the summary).

[0004] Also, other techniques related to the technique for selecting learning data are disclosed, for example, in Japanese Unexamined Patent Application Publication No. 2020-052999 (Patent Document 2).

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] According to the techniques disclosed in Patent Documents 1 and 2, feature amounts are extracted from an image without considering the objects and background shown in the image. When learning data is selected based on such feature amounts, in the learning data, elements of either the object or the background may be biased. Therefore, there is a need for a technique for selecting learning data based on feature amounts individually extracted from each of the objects and background shown in the image.

[0007] The present disclosure has been made in view of the above background, and an object in one aspect is to provide a technique for selecting learning data based on feature amounts individually extracted from each of the objects and background shown in the image.

Means for Solving the Problems

[0008] According to an embodiment, a learning data selection device is provided. The learning data selection device includes an object feature amount extraction unit that extracts an object feature amount from each of one or more images, a background feature amount extraction unit that extracts a background feature amount from each of one or more images, and a selection unit that selects one or more learning data for machine learning from among the one or more images. The selection unit calculates a final feature amount for each of the one or more images based on the object feature amount and the background feature amount, and selects one or more learning data based on the final feature amount for each of the one or more images.

[0009] In one aspect, calculating the final feature amount for each of the one or more images includes multiplying the object feature amount by an object importance to obtain a first calculation result, multiplying the background feature amount by a background importance to obtain a second calculation result, and concatenating the first calculation result and the second calculation result.

[0010] In a certain situation, calculating the final feature amount of each of one or more images includes multiplying the object feature amount by a first coefficient for correcting the value of an attribute of a part of the object before calculating the first calculation result and the second calculation result, and multiplying the background feature amount by a second coefficient for correcting the value of an attribute of a part of the background.

[0011] In a certain situation, the object feature amount includes at least one of information regarding the size of the object, the color of the object, the posture of the object, and the orientation of the object.

[0012] In a certain situation, the learning data selection device further includes an object extraction unit that extracts an object image from each of one or more images. Extracting the object feature amount from each of one or more images includes extracting the object feature amount from the object images extracted from each of one or more images.

[0013] In a certain situation, the object feature amount includes at least one of the output result of an intermediate layer of a neural network, the size of the object, the color of the object, the posture of the object, and the orientation of the object.

[0014] In a certain situation, the background feature amount includes at least one of information regarding the shooting location of the background, the shooting time of the background, the camera position at the time of background shooting, and the weather of the background.

[0015] In a certain situation, the learning data selection device further includes a background extraction unit that extracts a background image from each of one or more images. Extracting the background feature amount from each of one or more images includes extracting the background feature amount from the background images extracted from each of one or more images.

[0016] In a certain situation, the background feature amount includes at least one of the output result of an intermediate layer of a neural network, the shooting location of the background, the shooting time of the background, the camera position at the time of background shooting, and the weather of the background.

[0017] In a certain situation, when each of one or more images is one or more videos, the object feature amount extraction unit extracts object feature amounts from objects shown in one or more frames constituting each of the one or more videos, and the background feature amount extraction unit extracts background feature amounts from one frame constituting each of the one or more videos.

[0018] According to another embodiment, a learning data selection method is provided. The learning data selection method includes extracting object feature amounts from each of one or more images, extracting background feature amounts from each of one or more images, calculating final feature amounts for each of one or more images based on the object feature amounts and the background feature amounts, and selecting one or more learning data for machine learning from among one or more images based on the final feature amounts for each of one or more images.

[0019] Furthermore, according to another embodiment, a learning data selection program is provided. The learning data selection program causes a computer to extract object feature amounts from each of one or more images, extract background feature amounts from each of one or more images, calculate final feature amounts for each of one or more images based on the object feature amounts and the background feature amounts, and select one or more learning data for machine learning from among one or more images based on the final feature amounts for each of one or more images.

Advantages of the Invention

[0020] According to an embodiment, learning data can be selected based on feature amounts individually extracted from an object and a background shown in an image.

[0021] The above and other objects, features, aspects and advantages of this disclosure will become apparent from the following detailed description of the disclosure understood in connection with the accompanying drawings.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0023] Hereinafter, embodiments of the technical idea according to the present disclosure will be described with reference to the drawings. In the following description, the same parts are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed descriptions thereof will not be repeated. Also, each embodiment, each modification example, each software configuration, each hardware configuration, each function, and each process, etc. may be selectively combined as appropriate.

[0024] <A. Outline of the Technology of the Present Disclosure> FIG. 1 is a diagram showing an example of the outline of the operation of the learning data selection device 100 according to the present embodiment. The learning data selection device 100 selects learning data for machine learning from one or more images. The learning data here is an image selected from one or more images. Alternatively, the learning data may be a feature amount of an image selected from one or more images. Hereinafter, the features of the image as learning data and the selection criteria for the learning data by the learning data selection device 100 will be described.

[0025] (a. Image as learning data) In recent years, in various fields, a model that has learned some data by machine learning has been used. The model obtained by machine learning may also be called a machine learning model or an AI (Artificial Intelligence) model. Machine learning is roughly divided into supervised learning and unsupervised learning. In supervised learning, a combination of learning data and correct answers is input to the model. The model learns based on the combination of learning data and correct answers. The correct answer here may also include a label or the like for classifying data. On the other hand, in unsupervised learning, only learning data is input to the model. The model learns based on the learning data. In any case, learning data is required for the learning of the model. The learning data varies depending on the use of the model. Generally, images or videos are used as learning data for models used in applications such as object classification, product inspection, and detection of human behavior.

[0026] When learning data is input to the model, the feature amount of the learning data is extracted. As an example, when the learning data is an image, the feature amount of the image is extracted based on some algorithm. The feature amount of the image can be calculated from the color of each pixel, the edge, any other arbitrary item, or a combination thereof. This feature amount is given to the model as learning data. Alternatively, the model may be equipped with a function of extracting feature amounts from data in any format.

[0027] Here, the feature amounts of the images will be specifically described. An image often includes a main subject and the background other than the main subject. Hereinafter, the main subject will be referred to as an "object", and the other part will be referred to as the "background". Since an image includes two elements, an object and a background, the feature amount of the entire image is a mixture of the feature amount of the object and the feature amount of the background. Therefore, the feature amount of the entire image is affected by both the object and the background. Hereinafter, the feature amount of the entire image will be referred to as an "image feature amount", the feature amount of the object will be referred to as an "object feature amount", and the feature amount of the background will be referred to as a "background feature amount".

[0028] In the prior art, when an image was used as learning data, the image feature amount was input to the model. That is, the learning data was data in which the object feature amount and the background feature amount were mixed. Therefore, in the selected learning data, one of the elements of the object or the background might be biased. As a result, a learning result unintended by the developer might be obtained.

[0029] (b. Outline of learning data selection device) The learning data selection device 100 according to the present embodiment individually extracts an object feature amount and a background feature amount from an image instead of the image feature amount. Then, the learning data selection device 100 selects learning data to be used for learning the model based on the object feature amount and the background feature amount. Thereby, compared with the prior art, the learning data selection device 100 can suppress one of the elements of the object or the background from being biased and can select the learning data more appropriately.

[0030] The learning data selection device 100 includes an object feature amount extraction unit 102, a background feature amount extraction unit 104, and a selection unit 106. The learning data selection device 100 receives an input of one or more images 110 that are candidates for learning data. The learning data selection device 100 may store the input one or more images 110 in a storage unit 406 (see FIG. 4).

[0031] The object feature amount extraction unit 102 extracts the object feature amount 120 from the image 110 input to the learning data selection device 100. In a certain aspect, when a plurality of objects are shown in the image 110, the object feature amount extraction unit 102 may extract the feature amounts of each object. In another aspect, when a plurality of objects are shown in the image 110, the object feature amount extraction unit 102 may regard the plurality of objects as one object and extract one object feature amount. The background feature amount extraction unit 104 extracts the background feature amount 130 from the image 110 input to the learning data selection device 100.

[0032] The selection unit 106 selects one or more learning data from among one or more images 110 based on the object feature amount 120 and the background feature amount 130. More specifically, the selection unit 106 receives an input from the user regarding the use or importance of the model. The importance refers to the importance of each of the object feature amount 120 and the background feature amount 130. Hereinafter, the importance of the object feature amount 120 is referred to as the "object importance". Similarly, the importance of the background feature amount 130 is referred to as the "background importance". The object importance is used as a coefficient to be multiplied by the object feature amount 120. The background importance is used as a coefficient to be multiplied by the background feature amount 130. In a certain aspect, the selection unit 106 may determine the values of the object importance and the background importance respectively based on the use of the input model. As an example, the selection unit 106 selects an image to be used for learning based on the concatenated value of the object feature amount 120 multiplied by the object importance and the background feature amount 130 multiplied by the background importance. Hereinafter, the said concatenated value is referred to as the "final feature amount" of the image 110 which is a candidate for learning data.

[0033] As an example, assume that the use of the model is to determine the expression, age, and gender of a person indoors. In this case, a photo or image of a person is used as learning data. Considering the application's use, the object becomes an important element in learning. Here, the object is a person. On the other hand, the background is not such an important element in learning. In this case, the value of the object importance becomes large, and the value of the background importance becomes small. As a result, the final feature amount is greatly influenced by the object feature amount 120. The selection unit 106 selects learning data from among one or more images 110 such that the variation in the feature amounts becomes large. As a result, an image including various objects is likely to be selected as learning data. That is, a photo showing various people is likely to be selected as learning data.

[0034] As another example, assume that the use of the model is the recognition of an object shown in an image captured by a camera installed outdoors for 24 hours. In this case, a photo taken outdoors is used as learning data. Considering the application's use, the background becomes an important element in learning. This is because the model needs to detect an object regardless of the time zone and weather such as morning, noon, night, sunny, and rainy. On the other hand, if a detailed classification of the detected object is not required, the object is not such an important element in learning. In this case, the value of the object importance becomes small, and the value of the background importance becomes large. As a result, the final feature amount is greatly influenced by the background feature amount 130. The selection unit 106 selects learning data from among one or more images 110 such that the variation in the feature amounts becomes large. As a result, an image including various backgrounds is likely to be selected as learning data.

[0035] As described above, the learning data selection device 100 includes an object feature amount extraction unit 102 that extracts an object feature amount 120 from each of one or more images 110, a background feature amount extraction unit 104 that extracts a background feature amount 130 from each of the one or more images 110, and a selection unit 106 that selects one or more learning data for machine learning from among the one or more images 110. The selection unit 106 calculates the final feature amount of each of the one or more images 110 based on the object feature amount 120 and the background feature amount 130, and selects one or more learning data based on the final feature amount of each of the one or more images 110.

[0036] In this way, the learning data selection device 100 can select learning data based on the individually calculated object importance and background importance. Thereby, the learning data selection device 100 can select appropriate learning data according to the use of the model. As a result, the generalization performance of the model after learning is improved.

[0037] (c. Terms) Next, definitions and examples of some important terms in this specification will be described. In this specification, the "device" includes any information processing device such as a personal computer, a workstation, a server, a tablet, or a smartphone. Also, the device may be a combination of these. The learning data selection device 100 can be realized by any device. Also, when the learning data selection device 100 is composed of a plurality of devices, the learning data selection device 100 may be renamed as a learning data selection system. In one aspect, the learning data selection device 100 may be connected to input / output devices such as a display and a keyboard and used by a user. In another aspect, the learning data selection device 100 may provide various functions to a user as a service or a web application via a network. In this case, the user can use the functions of the learning data selection device 100 via a browser or client software installed on their own terminal. Further, in another aspect, the learning data selection device 100 may be constructed as a virtual machine or an instance on a cloud environment.

[0038] In this specification, "learning data" refers to any data used as input data for machine learning. As an example, learning data includes images, videos, audio, text, binary data, output results of intermediate layers of neural networks, any other data, or combinations thereof. Further, learning data includes not only original data such as images, but also feature amounts extracted from the original data.

[0039] In this specification, a "feature amount" is data directly input to a model for machine learning. In a certain aspect, a feature amount can be expressed in any format such as a matrix, an array, a list, JSON (JavaScript (registered trademark) Object Notation), binary data, etc. An "image feature amount" is a feature amount obtained from an entire image. An "object feature amount" is a feature amount obtained from an object in an image. A "background feature amount" is a feature amount obtained from a background in an image. The "final feature amount" calculated by the selection unit 106 based on the object feature amount 120 and the background feature amount 130 can also be said to be an image feature amount.

[0040] In this specification, "importance" is data used as a coefficient for changing the value of a feature amount. As an example, the object importance changes the value of the object feature amount 120. Importance can be expressed in any format such as a matrix, an array, a list, JSON, binary data, etc. In a certain aspect, importance may change the values of some of the feature amounts. As an example, the feature amounts of an image include feature amounts of various attributes such as color, edges, size, object pose, and object orientation. Importance may change the values of some of the feature amounts of these attributes. For example, if the color of an image is an important element, importance may change only the values of the feature amounts related to the color included in the feature amounts of the image.

[0041] In this specification, "generalization performance" refers to the performance of being able to correctly predict not only learning data but also unknown data. For example, assume there is a model for determining the types of animals and plants from images. If the model can correctly determine the types of animals and plants with a high probability from images not used in the learning data, it can be said that the generalization performance of the model is high.

[0042] In this specification, "concatenation" is a computational process or processing operation that combines two or more data indicating feature quantities and converts them into one data. Concatenation is frequently used in array processing in programming languages. Also, "concatenated value" is the value obtained by concatenation. Usually, in a program, feature quantities are represented in formats such as matrices, arrays, lists, or objects. As an example, assume there are feature quantities 1 and 2 represented as arrays. And assume that feature quantity 1 = [[1, 1, 1], [1, 1, 1]]. Also assume that feature quantity 2 = [[2, 2, 2], [2, 2, 2]]. In this case, the concatenated value of feature quantities 1 and 2 = [[1, 1, 1], [1, 1, 1], [2, 2, 2], [2, 2, 2]]. The concatenation operator is different for each programming language. Operators such as "+" and "&" may be used as concatenation operators. Also, some programming languages have methods for concatenating arrays, etc.

[0043] Figure 2 is a diagram showing an example of a problem that may occur when selecting data for machine learning. Referring to Figure 2, the problems when extracting image feature quantities without separating object feature quantities and background feature quantities will be described. In the following explanations, several images may be shown as examples of learning data in each figure, but these are just examples for explanation. In actual machine learning, an enormous number of learning data are used.

[0044] Generally, when selecting training data, it is desirable to select training data such that the feature quantities vary as much as possible. Regarding the selection of training data, a model for discriminating dog breeds will be used as an example for explanation. Images of the same dog breed have similar feature quantities. Therefore, a model that inputs only images with similar feature quantities as training data can only discriminate specific dog breeds. Therefore, it is desirable that the training data includes data with different feature quantities as much as possible.

[0045] Figure 2 shows images 202, 204, 206, 208, 210, and 212. Images 202, 204, and 206 are photos of children in the wild. Images 208, 210, and 212 are photos of adults or children indoors.

[0046] As an example, assume that a developer wants to develop a model for identifying the presence or absence of people in various locations. In this case, it is desirable that the training data includes images with various backgrounds. That is, it is desirable that at least one image is selected from images 202, 204, and 206 and at least one image is selected from images 208, 210, and 212 as training data.

[0047] Here, assume that the developer directly extracts image feature quantities from images 202, 204, 206, 208, 210, and 212 using some tool. And assume that the developer selects several images such that the feature quantities vary as much as possible. However, with this method, it is not always the case that images with different backgrounds are selected. For example, in images 208, 210, and 212, the types and sizes of the human figures, which are the objects, are very different. That is, the feature quantities of images 208, 210, and 212 are very different. As a result, the learning engine may select images 208, 210, and 212 as training data. In this case, the backgrounds included in the training data will be biased, and the generalization performance of the model will deteriorate.

[0048] Thus, when the learning data is selected based only on the feature amounts of the entire image, there is a possibility that the objects or the background shown in the selected images will be biased. Therefore, the learning data selection device 100 separately extracts the object feature amount 120 and the background feature amount 130. Further, the learning data selection device 100 adjusts the weights of the object feature amount 120 and the background feature amount 130 using the importance to calculate the final feature amount. The learning data selection device 100 can select appropriate learning data according to the application by selecting the learning data so that the final feature amounts vary.

[0049] FIG. 3 is a diagram showing an example in which variations of an object or a background are important as learning data. Whether to emphasize an object or a background shown in an image as learning data is determined by the application of the model. Taking FIG. 3 as an example, the applications of a model that emphasizes background variations and a model that emphasizes object variations will be described.

[0050] First, an example of a model that requires many background variations as learning data will be described. As a model that requires many background variations as learning data, there is an object detection model mounted on a camera installed outdoors for 24 hours. For learning such an object detection model, images with various backgrounds are required. For example, assume that the object detection model has learned only images of a person walking on a sunny street during the day. In this case, the object detection model may not be able to detect a person walking on the street on a rainy day or at night.

[0051] For learning such an object detection model, an image set including various backgrounds such as group 300 is suitable. In the example of FIG. 3, group 300 includes images 302, 304, and 306. Image 302 is an image showing a scene where a man and a woman are walking on a sunny mountain path. Image 304 is an image showing a scene where a man is walking while holding an umbrella in the middle of a town on a rainy day. Image 306 is an image showing a scene where a child is running on a forest path in the evening. Actually, more images are used in the learning of the object detection model.

[0052] Next, an example of a model that requires a large number of variations of an object as learning data will be described. As a model that requires a large number of variations of an object as learning data, there is a person attribute determination model mounted on a camera installed indoors. For learning such an attribute determination model, various images of the object are required. More specifically, the attribute determination model requires images of various people with different beards, wrinkles, genders, presence or absence of glasses, hairstyles, etc. as learning data. In addition, since the camera equipped with the attribute determination model is installed indoors, it inevitably captures the same background. Therefore, in the learning of the attribute determination model, background variation is not important.

[0053] For learning such an attribute determination model, an image set that captures the characteristics of the object in detail, such as group 310, is suitable. In the example of FIG. 3, group 310 includes images 312 and 314. Image 312 is an image depicting a scene where a young man is walking indoors. Image 314 is an image depicting a scene where an elderly man and woman are indoors. Images 312 and 314 depict the characteristics of the person in more detail compared to the images included in group 300. In reality, more images are used in the learning of the attribute determination model.

[0054] As described with reference to FIG. 3, whether to emphasize the object or the background depicted in the image as learning data depends on the use of the model. The learning data selection device 100 can select an appropriate learning set including a plurality of images or videos according to the use of the model. That is, the learning data selection device 100 can select a learning set including a large number of variations of the object or a learning set including a large number of variations of the background according to the use of the model.

[0055] <B. Device Configuration> FIG. 4 is a diagram showing an example of the functional configuration of a learning data selection device 100 for selecting images to be used for machine learning from one or more images 110. The learning data selection device 100 includes an input unit 400, an object extraction unit 402, a background extraction unit 404, an object feature amount extraction unit 102, a background feature amount extraction unit 104, a selection unit 106, a storage unit 406, and an output unit 408.

[0056] The input unit 400 receives the input of one or more images 110 that are candidates for learning data. The input unit 400 outputs the one or more images 110 to either the object extraction unit 402 or the object feature amount extraction unit 102. Similarly, the input unit 400 outputs the one or more images 110 to either the background extraction unit 404 or the background feature amount extraction unit 104.

[0057] The object extraction unit 402 extracts an object image from each of the one or more images 110. The object image is an image of the portion where the main subject is depicted. As an example, assume that a person walking on a mountain path is depicted in a certain image. And assume that the main subject is the person, and the other portions are the background. In this case, the object extraction unit 402 extracts the portion where the person is depicted from the certain image as the object image. The object extraction unit 402 outputs the extracted object image to the object feature amount extraction unit 102. In a certain aspect, the object extraction unit 402 can obtain the contour information of the object from each of the one or more images 110 by means of a background difference method, optical flow, a neural network model, etc. And the object extraction unit 402 extracts an object image from each of the one or more images 110 based on the contour information of the object.

[0058] The background extraction unit 404 extracts a background image from each of one or more images 110. The background image is an image of only the background excluding the portion where the main subject is depicted. As an example, assume that a human walking on a mountain path is depicted in an image. And assume that the main subject is the human, and the other portions are the background. In this case, the background extraction unit 404 extracts, as the background image, the portion other than the portion where the human is depicted from the said image. The background extraction unit 404 outputs the extracted background image to the background feature amount extraction unit 104. In a certain aspect, the background extraction unit 404 can obtain contour information of an object from each of one or more images 110 by means of a background difference method, optical flow, a neural network model, or the like. Then, the background extraction unit 404 extracts or generates a background image from each of one or more images 110 based on the contour information of the object.

[0059] The object feature amount extraction unit 102 extracts an object feature amount 120 from each of one or more images 110. Alternatively, the object feature amount extraction unit 102 may extract the object feature amount 120 from the object image input from the object extraction unit 402. The object feature amount extraction unit 102 outputs the object feature amount 120 to the selection unit 106 or the storage unit 406.

[0060] The background feature amount extraction unit 104 extracts a background feature amount 130 from each of one or more images 110. Alternatively, the background feature amount extraction unit 104 may extract the background feature amount 130 from the background image input from the background extraction unit 404. The background feature amount extraction unit 104 outputs the background feature amount 130 to the selection unit 106 or the storage unit 406.

[0061] The selection unit 106 selects training data from one or more images 110 based on the object feature quantity 120 and the background feature quantity 130. At this time, the selection unit 106 receives an input from the user regarding the use or importance of the model. The selection unit 106 determines the object importance and the background importance based on the input regarding the use or importance of the model. Then, the selection unit 106 multiplies the object importance by the object feature quantity 120 to obtain a first calculation result. Similarly, the selection unit 106 multiplies the background importance by the background feature quantity 130 to obtain a second calculation result. The selection unit 106 calculates the final feature quantity of each of the one or more images 110 by concatenating the first calculation result and the second calculation result. Next, the selection unit 106 selects one or more training data from the one or more images 110 based on the final feature quantity of each of the one or more images 110. The selection unit 106 outputs the selected one or more training data to the output unit 408.

[0062] The storage unit 406 stores each of the one or more images 110 in association with the object feature quantity 120 and the background feature quantity 130. The selection unit 106 does not need to immediately select training data from the input one or more images 110. The selection unit 106 may select one or more training data from the one or more images 110 stored in the storage unit 406 in the past based on a request from the user.

[0063] The output unit 408 outputs the selected one or more training data. Assume that the training data selection device 100 is a stand-alone device. In this case, the output unit 408 can output the one or more training data to an arbitrary directory on the device or an external storage device. Assume that the training data selection device 100 is a web application or a service in a cloud environment. In this case, the output unit 408 can transmit the one or more training data to the user's terminal.

[0064] FIG. 5 is a diagram showing an example of the hardware configuration of the learning data selection device 100. In a certain aspect, by executing a program on the hardware shown in FIG. 5, each function shown in FIG. 4 or FIG. 10 can be realized. Depending on embodiments such as a stand-alone device and cloud services, the hardware required by the learning data selection device 100 may be different. Therefore, the learning data selection device 100 may include a plurality of some or all of the hardware shown in FIG. 5. Also, the learning data selection device 100 may not include some of the hardware shown in FIG. 5.

[0065] The learning data selection device 100 includes a processor 501, a primary storage device 502, a secondary storage device 503, an external device interface 504, an input interface 505, an output interface 506, and a communication interface 507. Also, these components are configured to be communicable with each other via a bus.

[0066] The processor 501 can execute a program for realizing various functions of the learning data selection device 100. The processor 501 is constituted by, for example, at least one integrated circuit. According to a certain embodiment, the integrated circuit may include at least one CPU (Central Processing Unit), at least one GPU (Graphics Processing Unit), at least one FPGA (Field Programmable Gate Array), at least one ASIC (Application Specific Integrated Circuit), at least one AI (Artificial Intelligence) chip, or a combination thereof, etc.

[0067] The primary storage device 502 functions as the workspace of the processor 501. For this purpose, the primary storage device 502 stores the programs executed by the processor 501 and the data referenced by the processor 501. In certain aspects, the primary storage device 502 can be implemented by DRAM (Dynamic Random Access Memory) or SRAM (Static Random Access Memory), etc.

[0068] The secondary storage device 503 is a non-volatile memory and stores the programs executed by the processor 501 and the data referenced by the processor 501. The processor 501 executes the programs read from the secondary storage device 503 into the primary storage device 502 and references the data read from the secondary storage device 503 into the primary storage device 502. In certain aspects, the secondary storage device 503 can be implemented by an HDD (Hard Disk Drive), an SSD (Solid State Drive), an EPROM (Erasable Programmable Read Only Memory), an EEPROM (Electrically Erasable Programmable Read Only Memory), or a flash memory, etc.

[0069] The external device interface 504 can be connected to any external device such as a printer, a scanner, and an external HDD. In certain aspects, the external device interface 504 can be implemented by a USB (Universal Serial Bus) terminal, etc.

[0070] The input interface 505 can be connected to any input device such as a keyboard, a mouse, a touch pad, or a game pad. In certain aspects, the input interface 505 can be implemented by a USB terminal, a PS / 2 terminal, and a Bluetooth (registered trademark) module, etc. In other aspects, the input interface 505 may be configured integrally with any input device.

[0071] The output interface 506 can be connected to any output device such as a cathode ray tube display, a liquid crystal display, or an organic EL (Electro-Luminescence) display. In one aspect, the output interface 506 can be realized by a USB terminal, a D-sub terminal, a DVI (Digital Visual Interface) terminal, an HDMI (registered trademark) (High-Definition Multimedia Interface) terminal, a display port terminal, etc. In other aspects, the output interface 506 may be configured integrally with any output device.

[0072] The communication interface 507 is connected to other devices via a wired network or a wireless network. In one aspect, the communication interface 507 can be realized by a wired LAN (Local Area Network) port, a Wi-Fi (registered trademark) (Wireless Fidelity) module, etc. In other aspects, the communication interface 507 can transmit and receive data using communication protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol) and UDP (User Datagram Protocol).

[0073] <C. Operating Example of the Device> FIG. 6 is a diagram showing an example of a procedure for extracting the object feature amount 120 from the image 110. As described above, the object feature amount extraction unit 102 may directly extract the object feature amount 120 from the image 110. In this case, a model learned to extract the object feature amount 120 from the image 110 is used as the object feature amount extraction unit 102.

[0074] Alternatively, the object feature amount extraction unit 102 may extract the object feature amount 120 from the object image 600. In this case, a model learned to extract the object feature amount 120 from the object image 600 is used as the object feature amount extraction unit 102.

[0075] The object extraction unit 402 can obtain the contour information of the object from the image 110 by means of background difference method, optical flow, neural network model, etc. Then, based on the contour information of the object, the object extraction unit 402 extracts the object image 600 of the image 110. In one aspect, the object image 600 may be an image 600A in which only the object is cut out. In another aspect, the object image 600 may be an image 600B in which the object and its surroundings are cut out in a square shape. Further, in another aspect, the object image 600 may be an image 600C in which the object and its surroundings are cut out in an arbitrary shape.

[0076] In one aspect, the object feature quantity 120 includes at least one of the size of the object, the color of the object, the posture of the object, the orientation of the object, and any other information. In another aspect, the object feature quantity 120 may be the output result of the intermediate layer of the neural network. The object feature quantity extraction unit 102 can be realized as a neural network model. In this case, the object feature quantity extraction unit 102 may output the output result of the intermediate layer of the object feature quantity extraction unit 102 as the object feature quantity 120. The output result of the intermediate layer cannot be understood by humans but includes a large number of parameters regarding the object. Therefore, the output result of the intermediate layer can be useful learning data. As an example, the output result of the intermediate layer may include at least one of information regarding the size of the object, the color of the object, the posture of the object, and the orientation of the object. When the object feature quantity extraction unit 102 outputs the output result of the intermediate layer, the object feature quantity extraction unit 102 needs to be input with the object image 600. This is because when the original image 110 is input to the object feature quantity extraction unit 102, the output result of the intermediate layer will include the features of the background.

[0077] As described above, the object feature amount 120 includes at least one of information regarding the size of the object, the color of the object, the posture of the object, and the orientation of the object. Further, the learning data selection device 100 includes an object extraction unit 402 that extracts an object image 600 from each of the one or more images 110. And extracting the object feature amount 120 from each of the one or more images 110 includes extracting the object feature amount 120 from the object image 600 extracted from each of the one or more images 110. Furthermore, the object feature amount 120 may include at least one of the output result of the intermediate layer of the neural network, information regarding the size of the object, the color of the object, the posture of the object, and the orientation of the object.

[0078] FIG. 7 is a diagram showing an example of a procedure for extracting the background feature amount 130 from the image 110. As described above, the background feature amount extraction unit 104 may directly extract the background feature amount 130 from the image 110. In this case, a model trained to extract the background feature amount 130 from the image 110 is used as the background feature amount extraction unit 104.

[0079] Alternatively, the background feature amount extraction unit 104 may extract the background feature amount 130 from the background image 700. In this case, a model trained to extract the background feature amount 130 from the background image 700 is used as the background feature amount extraction unit 104.

[0080] The background extraction unit 404 can obtain the contour information of the object from the image 110 by means of a background difference method, optical flow, a neural network model, or the like. And the background extraction unit 404 extracts or generates the background image 700 of the image 110 based on the contour information of the object. In one aspect, the background image 700 may be an image 700A with only the object removed. In another aspect, the background image 700 may be an image 700B in which the background synthesized at the location where the object was present is embedded. Furthermore, in another aspect, the background image 700 may be an image 700C in which no object photographed at the same location as the image 110 is shown.

[0081] In a certain situation, the background feature quantity 130 includes at least one of information regarding the background shooting location, the background shooting time, the camera position at the time of background shooting, the weather of the background, and any other arbitrary information. In another situation, the background feature quantity 130 may be the output result of an intermediate layer of a neural network. The background feature quantity extraction unit 104 can be realized as a neural network model. In this case, the background feature quantity extraction unit 104 may output the output result of the intermediate layer of the background feature quantity extraction unit 104 as the background feature quantity 130. Although the output result of the intermediate layer cannot be understood by humans, it includes a large number of parameters regarding the background. Therefore, the output result of the intermediate layer can become useful learning data. As an example, the output result of the intermediate layer may include at least one of the background shooting location, the background shooting time, the camera position at the time of background shooting, the weather of the background, and any other arbitrary information. When the background feature quantity extraction unit 104 outputs the output result of the intermediate layer, the background feature quantity extraction unit 104 needs to be input with the background image 700. This is because when the original image 110 is input to the background feature quantity extraction unit 104, the output result of the intermediate layer will include the features of the object. Furthermore, in another situation, the background feature quantity extraction unit 104 may directly obtain the feature quantity from the camera. In this case, the camera has a function of extracting the feature quantity from the image.

[0082] When obtaining the feature quantity from the image by a neural network model or image processing, the background feature quantity extraction unit 104 may extract other information from the image. The said other information may include any information such as the shooting location of the image, the shooting time, the position and posture of the camera at the time of shooting, etc. Also, when directly obtaining the feature quantity from the camera, the background feature quantity extraction unit 104 may obtain other information other than the image from the camera. The said other information may include any information such as the shooting location of the image, the shooting time, the position and posture of the camera at the time of shooting, etc.

[0083] The information on the image shooting location may include attributes such as indoor, outdoor, urban area, rural area, mountain, forest road, etc., latitude and longitude, and / or information regarding any other location. The information on the image shooting time may include attributes such as morning, noon, evening, night, etc., timestamp, and / or information regarding any other time. The information on the camera position and orientation may include the direction of the camera lens, the position of the camera viewpoint, the installation position of the camera such as the ceiling or wall, the inclination of the camera, etc. The camera viewpoint or installation position may be expressed by three-dimensional coordinates or the like. The inclination of the camera may be expressed by yaw, pitch, and roll. Further, the camera viewpoint and / or the direction of the lens may be expressed by a vector.

[0084] As described above, the background feature quantity 130 includes at least one of information regarding the background shooting location, the background shooting time, the camera position at the time of background shooting, and the background weather. Further, the learning data selection device 100 includes a background extraction unit 404 that extracts a background image 700 from each of one or more images 110. Extracting the background feature quantity 130 from each of the one or more images 110 includes extracting the background feature quantity 130 from the background image 700 extracted from each of the one or more images 110. Further, the background feature quantity 130 may include at least one of the output result of the intermediate layer of the neural network, information regarding the background shooting location, the background shooting time, the camera position at the time of background shooting, and the background weather.

[0085] FIG. 8 is a diagram showing an example of an equation used when selecting an image for machine learning. The final feature quantity of one or more images 110 is calculated using at least a part of equations 802 to 812. In one aspect, equations 802 to 812 are implemented as a program. In another aspect, equations 802 to 812 may be implemented as hardware. The processing corresponding to equations 802 to 812 may be executed by any one of the object feature quantity extraction unit 102, the background feature quantity extraction unit 104, or the selection unit 106.

[0086] Equation 802 is the final feature quantity Z of the image 110 iCalculate it. That is, Equation 802 calculates the image feature amount of Image 110. Equation 802 concatenates the first calculation result which is the multiplication result of the object importance and the object feature amount 120, and the second calculation result which is the multiplication result of the background importance and the background feature amount 130. More specifically, Equation 802 concatenates the first calculation result indicating the features of the object in a certain image and the second calculation result indicating the features of the background in a certain image. Thereby, Equation 802 can calculate the feature amount of the entire certain image. In Equation 802, "+" indicates concatenation.

[0087] Equation 804 is the first equation for calculating the object feature amount 120. Equation 804 calculates the object feature amount 120 by multiplying the output result of the object feature amount extraction unit 102 by a coefficient. The output data of the object feature amount extraction unit 102 may include a plurality of attributes such as the size of the object, the color of the object, the posture of the object, and the orientation of the object. The learning data selection device 100 can change the weight of a specific attribute in the object feature amount 120 by multiplying the output data of the object feature amount extraction unit 102 by a coefficient. By doing so, the final feature amount is likely to vary based on a specific attribute of the object. As a result, the learning data selection device 100 can select various images as learning data in a specific attribute of the object. In a certain aspect, the object feature amount extraction unit 102 may calculate the object feature amount 120 using Equation 804. In another aspect, the selection unit 106 may calculate the object feature amount 120 using Equation 804.

[0088] Equation 806 is the second equation for calculating the object feature amount 120. Equation 806 outputs the output data of the object feature amount extraction unit 102 as the object feature amount 120 as it is. The learning data selection device 100 can use either Equation 804 or Equation 806 based on the received user input. The user input here may be an input regarding the use or importance of the model. Or, the user input may be an input regarding the weight of a specific attribute of the object. In a certain aspect, the object feature amount extraction unit 102 may calculate the object feature amount 120 using Equation 806. In another aspect, the selection unit 106 may calculate the object feature amount 120 using Equation 806.

[0089] Expression 808 is an expression that converts the feature amounts of a plurality of objects into object feature amounts per image when a plurality of objects are captured in one image 110. In a certain aspect, the dimension of the object feature amount per image may be padded according to the maximum number of objects per image.

[0090] In a certain aspect, when a plurality of objects are captured in one image 110, the object feature amount extraction unit 102 may calculate the object feature amount 120 based on Expression 804 and Expression 808. Alternatively, when a plurality of objects are captured in one image 110, the object feature amount extraction unit 102 may calculate the object feature amount 120 based on Expression 806 and Expression 808. In another aspect, when a plurality of objects are captured in one image 110, the selection unit 106 may calculate the object feature amount 120 based on Expression 804 and Expression 808. Alternatively, when a plurality of objects are captured in one image 110, the selection unit 106 may calculate the object feature amount 120 based on Expression 806 and Expression 808.

[0091] Expression 810 is a first expression for calculating the background feature amount 130. Expression 810 calculates the background feature amount 130 by multiplying the output result of the background feature amount extraction unit 104 by a coefficient. The output data of the background feature amount extraction unit 104 may include a plurality of attributes such as information on the background shooting location, the background shooting time, the camera position at the time of background shooting, and the weather of the background. The learning data selection device 100 can change the weight of a specific attribute in the background feature amount 130 by multiplying the output data of the background feature amount extraction unit 104 by a coefficient. By doing so, the final feature amount is likely to vary based on a specific attribute of the background. As a result, the learning data selection device 100 can select various images as learning data in a specific attribute of the background. In a certain aspect, the background feature amount extraction unit 104 may calculate the background feature amount 130 using Expression 810. In another aspect, the selection unit 106 may calculate the background feature amount 130 using Expression 810.

[0092] Expression 812 is the second expression for calculating the background feature quantity 130. Expression 812 outputs the output data of the background feature quantity extraction unit 104 as the background feature quantity 130 as it is. The selection unit 106 can use either Expression 810 or Expression 812 based on the received user input. Here, the user input may be an input regarding the use or importance of the model. Or, the user input may be an input regarding the specific attribute weights of the object. In a certain aspect, the background feature quantity extraction unit 104 may calculate the background feature quantity 130 using Expression 812. In another aspect, the selection unit 106 may calculate the background feature quantity 130 using Expression 812.

[0093] As described above, calculating the final feature quantity Z of each of the one or more images 110 i includes multiplying the object feature quantity 120 by the object importance to obtain a first calculation result, multiplying the background feature quantity 130 by the background importance to obtain a second calculation result, and concatenating the first calculation result and the second calculation result, as shown in Expression 802. Here, the object importance is W obj and the background importance here is W back

[0094] Also, calculating the final feature quantity Z of each of the one or more images 110 i includes multiplying the object feature quantity 120 by a first coefficient for correcting the values of some attributes of the object before calculating the first calculation result and the second calculation result, as shown in Expression 804. Here, the first coefficient is W obj_param and further, calculating the final feature quantity of each of the one or more images 110 includes multiplying the background feature quantity 130 by a second coefficient for correcting the values of some attributes of the background before calculating the first calculation result and the second calculation result, as shown in Expression 810. Here, the second coefficient is W back_param

[0095] ​​FIG. 9 is a diagram showing an example of the operation procedure of the learning data selection device 100. In a certain aspect, the processor 501 may read a program for performing the processing of FIG. 9 from the secondary storage device 503 into the primary storage device 502 and execute the program. In other aspects, part or all of the processing may also be realized as a combination of circuit elements configured to execute the processing.

[0096] In step S910, the learning data selection device 100 acquires one or more images 110. In step S920, the learning data selection device 100 separates an object image and a background image from each of the one or more images 110. The processing of this step corresponds to the processing described with reference to FIGS. 6 and 7. When the one or more images 110 are directly input to the object feature amount extraction unit 102 and the background feature amount extraction unit 104, the processing of this step may not be executed.

[0097] In step S930, the learning data selection device 100 extracts object feature amounts 120 from each of the one or more images 110. In a certain aspect, the learning data selection device 100 may extract object feature amounts 120 from each of the one or more object images. In other aspects, the learning data selection device 100 may directly extract object feature amounts 120 from each of the one or more images 110.

[0098] In step S940, the learning data selection device 100 extracts background feature amounts 130 from each of the one or more images 110. In a certain aspect, the learning data selection device 100 may extract background feature amounts 130 from each of the one or more background images. In other aspects, the learning data selection device 100 may directly extract background feature amounts 130 from each of the one or more images 110.

[0099] In step S950, the learning data selection device 100 receives an input regarding the use or importance of the model. Based on the input, the learning data selection device 100 can determine which of the formulas shown in FIG. 8 to use. Also, based on the input, the learning data selection device 100 can determine the importance and coefficient values to be applied to each of the object feature amount 120 and the background feature amount 130. In step S960, the learning data selection device 100 calculates the final feature amount of each of the one or more images 110. The final feature amount is the value of Z in Equation 802 i of.

[0100] In step S970, the learning data selection device 100 selects one or more learning data based on the final feature amount of each of the one or more images 110. The learning data selection device 100 selects one or more learning data so that the final feature amounts vary as much as possible. The learning data here is one or more images selected from among the one or more images 110 or their feature amounts.

[0101] In step S980, the learning data selection device 100 outputs the selected one or more learning data. In one aspect, the learning data selection device 100 may output one or more learning data to a directory in the local environment. In another aspect, the learning data selection device 100 may transmit one or more learning data to another device.

[0102] As described above, the learning data selection device 100 can execute the following method by executing a program. The method includes extracting the object feature amount 120 from each of the one or more images 110, extracting the background feature amount 130 from each of the one or more images 110, calculating the final feature amount of each of the one or more images 110 based on the object feature amount 120 and the background feature amount 130, and selecting one or more learning data for machine learning from among the one or more images 110 based on the final feature amount of each of the one or more images 110.

[0103] <D. Application Example> FIG. 10 is a diagram showing an example of the functional configuration of a learning data selection device 1000 for selecting videos for machine learning from one or more videos. Different from the learning data selection device 100, the learning data selection device 1000 selects one or more learning data from among one or more videos. The learning data selection device 1000 can utilize the various formulas in FIG. 8 in calculating various feature amounts. In this case, the final feature amount is a video feature amount rather than an image feature amount. Similar to the image feature amount, the video feature amount can be expressed in any format such as a matrix, an array, a list, JSON, binary data, etc.

[0104] The learning data selection device 1000 includes an input unit 1001, an object extraction unit 1002, a background extraction unit 1004, an object feature amount extraction unit 1006, a background feature amount extraction unit 1008, a selection unit 1010, a storage unit 1012, and an output unit 1014.

[0105] The input unit 1001 receives the input of one or more videos that are candidates for learning data. The input unit 1001 outputs one or more videos to either the object extraction unit 1002 or the object feature amount extraction unit 1006. Similarly, the input unit 1001 outputs one or more videos to either the background extraction unit 1004 or the background feature amount extraction unit 1008.

[0106] The object extraction unit 1002 extracts one or more objects shown in each frame from each of the one or more videos. The object extraction unit 1002 can extract the objects shown in each frame as object images. In a certain aspect, the object extraction unit 1002 may track a certain object that appears over a plurality of frames and group the images of the certain object in multiple sheets. The object extraction unit 1002 outputs the one or more extracted object images to the object feature amount extraction unit 1006.

[0107] The background extraction unit 1004 extracts one background image from each of one or more videos. If the video is captured by a fixed camera, the background does not change. Therefore, the background extraction unit 1004 extracts or generates a background image from one frame in the video. In a certain aspect, the background extraction unit 1004 may extract a background image from a frame in which no object is shown.

[0108] The object feature extraction unit 1006 extracts object features from each of one or more videos. Further, the object feature extraction unit 1006 may extract object features from one or more object images input from the object extraction unit 1002. In a certain aspect, the object feature extraction unit 1006 may track a certain object that appears over a plurality of frames and calculate one object feature from the tracking result. As an example, the object feature extraction unit 1006 may output, as an object feature, an integrated value or an average value of the feature amounts of the object images obtained from each frame. In another aspect, the object feature extraction unit 1006 may calculate one object feature from a group of a plurality of images acquired from the background extraction unit 1004. As an example, the object feature extraction unit 1006 may output, as an object feature, an integrated value or an average value of the feature amounts of the object images within the group. The object features of the video may have the same data structure as the object features 120 of the image.

[0109] The background feature extraction unit 1008 extracts background features from each of one or more videos. Further, the background feature extraction unit 1008 may extract background features from the background image input from the background extraction unit 1004. The background feature extraction unit 1008 outputs the background features to the selection unit 1010 or the storage unit 1012. The background features of the video may have the same data structure as the background features 130 of the image.

[0110] The selection unit 1010 selects training data from one or more videos based on object feature quantities and background feature quantities. At this time, the selection unit 1010 receives an input regarding the use or importance of the model from the user. The selection unit 1010 determines object importance and background importance based on the input regarding the use or importance of the model. Then, the selection unit 1010 multiplies the object importance by the object feature quantity to obtain a first calculation result. Similarly, the selection unit 1010 multiplies the background importance by the background feature quantity to obtain a second calculation result. The selection unit 1010 calculates the final feature quantity of each of the one or more videos by concatenating the first calculation result and the second calculation result. Next, the selection unit 1010 selects one or more training data from the one or more videos based on the final feature quantity of each of the one or more videos. The selection unit 1010 outputs the selected one or more training data to the output unit 1014.

[0111] The storage unit 1012 stores each of the one or more videos in association with object feature quantities and background feature quantities. The selection unit 1010 does not need to immediately select training data from the input one or more videos. The selection unit 1010 may select one or more training data from the one or more videos stored in the storage unit 1012 in the past based on a request from the user.

[0112] The output unit 1014 outputs the selected one or more training data. Assume that the training data selection device 100 is a stand-alone device. In this case, the output unit 1014 can output the one or more training data to an arbitrary directory on the device or an external storage device. Assume that the training data selection device 100 is a web application or a service in a cloud environment. In this case, the output unit 1014 can transmit the one or more training data to the user's terminal.

[0113] In a certain aspect, the training data selection device according to the present embodiment may have the functions of both the training data selection devices 100 and 1000. That is, the training data selection device may have a function of selecting training data from one or more images and a function of selecting training data from one or more videos.

[0114] As described with reference to FIG. 10, when each of the one or more images 110 is one or more videos, the object feature quantity extraction unit 1006 extracts object feature quantities from the objects shown in one or more frames constituting each of the one or more videos. Further, the background feature quantity extraction unit 1008 extracts background feature quantities from any one of the one or more frames constituting each of the one or more videos.

[0115] <E. Summary> As described above, the learning data selection device 100 according to the present embodiment individually extracts the object feature quantity 120 and the background feature quantity 130 from an image. Further, the learning data selection device 100 adjusts the weights of each of the object feature quantity 120 and the background feature quantity 130 using importance to calculate a final feature quantity. The learning data selection device 100 can select appropriate learning data according to the application by selecting learning data so that the final feature quantity varies. As a result, the generalization performance of the model after learning is improved.

[0116] Similarly, the learning data selection device 1000 individually extracts object feature quantities and background feature quantities from a video. Further, the learning data selection device 1000 adjusts the weights of each of the object feature quantity and the background feature quantity using importance to calculate a final feature quantity. The learning data selection device 100 can select appropriate learning data according to the application by selecting learning data so that the final feature quantity varies. As a result, the generalization performance of the model after learning is improved.

[0117] It should be considered that the embodiments disclosed this time are illustrative in all respects and not restrictive. The scope of the present disclosure is shown not by the above description but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims are included. Also, the disclosed content described in the embodiments and each modification example is intended to be implemented alone or in combination as much as possible.

Description of Reference Numerals

[0118] 100,1000 learning data selection device, 102,1006 object feature quantity extraction unit, 104,1008 background feature quantity extraction unit, 106,1010 selection unit, 110,202,204,206,208,210,212,302,304,306,312,314,600A,600B,600C,700A,700B,700C images, 120 object feature quantities, 130 background feature quantities, 300,310 groups, 400,1001 input unit, 402,1002 object extraction unit, 404,1004 background extraction unit, 406,1012 memory unit, 408,1014 output unit, 501 processor, 502 primary storage device, 503 secondary storage device, 504 external device interface, 505 input interface, 506 output interface, 507 communication interface, 600 object images, 700 background images.

Claims

1. An object feature amount extraction unit that extracts object feature amounts from each of one or more images, A background feature amount extraction unit that extracts background feature amounts from each of the one or more images, And a selection unit that selects one or more learning data for machine learning from among the one or more images, The selection unit, Based on the object feature amount and the background feature amount, calculates the final feature amount of each of the one or more images, A learning data selection device that selects the one or more learning data based on the final feature amount of each of the one or more images.

2. Calculating the final feature amount of each of the one or more images includes: Multiplying the object feature amount by an object importance to obtain a first calculation result, Multiplying the background feature amount by a background importance to obtain a second calculation result, And concatenating the first calculation result and the second calculation result. The learning data selection device according to claim 1.

3. Calculating the final feature amount of each of the one or more images includes, before calculating the first calculation result and the second calculation result, Multiplying the object feature amount by a first coefficient that corrects the value of an attribute of a part of the object, And multiplying the background feature amount by a second coefficient that corrects the value of an attribute of a part of the background. The learning data selection device according to claim 2.

4. The object feature amount includes at least one of information regarding the size of the object, the color of the object, the posture of the object, and the orientation of the object. The learning data selection device according to any one of claims 1 to 3.

5. Further comprising an object extraction unit that extracts object images from each of the one or more images, Extracting the object feature amount from each of the one or more images includes extracting the object feature amount from the object images extracted from each of the one or more images. The learning data selection device according to any one of claims 1 to 3.

6. The object feature amount includes at least one of the output result of an intermediate layer of a neural network, the size of the object, the color of the object, the posture of the object, and the orientation of the object. The learning data selection device according to claim 5.

7. The background feature amount includes at least one of information regarding the shooting location of the background, the shooting time of the background, the camera position at the time of background shooting, and the weather of the background. The learning data selection device according to any one of claims 1 to 3.

8. The apparatus further includes a background extraction unit that extracts a background image from each of the one or more images. The learning data selection apparatus according to any one of claims 1 to 3, wherein extracting the background feature amount from each of the one or more images includes extracting the background feature amount from the background image extracted from each of the one or more images.

9. The learning data selection apparatus according to claim 8, wherein the background feature amount includes at least one of an output result of an intermediate layer of a neural network, a background shooting location, a background shooting time, a camera position at the time of background shooting, and information regarding the weather of the background.

10. When each of the one or more images is one or more videos, the object feature amount extraction unit extracts the object feature amount from an object shown in one or more frames constituting each of the one or more videos, the learning data selection apparatus according to any one of claims 1 to 3, wherein the background feature amount extraction unit extracts the background feature amount from one frame constituting each of the one or more videos.

11. extracting an object feature amount from each of one or more images; extracting a background feature amount from each of the one or more images; calculating a final feature amount of each of the one or more images based on the object feature amount and the background feature amount; A learning data selection method, comprising: selecting one or more learning data for machine learning from the one or more images based on the final feature amount of each of the one or more images.

12. extracting an object feature amount from each of one or more images; extracting a background feature amount from each of the one or more images; calculating a final feature amount of each of the one or more images based on the object feature amount and the background feature amount; A learning data selection program that causes a computer to select one or more learning data for machine learning from the one or more images based on the final feature amount of each of the one or more images.

Citation Information

Patent Citations

  • Method, apparatus, and program for sampling learning target frame image of video for ai image learning, and method of same image learning

    JP2020052999A

  • Image dataset generation apparatus, learning apparatus, in-vehicle system, and image dataset generation method

    JP2022112819A