Machine learning model generation device
The machine learning model generation device addresses the imbalance in training datasets by applying a data oversampling process that focuses on increasing target image data with feature parts, resulting in high-precision estimation results.
Patent Information
- Application Number
- JP2021072450
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-22
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-04-22
AI Technical Summary
Existing machine learning models for image recognition face challenges in achieving high accuracy due to the imbalance in training datasets, particularly the scarcity of abnormal image data in industrial inspection contexts.
A machine learning model generation device that applies a data oversampling process to increase target image data by excluding image data without feature parts during augmentation, generating and storing valid image data with feature parts, and using these enhanced datasets to train models.
This approach allows for the generation of learned models capable of producing high-precision estimation results by increasing the amount of image data with feature parts, thereby addressing the imbalance in training datasets.
Smart Images

Figure 0007694128000001 
Figure 0007694128000002 
Figure 0007694128000003
Abstract
Description
Technical Field
[0001] The present invention relates to a machine learning model generation device.
Background Art
[0002] In recent years, attempts have been made to perform inspections using appearance images of inspection objects by applying image recognition technology based on machine learning. First, as a learning phase, a trained model is generated by performing machine learning using a previously prepared training dataset. Then, as an estimation phase, new image data of the inspection object is input into the trained model to determine whether the inspection object is normal or abnormal.
[0003] In machine learning, it is known that in order to obtain highly accurate inspection results, it is advisable to prepare a large number of image data for the training dataset used in the learning phase. In particular, preparing a large number of normal image data and abnormal image data respectively can obtain highly accurate inspection results. However, generally, in industrial products, it is not easy to prepare a large number of actual abnormal objects. Therefore, the image data of abnormal objects is less than that of normal objects, and it is not easy to obtain highly accurate inspection results.
[0004] By the way, Patent Document 1 describes performing data augmentation processing (data augmentation, data expansion) on image data. Data augmentation processing includes rotation, translation, scaling, inversion, shear processing, brightness adjustment, noise addition, focus adjustment, etc. of image data.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Here, in order to be used in the learning phase, it is necessary to assign labels such as normal or abnormal labels, and labels indicating the presence or absence of feature parts such as scratches and dirt to the image data subjected to the oversampling process of the image data. However, for example, even if the image data before processing includes feature parts such as scratches and dirt, the processed image data may not include the feature parts after the data oversampling process.
[0007] For example, when the image data is rotated, the image area of the processed image data does not include a part of the image data before processing. When the target feature part exists around the image area, for example, at the corner, in the image data before processing, by rotating the image data, the feature part may move outside the image area in the processed image data. If the processed image data is labeled as including the feature part because the image data before processing includes the feature part, it will lead to a decrease in the estimation accuracy of machine learning.
[0008] Also, as described above, in industrial products, it is not easy to prepare a large number of actual abnormal objects, so there is a desire to increase the image data of abnormal objects. In other words, since there is a sufficient amount of image data of normal objects, there is no need to increase the image data of normal objects by the data oversampling process.
[0009] The present invention has been made in view of such problems, and aims to provide a machine learning model generation device that can generate a learned model capable of obtaining highly accurate estimation results by applying a data oversampling process to increase target image data.
Means for Solving the Problems
[0010] One aspect of the present invention is a machine learning model generation device configured by a computer device including an arithmetic processing unit and a storage device, wherein the storage device A training data set storage unit that stores a training data set including a plurality of first image data including a feature part, which is image data obtained by imaging an object to be inspected, a plurality of second image data not including the feature part, which is image data obtained by imaging the object to be inspected, and labels regarding the presence or absence of the feature part in the first image data and the second image data. Comprising The arithmetic processing unit Excludes image data that does not include the feature part when data augmentation processing is performed on the first image data, generates a plurality of effective image data including the feature part by performing the data augmentation processing, and adds and stores the plurality of effective image data and a label indicating the presence of the feature part as the training data set in the training data set storage unit; a data augmentation processing unit A model generation unit that generates a learned model by machine learning using the training data set Comprising 、 The data augmentation processing unit a residual original image data extraction unit that extracts residual original image data, which is the first image data in a state where the feature part remains within the image area of the processed image data when the data augmentation processing is performed on the first image data; a valid image data generation unit that generates a plurality of the valid image data by performing the data augmentation processing on the extracted residual original image data; a storage processing unit that additionally stores the valid image data and a label indicating that the feature part is included as the training data set in the training data set storage unit; and includes In a machine learning model generation device. Also, another aspect of the present invention is a machine learning model generation device configured by a computer device including an arithmetic processing unit and a storage device, wherein the storage device stores a training data set including a plurality of first image data that are image data obtained by imaging an inspection object and include a feature part, a plurality of second image data that are image data obtained by imaging the inspection object and do not include the feature part, and labels regarding the presence or absence of the feature part in the first image data and the second image data, in a training data set storage unit; and includes wherein the arithmetic processing unit excludes image data that does not include the feature part when performing data augmentation processing on the first image data, generates a plurality of valid image data including the feature part by performing the data augmentation processing, and additionally stores the plurality of valid image data and a label indicating that the feature part is included as the training data set in the training data set storage unit, as a data augmentation processing unit; a model generation unit that generates a learned model by machine learning using the training data set; and includes wherein the data augmentation processing unit a preliminary image data generation unit that generates a plurality of preliminary image data by performing the data augmentation processing on the first image data; a coordinate acquisition unit that acquires the coordinates of the feature part in the image area of the first image data; a valid image data extraction unit that extracts, as the valid image data, image data in which the feature part remains in the image area of the preliminary image data from among the plurality of preliminary image data; a storage processing unit that additionally stores the valid image data and a label indicating that the feature part is included as the training data set in the training data set storage unit; and includes wherein the valid image data extraction unit Based on the coordinates of the feature part in the first image data acquired by the coordinate acquisition unit, calculate the coordinates of the feature part in the preliminary image data. Extract the image data in which the coordinates of the feature part in the preliminary image data are located within the image area in the preliminary image data as the valid image data. It is in a machine learning model generation device comprising the above.
Advantages of the Invention
[0011] According to the above aspect, the data augmentation processing unit of the arithmetic processing device generates a plurality of valid image data including the feature part by performing data augmentation processing on the first image data. On the other hand, the data augmentation processing unit excludes the image data that does not include the feature part when performing data augmentation processing on the first image data. Then, the data augmentation processing unit additionally stores the generated plurality of valid image data and the label indicating that it has the feature part as a training data set.
[0012] Therefore, the training data set includes the first image data including the feature part, the second image data not including the feature part, and the valid image data including the feature part, and further includes a label regarding the presence or absence of the feature part in each image data. Then, a learned model is generated by machine learning using the training data set. Therefore, in the training data set, it is possible to increase the image data including the feature part, which is the target image data, and generate a learned model capable of obtaining high-precision estimation accuracy.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0014] (1. Machine Learning Model) The machine learning model is a model that outputs whether an object to be inspected has a feature part when image data obtained by imaging the object to be inspected is input during the estimation phase of machine learning. The machine learning model is a trained model generated by performing machine learning using a training data set during the learning phase of machine learning. The training data set includes first image data including a feature part, second image data not including a feature part, and labels (teacher labels) regarding the presence or absence of the feature part in the first image data and the second image data.
[0015] Here, the object to be inspected is preferably an industrial product, but can also be a consumer product. Generally, since the defective rate of industrial products is extremely low, it is useful to use industrial products as the object to be inspected. Industrial products are, for example, parts of industrial equipment such as vehicles, production machines, and conveying devices. The image data of the object to be inspected is the appearance inspection image data of the above object to be inspected. For example, the image data of the object to be inspected is image data obtained by imaging the appearance of an automotive part. The feature part is, for example, an abnormal part in an industrial product. The abnormal part is, for example, a scratch, stain, foreign matter adhesion, or abnormal shape part with unevenness. Therefore, the machine learning model becomes a model that outputs whether there is an abnormal part in the automotive part, for example, using the appearance inspection image data of the automotive part.
[0016] (2. Basic Configuration of Machine Learning Model Generation Device 1) The basic configuration of the machine learning model generation device 1 will be described with reference to FIG. 1. The machine learning model generation device 1 is configured by a computer device including a storage device 2 and an arithmetic processing device 3. In this embodiment, the machine learning model generation device 1 further includes an input device 4 and a display device 5.
[0017] The memory device 2 includes a training data set storage unit 11 that stores a training data set used in the learning phase of machine learning, and a learned model storage unit 12 that stores a learned model. The training data set includes at least a plurality of first image data including a feature portion, a plurality of second image data not including the feature portion, and labels regarding the presence or absence of the feature portion in the first image data and the second image data.
[0018] The arithmetic processing unit 3 includes a data augmentation processing unit 21 and a model generation unit 22. The data augmentation processing unit 21 generates a plurality of effective image data including a feature portion by performing data augmentation processing on the first image data that is a part of the training data set. However, the data augmentation processing unit 21 excludes image data that does not include the feature portion when performing data augmentation processing on the first image data. Here, the meaning of excluding image data that does not include the feature portion is either not generating the image data that does not include the feature portion itself, or removing the image data once it is generated that does not include the feature portion.
[0019] Furthermore, the data augmentation processing unit 21 additionally stores the plurality of effective image data and a label indicating that it has a feature portion in the training data set storage unit 11 as a training data set. That is, the additionally stored training data set includes the first image data, the second image data, and the effective image data, and further includes labels regarding the presence or absence of the feature portion in the first image data, the second image data, and the effective image data.
[0020] The input device 4 is a keyboard, a pointer input device, a touch panel, etc. The input device 4 inputs a training data set used in the learning phase, image data used in the estimation phase, and other various types of information. The display device 5 displays image data or displays the result of the presence or absence of the feature portion output in the estimation phase. Of course, the display device 5 may also display other various types of information.
[0021] (3. Machine Learning Model Generation Device of the First Embodiment) The machine learning model generation device 1 according to the first embodiment will be described with reference to FIGS. 2 to 10. Hereinafter, the training data set storage unit 11 and the data augmentation processing unit 21 that constitute the machine learning model generation device 1 will be described in detail. Note that the configurations of the machine learning model generation device 1 other than those described above are as shown in FIG. 1.
[0022] As shown in FIGS. 2 and 3, in the initial state, the training data set storage unit 11 includes first image data 31 including a feature portion 31b, second image data 32 not including the feature portion 31b, and a label 33 regarding the presence or absence of the feature portion 31b in the first image data 31 and the second image data 32. The first image data 31, the second image data 32, and the label 33 are input by the input device 4 shown in FIG. 1.
[0023] Here, the image data is, for example, appearance inspection image data of industrial products. And as shown in FIG. 4, the first image data 31 has an image area 31a composed of a plurality of pixels. The image area 31a is, for example, a rectangle, but can be any shape other than a rectangle.
[0024] The first image data 31 is image data including defects, dirt, attached foreign matters, etc. as the feature portion 31b. That is, the first image data 31 is image data including an abnormal portion as the feature portion 31b. In FIG. 4, an example in which the first image data 31 includes a linear feature portion 31b is given. That is, the first image data 31 includes a feature portion 31b of a plurality of pixels. However, the first image data 31 only needs to include a feature portion 31b of at least one pixel. Note that the feature portion 31b is not limited to a linear shape, and can be any shape such as a dot shape or a shape having an area (region).
[0025] The second image data 32 has an image area (not shown) having the same shape as the first image data 31. However, the second image data 32 is image data not including the feature portion 31b (shown in FIG. 4). That is, the second image data 32 is image data not including an abnormal portion, that is, image data of a normal industrial product.
[0026] As shown in FIG. 5, after the data augmentation process by the data augmentation processing unit 21, the training dataset storage unit 11 includes first image data 31, second image data 32, valid image data 34 including a feature portion 31b, and a label 33 regarding the presence or absence of the feature portion 31b in the first image data 31, the second image data 32, and the valid image data 34.
[0027] Since the valid image data 34 is image data including the feature portion 31b, it is the same as the first image data 31. That is, after the data augmentation process, the number of image data including the feature portion 31b in the training dataset increases.
[0028] As shown in FIG. 2, the data augmentation processing unit 21 generates a plurality of valid image data 34 including a feature portion 31b by performing a data augmentation process on the first image data 31 which is a part of the training dataset. The data augmentation processing unit 21 includes a partitioning unit 41, a coordinate acquisition unit 42, a remaining original image data extraction unit 43, a valid image data generation unit 44, and a storage processing unit 45.
[0029] The partitioning unit 41 will be described with reference to FIGS. 6 to 8. As shown in FIG. 6, the partitioning unit 41 partitions the image area 61 of the image data 60 into a remaining pixel area 62 and a lost pixel area 63. The remaining pixel area 62 is the white area in FIG. 6 and often includes the central portion of the image area 61. The lost pixel area 63 is the black area in FIG. 6 and is often located at the periphery of the image area 61.
[0030] The partitioning unit 41 partitions the pre - processing image data 60 using the image data 60 and 70 before and after the data augmentation process when the data augmentation process is applied to the image data 60. As shown in FIG. 7, there are multiple types of data augmentation processes. The types of data augmentation processes include single - type processes and composite processes. As single - type processes, there are multiple types such as rotation, enlargement, reduction, translation, and shear processing. For example, by rotating the image data, the post - processing image data is generated. As composite processes, there are multiple types such as a process that combines rotation and enlargement, or a process that combines shear processing and enlargement or reduction.
[0031] Furthermore, for each type of data augmentation process, processing conditions for performing the data augmentation process in multiple states are defined. The processing conditions are, for example, in the case of rotation among single - type processes, the maximum and minimum values of rotation, and the rotation angle processed between the maximum and minimum values of rotation. For example, in the case of rotation, when the minimum value of rotation is - 20° (left rotation of 20° assuming right rotation is positive), the maximum value of rotation is + 20°, and the rotation angle between them is defined as every 0.1°, the data augmentation process is performed for 400 rotation - angle states.
[0032] The processing conditions in the case of translation are the maximum value of right - hand translation, the number of moving pixels processed until reaching the maximum value of right - hand translation, etc., and the same applies to left - hand translation, upward translation, and downward translation. For example, when the maximum value of right - hand translation is 20,000 pixels and the number of moving pixels until reaching the maximum value is defined as every 100 pixels, the data augmentation process is performed for 200 right - hand translation states.
[0033] Note that "type" means rotation, enlargement, reduction, translation, shear processing, rotation and enlargement, rotation and reduction, shear processing and enlargement, etc., and "state" means the mode applied in each type of data augmentation process.
[0034] FIG. 8 shows, for each of a plurality of types of data augmentation processing, the pre-processing image data 60, the post-processing image data 70, and the regions 62 and 63 partitioned in the pre-processing image data 60. However, FIG. 8 shows only one state out of a plurality of states for each of the plurality of types of data augmentation processing. In FIG. 8, for example, the line located at the upper right in the image region 61 of the pre-processing image data 60 corresponds to the feature portion 31b shown for reference.
[0035] For example, in the case of rotation as the type of data augmentation processing, in the state of the maximum value of counterclockwise rotation, as shown by the dashed line in the central column of the first row of FIG. 8, the portion of the pre-processing image data 60 is rotated counterclockwise with respect to the image region 71 of the post-processing image data 70. In this case, a part of the region indicated by the dashed line is located within the image region 71 of the post-processing image data 70, but another part of the region indicated by the dashed line is located outside the image region 71 of the post-processing image data 70.
[0036] In the case of enlargement / reduction as the type of data augmentation processing, in the state of the maximum value of enlargement, as shown by the dashed line in the central column of the second row of FIG. 8, the portion of the pre-processing image data 60 is enlarged with the same center with respect to the image region 71 of the post-processing image data 70. In the case of translation as the type of data augmentation processing, in the state of the maximum value of upward translation, as shown by the dashed line in the central column of the third row of FIG. 8, the portion of the pre-processing image data 60 is translated upward with respect to the image region 71 of the post-processing image data 70.
[0037] In the case of shearing as a type of data augmentation processing, in the state of the maximum value of right shearing, as shown by the broken line in the central column of the fourth row of FIG. 8, the part of the pre-processing image data 60 is deformed into a parallelogram shape in which the upper side moves to the right and the lower side moves to the left with respect to the image area 71 of the post-processing image data 70. In the case of a composite process of rotation and enlargement as a type of data augmentation processing, in the state of the maximum value of left rotation and the maximum value of enlargement, as shown by the broken line in the central column of the fifth row of FIG. 8, the part of the pre-processing image data 60 is enlarged with the same center while rotating left with respect to the image area 71 of the post-processing image data 70.
[0038] And as shown in the right column of FIG. 8, the image area 61 of the pre-processing image data 60 is partitioned into a remaining pixel area 62 and a lost pixel area 63. The remaining pixel area 62 is the white area in the right column of FIG. 8 and is the area located within the image area 71 of the post-processing image data 70 among the areas indicated by the broken lines in the central column of FIG. 8. That is, the remaining pixel area 62 is the area where the pixels of the pre-processing image data 60 remain within the image area 71 of the post-processing image data 70.
[0039] On the other hand, the lost pixel area 63 is the black area in the right column of FIG. 8 and is the area located outside the image area 71 of the post-processing image data 70 among the areas indicated by the broken lines in the central column of FIG. 8. That is, the lost pixel area 63 is the area where the pixels of the pre-processing image data 60 are lost outside the image area 71 of the post-processing image data 70.
[0040] In the case of left rotation, the lost pixel region 63 is located at the four corners in the image region 61 of the pre - processed image data 60. In the case of enlargement, the lost pixel region 63 is located over the entire periphery in the image region 61 of the pre - processed image data 60. In the case of upward translation, the lost pixel region 63 is located along the upper side in the image region 61 of the pre - processed image data 60. In the case of right - shearing processing, the lost pixel region 63 is located at the upper - right corner and the lower - left corner in the image region 61 of the pre - processed image data 60. In the case of the combined processing of left rotation and enlargement, the lost pixel region 63 is located in most of the entire periphery in the image region 61 of the pre - processed image data 60 and has an inclined inner - periphery shape.
[0041] As shown in FIG. 8, in each of the plurality of types of data up - sampling processing, many of the corners in the pre - processed image data 60 belong to the lost pixel region 63. Therefore, when the feature portion 31b is located at the corner of the pre - processed image data 60, the feature portion 31b is likely to be located in the lost pixel region 63. On the contrary, in each of the plurality of types of data up - sampling processing, many of the central portions in the pre - processed image data 60 belong to the remaining pixel region 62. Therefore, when the feature portion 31b is located at the central portion of the pre - processed image data 60, the feature portion 31b is likely to be located in the remaining pixel region 62.
[0042] And, as shown in the left column of FIG. 9, in each of the plurality of types of data up - sampling processing, the remaining pixel region 62 and the lost pixel region 63 partitioned when processing each of the plurality of states in the entire range of the processing operation are shown. That is, in FIG. 8, one state in each type is shown, while in the left column of FIG. 9, the image data 60 is synthesized for a plurality of states in each type.
[0043] For example, in the case of rotation, for each of a plurality of states within the entire angular range of rotation (including left and right rotations), i.e., for each of the plurality of states in the entire angular range (e.g., at 0.1° angular intervals in the range of -20 to +20°), the area where the pixels of the pre - processed image data 60 remain within the image area 71 of the post - processed image data 70 during each rotation of the plurality of states in the entire angular range is defined as the remaining pixel area 62.
[0044] On the other hand, in at least one of the plurality of states in rotation, the area where the pixels of the pre - processed image data 60 are lost outside the image area 71 of the post - processed image data 70 is defined as the lost pixel area 63. Therefore, as described in the first paragraph from the top in the left column of FIG. 9, in the case of rotation, the lost pixel area 63 is located along the entire periphery of the image area 61 and is mainly located at the four corners. The same applies to other data augmentation processes. In the left column of FIG. 9, the same applies to other types.
[0045] And as shown in the right column of FIG. 9, in this embodiment, the partitioning unit 41 comprehensively processes all types of data augmentation processes and partitions the remaining pixel area 62 and the lost pixel area 63. That is, in all types of data augmentation processes such as rotation and enlargement, the partitioning unit 41 defines the area where the pixels of the pre - processed image data 60 remain within the image area 71 of the post - processed image data 70 as the remaining pixel area 62. Also, in at least one type of data augmentation process such as rotation and enlargement, the partitioning unit 41 defines the area where the pixels of the pre - processed image data 60 are lost outside the image area 71 of the post - processed image data 70 as the lost pixel area 63. In this way, the partitioning unit 41 partitions the image area 61 of the pre - processed image data 60 into a common remaining pixel area 62 and a lost pixel area 63 for a plurality of types of data augmentation processes.
[0046] In the coordinate acquisition unit 42 in the data expansion processing unit 21 shown in FIG. 2, the coordinates of the feature part 31b in the image area 31a (shown in FIG. 4) of the first image data 31 are acquired. The coordinates of the feature part 31b can also be automatically acquired from the first image data 31 by applying image processing technology. Further, the coordinates of the feature part 31b can also be acquired by a human input operation. For example, a person can acquire the coordinates of the feature part 31b by selecting or describing the position of the feature part 31b with respect to the first image data 31 displayed on the display device 5.
[0047] In the remaining original image data extraction unit 43 in the data expansion processing unit 21 shown in FIG. 2, using the information on the remaining pixel area 62 and the lost pixel area 63 partitioned by the partitioning unit 41 and the information on the coordinates of the feature part 31b acquired by the coordinate acquisition unit 42, the remaining original image data is extracted from among the plurality of first image data 31. The remaining original image data is the first image data 31 that, when the data expansion processing is performed on the first image data 31, the feature part 31b remains within the image area 71 of the processed image data 70 (shown in FIG. 8).
[0048] The processing of the remaining original image data extraction unit 43 will be described with reference to FIG. 10. In FIG. 10, in the image area 31a of the first image data 31, the inside of the broken line is the remaining pixel area 62, and the outside of the broken line is the lost pixel area 63.
[0049] The first image data 31 shown in FIG. 10(A) has a linear feature part 31b. All of the feature part 31b is included in the remaining pixel area 62. That is, in the first image data 31, the lost pixel area 63 does not include the feature part 31b. When the data expansion processing is performed on the first image data 31, the feature part 31b will be included in the image area 71 of the processed image data 70 (shown in FIG. 8). Therefore, the remaining original image data extraction unit 43 extracts the first image data 31 shown in FIG. 10(A) as the remaining original image data.
[0050] In particular, in the first image data 31 shown in FIG. 10(A), when all data interpolation processing is performed on the first image data 31, the feature portion 31b is included in the image area 71 of all the processed image data 70. In particular, all of the linear feature portions 31b are included in the image area 71 of all the processed image data 70. That is, all of the plurality of pixels constituting the feature portion 31b are included in the image area 71 of the processed image data 70.
[0051] Next, the first image data 31 shown in FIG. 10(B) has a linear feature portion 31b. However, a part of the feature portion 31b, that is, a part of the plurality of pixels constituting the feature portion 31b, is included in the remaining pixel area 62. On the other hand, the remaining part of the feature portion 31b, that is, the remaining part of the plurality of pixels constituting the feature portion 31b, is included in the lost pixel area 63. When data interpolation processing is performed on the first image data 31, at least a part of the feature portion 31b is included in the image area 71 of the processed image data 70. Therefore, the remaining original image data extraction unit 43 extracts the first image data 31 shown in FIG. 10(B) as the remaining original image data.
[0052] In the first image data 31 shown in FIG. 10(B), when a part of the data interpolation processing is performed, all of the linear feature portions 31b may be included in the image area 71 of the processed image data 70. However, when another part of the data interpolation processing is performed, only a part of the linear feature portion 31b is included in the image area 71 of the processed image data 70. That is, only a part of the plurality of pixels constituting the feature portion 31b is included in the image area 71 of the processed image data 70. However, even a part of the feature portion 31b is included in the processed image data 70. Therefore, the first image data 31 is extracted as the remaining original image data.
[0053] Next, the first image data 31 shown in FIG. 10(C) has a linear feature portion 31b. However, all of the feature portion 31b is included in the lost pixel region 63. That is, in the first image data 31, the remaining pixel region 62 does not include the feature portion 31b. All of the plurality of pixels constituting the feature portion 31b are not included in the image region 71 of the processed image data 70. When data expansion processing is performed on the first image data 31, there may be a case where the feature portion 31b is not included at all in the image region 71 of the processed image data 70. Therefore, the remaining original image data extraction unit 43 does not extract the first image data 31 as the remaining original image data and excludes it.
[0054] Here, depending on the type of data expansion processing, the first image data 31 shown in FIG. 10(C) may include at least a part of the feature portion 31b in the image region 71 of the processed image data 70. However, since there is a case where the feature portion 31b is not included at all in the image region 71 of the processed image data 70 in a part of the data expansion processing, the first image data 31 is excluded.
[0055] The valid image data generation unit 44 of the data expansion processing unit 21 shown in FIG. 2 generates a plurality of valid image data 34 by performing data expansion processing on the extracted remaining original image data. The valid image data 34 is image data after data expansion processing and includes at least a part of the feature portion 31b in the image region.
[0056] Here, as described above, the remaining original image data is image data that includes at least a part of the feature portion 31b in the image region 71 of the processed image data 70 regardless of which of the plurality of data expansion processes is performed on the remaining original image data. Therefore, by performing data expansion processing using the remaining original image data as described above by the valid image data generation unit 44, the generated valid image data 34 will necessarily include at least a part of the feature portion 31b.
[0057] The storage processing unit 45 of the data augmentation processing unit 21 shown in FIG. 2 additionally stores the valid image data 34 and the label 33 indicating that it has the feature unit 31b in the training data set storage unit 11 as a training data set. The additionally stored training data set is as shown in FIG. 5. Then, the additionally stored training data set is used for generating a learned model by the model generation unit 22.
[0058] (4. Effects of the machine learning model generation apparatus 1 according to the first embodiment) According to this embodiment, the data augmentation processing unit 21 of the arithmetic processing unit 3 generates a plurality of valid image data 34 including the feature unit 31b by performing data augmentation processing on the first image data 31. On the other hand, the data augmentation processing unit 21 excludes the image data that does not include the feature unit 31b when performing data augmentation processing on the first image data 31. Then, the data augmentation processing unit 21 additionally stores the generated plurality of valid image data 34 and the label 33 indicating that it has the feature unit 31b as a training data set.
[0059] Therefore, as shown in FIG. 5, the training data set includes the first image data 31 including the feature unit 31b, the second image data 32 not including the feature unit 31b, and the valid image data 34 including the feature unit 31b, and further includes the label 33 regarding the presence or absence of the feature unit 31b in each of the image data 31, 32, 34. That is, the image data including the feature unit 31b becomes the first image data 31 and the valid image data 34, and the image data not including the feature unit 31b becomes the second image data 32.
[0060] In this way, through data augmentation processing, it is possible to increase the data specifically for the image data including the feature portion 31b. Particularly when performing appearance inspection of industrial products, if abnormal portions are regarded as feature portions, generally, the number of the first image data 31 including the feature portion 31b is extremely small compared to the number of the second image data 32 not including the feature portion 31b. However, since effective image data 34 including the feature portion 31b can be generated through data augmentation processing, the total ratio of the first image data 31 and the effective image data 34 including the feature portion 31b can be made sufficiently larger than before the processing.
[0061] Then, the model generation unit 22 generates a learned model by machine learning using the training data set. As described above, in the training data set, since it is possible to increase the image data 31, 34 including the feature portion 31b which is the target image data, it is possible to generate a learned model capable of obtaining high-precision estimation accuracy.
[0062] Also, by adjusting the number of data augmentation processes and the number of the first image data 31 used for the data augmentation process, the number of the effective image data 34 can be freely and easily set.
[0063] In this embodiment, the data augmentation processing unit 21 extracts residual original image data from among a plurality of the first image data 31 using the partitioned residual pixel region 62 and lost pixel region 63, and then generates the effective image data 34 by performing data augmentation processing on the residual original image data. That is, since the residual original image data is extracted in advance before performing the data augmentation processing, there is no need to perform any processing after the data augmentation.
[0064] However, before the data augmentation process, it is necessary to extract the original remaining image data from among the plurality of first image data 31. Here, when comparing the number of image data before the data augmentation process with the number of image data after the data augmentation process, naturally, the former is less. Therefore, by performing the process of extracting the original remaining image data before the data augmentation process, the computational processing load can be reduced, and as a result, the processing time can be shortened.
[0065] Then, by partitioning into a remaining pixel region 62 and a lost pixel region 63 by the partitioning unit 41, it is possible to extract the original remaining image data before the data augmentation process. Further, in this embodiment, the partitioning unit 41 partitions the image region 61 of the image data 60 before processing into a common remaining pixel region 62 and a lost pixel region 63 for a plurality of types of data augmentation processes. Therefore, the process by the original remaining image data extraction unit 43 is performed using one type of remaining pixel region 62 and lost pixel region 63. Further, the process by the effective image data generation unit 44 can also be performed without distinguishing the types of data augmentation processes. Therefore, the process becomes easy, and the man-hours of the operator required for the process and the processing time by the arithmetic processing device 3 can be shortened.
[0066] (5. Machine Learning Model Generation Device of the Second Embodiment) The machine learning model generation device 1 of the second embodiment will be described with reference to FIGS. 9 and 11. The machine learning model generation device 1 of the second embodiment is different in part of the process of the data augmentation processing unit 21 from the first embodiment. In this embodiment, the partitioning unit 41 partitions the image region 61 of the image data 60 before processing into separate remaining pixel regions 62 and lost pixel regions 63 for a plurality of types of data augmentation processes. For example, as shown in the left column of FIG. 9, for each of rotation, enlargement, translation, shear processing, rotation and enlargement, and shear processing and enlargement, a remaining pixel region 62 and a lost pixel region 63 are partitioned. That is, in the first embodiment, a common partitioning process was performed for these plurality of types of data augmentation processes, but in this embodiment, no common partitioning process is performed.
[0067] In this case, the remaining original image data extraction unit 43 performs extraction processing of the remaining original image data for each type of data augmentation processing partitioned by the partitioning unit 41. For example, using the remaining pixel region 62 and the lost pixel region partitioned in the case of rotation as the data augmentation processing, the remaining original image data for rotation processing is extracted. Also, using the remaining pixel region 62 and the lost pixel region partitioned in the case of enlargement or reduction as the data augmentation processing, the remaining original image data for enlargement or reduction processing is extracted. The same applies to others.
[0068] And the effective image data generation unit 44 also generates the effective image data 34 for each type of data augmentation processing. For example, the effective image data generation unit 44 generates the effective image data 34 by performing rotation as the data augmentation processing on the remaining original image data for rotation processing. Also, the effective image data generation unit 44 generates the effective image data 34 by performing enlargement or reduction as the data augmentation processing on the remaining original image data for enlargement or reduction processing.
[0069] According to this embodiment, the first image data 31 can be effectively utilized, and a large number of effective image data 34 can be generated. However, the processing of the partitioning unit 41, the processing of the remaining original image data extraction unit 43, and the processing of the effective image data generation unit 44 need to be performed for each type of data augmentation processing.
[0070] (6. Machine Learning Model Generation Device of the Third Embodiment) The machine learning model generation device 1 of the third embodiment will be described with reference to FIGS. 12 and 13. The machine learning model generation device 1 of this embodiment is different from the first embodiment in that the data augmentation processing unit 21 is different. The data augmentation processing unit 21 in this embodiment includes a preliminary image data generation unit 81, a coordinate acquisition unit 82, an effective image data extraction unit 83, and a storage processing unit 84 as shown in FIG. 12. Here, since the coordinate acquisition unit 82 and the storage processing unit 84 are substantially the same as the coordinate acquisition unit 42 and the storage processing unit 45 of the first embodiment, the description thereof is omitted.
[0071] The preliminary image data generation unit 81 generates a plurality of preliminary image data 90 by performing data augmentation processing on the first image data 31. Here, the data augmentation processing is various types of processing shown in FIG. 7. The preliminary image data 90 is image data obtained by performing data augmentation processing, for example, left rotation and enlargement processing, on the first image data 31, as shown in FIGS. 13(A), (B), and (C).
[0072] The first image data 31 shown in FIG. 13(A) has a feature portion 31b near the center of the image region 31a. When left rotation and enlargement processing are performed on the first image data 31, it becomes as shown by the broken line in the right column of FIG. 13(A). At this time, the image region 91 of the preliminary image data 90 includes all of the feature portion 31b.
[0073] The first image data 31 shown in FIG. 13(B) has a feature portion 31b near the center of the upper side of the image region 31a. When left rotation and enlargement processing are performed on the first image data 31, it becomes as shown by the broken line in the right column of FIG. 13(B). At this time, only a part of the feature portion 31b is included in the image region 91 of the preliminary image data 90, and the other part is not included.
[0074] The first image data 31 shown in FIG. 13(C) has a feature portion 31b near the upper right corner of the image region 31a. When left rotation and enlargement processing are performed on the first image data 31, it becomes as shown by the broken line in the right column of FIG. 13(C). At this time, the image region 91 of the preliminary image data 90 is in a state where all of the feature portion 31b is not included.
[0075] Here, the preliminary image data generation unit 81 performs a plurality of data augmentation processes on, for example, all of the first image data 31. Then, among the generated preliminary image data 90, image data including the feature portion 31b and image data not including the feature portion 31b are mixed.
[0076] The valid image data extraction unit 83 extracts, as valid image data 34, the image data in which the feature part 31b remains in the image area 91 of the preliminary image data 90 from among the plurality of preliminary image data 90. Specifically, the valid image data extraction unit 83 first calculates the coordinates of the feature part 31b in the preliminary image data 90 based on the coordinates of the feature part 31b in the first image data 31 acquired by the coordinate acquisition unit 82. For example, the coordinates of the feature part 31b in the preliminary image data 90 are calculated by performing coordinate conversion processing.
[0077] Subsequently, the valid image data extraction unit 83 extracts, as valid image data 34, the image data in which the coordinates of the feature part 31b in the preliminary image data 90 are located within the image area 91 in the preliminary image data 90. That is, the preliminary image data 90 shown in FIGS. 13(A) and (B) are extracted as valid image data 34. On the other hand, the preliminary image data 90 shown in FIG. 13(C) is excluded from the valid image data 34 and is not extracted as the valid image data 34.
[0078] Then, the storage processing unit 84 additionally stores, in the training data set storage unit 11 as a training data set, the valid image data 34 extracted by the valid image data extraction unit 83 and the label 33 indicating that it has the feature part 31b.
[0079] Also in this embodiment, similar to the first embodiment, the image data including the feature part 31b can be increased. Therefore, a learned model with high estimation accuracy can be generated.
Explanation of Reference Numerals
[0080] 1 Machine learning model generation device 2 Storage device 3 Arithmetic processing device 11 Training data set storage unit 21 Data augmentation processing unit 22 Model generation unit 31 First image data 31a Image area 31b Feature part 32 Second image data 33 Label 34 Valid image data
Claims
1. A machine learning model generation device configured by a computer device including an arithmetic processing unit and a storage device, The storage device includes: A training data set storage unit that stores a training data set including a plurality of first image data including a feature part, which is image data obtained by imaging an inspection object, a plurality of second image data obtained by imaging the inspection object and not including the feature part, and labels regarding the presence or absence of the feature part in the first image data and the second image data, and includes, The arithmetic processing unit includes: An image data excluding unit that excludes image data that does not include the feature part when data augmentation processing is performed on the first image data, and generates a plurality of effective image data including the feature part by performing the data augmentation processing, and additionally stores the plurality of effective image data and a label indicating the presence of the feature part as the training data set in the training data set storage unit; A model generation unit that generates a learned model by machine learning using the training data set; and includes, The data augmentation processing unit includes: A remaining original image data extraction unit that extracts remaining original image data, which is the first image data in a state where the feature part remains within the image area of the image data after the data augmentation processing is performed on the first image data; An effective image data generation unit that generates a plurality of the effective image data by performing the data augmentation processing on the extracted remaining original image data; A storage processing unit that additionally stores the effective image data and a label indicating the presence of the feature part as the training data set in the training data set storage unit; A machine learning model generation device comprising.
2. The data augmentation processing unit further includes: A coordinate acquisition unit that acquires coordinates of the feature part in the image area of the first image data, When the data augmentation process is applied to the image data, the image area of the pre - processing image data is divided into a remaining pixel area where the pixels of the pre - processing image data remain within the image area of the post - processing image data, and a lost pixel area where the pixels of the pre - processing image data are lost outside the image area of the post - processing image data. A partitioning unit for partitioning, comprising The remaining original image data extraction unit extracts the remaining original image data in which the feature part exists in the remaining pixel area from the first image data by using the coordinates acquired by the coordinate acquisition unit. The machine learning model generation device according to claim 1.
3. The partitioning unit partitions the image area of the pre - processing image data into the common remaining pixel area and the lost pixel area for a plurality of types of the data augmentation processes. The machine learning model generation device according to claim 2.
4. When each of a plurality of types of the data augmentation processes is applied to the image data, the partitioning unit divides the image area of the pre - processing image data into a remaining pixel area where the pixels of the pre - processing image data remain within the image area of the post - processing image data in all types of the data augmentation processes, and a lost pixel area where the pixels of the pre - processing image data are lost outside the image area of the post - processing image data in at least one type of the data augmentation processes. The machine learning model generation device according to claim 3.
5. The partitioning unit partitions the image area of the pre - processing image data into separate remaining pixel areas and lost pixel areas for a plurality of types of the data augmentation processes. The machine learning model generation device according to claim 2. A machine learning model generation device constituted by a computer device including an arithmetic processing device and a storage device, The storage device Image data obtained by imaging an object to be inspected, including a plurality of first image data including feature portions, a plurality of second image data obtained by imaging the object to be inspected and not including the feature portions, and a training data set storage unit that stores a training data set including labels regarding the presence or absence of the feature portions in the first image data and the second image data. Comprising The arithmetic processing unit Excludes image data that does not include the feature portion when data augmentation processing is performed on the first image data, generates a plurality of valid image data including the feature portion by performing the data augmentation processing, and adds and stores the plurality of valid image data and a label indicating the presence of the feature portion as the training data set in the training data set storage unit. A data augmentation processing unit A model generation unit that generates a learned model by machine learning using the training data set Comprising The data augmentation processing unit A preliminary image data generation unit that generates a plurality of preliminary image data by performing the data augmentation processing on the first image data A coordinate acquisition unit that acquires the coordinates of the feature portion in the image region of the first image data An effective image data extraction unit that extracts, as the effective image data, image data in which the feature portion remains in the image region of the preliminary image data from among the plurality of preliminary image data A storage processing unit that adds and stores the effective image data and a label indicating the presence of the feature portion as the training data set in the training data set storage unit Comprising The effective image data extraction unit Calculates the coordinates of the feature portion in the preliminary image data based on the coordinates of the feature portion in the first image data acquired by the coordinate acquisition unit Extracts, as the effective image data, image data in which the coordinates of the feature portion in the preliminary image data are located within the image region of the preliminary image data A machine learning model generation device comprising
7. The first image data includes the feature portions of a plurality of pixels, The effective image data is image data including the feature portion of at least one pixel among the feature portions of the plurality of pixels by subjecting the first image data to the data augmentation process. The machine learning model generation device according to any one of claims 1 to 6.
8. The data augmentation process includes at least one type of rotation, enlargement, translation, and shear process of image data. The machine learning model generation device according to any one of claims 1 to 7.
9. The first image data and the second image data are appearance inspection image data of industrial products, The feature portion is an abnormal portion. The machine learning model generation device according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data generation device, data generation method, and data generation program
JP2019114116A
Fusion splicing system, fusion splicing machine, and optical fiber category discrimination method
JP2020020997A
Machine learning device, machine learning method, and program
JP2021036969A