Method for training machine learning model

By using truth annotation and ignoring masks when training machine learning models, the problems of inconsistent annotation and high false positive rates in fuzzy dent recognition are solved, and higher recognition accuracy and training quality are achieved.

CN120088592APending Publication Date: 2025-06-03AIRBUS (SAS)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411727825.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2024-11-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, when training machine learning models to identify image features, especially fuzzy dents, there are problems such as inconsistent labeling, high false positive rates, and contradictions in the optimization process.

Method used

By providing training image data, assigning truth value annotations to each pixel, and generating an ignorance mask, only pixels that do not ignore flags are trained, and the predicted values ​​of pixels with ignorance flags are ignored.

Benefits of technology

This improves the accuracy and consistency of the model when identifying fuzzy image features, reduces the false positive rate, and improves the quality and stability of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088592A_ABST
    Figure CN120088592A_ABST
Patent Text Reader

Abstract

A method of training a machine learning model to identify image features, the method comprising: a. Providing training image data comprising a plurality of pixels; b. Assigning a truth value label to each pixel, where each truth value label is related to a respective one of the pixels, and each truth value label indicates whether the pixel corresponds to the image feature; c. Providing a neglect mask comprising a set of neglect flags, where each neglect flag is associated with a respective one of the pixels and each neglect flag provides an indication that a pixel should be neglected; d. For each pixel, receiving a prediction value from the machine learning model, where each prediction value provides an indication of the probability that the pixel corresponds to the image feature; e. For each pixel without the ignore flag, determining a loss value based on the predicted value and the truth value label of the pixel, and f. Training a machine learning model based on the loss value; and for each pixel with the ignoring flag, ignoring the predicted value of the pixel so that the pixel is not used for training the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods for training machine learning models to identify image features such as dents or other surface defects. The present invention also relates to computer systems and computer software configured to train machine learning models. Background Art

[0002] Object detection is a known method for localizing objects within an image. Modern deep learning algorithms solve this problem by collecting a large number of images in which the desired object is present. A person marks a bounding box around it. These boxes are called ground truth boxes or annotations, and these are the values to which the algorithm converges during the iterative learning process. The person doing the annotation is called an annotator.

[0003] The deep learning algorithm then attempts to learn this through optimization techniques to converge to the exact values given by the annotator. Through processes called backpropagation and gradient descent, the process of converging to the annotation values occurs through multiple iterations. Gradient descent calculates the error between the model prediction and the ground truth and attempts to minimize this error.

[0004] In the case of detecting rigid objects with clear contours, it may be easy for an annotator to know the boundaries of the object. In the case of features (such as dents) within the object, this may be more difficult.

[0005] Due to the ambiguous nature of dents and the frames being randomly fed for annotation, annotators sometimes can annotate or skip ambiguous dents according to their personal preferences. In the case where an ambiguous dent is annotated, the machine learning model attempts to learn it as "yes", and in the case where it is skipped, the machine learning model attempts to learn it as "no". Thus, the optimizer is also in a contradictory and inconsistent situation of sometimes learning and sometimes not learning this pattern. Sometimes, the ambiguous dents are so insignificant that if the machine learning model is forced to learn them, the machine learning model may end up detecting many false positives, making it very sensitive.

[0006] Under normal circumstances, there is a specific class called the background class that the algorithm uses to distinguish and learn such false detections from actual detections. In ambiguous cases, the background class will contradict the dent class and lead to inconsistencies.

[0007] When annotating frames, they may be randomly selected from a video, so annotators usually do not get a time series as a hint for localizing dents and marking the correct bounding boxes. Since this is already a long manual process, providing a time series for reference to the annotator will cause more difficulties for the annotator and does not really solve the problem.

[0008] The annotation process may also require giving the same frame to multiple annotators to have better consistency in the bounding boxes. This method is introduced mainly to reduce errors caused by sloppiness rather than to solve problems caused by the ambiguity of not being able to determine the boundaries. Similar voting techniques like these still retain the ambiguity of the boxes, and consecutive frames will still be inconsistent.

[0009] Since the optimization process has no knowledge of the real world and only focuses on converging to the ground truth boxes provided by the annotators, similar indentations are labeled differently and are in a mess because it is forced to converge to different values for similar or the same patterns, which will cause contradictions. Summary of the Invention

[0010] Aspects of the present invention provide a method for training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data including a plurality of pixels; b. assigning a ground truth annotation to each pixel, wherein each ground truth annotation is related to the corresponding pixel in the pixel, and each ground truth annotation indicates whether the pixel corresponds to an image feature; c. providing an ignore mask including a set of ignore flags, wherein each ignore flag is related to the corresponding pixel in the pixel, and each ignore flag provides an indication that the pixel should be ignored; d. for each pixel, receiving a predicted value from the machine learning model, wherein each predicted value provides an indication of the probability that the pixel corresponds to an image feature; e. for each pixel without an ignore flag, determining a loss value based on the predicted value and the ground truth annotation of the pixel, and f. training the machine learning model based on the loss value; and for each pixel with an ignore flag, ignoring the predicted value of the pixel such that it is not used to train the machine learning model.

[0011] Optionally, c. includes inspecting the object and generating an ignore mask based on the inspection.

[0012] Optionally, c. includes receiving an input from a manual inspection of the object and generating an ignore mask based on the input.

[0013] Optionally, c. includes inspecting the object with a sensor to generate three-dimensional inspection data and generating an ignore mask based on the three-dimensional inspection data.

[0014] Optionally, the training image data includes one or more images of an object.

[0015] Optionally, the method further includes generating the training image data by imaging the object.

[0016] Optionally, imaging the object with light.

[0017] Optionally, the training image data includes a series of images of an object containing the same features as viewed from different perspectives.

[0018] Optionally, the method further includes generating training image data by imaging an object from a series of different perspectives.

[0019] Optionally, b. includes displaying the training image data to a human annotator; and receiving a ground truth mask via input from the human annotator, the ground truth mask providing an indication of regions of the training image data that contain image features.

[0020] Optionally, steps d. to f. are repeated, each repetition including a corresponding training round.

[0021] Optionally, the image features include surface defects.

[0022] Optionally, the image features include surface defects of an aircraft.

[0023] Optionally, the image features include dents.

[0024] Optionally, the loss value is determined by the following algorithm:

[0025] -y k ln p k -(1 - y k ) ln(1 - p k ).

[0026] Where y k is the ground truth annotation of the pixel; p k is the predicted value of the pixel, and the pixels corresponding to the image features have a ground truth annotation y k of 1, and the pixels not corresponding to the image features have a ground truth annotation y k of 0.

[0027] Optionally, after the machine learning model has been trained, the machine learning model is used to segment an image in an inference phase.

[0028] Optionally, c. includes creating an ignore mask based on the ground truth mask, where the ignore mask includes a loop around the ground truth mask that has an inner edge and an outer edge.

[0029] Another aspect of the present invention provides a computer system configured to train a machine learning model by the method of the foregoing aspect.

[0030] Another aspect of the present invention provides a computer software configured to train a machine learning model by the method of the foregoing aspect.

[0031] Another aspect of the present invention provides a method for training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data including a plurality of pixels; b. providing a ground truth mask that provides an indication of regions of the training image data that contain image features; c. creating one or more ignore masks based on the ground truth mask, each ignore mask including a loop at the periphery of the ground truth mask, the loop having an inner edge and an outer edge; d. for each pixel, receiving a predicted value from the machine learning model, wherein each predicted value provides an indication of the probability that the pixel corresponds to an image feature; e. for each pixel that overlaps with the ground truth mask and does not overlap with the ignore mask, determining a loss value based on the predicted value of the pixel and training the machine learning model based on the loss value; and f. for each pixel located between the inner edge and the outer edge of the ignore mask, ignoring the predicted value of the pixel such that it is not used to train the machine learning model.

[0032] Optionally, the periphery of the ground truth mask includes a margin region extending to the edge; and all or part of the ignore mask is inside the edge such that it overlaps with the margin region of the ground truth mask.

[0033] Optionally, the periphery of the ground truth mask includes a margin region extending to the edge of the ground truth mask; and all or part of the ignore mask is outside the edge of the ground truth mask such that it does not overlap with the ground truth mask.

[0034] Optionally, the periphery of the ground truth mask includes a margin region extending to the edge of the ground truth mask; a first part of the ignore mask is inside the edge of the ground truth mask such that it overlaps with the margin region; and a second part of the ignore mask is outside the edge of the ground truth mask such that it does not overlap with the ground truth mask.

[0035] Optionally, the ignore mask is created based on the ground truth mask by dilation of a line following the edge of the ground truth mask.

[0036] Optionally, for each ignore mask, each ignore mask is created based on the ground truth mask by analyzing the ground truth mask by an automatic edge detection process to detect the edge of the ground truth mask; and creating the ignore mask such that it has the same shape as the edge of the ground truth mask.

[0037] Optionally, the periphery of the ground truth mask includes a margin region extending to the edge of the ground truth mask; and the inner edge and the outer edge of the ignore mask each have the same shape as the edge of the ground truth mask.

[0038] Optionally, for each ignore mask, the radial distance between the inner edge and the outer edge of the ignore mask is constant throughout the ignore mask.

[0039] Optionally, the method further includes assigning a ground truth annotation to each pixel, where each ground truth annotation is associated with a corresponding pixel in the pixel, and each ground truth annotation indicates whether the pixel corresponds to an image feature; and for each pixel that does not overlap with the ignore mask, a loss value is determined based on the predicted value and the ground truth annotation of the pixel.

[0040] Optionally, the loss value is determined by the following algorithm:

[0041] -y k ln p k -(1 - y k ) ln(1 - p k ).

[0042] Where y k is the ground truth annotation of the pixel; p k is the predicted value of the pixel, and the pixel corresponding to the image feature has a ground truth annotation y k of 1, and the pixel not corresponding to the image feature has a ground truth annotation y k of 0.

[0043] Optionally, the training image data includes a series of images of an object each containing the same feature observed from different perspectives.

[0044] Optionally, the method further includes generating training image data by imaging the object from a series of different perspectives.

[0045] Optionally, the object is imaged with light.

[0046] Optionally, b. includes displaying the training image data to a human annotator and receiving a ground truth mask via input from the human annotator.

[0047] Optionally, steps d. to f. are repeated, each repetition including a corresponding training round.

[0048] Optionally, the image feature includes a surface defect.

[0049] Optionally, the image feature includes a surface defect of an aircraft.

[0050] Optionally, the image feature includes a dent.

[0051] Optionally, after the machine learning model has been trained, the machine learning model is used to segment an image in an inference phase.

[0052] Another aspect of the present invention provides a computer system configured to train a machine learning model by the method of the foregoing aspect.

[0053] Another aspect of the present invention provides a computer software configured to train a machine learning model by the method of the foregoing aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Embodiments of the present invention will now be described with reference to the accompanying drawings, in which:

[0055] Figure 1 a computer system is shown;

[0056] Figure 2 an aircraft scanned by an imaging device to generate training image data is shown;

[0057] Figure 3 a training phase is shown, in which a machine learning model is trained by a training engine;

[0058] Figure 4 is a flowchart showing the steps of the training phase;

[0059] Figure 5 a time series of training images is shown;

[0060] Figure 6 generation of a ground truth mask is shown;

[0061] Figure 7 a ground truth mask on a frame of a training image is shown;

[0062] Figure 8 a ground truth mask on a frame of a training image, and regions with high prediction values are shown;

[0063] Figure 9 a ground truth mask is shown;

[0064] Figure 10 a line following Figure 9 the outer edge of the ground truth mask is shown;

[0065] Figure 11 an ignore mask based on Figure 9 the ground truth mask is shown;

[0066] Figure 12 an Figure 11 ignore mask and associated regions with high prediction values are shown;

[0067] Figure 13 another ignore mask based on Figure 9 the ground truth mask is shown;

[0068] Figure 14 another ignore mask based on Figure 9 the ground truth mask is shown;

[0069] Figure 15Shows a manual method for generating an ignore mask;

[0070] Figure 16 Shows Figure 9 the ignore mask on the frame of the training image;

[0071] Figure 17 Shows an aircraft scanned by a LIDAR sensor;

[0072] Figure 18 Shows an automatic method for generating an ignore mask;

[0073] Figure 19 Shows a pre - processing set of ignore masks; and

[0074] Figure 20 Shows the ground truth mask on the frame of the training image, as well as the ignore mask and regions with high prediction values. Detailed Description

[0075] Figure 1 Is a schematic diagram showing various elements of a computer system for identifying image features. The computer system includes a memory 1, a machine learning model 2, a training engine 3, and various input / outputs 4.

[0076] The computer system includes computer software configured to train the machine learning model 2 by the methods described below.

[0077] In this example, the machine learning model 2 is used to identify features within an object, such as Figure 2 the surface defects (such as dents or scratches) of the aircraft 10 shown. The aircraft needs to be inspected regularly, and the identification of surface defects can be part of such an inspection. Once a defect is automatically identified by the machine learning model 2, the defect can be repaired or inspected manually.

[0078] Before being used to identify surface defects in the inference phase, the machine learning model 2 must be trained in the training phase. The training phase described below broadly involves: a pre - training process that provides ground truth masks and ignore masks; and a training process that involves receiving a set of prediction values from the machine learning model 2 and training the machine learning model 2 accordingly.

[0079] Figure 3 Shows a method for training the machine learning model 2 to identify image features through deep learning using the above - mentioned data. Training image data 42 is provided and fed into the machine learning model 2. The training engine 3 receives a set 60 of prediction values, a set 50 of ground truth masks, and a set 20 of ignore masks output by the machine learning model 2. The training phase includes a series of training rounds. For each training round, the training engine 3 applies a training update that adjusts the weights of the neural network of the machine learning model 2.

[0080] Figure 4 is a flowchart showing the various steps of the training phase, where blurred dent edges are processed using a surrounding ignore mask.

[0081] In a first step a., training image data 42 is generated by imaging the aircraft 10 from a series of different viewpoints (as Figure 2 shown). The aircraft 10 can be imaged in visible light by an imaging device 40 that scans the entire outer surface of the aircraft 10. As an example, a video of the aircraft can be acquired as the imaging device 40 moves continuously around the aircraft. The imaging device 40 can be held at the end of a robotic arm or carried by an aerial drone. Thus, the video includes a series of frames, each frame including a training image containing a respective part of the aircraft viewed from a different viewpoint. These training images together provide the Figure 3 complete set of training image data 42 shown. Each training image includes a plurality of pixels.

[0082] In the above example, the imaging device 40 senses visible light, but in other examples, the imaging device 40 can be replaced by an infrared imaging device (with an active source of infrared light), a thermal imaging device, or an ultrasonic imaging device.

[0083] In the above example, images are acquired from a single aircraft 10. In other examples, the complete set of training image data 42 can be acquired by imaging multiple aircraft or by imaging multiple components.

[0084] Any visible defect will be present in more than one training image, and each training image may contain more than one defect. Here, the term "visible defect" means a defect that is visible in the image and may (or may not) be visible on the aircraft 10. Figure 5 An example of a time series of twenty cropped training images 101 to 120 is given, each containing the same feature (in this case, a dent) viewed from different observation positions or different viewpoints.

[0085] In training images 104 to 112, the image of the dent is clearly visible. In training images 101 to 102 and 115 to 120, the dent is barely visible - for example, it may be blurred by glare or less visible due to the viewpoint of the imaging device or the angle of the light. In training images 103, 113, and 114, the dent is visible, but the image of the dent is blurred.

[0086] Note that only a single dent is shown per frame in each image 101 to 120, but multiple dents may be visible in each frame.

[0087] In Figure 4In step b., a set 50 of ground truth masks is provided. Each ground truth mask provides an indication of a region of the training image data 42 that contains an image feature (in this case, a dent).

[0088] Each ground truth mask may include a set of ground truth annotations, with one annotation for each pixel within the region. Each ground truth annotation is associated with a corresponding one of the pixels in the pixel and provides a binary indication or flag (value 1) that the pixel corresponds to the image feature. Alternatively, each ground truth mask may consist only of an indication of the edge of the region containing the image feature, rather than a "per-pixel" set of ground truth annotations.

[0089] As Figure 6 shown, a set 50 of ground truth masks can be generated by displaying the training image data 42 to a human annotator via a display device 51 and receiving the ground truth masks as input from the human annotator via an annotation input device 52. Each ground truth mask can be created by marking the boundary of the feature in the training image data 42. For example, the annotator can use a mouse, a drawing tool, or a similar input device to mark the outer edge of the feature. The outer edge can be rectangular or any other shape (including an irregular shape). Alternatively, the annotator can provide each ground truth mask by "painting" over the feature in the training image using a digital painting tool, thereby providing a "per-pixel" set of ground truth annotations that indicate the pixel blocks associated with the feature.

[0090] If the annotator only marks the outer edge of the image feature, then optionally, the interior of the image feature is "filled" with ground truth annotations, so that each pixel within the interior of the ground truth mask is assigned a ground truth annotation of 1.

[0091] Any pixel that does not fall within the ground truth mask is assigned a ground truth annotation of 0. Thus, each pixel in each image is assigned a ground truth annotation y of 0 or 1 k 。

[0092] Figure 7 Frames of twenty training images 101 to 120 of an image with the dents removed are shown. Frames 104 to 112 where the image with the dents annotated with the ground truth masks are clearly visible - in this case, by using a digital painting tool to "paint" over the feature in the training image, thereby indicating the pixel blocks associated with the feature.

[0093] Note that the human annotator is instructed to generate the ground truth masks only when the feature is clearly visible. Thus, no ground truth masks were generated for frames 101 to 103 or frames 113 to 120.

[0094] There may be 1000 images to be labeled, and each frame is randomly presented to different annotators. Therefore, frames with blurred images can be marked with ground truth masks, while other frames (such as frames 103, 113, and 114) can be left unmarked.

[0095] In Figure 4 step d. of, for each pixel of each training image, a predicted value p is generated by the machine learning model 2 k . The machine learning model 2 includes a neural network that identifies and classifies features in the training image data 42. For example, the features can be identified and classified as dents.

[0096] Each predicted value output by the machine learning model 2 provides an indication of the probability that the pixel corresponds to an image feature (e.g., a dent). In the sense that each predicted value can take any decimal value p between 0 and 1 k , each predicted value is non-binary. In Figure 3 , the output of the machine learning model 2 is indicated as the data set 60.

[0097] If the machine learning model 2 is 100% certain that a pixel should be classified as an image feature, the predicted value of that pixel is output as 1; and if the machine learning model 2 is 100% certain that a pixel should not be classified as an image feature, the predicted value of that pixel is output as 0. In most cases, the predicted value will take an intermediate value between 0 and 1. For training images 103 to 114, the machine learning model 2 predicts dents as regions with high predicted values. The edges of the regions with high predicted values (∼1) are indicated by solid lines in Figure 8 — some of these regions are numbered 62, 63.

[0098] As an example, take the simple case of image segmentation with only two classes. Each pixel can be classified as a dent or classified as the background. If the predicted value is high (e.g., greater than 50%), the pixel is classified as a dent; and if the predicted value is low (e.g., less than 50%), it is classified as the background (or "non-dent").

[0099] The regions with high predicted values in frames 104 to 112 all overlap with the ground truth mask 64: for example, region 63. The regions 62 with high predicted values in frames 102, 113, 114 do not overlap with the ground truth mask 64. This is because images 103, 113, 114 contain blurred images of dents that are visible enough to be detected by the machine learning model 2 but not clear enough to be labeled with a ground truth mask.

[0100] In Figure 4 step c. of, a set 20 of ignore masks is generated. Figures 10 to 14 shows based on Figure 9The truth mask 64 shown creates three methods of creating an ignore mask. An ignore mask is generated for each one in the truth mask.

[0101] Figure 9 The truth mask 64 shown includes a core 65 surrounded by outer peripheries 66, 67. The outer peripheries include an outer margin region 66 extending to the outer edge 67 of the truth mask 64. Each truth annotation within the truth mask 64 is set to 1.

[0102] In this example, the truth mask 64 is created by a painting tool. An automatic edge detection process analyzes the truth mask 64 to detect the outer edge 67 and generates Figure 10 the line 68 of pixels shown, which follows the detected outer edge. Then, by Figures 11 to 14 one of the methods shown, a surrounding ignore mask having the same shape as the line 68 is created.

[0103] In all cases, the ignore mask includes a closed loop (i.e., a closed path whose starting point coincides with its ending point) at the outer peripheries 66, 67 of the truth mask 64.

[0104] In Figure 11 this case, an ignore mask 75 is created based on the truth mask 64 by expanding the line 68 both inwards and outwards. The ignore mask 75 has a first (inner) portion 70 inside the outer edge 67 of the truth mask such that it overlaps with the outer margin region 66; and a second (outer) portion 71 outside the outer edge 67 of the truth mask such that it does not overlap with the truth mask 64.

[0105] The radial distance between the inner edge 72 and the outer edge 73 of the ignore mask 75 is constant around the ignore mask 75 and can be selected by design.

[0106] In an alternative embodiment, the radial distance between the inner edge 72 and the outer edge 73 of the ignore mask 75 can vary in a predetermined manner around the ignore mask 75 - for example, the radial distance at the top of the ignore mask is greater than the radial distance at the bottom.

[0107] The ignore mask 75 can include a set of ignore flags, each ignore flag being associated with a corresponding one of the pixels in the pixels between the inner edge 72 and the outer edge 73 of the ignore mask 75. Each ignore flag provides a binary indication that the pixel should be ignored. Each pixel consistent with the ignore mask is given an ignore flag.

[0108] Alternatively, the ignore mask 75 can consist only of the indication of the inner edge 72 and the outer edge 73, rather than a "per - pixel" set of ignore flags. In this case, optionally, the interior of the ignore mask 75 is "filled" with ignore flags so that each pixel in the interior has an associated ignore flag.

[0109] Optionally, a preprocessing step is performed such that for all pixels in the inner part 70 of the ignore mask 75 where both the ground truth annotation and the ignore flag are 1, the ground truth annotation is set to 0.

[0110] Returning to Figure 3 and Figure 4 : Once the ground truth mask and the ignore mask have been generated, the training engine 3 trains the machine learning model 2 in a series of training rounds.

[0111] For each pixel that does not have an ignore flag (i.e., it does not coincide with the ignore mask), the training engine 3 determines a loss value based on the predicted value and the ground truth annotation of that pixel. This generates a loss value for each pixel, except for pixels that have an ignore flag.

[0112] As an example, the loss value for each pixel can be determined by a logistic regression function, such as the function:

[0113] -y k ln p k -(1 - y k ) ln(1 - p k ).

[0114] where y k is the ground truth annotation, and p k is the predicted value.

[0115] Pixels corresponding to image features have a ground truth annotation y k of 1, while pixels not corresponding to image features have a ground truth annotation y k of 0. Thus, y k can be 0 or 1, and p k can be any decimal value between 0 and 1. If the predicted value does not match the ground truth annotation, the loss value will be large, and low otherwise.

[0116] Then, the machine learning model 2 is trained based on the loss value for each pixel. For example, the loss values for each pixel can be summed to determine an average loss value for training the machine learning model 2.

[0117] For each pixel that has an ignore flag, the training engine 3 ignores the predicted value of that pixel such that it does not contribute to the calculation of the average loss value and thus is not used to train the machine learning model 2.

[0118] Figure 12 Shows the outer edge of the region 63 with a high predicted value (∼1) that overlaps with the ignore mask 75. A loss value is determined for each pixel that coincides with the ground truth mask 64 and is also within the inner edge 72 of the ignore mask 75 (thus it does not coincide with the ignore mask).

[0119] Thus, all pixels in the core 65 of the ground truth mask are used to train the machine learning model 2. For each pixel located between the inner edge 72 and the outer edge 73 of the ignore mask 75 (and thus consistent with the ignore mask 75), the predicted value of that pixel is ignored such that it is not used to train the machine learning model 2. Thus, the pixels at the outer boundary 76 of the region 63 are ignored.

[0120] Figure 13 and Figure 14 Shows alternative surrounding ignore masks 80, 81. Like the ignore mask 75, each ignore mask 80, 81 includes a closed loop at the outer periphery 66, 67 of the ground truth mask 64.

[0121] Created by expanding the line 68 outwards instead of inwards Figure 13 the ignore mask 80. In this case, all the ignore masks 80 are located outside the outer edge of the ground truth mask 64 such that it does not overlap with the ground truth mask 64. The inner edge of the ignore mask 80 is adjacent to the line 68.

[0122] Created by expanding the line 68 inwards instead of outwards Figure 14 the ignore mask 81. In this case, all the ignore masks 81 are located inside the outer edge of the ground truth mask 64 such that it overlaps with the outer margin region 66 of the ground truth mask 64. The outer edge of the ignore mask 81 is the line 68.

[0123] Table 1 gives eight scenarios that will be encountered at the pixel level, and the corresponding loss values.

[0124] Table 1

[0125] Scene True value annotation Predicted value Ignore flag Loss value #1 0 ~0 0 Low #2 0 ~1 0 High (false positive) #3 0 ~0 1 Ignore #4 0 ~1 1 Ignore #5 1 ~0 0 High (false negative) #6 1 ~1 0 Low #7 1 ~0 1 Ignore #8 1 ~1 1 Ignore

[0126] For each pixel with a ground truth annotation of zero and no ignore flag (i.e., Scenarios #1 and #2), the loss value is determined based on the predicted value and the ground truth annotation of that pixel, and the machine learning model 2 is trained based on that loss value to learn that pixel as the background. Figure 12 Indicates pixel #1 and pixel #2 corresponding to Scenarios #1 and #2 respectively.

[0127] For each pixel with an ignore flag (i.e., it is consistent with the ignore mask) and a ground truth annotation of zero (i.e., Scenarios #3 and #4), the training engine 3 ignores the predicted value of that pixel such that it is not used to train the machine learning model 2. Figure 12 Indicates pixel #3 corresponding to Scenario #3 and pixel #4 corresponding to Scenario #4.

[0128] For each pixel that has a ground truth annotation of 1 (i.e., it is consistent with the ground truth mask) and does not have an ignore flag (i.e., Scenario #5 and Scenario #6), a loss value is determined based on the predicted value and the ground truth annotation of the pixel, and a machine learning model 2 is trained based on the loss value to learn the pixel as a dent. Figure 12 Indicates pixel #5 and pixel #6 corresponding to Scenario #5 and Scenario #6 respectively. These pixel #5 and pixel #6 each overlap with the core 65 of the ground truth mask and are located inside the inner edge 72 of the ignore mask (so it does not overlap with the ignore mask). Therefore, pixel #5 and pixel #6 do not overlap with the ignore mask 75 and are not ignored.

[0129] Figure 12 Indicates pixel #7 corresponding to Scenario #7 and pixel #8 corresponding to Scenario #8. Note that if the above-mentioned preprocessing steps are performed, Scenario #7 and Scenario #8 will not occur because the preprocessing steps set the ground truth annotation to 0.

[0130] If the above-mentioned preprocessing steps are not performed, the training engine 3 processes pixel #7 and pixel #8 in the same way as pixel #3 and pixel #4. Therefore, for each pixel that has an ignore flag and a ground truth annotation of 1 (i.e., Scenario #7 and Scenario #8), the training engine 3 ignores the predicted value of the pixel such that it is not used to train the machine learning model 2.

[0131] Figure 7 Shows two consecutive image frames 111, 112, both of which contain images of the same dent at approximately the same location, but where the ground truth masks are completely different. For example, such an inconsistency may be caused by ground truth masks created by different people. Alternatively, even if the same annotator generates two ground truth masks, the annotator is not presented with consecutive frames 111, 112 one after the other, and this can lead to inconsistencies.

[0132] Therefore, pixels 65, 66 at the same location at the edge of the dent are inside the ground truth mask in one image 112, but outside the ground truth mask in an adjacent image 111 in the time series.

[0133] Such an inconsistency causes the machine learning model 2 to be forced to classify pixel 65 as a dent and pixel 66 as the background. Such contradictory inputs lead to poor and inconsistent training. The above-mentioned surrounding ignore masks 75, 80, 81 can prevent such ambiguous pixels from being used to train the machine learning model 2, and this improves the quality and consistency of the training.

[0134] In the above example, Figure 9The true value mask 64 includes a continuous core 65 surrounded by outer peripheries 66, 67. Thus, the true value mask 64 has outer peripheries 66, 67 but no inner periphery, and thus only generates first (outer) periphery ignore masks 75, 80, 81 that follow the outer edge 67.

[0135] In other examples, the true value mask may contain holes, and thus it also has an inner periphery with an inner edge. In such a case, a second (inner) periphery ignore mask can also be generated, and such an ignore mask includes a loop at the inner periphery of the true value mask, which loop has an inner edge and an outer edge.

[0136] The second (inner) periphery ignore mask can be generated in the same manner as the first (outer) periphery ignore mask. That is, the second (inner) periphery ignore mask can be created based on the true value mask by the following means: the true value mask is analyzed by an automatic edge detection process to detect the inner edge of the true value mask; and an ignore mask is created such that it has the same shape as the inner edge of the true value mask.

[0137] In such a case, the pixels within the inner edge of the loop will not overlap with the true value mask (since they are located in the holes in the true value mask). Then, such a second (inner) periphery ignore mask can be used in the same manner as the first (outer) ignore mask. In such a case, for each pixel that overlaps with the true value mask and does not overlap with the ignore mask (i.e., does not overlap with the first periphery ignore mask or the second periphery ignore mask), the loss value is determined based on the predicted value of the pixel, and the machine learning model is trained based on the loss value; and for each pixel located between the inner edge and the outer edge of the ignore mask (i.e., it overlaps with the first periphery ignore mask or the second periphery ignore mask), the predicted value of the pixel is ignored such that it is not used to train the machine learning model.

[0138] In Figure 5 the time series, the image of the dent starts to fade, becomes prominent, and then fades again. This is a continuous process, and there is no strict line between the dent and non-dent cases. Due to its robustness, the machine learning model 2 will start to detect the dents in the blurred images 103, 113, 114 - see the region 62 in Figure 8 Since the machine learning model 2 does not have an intermediate category, it considers any pixel to either fall into the dent category or the non-dent category. Since the frames 103, 113, 114 with blurred images are not labeled with a true value mask, the machine learning model 2 may try to "forget" the region 62 as a dent, and this may affect the detection of dents in other images. This inconsistency leads to a contradictory situation for the machine learning algorithm.

[0139] For the training images 101 to 102 and 115 to 120 where the dents are barely visible, it is undesirable to force the machine learning model 2 to learn the areas of the dents as background simply because of low visibility, since these are areas where the dents physically exist.

[0140] Figures 15 to 20 A method of training a machine learning model 2 to recognize image features is shown, which uses a box ignore mask to handle such a problem. Figures 15 to 20 The process uses Figure 3 The difference is that the box-ignoring masks are not based on a set of ground-truth masks50.

[0141] The term "box" ignore mask is used to refer to an ignore mask having an outer edge but no inner edge, as opposed to the "surrounding" ignore mask described above which is a loop having both an inner edge and an outer edge.

[0142] Figure 15 A manual process for generating a set 20 of box ignore masks is shown. Each box ignore mask is generated by receiving input from a manual inspection of the aircraft 10 via an input device 22. That is, a human inspector visually inspects the aircraft 10 and enters the location of any visible defects into the input device. A digital model (DMU) 23 of the aircraft 10 provides a 3D representation of the aircraft 10. The ignore mask generator 24 (which may be a Figure 1 The computer system or part of another computer system) receives input from the input device 22, superimposes it on the DMU 23, and generates a set of box ignore masks 20 accordingly. Each box ignore mask can include the outer boundary of the defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each box ignore mask can include a set of pixels, each of which is associated with a defect, i.e., it is located within the boundary of the defect.

[0143] Figure 17 and Figure 18 An alternative automatic process for generating a set 20 of box ignore masks is shown. In this case, the ignore masks are provided by inspecting the aircraft with a sensor 30 (e.g., a LIDAR sensor) to generate three-dimensional inspection data 31 (e.g., LIDAR data) and generating the ignore masks based on the inspection data 31. The ignore masks are generated by an automatic ignore mask generator 32 (which may be Figure 1 The computer system or part of another computer system) automatically generates the ignore mask based on the inspection data 31 and the DMU 23. Figure 15 In the embodiment of the present invention, each box-ignore mask may include the outer boundary of the defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each box-ignore mask may include a set of pixels, each of which is associated with a defect, i.e., it is located within the boundary of the defect.

[0144] Figure 16 Frames 101 to 120 with box ignore masks created by one of the above methods are shown.

[0145] Figure 16 Each box ignore mask in may include a set of ignore flags, each flag being associated with a respective one of the pixels inside the outer edge of the ignore mask. Each ignore flag provides a binary indication that the pixel should be ignored.

[0146] Alternatively, each box ignore mask may consist only of an indication of the rectangular outer edge, rather than a "per-pixel" set of ignore flags. In this case, optionally, the interior of the box ignore mask may be "filled" with ignore flags such that each pixel in the interior has an associated ignore flag.

[0147] In the preprocessing stage, all box ignore masks that overlap with the ground truth mask are removed. Figure 19 Shows how the preprocessing stage changes Figure 16 the data shown. The box ignore masks that overlap with the ground truth mask 64 shown in have been removed, leaving only eleven box ignore masks. Figure 7 The box ignore masks of are superimposed on the ground truth mask 64 and the predictions 62, 63 of.

[0148] Figure 20 The Figure 19 box ignore masks are superimposed on the ground truth mask 64 and Figure 8 the predictions 62, 63 of.

[0149] The training engine follows the process of Table 1 to use the Figure 20 box ignore masks.

[0150] For each pixel with an ignore flag (i.e., scenarios #3 and #4 where the pixel coincides with the box ignore mask), the training engine 3 ignores the predicted value of the pixel such that it is not used for training the machine learning model 2. All pixels within the box ignore masks in frames 101, 102, and frames 115 to 120 correspond to scenario #3 or scenario #4. Due to the preprocessing stage mentioned above, scenarios #7 and #8 will not occur.

[0151] For each pixel without an ignore flag (i.e., scenarios #1, #2, #5, and #6), a loss value is determined based on the predicted value and the ground truth annotation of the pixel, and the machine learning model 2 is trained based on the loss value. Figure 20 All pixels within the ground truth mask of correspond to scenario #5 or scenario #6.

[0152] Note that there are three regions 62 with high predictive values that do not overlap with the ground truth mask 64. In previous methods of generating surrounding ignore masks 75, 80, 81 based on the ground truth mask 64, the pixels in the core of these regions 62 corresponded to Scene #2 and were used to train the machine learning model 2 as false positives. In Figure 20 cases, all pixels in these regions 62 fall within the block ignore mask, so they correspond to Scene #4 and are ignored.

[0153] In Figure 20 the training process shown, the machine learning model 2 is forced to learn about dents from images 104 to 112, and the block ignore mask causes other images to not be determined. This avoids the problem of poor visibility of dents as described above.

[0154] Figures 11 to 14 The surrounding ignore masks 75, 80, 81 provide a solution to the problem of blurring of the exact boundaries of dents in images 104 to 112; while Figure 20 the larger block ignore mask in

[0155] provides a solution to the problems associated with training the machine learning model 2 based on images 101 to 103 and images 115 to 120, where the dents are blurred or almost invisible. Figure 20 Optionally, these two solutions can be used together in the same training process: that is, in addition to Figures 11 to 14 the block ignore mask, surrounding ignore masks 75, 80, 81 can also be created and applied to frames 104 to 112 with a ground truth mask.

[0156] In the above example, a single prediction value is determined for each pixel by the machine learning model 2. This prediction value provides an indication of the probability that the pixel corresponds to an image feature (such as a dent). In other embodiments of the present invention, the machine learning model 2 can output multiple prediction values for each pixel, each prediction value being associated with a different defect category, such as a dent or a scratch. These multiple prediction values can then be used to classify each pixel as, for example, a dent and / or a scratch.

[0157] In the case of more than two defect categories, the annotator can also have the ability to generate ground truth masks associated with each defect category. Ignore masks for these multiple defect categories can also be generated and used as described above.

[0158] In other embodiments of the present invention, the machine learning model 2 can output a single prediction value for each pixel, and this single prediction value is used to identify the pixel as background, unclear, or defective: for example, 0 - 30% = background class; 30% - 70% = intermediate class; 70% - 100% = dent class.

[0159] After the machine learning model 2 has been trained to identify image features through the above process, it can then be used to segment an image during the inference phase. During the inference phase, each pixel is classified based on the predicted values output by the trained machine learning model 2.

[0160] In the case where the word "or" appears, this is to be interpreted as meaning "and / or", such that the items involved are not necessarily mutually exclusive and may be used in any suitable combination.

[0161] Although the invention has been described above with reference to one or more preferred embodiments, it should be understood that various changes or modifications can be made without departing from the scope of the invention as defined in the appended claims.

Claims

1. A method for training a machine learning model to recognize image features, the method comprising: a. Providing training image data, wherein the training image data comprises a plurality of pixels; b. assigning a true value annotation to each pixel, wherein each true value annotation is associated with a corresponding pixel in the pixels and each true value annotation indicates whether the pixel corresponds to an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag is associated with a corresponding pixel in the pixels, and each ignore flag provides an indication that the pixel should be ignored; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability that the pixel corresponds to an image feature; e. for each pixel without the ignored flag, determining a loss value based on the predicted value and the true value label of the pixel, and training the machine learning model based on the loss value; and f. For each pixel with an ignore flag, ignore the predicted value of the pixel so that it is not used to train the machine learning model.

2. The method according to claim 1, wherein: c. Including inspecting an object, and generating the ignore mask based on the inspecting.

3. The method according to claim 1, wherein: c. comprising providing receiving input from manual inspection of an object, and generating said ignore mask based on said input.

4. The method according to claim 1, wherein: c. comprising inspecting an object with a sensor to generate three-dimensional inspection data, and generating the ignore mask based on the three-dimensional inspection data.

5. The method according to any one of claims 2 to 4, wherein: The training image data includes one or more images of the object. The method of claim 5 , further comprising generating the training image data by imaging the object.

7. The method according to claim 6, wherein: The object is imaged with light.

8. The method according to any one of claims 2 to 7, wherein: The training image data includes a series of images of the object observed from different viewing angles, each containing the same features.

9. The method of claim 8, further comprising generating the training image data by imaging the object from a series of different viewing angles.

10. The method according to claim 9, wherein: The object is imaged with light.

11. A method according to any preceding claim, wherein: The training image data includes a series of images of an object each containing the same features observed from different viewing angles.

12. The method of claim 11, further comprising generating the training image data by imaging the object from a series of different viewing angles.

13. The method according to claim 12, wherein: The object is imaged with light.

14. A method according to any preceding claim, wherein b. comprises displaying the training image data to a human annotator; and receiving a ground truth mask via input from the human annotator, the ground truth mask providing an indication of regions of the training image data containing image features.

15. A method according to any preceding claim, wherein: Repeat d. to f., with each repetition including a corresponding training round.

16. A method according to any preceding claim, wherein: The image features include surface defects.

17. The method according to claim 16, wherein: The image features include surface defects of the aircraft.

18. The method according to claim 16 or 17, wherein: The image features include indentations.

19. A method according to any preceding claim, wherein: The loss value is determined by the following algorithm: -y k ln p k -(1-y k )ln(1-p k ). Among them, y k is the true value label of the pixel; p k is the predicted value of the pixel, and the pixel corresponding to the image feature has a true value label y of 1 k , and pixels that do not correspond to image features have a true value label y of 0 k .

20. A method according to any preceding claim, wherein: After the machine learning model has been trained, the machine learning model is used to segment images in the inference phase.

21. A computer system configured to train a machine learning model by the method of any preceding claim.

22. Computer software configured to train a machine learning model using the method of any one of claims 1 to 20.