Training machine learning models with surrounding ignoring masks

By introducing an ignore mask during the training process and adjusting the loss function to ignore unstable pixels, the problems of labeling inconsistency and large errors in machine learning models when identifying fuzzy image features are solved, achieving higher recognition accuracy and stability.

CN120707913APending Publication Date: 2025-09-26AIRBUS (SAS)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357774.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-03-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When training machine learning models to identify fuzzy image features, especially dents, existing technologies suffer from labeling inconsistencies and large errors, as well as unstable model learning. This leads to frequent false positive detections, and the optimization process lacks real-world knowledge guidance.

Method used

The ignore mask training method is adopted. By providing training image data, ground truth annotations and ignore masks, the loss function is adjusted to ignore unstable pixels, and the machine learning model is trained to generate ignore masks using loops around the ground truth mask to reduce the impact of inconsistency.

Benefits of technology

It improves the accuracy and stability of machine learning models in identifying blurred image features, reduces false positive detections, and improves the robustness and consistency of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707913A_ABST
    Figure CN120707913A_ABST
Patent Text Reader

Abstract

A method of training a machine learning model to identify image features, the method comprising: a. Providing training image data comprising a plurality of pixels; b. Assigning a truth value label to each pixel, where each truth value label is related to a respective one of the pixels, and each truth value label indicates whether the pixel corresponds to the image feature; c. Providing a neglect mask comprising a set of neglect flags, where each neglect flag is associated with a respective one of the pixels and each neglect flag provides an indication that a pixel should be neglected; d. For each pixel, receiving a prediction value from the machine learning model, where each prediction value provides an indication of the probability that the pixel corresponds to the image feature; e. For each pixel without the ignore flag, determining a loss value based on the predicted value and the truth value label of the pixel, and f. Training a machine learning model based on the loss value; and for each pixel with the ignoring flag, ignoring the predicted value of the pixel so that the pixel is not used for training the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods for training machine learning models to recognize image features, such as dents or other surface defects. The present invention also relates to computer systems and computer software configured to train machine learning models, as well as computer systems and computer-implemented methods for recognizing image features. Background Art

[0002] Object detection is a well-known method for locating objects within an image. Modern deep learning algorithms address this problem by collecting a large number of images with the desired object in them. A human then labels bounding boxes around the object. These boxes are called ground truth boxes or annotations, and they are the values ​​that the algorithm converges to during the iterative learning process. The person who creates these annotations is called an annotator.

[0003] The deep learning algorithm then attempts to learn from this using optimization techniques to converge to the exact value given by the annotator. This convergence to the annotated value occurs over multiple iterations through a process called backpropagation and gradient descent. Gradient descent calculates the error between the model's prediction and the true value and attempts to minimize this error.

[0004] In the case of detecting rigid objects with clear outlines, it may be easy for the annotator to know the boundaries of the object. In the case of features within the object (such as dents), this may be more difficult.

[0005] Due to the fuzzy nature of dents and the randomness of the frames fed for annotation, annotators can sometimes annotate or skip fuzzy dents based on their personal preferences. When a fuzzy dent is annotated, the machine learning model attempts to learn it as a "yes," and when it is skipped, it attempts to learn it as a "no." This creates a conflicting and inconsistent situation for the optimizer, where it sometimes learns and sometimes doesn't. Sometimes, fuzzy dents are so subtle that if the machine learning model is forced to learn them, it might end up detecting many false positives, making it very sensitive.

[0006] In the usual case, there is a special class called background class that the algorithm uses to distinguish and learn such false detections from real detections. In the case of blur, the background class will contradict the dent class and lead to inconsistency.

[0007] When annotating frames, they may be randomly selected from the video, so the annotator usually does not get the time series as a hint to locate the indentation and label the correct bounding box. Since this is already a tedious manual process, providing the annotator with the time series for reference will cause more difficulties for the annotator and does not truly solve the problem.

[0008] The annotation process may also require giving the same frame to multiple annotators to achieve better agreement on the bounding boxes. This approach was introduced to primarily reduce errors due to laziness, rather than addressing issues caused by ambiguity in the inability to decide on bounds. Voting techniques similar to these still retain box ambiguity, and consecutive frames can still be inconsistent.

[0009] Since the optimization process has no knowledge of the real world and only focuses on converging to the ground truth boxes provided by the annotator, similar dents are annotated differently which puts it in a confusing state as it is forced to converge to different values ​​for similar patterns or the same pattern, which will cause contradictions. Summary of the Invention

[0010] Aspects of the present invention provide a method for training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a true value label to each pixel, wherein each true value label is associated with a corresponding pixel in the pixels and each true value label indicates whether the pixel corresponds to an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag is associated with a corresponding pixel in the pixels and each ignore flag provides an indication that the pixel should be ignored; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability that the pixel corresponds to the image feature; e. for each pixel without an ignore flag, determining a loss value based on the predicted value and the true value label for the pixel, and f. training the machine learning model based on the loss value; and for each pixel with an ignore flag, ignoring the predicted value for the pixel so that it is not used to train the machine learning model.

[0011] Optionally, c. includes inspecting the object, and generating an ignore mask based on the inspection.

[0012] Optionally, c. includes providing for receiving input from manual inspection of the object, and generating the ignore mask based on the input.

[0013] Optionally, c. includes inspecting the object with a sensor to generate three-dimensional inspection data, and generating a ignore mask based on the three-dimensional inspection data.

[0014] Optionally, the training image data comprises one or more images of an object.

[0015] Optionally, the method further comprises generating training image data by imaging the object.

[0016] Optionally, the object is imaged with light.

[0017] Optionally, the training image data includes a series of images of an object each containing the same features, observed from different viewing angles.

[0018] Optionally, the method further comprises generating training image data by imaging the object from a series of different viewpoints.

[0019] Optionally, b. includes displaying the training image data to a human annotator; and receiving a ground truth mask via input from the human annotator, the ground truth mask providing an indication of regions of the training image data containing image features.

[0020] Optionally, d. to f. are repeated, each repetition comprising a corresponding training epoch.

[0021] Optionally, the image features include surface defects.

[0022] Optionally, the image features include surface defects of the aircraft.

[0023] Optionally, the image features include indentations.

[0024] Optionally, the loss value is determined by the following algorithm:

[0025] -y k lnp k -(1-y k )ln(1-p k ).

[0026] Among them, y k is the true value label of the pixel; p k is the predicted value of the pixel, and the pixel corresponding to the image feature has a true value label y of 1 k , and pixels that do not correspond to image features have a true value label y of 0 k .

[0027] Optionally, after the machine learning model has been trained, the machine learning model is used to segment the image in an inference phase.

[0028] Optionally, c. includes creating a ignore mask based on the truth mask, wherein the ignore mask includes a loop around the truth mask, the loop having an inner edge and an outer edge.

[0029] Another aspect of the present invention provides a computer system configured to train a machine learning model using the method of the aforementioned aspect.

[0030] Another aspect of the present invention provides computer software configured to train a machine learning model using the method of the aforementioned aspect.

[0031] Another aspect of the present invention provides a method for training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. providing a true value mask, which provides an indication of areas of the training image data containing image features; c. creating one or more ignore masks based on the true value mask, each ignore mask comprising a loop around the true value mask, the loop having an inner edge and an outer edge; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability that the pixel corresponds to the image feature; e. for each pixel that overlaps with the true value mask and does not overlap with the ignore mask, determining a loss value based on the prediction value of the pixel, and training the machine learning model based on the loss value; and f. for each pixel located between the inner edge and the outer edge of the ignore mask, ignoring the prediction value of the pixel so that it is not used to train the machine learning model.

[0032] Optionally, the periphery of the true value mask includes a margin region extending to the edge; and all or part of the neglect mask is inside the edge such that it overlaps with the margin region of the true value mask.

[0033] Optionally, the periphery of the true value mask includes a margin area extending to the edge of the true value mask; and all or part of the neglect mask is outside the edge of the true value mask so that it does not overlap with the true value mask.

[0034] Optionally, the periphery of the true value mask includes a margin region extending to the edge of the true value mask; a first portion of the ignored mask is inside the edge of the true value mask so that it overlaps with the margin region; and a second portion of the ignored mask is outside the edge of the true value mask so that it does not overlap with the true value mask.

[0035] Optionally, a ignore mask is created based on the ground-truth mask by dilation of lines that follow the edges of the ground-truth mask.

[0036] Optionally, each ignore mask is created based on the ground truth mask by analyzing the ground truth mask by an automatic edge detection process to detect edges of the ground truth mask; and the ignore masks are created so that they have the same shape as the edges of the ground truth mask.

[0037] Optionally, the periphery of the true value mask includes a margin area extending to an edge of the true value mask; and the inner edge and the outer edge of the neglect mask each have the same shape as an edge of the true value mask.

[0038] Optionally, for each neglect mask, a radial distance between an inner edge and an outer edge of the neglect mask is constant across the neglect mask.

[0039] Optionally, the method further includes assigning a true value label to each pixel, wherein each true value label is associated with a corresponding pixel in the pixels and each true value label indicates whether the pixel corresponds to an image feature; and for each pixel that does not overlap with the ignore mask, determining a loss value based on the predicted value and the true value label of the pixel.

[0040] Optionally, the loss value is determined by the following algorithm:

[0041] -y k lnp k -(1-y k )ln(1-p k ).

[0042] Among them, y k is the true value label of the pixel; p k is the predicted value of the pixel, and the pixel corresponding to the image feature has a true value label y of 1 k , and pixels that do not correspond to image features have a true value label y of 0 k .

[0043] Optionally, the training image data includes a series of images of an object each containing the same features, observed from different viewing angles.

[0044] Optionally, the method further comprises generating training image data by imaging the object from a series of different viewpoints.

[0045] Optionally, the object is imaged with light.

[0046] Optionally, b. includes displaying the training image data to a human annotator, and receiving the ground truth masks via input from the human annotator.

[0047] Optionally, d. to f. are repeated, each repetition comprising a corresponding training epoch.

[0048] Optionally, the image features include surface defects.

[0049] Optionally, the image features include surface defects of the aircraft.

[0050] Optionally, the image features include indentations.

[0051] Optionally, after the machine learning model has been trained, the machine learning model is used to segment the image in an inference phase.

[0052] Another aspect of the present invention provides a computer system configured to train a machine learning model using the method of the aforementioned aspect.

[0053] Another aspect of the present invention provides computer software configured to train a machine learning model using the method of the aforementioned aspect.

[0054] Another aspect of the present invention provides a computer system configured to recognize image features, the computer system comprising a machine learning model trained according to the method of the preceding aspect.

[0055] Another aspect of the present invention provides a computer-implemented method for identifying image features, which includes using a machine learning model trained according to the method of the previous aspect to identify image features. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Embodiments of the present invention will now be described with reference to the accompanying drawings, in which:

[0057] Figure 1 A computer system is shown;

[0058] Figure 2 An aircraft is shown being scanned by a camera device to generate training image data;

[0059] Figure 3 shows the training phase, where the machine learning model is trained by a training engine;

[0060] Figure 4 is a flow chart showing the various steps of the training phase;

[0061] Figure 5 A time series of training images is shown;

[0062] Figure 6 The generation of the ground truth mask is shown;

[0063] Figure 7 The ground truth masks on the frames of the training images are shown;

[0064] Figure 8 The ground truth masks on the frames of the training image are shown, along with the regions with high predicted values.

[0065] Figure 9 The ground truth mask is shown;

[0066] Figure 10 Shows the follow Figure 9 The line of the outer edge of the ground truth mask;

[0067] Figure 11 Shown based on Figure 9 The ignore mask of the true value mask;

[0068] Figure 12 Shown Figure 11 The ignore mask and the relevant regions with high prediction values;

[0069] Figure 13 Shown based on Figure 9 Another ignore mask of the true value mask;

[0070] Figure 14 Shown based on Figure 9 Another ignore mask of the true value mask;

[0071] Figure 15 A manual method for generating ignore masks is shown;

[0072] Figure 16 Shown Figure 9 Ignore masks on frames of training images;

[0073] Figure 17 An aircraft is shown being scanned by a LIDAR sensor;

[0074] Figure 18 An automatic method for generating ignore masks is shown;

[0075] Figure 19 The pre-processing set of ignore masks is shown; and

[0076] Figure 20 The ground truth masks on frames of the training images are shown, along with the ignored masks and regions with high predicted values. DETAILED DESCRIPTION

[0077] Figure 1 FIG2 is a schematic diagram showing various elements of a computer system for identifying image features. The computer system includes a memory 1, a machine learning model 2, a training engine 3, and various input / outputs 4.

[0078] The computer system includes computer software configured to train the machine learning model 2 by the method described below.

[0079] In this example, a machine learning model 2 is used to identify features within an object, such as Figure 2 Surface defects (such as dents or scratches) of an aircraft 10 are shown. Aircraft need to be inspected regularly, and the identification of surface defects can be part of such inspections. Once a defect is automatically identified by the machine learning model 2, the defect can be repaired or manually inspected.

[0080] Before being used to identify surface defects in the inference phase, the machine learning model 2 must be trained in the training phase. The training phase described below broadly involves: a pre-training process of providing a true value mask and an ignore mask; and a training process of receiving a set of predicted values ​​from the machine learning model 2 and training the machine learning model 2 accordingly.

[0081] Figure 3 A method for training a machine learning model 2 to recognize image features through deep learning using the above data is shown. Training image data 42 is provided and fed into the machine learning model 2. A training engine 3 receives a set of predicted values ​​60, a set of ground truth masks 50, and a set of ignore masks 20 output by the machine learning model 2. The training phase includes a series of training epochs. For each training epoch, the training engine 3 applies a training update that adjusts the weights of the neural network of the machine learning model 2.

[0082] Figure 4 is a flowchart showing the various steps of the training phase, where a surrounding ignore mask is used to handle blurred dent edges.

[0083] In a first step a., the aircraft 10 is imaged from a series of different viewpoints (e.g. Figure 2 The training image data 42 is generated by scanning the entire exterior surface of the aircraft 10 using a camera 40 that scans the entire exterior surface of the aircraft 10. As an example, a video of the aircraft can be acquired as the camera 40 continuously moves around the aircraft. The camera 40 can be held on the end of a robotic arm or carried by an aerial drone. Thus, the video includes a series of frames, each frame including a training image containing a corresponding portion of the aircraft viewed from a different perspective. These training images together provide a Figure 3 Shown is a complete set of training image data 42. Each training image comprises a plurality of pixels.

[0084] In the above example, the camera 40 senses visible light, but in other examples, the camera 40 may be replaced by an infrared camera (with an active source of infrared light), a thermal camera, or an ultrasound camera.

[0085] In the above examples, images were acquired from a single aircraft 10. In other examples, a complete set of training image data 42 may be acquired by imaging multiple aircraft or by imaging multiple parts.

[0086] Any visible defect will be present in more than one training image, and each training image may contain more than one defect.Herein, the term "visible defect" means a defect that is visible in an image and may (or may not) be visible on the aircraft 10. Figure 5 An example is given of a time series of twenty cropped training images 101 to 120 each containing the same feature (in this case, a dent) observed from different observation positions or different viewing angles.

[0087] In training images 104 to 112, the image of the dent is clearly visible. In training images 101 to 102 and 115 to 120, the dent is barely visible—for example, it may be blurred by glare or less visible due to the viewing angle of the camera or the angle of the light. In training images 103, 113, and 114, the dent is visible, but the image of the dent is blurred.

[0088] Note that only a single indentation per frame is shown in each of images 101 to 120 , but multiple indentations may be visible in each frame.

[0089] exist Figure 4 In step b. of , a set of ground truth masks 50 is provided. Each ground truth mask provides an indication of a region of the training image data 42 that contains an image feature, in this case a dent.

[0090] Each ground truth mask can include a set of ground truth labels, one for each pixel in the region. Each ground truth label is associated with a corresponding one of the pixels and provides a binary indication or flag (value 1) that the pixel corresponds to the image feature. Alternatively, each ground truth mask can consist of only indications of the edges of the region containing the image feature, rather than a "pixel-by-pixel" set of ground truth labels.

[0091] like Figure 6 As shown, a set 50 of ground truth masks can be generated by displaying training image data 42 to a human annotator via a display device 51 and receiving ground truth masks as input from the human annotator via an annotation input device 52. Each ground truth mask can be created by marking the boundaries of a feature in the training image data 42. For example, the annotator can use a mouse, a drawing tool, or a similar input device to mark the outer edge of the feature. The outer edge can be a rectangle or any other shape (including irregular shapes). Alternatively, the annotator can provide each ground truth mask by "painting" over the feature in the training image using a digital painting tool, thereby providing a "pixel-by-pixel" set of ground truth annotations indicating the pixel block associated with the feature.

[0092] If the annotator only marked the outer edges of the image feature, then the interior of the image feature can optionally be “filled” with the ground truth annotation, so that every pixel within the interior of the ground truth mask is assigned a ground truth annotation of 1.

[0093] Any pixel that does not fall within the ground truth mask is assigned a ground truth label of 0. Thus, each pixel in each image is assigned a ground truth label y of either 0 or 1. k .

[0094] Figure 7Twenty frames of training images 101 to 120 are shown, with the dent removed. Frames 104 to 112, where the dent is clearly visible, have been annotated with a ground truth mask 64—in this case, by "painting" over the feature in the training image using a digital painting tool to indicate the patch of pixels associated with that feature.

[0095] Note that the human annotators were instructed to generate ground truth masks only when features were clearly visible. Therefore, no ground truth masks were generated for frames 101 to 103 or frames 113 to 120.

[0096] There may be 1000 images to be annotated, where each frame is randomly presented to different annotators. Therefore, frames with blurred images can be labeled with the ground truth mask, while other frames (such as frames 103, 113 and 114) may not be labeled.

[0097] exist Figure 4 In step d., for each pixel of each training image, the machine learning model 2 generates a prediction value p k The machine learning model 2 includes a neural network that identifies and classifies features in the training image data 42. For example, a feature may be identified and classified as a dent.

[0098] Each prediction value output by the machine learning model 2 provides an indication of the probability that a pixel corresponds to an image feature (e.g., a dent). Each prediction value can take any decimal value p between 0 and 1. k Each predicted value is non-binary in the sense that Figure 3 , the output of the machine learning model 2 is indicated as dataset 60.

[0099] If the machine learning model 2 is 100% sure that the pixel should be classified as an image feature, the predicted value output for that pixel is 1; and if the machine learning model 2 is 100% sure that the pixel should not be classified as an image feature, the predicted value output for that pixel is 0. In most cases, the predicted value will take an intermediate value between 0 and 1. For training images 103 to 114, the machine learning model 2 predicts the dent as an area with a high predicted value. The edge of the area with a high predicted value (~1) is at Figure 8 Indicated by solid lines in the middle - some of these areas are numbered 62, 63.

[0100] As an example, take a simple image segmentation case with only two classes. Each pixel can be classified as either a dent or as background. If the predicted value is high (e.g., greater than 50%), the pixel is classified as a dent; and if the predicted value is low (e.g., less than 50%), it is classified as background (or "non-dent").

[0101] Regions with high predicted values ​​in frames 104 to 112 all overlap with the ground truth mask 64: for example, region 63. Region 62 with high predicted value in frames 102, 113, and 114 does not overlap with the ground truth mask 64. This is because images 103, 113, and 114 contain blurred images of the dents, which are visible enough to be detected by the machine learning model 2, but not clear enough to be annotated with the ground truth mask.

[0102] exist Figure 4 In step c., a set 20 of ignore masks is generated. Figures 10 to 14 Shown based on Figure 9 Three methods of creating ignore masks are shown for the truth mask 64. An ignore mask is generated for each of the truth masks.

[0103] Figure 9 The illustrated ground truth mask 64 includes a core 65 surrounded by an outer periphery 66, 67. The outer periphery includes an outer margin region 66 that extends to an outer edge 67 of the ground truth mask 64. Each ground truth label within the ground truth mask 64 is set to one.

[0104] In this example, the ground truth mask 64 is created by a paint tool. The automatic edge detection process analyzes the ground truth mask 64 to detect outer edges 67 and generates Figure 10 The line 68 of pixels shown follows the detected outer edge. Then, by Figures 11 to 14 One of the methods shown creates a surrounding ignore mask having the same shape as line 68 .

[0105] In all cases, the ignore mask includes a closed loop (ie, a closed path whose starting point coincides with its ending point) at the outer perimeter 66 , 67 of the truth mask 64 .

[0106] exist Figure 11 In the case of , a ignore mask 75 is created based on the true value mask 64 by expanding the line 68 both inwardly and outwardly. The ignore mask 75 has a first (inner) portion 70 that is inside the outer edge 67 of the true value mask so that it overlaps with the outer margin area 66; and a second (outer) portion 71 that is outside the outer edge 67 of the true value mask so that it does not overlap with the true value mask 64.

[0107] The radial distance between the inner edge 72 and the outer edge 73 of the ignore mask 75 is constant around the ignore mask 75 and may be selected by design.

[0108] In alternative embodiments, the radial distance between the inner edge 72 and the outer edge 73 of the ignore mask 75 may vary in a predetermined manner around the ignore mask 75—eg, the radial distance at the top of the ignore mask is greater than the radial distance at the bottom.

[0109] The ignore mask 75 may include a set of ignore flags, each ignore flag being associated with a corresponding one of the pixels between the inner edge 72 and the outer edge 73 of the ignore mask 75. Each ignore flag provides a binary indication that the pixel should be ignored. Each pixel that is consistent with the ignore mask is assigned an ignore flag.

[0110] Alternatively, rather than a "pixel-by-pixel" set of ignore flags, the ignore mask 75 may consist solely of an indication of the inner edge 72 and the outer edge 73. In this case, the interior of the ignore mask 75 is optionally "filled" with ignore flags so that each pixel in the interior has an associated ignore flag.

[0111] Optionally, a pre-processing step is performed such that for all pixels in the inner portion 70 of the ignore mask 75 (where both the truth label and the ignore flag are 1), the truth label is set to 0.

[0112] Back to Figure 3 and Figure 4 : Once the ground truth masks and ignore masks have been generated, the training engine 3 trains the machine learning model 2 in a series of training epochs.

[0113] For each pixel that does not have an ignore flag (i.e., it does not conform to the ignore mask), the training engine 3 determines a loss value based on the predicted value and the true value label of the pixel. This generates a loss value for each pixel unless the pixel has an ignore flag.

[0114] As an example, the loss value for each pixel can be determined by a logistic regression function, such as the function:

[0115] -y k lnp k -(1-y k )ln(1-p k ).

[0116] Among them, y k is the true value annotation, and p k is the predicted value.

[0117] The pixels corresponding to the image features have a ground truth label y of 1 k , and pixels that do not correspond to image features have a true value label y of 0 k Therefore, y k can be 0 or 1, and p k Can be any decimal value between 0 and 1. The loss will be huge if the predicted value does not match the ground truth label, and low otherwise.

[0118] Then, the machine learning model 2 is trained based on the loss value of each pixel. For example, the loss value of each pixel can be summed to determine the average loss value used to train the machine learning model 2.

[0119] For each pixel with an ignore flag, the training engine 3 ignores the predicted value of that pixel so that it does not contribute to the calculation of the average loss value and is therefore not used to train the machine learning model 2.

[0120] Figure 12 The outer edge of a region 63 with high predicted values ​​(~1) is shown, which overlaps the ignore mask 75. A loss value is determined for each pixel that is consistent with the true value mask 64 and also lies within the inner edge 72 of the ignore mask 75 (so it does not coincide with the ignore mask).

[0121] Therefore, all pixels in the core 65 of the ground truth mask are used to train the machine learning model 2. For each pixel that lies between the inner edge 72 and the outer edge 73 of the ignore mask 75 (and therefore coincides with the ignore mask 75), the predicted value for that pixel is ignored so that it is not used to train the machine learning model 2. Therefore, pixels at the outer boundary 76 of the region 63 are ignored.

[0122] Figure 13 and Figure 14 Alternative surrounding ignore masks 80, 81 are shown. Like the ignore mask 75, each ignore mask 80, 81 comprises a closed loop at the outer periphery 66, 67 of the truth mask 64.

[0123] Created by expanding line 68 outwards rather than inwards Figure 13 In this case, all of the ignore masks 80 are located outside the outer edge of the true value mask 64 so that they do not overlap with the true value mask 64. The inner edge of the ignore mask 80 is adjacent to the line 68.

[0124] Created by expanding line 68 inwards rather than outwards Figure 14 In this case, all of the ignore masks 81 are located inside the outer edge of the true value mask 64 so that they overlap with the outer margin area 66 of the true value mask 64. The outer edge of the ignore mask 81 is line 68.

[0125] Table 1 shows the eight scenarios that will be encountered at the pixel level and the corresponding loss values.

[0126] Table 1

[0127]

[0128]

[0129] For each pixel with a true value label of zero and no ignored flag (i.e., scene #1 and scene #2), a loss value is determined based on the predicted value and the true value label of the pixel, and the machine learning model 2 is trained based on the loss value to learn the pixel as background. Figure 12 Indicates pixel #1 and pixel #2 corresponding to scene #1 and scene #2, respectively.

[0130] For each pixel with an ignore flag (i.e., it is consistent with the ignore mask) and a true value label of zero (i.e., scene #3 and scene #4), the training engine 3 ignores the predicted value of that pixel so that it is not used to train the machine learning model 2. Figure 12 Indicates pixel #3 corresponding to scene #3 and pixel #4 corresponding to scene #4.

[0131] For each pixel with a true value label of 1 (i.e., it is consistent with the true value mask) and no ignored flag (i.e., scene #5 and scene #6), a loss value is determined based on the predicted value and the true value label of the pixel, and the machine learning model 2 is trained based on the loss value to learn that the pixel is a dent. Figure 12 Pixels #5 and #6 correspond to scenes #5 and #6, respectively. These pixels #5 and #6 each overlap with the kernel 65 of the true value mask and lie within the inner edge 72 of the ignore mask (thus, they do not overlap with the ignore mask). Therefore, pixels #5 and #6 do not overlap with the ignore mask 75 and are not ignored.

[0132] Figure 12 Indicates pixel #7 corresponding to scene #7 and pixel #8 corresponding to scene #8. Note that if the pre-processing steps mentioned above are performed, scene #7 and scene #8 will not occur because the pre-processing steps set the ground truth labels to 0.

[0133] If the above-mentioned pre-processing steps are not performed, the training engine 3 processes pixels #7 and #8 in the same manner as pixels #3 and #4. Therefore, for each pixel with an ignore flag and a true value label of 1 (i.e., scene #7 and scene #8), the training engine 3 ignores the predicted value of the pixel so that it is not used to train the machine learning model 2.

[0134] Figure 7 Two consecutive image frames 111 and 112 are shown, both containing images of the same indentation at approximately the same location, but with completely different ground truth masks. This inconsistency could be caused, for example, by different people creating the ground truth masks. Alternatively, even if the same annotator generated both ground truth masks, they were not presented with the consecutive frames 111 and 112 immediately following one another, which could lead to inconsistencies.

[0135] Therefore, pixels 65, 66 at the same position at the edge of the dent are inside the ground truth mask in one image 112, but outside the ground truth mask in the adjacent image 111 of the time series.

[0136] This inconsistency can cause the machine learning model 2 to be forced to classify pixel 65 as a dent and pixel 66 as background. Such contradictory inputs can lead to poor and inconsistent training. The surrounding ignore masks 75, 80, and 81 described above can prevent such blurry pixels from being used to train the machine learning model 2, thereby improving the quality and consistency of training.

[0137] In the example above, Figure 9 The ground truth mask 64 comprises a continuous core 65 surrounded by outer perimeters 66, 67. Thus, the ground truth mask 64 has outer perimeters 66, 67 but no inner perimeter, so only a first (outer) perimeter ignore mask 75, 80, 81 following the outer edge 67 is generated.

[0138] In other examples, the true value mask may contain holes, so it also has an inner periphery with an inner edge. In such a case, a second (inner) periphery ignore mask may also be generated, such ignore mask including a loop at the inner periphery of the true value mask, the loop having an inner edge and an outer edge.

[0139] The second (inner) surrounding ignore mask can be generated in the same manner as the first (outer) surrounding ignore mask. That is, the second (inner) surrounding ignore mask can be created based on the true value mask by analyzing the true value mask by an automatic edge detection process to detect the inner edge of the true value mask; and creating the ignore mask so that it has the same shape as the inner edge of the true value mask.

[0140] In such a case, the pixels within the inner edge of the loop will not overlap with the true value mask (because they are located in a hole in the true value mask). Such a second (inner) surrounding ignore mask can then be used in the same manner as the first (outer) ignore mask. In such a case, for each pixel that overlaps with the true value mask and does not overlap with the ignore mask (i.e., does not overlap with the first surrounding ignore mask or the second surrounding ignore mask), a loss value is determined based on the predicted value of the pixel, and the machine learning model is trained based on the loss value; and for each pixel located between the inner and outer edges of the ignore mask (i.e., it overlaps with the first surrounding ignore mask or the second surrounding ignore mask), the predicted value of the pixel is ignored so that it is not used to train the machine learning model.

[0141] exist Figure 5In the time series of , the image of the dent starts to fade, becomes prominent, and then fades again. This is a continuous process, and there is no strict line between the dent and no dent cases. Due to its robustness, the machine learning model 2 will start to detect dents in the blurred images 103, 113, 114 - see Figure 8 Since machine learning model 2 has no intermediate classes, it considers any pixel to be either a dent or a non-dent. Since frames 103, 113, and 114 with blurred images are not labeled with ground truth masks, machine learning model 2 may attempt to "forget" that region 62 is a dent, potentially affecting dent detection in other images. This inconsistency can lead to inconsistencies in the machine learning algorithm.

[0142] For the training images 101 to 102 and 115 to 120 where the dents are barely visible, it is undesirable to force the machine learning model 2 to learn the areas of the dents as background simply because of low visibility, since these are areas where the dents physically exist.

[0143] Figures 15 to 20 A method for training a machine learning model 2 to recognize image features is shown, which uses box-ignoring masks to handle such a problem. Figures 15 to 20 The process of using Figure 3 The difference is that the box-ignoring mask is not based on the set of ground-truth masks50.

[0144] The term "box" ignore mask is used to refer to an ignore mask having an outer edge but no inner edge, as opposed to the "surrounding" ignore mask described above, which is a loop having both an inner edge and an outer edge.

[0145] Figure 15 A manual process for generating a set 20 of box ignore masks is shown. Each box ignore mask is generated by receiving input from a manual inspection of the aircraft 10 via an input device 22. That is, a human inspector visually inspects the aircraft 10 and enters the location of any visible defects into the input device. A digital mock-up (DMU) 23 of the aircraft 10 provides a 3D representation of the aircraft 10. An ignore mask generator 24 (which may be Figure 1 The computer system (or part of another computer system) receives input from the input device 22, superimposes it on the DMU 23, and accordingly generates a set of box-ignore masks 20. Each box-ignore mask can include the outer boundary of the defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each box-ignore mask can include a set of pixels, each of which is associated with a defect, i.e., it is located within the boundary of the defect.

[0146] Figure 17 and Figure 18 An alternative automated process for generating a set 20 of box ignore masks is shown. In this case, the ignore masks are provided by inspecting the aircraft with a sensor 30 (e.g., a LIDAR sensor) to generate three-dimensional inspection data 31 (e.g., LIDAR data) and generating the ignore masks based on the inspection data 31. The ignore masks are generated by an automated ignore mask generator 32 (which may be Figure 1 The computer system or part of another computer system) automatically generates the ignore mask based on the inspection data 31 and the DMU 23. Figure 15 In , each box-ignore mask may include the outer boundary of the defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each box-ignore mask may include a set of pixels, each of which is associated with a defect, i.e., it is located within the boundary of the defect.

[0147] Figure 16 Frames 101 to 120 are shown with box-ignore masks created by one of the methods described above.

[0148] Figure 16 Each box in the ignore mask may include a set of ignore flags, each flag being associated with a corresponding one of the pixels within the outer edge of the ignore mask. Each ignore flag provides a binary indication that the pixel should be ignored.

[0149] Alternatively, each box ignore mask may consist of only an indication of the outer edges of the rectangle, rather than a "pixel-by-pixel" set of ignore flags. In this case, the interior of the box ignore mask may optionally be "filled" with ignore flags so that each pixel in the interior has an associated ignore flag.

[0150] In the preprocessing stage, all boxes that overlap with the ground-truth mask are removed and the mask is ignored. Figure 19 Shows how the preprocessing stage changes Figure 16 The data shown. Figure 7 The truth mask 64 is shown with the box-ignore masks that overlap it removed, leaving only eleven box-ignore masks.

[0151] Figure 20 Will Figure 19 The box ignore mask is superimposed on the ground truth mask 64 and Figure 8 Predictions 62 and 63.

[0152] The training engine follows the process in Table 1, thus using Figure 20 The box ignores the mask.

[0153] For each pixel with an ignore flag (i.e., scene #3 and scene #4 where the pixel is consistent with the box ignore mask), the training engine 3 ignores the predicted value of that pixel so that it is not used to train the machine learning model 2. All pixels within the box ignore mask in frames 101, 102, and frames 115 to 120 correspond to scene #3 or scene #4. Due to the pre-processing stage mentioned above, scene #7 and scene #8 will not occur.

[0154] For each pixel without the ignored flag (i.e., scene #1, scene #2, scene #5, and scene #6), a loss value is determined based on the predicted value and the true value label of the pixel, and the machine learning model 2 is trained based on the loss value. Figure 20 All pixels within the ground-truth mask of correspond to scene #5 or scene #6.

[0155] Note that there are three regions 62 with high prediction values ​​that do not overlap with the ground truth mask 64. In the previous method of generating surrounding ignore masks 75, 80, 81 based on the ground truth mask 64, the pixels in the core of these regions 62 correspond to scene #2 and are used as false positives to train the machine learning model 2. Figure 20 In this case, all pixels in these regions 62 fall within the block ignore mask, so they correspond to scene #4 and are ignored.

[0156] exist Figure 20 During the training process shown, the machine learning model 2 is forced to learn the dents from images 104 to 112, and the block ignore mask causes the other images to be undecided. This avoids the problem of poor visibility of the dents as described above.

[0157] Figures 11 to 14 The surrounding ignore masks 75, 80, 81 provide a solution to the problem of blurring the exact boundaries of the dents in the images 104 to 112; Figure 20 The larger block ignore mask in provides a solution to the problem associated with training the machine learning model 2 based on images 101 to 103 and images 115 to 120, where the dents are blurred or barely visible.

[0158] Optionally, both solutions can be used together in the same training process: i.e., in addition to Figure 20 In addition to the block ignore mask, you can also create Figures 11 to 14 The surrounding masks 75, 80, 81 are ignored and applied to frames 104 to 112 with true value masks.

[0159] In the above example, a single prediction value is determined for each pixel by the machine learning model 2. This prediction value provides an indication of the probability that the pixel corresponds to an image feature (e.g., a dent). In other embodiments of the present invention, the machine learning model 2 may output multiple prediction values ​​for each pixel, each associated with a different defect category, such as a dent or a scratch. These multiple prediction values ​​can then be used to classify each pixel as, for example, a dent and / or a scratch.

[0160] In the case where there are more than two defect classes, the annotator may also have the ability to generate a ground truth mask associated with each defect class. Ignore masks for these multiple defect classes may also be generated and used as described above.

[0161] In other embodiments of the present invention, the machine learning model 2 can output a single prediction value for each pixel, and the single prediction value is used to identify the pixel as background, unclear or defective: for example, 0-30% = background class; 30%-70% = intermediate class; 70%-100% = dent class.

[0162] After the machine learning model 2 has been trained to recognize image features through the above process, it can then be used to segment the image in the inference phase. In the inference phase, each pixel is classified based on the predicted value output by the trained machine learning model 2.

[0163] Where the word "or" appears, this is to be interpreted as meaning "and / or" such that the items involved are not necessarily mutually exclusive and can be used in any appropriate combination.

[0164] Although the invention has been described above with reference to one or more preferred embodiments, it will be appreciated that various changes or modifications may be made without departing from the scope of the invention as defined in the appended claims.

Claims

1. A method for training a machine learning model to recognize image features, the method comprising: a. Providing training image data, wherein the training image data comprises a plurality of pixels; b. providing a ground truth mask, the ground truth mask providing an indication of an area of ​​the training image data including image features; c. creating one or more ignore masks based on the truth mask, each ignore mask comprising a loop around the truth mask, the loop having an inner edge and an outer edge; d. For each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability that the pixel corresponds to an image feature; e. for each pixel that overlaps with the true value mask and does not overlap with the ignore mask, determining a loss value based on the predicted value of the pixel, and training the machine learning model based on the loss value; and f. For each pixel located between the inner and outer edges of the ignore mask, ignore the predicted value of that pixel so that it is not used to train the machine learning model.

2. The method according to claim 1, wherein The periphery of the true value mask includes a margin area extending to an edge; and all or part of the ignore mask is inside the edge, so that all or part of the ignore mask overlaps with the margin area of ​​the true value mask.

3. The method according to claim 1 or 2, wherein: The periphery of the true value mask includes a margin area extending to the edge of the true value mask; and all or part of the ignore mask is outside the edge of the true value mask, so that all or part of the ignore mask does not overlap with the true value mask.

4. The method according to claim 1, wherein The periphery of the true value mask includes a margin area extending to the edge of the true value mask; the first portion of the ignore mask is inside the edge of the true value mask, so that the first portion of the ignore mask overlaps with the margin area; and the second portion of the ignore mask is outside the edge of the true value mask, so that the second portion of the ignore mask does not overlap with the true value mask, and optionally, The ignore mask is created based on the ground-truth mask by dilating lines along the edges of the ground-truth mask.

5. The method according to claim 1, wherein creating each ignore mask based on the ground truth mask by analyzing the ground truth mask by an automatic edge detection process to detect edges of the ground truth mask; And creating the ignore mask so that the ignore mask has the same shape as the edge of the true value mask.

6. The method according to claim 1, wherein The periphery of the true value mask includes a margin area extending to an edge of the true value mask; and the inner edge and the outer edge of the ignore mask each have the same shape as the edge of the true value mask.

7. The method according to claim 1, wherein For each neglect mask, the radial distance between the inner edge and the outer edge of the neglect mask is constant across the neglect mask.

8. The method of claim 1, further comprising assigning a ground truth label to each pixel, wherein Each ground truth annotation is associated with a corresponding pixel in the pixels, and each ground truth annotation indicates whether the pixel corresponds to an image feature; And for each pixel that does not overlap with the ignore mask, a loss value is determined based on the predicted value and the true value label of the pixel.

9. The method according to claim 8, wherein the loss value is determined by the following algorithm: -and k lnp k -(1-and k )ln(1-p k ). in, y k is the true value label of the pixel; p k is the predicted value of the pixel, and the pixel corresponding to the image feature has a true value label y of 1 k , and pixels that do not correspond to image features have a true value label y of 0 k .

10. The method according to claim 1, wherein The training image data includes a series of images of an object each containing the same features, observed from different viewing angles.

11. The method of claim 10, further comprising generating the training image data by imaging the object from a series of different viewing angles.

12. The method according to claim 11, wherein The object is imaged with light.

13. The method of claim 1, wherein b. comprises displaying the training image data to a human annotator; and receiving the ground truth mask via input from the human annotator.

14. The method according to claim 1, wherein Repeat d. to f., with each repetition including the corresponding training period.

15. The method according to claim 1, wherein The image feature comprises a surface defect, optionally wherein the image feature comprises a surface defect of an aircraft, and further optionally wherein the image feature comprises a dent.

16. The method according to claim 1, wherein After the machine learning model has been trained, the machine learning model is used to segment images in the inference phase.

17. A computer system configured to train a machine learning model using the method of claim 1.

18. Computer software configured to train a machine learning model using the method of claim 1.

19. A computer system configured to recognize image features, the computer system comprising a machine learning model trained according to the method of claim 1.

20. A computer-implemented method for identifying image features, comprising identifying image features using a machine learning model trained according to the method of claim 1.