Method for training machine learning model

By generating similarity coefficients and threshold-based loss values, the machine learning model is trained, and the problems of inconsistent fuzzy dent labeling and confusion in the optimization process during training are solved, which reduces the false positive rate and improves the recognition accuracy of the model.

CN120088593APending Publication Date: 2025-06-03AIRBUS (SAS)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411727903.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-11-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, when training machine learning models to identify image features, especially blurred dents, there are problems such as inconsistent labeling, confusing optimization process, and high false positive rates.

Method used

By receiving a set of predicted feature regions, a set of similarity coefficients is generated, the loss value is determined based on the similarity coefficient and the threshold, and the machine learning model is trained based on the loss value. This method introduces similarity coefficients less than and greater than the threshold in the training round to avoid overtraining the model.

Benefits of technology

It effectively solves the problems of inconsistent fuzzy dent labeling and confusing optimization process, reduces the false positive rate, and improves the identification accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088593A_ABST
    Figure CN120088593A_ABST
Patent Text Reader

Abstract

A method of training a machine learning model to identify image features, the method comprising: a. Receiving a set of predicted feature regions from the machine learning model, each predicted feature region comprising a prediction of features in training image data; b, generating a set of similarity coefficients, wherein each similarity coefficient indicates the similarity between the predicted feature region and a corresponding true value region overlapped with the predicted feature region; c, determining a loss value based on the similarity coefficient and a threshold value; and d. Training the machine learning model based on the loss value, where a. To d. Are repeated, each repetition comprising a respective training round; in one or more of the training rounds, the set of similarity coefficients includes one or more similarity coefficients less than a threshold, and the loss value is based on a difference between the threshold and each similarity coefficient less than the threshold; and in one or more of the training rounds, the set of similarity coefficients includes similarity coefficients greater than a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods for training machine learning models to identify image features, such as dents or other surface defects. The present invention also relates to computer systems and computer software configured to train machine learning models. Background Art

[0002] Object detection is a known method of locating objects within an image. Modern deep learning algorithms solve this problem by collecting a large number of images in which the desired object is present. A person marks bounding boxes around it. These boxes are called ground truth boxes or annotations, and these are the values to which the algorithm converges during the iterative learning process. The person doing the annotation is called an annotator.

[0003] The deep learning algorithm then attempts to learn this through optimization techniques to converge to the exact values given by the annotator. Through processes called backpropagation and gradient descent, the process of converging to the annotation values occurs through multiple iterations. Gradient descent calculates the error between the model prediction and the ground truth and attempts to minimize this error.

[0004] In the case of detecting rigid objects with clear contours, it may be easy for an annotator to know the boundaries of the object. In the case of features within the object (such as dents), this may be more difficult.

[0005] Due to the ambiguous nature of dents and the frames being randomly fed for annotation, annotators can sometimes annotate or skip ambiguous dents according to their personal preferences. In the case where an ambiguous dent is annotated, the machine learning model attempts to learn it as "yes", and in the case where it is skipped, the machine learning model attempts to learn it as "no". Thus, the optimizer is also in a contradictory and inconsistent situation of sometimes learning and sometimes not learning this pattern. Sometimes, the ambiguous dents are so insignificant that if the machine learning model is forced to learn them, the machine learning model may end up detecting many false positives, making it very sensitive.

[0006] Under normal circumstances, there is a specific class called the background class, which the algorithm uses to distinguish and learn such false detections from actual detections. In ambiguous cases, the background class will contradict the dent class and lead to inconsistencies.

[0007] When annotating frames, they may be randomly selected from a video, so annotators usually do not get a time series as a hint for locating dents and marking the correct bounding boxes. Since this is already a long manual process, providing a time series for reference to the annotator will cause more difficulties for the annotator and does not really solve the problem.

[0008] The annotation process may also require giving the same frame to multiple annotators to have better consistency in the bounding boxes. This method is introduced mainly to reduce the errors caused by laziness rather than to solve the problems caused by the ambiguity of not being able to determine the boundaries. Similar voting techniques like these still retain the ambiguity of the boxes, and consecutive frames will still be inconsistent.

[0009] Since the optimization process has no knowledge of the real world and only focuses on converging to the ground truth boxes provided by the annotators, similar indentations are labeled differently, causing confusion because it is forced to converge to different values for similar or the same patterns, which will cause contradictions. Summary of the Invention

[0010] A first aspect of the present invention provides a method for training a machine learning model to identify image features, the method comprising: a. receiving a set of predicted feature regions from the machine learning model, each predicted feature region including a prediction of a feature in the training image data; b. generating a set of similarity coefficients, each similarity coefficient indicating the similarity between the predicted feature region and the corresponding ground truth region overlapping the predicted feature region; c. determining a loss value based on the similarity coefficients and a threshold; and d. training the machine learning model based on the loss value, wherein steps a. to d. are repeated, each repetition including a corresponding training round; in one or more training rounds of the training rounds, the set of similarity coefficients includes one or more similarity coefficients less than the threshold, and the loss value is based on the difference between the threshold and each similarity coefficient less than the threshold; and in one or more training rounds of the training rounds, the set of similarity coefficients includes similarity coefficients greater than the threshold.

[0011] Optionally, in one or more training rounds of the training rounds, the set of similarity coefficients includes one or more similarity coefficients greater than the threshold, and the loss value is based on the difference between the threshold and each similarity coefficient greater than the threshold.

[0012] Optionally, the threshold varies between at least two training rounds of the training rounds.

[0013] Optionally, the threshold decreases between at least two training rounds of the training rounds.

[0014] Optionally, the training image data includes a series of images of an object containing the same feature observed from different perspectives.

[0015] Optionally, the method further includes generating the training image data by imaging the object from a series of different perspectives.

[0016] Optionally, the object is imaged with visible light.

[0017] Optionally, each feature includes a surface defect.

[0018] Optionally, each feature includes a surface defect of the aircraft.

[0019] Optionally, each feature includes a dent.

[0020] Optionally, the similarity coefficient is the Jaccard index.

[0021] Optionally, the method further includes generating a ground truth region by: displaying training image data to a human annotator and receiving a ground truth region as input from the human annotator, each ground truth region including an annotation of a feature in the training image data.

[0022] Another aspect of the present invention provides a computer system configured to train a machine learning model by the method of the first aspect.

[0023] Another aspect of the present invention provides a computer software configured to train a machine learning model by the method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Embodiments of the present invention will now be described with reference to the accompanying drawings, in which:

[0025] Figure 1 A computer system is shown;

[0026] Figure 2 An aircraft is shown;

[0027] Figure 3 A manual method of generating an ignore region is shown;

[0028] Figure 4 An aircraft scanned by a LIDAR sensor is shown;

[0029] Figure 5 An automatic method of generating an ignore region is shown;

[0030] Figure 6 An aircraft scanned by a camera device to generate training image data is shown;

[0031] Figure 7 Four exemplary training images are shown;

[0032] Figure 8 An ignore region superimposed on an exemplary training image is shown;

[0033] Figure 9 Generation of a ground truth region is shown;

[0034] Figure 10 An ignore region and a ground truth region superimposed on an exemplary training image are shown;

[0035] Figure 11 Shows the removal of the ignored regions in preprocessing;

[0036] Figure 12 Shows the ignored regions and the ground truth regions superimposed on an exemplary training image after preprocessing;

[0037] Figure 13 Shows the generation of the predicted feature regions;

[0038] Figure 14 Shows the predicted feature regions superimposed on an exemplary training image;

[0039] Figure 15 Shows all the regions superimposed on an exemplary training image;

[0040] Figure 16 Shows the training phase in which the machine learning model is trained by a training engine; and

[0041] Figure 17 Shows the selected key steps of the training phase. Detailed Description of the Invention

[0042] Figure 1 Is a schematic diagram showing various elements of a computer system for identifying image features. The computer system includes a memory 1, a machine learning model 2, a training engine 3, and various input / outputs 4.

[0043] The computer system includes computer software configured to train the machine learning model by the method described below.

[0044] In this example, the machine learning model 2 is used to identify features within an object, such as Figure 2 The surface defects (such as dents or scratches) of the aircraft 10 shown. The aircraft needs to be inspected regularly, and the identification of surface defects can be part of such an inspection. Once a defect is automatically identified by the machine learning model 2, the defect can be repaired or manually inspected.

[0045] Before being used to identify surface defects in the inference phase, the machine learning model 2 must be trained in the training phase. The training phase described below broadly involves: a pre-training process of providing ground truth regions and ignored regions; and a training process of receiving a set of predicted feature regions from the machine learning model and training the machine learning model 2 accordingly.

[0046] Figure 3A manual process for generating a set 20 of ignore regions 21 is shown. The ignore regions 21 are generated by receiving input from a manual inspection of the aircraft 10 via an input device 22. That is, a human inspector visually inspects the aircraft 10 and inputs the location of any visible defects into the input device. A digital mock-up (DMU) 23 of the aircraft 10 provides a 3D representation of the aircraft 10. An ignore region generator 24 (which can be a Figure 1 computer system or part of another computer system) receives the input from the input device 22, superimposes the input onto the DMU 23, and accordingly generates the set of ignore regions 21. Each ignore region 21 can include the boundary of a defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each ignore region 21 can include a set of pixels each associated with the defect, that is, they are located within the boundary of the defect.

[0047] Figure 4 and Figure 5 An alternative automatic process for generating the set 20 of ignore regions 21 is shown. In this case, the ignore regions 21 are provided by: inspecting the aircraft with a sensor 30 (e.g., a LIDAR sensor) to generate three-dimensional inspection data 31 (e.g., LIDAR data) and generating the ignore regions 21 based on the inspection data 31. The ignore regions 21 are automatically generated by an automatic ignore region generator 32 (which can be a Figure 1 computer system or part of another computer system) based on the inspection data 31 and the DMU 23. As Figure 3 in, each ignore region 21 can include the boundary of a defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each ignore region 21 can include a set of pixels each associated with the defect, that is, they are located within the boundary of the defect.

[0048] As Figure 6 shown, training image data 42 is also generated by imaging the aircraft 10 from a series of different viewpoints. The aircraft 10 can be imaged with visible light by an imaging device 40 that scans the entire outer surface of the aircraft 10. As an example, a video of the aircraft can be acquired when the imaging device 40 moves continuously around the aircraft. The imaging device 40 can be held at the end of a robotic arm or carried by an aerial drone. Thus, the video includes a series of frames, each frame including a training image 41 that contains the corresponding part of the aircraft viewed from a different viewpoint. These training images 41 together provide a set of training image data 42. Note that the object (in this case, the aircraft 10) in the training image data 42 is the same as the object being inspected to generate the ignore regions 21.

[0049] Any visible defect will be present in more than one training image, and each training image may contain more than one defect. Figure 7 An example is given of a series of four training images 41a to 41d, each containing the same pair of features, viewed from different perspectives. In the first training image 41a, two features (designated X and Y) are clearly visible on the right side of the frame of the training image. In the next training image 41b, the two features have moved slightly to the left and are slightly less visible (since they are viewed from a different angle and under different lighting conditions). In the next training image 41c, the two features are still just visible, but only faintly. In the final training image 41d, the features are almost invisible - for example, they may be completely blurred by glare.

[0050] The two features (designated as X and Y in Figure 7 will each have an associated ignore region 21, and Figure 8 shows the positions of these ignore regions in each of the training images 41a to 41d superimposed on the training images (designated as C1, C3, C5, C7 for feature X and C2, C4, C6, C8 for feature Y). Note that there is a pair of ignore regions in each training image (including training image 41d where the features are invisible).

[0051] A set of ground truth regions 50 is also provided, each ground truth region including an annotation of the feature in the training image data 42. As Figure 9 shown, the set of ground truth regions 50 can be generated by displaying the training image data 42 to a human annotator via a display device 51 and receiving the ground truth regions as input from the human annotator via an annotation input device 52. Each ground truth region can include an annotation of the boundary of the feature in the training image data 42. For example, the annotator can use a mouse or a similar input device to mark the boundary around the feature. The boundary can be rectangular or any other shape (including an irregular shape). Alternatively, the annotator can annotate the feature by "painting" above the feature in the training image using a digital painting tool, thereby indicating the set of pixels associated with the feature.

[0052] In Figure 9 two ground truth regions B1, B2 superimposed on the training image 41a are shown. The ground truth region B1 encloses the feature X, and the ground truth region B2 encloses the feature Y. In this case, each ground truth region includes a rectangular bounding box enclosing the defect.

[0053] Note that not all of the training images in the training image data 42 can be shown to the human annotator, and the training images may not be shown in a continuous order.

[0054] Figure 10shows the true value regions B1 to B5 superimposed on Figure 8 the ignored regions. Note that the three ignored regions C6 to C8 do not have overlapping true value regions.

[0055] In Figure 11 the preprocessing stage shown, the preprocessing engine 55 generates a preprocessing set 20a of ignored regions. In the preprocessing set 20a, all ignored regions that overlap with the true value regions are removed. Figure 12 shows how the preprocessing stage changes Figure 10 the data shown. The ignored boxes C1 to C5 that overlap with the true value regions B1 to B5 have been removed, leaving only the three ignored regions C6 to C8.

[0056] Once the true value regions and the ignored regions have been generated and preprocessed, the training engine 3 trains the machine learning model 2 in a series of training rounds. In each training round, the training image data 42 is input into the machine learning model 2, as Figure 13 shown.

[0057] The machine learning model 2 includes a neural network that identifies and classifies features in the training image data 42. For example, the features can be identified and classified as dents. Other parts of the training image data can be classified as background.

[0058] The output of the machine learning model 2 is a set 60 of predicted feature regions. Each predicted feature region can include the boundary of a defect, such as a rectangular bounding box, or any other shape (regular or irregular) surrounding the defect. Alternatively, each predicted feature region can include a set of pixels, each of which is associated with the defect, i.e., they are located within the boundary of the defect.

[0059] Figure 14 Examples of six predicted feature regions A1 to A7 superimposed on the training images 41a to 41d are given. In this example, each predicted feature region A1 to A7 is a rectangular bounding box that encloses the features that have been identified and classified as defects by the machine learning model 2.

[0060] Figure 15 shows six predicted feature regions A1 to A7 superimposed on Figure 12 the ignored regions and the true value regions. Each of the predicted feature regions A1 to A5 overlaps with the true value regions B1 to B5. The predicted feature region A6 overlaps with the ignored region C8, but does not overlap with any of the true value regions in the true value regions. The predicted feature region A7 does not overlap with the ignored regions or the true value regions. The ignored regions C6, C7 do not overlap with the predicted feature regions or the ignored regions.

[0061] Note that the ignored region has a larger area compared to its associated ground truth region or its associated predicted feature region. For example, region C8 is larger than region A6, and region C5 is larger than region B5.

[0062] Figure 16 A method for training a machine learning model to identify image features through deep learning using the above data is shown. The training engine 3 receives a set 60 of predicted feature regions, a set 50 of ground truth regions, and a preprocessed set 20a of ignored regions from the machine learning model 2. For each training round, the training engine 3 applies a training update that adjusts the weights in the neural network of the machine learning model 2.

[0063] Table 1 below shows how the method works in different scenarios. In Table 1, the number 1 indicates the presence of a region (before preprocessing), and the number 0 indicates the absence of a region (before preprocessing).

[0064] Table 1

[0065]

[0066] In the first scenario, labeled #1 in Table 1, there are no ignored regions, no predicted feature regions, and no ground truth regions. In scenario #1, the action of the training engine 3 depends on whether the machine learning model 2 has a background class. If the machine learning model 2 has a background class, the training engine 3 trains the machine learning model 2 to classify the relevant part of the training image data 60 as background. If the machine learning model 2 does not have a background class, the training engine 3 takes no action.

[0067] In the second scenario, labeled #2 in Table 1, there are ignored regions that do not overlap with any of the predicted feature regions in the predicted feature region set and do not overlap with any of the ground truth regions in the ground truth region set. See Figure 15 ignored regions C6 and C7 in as examples. This may be caused by unclear features in the training image data or errors in the generation of the ignored regions. In scenario #2, the training engine 3 ignores the data, so it does not train the machine learning model to classify the relevant part of the training image data as background. In other words, there is no feedback via the training update path to the machine learning model 2.

[0068] In the third scenario, labeled #3 in Table 1, there are ground truth regions that do not overlap with any of the predicted feature regions in the predicted feature region set and do not overlap with any of the ignored regions in the ignored region set. In scenario #3, the training engine 3 can feed the ground truth regions into the machine learning model 2 to train the machine learning model 2 to classify the relevant part of the training image data as a defect.

[0069] In the fourth scenario, labeled as #4 in Table 1, there is a ground truth region that overlaps with the ignore region but does not overlap with any of the predicted feature regions in the predicted feature regions. In scenario #4, Figure 11 preprocessing of Figure 11 removes the ignore region, and the training engine 3 feeds the ground truth region into the machine learning model 2 to train the machine learning model 2 to classify the relevant part of the training image data as a defect. In scenario #4, it is certain that a defect exists because both an ignore region and a ground truth region exist. Therefore, it is certain that the absence of a predicted feature region is a false negative, and thus the machine learning model 2 needs to learn to identify the defect.

[0070] In the fifth scenario, labeled as #5 in Table 1, there is a predicted feature region that does not overlap with any of the ignore regions in the ignore region and does not overlap with any of the ground truth regions in the ground truth region. See predicted feature region A7 as an example. This may be a false positive, which requires training the machine learning model 2 to avoid the same false positive in future training rounds. Therefore, in scenario #5, the training engine 3 trains the machine learning model 2 based on the predicted feature region.

[0071] In scenario #5, the training engine 3 can be trained in different ways based on the predicted feature region. In one example, in the case where the machine learning model 2 has a background class, the training engine 3 can train the machine learning model 2 to classify the relevant part of the training image data as the background. In another example, for example, in the case where the machine learning model 2 does not have a background class, the training engine 3 can train the machine learning model 2 to forget the predicted feature region A7. For example, the machine learning model 2 may have predicted the predicted feature region A7 with 80% confidence, and the training engine 3 trains the machine learning model 2 to forget the predicted feature region A7 by setting this confidence to 0%. In another example, in scenario #5, the training engine 3 can do both: that is, classify the relevant part of the training image data as the background and forget the predicted feature region.

[0072] In the sixth scenario, labeled as #6 in Table 1, there is a predicted feature region that overlaps with the corresponding ignore region and does not overlap with any of the ground truth regions in the ground truth region. See Figure 15 predicted feature region A6 in Figure 15 as an example. This is a fuzzy prediction, which may be caused by unclear features in the training image data 42. The fuzzy prediction in scenario #6 is not used to train the machine learning model 2. Therefore, in scenario #6, the training engine 3 ignores the predicted feature region A6 so that it is not used to train the machine learning model 2. In other words, the existence of the ignore region prevents any feedback via the training update path to the machine learning model 2.

[0073] In the seventh scenario, labeled #7 in Table 1, there is a predicted feature region that overlaps with the corresponding ground truth region and does not overlap with any of the ignored regions in the ignored region. Scenario #7 may be caused by an error in the generation of the ignored region and will be rare. In scenario #7, the training engine 3 generates a similarity coefficient indicating the similarity between the predicted feature region and the corresponding ground truth region.

[0074] In the eighth scenario, labeled #8 in Table 1, there is a predicted feature region that overlaps with the corresponding ground truth region and also overlaps with the ignored region. See Figure 15 predicted feature regions A1 to A5 in as an example. In scenario #8, Figure 11 preprocessing removed the ignored region. Therefore, in scenario #8, the operation of the training engine 3 is the same as in scenario #7, and a similarity coefficient indicating the similarity between the predicted feature region and the corresponding ground truth region is generated.

[0075] Figure 17 is a flowchart describing how the training engine 3 trains the machine learning model 2. After the generation and preprocessing of the ground truth region and the ignored region, a series of steps a. to d. are repeatedly executed, and each repetition includes a corresponding training round.

[0076] In step a., the training engine 3 receives a set of predicted feature regions from the machine learning model 2, and each predicted feature region includes a prediction of the boundary of a feature in the training image data.

[0077] In step b., the training engine 3 generates a set of similarity coefficients, and each similarity coefficient indicates the similarity between the predicted feature region and the corresponding ground truth region that overlaps with the predicted feature region. As an example, each similarity coefficient generated in step b. can be the Jaccard index: IoU = │A∩B│ / │AUB│. Alternatively, the similarity coefficient can be any other suitable indicator.

[0078] In step c., the training engine 3 determines a loss value based on the similarity coefficient.

[0079] In step d., the training engine 3 trains the machine learning model 2 based on the loss value. In step d., the training engine can also apply other required training updates - such as background classification in scenarios #1 and #5, or feeding the ground truth into the model in scenarios #3 and #4.

[0080] After the machine learning model 2 has been trained, it can be used later to identify surface defects of the aircraft 10 (or another aircraft) in the inference phase.

[0081] In step b., for each predicted feature region that overlaps with the corresponding ground truth region, in other words, for scenarios #7 and #8 in Table 1, a similarity coefficient is generated. In scenarios #7 and #8, there is no ambiguity because the defect has been recognized by machine learning model 2 and the process of generating the ignored regions.

[0082] The ignored regions enable the computer system to distinguish the false positives in scenario #5 (where the predicted feature region A7 does not overlap with any of the ignored regions in the ignored regions and does not overlap with any of the ground truth regions in the ground truth regions) from the ambiguous situation in scenario #6 (where the predicted feature region A6 overlaps with the corresponding ignored region and does not overlap with any of the ground truth regions in the ground truth regions).

[0083] In the false positive case of scenario #5, the training engine 3 trains the machine learning model 2 based on the predicted feature region A7, for example, by training the machine learning model 2 to classify the predicted feature region A7 as the background class and / or by training the machine learning model 2 to forget the predicted feature region A7.

[0084] In the ambiguous case of scenario #6, the training engine 3 ignores the predicted feature region A6. Therefore, the predicted feature region A6 is not used to train the machine learning model 2, unlike in scenario #5. That is, the predicted feature region A6 is not used to forget the predicted feature region A6, and it is not used to train the machine learning model 2 to classify the predicted feature region A6 as the background.

[0085] Now, different ways of generating the loss value in step c. will be described.

[0086] In one example, the loss value for each training round can be calculated in step c. by the following equation (1):

[0087]

[0088] where, l box (j) is the loss value for the j-th training round (the loss value is updated via the training update path for each training round); iou i is the Jaccard index of the i-th predicted feature region in the set 60 of predicted feature regions; iou j thresh is the threshold for the j-th training round; and N is the total number of predicted feature regions in the set 60 of predicted feature regions.

[0089] Optionally, for all training rounds, the threshold iou j threshIt can be set to the value 1, in which case the loss value simply takes the average of 1 and the IoU. The backpropagation algorithm will continuously attempt to improve the parameters of the machine learning model 2 to ensure that the loss value is zero, which means that all predicted feature regions are precisely aligned with the ground truth regions.

[0090] More preferably, the threshold iou j thresh is set to an intermediate value (i.e., a value less than 1 and greater than 0). In this case, some of the similarity coefficients in the similarity coefficient iou i will be less than the threshold (so the difference is positive), while other similarity coefficients will be greater than the threshold (so the difference is negative). In equation (1), the loss value is based on the absolute difference between the threshold and each similarity coefficient.

[0091] In the following discussion, similarity coefficients less than the threshold are referred to as "low similarity coefficients", while similarity coefficients greater than the threshold are referred to as "high similarity coefficients".

[0092] Within the complete set of training rounds, a mixture of low and high similarity coefficients will be generated. Generally, in one or more of the training rounds, the set of similarity coefficients will include one or more low similarity coefficients, and in one or more of the training rounds, the set of similarity coefficients will include one or more high similarity coefficients.

[0093] Depending on the value of the threshold, in some training rounds (e.g., earlier training rounds), only low similarity coefficients may be generated, in other training rounds, only high similarity coefficients may be generated, and in other training rounds, a mixture of low and high similarity coefficients may be generated.

[0094] Optionally, the threshold varies between at least two of the training rounds. For example, in a set of 100 training rounds, the threshold can be initially set to 1.0, then decreased until it reaches 0.8 at the 66th training round, and then fixed at 0.8 until the last (100th) training round.

[0095] The threshold can be decreased in each round until the 66th training round, at which point it remains fixed at 0.8 until the last (100th) training round. In this example, the threshold can be decreased linearly (e.g., by the same amount from one round to the next) or non-linearly (e.g., following a cosine curve).

[0096] Instead of decreasing in each round, the threshold can be decreased in a stepwise manner, e.g., every five rounds, e.g.: 1.0, 1.0, 1.0, 1.0, 1.0, 0.992, 0.992, 0.992, 0.992, 0.992, etc.

[0097] In another example, the threshold can start at a mid-value (e.g., 0.8), then increase to 1.0, and then decrease.

[0098] Suppose the threshold is set to 0.7, and during a particular training round there are ten predicted feature regions to be learned, and the Jaccard indices of these ten predicted feature regions are 0.1, 0.3, 0.7, 0.9, 0.88, 0.95, 0.45, 0.67, 0.85, 0.58. Then the loss calculation would be:

[0099] Loss = (0.6 + 0.4 + 0.0 + 0.2 + 0.18 + 0.25 + 0.25 + 0.03 + 0.15 + 0.12) / 10

[0100] In this case the loss value would be 2.18 / 10 = 0.218.

[0101] Experiments have shown that setting the threshold to a mid-value results in more accurate performance of Machine Learning Model 2 compared to a baseline where the threshold is set to 1.

[0102] In Equation (1), both the high similarity coefficient and the low similarity coefficient are used to determine the loss value. In other embodiments, the high similarity coefficient and the low similarity coefficient can be treated differently. In other embodiments, the loss value can be based on Equations (2) and (3):

[0103] If iou >= iou th , then iou = 1 (2)

[0104] Loss = ∑(1 - iou) / N (3)

[0105] According to Equation (2), whenever the Jaccard index crosses the threshold iou th it is clamped to 1, such that the high similarity coefficient does not contribute to the loss value in Equation (3).

[0106] Thus, the loss value from Equation (3) is based on the difference between the threshold and each low similarity coefficient, and any high similarity coefficients are ignored (by clamping them to 1).

[0107] Depending on the amount of fuzziness, the threshold iou th can be a user-defined value. It can be set to a value below and close to 1, such as 0.7, 0.8, 0.9, etc. Thus, any prediction with a low similarity coefficient where iou < iou th will be fed back for optimization while other predictions are ignored.

[0108] Assume that during a specific training round, there are ten predicted feature regions to be learned, and the Jaccard indices of these ten predicted feature regions are 0.1, 0.3, 0.7, 0.9, 0.88, 0.95, 0.45, 0.67, 0.85, 0.58. Then the usual loss calculation will take the average of all ten Jaccard indices:

[0109] Loss = (0.9 + 0.7 + 0.3 + 0.1 + 0.12 + 0.05 + 0.55 + 0.33 + 0.15 + 0.42) / 10

[0110] In this case, the loss value will be 3.62 / 10 = 0.362.

[0111] Assume the threshold iou th is set to 0.7. Then, modifying the loss value through Equation (2) results in a modified loss value:

[0112] Loss = (0.9 + 0.7 + 0.0 + 0.0 + 0.0 + 0.0 + 0.55 + 0.33 + 0.0 + 0.42) / 10

[0113] The modified loss value will be lower than the average value of 0.362, i.e., it will be 2.9 / 10 = 0.29.

[0114] In summary, two alternative methods for determining the loss value for training the machine learning model 2 are described above: the first method using Equation (1) and the second method using Equation (2) and Equation (3). Both methods determine the loss value based on the similarity coefficient and the threshold, which avoids training the machine learning model 2 to generate predicted feature regions that are precisely aligned with the ground truth regions. This avoids "overtraining" the machine learning model 2 based on fuzzy features by not strictly converging the predictions to the perfect IoU value of 1.

[0115] The above-described embodiments include two methods. The first method illustrated in Table 1 uses the ignore region to ignore fuzzy predictions. The second method illustrated in Equation (1) and Equation (2) uses the threshold to determine the loss value that is less likely to cause "overtraining" of the machine learning model.

[0116] As described above, these two methods can be implemented together or separately. That is, the first method can be implemented with a conventional loss value based on the average of 1 to IoU of all similarity coefficients, as in Equation (4) below.

[0117] Loss = ∑(1 - iou) / N (4)

[0118] Conversely, as in Table 1, the second method can be implemented in a conventional training process without using the ignore region to ignore fuzzy predictions.

[0119] Where the term "or" appears, this is to be construed as meaning "and / or", such that the items involved are not necessarily mutually exclusive and may be used in any suitable combination.

[0120] Although the present invention has been described above with reference to one or more preferred embodiments, it should be understood that various changes or modifications can be made without departing from the scope of the invention as defined in the appended claims.

Claims

1. A method for training a machine learning model to recognize image features, the method comprising: a. receiving a set of predicted feature regions from the machine learning model, each predicted feature region comprising a prediction of a feature in the training image data; b. generating a set of similarity coefficients, each similarity coefficient indicating the similarity between a predicted feature region and a corresponding true value region overlapping with the predicted feature region; c. determining a loss value based on the similarity coefficient and a threshold; and d. training the machine learning model based on the loss value, Wherein, a. to d. are repeated, and each repetition includes a corresponding training round; In one or more of the training rounds, the set of similarity coefficients includes one or more similarity coefficients that are less than a threshold, and the loss value is based on a difference between the threshold and each similarity coefficient that is less than the threshold; and In one or more of the training rounds, the set of similarity coefficients includes similarity coefficients greater than the threshold.

2. The method according to claim 1, wherein: In one or more of the training rounds, the set of similarity coefficients includes one or more similarity coefficients greater than the threshold, and the loss value is based on a difference between the threshold and each similarity coefficient greater than the threshold.

3. A method according to any preceding claim, wherein: The threshold value varies between at least two of the training rounds.

4. The method according to claim 3, wherein: The threshold value decreases between at least two of the training rounds.

5. A method according to any preceding claim, wherein: The training image data includes a series of images of an object each containing the same features observed from different viewing angles.

6. The method of claim 5, further comprising generating the training image data by imaging the object from a series of different viewing angles.

7. The method according to claim 6, wherein: The object is imaged using visible light.

8. A method according to any preceding claim, wherein: Each feature includes surface defects.

9. A method according to any preceding claim, wherein: Each feature comprises a surface defect of the aircraft.

10. A method according to any preceding claim, wherein: Each feature includes an indentation.

11. A method according to any preceding claim, wherein: The similarity coefficient is the Jaccard index.

12. The method of any preceding claim, further comprising generating the ground truth regions by displaying the training image data to a human annotator, and receiving the ground truth regions as input from the human annotator, each ground truth region comprising an annotation of a feature in the training image data.

13. A computer system configured to train a machine learning model by the method of any preceding claim.

14. Computer software configured to train a machine learning model by the method of any one of claims 1 to 12.