A YOLO-based automatic detection method for aircraft tank excess
Patent Information
- Application Number
- CN202311370535.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-10-20
AI Technical Summary
飞机的内部多余物清除是一项任务量大、要求精度高的工作,而在航空部件的密闭、狭小、复杂的空间内,多余物具有体积小、形状非规则、散落位置随机等特点,清除工作非常困难,传统的手工清除无法满足飞机检修维护对精度和效率的要求
[0012]由上述本发明提供的技术方案可以看出,上述方法利用YOLO目标检测算法对飞机油箱内部进行自动检测,能够高效、准确地检测出油箱内部的多余物,提升了飞机油箱的安全性和可靠性,从而提高了工作效率。
Smart Images

Figure CN117612081B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an automatic detection method for excess materials in aircraft fuel tanks based on YOLO (You Only Look Once). Background Technology
[0002] During the maintenance of aircraft components, various types of debris, such as chips and rivets, are generated. If these debris are not removed in a timely manner, they will seriously affect the normal operation of the aircraft and increase the risk of accidents. Therefore, the removal of debris from aircraft is extremely important. Removing debris from the interior of an aircraft is a large-scale and highly precise task. Within the confined, narrow, and complex spaces of aircraft components, debris is characterized by its small size, irregular shape, and random location, making removal very difficult. Traditional manual removal methods cannot meet the precision and efficiency requirements of aircraft maintenance.
[0003] Currently, there is no effective solution to the problem of how to efficiently and accurately detect foreign objects from aircraft fuel tank image data. Therefore, it is necessary to develop an advanced automated detection and removal technology to improve the quality of aircraft maintenance. Summary of the Invention
[0004] The purpose of this invention is to provide an automatic detection method for foreign objects in aircraft fuel tanks based on YOLO. This method uses the YOLO target detection algorithm to automatically detect the interior of aircraft fuel tanks, which can efficiently and accurately detect foreign objects inside the fuel tanks, improve the safety and reliability of aircraft fuel tanks, and thus improve work efficiency.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] An automatic detection method for foreign matter in aircraft fuel tanks based on YOLO, the method comprising:
[0007] Step 1: After attaching the binocular endoscope camera to the arm robot, place it inside the aircraft fuel tank to take fixed-point pictures, collect image datasets of foreign objects inside the aircraft fuel tank, and label the obtained image datasets.
[0008] Step 2: Preprocess the acquired image dataset. The preprocessing process includes image resizing, data partitioning, image data cleaning, and image data enhancement.
[0009] Step 3: Then, construct a neural network model suitable for detecting foreign objects inside aircraft fuel tanks, and design the loss function of the neural network model;
[0010] Step 4: Use the preprocessed image data from Step 2 to train the constructed neural network model, and optimize the neural network model based on the training and testing results;
[0011] Step 5: Use the neural network model trained and optimized in Step 4 to detect the actual target image set, thereby realizing the automatic detection of excess objects in the aircraft fuel tank.
[0012] As can be seen from the technical solution provided by the present invention, the above method uses the YOLO target detection algorithm to automatically detect the inside of the aircraft fuel tank, which can efficiently and accurately detect foreign objects inside the fuel tank, improve the safety and reliability of the aircraft fuel tank, and thus improve work efficiency. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the automatic detection method for excess material in aircraft fuel tanks based on YOLO, provided in an embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments, and do not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0016] like Figure 1 This is a schematic flowchart of an automatic detection method for foreign matter in aircraft fuel tanks based on YOLO, provided in an embodiment of the present invention. The method includes:
[0017] Step 1: After mounting the binocular endoscope, place it inside the aircraft fuel tank to take fixed-point pictures, collect image datasets of foreign objects inside the aircraft fuel tank, and label the obtained image datasets.
[0018] In this step, before taking pictures using the binocular endoscope, the binocular endoscope needs to be calibrated. The specific process is as follows:
[0019] First, create two folders to store calibration images taken using the left and right lenses of a binocular endoscope;
[0020] Set the resolution and automatic photo capture interval parameters of the binocular endoscope according to your needs;
[0021] Prepare a black and white checkered calibration board for shooting. During the shooting process, the angle of the calibration board should be continuously adjusted to ensure that various angles can be captured, but the offset angle in each direction should not exceed 45 degrees, and the grid inside the board should appear in the left and right images captured by the binocular endoscope. At the same time, ensure sufficient light to ensure image clarity.
[0022] Import the captured calibration images into the stereo camera calibrator toolbox in MATLAB, set the calibration plate grid size to 6mm, and after calibration, export and update the corresponding intrinsic parameter information into the Python code. The intrinsic parameter information includes: intrinsic matrix of left and right cameras; radial distortion coefficients of left and right cameras; tangential distortion coefficients of left and right cameras; rotation matrix; and translation matrix.
[0023] In practice, the collected image dataset includes images of unwanted objects from different angles, distances, densities, and types. Damaged or blurry images are deleted after shooting. For example, there are approximately twenty types of unwanted objects, including: yellow tape, steel tape measure, measuring tape, wire harness, wrench, screwdriver, pliers, bolt, nut, washer, rag, etc. Approximately 2000 images were taken for each type of unwanted object.
[0024] The process of annotating the obtained image dataset is as follows:
[0025] The LabelImg tool was used to annotate each of the filtered image datasets. The annotations included the location, size, and category of the extraneous objects in the aircraft fuel tanks. Each image pair had a corresponding image label.
[0026] LabelImg is a tool for annotating target detection data. It can generate two formats through annotation: 1. VOC label format, where the labeled labels are stored in an XML file; the key information in the XML includes: the image name, and the coordinates of the bounding box for each target: namely, the coordinates of the top left corner and the bottom right corner (xmin, ymin, xmax, ymax); 2. YOLO label format, where the labeled labels are stored in a TXT file; the information in the TXT file includes: each line represents a labeled target and its information: the first column represents the label of the labeled target, and the following four numbers represent the center coordinates of the bounding box and the relative width and height of the bounding box, respectively.
[0027] Step 2: Preprocess the acquired image dataset. The preprocessing process includes image resizing, data partitioning, image data cleaning, and image data enhancement.
[0028] In this step, the preprocessing of the acquired image dataset specifically involves:
[0029] Image resizing includes: adjusting image dimensions (since YOLO models typically require input images to have fixed dimensions, the width and height of all images need to be adjusted to the same value for subsequent processing); maintaining aspect ratio (when resizing images, the original aspect ratio is usually maintained to prevent image distortion). This means that if the image size needs to be reduced to fit the model requirements, the width and height will be reduced proportionally or padding will be added to maintain consistent image dimensions; and normalization (some YOLO variants require image pixel values to be normalized to between 0 and 1). Therefore, the range of image pixel values from 0 to 255 is usually normalized to between 0 and 1.
[0030] The data partitioning operation includes: dividing the labeled image dataset into training, validation, and test sets for subsequent training, tuning, and model performance evaluation; since the images before and after data augmentation are highly correlated, the training, validation, and test sets are divided in a 7:1:2 ratio, and then data augmentation is performed on the partitioned dataset, and the few redundant object category images are augmented to ensure the robustness of the neural network model.
[0031] Image cleaning operations include: denoising and cropping the acquired image data. Cropping reduces the image size, thereby reducing computational complexity and improving data quality.
[0032] Image data augmentation operations include: rotating, flipping, scaling and translating the acquired image data; performing regional elastic deformation data augmentation on the divided training set; generating more training data; and improving the robustness of the model.
[0033] The steps for performing regional elastic deformation data augmentation on the divided training set are as follows:
[0034] First, the training set image is randomly divided into several sub-regions; then, random elastic deformation is performed on each sub-region, that is, the pixels in each sub-region are randomly shifted, scaled, and rotated; finally, the transformed sub-regions are stitched together to form a new image, which serves as the data-augmented training set image.
[0035] The purpose of this is to increase the diversity of the training set and improve the generalization ability of the model. The above preprocessing operations can improve the accuracy of the model and reduce the computational complexity.
[0036] Step 3: Then, construct a neural network model suitable for detecting foreign objects inside aircraft fuel tanks, and design the loss function of the neural network model;
[0037] In this step, the selection of the network model should consider the following factors:
[0038] (1) Types and sizes of extraneous objects in the aircraft fuel tank to be detected: The selected model should be able to accurately detect the type and size of the target object. For oddly shaped or small objects, a model that can handle small targets should be selected.
[0039] (2) Response time and accuracy of detection: It is necessary to quickly detect objects inside the aircraft fuel tank to ensure that there are no extraneous objects inside the fuel tank. A model with fast speed and high accuracy should be selected.
[0040] (3) Training data and model size: Select a model size that is appropriate for the size of the training dataset and computing resources;
[0041] (4) Others: such as multi-target detection, real-time detection, etc.
[0042] This embodiment demonstrates the construction of a YOLO v5 neural network model for detecting foreign objects inside aircraft fuel tanks. This YOLO v5 neural network model boasts advantages such as high accuracy, fast speed, simplicity, strong adaptability, and good robustness. There are four versions of YOLO v5: YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x.
[0043] The YOLO v5 neural network model comprises four parts: an input layer, a backbone network, a neck network, and a detection layer. A C3SE attention mechanism is added to the backbone network. This C3SE attention mechanism combines the following three key aspects of attention to enhance the expressive power of the feature map: 1. Channel-wise Attention: This focuses on the importance of each channel in the feature map. Different channels capture different features of the image. By learning weights, C3SE can enhance the response to important channels and suppress the response to unimportant channels; 2. Spatial Attention: This focuses on the spatial location of the feature map. It allows the network to focus on specific regions in the image, helping to locate and capture the spatial location information of the target; 3. Spectral Attention: This focuses on the spectral characteristics of the feature map. It can be used to process frequency domain information in the image, such as texture and patterns.
[0044] The C3SE attention mechanism is used to improve the model's detection performance for small targets and partially occluded targets. In the specific implementation, the attention structure code can be put into a Python file, and the four C3 modules in the backbone network can be changed to C3SE.
[0045] Furthermore, during the loss function design process, because DIOU_loss converges faster during model training, the original GIOU_loss used to calculate the target box regression loss function in the network is replaced with DIOU_loss as the loss function for predicting the box. DIOU_loss (Distance IoU Loss) is a target box regression loss function that adds a distance metric to GIOU_loss. The expression for DIOU_loss is as follows:
[0046]
[0047] Where IoU is the Intersection over Union (IoU) between the predicted bounding box and the ground truth bounding box, which is the same as GIOU; d is the Euclidean distance between the center points of the two boxes; and c is the radius of the minimum circumcircle used to normalize the distance.
[0048] The introduction of DIOU_loss is mainly to solve the problem of misaligned target boxes, and further improves the regression loss function by adding a distance metric.
[0049] Step 4: Use the preprocessed image data from Step 2 to train the constructed neural network model, and optimize the neural network model based on the training and testing results;
[0050] The optimization process includes: data augmentation, gradient accumulation, learning rate adjustment, and loss function adjustment.
[0051] The specific process is as follows:
[0052] First, the training set is augmented and then input into the constructed neural network model to obtain the object detection prediction value;
[0053] The loss function value is calculated using the actual and predicted object detection values in the training set. Specifically, it is calculated using the Mean Squared Error Loss (MSE), a commonly used loss function in regression tasks. MSE measures the average squared error between the model's predicted and actual values. The formula for calculating MSE is as follows:
[0054]
[0055] Where n is the number of samples; x i y is the true value of the i-th sample; i It is the model's prediction for the i-th sample; the smaller the MSE value, the smaller the difference between the model's prediction and the true value, and the better the model's performance.
[0056] The network model parameters are updated based on the loss function value MSE;
[0057] Then, the test set is augmented and input into the constructed neural network model to obtain the target detection prediction value;
[0058] Next, the loss function value (using MSE, formula as above) and test set accuracy are calculated using the true and predicted object detection values in the test set. Test set accuracy is calculated by inputting samples from the test set into the trained neural network model and then comparing the model's predictions with the true labels in the test set. Typically, accuracy is the number of correctly classified samples divided by the total number of samples in the test set, calculated as: Accuracy = Total number of samples in the test set ÷ Number of correctly classified samples.
[0059] Determine whether the obtained test set accuracy is greater than the maximum accuracy M. If so, save the neural network model and update the maximum accuracy M. If not, neither save the neural network model nor update the maximum accuracy M. Here, the maximum accuracy M represents the minimum threshold of accuracy.
[0060] The process involves determining whether the neural network model has converged. If it has, proceed to the next step; otherwise, reduce the learning rate and continue training. The learning rate is adjusted by modifying parameters in the Python code. Reducing the learning rate uses a fixed learning rate scheduling strategy: a predefined learning rate schedule is used, and the learning rate is decreased at regular training epochs or based on changes in model performance. A common strategy is learning rate decay, which gradually reduces the learning rate.
[0061] Determine if the maximum number of training rounds has been reached. If so, output the trained neural network model and use it as the object detection model.
[0062] Furthermore, the process described above—calculating the loss function value using the true and predicted object detection values in the training set, and updating the network model parameters based on the loss function value—is as follows:
[0063] The YOLO v5 neural network model performs forward propagation on the input training set images: the input images pass through a series of convolutional layers, activation functions and pooling layers in the forward propagation process of the neural network to gradually extract feature information from the images. YOLO v5 uses a backbone network to implement this step, such as CSPDarknet53, and finally obtains the results of redundant object detection, including bounding box position, class probability and bounding box confidence.
[0064] Based on the labeled data, the bounding box of each redundant object is mapped to its corresponding grid cell. The overlap between the predicted bounding box and the ground truth bounding box (such as IoU) is calculated to determine which object each grid cell is responsible for predicting.
[0065] The predicted bounding box is the bounding box of the target predicted by the neural network model based on the input image and the parameters learned during training. It usually includes the target's location (x, y coordinates, width, height), class probability (indicating the probability that the object belongs to different classes), and bounding box confidence (indicating the confidence that the bounding box contains the object). The ground truth bounding box is the annotation information provided during training, representing the bounding box of the actual object in the image. The information of the ground truth bounding box is usually provided by the annotator of the dataset, including the target's location and class.
[0066] The degree of overlap is usually calculated using a metric called Intersection over Union (IoU), which is calculated as follows:
[0067] IoU=Area of Overlap / Area of Union
[0068] Among them, "Area of Overlap" is the area of the region where the predicted bounding box and the ground truth bounding box intersect; "Area of Union" is the area of their union; when calculating IoU, the IoU value is usually limited to between 0 and 1.
[0069] For each predicted bounding box, a loss function is typically used to calculate the difference between the predicted and ground truth bounding boxes. For example, the positional difference can be calculated using the squared error (MSE); the class difference can be calculated using cross-entropy loss.
[0070] The formula for calculating cross-entropy loss is:
[0071] L = -∑yi*log(pi)
[0072] Where yi represents the true label; pi represents the probability value predicted by the model;
[0073] Since the YOLOv5 neural network model makes predictions at different scales, a weighted adjustment is used to balance the size of redundant objects at different scales to adjust the loss contribution at different scales. Specifically, the loss of all grid cells is weighted and summed to obtain the final loss function, which is expressed as a weighted sum, as shown below:
[0074] Loss=λ loc ·LocalizationLoss+λ cls ·ClassLoss+λ conf ·ConfidenceLoss+λ iou ·IoULoss
[0075] Among them, Localization Loss measures the difference between the predicted bounding box location and the true bounding box location, including the center coordinates, width, and height of the bounding box, and is usually expressed as Squared Error (MSE) or Smooth L1 Loss; Class Loss measures the difference between the predicted class and the true class, and is usually expressed as Cross Entropy Loss; Confidence Loss measures the difference between the bounding box confidence and the true situation; and IoU Loss encourages the model to better predict the location of the bounding box.
[0076] By weighting the loss terms of the above components, object sizes at different scales can be taken into account; λ loc , λ cls , λ conf and λ iouThese are weight parameters used to adjust the importance of each loss term. The specific values of the weight parameters are adjusted according to the task and network structure.
[0077] Then, the backpropagation algorithm is used to update the network model parameters in order to minimize the loss function;
[0078] The backpropagation algorithm is a fundamental algorithm for updating neural network model parameters. Specifically, it calculates the gradient of the loss function with respect to the model parameters, and then uses gradient descent or its variants to update the parameters to reduce the value of the loss function. The specific process is as follows:
[0079] 1. Forward Propagation: First, the input data is fed into the neural network, and forward propagation is performed to calculate the output value of the model. This includes calculating the activation value of each layer, such as using activation functions like ReLU, Sigmoid, etc.
[0080] 2. Calculate the loss function (Compute Loss): Using the model's output values and the true labels, calculate the value of the loss function. The loss function is usually a function of the model parameters and is used to measure the model's performance.
[0081] 3. Backpropagation Gradients: Starting with the loss function, the gradient of each model parameter with respect to the loss function is calculated using the chain rule. This is the core step of backpropagation. For each parameter θ, its gradient can be expressed as:
[0082] in, This represents the partial derivative of the loss function with respect to the parameter θ. This partial derivative can be estimated using gradient calculation methods, such as numerical gradient or automatic differentiation.
[0083] 4. Parameter Update: Using the calculated parameter gradients, the model parameters can be updated using gradient descent or its variants. The general form of parameter update is as follows: Where θ is the parameter to be updated, and θ is the learning rate; It's the gradient. The learning rate controls the step size for parameter updates;
[0084] 5. Repeat Training: Repeat the process of forward propagation, calculating the loss function, backpropagating gradients, and updating parameters until the loss function converges or the stopping condition is met.
[0085] Step 5: Use the neural network model trained and optimized in Step 4 to detect the actual target image set, thereby realizing the automatic detection of excess objects in the aircraft fuel tank.
[0086] In this step, based on the optimized neural network model, the input data is the actual set of target images (e.g., images of extraneous objects in an airplane fuel tank); the output detection results include: 1. Bounding box coordinates: For each detected extraneous object, it provides its location information, usually represented by the coordinates of the upper left and lower right corners; 2. Extraneous object category: Identifies the category to which each detected extraneous object belongs, such as extraneous objects in an airplane fuel tank; 3. Confidence score: Each detection result is also accompanied by a confidence score, which represents the model's confidence in the detection result and can be used to filter out detection results with high confidence.
[0087] The detection results are transformed into a visual form, specifically by marking the location of each extraneous object with a rectangle and providing a corresponding text description. The detection results are displayed on a monitor or presented through audio broadcast, improving detection efficiency and convenience, so that users or the system can further analyze or take appropriate actions.
[0088] It is worth noting that the contents not described in detail in the embodiments of the present invention belong to the prior art known to those skilled in the art.
[0089] In summary, the method described in this embodiment of the invention inherits the high detection speed and accuracy of the YOLO network model, and has good accuracy for small target objects or different shapes of the same object. In addition, it can achieve real-time detection by running pre-programmed real-time shooting code, and the neural network model structure can also be adjusted according to the dataset.
[0090] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. An automatic detection method for foreign matter in aircraft fuel tanks based on YOLO, characterized in that, The method includes: Step 1: After mounting the binocular endoscope, place it inside the aircraft fuel tank to take fixed-point pictures, collect image datasets of foreign objects inside the aircraft fuel tank, and label the obtained image datasets. Step 2: Preprocess the acquired image dataset. The preprocessing process includes image resizing, data partitioning, image data cleaning, and image data enhancement. Step 3: Then, construct a neural network model suitable for detecting foreign objects inside aircraft fuel tanks, and design the loss function of the neural network model; In step 3, a YOLO v5 neural network model is specifically constructed to detect foreign objects inside the aircraft fuel tank. The YOLO v5 neural network model comprises four parts: an input layer, a backbone network, a neck network, and a detection layer. A C3SE attention mechanism is added to the backbone network to improve the model's detection performance for small targets and partially occluded targets. The C3SE attention mechanism includes: channel attention (focusing on the importance of each channel in the feature map); spatial attention (focusing on the spatial location in the feature map); and spectral attention (focusing on the spectral characteristics of the feature map). In the loss function design process, because DIOU_loss converges faster during model training, the original GIOU_loss used to calculate the target box regression loss function in the network is replaced with DIOU_loss as the loss function for predicting the box. DIOU_loss is a target box regression loss function that adds a distance metric to GIOU_loss. The expression for DIOU_loss is as follows: ; Where IoU is the intersection-union ratio between the predicted bounding box and the ground truth bounding box, which is the same as GIOU; d is the Euclidean distance between the center points of the two boxes; and c is the radius of the minimum circumcircle used to normalize the distance. Step 4: Use the preprocessed image data from Step 2 to train the constructed neural network model, and optimize the neural network model based on the training and testing results; Step 5: Use the neural network model trained and optimized in Step 4 to detect the actual target image set, thereby realizing the automatic detection of excess objects in the aircraft fuel tank.
2. The automatic detection method for excess material in aircraft fuel tanks based on YOLO according to claim 1, characterized in that, In step 1, before taking pictures using the binocular endoscope, the binocular endoscope needs to be calibrated. The specific process is as follows: First, create two folders to store calibration images taken using the left and right lenses of a binocular endoscope; Set the resolution and automatic photo capture interval parameters of the binocular endoscope according to your needs; Prepare a black and white checkered calibration board for shooting. During the shooting process, the angle of the calibration board should be continuously adjusted to ensure that various angles can be captured, but the offset angle in each direction should not exceed 45 degrees, and the grid inside the board should appear in the left and right images captured by the binocular endoscope. At the same time, ensure sufficient light to ensure image clarity. Import the captured calibration images into the Stereo Camera calibrator toolbox in MATLAB, set the calibration plate grid size to 6mm, and after calibration, export and update the corresponding intrinsic parameter information into the Python code. The intrinsic parameter information includes: left and right camera intrinsic parameters; left and right camera radial distortion coefficients; left and right camera radial distortion coefficients; rotation matrix; and translation matrix.
3. The automatic detection method for excess material in aircraft fuel tanks based on YOLO according to claim 1, characterized in that, In step 1, the acquired image dataset includes images of redundant objects from different angles, distances, densities, and types, and invalid images that are damaged or blurred are deleted from the image dataset after the images are taken. The process of annotating the obtained image dataset is as follows: The LabelImg tool was used to annotate each of the filtered image datasets. The annotations included the location, size, and category of the extraneous objects in the aircraft fuel tanks. Each image pair had a corresponding image label.
4. The automatic detection method for excess material in aircraft fuel tanks based on YOLO according to claim 1, characterized in that, In step 2, the preprocessing of the acquired image dataset is specifically as follows: Image resizing includes: adjusting image dimensions to make the width and height of all images the same for subsequent processing; maintaining aspect ratio to prevent image distortion; and normalization to normalize the pixel values of the image from the range of 0 to 255 to between 0 and 1. The data partitioning operation includes: dividing the labeled image dataset into training, validation, and test sets for subsequent training, tuning, and model performance evaluation; specifically, the training, validation, and test sets are divided in a 7:1:2 ratio, and then data augmentation is performed on the partitioned dataset, and the fewer redundant object category images are augmented to ensure the robustness of the neural network model. Image cleaning operations include: denoising and cropping the acquired image data. Cropping reduces the image size, thereby reducing computational complexity and improving data quality. Image data augmentation operations include: rotating, flipping, scaling and translating the acquired image data, performing regional elastic deformation data augmentation on the divided training set, generating more training data to improve the robustness of the model; The process of performing regional elastic deformation data augmentation on the divided training set is as follows: First, the training set image is randomly divided into several sub-regions; then, random elastic deformation is performed on each sub-region, that is, the pixels in each sub-region are randomly shifted, scaled, and rotated; finally, the transformed sub-regions are stitched together to form a new image, which serves as the data-augmented training set image.
5. The automatic detection method for excess material in aircraft fuel tanks based on YOLO according to claim 4, characterized in that, The process of step 4 is as follows: First, the training set is augmented and then input into the constructed neural network model to obtain the object detection prediction value; The loss function value is calculated using the true and predicted object detection values in the training set. Specifically, it is calculated using the mean squared error (MSE) loss, and the formula for calculating MSE is as follows: ; Where n is the number of samples; It is the true value of the i-th sample; It is the model's prediction for the i-th sample; the smaller the MSE value, the smaller the difference between the model's prediction and the true value, and the better the model's performance. Update the network model parameters based on the loss function value; Then, the test set is augmented and input into the constructed neural network model to obtain the target detection prediction value; Then, the loss function value and test set accuracy are calculated using the true values and predicted values of object detection in the test set. The test set accuracy is calculated by inputting the samples in the test set into the trained neural network model and then comparing the model's prediction results with the true labels in the test set. The formula is: Accuracy = Number of correctly classified samples ÷ Total number of samples in the test set. Determine whether the obtained test set accuracy is greater than the maximum accuracy M. If yes, save the neural network model and update the maximum accuracy M; otherwise, neither save the neural network model nor update the maximum accuracy M. Determine if the neural network model has converged. If it has, proceed to the next step; otherwise, reduce the learning rate and continue training. The reduction of the learning rate uses a fixed learning rate scheduling strategy: a learning rate scheduling plan is defined in advance, and the learning rate is reduced every certain number of training rounds or according to changes in model performance. Determine if the maximum number of training rounds has been reached. If so, output the trained neural network model and use it as the object detection model.
6. The automatic detection method for excess material in aircraft fuel tanks based on YOLO according to claim 5, characterized in that, The process of calculating the loss function value using the true and predicted object detection values in the training set, and updating the network model parameters based on the loss function value, is as follows: The YOLO v5 neural network model is used to perform forward propagation on the input training set images to obtain the results of redundant object detection, including bounding box location, class probability, and bounding box confidence. Based on the labeled data, the bounding box of each redundant object is mapped to its corresponding grid cell. The degree of overlap between the predicted bounding box and the actual bounding box is calculated to determine which object each grid cell is responsible for predicting. The predicted bounding box is the target bounding box predicted by the neural network model based on the input image and the parameters learned during training. It includes the target's location, class probability, and bounding box confidence. The ground truth bounding box is the annotation information provided during training, representing the bounding box of the actual object in the image. It is provided by the annotator of the dataset and includes the target's location and class. The degree of overlap is calculated using a metric called the Intersection over Union (IoU), which is calculated as follows: ; Among them, "Area of Overlap" is the area of the region where the predicted bounding box and the ground truth bounding box intersect; "Area of Union" is the area of their union; when calculating IoU, the IoU value is usually limited to between 0 and 1. For each predicted bounding box, a loss function is used to calculate the difference between the predicted bounding box and the ground truth bounding box; the class difference between the two is calculated using cross-entropy loss, where the formula for cross-entropy loss is: ; Where yi represents the true label; pi represents the probability value predicted by the model; Weighted adjustment is used to adjust the loss contribution at different scales. Specifically, the final loss function is obtained by weighted summation of the losses of all grid cells, as shown below: ; Among them, Localization Loss measures the difference between the predicted bounding box location and the true bounding box location, including the center coordinates, width, and height of the bounding box; Class Loss measures the difference between the predicted class and the true class; Confidence Loss measures the difference between the bounding box confidence and the true situation; and IoU Loss encourages the model to predict the location of the bounding box. , , and These are weight parameters used to adjust the importance of each loss term. The specific values of the weight parameters are adjusted according to the task and network structure. Then, the backpropagation algorithm is used to update the network model parameters to minimize the loss function. The backpropagation algorithm is a basic algorithm for updating the parameters of a neural network model. Specifically, it calculates the gradient of the loss function with respect to the model parameters and then uses gradient descent to update the parameters to reduce the value of the loss function.
7. The automatic detection method for foreign matter in aircraft fuel tanks based on YOLO according to claim 1, characterized in that, In step 5, based on the optimized neural network model, the input data is the actual set of target images collected; The output detection results include: 1) Bounding box coordinates: For each detected redundant object, provide its position information, represented by the coordinates of the top left and bottom right corners; 2) Redundant Object Category: Identifies the category to which each detected redundant object belongs; 3) Confidence score: Each detection result is also accompanied by a confidence score, which represents the model's confidence in the detection result and is used to filter out detection results with high confidence. The detection results are transformed into a visual form, specifically by marking the location of each extraneous object with a rectangle and providing a corresponding text description; The test results are displayed on a monitor or presented via audio broadcast.
Citation Information
Patent Citations
Aircraft fuel tank remainder automatic detection method based on deep learning
CN111145239A
Method and system for automated target recognition
US20220284703A1