Aircraft power distribution equipment fastener wrong and neglected loading detection method related to position types
Through rotating table and multi-angle light acquisition, combined with exposure fusion and denoising diffusion processing, the YOLO11-AEDSF model was constructed, which solved the detection accuracy and insufficient sample data in the detection of fault-missing assembly defects of aircraft distribution equipment, and achieved efficient and accurate detection of fault-missing assembly of fasteners in fasteners in aircraft distribution equipment.
Patent Information
- Application Number
- CN202510358547.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing deep learning object detection methods cannot effectively identify missed assembly defects in the fastener detection of aircraft distribution equipment, and the imaging angle and lighting conditions interfere with the detection results, insufficient sample data and category imbalance, resulting in insufficient detection accuracy and generalization capabilities.
Through the detection device coordinated by the rotary table, images of multiple angles and different light intensities are collected, exposure fusion and denoising diffusion are performed, and the YOLO11-AEDSF model is constructed, combining the C3k2 Faster CGLU module and Focaler-SIoU loss function are used to detect fastener error-missing assembly defects.
It realizes efficient, complete and accurate detection under multi-view conditions, improves detection accuracy and calculation efficiency, solves the problem that a single object detection model cannot verify the correspondence between fastener position and type, and enhances the detection ability of small sample targets.
Smart Images

Figure CN120298774A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of surface assembly defect detection of aircraft power distribution equipment, and more specifically, relates to a method for detecting missing or misassembled fasteners of aircraft power distribution equipment related to position types. Background Art
[0002] With the development of the aerospace manufacturing industry, the structure of the aircraft power distribution system has changed from centralized to distributed, becoming increasingly complex, and at the same time, the manufacturing precision requirements have increased. The distributed power distribution system covers power distribution equipment with various functions, and fasteners composed of various screws and washers are assembled on the surface of its equipment shell. However, during the assembly process, these fasteners may have technological defects such as misassembly and missing assembly. If these defects cannot be discovered and properly handled in a timely manner, it may lead to a decrease in the structural strength of the equipment, electrical performance failures, and even major accidents, resulting in serious consequences. Therefore, in order to test whether the components of the aircraft power distribution equipment meet the assembly requirements, it is necessary to detect the missing or misassembled defects of the fasteners on the surface of the power distribution equipment.
[0003] Vision detection technology relies on a high-resolution imaging system and advanced image processing algorithms to achieve fast and accurate detection without physical contact and material limitations. Currently, the mainstream vision detection technology solutions use deep learning object detection algorithms such as RCNN, YOLO, and RTDETR to identify the types of fasteners, and significant progress has been made in terms of detection speed and accuracy. However, this method has certain deficiencies:
[0004] 1. For the missing assembly defect where no fastener is assembled at the target position and the misassembly defect of assembling the wrong type of fastener in the actual assembly scenario. A single deep learning object detection method can only determine the existence and category attributes of the fasteners, lacking an analysis mechanism for the correlation between the assembly position and the fastener type, unable to verify the corresponding relationship between the target position and the fastener type in the preset assembly process, and it is difficult to effectively identify the missing or misassembled defects.
[0005] 2. During the detection process, the imaging angle and lighting conditions interfere with the detection results. Due to the limitation of the equipment shell structure, some fasteners are blocked during shooting and acquisition; there is surface reflection under strong light conditions, and image details are lost under low light conditions.
[0006] 3. The detection method based on deep learning is restricted by the scarcity of research samples in the field of surface defect detection related to aerospace, especially for the problem of insufficient effective sample data of the assembly state of power distribution equipment fasteners and category imbalance. Existing solutions either rely on artificial modeling to synthesize samples, which have a large difference from the original images; or are limited to the original image enhancement technology, restricting the improvement of the model generalization performance. Summary of the Invention
[0007] The object of the present invention is to overcome the deficiencies of the prior art and provide a method for detecting missing or misassembled fasteners of aircraft power distribution equipment related to position types. Through the intelligent detection of surface fastener assembly defects of aircraft power distribution equipment coordinated by a rotating table, it is possible to efficiently, completely, and accurately detect missing or misassembled fastener defects in multi-view detection.
[0008] To achieve the above object of the invention, a method for detecting missing or misassembled fasteners of aircraft power distribution equipment related to position types according to the present invention is characterized by including the following steps:
[0009] (1), Commissioning of the detection device and data acquisition;
[0010] (1.1), Commissioning of the detection device;
[0011] Fix the object to be measured at the center of the rotating table; fix the grayscale camera at the center of the image acquisition device, and fix the light supplement device on the left side of the grayscale camera;
[0012] Adjust the positions of the rotating table and the image acquisition device so that the grayscale camera can clearly capture the object to be measured placed on the rotating platform, and the projected light of the light supplement device can cover the surface of the object to be measured, enabling the image acquisition device to capture clear images;
[0013] (1.2), Acquisition of template images;
[0014] Take the standard sample of the aircraft power distribution equipment as the object to be measured and fix it at the center of the rotating table. Under the light supplement condition of normal light intensity, control the rotation angle of the rotating table, and then acquire 1 template image at each angle, denoted as T i , where i is the angle index;
[0015] (1.3), Acquisition of images to be measured;
[0016] Take the aircraft power distribution equipment to be measured as the object to be measured and fix it at the center of the rotating table. Under the light supplement conditions of different light intensities, control the rotation angle of the rotating table, and then acquire images to be measured at each angle. Among them, the image to be measured under the nth light supplement intensity at the i-th angle is denoted as I i,n , where i is the angle index; n is the light supplement intensity index, n = 1, 2, 3, n = 1 represents low light intensity, n = 2 represents normal light intensity, n = 3 represents high light intensity;
[0017] (1.4), Acquisition of the sample data set;
[0018] Take the aircraft power distribution equipment to be measured as the object under test and fix it at the center of the rotating table. Adjust the distance between the image acquisition device and the rotating table. On the premise that the grayscale camera can clearly capture the object under test placed on the rotating platform, control the rotation angle of the rotating table under different distances and different light intensity fill light conditions, and then collect the sample images of the object under test at each angle. Among them, the sample image collected at the k-th distance, the i-th angle, and the n-th fill light intensity is denoted as F k,i,n , where k is the distance index between the rotating table and the image acquisition device, k = 1, 2, 3. k = 1 represents a short distance, k = 2 represents a normal distance, and k = 3 represents a long distance;
[0019] (2) Expose and fuse the test images with different light intensities taken at each angle;
[0020] (2.1) Calculate the contrast, saturation, and exposure of each pixel point in the test image;
[0021] For each pixel point (x, y) in the test image I i,n with different light intensities at each angle, use the Laplace operator to measure the contrast C i,n (x, y) of the pixel point, use the standard deviation of the RGB space channel values of the image to measure the saturation S i,n (x, y) of the pixel point, and use the Gaussian weighted distance to measure the exposure E i,n (x, y);
[0022] (2.2) Calculate the fusion weight of each pixel point;
[0023] Combine the contrast measure C i,n (x, y), the saturation measure S i,n (x, y) and the exposure measure E i,n (x, y) to calculate the fusion weight W i,n (x, y) of each pixel point in the test image with different light intensities at each angle:
[0024]
[0025] Among them, ω C , ω S , ω E are adjustable weight coefficients, which respectively control detail enhancement, color fidelity, and exposure balance;
[0026] (2.3) Image exposure fusion;
[0027] Based on the fusion weight W i,n (x, y) of each pixel point, perform exposure fusion processing on the test image I i,n (x, y) with different light intensities at each angle to obtain the test image after exposure fusion
[0028]
[0029] (3) Synthesize diverse sample images through a denoising diffusion probabilistic model to expand the dataset;
[0030] (3.1) Use the denoising diffusion probabilistic model to perform t-step noise addition processing on each sample image F in the sample dataset k,i,n to obtain a noisy image N k,i,n,t , where t is the noise addition step index, t = 1, 2,..., T, and T is the total number of steps in the noise addition process;
[0031]
[0032] Among them, the parameter α t is a parameter preset in the t-th iteration, and ε t is the Gaussian noise added in the t-th time and follows a distribution with a mean of 0 and a variance of 1;
[0033] (3.2) Use the denoising diffusion probabilistic model to perform denoising processing on the noisy image N k,i,n,t :
[0034]
[0035] Among them, β t = 1 - α t , and ζ t is the noise amount predicted in the t-th iteration;
[0036] (3.3) Each sample image F k,i,n synthesizes diverse sample images after T rounds of forward noise addition and reverse denoising;
[0037] (4) Construct and train the YOLO11-AEDSF model;
[0038] (4.1) Construct the YOLO11-AEDSF model;
[0039] Taking the YOLO11 network as the basic model, introducing the ConvGLU convolutional gated linear unit and the FasterNet Block sparse computing module to construct the C3k2 Faster CGLU module, and then using the C3k2 Faster CGLU module to replace the C3k2 module in the YOLO11 network; introducing the SIoU loss function and the Focaler-IoU mechanism to construct the Focaler-SIoU loss function, and then using the Focaler-SIoU loss function to replace the CIoU loss function in the YOLO11 network, thus constructing the YOLO11-AEDSF model;
[0040] (4.2), Train the YOLO11-AEDSF model;
[0041] (4.2.1), Combine the sample images collected in step (1.4) and the diversified sample images synthesized in step (3) to form a training image set;
[0042] Add annotation boxes to the fasteners in the training images through manual annotation, and mark the category and position coordinates of each fastener;
[0043] (4.2.2), YOLO11-AEDSF model training;
[0044] Divide the training image set into several batches, input the training images of each batch into the YOLO11-AEDSF model, and obtain the detection result images corresponding to each training image in the input batch Among them, each detection result image also contains prediction boxes for multiple fasteners, as well as the predicted categories and predicted position coordinates of each fastener;
[0045] Taking the prediction box as the reference, calculate the matching degree between the prediction box and each annotation box;
[0046]
[0047] Among them, represents the matching degree between the τ-th prediction box and the τ-th annotation box b τ , B loro represents the spatial prior probability of whether the anchor point is inside the annotation box, P is the classification category of the prediction box, is the intersection over union between the prediction box and the annotation box, and α, β are hyperparameters for balancing object classification and regression localization;
[0048] Select the annotation box corresponding to the maximum matching degree as the annotation box of the prediction box and then according to the prediction box and the corresponding annotation box b τ Calculate the loss function value:
[0049]
[0050] Among them, L Focaler-SIoU is the Focaler-SIoU loss, L dfl is the distribution focus loss, L cls is the classification loss, η1, η2, η3 are the gain coefficients corresponding to the three losses respectively;
[0051] Each detection result image The loss function values of each prediction box and the corresponding annotation box are superimposed to obtain the total loss value of the current batch of training images. Then, based on the total loss value, the network model parameters are updated by back propagation using the gradient descent method. Then, the next batch of training images is input repeatedly in this way until the YOLO11-AEDSF model converges to obtain the trained YOLO11-AEDSF model.
[0052] (4.2.3) Use the layer adaptive pruning algorithm to prune and optimize the trained YOLO11-AEDSF model, remove redundant network connections and weights, and obtain the optimized YOLO11-AEDSF model;
[0053] (5) Detection of incorrect or missing assembly defects of surface fasteners of aircraft power distribution equipment;
[0054] (5.1) Annotate each template image T i The annotation box of each fastener in , where the bth annotation box is denoted as gt i,b Each annotation box contains the annotation box coordinates gtbox i,b and the annotation box category gtCls i,b , where i is the angle index and b is the prediction box index;
[0055] (5.2) Use the YOLO11-AEDSF model to detect the image to be tested after exposure fusion The type and coordinate information of the fastener;
[0056] (5.2.1) The image to be tested Perform pre-processing;
[0057] Keep the image to be tested The aspect ratio is scaled to the input size of the YOLO11-AEDSF model, and then the pixel values are normalized to obtain the preprocessed image.
[0058] (5.2.2), preprocess the image Input the backbone network of the YOLO11 - AEDSF model. Through the Conv convolution module, C3k2 - Faster - CGLU module, SPPF module, and C2PSA module, feature maps of three scales are output, denoted as the shallow - layer feature map P3 i , the middle - layer feature map P4 i , and the deep - layer feature map P5 i ;
[0059] (5.2.3), Input the feature maps P3 i , P4 i , P5 i into the neck network of the YOLO11 - AEDSF model. Through upsampling, Conv convolution module, and C3k2 Faster CGLU module, optimized feature maps of three scales are output, denoted as the shallow - layer optimized feature map P3' i , the middle - layer optimized feature map P4' i , and the deep - layer optimized feature map P5' i ;
[0060] (5.2.4), Input the fused feature maps P3' i , P4' i , P5' i into the detection - head network of the YOLO11 - AEDSF model to obtain a prediction result image containing multiple prediction boxes
[0061] (5.3), Use the Linemod - 2d method to match 1 template image with the highest similarity to the prediction result image as the valid template image, and then obtain the rotation and translation parameters (x , y , θ ) that map the valid template image i , y i , θ i ) to the prediction result image, where (x i , y i ) are the translation parameters, corresponding to the pixel translation amounts along the horizontal and vertical directions respectively, and θ i is the rotation parameter, representing the radian value of the counter - clockwise rotation of the template image around its geometric center;
[0062] (5.4), Rotate and translate the valid template image and each annotation box using the rotation and translation parameters (x i , y i , θ i ) to obtain an image that is consistent with the prediction result image Aligned calibration template image and calibration annotation box
[0063] (5.5) Detection of missing or misassembled defects;
[0064] (5.5.1) Calculate the prediction result image using the intersection over union formula for each prediction box in the prediction result image and the annotation box in the calibration template image
[0065]
[0066] Among them, is the area of the prediction box, is the area of the annotation box, A intersection is the intersection area of the prediction box and the annotation box;
[0067] (5.5.2) Effective matching determination;
[0068] Set the intersection over union threshold T IoU , if it satisfies then it is determined that the prediction box and the annotation box are effectively matched, otherwise it is an invalid match;
[0069] (5.5.3) Missing or misassembled defect determination;
[0070] If the category of the effectively matched prediction box is the same as the category of the annotation box, it is determined that the fastener is assembled correctly;
[0071] If the category of the effectively matched prediction box is not the type of assembly hole and is different from the category of the annotation box, it is determined that the fastener is misassembled;
[0072] If the category of the effectively matched prediction box is the type of assembly hole and is different from the category of the annotation box, it is determined that the fastener is missing;
[0073] If there is an annotation box with an invalid match but no prediction box, it is determined that the fastener is missing.
[0074] The object of the present invention is achieved as follows:
[0075] A method for detecting missing or misassembled fasteners of aircraft power distribution equipment related to position types in the present invention. First, the detection device is debugged. After the detection device is debugged, the standard sample and the sample to be measured of the aircraft power distribution equipment are collected respectively to obtain the template image, the sample image and the test image. Then, the image to be measured is subjected to exposure fusion, the sample image is subjected to forward noise addition and reverse denoising, and diverse sample images are synthesized. The collected sample images and the synthesized diverse sample images together constitute a training image set for training the built YOLO11-AEDSF model. Finally, the trained YOLO11-AEDSF model is used to detect the image to be measured after exposure fusion, and the missing or misassembled defects of the fasteners on the surface of the aircraft power distribution equipment are identified.
[0076] Meanwhile, a method for detecting missing or misassembled fasteners of aircraft power distribution equipment related to position types in the present invention also has the following beneficial effects:
[0077] (1) In the present invention, the C3k2 Faster CGLU module is designed to replace the original C3k2 structure to achieve lightweight design. The Focaler-SIoU loss function is introduced to optimize the detection of difficult samples, and the model structure is optimized through the layer adaptive pruning algorithm to remove redundant network connections and weights, improving the calculation efficiency and detection accuracy.
[0078] (2) In the present invention, diverse sample images are generated through the denoising diffusion probability model to solve the problem of few-sample data. Specifically, noise data is sampled from the standard Gaussian distribution, and various aircraft power distribution equipment surface fastener assembly state image data is gradually denoised through the reverse process to expand the original data set, balance the defect category distribution, and enhance the model's detection ability for few-sample targets.
[0079] (3) Based on multi-angle rotation acquisition, the detection blind area is eliminated. For the images to be measured with different light intensities collected at the same angle, the interference of reflection and dark areas is eliminated through the weighted fusion of the three elements of contrast, saturation, and exposure.
[0080] (4) In the present invention, the effective template image of the predicted image is matched and spatially aligned through the Linemod-2D method. Finally, through the intersection over union matching mechanism between the template annotation box and the result prediction box, the correlation analysis of position and type is carried out to realize the detection of missing or misassembled fastener defects, and solve the pain point that a single target detection model cannot verify the corresponding relationship between the target position and the fastener type in the preset assembly process. Description of the Drawings
[0081] Figure 1 is the flow chart of a method for detecting missing or misassembled fasteners of aircraft power distribution equipment related to position types in the present invention;
[0082] Figure 2 is the architecture diagram of the detection device;
[0083] Figure 3 is the image after fusing test images under different exposure intensities;
[0084] Figure 4 is the synthesized diversified sample image after adding noise to the sample image in the forward direction and removing noise in the reverse direction;
[0085] Figure 5 is the schematic diagram of the YOLO11-AEDSF model structure;
[0086] Figure 6 is the schematic diagram of the C3k2 Faster CGLU module structure;
[0087] Figure 7 is the schematic diagram for comparing the convergence process of the YOLO11-AEDSF model;
[0088] Figure 8 is the schematic diagram of the inspection result of the surface fasteners of the aircraft power distribution equipment. Specific Embodiments
[0089] The following describes the specific embodiments of the present invention with reference to the accompanying drawings so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.
[0090] Embodiment
[0091] In this embodiment, as Figure 1 shown, a method for detecting missing and misinstalled fasteners of aircraft power distribution equipment related to position types according to the present invention includes the following steps:
[0092] S1. Commissioning of the detection device and data acquisition;
[0093] S1.1. Commissioning of the detection device;
[0094] In this embodiment, as Figure 2 shown, the object to be measured is fixed at the center of the rotating table; the grayscale camera is fixed at the center of the image acquisition device, and the supplementary light device is fixed on the left side of the grayscale camera;
[0095] Adjust the positions of the rotating table and the image acquisition device so that the grayscale camera can clearly capture the object to be measured placed on the rotating platform, and the projected light of the supplementary light device can cover the surface of the object to be measured, enabling the image acquisition device to capture clear images;
[0096] S1.2. Acquisition of template images;
[0097] Take the standard sample of the aircraft power distribution equipment as the object to be measured and fix it at the center of the rotating table. Under the supplementary lighting condition of normal light intensity, control the rotation angle of the rotating table, and then collect 1 template image at each angle, denoted as T i , where i is the angle index;
[0098] S1.3. Collect the images to be measured;
[0099] Take the aircraft power distribution equipment to be measured as the object to be measured and fix it at the center of the rotating table. Under the supplementary lighting conditions of different light intensities, control the rotation angle of the rotating table, and then collect the images to be measured at each angle. Among them, the image to be measured collected at the i-th angle under the n-th supplementary lighting intensity is denoted as I i,n , where i is the angle index; n is the supplementary lighting intensity index, n = 1, 2, 3, n = 1 represents low light intensity, n = 2 represents normal light intensity, and n = 3 represents high light intensity;
[0100] S1.4. Collect the sample data set;
[0101] Take the aircraft power distribution equipment to be measured as the object to be measured and fix it at the center of the rotating table. Adjust the distance between the image acquisition device and the rotating table. On the premise that the grayscale camera can clearly capture the object to be measured placed on the rotating platform, under the supplementary lighting conditions of different distances and different light intensities, control the rotation angle of the rotating table, and then collect the sample images of the object to be measured at each angle. Among them, the sample image collected at the k-th distance, the i-th angle, and the n-th supplementary lighting intensity is denoted as F k,i,n , k is the distance index between the rotating table and the image acquisition device, k = 1, 2, 3, k = 1 represents a short distance, k = 2 represents a normal distance, and k = 3 represents a long distance;
[0102] In this embodiment, taking the i-th angle as an example, the projection intensity of the supplementary lighting device is L n = L base +(n - 1)L Δ of light onto the surface of the object to be measured, and then collect the grayscale images of the aircraft power distribution equipment to be measured with different supplementary lighting intensities. Among them, L base is the initial supplementary lighting intensity, and L Δ is the incremental supplementary lighting intensity each time;
[0103] S2. Perform exposure fusion on the images to be measured with different light intensities taken at each angle;
[0104] S2.1. Calculate the contrast, saturation, and exposure of each pixel point in the images to be measured;
[0105] For each pixel point (x, y) in the images to be measured with different light intensities at each angle I i,n , use the Laplacian operator to measure the contrast C of the pixel point i,n(x, y), the standard deviation of the image RGB spatial channel value is used to measure the saturation S of the pixel point i,n (x, y), using Gaussian weighted distance to measure the exposure E of the pixel i,n (x,y);
[0106] S2.2, calculate the fusion weight of each pixel;
[0107] Combined with contrast metric C i,n (x,y), saturation measure S i,n (x,y) and the exposure metric E i,n (x, y), calculate the fusion weight W of each pixel in the image to be tested at different light intensities at each angle i,n (x,y):
[0108]
[0109] Among them, ω C ,ω S ,ω E The weight coefficients are adjustable, controlling detail enhancement, color fidelity, and exposure balance respectively;
[0110] S2.3, image exposure fusion;
[0111] Based on the fusion weight W of each pixel i,n (x, y), for each angle, the image I with different light intensities i,n (x, y) is subjected to exposure fusion processing to obtain the image to be tested after exposure fusion
[0112]
[0113] In this embodiment, Figure 3 The dark light image I collected at angle i i,1 (x,y), normal light image to be tested I i,2 (x,y), strong light image to be tested I i,3 (x, y) is subjected to exposure fusion processing to obtain the image to be tested after exposure fusion Comparison chart, where (a) is the dark light image I i,1 (x, y), (b) is the strong light image I i,3 (x, y), (c) is the strong light image I i,3 (x, y), (d) are the images to be tested after exposure fusion from Figure 3It can be seen that a large amount of dark details in the low-light intensity image are lost, the details in the normal-light intensity image are also insufficient, and there is a serious surface reflection phenomenon in the high-light intensity image. After exposure fusion, the dark details of the image are significantly enhanced and the reflection phenomenon is significantly reduced.
[0114] S3. Synthesize diverse sample images through a denoising diffusion probabilistic model;
[0115] S3.1. Use the denoising diffusion probabilistic model to perform t-step noise addition processing on each sample image F in the sample dataset k,i,n to obtain a noisy image N k,i,n,t , where t is the noise addition step index, t = 1, 2,..., T, and T is the total number of steps in the noise addition process;
[0116]
[0117] Among them, the parameter α t is the parameter preset at the t-th iteration, ε t is the Gaussian noise added at the t-th time and follows a distribution with a mean of 0 and a variance of 1;
[0118] S3.2. Use the denoising diffusion probabilistic model to perform denoising processing on the noisy image N k,i,n,t :
[0119]
[0120] Among them, β t = 1 - α t , ζ t is the amount of noise predicted at the t-th iteration;
[0121] S3.3. Each sample image F k,i,n is synthesized into diverse sample images after T rounds of forward noise addition and reverse denoising;
[0122] In this embodiment, Figure 4 is a schematic diagram of synthesizing diverse sample images after the sample image F k,i,n has undergone T rounds of forward noise addition and reverse denoising, where (a) is the sample image F k,i,n , and (b) is a schematic diagram of the synthesized sample image;
[0123] S4. Construct and train the YOLO11-AEDSF model;
[0124] S4.1. Construct the YOLO11-AEDSF model;
[0125] In this embodiment, as Figure 5As shown in the figure, the YOLO11-AEDSF model uses the YOLO11 network as the basic model, introduces the ConvGLU convolutional gated linear unit and the FasterNet Block sparse computing module to construct the C3k2 Faster CGLU module, and then uses the C3k2 Faster CGLU module to replace the C3k2 module in the YOLO11 network; introduces the SIoU loss function and the Focaler-IoU mechanism to construct the Focaler-SIoU loss function, and then uses the Focaler-SIoU loss function to replace the CIoU loss function in the YOLO11 network, thus constructing the YOLO11-AEDSF model;
[0126] In this embodiment, as Figure 6 shown, the structure of the C3k2 Faster CGLU module is:
[0127] Replace the pointwise convolution in the FasterNet Block sparse computing module with the ConvGLU convolutional gated linear unit to construct the Faster CGLU Block module.
[0128] Use the Faster CGLU Block module to replace the Bottleneck bottleneck block of the C3k module in the YOLO11 basic model to construct the C3k Faster CGLU module.
[0129] In the C3k2 module of the YOLO11 basic model, when the module configuration parameter C3k = True, use the C3k Faster CGLU module to replace the C3k module; when the module configuration parameter C3k = False, use the Faster CGLU Block to replace the Bottleneck bottleneck block to construct the C3k2 Faster CGLU module.
[0130] S4.2. Train the YOLO11-AEDSF model;
[0131] S4.2.1. Combine the sample images collected in step S1.4 and the diversified sample images synthesized in step S3 to form a training image set;
[0132] Add annotation boxes to the fasteners in the training images through manual annotation, and annotate the category and position coordinates of each fastener;
[0133] Table 1 shows the specific information of the training image set composed of the sample images collected in step S1.4 and the diversified sample images synthesized in step S3. Among them, there are 2,100 collected sample images and 700 synthesized diversified sample images. The collected sample image data is divided in the ratio of training images: validation images: test images = 8:1:1. To ensure the authenticity of the test set, the synthesized diversified sample image data is divided in the ratio of training images: validation images = 8:2. Finally, the training image set contains 2,240 training images, 350 validation images, and 210 test images.
[0134] Table 1, Specific information of the training image set
[0135] Image dataset Training image Validation image Test image Total Collected sample images 1680 210 210 2100 Synthesized diverse sample images 560 140 0 700 Training image set 2240 350 210 2800
[0136] S4.2.2. Training of the YOLO11-AEDSF model;
[0137] Divide the training image set into several batches, and input the training images of each batch into the YOLO11-AEDSF model to obtain the detection result images corresponding to each training image in the input batch. Among them, each detection result image also contains prediction boxes of multiple fasteners, as well as the predicted categories and predicted position coordinates of each fastener;
[0138] Based on the prediction box calculate the matching degree between the prediction box and each annotation box;
[0139]
[0140] Among them, represents the matching degree between the τ-th prediction box and the τ-th annotation box b τ , B loro represents the spatial prior probability of whether the anchor point is inside the annotation box, P is the classification category of the prediction box, is the intersection over union between the prediction box and the annotation box, and α, β are hyperparameters for balancing object classification and regression localization;
[0141] Select the annotation box corresponding to the maximum matching degree as the annotation box of the prediction box , and then calculate the loss function value according to the prediction box and the corresponding annotation box b τ :
[0142]
[0143] Among them, L Focaler-SIoU is the Focaler-SIoU loss, Ldfl is the distribution focal loss, L cls is the classification loss, and η1, η2, and η3 are the gain coefficients corresponding to the three losses respectively;
[0144] Among them, the Focaler - SIoU loss is:
[0145] L Focaler-SIoU = L SIoU +(IoU - IoU Focaler )
[0146] Among them, L SIoU is the SIoU loss function, IoU is the intersection - over - union of the predicted bounding box and the annotated bounding box, and IoU Focaler is the IoU reconstructed by using the linear interval mapping method, which is expressed as:
[0147]
[0148] Among them, the thresholds g and u are used to control the IoU weight assignment.
[0149] In this embodiment, the SIoU loss function is introduced, and the influence of the angle between the bounding boxes on the bounding box regression is further considered. The convergence process is accelerated by reducing the angle between the anchor box and the annotated box in the horizontal or vertical direction; in addition, by setting the dynamic thresholds g and u to adjust the loss weights, the loss contribution of the easy - to - detect samples is reduced, and the risk of overfitting is reduced; the difficult samples are focused, and the complete loss gradient is retained;
[0150] The distribution focal loss is:
[0151] L dfl = L top + L bottom + L left + L right
[0152]
[0153] Among them, L top , L bottom , L left , L right respectively represent the distribution focal losses of the predicted bounding box and the annotated bounding box at the upper boundary, lower boundary, left boundary, and right boundary, and d top , d bottom , d left , d right represent the distances from the upper boundary, lower boundary, left boundary, and right boundary of the annotated bounding box to the center of the annotated bounding box; represents the integer part of rounding down the distance d, and k ∈ {k top , k bottom , k left , kright}, d ∈ {d top , d bottom , d left , d right}; is the residual weight of the distance d to the ceiling; S k represents the discrete probability that the distance from the prediction box boundary to the center of the annotation box is an integer k; S k+1 represents the discrete probability that the distance from the prediction box boundary to the center of the annotation box is an integer k + 1, represents rounding up, represents rounding down;
[0154] The classification loss is:
[0155]
[0156] where y j represents the probability value that the annotation box is of the j-th category, y j = 0, 1, represents the probability value that the prediction box is of the j-th category, N is the number of fastener categories.
[0157] For each detection result image the loss function values of each prediction box and the corresponding annotation box are superimposed to obtain the total loss value of the current batch of training images. Then, based on the obtained total loss value, the network model parameters are updated by backpropagation using the gradient descent method. Then, by analogy, the next batch of training images is input repeatedly until the YOLO11 - AEDSF model converges to obtain the trained YOLO11 - AEDSF model;
[0158] In this embodiment, the model training platform is built based on the Windows system, the experimental environment is configured with PyTorch1.20.1 and CUDA11.8, and all experiments are carried out on GeForce RTX 4060;
[0159] Based on the training image set in step S4.2.1, the training parameters of this experiment are set: momentum factor 0.937, weight decay coefficient 0.0005, initial learning rate 0.01, optimized using the SGD function, the number of input images per batch is 16, and the number of training epochs is 200 to train the YOLO11 - AEDSF model.
[0160] Figure 7 is a schematic diagram of the convergence process comparison between the CIoU loss function of the conventional YOLO11 network and the Focaler - SIoU loss function of the YOLO11 - AEDSF model of the present invention during the training process of this embodiment; From Figure 7It can be seen that the YOLO11-AEDSF model constructed by the present invention has a faster convergence speed.
[0161] S4.2.3. Use the layer adaptive pruning algorithm to prune and optimize the trained YOLO11-AEDSF model, remove redundant network connections and weights, and obtain the optimized YOLO11-AEDSF model.
[0162] Table 2 is the experimental results of the comparison between the optimized YOLO11-AEDSF model and the conventional YOLO11 model in this embodiment. The experimental results show that the floating point operation Flops (G) of the optimized YOLO11-AEDSF model is reduced by 60% compared with the conventional YOLO11 model, the time for reasoning an image is shortened from 33ms to 21ms, and the mAP is improved by 2.8%. The above results show that the optimized YOLO11-AEDSF model requires less computing resources, and has higher computing efficiency and detection accuracy.
[0163] Table 2, comparative experimental results
[0164] Model Infer_time(ms) mAP@0.5(%) Flops(G) YOLO11 33 94.8 6.4 YOLO11 - AEDSF 30 97.3 6.0 Optimized YOLO11 - AEDSF model 21 97.6 2.4
[0165] Among them, Flops (G) represents the floating point number operation of the model, and infer_time represents the time it takes for the model to infer an image. These two indicators are used to measure the detection speed performance of the model. The lower the value, the better the model performance. mAP@0.5 represents the average precision of the model in all categories when the IoU threshold is 0.5. Usually, this indicator is used to measure the performance of the model in the target detection task. The higher the value, the better the model performance.
[0166] S5. Detection of incorrect or missing assembly defects of surface fasteners of aircraft power distribution equipment;
[0167] S5.1. Label each template image T i The annotation box of each fastener in , where the bth annotation box is denoted as gt i,b Each annotation box contains the annotation box coordinates gtbox i,b and the annotation box category gtCls i,b , where i is the angle index and b is the prediction box index;
[0168] S5.2. Use the optimized YOLO11-AEDSF model to detect the image to be tested after exposure fusion The type and coordinate information of the fastener;
[0169] S5.2.1. The image to be tested Perform pre-processing;
[0170] Keep the image to be tested According to the aspect ratio, scale it to the input size of the YOLO11-AEDSF model, and then perform pixel value normalization to obtain the preprocessed image
[0171] S5.2.2. Input the preprocessed image into the backbone network of the YOLO11-AEDSF model, and output feature maps of three scales through the Conv convolution module, C3k2-Faster-CGLU module, SPPF module and C2PSA module, which are respectively denoted as the shallow feature map P3 i , the middle feature map P4 i , and the deep feature map P5 i ;
[0172] S5.2.3. Input the feature maps P3 i , P4 i , and P5 i into the neck network of the YOLO11-AEDSF model, and output optimized feature maps of three scales through upsampling, Conv convolution module and C3k2 Faster CGLU module, which are respectively denoted as the shallow optimized feature map P3' i , the middle optimized feature map P4' i , and the deep optimized feature map P5' i ;
[0173] S5.2.4. Input the fused feature maps P3' i , P4' i , and P5' i into the detection head network of the YOLO11-AEDSF model to obtain a prediction result image containing multiple prediction boxes
[0174] S5.3. Use the Linemod-2d method to match the 1 template image with the highest similarity to the prediction result image as the effective template image, and then obtain the rotation and translation parameters (x , y ) that map the effective template image to the prediction result image , where (x i , y i ) are the translation parameters, corresponding to the pixel translation amounts along the horizontal and vertical directions respectively, and θ i is the rotation parameter, representing the radian value of the counterclockwise rotation of the template image around its geometric center; i , y i ) are the translation parameters, corresponding to the pixel translation amounts along the horizontal and vertical directions respectively, and θ i is the rotation parameter, representing the radian value of the counterclockwise rotation of the template image around its geometric center;
[0175] S5.4. The effective template image and each annotation box Adopt rotation and translation parameters (x i , y i , θ i ) to perform rotation and translation transformation, and obtain a corrected template image aligned with the predicted result image and corrected annotation boxes
[0176] S5.5. Detection of missing and misaligned assembly defects;
[0177] S5.5.1. Calculate the geometric overlap degree between each predicted box in the predicted result image and the annotation box in the corrected template image :
[0178]
[0179] wherein, is the area of the predicted box, is the area of the annotation box, and A intersection is the intersection area between the predicted box and the annotation box;
[0180] S5.5.2. Effective matching determination;
[0181] Set the intersection over union threshold T IoU . If is satisfied, then it is determined that the predicted box and the annotation box are effectively matched, otherwise it is an invalid match;
[0182] S5.5.3. Missing and misaligned assembly defect determination;
[0183] If the category of the effectively matched predicted box is the same as the category of the annotation box, it is determined that the fastener assembly is correct;
[0184] If the category of the effectively matched predicted box is not the assembly hole type and is different from the category of the annotation box, it is determined that the fastener is misassembled;
[0185] If the category of the effectively matched predicted box is the assembly hole type and is different from the category of the annotation box, it is determined that the fastener is missing;
[0186] If the annotation box with invalid match exists but the predicted box does not exist, it is determined that the fastener is missing.
[0187] Figure 8 This is the method for detecting the image to be measured after exposure fusion through step S5 in this embodiment Schematic diagram of experimental results of misassembly defects of medium fasteners, where (A) is the image to be measured after exposure fusion detected by the optimized YOLO11-AEDSF model The obtained predicted result image (B) is the effective template image with the highest similarity matched and predicted result image by the Linemod-2d method The effective template image with the highest similarity (C) is the image to be measured after exposure fusion Detection results of misassembly defects of medium fasteners. Among them, the white box is the misassembly defect, that is, the category of the effectively matched prediction box is not the assembly hole type and is inconsistent with the category of the annotation box. The gray box is the missing assembly defect, that is, the category of the effectively matched prediction box is the assembly hole type and is inconsistent with the category of the annotation box;
[0188] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate the understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A method for detecting missing or incorrect installation of fasteners of aircraft power distribution equipment related to position types, characterized in that, It includes the following steps: (1) Debug the detection device and collect data; (1.1) Debug the detection device; Fix the object to be measured at the center of the rotary table; fix the grayscale camera at the center of the image acquisition device, and fix the supplementary light device on the left side of the grayscale camera; Adjust the positions of the rotary table and the image acquisition device so that the grayscale camera can clearly capture the object to be measured placed on the rotary platform, and the projected light of the supplementary light device can cover the surface of the object to be measured, enabling the image acquisition device to capture clear images; (1.2) Collect template images; Take the standard sample of the aircraft power distribution equipment as the object to be measured and fix it at the center of the rotating table. Under the supplementary light condition of normal light intensity, control the rotation angle of the rotating table, and then collect 1 template image at each angle, denoted as T i , where i is the angle index; (1.3) Collect images to be measured; Take the aircraft power distribution equipment to be measured as the object under test and fix it at the center of the rotary table. Under the supplementary lighting conditions of different light intensities, control the rotation angle of the rotary table, and then collect the images to be measured at each angle. Among them, the image to be measured under the nth supplementary lighting intensity at the ith angle is denoted as I i,n , where i is the angle index; n is the supplementary lighting intensity index, n = 1, 2, 3, n = 1 represents low light intensity, n = 2 represents normal light intensity, and n = 3 represents high light intensity; (1.4) Collect the sample data set; Take the aircraft power distribution equipment to be measured as the object under test and fix it at the center of the rotary table. Adjust the distance between the image acquisition device and the rotary table. On the premise that the gray-scale camera can clearly capture the object under test placed on the rotary platform, control the rotation angle of the rotary table under the supplementary light conditions of different distances and different light intensities, and then collect the sample images of the object under test at each angle. Among them, the sample image collected at the k-th distance, the i-th angle, and the n-th supplementary light intensity is denoted as F k,i,n , where k is the distance index of the rotary table and the image acquisition device, k = 1, 2, 3. k = 1 represents a short distance, k = 2 represents a normal distance, and k = 3 represents a long distance; (2) Perform exposure fusion on the images to be measured with different light intensities taken at each angle; (2.1) Calculate the contrast, saturation, and exposure of each pixel point in the images to be measured; For each image I of different light intensities to be measured at each angle i,n For each pixel point (x, y) in i,n , the Laplacian operator is used to measure the contrast C of the pixel point i,n (x, y), the standard deviation of the RGB space channel values of the image is used to measure the saturation S of the pixel point i,n (x, y), the Gaussian weighted distance is used to measure the exposure E of the pixel point i,n (x, y); (2.2) Calculate the fusion weight of each pixel point; Combined with the contrast metric C i,n (x, y), the saturation metric S i,n (x, y) and the exposure metric E i,n (x, y), calculate the fusion weight W of each pixel in the image to be measured with different light intensities at each angle i,n (x, y): Among them, ω C , ω S , ω E are adjustable weight coefficients, which respectively control detail enhancement, color fidelity, and exposure equalization; (2.3) Perform image exposure fusion; Based on the fusion weight W of each pixel point i,n (x, y), for the measured images I of different light intensities at each angle i,n (x, y), perform exposure fusion processing to obtain the measured image after exposure fusion (3) Synthesize diverse sample images through the denoising diffusion probability model; (3.1) Apply the denoising diffusion probability model to each sample image F in the sample dataset k,i,n to perform t-step noise addition processing to obtain a noisy image N k,i,n,t , where t is the noise addition step index, t = 1, 2,..., T, and T is the total number of steps in the noise addition process; Among them, the parameter α t is the preset parameter at the t-th iteration, and ε t is the Gaussian noise added at the t-th time and follows a distribution with a mean of 0 and a variance of 1; (3.2) Then, use the denoising diffusion probabilistic model to further denoise the noisy image N k,i,n,t for denoising processing: where β t = 1 - α t , ζ t is the amount of noise predicted in the t-th iteration; (3.3), each sample image F k,i,n After T rounds of forward noise addition and reverse denoising, diverse sample images are synthesized; (4) Construct and train the YOLO11-AEDSF model; (4.1) Construct the YOLO11-AEDSF model; Using the YOLO11 network as the basic model, introduce the ConvGLU convolutional gated linear unit and the FasterNet Block sparse calculation module to construct the C3k2 Faster CGLU module, and then use the C3k2 Faster CGLU module to replace the C3k2 module in the YOLO11 network; introduce the SIoU loss function and the Focaler-IoU mechanism to construct the Focaler-SIoU loss function, and then use the Focaler-SIoU loss function to replace the CIoU loss function in the YOLO11 network, thereby constructing the YOLO11-AEDSF model; (4.2) Train the YOLO11-AEDSF model; (4.2.1) jointly form a training image set with the sample images collected in step (1.4) and the diverse sample images synthesized in step (3); Add annotation boxes to the fasteners in the training images through manual annotation, and mark the category and position coordinates of each fastener; (4.2.2) Train the YOLO11-AEDSF model; Divide the training image set into several batches, and input the training images of each batch into the YOLO11-AEDSF model to obtain the detection result images corresponding to each training image in the input batch Among them, each detection result image also contains prediction boxes of multiple fasteners, as well as the predicted categories and predicted position coordinates of each fastener; Taking the prediction box as a reference, calculate the matching degree between the prediction box and each annotation box; Among them, represents the τ-th predicted bounding box and the matching degree with the τ-th labeled bounding box b τ , B loro represents the spatial prior probability of whether the anchor point is within the labeled bounding box, P is the classification category of the predicted bounding box, is the intersection over union between the predicted bounding box and the labeled bounding box, and α, β are hyperparameters for balancing object classification and regression localization; Select the bounding box corresponding to the maximum matching degree as the prediction box of the bounding box, and then according to the prediction box and the corresponding bounding box b τ Calculate the loss function value: Among them, L Focaler-SIoU is the Focaler-SIoU loss, L dfl is the distribution focal loss, L cls is the classification loss, and η1, η2, and η3 are the gain coefficients corresponding to the three losses respectively; For each detected result image the loss function values of each prediction box and the corresponding annotation box are superimposed to obtain the total loss value of the training images in the current batch. Then, based on the obtained total loss value, the network model parameters are updated by backpropagation using the gradient descent method. Then, the next batch of training images is input repeatedly in this way until the YOLO11-AEDSF model converges to obtain the trained YOLO11-AEDSF model; (4.2.3) Use the layer adaptive pruning algorithm to prune and optimize the trained YOLO11-AEDSF model, removing redundant network connections and weights to obtain the optimized YOLO11-AEDSF model; (5) Detect the misassembly and missing assembly defects of the fasteners on the surface of the aircraft electrical distribution equipment; (5.1) Label each template image T i with a bounding box for each fastener therein, where the b-th bounding box is denoted as gt i,b , and each bounding box contains the bounding box coordinates gtbox i,b and the bounding box category gtCls i,b , where i is the angle index and b is the prediction box index; (5.2) Use the optimized YOLO11-AEDSF model to detect the image to be measured after exposure fusion. The type and coordinate information of the fasteners in it; (5.2.1), preprocess the image to be measured ; Keep the aspect ratio of the image to be measured, scale it to the input size of the YOLO11-AEDSF model, and then perform pixel value normalization to obtain the preprocessed image Then, perform pixel value normalization to obtain the preprocessed image (5.2.2), Input the preprocessed image into the backbone network of the YOLO11-AEDSF model, and output feature maps of three scales through the Conv convolution module, C3k2-Faster-CGLU module, SPPF module and C2PSA module, which are respectively denoted as the shallow feature map P3 i , the middle feature map P4 i , and the deep feature map P5 i ; (5.2.3), input the feature maps P3 i , P4 i , P5 i into the neck network of the YOLO11-AEDSF model, and output optimized feature maps of three scales through upsampling, Conv convolutional module and C3k2 Faster CGLU module, which are respectively denoted as the shallow optimized feature map P3' i , the middle optimized feature map P4' i , and the deep optimized feature map P5' i ; (5.2.4), input the fused feature maps P3' i , P4' i , P5' i into the detection head network of the YOLO11-AEDSF model to obtain a prediction result image containing multiple prediction boxes (5.3) Use the Linemod-2d method to match and predict the result image The template image with the highest similarity As the effective template image, and then obtain the effective template image Mapped to the predicted result image The rotation and translation parameters (x i , y i , θ i ), where (x i , y i ) are the translation parameters, corresponding to the pixel translation amounts along the horizontal and vertical directions respectively, and θ i Is the rotation parameter, representing the radian value of the counterclockwise rotation of the template image around its geometric center; (5.4), the effective template image and each annotation box are subjected to rotation and translation transformation using the rotation and translation parameters (x i , y i , θ i ) to obtain a corrected template image aligned with the predicted result image and corrected annotation boxes (5.5) Detect misassembly and missing assembly defects; (5.5.1) Calculate the geometric overlap degree of each prediction box in the predicted result image and the annotation box in the corrected template image using the intersection over union formula: Among them, is the area of the prediction box, is the area of the annotation box, A intersection is the intersection area between the prediction box and the annotation box; (5.5.2) Determine effective matching; Set the intersection over union threshold T IoU , if it satisfies then it is determined that the predicted bounding box and the annotated bounding box are effectively matched, otherwise it is an invalid match; (5.5.3) Determine misassembly and missing assembly defects; If the category of the predicted box with effective matching is the same as the category of the annotation box, it is determined that the fastener is assembled correctly; If the category of the predicted box with effective matching is not the type of assembly hole and is inconsistent with the category of the annotation box, it is determined that the fastener is misassembled; If the category of the predicted box with effective matching is the type of assembly hole and is inconsistent with the category of the annotation box, it is determined that the fastener is missing; If the labeled bounding box for an invalid match exists but the predicted bounding box does not, it is determined that the fastener is not installed.
2. The method for detecting missing or misinstalled fasteners of an aircraft power distribution equipment of an associated position type according to claim 1, characterized in that, The structure of the C3k2 Faster CGLU module is as follows: Replace the pointwise convolution in the FasterNet Block sparse calculation module with the ConvGLU convolutional gated linear unit to construct the Faster CGLU Block module; Replace the Bottleneck bottleneck block in the C3k module of the YOLO11 basic model with the Faster CGLU Block module to construct the C3k Faster CGLU module; In the C3k2 module of the YOLO11 basic model, when the module configuration parameter C3k = True, replace the C3k module with the C3k Faster CGLU module; when the module configuration parameter C3k = False, replace the Bottleneck bottleneck block with the Faster CGLU Block to construct the C3k2 Faster CGLU module.
3. A method for detecting missing or incorrect installation of fasteners in an aircraft power distribution equipment of an associated position type according to claim 1, characterized in that, The Focaler-SIoU loss is as follows: L Focaler-SIoU = L SIoU +(IoU - IoU Focaler ) Among them, L SIoU is the SIoU loss function, IoU is the intersection over union of the predicted bounding box and the annotated bounding box, and IoU Focaler is the IoU reconstructed by using the linear interval mapping method, which is expressed as: Among them, the thresholds g and u are used to control the IoU weight allocation.
4. A method for detecting missing or incorrect installation of fasteners in an aircraft power distribution equipment of an associated position type according to claim 1, characterized in that, The distribution focal loss is as follows: L dfl = L top + L bottom + L left + L right Among them, d top , d bottom , d left , d right represent the distances of the upper boundary, lower boundary, left boundary, and right boundary of the annotation box relative to the center of the annotation box; represents the integer part k ∈ {k top , k bottom , k left , k right} obtained by rounding down the distance d, where d ∈ {d top , d bottom , d left , d right}; is the residual weight of the distance d to the ceiling; S k represents the discrete probability that the distance from the boundary of the prediction box to the center of the annotation box is the integer k; S k+1 represents the discrete probability that the distance from the boundary of the prediction box to the center of the annotation box is the integer k + 1, represents rounding up, represents rounding down.
5. A method for detecting missing or incorrect installation of fasteners in an aircraft power distribution equipment of an associated position type according to claim 1, characterized in that The classification loss is as follows: Among them, y j represents the probability value that the annotation box belongs to the j-th category, and y j = 0, 1, represents the probability value that the prediction box belongs to the j-th category, where j = 1, 2,..., N; N is the number of fastener categories.
Citation Information
Cited By
Aircraft body part wrong and neglected loading detection method based on deep learning
CN121095583A
An aircraft body part missing and wrong installation detection method based on deep learning
CN121095583B