Air turbine starter appearance quality detection method based on multi-scale characteristics
By building a multi-scale feature processing module and a defect detection module, combining high-precision image acquisition and improved bounding box loss function, the timely and accurate problem of air turbine starter appearance quality detection is solved, efficient and accurate automated detection is achieved, and product reliability and operating performance is improved.
Patent Information
- Application Number
- CN202510000532.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
It is difficult for the prior art to achieve timely and accurate detection of the appearance quality of air turbine starters, especially in the case of complex structures and harsh working environments, traditional manual detection has problems such as misjudgment, long cycles and inability to achieve real-time monitoring.
Using a detection method based on multi-scale features, the appearance image data of the air turbine starter is obtained through a high-precision image acquisition device, a multi-scale feature processing module and a defect detection module are constructed, and a multi-scale feature processing module is extracted and fused with the MSConvFormerNet backbone network and feature fusion network, combined with the decoupling head structure and improved bounding box loss function, the object detection model is trained to achieve high accuracy and high reliability detection.
It realizes efficient and accurate automated inspection of the appearance elements of the air turbine starter, and can quickly complete on-site inspections, help assembly workers to discover problems in a timely manner, improve product reliability and operating performance, and has high accuracy, high stability and easy traceability.
Smart Images

Figure CN119941657A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of air turbine starter detection, and in particular relates to an air turbine starter appearance quality detection method based on multi-scale features. Background Art
[0002] With the growing demand for efficient, reliable and safe operation in the aviation industry, air turbine starters, as a key aviation equipment, ensure the appearance quality is an important part of ensuring the quality of equipment. Traditional turbine starter inspection methods mainly rely on manual observation and experience judgment. Manual inspection is easily affected by subjective factors, the results may be misjudged, the inspection cycle is long, and it is impossible to achieve real-time monitoring of the equipment. The detection effect of hidden faults or minor defects is poor. Due to the complex structure and harsh working environment of air turbine starters, some low-level appearance quality problems may also cause failures of air turbine starters, which in turn cause serious equipment damage and production stoppages. Therefore, it is of great significance to achieve timely and accurate inspection of air turbine starters.
[0003] In recent years, machine vision inspection technology has developed rapidly and is widely used in intelligent inspection in precision manufacturing production lines, online automated inspection of industrial product quality, intelligent robots, fine operations, engineering machinery and other fields. It has the characteristics of high intelligence, fast inspection speed and high accuracy. There are many assembly elements in air turbine starters, and the main appearance inspection elements involved are nameplates, rivets, locks, redundant objects, fuses, pins, etc. In order to ensure the reliability of the installation quality of the turbine starter, it is necessary to conduct assembly quality inspection on the multiple appearance elements of the air turbine starter. Due to the large number of inspections of each element, large shape differences, and diverse structures, there is currently a lack of effective appearance quality inspection methods for air turbine starters. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide an air turbine starter appearance quality detection method based on multi-scale features in view of the deficiencies of the above-mentioned prior art. By acquiring image data through a high-precision image acquisition device and constructing a detection model algorithm based on multi-scale features, high-accuracy and high-reliability detection of multiple appearance elements of the air turbine starter can be quickly achieved.
[0005] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0006] A method for detecting the appearance quality of an air turbine starter based on multi-scale features comprises the following steps:
[0007] Acquire appearance image data of the air turbine starter and perform data preprocessing and data enhancement processing, and divide the processed data into a training set, a verification set, and a test set to train, verify, and test an appearance quality detection model for the air turbine starter, wherein the detection model includes a multi-scale feature processing module and a defect detection module;
[0008] The multi-scale feature processing module extracts multi-scale features from the input image data through the MSConvFormerNet backbone network, fuses the extracted multi-scale features through the feature fusion network, and generates a feature map with multi-scale fusion information;
[0009] The defect detection module decouples the regression and classification tasks of the detection through a decoupling head structure, detects feature maps of fused information at different scales, and outputs the detection classification, bounding box prediction and confidence results;
[0010] During the detection model training process, an improved bounding box loss function is used to update parameters and weights to obtain a final detection model;
[0011] The trained detection model is used to detect the appearance quality of the air turbine starter and the detection results are output.
[0012] To optimize the above technical solutions, the specific measures taken also include:
[0013] The specific operation process of obtaining the appearance image data of the air turbine starter, performing data preprocessing and data enhancement processing, and dividing the processed data into a training set, a verification set, and a test set includes:
[0014] 1) Use a high-precision image acquisition device to obtain image data of the appearance of the air turbine starter. The image data contains the following detection elements: nameplate, rivets, lock plate, excess objects, fuses, and pins;
[0015] 2) Preprocessing the collected image data, including image denoising and histogram equalization, and performing data enhancement processing on the preprocessed image, including rotation, cropping, and flipping;
[0016] 3) Divide the processed image data into training set, verification set, and test set in a ratio of 8:1:1.
[0017] The above-mentioned high-precision image acquisition device includes a visual camera, a collaborative robotic arm, a servo turntable, a product tray, a base, and a detection background whiteboard. The detection background whiteboard is fixed on the servo turntable. The working environment is a darkroom. During work, the air turbine starter is moved to the servo turntable using the product tray. The collaborative robotic arm carries the visual camera and light source for shooting. The servo motor drives the servo turntable to rotate at any angle to adapt to the shooting requirements of different product positions and angles.
[0018] The above-mentioned MSConvFormerNet backbone network includes multiple groups of MSConvFormerBlocks, each group of MSConvFormerBlocks includes 2 MSConvFormerBlocks; MSConvFormerBlock is built based on the Transformer architecture, retaining the channel MLP, residual connection and normalization of the Transformer architecture, and replacing the self-attention module of mixed token information with the MSConv module for local feature extraction.
[0019] The above-mentioned EMFA attention mechanism is inserted before the MSConvFormerNet backbone network extracts the second-layer features and before inputting the defect detection module. The MSConvFormerNet backbone network introduces the EMFA attention mechanism before extracting the second-layer features to further increase the model's representation ability for multi-scale structural features. The extracted multi-scale features are fused using a feature fusion network with a feature pyramid structure to generate a feature map with multi-scale fusion information. The EMFA attention mechanism is used before inputting the defect detection module to increase the network's learning and attention to key features.
[0020] The formula used by the above MSConvFormerBlock is:
[0021] X=InEmb(I)
[0022] Y = MSConv(Norm(X)) + X
[0023] Z=σ(Norm(Y)W1)W2+Y
[0024] Where I represents the input image, X represents the input after batch embedding of the image, Y represents the output of feature extraction by multi-scale convolution, Z represents the output of MSConvFormerBlock, W1 and W2 represent learnable parameters, Norm represents the normalization layer, MSConv represents the MSConv module layer, InEmb represents batch embedding processing, and σ represents the nonlinear activation function.
[0025] The above-mentioned MSConv module uses a depth-wise convolution with a convolution kernel of 3×3 to extract local information, and then divides it into three groups to perform depth-wise convolution processing with convolution kernels of 1×5 and 5×1, 1×7 and 7×1, 1×13 and 13×1 respectively. The processing results of the three groups are added together and sent to the convolution with a convolution kernel of 1×1 for channel number adjustment. After adjustment, it is multiplied with the original input and the resulting feature map is output.
[0026] The above-mentioned EMFA attention mechanism divides the input feature map into n groups of sub-features along the channel direction, inputs the sub-features into three parallel paths to obtain the corresponding attention weights, and obtains the feature map with multi-scale convolutional attention through feature assignment; the three parallel paths include two 1×1 parallel branches and one 3×3 parallel branch; the 1×1 branch uses two one-dimensional global average pooling operations XAvgPool and YAvgPool to encode the channel information along the X and Y directions of space respectively, concatenates the encoded features, and shares the same 1×1 convolution, and then decomposes the features into two vectors, and performs Sigmo id operation is performed to assign channel weights to sub-features; the 3×3 parallel branches use 3×3DWConv depth-wise separable convolution to extract spatial features, and then use 1×1 point-by-point convolution to coordinate channel information to obtain features extracted by depth-wise separable convolution; the assigned sub-features of the 1×1 branch are interacted with the features extracted by the depth-wise separable convolution of the 3×3 branch across channels, and the multi-scale spatial structure information of the channel direction in each group is aggregated through multiplication to obtain different cross-channel interaction features between the 1×1 branch and the 3×3 branch parallel paths, and feature assignment is performed again with the grouped sub-features to output a feature map with multi-scale convolution attention.
[0027] The above-mentioned decoupling head structure includes a classification branch, a regression branch and a target branch; among them, the classification branch is aimed at the classification task and integrates the high-level feature map P l+1 The regression branch uses the rich semantic information in the low-level feature map P for regression tasks. l-1 The target branch is used to predict the rich edge texture information in the target recognition task, and the high-level feature map P is integrated. l+1 The prediction formula used by the classification branch, regression branch and target branch is:
[0028]
[0029] Where Concat(·), Upsample(·), and DConv(·) represent aggregation, upsampling, and downsampling operations, respectively. l is the input of the l-layer feature map, P cls and P reg Corresponding to the classification feature map and regression feature map output, P cls Represents the target feature map output.
[0030] The above improved bounding box loss function is:
[0031]
[0032] in, is the improved bounding box loss function; is the focusing factor; is the EIoU loss function regulated by the intensity factor; IoU is the intersection-over-union ratio of the predicted box and the real box; ω,ω gt is the width of the predicted box and the real box; h,h gt is the height of the predicted box and the real box; b and b gt denote the center points of the predicted box and the real box respectively, ρ(·) denotes the Euclidean distance, c denotes the diagonal length of the minimum outer bounding box, and c w and c h denote the width and height of the minimum outer bounding box respectively, and ∝ denotes an adjustable strength factor.
[0033] The present invention has the following beneficial effects:
[0034] The present invention uses a high-precision image acquisition device to collect appearance image data of an air turbine starter, and performs preprocessing and data enhancement on the data; constructs a multi-scale feature processing module of a target detection model, designs a backbone network to extract an input image as multi-scale features, uses a feature fusion network to perform feature fusion on features of different scales, and aggregates feature information of different dimensions; constructs a defect detection module, designs a decoupling detection head, and performs classification and positioning of targets on different feature maps; constructs a training parameter optimization module, designs an improved bounding box loss function, introduces an adjustable intensity factor α and a focusing mechanism in the training process, optimizes training parameters, and obtains a training model; uses the training model to perform automatic appearance detection on the air turbine starter; successfully realizes efficient and accurate automatic detection of appearance elements of air turbine starters at production sites, can quickly complete automatic detection at the site, helps assembly workers quickly discover problems, and takes timely measures to repair or replace them, thereby improving the reliability and operating performance of the product, and has the characteristics of high accuracy, high stability, and easy traceability.
[0035] The present invention can realize the appearance quality inspection of a type of air turbine starter after assembly, issue early warnings for identified defects, help to promptly discover appearance quality problems on the product surface on site, ensure assembly quality, and promote the realization of scientific research goals of advanced manufacturing and fault diagnosis of starters, and has great engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flow chart of an air turbine starter appearance quality inspection method based on multi-scale features;
[0037] Figure 2 is a schematic diagram of a high-precision image acquisition device proposed by the present invention;
[0038] Figure 3 It is the structure diagram of MSConvFormer Block proposed by the present invention;
[0039] Figure 4 This is the structural diagram of the EMFA attention mechanism proposed in the present invention;
[0040] Figure 5 It is a structural diagram of the decoupling detection head proposed by the present invention;
[0041] Figure 6 It is a structural diagram of the detection model proposed by the present invention;
[0042] Figure 7 is an operation flow chart of an embodiment of the present invention;
[0043] Figure 8 It is a rendering of the appearance quality inspection effect of the air turbine starter in the embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] Although the steps in the present invention are arranged with numbers, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used in this article involves and covers any and all possible combinations of one or more of the associated listed items.
[0046] like Figure 1 As shown, a method for detecting the appearance quality of an air turbine starter based on multi-scale features of the present invention comprises the following steps:
[0047] S1. Use a high-precision image acquisition device to obtain appearance image data of the air turbine starter, pre-process the data and perform data enhancement processing, and divide the processed data into a training set, a verification set, and a test set;
[0048] S2. Construct a multi-scale feature processing module, design the MSConvFormerNet backbone network, perform multi-scale feature extraction on the input image, use the feature fusion network to fuse feature maps of different scales, and generate a feature map with multi-scale information fusion.
[0049] S3. Build a defect detection module and design a decoupling head structure to decouple the regression and classification tasks of detection, detect the input features of targets of different scales, and output the detection classification, bounding box prediction, and confidence results.
[0050] S4. Construct a training process parameter optimization module, design an improved bounding box loss function, update the parameters and weights during the model training process, and obtain the final detection model.
[0051] S5. Use the trained model to test the appearance quality of the air turbine starter and output the test results.
[0052] In the embodiment, the specific operation process of step S1 includes:
[0053] 1) Combination Figure 2 As shown, a high-precision image acquisition device is used to obtain image data of the appearance of the air turbine starter, and the image data includes detection elements such as nameplate, rivets, lock plates, redundant objects, fuses, pins, etc.;
[0054] The high-precision image acquisition device is composed of hardware such as a visual camera, a collaborative robotic arm, a servo turntable, a product tray, and a base. The photo-taking process is carried out in a darkroom built with a sheet metal shell using a collaborative robotic arm carrying a camera and a dome light source. The starter product is manually carried to the turntable using a pallet. The servo motor can drive the turntable to rotate at any angle. Combined with the six-degree-of-freedom motion of the robotic arm, it can adapt to the different positions and angles of the product. In order to avoid interference between the visual camera and the product when taking pictures, a grating protection fence can be set in front of the circular dome light source. It is mainly composed of a ring of multiple groups of laser sensors, which controls the power supply of the equipment by sensing the distance between the product and the equipment; in order to further ensure the image quality, a detection background whiteboard can be installed on the turntable circular guide rail, fixed by an electromagnet, and can be disassembled and installed according to needs to reduce interference from reflections and background.
[0055] 2) Preprocess the collected image data, using operations such as image denoising and histogram equalization, and perform data enhancement on the preprocessed image, using operations such as rotation, cropping, flipping, etc.;
[0056] 3) After processing the images, the data set is divided into training set, verification set, and test set in a ratio of 8:1:1.
[0057] In the embodiment, step S2 constructs a multi-scale feature processing module, designs a MSConvFormerNet backbone network, uses the MSConv convolution module as the main structure to extract multi-scale features, enhances the feature extraction capability of the backbone network, and combines the long-distance feature modeling capability of Transformer.
[0058] The data set divided in step S1 is input into the multi-scale feature processing module, and the MSConvFormerNet backbone network is used to extract multi-scale features. The feature fusion network is used to fuse feature maps of different scales to generate a feature map with multi-scale information fusion.
[0059] An EMFA attention mechanism is introduced before the backbone network extracts the second-layer features to further enhance the model's ability to represent multi-scale structural features.
[0060] The extracted multi-scale features are fused using a feature pyramid structure to generate a feature map with multi-scale information fusion. Before entering the defect detection module, it passes through the EMFA attention mechanism to increase the network's learning and attention to key features.
[0061] The MSConvFormerNet backbone network consists of multiple groups of MSConvFormerBlocks, and each group of MSConvFormerBlocks consists of 2 MSConvFormerBlocks.
[0062] Combination Figure 3 As shown in the figure, MSConvFormerBlock is mainly composed of MSConv convolution block and channel MLP component. The input image I is batch embedded to obtain X, X is input to the MSConv convolution block to extract multi-scale features and then residually connected with X to obtain Y, which is input to the channel MLP component to obtain the output Z.
[0063] MSConvFormerBlock retains the channel MLP and residual connection of the Transformer architecture. Based on the Transformer architecture, the self-attention module of mixed token information is replaced with the MSConv convolution module with stronger local feature extraction capability. The MSConv convolution block consists of a DWConv depth-wise convolution, 3 multi-scale depth-wise convolution branches, and a Conv convolution with a convolution kernel of 1×1. This structure can significantly reduce the amount of model calculation while ensuring the effectiveness of multi-scale feature extraction. At the same time, MSConvFormerBlock retains the Transformer architecture's long-distance perception of input sequence elements and has the Transformer architecture's long-distance feature modeling capability.
[0064] The formula used by MSConvFormerBlock is:
[0065] X=InEmb(I)
[0066] Y = MSConv(Norm(X)) + X
[0067] Z=σ(Norm(Y)W1)W2+Y
[0068] Where I represents the input image, X represents the input after image batch embedding processing, Y represents the output of feature extraction by multi-scale convolution, Z represents the output of the MSConvFormer module, and W1 and W2 represent learnable parameters.
[0069] The MSConv module uses a depth-wise convolution with a convolution kernel of 3×3 to extract local information, and then divides it into three groups for depth-wise convolution with convolution kernels of 5×5, 7×7, and 13×3. In order to reduce model parameters and increase the ability of convolution to extract band features, the depth-wise convolutions of 5×5, 7×7, and 13×13 are split into two strip convolutions of 1×5 and 5×1, 1×7 and 7×1, and 1×13 and 13×1 respectively. The results of the three groups of scales are added together and sent to the convolution with a convolution kernel of 1×1 for channel number adjustment. After adjustment, they are multiplied with the original input to output the resulting feature map.
[0070] The MSConv convolution consists of a DWConv depth-wise convolution, three multi-scale depth-wise convolution branches, and a Conv convolution with a convolution kernel of 1×1. First, a DWConv with a convolution kernel of 3×3 is used to extract local information from the input, and then it is divided into three groups and input into three branches for depth-wise convolution to extract multi-scale features. The three branches use two strip convolutions of 1×5 and 5×1, 1×7 and 7×1, and 1×13 and 13×1, respectively. The results of the three groups of different scales are added together, and then input into the Conv convolution with a convolution kernel of 1×1 to coordinate the information between channels. Finally, the adjusted information is multiplied with the original input to output the result feature map. The convolution structure of MSConv fully represents and processes the features at each scale, and the design of deep convolution can reduce the amount of calculation of the model.
[0071] The EMFA attention mechanism uses three parallel paths to obtain the attention weights of grouped features. Two of the parallel paths are 1×1 branches, which use two one-dimensional global average pooling operations to encode channel information along two spatial directions; the third path is a 3×3 branch, which uses a depth-wise separable convolution to encode spatial information along the channel direction. The outputs of the two branches are fused through multiplication operations to generate attention weights, which are assigned to the input features to obtain output features with convolutional attention.
[0072] Combination Figure 4 As shown, the EMFA attention mechanism takes the input feature map Divide into n groups of sub-features X=[X0, X1, ..., X n ], The sub-features are input into three parallel paths to obtain the corresponding attention weights, two of which are located in the 1×1 branch and the third is located in the 3×3 parallel branch.
[0073] The 1×1 branch uses two one-dimensional global average pooling operations XAvgPool and YAvgPool to encode channel information along the X and Y directions of space respectively, concatenate the encoded features, and share the same 1×1 convolution, perform the same 1×1 convolution operation, balance the number of channels of the 1×1 branch, and ensure that the dimension of the 1×1 branch does not decrease. Then the features are decomposed into two vectors, and the Sigmoid operation is performed respectively, and the channel weights are assigned to the sub-feature map; the 3×3DWConv depth-wise convolution of the 3×3 branch is used to extract spatial features, and the 1×1 point-wise convolution is equivalent to coordinating channel information. Then the assigned sub-features of the 1×1 channel interact with the features extracted by the depth-wise separable convolution of the 3×3 branch across channels, and the multi-scale spatial structure information in the channel direction of each group is aggregated through multiplication operations to obtain different cross-channel interaction features between the two parallel paths, and the features are assigned again with the grouped sub-features, and the feature map with multi-scale convolution attention is output.
[0074] The EMFA attention mechanism processes features from the spatial and channel directions through two parallel branches, which increases the model's ability to capture spatial and channel information, pays attention to key features in the image, and improves the model's detection performance.
[0075] S2 uses a feature fusion network to fuse feature maps of different scales for the multi-scale features extracted by the backbone network, generating a feature map with multi-scale information fusion. The selectable neck networks include feature pyramid (FPN) and its variant structures, path aggregation network + feature pyramid network (PAN+FPN), etc. This module uses PAN+FPN to enhance the fused features through a series of bottom-up and top-down paths, providing richer feature information for subsequent defect detection.
[0076] In the embodiment, step S3 constructs a defect detection module, designs a decoupling head structure, decouples the regression and classification tasks of the detection, and uses parameters that are no longer shared to predict the feature maps that integrate different features, and outputs the detection classification, bounding box prediction and confidence results. This decoupling method essentially avoids the joint optimization of two competing tasks at the same time, and instead operates on different feature maps, which can improve the classification score and positioning accuracy.
[0077] For classification tasks, the rich semantic information in the high-level feature map is integrated for prediction; for regression tasks, the rich edge texture information in the low-level feature map is used for prediction. The formula used is:
[0078]
[0079] Where Concat(·), Upsample(·), and DConv(·) represent aggregation, upsampling, and downsampling operations, respectively. l Corresponding to the l-layer feature map input, P cls and P reg Corresponding to the classification feature map and regression feature map output, P cls Represents the target feature map output.
[0080] Furthermore, combined with Figure 5 As shown, for a certain layer of features P l (Feature P5 of large-scale targets) is used for classification tasks. l+1 (P6) is upsampled by 2 times, and the number of channels is adjusted through a convolution layer with a convolution kernel of 1×1, and then combined with P l (P5) is concatenated to obtain the feature map for the classification task This feature map integrates the semantic information of high-level feature maps, which can improve the accuracy of classification tasks. Finally, the decoding operation is performed to obtain the classification result. The formula used is:
[0081]
[0082] Furthermore, for the layer feature P l When performing regression tasks, P l (P5) features are upsampled by 2 times and then a convolution layer with a convolution kernel of 1×1 is used to adjust the number of channels, which is combined with the features P from the lower layer. l-1 (P4) performs an addition operation, then performs a 2-fold downsampling process, and finally combines the feature with P l (P5) Perform addition operation to obtain the feature map for regression task This feature map incorporates the richer texture feature details of the low-level feature map, which can improve the accuracy of the regression task. Finally, the decoding operation is performed to obtain the regression result. cls (·),f reg (·) is the feature mapping function used for classification and regression, which contains two 3×3 convolution operations. It is the last layer of the classification, regression, and target branches, used to decode features into classification, regression, and confidence scores. The formula used by the decoupling head is:
[0083]
[0084] The target branch and the classification branch have the same task objectives and use the same structure and parameters. The formula used is:
[0085]
[0086] In the embodiment, S4 constructs a training process parameter optimization module, designs an improved bounding box loss function, updates the parameters and weights in the model training process, and obtains the final detection model. It comprehensively considers the differences in overlapping areas, center points and side length factors, and introduces an adjustable strength factor ∝ on the basis of clearly measuring the differences in overlapping areas, center points and side length factors during target box regression, that is, an adjustable strength factor ∝ is introduced on the basis of the bounding box loss function EIOU. By adjusting ∝, the regression accuracy of the model at different levels of target boxes can be improved, and the flexibility and accuracy of the model bounding box regression can be increased. At the same time, in order to reduce the impact of low-quality samples on training, the focusing mechanism is optimized, and the prediction boxes of medium quality during training are better focused, solving the problem of imbalanced bounding box regression of samples of different qualities during training, and improving the overall performance of the model. The formula used is:
[0087]
[0088] in, is the improved bounding box loss function; is the focusing factor; is the EIoU loss function regulated by the intensity factor; IoU is the intersection-over-union ratio of the predicted box and the real box; ω,ω gt is the width of the predicted box and the real box; h,h gt is the height of the predicted box and the real box; b and b gt denote the center points of the predicted box and the real box respectively, ρ(·) denotes the Euclidean distance, c denotes the diagonal length of the minimum outer bounding box, and c w and c h Respectively represent the width and height of the minimum outer bounding box, and ∝ represents an adjustable strength factor. Specifically, ∝=3.
[0089] In the embodiment, step S5 uses the trained model to detect the appearance quality of the air turbine starter and outputs the detection result, specifically including: combining Figure 6As shown, through the previous steps, each module is designed to build a training model. The training model is mainly composed of the MSConvFormer backbone network, the neck network, and the detection decoupling head. The backbone network uses 4 groups of MSConvFormerBlocks for multi-scale feature extraction, and uses a group of Conv convolution + C2F components for feature extraction; the feature processing module uses PAN + FPN as the neck network to fuse and enhance the feature map of the backbone network; the EMFA attention mechanism is inserted before the backbone network extracts the second layer of features and before the input defect detection module, which can strengthen the focus on key features before feature fusion and improve the expressiveness of features; the defect detection module uses a structure composed of decoupling heads to output multi-scale results; the parameters of the model training process are optimized through the designed loss function to obtain a trained detection model. Combined with Figure 7 As shown, the air turbine starter appearance quality inspection process is as follows: at the beginning of the inspection task, first control the robot arm and turntable to return to zero position, then move the air turbine starter to be inspected onto the turntable, close the safety door, and work in a timely environment; the turntable first turns to the set position, and then the robot arm drives the camera to take pictures, repeating the photo action, the product completes a 360° rotation, the inspection model self-identifies and judges the photos taken, and outputs the inspection report; finally, control the equipment to return to zero position, open the safety door, move the product down, and the inspection task is completed. Figure 8 As shown, the detection model is used to detect the processed high-precision image to obtain an example of the result of the appearance detection elements.
[0090] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
[0091] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. A method for detecting the appearance quality of an air turbine starter based on multi-scale features, characterized in that: The following steps are involved: Acquire appearance image data of the air turbine starter and perform data preprocessing and data enhancement processing, and divide the processed data into a training set, a verification set, and a test set to train, verify, and test an appearance quality detection model for the air turbine starter, wherein the detection model includes a multi-scale feature processing module and a defect detection module; The multi-scale feature processing module extracts multi-scale features from the input image data through the MSConvFormerNet backbone network, fuses the extracted multi-scale features through the feature fusion network, and generates a feature map with multi-scale fusion information; The defect detection module decouples the regression and classification tasks of the detection through a decoupling head structure, detects feature maps of fused information at different scales, and outputs the detection classification, bounding box prediction and confidence results; During the detection model training process, an improved bounding box loss function is used to update parameters and weights to obtain a final detection model; The trained detection model is used to detect the appearance quality of the air turbine starter and the detection results are output.
2. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 1, characterized in that: The specific operation process of obtaining the appearance image data of the air turbine starter, performing data preprocessing and data enhancement processing, and dividing the processed data into a training set, a validation set, and a test set includes: 1) Use a high-precision image acquisition device to obtain image data of the appearance of the air turbine starter. The image data contains the following detection elements: nameplate, rivets, lock plate, excess objects, fuses, and pins; 2) Preprocessing the collected image data, including image denoising and histogram equalization, and performing data enhancement processing on the preprocessed image, including rotation, cropping, and flipping; 3) Divide the processed image data into training set, verification set, and test set in a ratio of 8:1:
1.
3. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 2, characterized in that: The high-precision image acquisition device includes a visual camera, a collaborative robotic arm, a servo turntable, a product tray, a base, and a detection background whiteboard. The detection background whiteboard is fixed on the servo turntable. The working environment is a darkroom. During work, the air turbine starter is moved to the servo turntable using the product tray. The collaborative robotic arm carries the visual camera and light source for shooting. The servo motor drives the servo turntable to rotate to any angle to adapt to the shooting requirements of different product positions and angles.
4. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 1, characterized in that: The MSConvFormerNet backbone network includes multiple groups of MSConvFormerBlocks, each group of MSConvFormerBlocks includes 2 MSConvFormerBlocks; MSConvFormerBlock is built based on the Transformer architecture, retains the channel MLP, residual connection and normalization of the Transformer architecture, and replaces the self-attention module of mixed token information with the MSConv module for local feature extraction.
5. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 4, characterized in that: The EMFA attention mechanism is inserted before the MSConvFormerNet backbone network extracts the second-layer features and before inputting the defect detection module. The MSConvFormerNet backbone network introduces the EMFA attention mechanism before extracting the second-layer features, further increasing the model's ability to represent multi-scale structural features, and fuses the extracted multi-scale features using a feature fusion network with a feature pyramid structure to generate a feature map with multi-scale fusion information. The EMFA attention mechanism is used before inputting the defect detection module to increase the network's learning and attention to key features.
6. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 4, characterized in that: The formula used by the MSConvFormerBlock is: X=InEmb(I) Y = MSConv(Norm(X)) + X Z=σ(Norm(Y)W1)W2+Y Where I represents the input image, X represents the input after batch embedding of the image, Y represents the output of feature extraction by multi-scale convolution, Z represents the output of MSConvFormerBlock, W1 and W2 represent learnable parameters, Norm represents the normalization layer, MSConv represents the MSConv module layer, InEmb represents batch embedding processing, and σ represents the nonlinear activation function.
7. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 4, characterized in that: The MSConv module uses a depth-wise convolution with a convolution kernel of 3×3 to extract local information, and then divides it into three groups to perform depth-wise convolution processing with convolution kernels of 1×5 and 5×1, 1×7 and 7×1, 1×13 and 13×1 respectively. The processing results of the three groups are added together and sent to the convolution with a convolution kernel of 1×1 for channel number adjustment. After adjustment, it is multiplied with the original input to output the result feature map.
8. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 5, characterized in that: The EMFA attention mechanism divides the input feature map into n groups of sub-features along the channel direction, inputs the sub-features into three parallel paths to obtain the corresponding attention weights, and obtains the feature map with multi-scale convolutional attention through feature assignment; the three parallel paths include two 1×1 parallel branches and one 3×3 parallel branch; the 1×1 branch uses two one-dimensional global average pooling operations XAvgPool and YAvgPool to encode the channel information along the X and Y directions of space respectively, concatenates the encoded features, and shares the same 1×1 convolution, then decomposes the features into two vectors, performs Sigmoid operations respectively, and assigns the channel weights to the sub-features; the 3×3 parallel branch uses 3×3DWConv depth-wise separable convolution to extract spatial features, and then uses 1×1 point-wise convolution to coordinate channel information to obtain features extracted by depth-wise separable convolution; The assigned sub-features of the 1×1 branch are interacted with the features extracted by the depthwise separable convolution of the 3×3 branch across channels, and the multi-scale spatial structure information in the channel direction of each group is aggregated through multiplication operation to obtain different cross-channel interaction features between the 1×1 branch and the 3×3 branch parallel paths, which are then assigned features again with the grouped sub-features to output a feature map with multi-scale convolutional attention.
9. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 1, characterized in that: The decoupling head structure includes a classification branch, a regression branch and a target branch; wherein the classification branch is for classification tasks, integrating high-level feature maps P l+1 The regression branch uses the rich semantic information in the low-level feature map P for regression tasks. l-1 The target branch is used to predict the rich edge texture information in the target recognition task, and the high-level feature map P is integrated. l+1 The prediction formula used by the classification branch, regression branch and target branch is: Where Concat(·), Upsample(·), and DConv(·) represent aggregation, upsampling, and downsampling operations, respectively. l is the input of the l-layer feature map, P cls and P reg Corresponding to the classification feature map and regression feature map output, P cls Represents the target feature map output.
10. The method for detecting the appearance quality of an air turbine starter based on multi-scale features according to claim 1, characterized in that: The improved bounding box loss function is: in, is the improved bounding box loss function; is the focusing factor; is the EIoU loss function regulated by the intensity factor; IoU is the intersection-over-union ratio of the predicted box and the real box; ω,ω gt is the width of the predicted box and the real box; h,h gt is the height of the predicted box and the real box; b and b gt denote the center points of the predicted box and the real box respectively, ρ(·) denotes the Euclidean distance, c denotes the diagonal length of the minimum outer bounding box, and c w and c h denote the width and height of the minimum outer bounding box respectively, and ∝ denotes an adjustable strength factor.
Citation Information
Cited By
Method, device and equipment for detecting port state of optical cable cross connecting cabinet and storage medium
CN121053138A