Rotor blade size measurement method based on computer vision
Through a computer vision-based method, a multi-scale CNN network using adaptive Canny edge detection and convolutional block attention modules solves the efficiency and accuracy problems in rotor size measurement, realizes high-precision measurement of complex shape rotors, and improves the efficiency and quality of industrial production.
Patent Information
- Application Number
- CN202510649695.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing rotor size measurement technology has obvious defects in efficiency, accuracy, anti-interference ability and adaptability to complex-shaped rotors, and it is difficult to meet the high-precision and high-efficiency industrial production needs.
Using a computer vision-based method, the rotor blade image is obtained by X-rays, edge detection is performed through the adaptive Canny edge detection module combined with the multi-scale CNN network of the convolutional block attention module, and the composite loss function and weighted cross entropy loss function are optimized to optimize the model training, the edge features of the rotor blade are extracted, and the size is calculated using the conversion formula.
It improves image quality, enhances the accuracy of edge detection and generalization capabilities of the model, reduces noise interference, and improves the accuracy and efficiency of rotor blade size measurement.
Smart Images

Figure CN120411069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of non-contact measurement of rotor blade dimensions, and particularly to a method for measuring rotor blade dimensions based on computer vision. Background Art
[0002] In the current industrial production and manufacturing fields, as the core component of many mechanical devices, the size accuracy of the rotor has a crucial impact on the overall performance, operating stability, and service life of the device. Accurately measuring the rotor size is a key link to ensure the quality of the device, improve production efficiency, and achieve automated manufacturing.
[0003] Existing rotor size measurement methods mainly include two categories: contact measurement and non-contact measurement. Contact measurement often uses tools such as micrometers, screw micrometers, coordinate measuring machines, or special measuring tools. Such methods rely on manual operation, not only with low measurement efficiency, but also the measurement personnel are prone to fatigue during long-term work, resulting in increased work intensity. Moreover, manual operation is difficult to avoid subjective measurement errors, and the differences in operation habits and techniques of different measurement personnel will also lead to large discreteness of measurement results.
[0004] In the aspect of non-contact measurement, technologies such as laser displacement sensors are mostly used. Although this method avoids some disadvantages of contact measurement, it has strict requirements for the measurement environment and is easily interfered by factors such as light, temperature, and humidity in the environment. For example, in a high-temperature environment, the propagation path of the laser will shift, resulting in large deviations in the measurement results. In addition, when measuring rotors with complex shapes, the laser displacement sensor has measurement blind spots, and for rotors with structures such as concave and deep grooves, complete size data cannot be obtained.
[0005] There are also some measurement methods based on machine vision that have been applied, but most of them have deficiencies. The image acquisition accuracy of some machine vision measurement systems is limited, and it is difficult to provide sufficiently accurate image data for size calculation when measuring small rotors or rotors with extremely high size accuracy requirements. When some algorithms process image edge detection, for rotor images in complex backgrounds, situations such as edge misjudgment and missed judgment are likely to occur. For example, when measuring a rotor with stains, scratches, and other interferences on its surface, traditional edge detection algorithms will misjudge the edge of the stain as the actual edge of the rotor, resulting in an increase in size measurement error.
[0006] In summary, the existing rotor size measurement technologies have obvious defects in terms of efficiency, accuracy, anti-interference ability, and adaptability to rotors with complex shapes, and it is difficult to meet the growing high-precision and high-efficiency industrial production requirements. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for measuring rotor blade dimensions based on computer vision to solve the above technical problems.
[0008] To achieve the above object, the present invention provides a method for measuring the size of a rotor blade based on computer vision, comprising the following steps:
[0009] S1. Transmit the engine using X-rays to obtain an image of the rotor blade and perform preprocessing;
[0010] S2. Use a multi-scale CNN network with an adaptive Canny edge detection module combined with a convolutional block attention module to perform edge detection on the preprocessed rotor blade image to obtain a binary image;
[0011] S3. Statistically analyze the gray values of the rotor blade and draw a gray curve;
[0012] S4. Take the mutation points in the gray curve as edge points and use a conversion formula to calculate the actual distance between adjacent edge points, thereby obtaining the size of the rotor blade.
[0013] Preferably, the preprocessing described in step S1 includes using median filtering to remove white noise in the rotor blade image and adopting an unsharp masking algorithm to enhance the detail features of the rotor blade image.
[0014] Preferably, step S2 specifically includes the following steps:
[0015] S21. Construct a CACnet network: The CACnet network is based on a multi-scale CNN network, and the multi-scale CNN network includes an encoder, a neck, and a decoder, where an adaptive Canny edge detection module combined with a convolutional block attention module is added to the neck;
[0016] S22. Train the CACnet network and use a composite loss function to supervise the training process of the model until convergence;
[0017] S23. Input the preprocessed rotor blade image into the encoder of the CACnet network, and extract multi-scale features through a combination of multiple convolutional layers and pooling layers;
[0018] At the same time, use the adaptive Canny edge detection module to extract the edge features of the preprocessed rotor blade image;
[0019] S24. Input the multi-scale features extracted by the encoder and the edge features extracted by the adaptive Canny edge detection module into the convolutional block attention module for fusion to obtain a fused feature map;
[0020] S25. Input the fused feature map into the decoder, and use the decoder to perform upsampling and feature fusion on the fused feature map to gradually restore the spatial resolution of the rotor blade image and generate a final edge image;
[0021] S26. Convert the finally generated edge image into a binary image according to the detection result of the adaptive Canny edge detection module.
[0022] Preferably, the composite loss function L described in step S22 floss has the following expression:
[0023]
[0024] In the formula, L1 represents the loss function for measuring the gap between non-edge pixel points and edge pixel points; L2 represents the loss function for measuring the edge connection loss; represents the i-th prediction result; y represents the true label; represents the prediction result after fusion at different stages;
[0025] Among them,
[0026] L1 = w side ·L side + w fuse L fuse (P, Y) (2);
[0027]
[0028] In the formula, w side and w fuse respectively represent the weights of the individual outputs of each stage and the weight of the final output; L side and L fuse respectively represent the loss functions of the individual outputs of each stage and the loss function of the final output; P and Y respectively represent the prediction result and the true label of the model; λ1 and λ2 both represent weight factors; L bdry and L tex both represent the boundary tracking fusion function; E represents the set of edge pixel points among all pixel points; and both represent image patches; L p represents the edge point geometry in the image patch ; L ce , and are all weighted cross-entropy loss functions L(·);
[0029] And the expression of the weighted cross-entropy loss function L(·) is as follows:
[0030]
[0031] Y + = {i|y i ∈Y, y i > δ} (7);
[0033] Y - ={i|y i ∈Y, y i =0} (8);
[0034] In the formula, Y represents all pixel points, that is, the set Y + ∪Y - ; y i represents the i-th pixel point; Y + and Y - represent edge pixel points and non-edge pixel points respectively; α represents the proportion of negative samples in the set Y + ∪Y - ; β represents the hyperparameter pixel point, which is used to balance the importance of edge and non-edge pixel points; δ represents the set threshold value.
[0035] Preferably, the step of using the adaptive Canny edge detection module to extract the edge features of the preprocessed rotor blade image in step S23 specifically includes the following steps:
[0036] Step 1: Use the Gaussian filter G(x, y) to denoise the input rotor blade image I(x, y) to obtain a smoothed image I Smooth (x, y):
[0037] G(x, y)=exp[-(x 2 +y 2 ) / 2σ1] / 2πσ1 2 (9);
[0038] I Smooth (x, y)=I(x, y)*G(x, y) (10);
[0039] In the formula, σ1 represents the standard deviation; (x, y) represents the position coordinates of the pixel point;
[0040] Step 2: Use the first-order partial derivative operator to calculate the gradient value G Smooth in the horizontal direction and the gradient value G x in the vertical direction for each pixel point on the smoothed image I y ;
[0041] Step 3: Based on the gradient value G x in the horizontal direction and the gradient value G y in the vertical direction, calculate the gradient amplitude G and the gradient direction θ to replace the original gray value to obtain the image E(G, θ):
[0042]
[0043] Step 4: Use non-maximum suppression to perform edge localization on the image E(G,θ);
[0044] Step 5: Divide the image E(G,θ) into multiple sub-intervals, statistically analyze the probability distribution P(i) of the gray values within each sub-interval, and by setting a high threshold T high divide the pixel points into two categories C1 and C2:
[0045]
[0046] where G (i,y) represents the gradient magnitude of the pixel point with coordinates (x,y);
[0047] Step 6: Calculate the probabilities and means of the two categories of pixel points C1 and C2:
[0048]
[0049] where ω1 and ω2 respectively represent the probabilities of the two categories of pixel points C1 and C2; L represents the total number of image gray levels; μ1 and μ2 respectively represent the means of the two categories of pixel points C1 and C2;
[0050] Step 7: Calculate the variance υ between the two categories of pixel points C1 and C2 based on the probabilities and means of the two categories of pixel points:
[0051] υ = ω1(μ1 - μ2) 2 + ω2(μ2 - μ1) 2 (18);
[0052] Step 8: Traverse all the set high thresholds T high such that the variance υ is the largest, and take the set high threshold T high at this time as the optimal high threshold T best-high ;
[0053] Step 9: Set a low threshold T low :
[0054]
[0055] Step 10: Compare the pixels in the image E(G,θ) with the optimal high threshold T best-high and the low threshold T low respectively to classify the pixels in the image E(G,θ) and obtain edge features.
[0056] Preferably, in step S24, the use of the convolutional block attention module includes a channel attention module and a spatial attention module, where the expression of the channel attention module is as follows:
[0057]
[0058] where M c (F) represents the output of the channel attention module; W1 and W both represent weight matrices in the multi-layer perceptron; σ represents the Sigmoid activation function; and represent the spatial features generated through the max-pooling channel and the average-pooling channel respectively;
[0059] The expression of the spatial attention module is as follows:
[0060]
[0061] where M s (F) represents the output of the spatial attention module; f 7×7 represents a 7×7 convolutional kernel; and represent the features generated through max-pooling and average-pooling respectively;
[0062] It specifically includes the following steps:
[0063] S241. Sequentially infer the 1D channel attention map M c ∈R C×1×1 and the 2D spatial attention map M s ∈R 1×H×W :
[0064]
[0065] where F' represents the feature map processed and weighted by the channel attention module;
[0066] Among them, in the channel attention module, the input feature F ∈ R C×H×W passes through the max-pooling channel and the average-pooling channel respectively to generate the spatial features and Then, and are passed into the shared multi-layer perceptron to generate the 1D channel attention map M c ∈R C×1×1 , and M c is multiplied by the input feature F to obtain the feature F c after channel dimension fusion;
[0067] In the spatial attention module, the input feature F ∈ R C×H×W propagates along the axis and sequentially passes through max-pooling and average-pooling to generate the features and Then, it passes through the convolutional layer to generate the 2D spatial attention map M s ∈R 1 ×H×W ;
[0068] S242, the feature F after channel dimension fusion c With 2D spatial attention map M s Perform element-by-element multiplication to obtain the final fusion feature map.
[0069] Preferably, the decoder described in step S25 includes a DBLOCK decoding module, a UBlock module and a ConCat module, wherein the DBLOCK decoding module mainly consists of a Conv2d convolution layer and a Gaussian error linear unit activation function, the Conv2d convolution layer is used to extract edge features of the rotor blade image, and the Gaussian error linear unit activation function is used to perform normalization processing;
[0070] The output feature expression of the DBLOCK decoding module is as follows:
[0071]
[0072] Where, f n+1 represents the output features of the DBLOCK decoding module at the n+1th iteration; and Respectively represent the height and width of the fused feature map; C represents the number of channels of the fused feature map;
[0073] The UBlock module uses the Conv2d convolution layer to learn the fused feature map, and then uses the ConvTranspose2d transposed convolution to upsample the fused feature map and restore the resolution to gradually generate the edge image;
[0074] The ConCat module is used to rearrange the edge map pixels generated by the UBlock module at different stages onto the new image using PixelShuffle to obtain the final edge image.
[0075] Preferably, the conversion formula in step S4 is as follows:
[0076] L = δ*N (25);
[0077] Where L represents the distance between the two nearest mutation points; δ represents the conversion ratio coefficient; and N represents the number of pixels between the two mutation points.
[0078] Therefore, the present invention adopts the above-mentioned rotor blade size measurement method based on computer vision, which has the following beneficial effects:
[0079] 1. Image quality optimization: Use a Gaussian filter to reduce noise on the input rotor blade image, effectively removing noise interference to obtain a smooth image, reducing the impact of noise on subsequent edge detection and other operations, improving the overall image quality, and laying the foundation for accurate blade feature extraction;
[0080] 2. Precise edge detection: The adaptive Canny edge detection module can accurately extract the edge features of the rotor blade through multi-step processing, precisely locate the blade edge, and improve the accuracy of edge detection.
[0081] 3. Enhanced feature attention: The convolutional block attention module adaptively refines the input feature map from both channel and spatial dimensions by sequentially inferring 1D channel attention maps and 2D spatial attention maps, enabling the model to pay more attention to features related to blade edge detection, suppressing irrelevant noise and background information, and enhancing the model's sensitivity and detection accuracy to blade features.
[0082] 4. Efficient model training: By using composite loss functions and weighted cross-entropy loss functions, etc., and reasonably setting weights, it can effectively measure the gap between non-edge and edge pixels, balance the importance of edge and non-edge pixel points, guide the model to pay more attention to key information during training, optimize model parameters, improve the generalization ability of the model and the detection accuracy of rotor blade images, and make the model training effect better.
[0083] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0084] Figure 1 It is a flowchart of a method for measuring the size of a rotor blade based on computer vision according to the present invention.
[0085] Figure 2 It is a schematic diagram of the image acquisition structure of a method for measuring the size of a rotor blade based on computer vision according to the present invention.
[0086] Figure 3 It is a CACnet network architecture diagram of a method for measuring the size of a rotor blade based on computer vision according to the present invention.
[0087] Figure 4 It is a comparison diagram of the simulation experiment according to the present invention.
[0088] Figure 5 It is a grayscale curve diagram according to the present invention. Detailed Embodiments
[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clearly understood, the following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining the embodiments of the present invention and are not intended to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout.
[0090] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0091] The following elaborates in detail on the implementation manners of the present invention in conjunction with the accompanying drawings.
[0092] As Figure 1 shown, a method for measuring the size of a rotor blade based on computer vision includes the following steps:
[0093] S1. Use X-rays as Figure 2 shown to transmit the engine, obtain the rotor blade image, and perform preprocessing;
[0094] The preprocessing described in step S1 includes using median filtering to remove white noise in the rotor blade image (Radius is selected as 2 pixels), and adopting an unsharp masking algorithm (the range of Raidus (Sigma) is set to 3 - 5, and the range of mask weight is set to 0.4 - 0.6) to enhance the detailed features of the rotor blade image.
[0095] S2. Use a multi-scale CNN network with an adaptive Canny edge detection module combined with a convolutional block attention module to perform edge detection on the preprocessed rotor blade image to obtain a binary image;
[0096] Step S2 specifically includes the following steps:
[0097] S21. Construct a CACnet network: The CACnet network has a multi-scale CNN network as the backbone, and the multi-scale CNN network includes an encoder, a neck, and a decoder, where an adaptive Canny edge detection module combined with a convolutional block attention module is added to the neck;
[0098] S22. Train the CACnet network and use the composite loss function to supervise the training process of the model until convergence;
[0099] The composite loss function L described in step S22 floss The expression is as follows:
[0100]
[0101] In the formula, L1 represents the loss function used to measure the gap between non-edge pixel points and edge pixel points; L2 represents the loss function used to measure the edge connection loss; represents the i-th prediction result; y represents the true label; represents the prediction result after fusion at different stages;
[0102] Among them,
[0103] L1 = w side ·L side +w fuse L fuse (P, Y) (2);
[0104]
[0105] In the formula, w side and w fuse respectively represent the weights of the individual outputs at each stage and the weight of the final output; L side and L fuse respectively represent the loss functions of the individual outputs at each stage and the loss function of the final output; P and Y respectively represent the prediction result and the true label of the model; λ1 and λ2 both represent weight factors; L bdry and L tex both represent the boundary tracking fusion function; E represents the set of edge pixel points among all pixel points; and both represent image patches; L p represents the edge point geometry in the image patch ; L ce , and are all weighted cross-entropy loss functions L(·);
[0106] And the expression of the weighted cross-entropy loss function L(·) is as follows:
[0107]
[0108] Y + = {i|y i ∈ Y, y i > δ} (7);
[0109] Y- = {i | y i ∈ Y, y i = 0} (8);
[0110] In the formula, Y represents all pixel points, that is, the set Y + ∪ Y - ; y i represents the i-th pixel point; Y + and Y - represent edge pixel points and non-edge pixel points respectively; α represents the proportion of negative samples in the set Y + ∪ Y - ; β represents the hyperparameter pixel point, which is used to balance the importance of edge and non-edge pixel points; δ represents the set threshold.
[0111] S23. Input the preprocessed rotor blade image into the encoder of the CACnet network, and extract multi-scale features through the combination of multiple convolutional layers and pooling layers;
[0112] At the same time, use the adaptive Canny edge detection module to extract the edge features of the preprocessed rotor blade image;
[0113] The using of the adaptive Canny edge detection module to extract the edge features of the preprocessed rotor blade image described in step S23 specifically includes the following steps:
[0114] Step 1. Use the Gaussian filter G(x, y) to denoise the input rotor blade image I(x, y) to obtain the smoothed image I Smooth (x, y):
[0115] G(x, y) = exp[-(x 2 + y 2 ) / 2σ1] / 2πσ1 2 (9);
[0116] I Smooth (x, y) = I(x, y) * G(x, y) (10);
[0117] In the formula, σ1 represents the standard deviation; (x, y) represents the position coordinates of the pixel point;
[0118] Step 2. Use the first-order partial derivative operator to calculate the gradient value G Smooth in the horizontal direction and the gradient value G x in the vertical direction for each pixel point on the smoothed image I y ;
[0119] Step 3. Based on the gradient value G x in the horizontal direction and the gradient value G y, calculate the gradient magnitude G and the gradient direction θ, and obtain the image E(G, θ) by replacing the original grayscale value:
[0120]
[0121] Step Four: Use the non-maximum suppression method to perform edge localization on the image E(G, θ);
[0122] Step Five: Divide the image E(G, θ) into multiple sub-intervals, statistically analyze the probability distribution P(i) of the grayscale values within each sub-interval, and by setting a high threshold T high divide the pixel points into two categories C1 and C2:
[0123]
[0124] where G (i,y) represents the gradient magnitude of the pixel point with coordinates (x, y);
[0125] Step Six: Calculate the probabilities and means of the two categories of pixel points C1 and C2:
[0126]
[0127] where ω1 and ω2 respectively represent the probabilities of the two categories of pixel points C1 and C2; L represents the total number of image gray levels; μ1 and μ2 respectively represent the means of the two categories of pixel points C1 and C2;
[0128] Step Seven: Calculate the variance υ between the two categories of pixel points C1 and C2 based on the probabilities and means of the two categories of pixel points:
[0129] υ = ω1(μ1 - μ2) 2 + ω2(μ2 - μ1) 2 (18);
[0130] Step Eight: Traverse all the set high thresholds T high such that the variance υ is maximized (the larger the variance υ, the more obvious the difference between the background and the edge), and take the set high threshold T high at this time as the optimal high threshold T best-high ;
[0131] Step Nine: Set a low threshold T low :
[0132]
[0133] Step Ten: Compare the pixels in the image E(G, θ) with the optimal high threshold T best-high and the low threshold T low respectively to classify the pixels in the image E(G, θ) and obtain the edge features.
[0134] In this embodiment, the classification strategy is as follows: Strong edge pixels: Pixels with a gradient magnitude greater than the optimal high threshold T best-high are classified as strong edge pixels, and the strong edge pixels are regarded as real edge pixels.
[0135] Weak edge pixels: Pixels with a gradient magnitude between the low threshold T low and strong edge pixels are classified as weak edge pixels. For weak edge pixels, if they are connected to strong edge pixels, they are considered edge pixels; otherwise, they are suppressed.
[0136] Non-edge pixels: Pixels with a gradient magnitude less than the low threshold T low are classified as non-edge pixels and are suppressed.
[0137] Based on the above edge pixels, edge features can be obtained.
[0138] S24. Input the multi-scale features extracted by the encoder and the edge features extracted by the adaptive Canny edge detection module into the convolutional block attention module for fusion to obtain a fused feature map;
[0139] In step S24, the convolutional block attention module includes a channel attention module and a spatial attention module. The expression of the channel attention module is as follows:
[0140]
[0141] In the formula, M c (F) represents the output of the channel attention module; W1 and W both represent weight matrices in the multi-layer perceptron; σ represents the Sigmoid activation function; and respectively represent the spatial features generated through the maximum pooling channel and the average pooling channel;
[0142] The expression of the spatial attention module is as follows:
[0143]
[0144] In the formula, M s (F) represents the output of the spatial attention module; f 7×7 represents a 7*×7 convolutional kernel; and respectively represent the features generated through maximum pooling and average pooling;
[0145] It specifically includes the following steps:
[0146] S241. Sequentially infer the 1D channel attention map M c ∈R C×1×1 and the 2D spatial attention map Ms ∈R 1×H×W :
[0147]
[0148] In the formula, F' represents the feature map processed and weighted by the channel attention module;
[0149] Among them, in the channel attention module, the input feature F ∈ R C×H×W passes through the max-pooling channel and the average-pooling channel respectively to generate spatial features and Then and are fed into a shared multi-layer perceptron to generate a 1D channel attention map M c ∈R C×1×1 , multiply M c with the input feature F to obtain the feature F after channel dimension fusion c ;
[0150] In the spatial attention module, the input feature F ∈ R C×H×W propagates along the axis and passes through max-pooling and average-pooling in sequence to generate features and Then, it passes through a convolutional layer to generate a 2D spatial attention map M s ∈R 1 ×H×W ;
[0151] S242. Multiply the feature F after channel dimension fusion c with the 2D spatial attention map M s element-wise to obtain the final fused feature map.
[0152] Preferably, the decoder described in step S25 includes a DBLOCK decoding module, a UBlock module, and a ConCat module. Among them, the DBLOCK decoding module mainly consists of a Conv2d convolutional layer and a Gaussian error linear unit activation function. The Conv2d convolutional layer is used to extract the edge features of the rotor blade image, and the Gaussian error linear unit activation function is used for normalization processing;
[0153] The output feature expression of the DBLOCK decoding module is as follows:
[0154]
[0155] In the formula, f n+1 represents the output feature of the DBLOCK decoding module in the (n + 1)-th iteration; and represent the height and width of the fused feature map respectively; C represents the number of channels of the fused feature map;
[0156] The UBlock module uses the Conv2d convolutional layer to learn the fused feature map, and then uses the ConvTranspose2d transposed convolution to upsample the fused feature map to restore the resolution, so as to gradually generate the edge image;
[0157] The ConCat module is used to rearrange the edge image pixels generated by the UBlock module at different stages to a new image using PixelShuffle to obtain the final edge image.
[0158] S25. Input the fused feature map into the decoder, and use the decoder to upsample and fuse the features of the fused feature map to gradually restore the spatial resolution of the rotor blade image and generate the final edge image;
[0159] S26. According to the detection result of the adaptive Canny edge detection module, convert the finally generated edge image into a binary image.
[0160] S3. Statistically analyze the gray values of the rotor blades and draw a gray curve as shown in Figure 5 shown;
[0161] S4. Take the mutation points in the gray curve as edge points, and use the conversion formula to calculate the actual distance between adjacent edge points, so as to obtain the rotor blade size.
[0162] The conversion formula described in step S4 is as follows:
[0163] L = δ * N (25);
[0164] In the formula, L represents the distance between the two nearest mutation points, with the unit of mm; δ represents the conversion ratio coefficient, and δ takes 0.127 in this embodiment; N represents the number of pixels between the two mutation points.
[0165] Simulation experiment
[0166] The hardware and software environment configurations of this experiment are shown in Table 1, and the hyperparameters during the experiment are shown in Table 2.
[0167] Table 1 Hardware and software environment configuration
[0168] Experimental equipment Type / Version number GPU Nvidia Geforce 3090*2 Experimental system Windows11 Development environment Pycharm2023 python 3.8.20 Pytorch + cuda 2.4.1+11.6 torchvision 0.19.1
[0169] Table 2 Hyperparameter settings
[0170] Parameter Value / Category Image size 300*300 Optimizer Adam Batch size 8 Initial learning rate 8e-4 Number of iterations 50 Learning rate adjustment method Piecewise learning rate decay strategy Random seed 1021
[0171] The segmentation accuracy of the multi-scale CNN network is evaluated using the Optimal Dataset Scale (ODS) and the Optimal Image Scale (OIS), and the performance of the multi-scale CNN network is evaluated using MSE. The expressions are as follows:
[0172]
[0173] In the formula, P θ and R θ represent the accuracy and recall rate under the global optimal threshold θ respectively; N represents the number of images in the dataset; and represent the accuracy and recall rate under the optimal threshold θ i of image i respectively; n represents the number of samples participating in the calculation;
[0174] The roles of the adaptive Canny edge detection module (CIA), the convolutional block attention module (CBAM), and the composite loss function L floss are verified through ablation experiments, and the results are shown in Table 3.
[0175] Table 3 Results of ablation experiments
[0176] Base CIA CBAM <![CDATA[L floss > ODS OIS MSE Params √ 0.724 0.753 0.574 59K √ √ 0.683 0.715 0.682 59K √ √ 0.753 0.786 0.258 63K √ √ √ 0.802 0.814 0.128 63K √ √ √ √ 0.825 0.831 0.088 63k
[0177] As can be seen from Table 3, the BASE model (basic model) that only contains an encoder and a decoder achieved 0.724 on ODS, but 0.753 on OIS. This shows that the basic model has a certain generalization ability and can also perform better when specialized for a single image. This proves the effectiveness of the multi-scale structure in the edge segmentation task, but the higher MSE indicates that the edge image generated by the model has a large error; when CIA is activated, it decreased by 0.041 and 0.038 on ODS and OIS respectively, and the MSE also increased by 0.108. This shows that the model is affected by the unlearned rough edge information and its performance has decreased comprehensively. This may be because the noise and misjudged pixel points in the rough edge information let the model learn too much useless information; when CBAM is activated, the model improved by 0.029 and 0.033 on ODS and OIS respectively, and the MSE decreased by 0.316. Compared with the BASE model, the model has an overall improvement. Whether on the overall dataset or for a single image, the model shows strong generalization and adaptation abilities. When both CIA and CBAM are activated, ODS and OIS increased by 0.078 and 0.061 respectively, and the MSE decreased by 0.446. The performance of the model has been greatly improved, and there is an obvious improvement both in the performance on the overall dataset and the adaptability to a single image. And when the composite loss function L flossAfter that, all the indicators of the model reached the optimal level, and compared with the Params (model parameters) of the BASE model, there was only a 6.78% increase, which indicates that the method proposed in the present invention achieved a high performance improvement with a small expenditure, proving the effectiveness of the method described in the present invention.
[0178] In order to better demonstrate the effect of the method described in the present invention on the edge segmentation of rotor blades, the following comparative experiments were carried out.
[0179] Table 4 Comparison of the performance of different algorithms
[0180] Algorithm ODS OIS MSE Params Time(s) DiffuseEdge 0.782 0.791 0.143 224.9M 34.276 Canny 0.754 0.765 0.283 N / A 0.006 Denxied 0.794 0.805 0.156 3.5M 1.282 Sobel 0.786 0.793 0.324 N / A 0.004 The present invention 0.825 0.831 0.101 63K 1.804
[0181] As Figure 4 shown, the traditional segmentation operator performs well in the edge segmentation of ordinary images (tigers, lions). Canny and Sobel can basically segment out the information of the key parts, and the detection speed is extremely fast. However, for the low-pressure turbine rotor blade images with large noise, it is very difficult to effectively segment the edges in the images. The edge images obtained by the traditional segmentation operator are either those that have lost a large amount of original information or those that contain a large amount of noise, and it is difficult to meet the accuracy requirements; Diffusion can segment out the most accurate and detailed edges and performs best in the case of low noise. However, in the face of complex regions in the image, it lacks the mining of potential information, and the inference speed is slow; The image inferred by Denxied is complete and the speed is very fast, but there are also problems such as the loss of edge details and the inability to better mine potential edges. The segmentation image obtained after the method proposed in the present invention segments the edge of the low-pressure turbine rotor blade is superior to the traditional algorithm both in the amount of information segmented and in the edge accuracy, and is also in the fast row in terms of inference speed, thus proving the superiority of the present invention.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for measuring the size of a rotor blade based on computer vision, characterized in that: It includes the following steps: S1. Use X-rays to transmit the engine, obtain the rotor blade image, and perform preprocessing; S2. Adopt a multi-scale CNN network with an adaptive Canny edge detection module combined with a convolutional block attention module to perform edge detection on the preprocessed rotor blade image to obtain a binary image; S3. Statistically analyze the gray values of the rotor blades and draw a gray curve; S4. Take the mutation points in the gray curve as edge points, and use the conversion formula to calculate the actual distance between adjacent edge points, thereby obtaining the rotor blade size.
2. The method for measuring the size of a rotor blade based on computer vision according to claim 1, wherein: The preprocessing described in step S1 includes using median filtering to remove white noise points in the rotor blade image and adopting an unsharp masking algorithm to enhance the detail features of the rotor blade image.
3. A method for measuring the size of a rotor blade based on computer vision according to claim 1, characterized in that: Step S2 specifically includes the following steps: S21. Construct the CACnet network: The CACnet network takes the multi-scale CNN network as the backbone, and the multi-scale CNN network includes an encoder, a neck, and a decoder. Among them, the neck is added with an adaptive Canny edge detection module combined with a convolutional block attention module; S22. Train the CACnet network and use the composite loss function to supervise the training process of the model until convergence; S23. Input the preprocessed rotor blade image into the encoder of the CACnet network, and extract multi-scale features through the combination of multiple convolutional layers and pooling layers; At the same time, use the adaptive Canny edge detection module to extract the edge features of the preprocessed rotor blade image; S24. Input the multi-scale features extracted by the encoder and the edge features extracted by the adaptive Canny edge detection module into the convolutional block attention module for fusion to obtain a fused feature map; S25. Input the fused feature map into the decoder, and use the decoder to perform upsampling and feature fusion on the fused feature map, gradually restoring the spatial resolution of the rotor blade image to generate the final edge image; S26. According to the detection result of the adaptive Canny edge detection module, convert the finally generated edge image into a binary image.
4. A method for measuring the size of a rotor blade based on computer vision according to claim 3, wherein: The composite loss function L described in step S22 floss The expression is as follows: Wherein, L1 represents a loss function for measuring the gap between non-edge pixel points and edge pixel points; L2 represents a loss function for measuring the edge connection loss; represents the i-th prediction result; y represents the true label; represents the prediction result after fusion in different stages; Among them, L1 = w side ·L side +w fuse L fuse (P,Y)(2); where, w side and w fuse represent the weights of the individual outputs of each stage and the weight of the final output respectively; L side and L fuse represent the loss functions of the individual outputs of each stage and the loss function of the final output respectively; P and Y represent the prediction result and the true label of the model respectively; λ1 and λ2 both represent weight factors; L bdry and L tex both represent the boundary tracking fusion function; E represents the set of edge pixels among all pixels; and both represent image patches; L p represents the edge point geometry in the image patch ; L ce , and are all weighted cross - entropy loss functions L(·); And the expression of the weighted cross-entropy loss function L(·) is as follows: Y + = {i | y i ∈ Y, y i > δ}(7); Y - = {i | y i ∈ Y, y i = 0}(8); Wherein, Y represents all pixel points, i.e., the set Y + ∪Y - ; y i represents the i-th pixel point; Y + and Y - represent edge pixel points and non-edge pixel points respectively; α represents the proportion of negative samples in the set Y + ∪Y - ; β represents the hyperparameter pixel point for balancing the importance of edge and non-edge pixel points; δ represents the set threshold value.
5. A method for measuring the size of a rotor blade based on computer vision according to claim 3, characterized in that: The step of using the adaptive Canny edge detection module described in step S23 to extract the edge features of the preprocessed rotor blade image specifically includes the following steps: Step 1: Use the Gaussian filter G(x, y) to denoise the input rotor blade image I(x, y) to obtain a smoothed image I Smooth (x, y): G(x,y) = exp[-(x 2 + y 2 ) / 2σ1] / 2πσ1 2 (9); I Smooth (x,y) = I(x,y) * G(x,y) (10); In the formula, σ1 represents the standard deviation; (x, y) represents the position coordinates of the pixel point; Step 2: Calculate the smoothed image I using the first-order partial derivative operator Smooth The gradient value G of each pixel point on (x, y) in the horizontal direction x and the gradient value G in the vertical direction y ; Step 3: Based on the gradient value G in the horizontal direction x and the gradient value G in the vertical direction y , calculate the gradient magnitude G and the gradient direction θ, and use them to replace the original gray value to obtain the image E(G, θ): Step four. Use the non-maximum suppression method to perform edge localization on the image E(G, θ); Step 5: Divide the image E(G,θ) into multiple sub-intervals, count the probability distribution P(i) of the gray values within each sub-interval, and by setting a high threshold T high Divide the pixel points into two categories C1 and C2: Where G (i,y) represents the gradient magnitude of the pixel point with coordinates (x, y); Step six. Calculate the probabilities and means of the two types of pixel points C1 and C2: In the formula, ω1 and ω2 respectively represent the probabilities of the two types of pixel points C1 and C2; L represents the total number of gray levels of the image; μ1 and μ2 respectively represent the means of the two types of pixel points C1 and C2; Step seven. Calculate the variance υ between the two types of pixel points based on the probabilities and means of the two types of pixel points C1 and C2; υ = ω1(μ1 - μ2) 2 + ω2(μ2 - μ1) 2 (18); Step VIII. Traverse all the set high thresholds T high to maximize the variance υ, and take the set high threshold T high at this time as the optimal high threshold T best-high ; Step Nine: Set the low threshold T low : Step Ten: Compare the pixels in the image E(G, θ) with the optimal high threshold T best-high and the low threshold T low respectively to classify the pixels in the image E(G, θ) and obtain edge features.
6. A method for measuring the size of a rotor blade based on computer vision according to claim 3, characterized in that: In step S24, the convolutional block attention module includes a channel attention module and a spatial attention module, and the expression of the channel attention module is as follows: where, M c (F) represents the output of the channel attention module; both W1 and W represent weight matrices in the multi-layer perceptron; σ represents the Sigmoid activation function; and respectively represent the spatial features generated through the maximum pooling channel and the average pooling channel; The expression of the spatial attention module is as follows: Where M s (F) represents the output of the spatial attention module; f 7×7 represents a 7*7 convolutional kernel; and represent the features generated by max pooling and average pooling respectively; It specifically includes the following steps: S241. Sequentially infer the 1D channel attention map M c ∈R C×1×1 and the 2D spatial attention map M s ∈R 1×H×W : In the formula, F' represents the feature map processed and weighted by the channel attention module; Among them, in the channel attention module, the input feature F ∈ R C×H×W passes through the max-pooling channel and the average-pooling channel respectively to generate spatial features and Then, and are fed into a shared multi-layer perceptron to generate a 1D channel attention map M c ∈ R C×1×1 . Multiply M c with the input feature F to obtain the feature F c after channel dimension fusion; In the spatial attention module, the input feature F ∈ R C×H×W propagates along the axis and generates features through max pooling and average pooling in sequence and then passes through a convolutional layer to generate a 2D spatial attention map M s ∈ R 1×H×W ; S242. Multiply the feature F after channel dimension fusion c element-wise with the 2D spatial attention map M s to obtain the final fused feature map.
7. A method for measuring the size of a rotor blade based on computer vision according to claim 3, characterized in that: The decoder described in step S25 includes a DBLOCK decoding module, a UBlock module, and a ConCat module. Among them, the DBLOCK decoding module mainly consists of a Conv2d convolutional layer and a Gaussian error linear unit activation function. The Conv2d convolutional layer is used to extract the edge features of the rotor blade image, and the Gaussian error linear unit activation function is used for normalization processing; The output feature expression of the DBLOCK decoding module is as follows: where f n+1 represents the output feature of the DBLOCK decoding module in the (n + 1)-th iteration; and represent the height and width of the fused feature map respectively; C represents the number of channels of the fused feature map; The UBlock module uses the Conv2d convolutional layer to learn the fused feature map, and then uses the ConvTranspose2d transposed convolution to upsample the fused feature map to restore the resolution, so as to gradually generate the edge image; The ConCat module is used to rearrange the edge image pixels generated by the UBlock module at different stages to a new image by using PixelShuffle to obtain the final edge image.
8. A method for measuring the size of a rotor blade based on computer vision according to claim 1, characterized in that: The conversion formula described in step S4 is as follows: L = δ * N(25); In the formula, L represents the distance between the two nearest mutation points; δ represents the conversion ratio coefficient; N represents the number of pixels between the two mutation points.
Citation Information
Patent Citations
Workpiece dimension measuring method and device based on machine vision
CN105865344A
Optical fiber diameter measuring method
CN110954002A
Plate glass size measurement method, system and device and readable storage medium
CN115760808A
Road scene three-dimensional reconstruction and defect detection method based on binocular vision
CN117036641A
Image denoising method based on multi-scale edge feature fusion
CN117315266A