A construction method for a super-resolution network model of rotating body images
By constructing an image super-resolution network model, using shallow and deep feature extraction modules to process low-resolution images, the problem of unclear image data characteristics in rotating body vibration measurement is solved, and more accurate vibration displacement signal extraction is achieved.
Patent Information
- Application Number
- CN202211361055.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-11-02
AI Technical Summary
In the vibration measurement of rotating body based on vision, problems such as low resolution, distortion and noise lead to unclear characteristics of the collected image data, which in turn affects the subsequent visual measurement results of vibration displacement.
The image super-resolution network model is constructed using shallow feature extraction module, deep feature extraction module and feature reconstruction module. Through multiple feature processing and feature fusion blocks in the deep feature extraction module, the resolution and feature clarity of the image are improved.
It effectively alleviates the deviation of visual vibration measurement task caused by low image resolution, improves the clarity of the acquired image data and feature extraction accuracy, and thus improves the accuracy of the vibration displacement signal.
Smart Images

Figure CN115619643B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a super-resolution network model for rotational body images, and belongs to the field of super-resolution reconstruction of rotational body images. Background Art
[0002] In industrial production, rotational machinery structures account for more than 40% of all mechanical structures. Therefore, it is essential to perform safety inspections on rotational bodies. The vision-based structural vibration displacement monitoring method has received increasing attention due to its many advantages such as long distance, non-contact, and multi-point measurement. However, in actual engineering applications, vision-based rotational body vibration measurement often results in low-resolution, distorted, and noisy rotational body vibration image data due to the low performance of the acquisition hardware or the constraints of the acquisition site environment. These phenomena will cause serious deviations between the rotational body displacement data regressed in the subsequent vibration displacement vision measurement step and the real data. How to effectively enhance the new boundary information of the low-resolution rotational body vibration displacement image data features and reduce the problem of low image data resolution caused by external factors is of great significance for vision-based engineering measurement projects. Summary of the Invention
[0003] The present invention provides a method for constructing a super-resolution network model for rotational body images. An image super-resolution network model is constructed through a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module, and is further used to achieve image super-resolution reconstruction.
[0004] The technical solution of the present invention is: a method for constructing a super-resolution network model for rotational body images, which uses a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module to construct an image super-resolution network model.
[0005] The shallow feature extraction module includes: using a convolution operation as the shallow feature extraction module, and taking the original low-resolution image input to the network as the input of the shallow feature extraction module to obtain shallow features.
[0006] The deep feature extraction module is composed of m Transformer blocks connected in series. The input of the first Transformer block is the shallow feature output by the shallow feature extraction module, and the input of the subsequent m - 1 Transformer blocks is the output of the previous Transformer block.
[0007] The Transformer block includes: performing Swin-Transformer operation on the input features to obtain feature S1, then performing downsampling operation on feature S1 to obtain feature D1, and at the same time inputting feature S1 into the attention block for operation to obtain feature C1; performing Swin-Transformer operation on the downsampled feature D1 to obtain feature S2, performing downsampling operation on feature S2 to obtain feature D2, and at the same time inputting feature S2 into the attention block for operation to obtain feature C2; performing Transformer operation on the downsampled feature D2 to obtain feature T1, then performing Transformer operation on feature T1 to obtain feature T2, inputting feature T2 and feature D2 into the feature fusion block to obtain feature F1, and then performing upsampling operation on feature F1 to obtain feature U1; adding the upsampled feature U1 and the output feature C2 of the attention block to obtain feature A2, then performing Swin-Transformer operation on feature A2 to obtain feature S3, inputting feature S3 and feature D1 into the feature fusion block to obtain feature F2, and then performing upsampling operation on feature F2 to obtain feature U2; adding the upsampled feature U2 and the output feature C1 of the attention block to obtain feature A1, then performing Swin-Transformer operation on feature A1 to obtain feature S4, and feature S4 is the output feature of a single Transformer block.
[0008] The attention block includes: the attention block sequentially performs convolution operation, PReLU activation function operation, and convolution operation on the input features to obtain feature P1, sequentially performs global average pooling operation, convolution operation, PReLU activation function operation, convolution operation, and Sigmoid activation function operation on feature P1 to obtain feature P2, multiplies feature P2 with the original input features of the attention block to obtain feature P3, and then adds feature P3 and feature P1 to obtain the output feature;
[0009] The feature fusion block includes: combining the output features D1 and D2 after downsampling operations in a single Transformer block, the output feature S3 after Swin-Transformer operations, and the output feature T2 after Transformer operations, and combining them pairwise into two pairs of inputs in the manner of (S3, D1) and (T2, D2). The feature in the front of each pair of inputs is used as the input feature 1 of the feature fusion block, and the feature in the back is used as the input feature 2 of the feature fusion block; performing two convolution operations with the same kernel size on the input feature 1 respectively to obtain the feature Co1 and the feature Co2, adding the feature Co2 to the input feature 2 to obtain the feature Ad1, performing a convolution operation and a Sigmoid activation function operation on the feature Ad1 in sequence to obtain the feature Si1, multiplying the feature Si1 by the feature Co1 to obtain the feature Mu1, and then adding the feature Mu1 to the input feature 1 to obtain the output feature of the feature fusion block.
[0010] The feature reconstruction module includes: using the feature output by the deep feature extraction module as the input feature of the feature reconstruction module, and performing a convolution operation, a LeakyReLU activation function operation, a convolution operation, a PixelShuffle operation, and a convolution operation on the input feature of the feature reconstruction module in sequence, and the output is the output of the feature reconstruction module.
[0011] The beneficial effects of the present invention are as follows: In view of the pain points of the visual vibration measurement task, the present invention proposes the concept of super-resolution reconstruction of the collected image data, which effectively alleviates the problem of excessive deviation in the visual vibration measurement task caused by too low image resolution. Further, the present invention proposes an image super-resolution network applicable to rotating bodies. By performing super-resolution reconstruction on the collected low-resolution image data of the rotating body, the problem of too low resolution of the collected image data caused by interference in the acquisition device or acquisition link can be well alleviated. Specifically: The image super-resolution process of the present invention is divided into three modules, namely, a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module. By performing multiple feature processes on the input low-resolution image, the important features in the image can be processed more effectively. Further, the deep feature extraction module proposed by the present invention based on the encoder-decoder architecture can extract feature information at multiple scales for a frame of image, and a feature fusion block is used in the deep feature extraction module to aggregate the feature information at different scales to achieve a better feature reconstruction effect. Moreover, by selecting Transformer to replace the traditional convolution operation, the image feature information can be learned more fully, and by using two Transformers, the feature detail information in the image can be extracted more fully. Furthermore, through the feature reconstruction module, the image features can be better refined and reconstructed. Through this model, sufficient image feature information can be obtained and further used to extract displacement signals, and the vibration displacement signals are relatively smooth. Based on the above, applying the model proposed by the present invention to the data collection stage of visual vibration measurement effectively reduces the performance requirements such as the resolution of the camera required for data acquisition. By comparing the vibration signal diagrams generated from the data processed by the present invention and the vibration signal diagrams generated from the unprocessed data, the practical engineering application value of the present invention is proved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram of the model and application process of the present invention;
[0013] Figure 2 is a structural diagram of the deep feature extraction module;
[0014] Figure 3 is a structural diagram of a single Transformer block;
[0015] Figure 4 is a structural diagram of the attention block in a single Transformer block;
[0016] Figure 5 is a structural diagram of the feature fusion block in a single Transformer block;
[0017] Figure 6 is a structural diagram of the feature reconstruction module;
[0018] Figure 7It is the application flow chart of the model of the present invention;
[0019] Figure 8 It is the comparison chart of high-resolution images of the rotor after reconstruction by different algorithms;
[0020] Figure 9 It is the displacement data of the rotor on the X-axis regressed from the image data after super-resolution reconstruction and the blurred image data;
[0021] Figure 10 It is the displacement data of the rotor on the Y-axis regressed from the image data after super-resolution reconstruction and the blurred image data. Specific implementation manners
[0022] Next, in combination with the accompanying drawings and embodiments, the invention will be further described, but the content of the present invention is not limited to the described scope.
[0023] As Figure 1 shown, a method for constructing a super-resolution network model for rotating body images uses a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module to construct an image super-resolution network model.
[0024] Further, the shallow feature extraction module includes: using a convolution operation as the shallow feature extraction module, taking the original low-resolution image input into the network as the input of the shallow feature extraction module to obtain shallow features. In this embodiment, the convolution kernel size of the convolution operation is 3×3.
[0025] Further, the deep feature extraction module includes: inputting the shallow features output by the shallow feature extraction module into the deep feature extraction module; wherein, the deep feature extraction module is composed of m Transformer blocks connected in series, the input of the first Transformer block is the shallow features output by the shallow feature extraction module, and the input of the subsequent m - 1 Transformer blocks are all the outputs of the previous Transformer block.
[0026] As Figure 2 shown, in the embodiment of the present invention, the deep feature extraction module is composed of 6 Transformer blocks connected in series, the input of the first Transformer block is the shallow features output by the shallow feature extraction module, and the input of the subsequent 5 Transformer blocks are all the outputs of the previous Transformer block.
[0027] Further, as Figure 3As shown in the figure, the single Transformer block includes: performing Swin-Transformer operation on the input features to obtain feature S1, then performing 2-fold downsampling operation on feature S1 to obtain feature D1, and at the same time inputting feature S1 into the attention block for operation to obtain feature C1; performing Swin-Transformer operation on the downsampled feature D1 to obtain feature S2, performing 2-fold downsampling operation on feature S2 to obtain feature D2, and at the same time inputting feature S2 into the attention block for operation to obtain feature C2; performing Transformer operation on the downsampled feature D2 to obtain feature T1, then performing Transformer operation on feature T1 to obtain feature T2, inputting feature T2 and feature D2 into the feature fusion block to obtain feature F1, and then performing 2-fold upsampling operation on feature F1 to obtain feature U1; adding the upsampled feature U1 and the output feature C2 of the attention block to obtain feature A2, then performing Swin-Transformer operation on feature A2 to obtain feature S3, inputting feature S3 and feature D1 into the feature fusion block to obtain feature F2, and then performing 2-fold upsampling operation on feature F2 to obtain feature U2; adding the upsampled feature U2 and the output feature C1 of the attention block to obtain feature A1, then performing Swin-Transformer operation on feature A1 to obtain feature S4, and feature S4 is the output feature of the single Transformer block. In addition, during the experiment, setting the downsampling multiple to 2 times can avoid the deficiency of being difficult to restore the real information during reconstruction.
[0028] Further, as Figure 4 shown in the figure, the attention block includes: the attention block takes the feature after Swin-Transformer processing in the single Transformer block as the input; the attention block sequentially performs convolution operation, PReLU activation function operation, and convolution operation on the input feature to obtain feature P1, performs global average pooling operation, convolution operation, PReLU activation function operation, convolution operation, and Sigmoid activation function operation on feature P1 in sequence to obtain feature P2, multiplies feature P2 and the original input feature of the attention block to obtain feature P3, and then adds feature P3 and feature P1 to be the output feature;
[0029] Further, as Figure 5As shown in the figure, the feature fusion block includes: combining the output features D1 and D2 after the downsampling operation in a single Transformer block, the output feature S3 after the Swin-Transformer operation, and the output feature T2 after the Transformer operation, and combining them pairwise into two pairs of inputs in the manner of (S3, D1) and (T2, D2). The feature in the front of each pair of inputs is used as the input feature 1 of the feature fusion block, and the feature in the back is used as the input feature 2 of the feature fusion block. That is, for the first pair of inputs, S3 is used as the input feature 1 and D1 is used as the input feature 2. For the second pair of inputs, T2 is used as the input feature 1 and D2 is used as the input feature 2. The input feature 1 is respectively subjected to two convolution operations with a convolution kernel size of 3×3 to obtain the feature Co1 and the feature Co2. The feature Co2 is added to the input feature 2 to obtain the feature Ad1. The feature Ad1 is sequentially subjected to a convolution operation with a convolution kernel size of 3×3 and a Sigmoid activation function operation to obtain the feature Si1. The feature Si1 is multiplied by the feature Co1 to obtain the feature Mu1. Then, the feature Mu1 is added to the input feature 1 to obtain the output feature of the feature fusion block.
[0030] Further, as Figure 6 shown in the figure, the feature reconstruction module includes: using the feature output by the deep feature extraction module as the input feature of the feature reconstruction module, and sequentially performing a convolution operation, a LeakyReLU activation function operation, a convolution operation, a PixelShuffle operation, and a convolution operation on the input feature of the feature reconstruction module, and the output is the output of the feature reconstruction module.
[0031] Further, as Figure 7 shown in the figure, the image super-resolution network model constructed by the method of the present invention is used for visual measurement of the vibration displacement of a rotating body. The specific steps are as follows: collecting a high-speed rotating rotor image dataset and dividing it into a training dataset and a validation dataset; constructing a prototype of the image super-resolution network model; before formal training, modifying the corresponding parameters in the configuration file to obtain training parameters; calling the training dataset and the configuration file to start training the selected image super-resolution network model. After training, screening out the optimal candidate weights; using the validation dataset to evaluate the performance of the optimal candidate weights to quantify the performance of the weights, and loading the optimal weights obtained according to the quantification result into the image super-resolution network model; using the image super-resolution network model loaded with the optimal weights to process the newly collected high-speed rotating rotor image dataset and extracting the vibration displacement signal of the high-speed vibrating rotor.
[0032] Specifically:
[0033] S1. On the ZT-3 rotor vibration simulation test bench, use a 5F01 high-speed camera to collect images of the high-speed rotating rotor at a collection speed of 2000 frames per second, construct an image dataset of the high-speed rotating rotor, with an image resolution of 512×512. Sequentially perform 4-fold downsampling on the collected high-speed rotating rotor images to obtain image data of 128×128, and obtain the original low-resolution image dataset. Then divide the original low-resolution image dataset obtained through downsampling into a training dataset and a validation dataset in proportion. In this embodiment, we collected 1000 frames of high-speed rotating rotor image data and the corresponding 1000 frames of low-resolution image dataset after downsampling, and performed dataset division. Among them, the training dataset is 700 frames of high-speed rotating rotor image data and the corresponding 700 frames of low-resolution image dataset, and the validation dataset is 300 frames of high-speed rotating rotor image data and the corresponding 300 frames of low-resolution image dataset (that is, the division ratio can also be 70% and 30%). In addition, when collecting images of the high-speed rotating rotor, use an EF-200 LED lamp for light compensation of the collection object; the division ratio can also be 90% and 10%, etc. It should be noted that other methods can also be used to obtain the image dataset of the high-speed rotating rotor, and then convert it into image data of 128×128.
[0034] S2. Use a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module to construct the prototype of an image super-resolution network model;
[0035] S3. Before formal training, modify the corresponding parameters in the configuration file to obtain training parameters;
[0036] S4. Call the training dataset and the configuration file to start training the selected image super-resolution network model. After training is completed, screen out the optimal candidate weights. In this embodiment, first set the training parameters in the configuration file as BatchSize = 8, the initial learning rate is set to 10 -4 ., the number of iterations is set to 500,000 times, until the learning rate drops to 10 -7 . Start training, load the network model of image super-resolution, and load images in batches according to the size of BatchSize for training. According to the set parameters, output the weights once when the number of iterations reaches the set number, and then screen out the optimal weights.
[0037] S5. Use the validation dataset to evaluate the performance of the optimal candidate weights to quantify the performance of the weights, and load the optimal weights obtained according to the quantification result into the image super-resolution network model;
[0038] S6. Use the image super-resolution network model loaded with the optimal weights to process the low-resolution high-speed rotating rotor image dataset to be reconstructed, and introduce the YOLOv5x network to identify the bounding boxes of the rotors in the reconstructed rotor image dataset. By summarizing the displacement changes of the center points of the rotor bounding boxes in each frame of an image sequence, the vibration displacement data of the high-speed vibrating rotor is obtained.
[0039] In this embodiment, the present invention takes 1000 frames of images with a resolution of 128×128 obtained through downsampling in S1 as the low-resolution high-speed rotating rotor image dataset to be reconstructed; divides the 1000 frames of low-resolution high-speed rotating rotor images to be reconstructed into a training set 1 for detection with a total of 700 frames of images and a validation set 1 with a total of 300 frames of images; inputs the training set 1 and the validation set 1 into the image super-resolution network model proposed by the present invention to obtain a training set 2 and a validation set 2, as Figure 8 As shown in the comparison chart of the high-resolution rotor images reconstructed by different algorithms, different algorithms are for the comparison results after reconstructing the same frame of picture. From the comparison results, it can be seen that the PNSR of the reconstructed images of the present invention is better than other traditional algorithms (it should be noted that several traditional algorithms given in the present invention not only cover traditional methods, but also include algorithms based on CNN and Swin-Transformer, further demonstrating the advantages of the present invention in high-speed rotating rotor image processing); use the Labelimg software to label the rotors in the images of the training set 2 and the validation set 2 with bounding boxes and generate txt files; input the training set 2, the validation set 2, and the txt files generated by the annotation into YOLOv5x for training to obtain a weight parameter model and get the optimal weight file in pt format; use the obtained weight file to detect the newly acquired super-resolution reconstructed image data to obtain the txt file of the center of gravity coordinates of the rotor in each frame of the image, and extract the center point coordinates in multiple txt files; plot the center point coordinates of multiple frames into a curve graph, which is the vibration displacement signal of the high-speed vibrating rotor.
[0040] When collecting image data of high-speed vibrating objects, problems such as hardware limitations of the collection device or interference in the collection site may occur. These problems can lead to blurred or low-resolution target displacement image data. Moreover, since the feature boundaries of image data without super-resolution reconstruction are relatively blurred, it is difficult to ensure that the boundaries of the bounding boxes marked during the annotation stage accurately fall on the target boundaries when extracting signals by detection methods, resulting in poor accuracy of the generated displacement vibration signals and a large deviation from the actual situation. The network model proposed in the present invention can well alleviate the problem of image clarity when collecting images by performing super-resolution reconstruction on the collected image data, thereby making the target boundaries clear and accurately marking the bounding boxes that fit the target boundaries during annotation. Further, the model trained with accurately marked labels can well calculate the coordinate information of the bounding boxes of the test set data, so that the generated signals are more in line with the actual vibration displacement signals.
[0041] By Figure 9 and Figure 10 Comparing, it can be seen that the vibration displacement signals collected from the image data after super-resolution reconstruction are smoother than those collected from the image data without super-resolution reconstruction.
[0042] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A construction method for a super-resolution network model of rotational body images, characterized in that: using a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module to construct an image super-resolution network model; the deep feature extraction module is composed of m Transformer blocks connected in series. The input of the first Transformer block is the shallow feature output by the shallow feature extraction module, and the inputs of the subsequent m - 1 Transformer blocks are all the outputs of the previous Transformer block; the Transformer block includes: performing Swin-Transformer operation on the input feature to obtain feature S1, then performing downsampling operation on feature S1 to obtain feature D1, and at the same time inputting feature S1 into the attention block for operation to obtain feature C1; performing Swin-Transformer operation on the downsampled feature D1 to obtain feature S2, performing downsampling operation on feature S2 to obtain feature D2, and at the same time inputting feature S2 into the attention block for operation to obtain feature C2; performing Transformer operation on the downsampled feature D2 to obtain feature T1, then performing Transformer operation on feature T1 to obtain feature T2, inputting feature T2 and feature D2 into the feature fusion block to obtain feature F1, and then performing upsampling operation on feature F1 to obtain feature U1; adding the upsampled feature U1 and the output feature C2 of the attention block to obtain feature A2, then performing Swin-Transformer operation on feature A2 to obtain feature S3, inputting feature S3 and feature D1 into the feature fusion block to obtain feature F2, and then performing upsampling operation on feature F2 to obtain feature U2; adding the upsampled feature U2 and the output feature C1 of the attention block to obtain feature A1, then performing Swin-Transformer operation on feature A1 to obtain feature S4, and feature S4 is the output feature of a single Transformer block.
2. The construction method for a super-resolution network model of rotational body images according to claim 1, characterized in that: the shallow feature extraction module includes using a single convolution operation as the shallow feature extraction module, taking the original low-resolution image input to the network as the input of the shallow feature extraction module, and obtaining the shallow feature.
3. The construction method for a super-resolution network model of rotational body images according to claim 1, characterized in that: the attention block includes: the attention block sequentially performs convolution operation, PReLU activation function operation, and convolution operation on the input feature to obtain feature P1, sequentially performs global average pooling operation, convolution operation, PReLU activation function operation, convolution operation, and Sigmoid activation function operation on feature P1 to obtain feature P2, multiplies feature P2 by the original input feature of the attention block to obtain feature P3, and then adds feature P3 and feature P1 to obtain the output feature.
4. The construction method for the super-resolution network model of rotational body images according to claim 1, characterized in that: the feature fusion block includes: combining the output features D1 and D2 after downsampling operations in a single Transformer block, the output feature S3 after Swin-Transformer operations, and the output feature T2 after Transformer operations, and combining them pairwise into two pairs of inputs in the manner of (S3, D1) and (T2, D2). The one in the front of each pair of inputs is used as the input feature 1 of the feature fusion block, and the one in the back is used as the input feature 2 of the feature fusion block; performing two convolution operations with the same convolution kernel size on the input feature 1 respectively to obtain the feature Co1 and the feature Co2, adding the feature Co2 to the input feature 2 to obtain the feature Ad1, performing a convolution operation and a Sigmoid activation function operation on the feature Ad1 in sequence to obtain the feature Si1, multiplying the feature Si1 by the feature Co1 to obtain the feature Mu1, and then adding the feature Mu1 to the input feature 1 to obtain the output feature of the feature fusion block.
5. The construction method for the super-resolution network model of rotational body images according to claim 1, characterized in that: the feature reconstruction module includes: using the feature output by the deep feature extraction module as the input feature of the feature reconstruction module, and performing convolution operation, LeakyReLU activation function operation, convolution operation, PixelShuffle operation, and convolution operation on the input feature of the feature reconstruction module in sequence, and the output is the output of the feature reconstruction module.
Citation Information
Patent Citations
Image super-resolution reconstruction model and method based on residual mixed attention network
CN115222601A