A high-frequency rotor vision vibration measurement method and system based on super-resolution reconstruction
By constructing a deep learning network model, using feature extraction, interpolation, pyramid alignment, ConvLSTM and series residual block modules, the blur and frame drop problems of high-frequency vibration video of low-resolution rotors are solved, and high-quality video super-resolution reconstruction and extraction of high-frequency vibration displacement signals are realized.
Patent Information
- Application Number
- CN202111589398.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-23
AI Technical Summary
In visual vibration measurement engineering, the high-frequency vibration video of low-resolution rotors collected by high-speed industrial cameras has problems such as blur, distortion, noise and frame loss, resulting in serious displacement errors during high-frequency visual vibration measurement. At the same time, traditional super-resolution reconstruction algorithms are difficult to effectively process time domain information between consecutive frames.
A deep learning-based method is adopted to build a deep learning network model, using feature extraction module, interpolation module, pyramid alignment module, ConvLSTM module and series residual block to realize video super-resolution reconstruction, and the high-frequency rotor vibration displacement signal is extracted through vibration measurement technology.
It effectively alleviates the frame drop problem when collecting images, enhances the time domain information of continuous frames, and the extracted vibration displacement signal is relatively smooth and has less noise, reducing the camera frame rate and resolution required for data acquisition.
Smart Images

Figure CN115293963B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a high-frequency rotor visual vibration measurement method based on super-resolution reconstruction, belonging to the fields of artificial intelligence video super-resolution reconstruction and computer vision. Background Art
[0002] In actual visual vibration measurement engineering applications, due to the constraint of the sampling theorem, the acquisition hardware will cause a large number of phenomena such as blur, distortion, noise, and frame loss in the low-resolution rotor high-frequency vibration video collected by high-speed industrial cameras. This information loss phenomenon will cause serious displacement errors during high-frequency visual vibration measurement. How to quickly and efficiently enhance the high-frequency detail information of the low-resolution rotor displacement image, reduce the vibration displacement error caused by the constraint of the sampling theorem, and solve the problem of the lack of inter-frame displacement correlation in the single-frame vibration image super-resolution method is of great significance to the development of engineering projects.
[0003] At the same time, in traditional super-resolution reconstruction, most algorithms are based on single-frame image reconstruction. However, single-frame image super-resolution reconstruction is not so excellent for processing continuous-frame images. The biggest problem is that single-frame image super-resolution reconstruction cannot well connect the temporal information between continuous frames. How to process the temporal information between continuous frames is the primary problem we face in obtaining high-quality reconstructed frames. Summary of the Invention
[0004] The present invention provides a high-frequency rotor visual vibration measurement method and system based on super-resolution reconstruction. By using a feature extraction module, an interpolation module, a pyramid alignment module, a ConvLSTM module, and a series of residual blocks, a prototype of a deep learning network model is constructed, and further used to implement video super-resolution reconstruction, and then the vibration of the reconstructed image is measured.
[0005] The technical solution of the present invention is: a high-frequency rotor visual vibration measurement method based on super-resolution reconstruction, including:
[0006] A division step of collecting a high-speed rotating rotor image dataset and dividing it into a training dataset and a validation dataset;
[0007] A construction step of constructing a prototype of a deep learning network model;
[0008] A determination step of performing ablation experiments on the prototype of the deep learning network model to determine two different deep learning network models;
[0009] An obtaining step of modifying the corresponding hyperparameters in the configuration file before formal training to obtain training parameters;
[0010] A screening step of calling the training dataset and the configuration file to start training the selected deep learning network model, and after the training is completed, screening out the optimal candidate weights;
[0011] Loading step: Use the validation dataset to evaluate the performance of the optimal candidate weights to quantify the performance of the weights, and load the optimal weights obtained according to the quantization results into the deep learning network model;
[0012] Extraction step: Use the deep learning network model loaded with the optimal weights to process the newly collected high-speed rotating rotor image dataset, and extract the vibration displacement signal of the high-speed vibrating rotor.
[0013] The specific division step is as follows: The image resolution in the collected high-speed rotating rotor image dataset is A×A. All images in the collected high-speed rotating rotor image dataset are successively subjected to 4-fold downsampling and Gaussian blur processing to obtain image data with a resolution of B×B, and the processed image data is divided into a training dataset and a validation dataset; A > B.
[0014] The specific construction step is as follows: Use a feature extraction module, an interpolation module, a pyramid alignment module, a ConvLSTM module, and a series of residual blocks to construct a prototype of the deep learning network model.
[0015] The interpolation module includes:
[0016] Take the features output by the feature extraction module Figure 1 and the features Figure 2 as the input of the interpolation module. After passing the features Figure 1 and the features Figure 2 through channel-level connection and convolution operations respectively according to different sequential arrangements, obtain convolution offset 1 and convolution offset 2;
[0017] Perform deformable convolution operations on convolution offset 1, features Figure 1 and convolution offset 2, features Figure 2 respectively to obtain sampled features Figure 1 and sampled features Figure 2 ; Give the two sampled feature maps two learnable weights through a mixing function for the sampled features Figure 1 and the sampled features Figure 2 respectively to generate an intermediate frame, that is, the feature Figure 3 .
[0018] The mixing function in the interpolation module is specifically:
[0019] F t =α*T t-1 +β*T t+1
[0020] where α and β both represent learnable convolution kernels of size 1×1, * represents convolution operation, and F t represents the interpolated intermediate frame, that is, the featureFigure 3 , T t-1 represents a feature Figure 1 , T t+1 represents a feature Figure 2 .
[0021] The pyramid alignment module includes:
[0022] Arrange the outputs of the feature extraction module and the frame interpolation module according to the features Figure 1 , feature Figure 3 and feature Figure 2 in order, and take every two frames as a group in turn as inputs, namely input 1 and input 2;
[0023] First, perform two - fold downsampling and four - fold downsampling on input 1 and input 2 respectively to obtain input 1 - 2, input 2 - 2, input 1 - 4, and input 2 - 4;
[0024] Perform channel - level connection on input 1 - 4 and input 2 - 4, then perform a convolution operation, calculate the offset for the result after convolution, that is, offset 1 of input 1 - 4 and input 2 - 4, perform a deformable convolution operation on offset 1 and input 2 - 4 to obtain reconstructed frame 1; perform channel - level connection on reconstructed frame 1 and input 2 - 4, and input the connected result into a dense residual module to obtain dense 1;
[0025] Perform channel - level connection on input 1 - 2 and input 2 - 2, perform two convolution operations on the connected result, perform an offset calculation operation on the convolution result, that is, offset 2 - 1 of input 1 - 2 and input 2 - 2; perform 2 - fold upsampling on offset 1 to obtain offset 1 - 2, perform channel - level connection on offset 2 - 1 and offset 1 - 2 to obtain offset 2; perform a deformable convolution operation on offset 2 and input 2 - 2 to obtain reconstructed frame 2 - 1; perform 2 - fold upsampling on dense 1 to obtain dense 1 - 2, perform channel - level connection on dense 1 - 2 and reconstructed frame 2 - 1, then input the connected result into a convolution operation to obtain reconstructed frame 2, perform channel - level connection on reconstructed frame 2 and input 2 - 2, and input the connected result into a dense residual module to obtain dense 2;
[0026] Perform channel-level connection on Input 1 and Input 2, input the connected result into two convolution operations, then perform an operation to calculate the offset on the convolved result to obtain the offset 3-1 of Input 1 and Input 2. Perform 2x upsampling on Offset 2 to obtain Offset 2-2. Perform channel-level connection on Offset 2-2 and Offset 3-1 to obtain Offset 3. Perform deformable convolution operation on Offset 3 and Input 2 to obtain Reconstructed Frame 3-1. Perform 2x upsampling on Dense 2 to get Dense 2-2. Perform channel-level connection on Dense 2-2 and Reconstructed Frame 3-1, and then input the connected result into a convolution operation to obtain Reconstructed Frame 3. Perform channel-level connection on Reconstructed Frame 3 and Input 2, and then input the connection result into the dense residual module to obtain Dense 3;
[0027] Perform channel-level connection on Input 2 and Dense 3, then perform convolution, deformable convolution, and operation to calculate the offset on the connected result to obtain the offset of Input 2 and Dense 3, that is, Offset 4. Use the result of performing deformable convolution operation on Offset 4 and Dense 3 as the output result of the pyramid alignment module.
[0028] The dense residual module in the pyramid alignment module takes the reconstructed frame after channel-level connection in the pyramid alignment module as the input; it includes: n residual blocks, where n is greater than or equal to 0 and is a positive integer;
[0029] The residual block includes a ReLU activation function and two convolution operations; among them, the relevant parameters of the convolution operations are all the same, where the number of input and output channels is 64, the convolution kernel size is 3×3, and the convolution stride is defaulted to 1;
[0030] If n is 0, the input is the output;
[0031] If n is 1, the connection of the residual block is: input the input into the first convolution, ReLU activation function, and the second convolution in sequence, and the output point is Node 1; at Node 1, add the input and the output through the above operations, which is the output at Node 1, and the result is the output;
[0032] If n takes a value greater than 1: For the structure of the first residual block: The input is successively fed into the first convolution, the ReLU activation function, and the second convolution, and the output point is node 1; at node 1, the input is added to the output obtained through the above operations, which is the output at node 1; for the structures of the second to the (n - 1)-th residual blocks, described by the i-th residual block: The output of node i - 1 is successively fed into the (2i - 1)-th convolution, the ReLU activation function, and the 2i-th convolution, and the output point is node i; at node i, the output result is added to the input of node i - 1, which is the output at node i; for the structure of the n residual blocks: The output of node n - 1 is successively fed into the (2n - 1)-th convolution, the ReLU activation function, and the 2n-th convolution, and the output point is node n; at node n, the output after the 2n-th convolution, the output after the 2(n - 1)-th convolution,..., and the output after the second convolution are added together as the output of node n; the output of node n is successively fed into the (2n + 1)-th convolution and the ReLU activation function, and the output point is node n + 1, and at node n + 1, the output is added to the original input, and the result is the output; where i = 2,... n - 1.
[0033] The ablation experiment includes:
[0034] Under the condition that other conditions are the same, by changing the number of residual blocks in the dense residual module in the pyramid alignment module, multiple different prototypes of deep learning network models are obtained;
[0035] Perform performance evaluation on multiple different prototypes of deep learning network models, and screen the optimal model under the current index according to multiple performance indicators.
[0036] The specific extraction steps are as follows: The specific implementation steps are as follows:
[0037] The newly collected high-speed rotating rotor image data is divided into training set 1 and validation set 1, and input into the deep learning network model to obtain training set 2 and validation set 2;
[0038] Use annotation software to perform bounding box annotation on the rotors in the images of training set 2 and validation set 2, and generate xml files;
[0039] Input training set 2, validation set 2, and the xml files generated by annotation into Fast-RCNN for training to obtain a weight parameter model; two corresponding pkl-1 and pkl-2 files are obtained; the pkl files contain the upper left and lower right diagonal vertex coordinates of the bounding boxes generated during detection;
[0040] Extract the Y-axis data of the diagonal coordinates in the pkl-1 and pkl-2 files respectively to calculate the Y-axis coordinates of the center points; Plot the multi-frame center point coordinates as a curve graph, which is the vibration displacement signal of the high-speed vibrating rotor.
[0041] According to another aspect of the embodiments of the present invention, there is also provided a high-frequency rotor visual vibration measurement system that integrates super-resolution reconstruction and video frame interpolation, including:
[0042] A division unit, configured to collect a high-speed rotating rotor image data set and divide it into a training data set and a validation data set;
[0043] A construction unit, configured to construct a prototype of a deep learning network model;
[0044] A determination unit, configured to perform ablation experiments on the prototype of the deep learning network model to determine two different deep learning network models;
[0045] An obtaining unit, configured to modify corresponding hyperparameters in the configuration file before formal training to obtain training parameters;
[0046] A screening unit, configured to call the training data set and the configuration file to start training the selected deep learning network model, and after the training ends, screen out the optimal candidate weights;
[0047] A loading unit, configured to perform performance evaluation on the optimal candidate weights using the validation data set to quantify the performance of the weights, and load the optimal weights obtained according to the quantification result into the deep learning network model;
[0048] An extraction unit, configured to use the deep learning network model loaded with the optimal weights to process a newly collected high-speed rotating rotor image data set and extract the vibration displacement signal of the high-speed vibrating rotor.
[0049] The beneficial effects of the present invention are as follows: The interpolation module proposed in the present invention supplements a frame of picture between two frames, which effectively alleviates the problem of frame loss during image acquisition; by integrating video interpolation into a framework of video super-resolution reconstruction tasks, the sharing of feature information between the two model tasks is realized; further, the present invention proposes a pyramid alignment module, which uses a pyramid structure for multi-stage information fusion to fully learn the spatial information within the frame, and uses deformable convolution to fully utilize the spatio-temporal domain information for feature-level adaptive alignment. By fusing the aligned latent frame with the reference frame, the temporal information of consecutive frames can be enhanced; furthermore, the ConvLSTM module and the cascaded residual blocks can better extract features; through this model, sufficient frame information can be obtained, which can be further used to extract displacement signals, and the vibration displacement signals are relatively smooth, and the generated signals have relatively less noise; furthermore, two methods with different performances are provided according to different requirements for model speed and accuracy priority, which can better adapt to the performance requirements of different devices. Based on the above, the model proposed in the present invention is used in the data collection stage of visual vibration measurement, effectively reducing the performance requirements such as the frame rate and resolution of the cameras required for data acquisition; by comparing the vibration signal diagrams generated from the data processed by the present invention and the vibration signal diagrams generated from the unprocessed data, the practical engineering application value of the present invention is demonstrated. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the present invention;
[0051] Figure 2 is a structural diagram of the interpolation module;
[0052] Figure 3 is a structural diagram of the pyramid alignment module;
[0053] Figure 4 is a structural diagram of the dense residual module;
[0054] Figure 5 is a structural diagram of the deep learning network model;
[0055] Figure 6 is a specific implementation diagram of the process of training the model;
[0056] Figure 7 is a comparison diagram of the present invention with other algorithms;
[0057] Figure 8 is the rotor visual vibration measurement data without super-resolution reconstruction image reconstruction;
[0058] Figure 9 is the rotor visual vibration measurement data of the super-resolution reconstructed image. DETAILED DESCRIPTION OF THE INVENTION
[0059] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but the content of the present invention is not limited to the described scope.
[0060] Embodiment 1: As Figures 1-9 shown, a high-frequency rotor visual vibration measurement method based on super-resolution reconstruction includes:
[0061] A dividing step of collecting a high-speed rotating rotor image data set and dividing it into a training data set and a validation data set;
[0062] A constructing step of constructing a prototype of a deep learning network model;
[0063] A determining step of performing ablation experiments on the prototype of the deep learning network model to determine two different deep learning network models;
[0064] An obtaining step of modifying corresponding hyperparameters in the configuration file before formal training to obtain training parameters;
[0065] A screening step of calling the training data set and the configuration file to start training the selected deep learning network model, and after the training is completed, screening out the optimal candidate weights;
[0066] A loading step of performing performance evaluation on the optimal candidate weights using the validation data set to quantify the performance of the weights, and loading the optimal weights obtained according to the quantification results into the deep learning network model;
[0067] An extracting step of using the deep learning network model loaded with the optimal weights to process a newly collected high-speed rotating rotor image data set and extracting the vibration displacement signal of the high-speed vibrating rotor.
[0068] Further, it can be set that the dividing step is specifically: the image resolution in the collected high-speed rotating rotor image data set is A×A, and all the images in the collected high-speed rotating rotor image data set are sequentially subjected to 4-fold downsampling and Gaussian blur processing to obtain image data with a resolution of B×B, and the processed image data is divided into a training data set and a validation data set; A > B.
[0069] Further, it can be set that the training data set and the validation data set respectively account for 70% and 30% of the image data set.
[0070] Further, it can be set that the constructing step is specifically: using a feature extraction module, an interpolation module, a pyramid alignment module, a ConvLSTM module, and a series of residual blocks to construct a prototype of a deep learning network model.
[0071] Further, it can be set that the feature extraction module is composed of 1 dimension expansion convolutional layer and 5 residual blocks; the input two images are processed and then output features Figure 1 and featureFigure 2 。
[0072] Furthermore, the interpolation module can be set to include:
[0073] Taking the features output by the feature extraction module Figure 1 and the features Figure 2 as the input of the interpolation module, and subjecting the features Figure 1 and the features Figure 2 to channel-level connection and convolution operations respectively according to different sequential arrangements, to obtain convolution offset 1 and convolution offset 2;
[0074] Subjecting convolution offset 1, the features Figure 1 and convolution offset 2, the features Figure 2 to deformable convolution operations respectively to obtain sampled features Figure 1 and sampled features Figure 2 ; giving two learnable weights to the two sampled feature maps through the mixing function for the sampled features Figure 1 and the sampled features Figure 2 respectively to generate an intermediate frame, that is, the features Figure 3 。
[0075] Furthermore, the mixing function in the interpolation module can be specifically set to:
[0076] F t =α*T t-1 +β*T t+1
[0077] where α and β both represent learnable convolution kernels of size 1×1, * represents the convolution operation, F t represents the interpolated intermediate frame, that is, the features Figure 3 , T t-1 represents the features Figure 1 , T t+1 represents the features Figure 2 。
[0078] Furthermore, the pyramid alignment module can be set as Figure 3 shown, including:
[0079] Taking the outputs of the feature extraction module and the interpolation module in the order of the features Figure 1 , the features Figure 3 and the features Figure 2 , with every two frames as a group, and sequentially as the input, that is, input 1 and input 2 (that is, when selecting the features Figure 1 , the features Figure 3 as a group, taking the features Figure 1 as input 1 and the features Figure 3 as input 2; when selecting the features Figure 3 and the features Figure 2When there is a group, the feature Figure 3 is used as Input 1, and the feature Figure 2 is used as Input 2);
[0080] First, perform two - fold downsampling and four - fold downsampling on Input 1 and Input 2 respectively to obtain Input 1 - 2, Input 2 - 2, Input 1 - 4, and Input 2 - 4;
[0081] Perform channel - level connection (torch.cat()) on Input 1 - 4 and Input 2 - 4, then perform a convolution operation, calculate the offset for the result after convolution, that is, Offset 1 of Input 1 - 4 and Input 2 - 4, and perform a deformable convolution operation on Offset 1 and Input 2 - 4 to obtain Reconstruction Frame 1; Connect Reconstruction Frame 1 and Input 2 - 4 at the channel level, and input the connected result into a dense residual module to obtain Dense 1;
[0082] Perform channel - level connection on Input 1 - 2 and Input 2 - 2, perform two convolution operations on the connected result, and calculate the offset operation for the convolution result, that is, Offset 2 - 1 of Input 1 - 2 and Input 2 - 2; Upsample Offset 1 by 2 times to obtain Offset 1 - 2, perform channel - level connection on Offset 2 - 1 and Offset 1 - 2 to obtain Offset 2; Perform a deformable convolution operation on Offset 2 and Input 2 - 2 to obtain Reconstruction Frame 2 - 1; Upsample Dense 1 by 2 times to obtain Dense 1 - 2, perform channel - level connection on Dense 1 - 2 and Reconstruction Frame 2 - 1, then input the connected result into a convolution operation to obtain Reconstruction Frame 2, connect Reconstruction Frame 2 and Input 2 - 2 at the channel level, and input the connected result into a dense residual module to obtain Dense 2;
[0083] Perform channel - level connection on Input 1 and Input 2, input the connected result into two convolution operations, then calculate the offset operation for the result after convolution to obtain Offset 3 - 1 of Input 1 and Input 2, upsample Offset 2 by 2 times to obtain Offset 2 - 2, perform channel - level connection on Offset 2 - 2 and Offset 3 - 1 to obtain Offset 3; Perform a deformable convolution operation on Offset 3 and Input 2 to obtain Reconstruction Frame 3 - 1, upsample Dense 2 by 2 times to obtain Dense 2 - 2, perform channel - level connection on Dense 2 - 2 and Reconstruction Frame 3 - 1, then input the connected result into a convolution operation to obtain Reconstruction Frame 3; Connect Reconstruction Frame 3 and Input 2 at the channel level, and then input the connected result into a dense residual module to obtain Dense 3;
[0084] Perform channel - level connection on Input 2 and Dense 3, then perform convolution, deformable convolution, and offset calculation operations on the connected result to obtain the offset between Input 2 and Dense 3, that is, Offset 4, and use the result of performing a deformable convolution operation on Offset 4 and Dense 3 as the output result of the pyramid alignment module.
[0085] The pyramid alignment module is characterized in that: this module is a module that uses deformable convolution based on a pyramid structure to fully utilize spatio-temporal domain information for feature-level adaptive alignment. When performing deformable convolution operations, the offsets of the previous layer are upsampled and fused with the offsets of the current layer to reduce the offset calculation error, and the feature map after alignment and fusion of the previous layer is fused with the deformable convolution of the current layer to reduce the deformable convolution alignment error; at the same time, the dense residual module is used multiple times, which can alleviate the problem that ordinary residual blocks are difficult to fully utilize local residual feature information. This multi-scale structure model can greatly reduce the calculation errors that occur in each stage, thereby improving the final inter-frame alignment accuracy.
[0086] The ConvLSTM module processes the 3 output frames processed by the pyramid alignment module frame by frame to obtain 3 processed frames. ConvLSTM can process the correlation of the time series of consecutive frames to connect the context information between consecutive frames. At the same time, ConvLSTM also has the spatial feature extraction ability of a general CNN structure.
[0087] Furthermore, the dense residual module in the pyramid alignment module can be set, and this module takes the reconstructed frame after channel-level connection in the pyramid alignment module as the input; it includes: n residual blocks, where n is greater than or equal to 0 and is a positive integer;
[0088] The residual block includes a ReLU activation function and two convolution operations; among them, the relevant parameters of the convolution operations are all the same, where the number of input and output channels is 64, the convolution kernel size is 3×3, and the convolution stride is defaulted to 1;
[0089] If n is 0, the input is the output;
[0090] If n is 1, the connection of the residual block is: the input is sent into the first convolution, the ReLU activation function, and the second convolution in sequence, and the output point is node 1; at node 1, the input is added to the output through the above operations, which is the output at node 1, and the result is the output;
[0091] If n takes a value greater than 1: For the structure of the first residual block: The input is successively fed into the first convolution, the ReLU activation function, and the second convolution, and the output point is node 1; at node 1, the input is added to the output obtained through the above operations, which is the output at node 1; for the structures of the second to the (n - 1)-th residual blocks, described by taking the i-th residual block as an example: The output of node i - 1 is successively fed into the (2i - 1)-th convolution, the ReLU activation function, and the 2i-th convolution, and the output point is node i; at node i, the output result is added to the input of node i - 1, which is the output at node i; for the structure of the n residual blocks: The output of node n - 1 is successively fed into the (2n - 1)-th convolution, the ReLU activation function, and the 2n-th convolution, and the output point is node n; at node n, the outputs after the 2n-th convolution, the 2(n - 1)-th convolution,..., and the second convolution are added as the output of node n; the output of node n is successively fed into the (2n + 1)-th convolution and the ReLU activation function, and the output point is node n + 1, and at node n + 1, the output is added to the original input, and the result is the output; where i = 2,... n - 1.
[0092] Taking 4 residual blocks as an example:
[0093] The dense residual module is specifically as follows: The input is successively fed into convolution 1, the ReLU activation function, and convolution 2, and the output point is node 1; at node 1, the input is added to the output obtained through the above operations, which is the output at node 1; the output of node 1 is successively passed through convolution 3, the ReLU activation function, and convolution 4, and the output point is node 2; at node 2, the output result is added to the input at node 1, which is the output of node 2; the output of node 2 is successively passed through convolution 5, the ReLU activation function, and convolution 6, and the output point is node 3; at node 3, the output result is added to the input of node 2, which is the output of node 3; the output of node 3 is successively passed through convolution 7, the ReLU activation function, and convolution 8, and the output point is node 4; at node 4, the outputs after convolution 8, convolution 6, convolution 4, and convolution 2 are added, which is the output of node 4; the output of node 4 is successively fed into convolution 9 and the ReLU activation function, and the output point is node 5, and at node 5, the output is added to the original input, and the result is the output.
[0094] The cascaded residual block is specifically as follows: A conventional residual block is introduced, and 40 conventional residual blocks are cascaded. The 3 processed frames output from the ConvLSTM are used as the input of this module. After being processed by this module, 3 final frames can be obtained. That is, the network output result.
[0095] Further, the ablation experiment can be set, including:
[0096] Under the condition that other conditions are the same, by changing the number of residual blocks in the dense residual module in the pyramid alignment module, multiple different prototypes of deep learning network models are obtained;
[0097] Perform performance evaluations on multiple different prototypes of deep learning network models, and screen the optimal model under the current indicators according to multiple performance indicators.
[0098] Use PSNR and screening derivation speed as screening indicators, and screen the model with the highest PSNR and the model with the fastest derivation speed from them; In the present invention, 8 models are compared, and the two selected models are specifically: One is a network with the highest peak signal-to-noise ratio (PSNR) among the 8 models but a slower network derivation speed. The other is a network with a poor PSNR performance, but the fastest derivation speed and a smaller number of parameters. By selecting two models with large contrasts, it is convenient to select a more suitable method according to different tasks. For example, in the case of insufficient device performance, a network with a faster derivation speed and a smaller number of parameters can be selected to achieve the best efficiency.
[0099] The method of the present invention cascades the video frame interpolation task and the video super-resolution reconstruction task into a framework to achieve end-to-end training.
[0100] Further, it can be set that the extraction step is specifically as follows: The specific implementation steps are as follows:
[0101] The newly collected high-speed rotating rotor image data is divided into training set 1 and validation set 1, and input into the deep learning network model to obtain training set 2 and validation set 2;
[0102] Use annotation software to perform bounding box annotation on the rotors in the images of training set 2 and validation set 2 to generate xml files;
[0103] Input training set 2, validation set 2 and the xml files generated by annotation into Fast-RCNN for training to obtain a weight parameter model; Obtain two corresponding pkl-1 and pkl-2 files; The pkl file contains the upper left and lower right two diagonal vertex coordinates of the bounding box generated during detection;
[0104] Respectively extract the Y-axis data of the diagonal coordinates in the pkl-1 and pkl-2 files to calculate the center point Y-axis coordinates; Plot the center point coordinates of multiple frames as a curve graph, which is the vibration displacement signal of the high-speed vibrating rotor.
[0105] Still further, the present invention provides the following:
[0106] Step 1: The resolution of the collected high-speed rotating rotor images is 512×512. The collected images are successively subjected to 4-fold downsampling and Gaussian blur processing to obtain image data of 128×128. The high-speed rotating rotor image dataset after downsampling and Gaussian blur processing is divided into a training dataset and a validation dataset; all the high-speed rotating rotor image datasets are collected on the ZT-3 rotor vibration simulation test bench. Use a 5F01M high-speed camera to collect images of the high-speed rotating rotor at a collection speed of 2000 frames per second, and use an EF-200LED lamp for optical compensation of the collection object. The resolution of the collected images is 512×512. The collected data is successively subjected to 4-fold downsampling and Gaussian blur processing to obtain training input image data with a resolution of 128×128. The training dataset and the validation dataset respectively account for 70% and 30% of the high-speed rotating rotor image dataset. In this embodiment, a total of 9000 frames of high-speed rotating rotor image datasets are collected; among them, there are 6300 frames in the training dataset and 2700 frames in the validation dataset; subsequently, images of the high-speed rotating rotor to be tested can be collected for further testing.
[0107] Step 2: Use a feature extraction module, an interpolation module, a pyramid alignment module, a ConvLSTM module, and a series of residual blocks to construct a prototype of a deep learning network model;
[0108] Step 3: Conduct ablation experiments on the prototype of the deep learning network model to determine two different deep learning network models;
[0109] The results of the ablation experiment are shown in Table 1. Ablation experiments are conducted on the presence or absence of the dense residual module and the number of dense residual blocks included in the dense residual module. The x in Dense-x represents the number of residual blocks included in the dense residual module. It can be found from Table 1 that different numbers of residual blocks in the dense residual module proposed by the method of the present invention exhibit different performances.
[0110] The method of the present invention is divided into two types according to the performance of the model. One is that Dense-8 has relatively high accuracy but a slow model derivation speed. The other is that Dense-0 has good performance in terms of accuracy, and has a fast derivation speed and small model parameters. The purpose of dividing it into two methods is to be able to select a more suitable method according to different tasks. For example, when the device performance is insufficient, we can choose the Dense-0 model to achieve the highest efficiency, and when the device performance is excellent, we can choose Dense-8 to achieve the best results. The specific values can be determined according to the actual situation.
[0111] Among them, PSNR-RGB represents the image data of the RGB channel, and PSNR-Y represents the grayscale image data. We take the value of PSNR-RGB as a reference.
[0112] Table 1 Results of Ablation Experiments
[0113]
[0114] Step 4: Before formal training, modify the corresponding hyperparameters in the configuration file to obtain training parameters;
[0115] Step 5: Call the training dataset and the configuration file to start training the selected deep learning network model. After training is completed, select the optimal candidate weights;
[0116] Step 6: Use the validation dataset to evaluate the performance of the optimal candidate weights to quantify the performance of the weights. Load the optimal weights obtained according to the quantification results into the deep learning network model;
[0117] Specifically, first set the number of extracted images Batch Size = 4 and the initial learning rate to 10 -4 、the number of iterations is set to 150,000 times, and the learning rate decreases linearly until it drops to 10 -7 after training ends. The operating environment and parameter settings of the other comparison methods used are all set to the default values of the corresponding methods. Start training, load the deep learning network model, load images in batches according to the size of BatchSize for training, and according to the set parameters, output the weights once when the number of iterations reaches the set number, and then select the optimal weights. Quantitatively evaluate the performance of the weights through the validation set, and view the evaluation results such as accuracy, recall rate, mean average precision, etc. to overall quantify the performance of the optimal weights. Load the obtained optimal weights into the deep learning network model, and the output image is obtained after the training image undergoes super-resolution reconstruction.
[0118] Step 7: Use the deep learning network model loaded with the optimal weights to process the newly collected high-speed rotating rotor image dataset, introduce the Fast-RCNN network to identify the bounding box of the rotor in the rotor image, and extract the vibration displacement signal of the high-speed vibrating rotor through the change in the Y-axis displacement of the center point of the rotor bounding box in each frame.
[0119] Specifically, the present invention takes 400 consecutive frame pictures with a resolution of 512×512 collected, and obtains 400 frame pictures with a resolution of 128×128 through downsampling processing. The 400 frame pictures obtained after downsampling processing are divided into 200 frames each for the training set 1 and the validation set 1 used for detection; the training set 1 and the validation set 1 are input into the network model proposed by the present invention to obtain the training set 2 and the validation set 2; the Labelimg software is used to perform bounding box annotation on the rotors in the images of the training set 2 and the validation set 2 to generate xml files; the training set 2, the validation set 2 and the xml files generated by the annotation are input into Fast-RCNN for training to obtain a weight parameter model, and two corresponding pkl-1 and pkl-2 files are obtained; the pkl files contain the coordinates of the upper left and lower right diagonal vertices of the bounding boxes generated during detection; the Y-axis data of the diagonal coordinates in the pkl-1 and pkl-2 files are respectively extracted to calculate the center point Y-axis coordinates; the center point coordinates of multiple frames are plotted as a curve graph, which is the vibration displacement signal of the high-speed vibrating rotor.
[0120] Through Figure 7 comparison, it can be seen that the lightweight version of the present invention, that is, the version with fewer residual blocks in the dense residual module, has better PSNR than other algorithms; from the experiment, it can be seen that even the lightweight version - that is, the worst experimental performance of the present invention, its PSNR is better than other algorithms. Through Figure 8 and Figure 9 comparison, it can be seen that the vibration displacement signal collected from the image data after super-resolution reconstruction is smoother than the data collected from the image data without super-resolution reconstruction.
[0121] The accuracy of the displacement signal of the vibrating object depends on the accuracy of the bounding box generated by the detection algorithm, and the accuracy of the detection algorithm in generating the target bounding box depends on the performance of the algorithm itself and the accuracy of the annotation of the training dataset. Whether the annotation of the training set is accurate seriously affects the accuracy of the bounding box generated by the detection algorithm, thus affecting the accuracy of the finally generated displacement signal of the high-speed vibrating rotor. However, the upper and lower boundaries of the image without the continuous frame restoration algorithm of the present invention are relatively blurred, and it is difficult to label an accurate bounding box. The present invention solves the above problems well. Specifically, due to insufficient camera frame rate when collecting images of high-speed vibrating bodies, frame loss will occur, resulting in partial loss of information in the vibration displacement signal and serious noise. And because the boundaries of the unprocessed images are relatively blurred, it is difficult to ensure that the boundaries of the labeled bounding boxes accurately fall on the target boundaries during the annotation stage when extracting signals by detection methods, resulting in poor accuracy and weak periodicity of the generated displacement vibration signals. The network model proposed by the present invention supplements a frame of picture between two frames, which well alleviates the problem of frame loss during image acquisition, and the sufficient frame information makes the vibration displacement signal relatively smooth and the generated signal has relatively less noise; and further focuses on the reconstruction of the high-frequency information of the frame, so that the target boundary can be clear, and the bounding box that fits the target boundary can be accurately marked during annotation. The model trained with accurately labeled tags can well calculate the bounding box coordinate information of the test set data, so that the generated signal is more in line with the real vibration displacement signal.
[0122] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A high-frequency rotor visual vibration measurement method based on super-resolution reconstruction, characterized in that: Including: A division step of collecting an image dataset of a high-speed rotating rotor and dividing it into a training dataset and a validation dataset; A construction step of constructing a prototype of a deep learning network model; A determination step of performing ablation experiments on the prototype of the deep learning network model to determine two different deep learning network models; An obtaining step of modifying corresponding hyperparameters in a configuration file before formal training to obtain training parameters; A screening step of calling the training dataset and the configuration file to start training the selected deep learning network model, and after the training ends, screening out the optimal candidate weights; A loading step of performing performance evaluation on the optimal candidate weights using the validation dataset to quantify the performance of the weights, and loading the optimal weights obtained according to the quantification result into the deep learning network model; An extraction step of using the deep learning network model loaded with the optimal weights to process a newly collected image dataset of a high-speed rotating rotor and extracting the vibration displacement signal of the high-speed vibrating rotor; The specific construction step is: using a feature extraction module, an interpolation module, a pyramid alignment module, a ConvLSTM module, and a series of residual blocks to construct a prototype of a deep learning network model; The interpolation module includes: Taking the feature map 1 and feature map 2 output by the feature extraction module as the input of the interpolation module, and respectively passing the feature map 1 and feature map 2 through channel-level connection and convolution operations according to different order arrangements to obtain convolution offset 1 and convolution offset 2; Performing deformable convolution operations on convolution offset 1, feature map 1, convolution offset 2, and feature map 2 respectively to obtain sampled feature map 1 and sampled feature map 2; giving two learnable weights to the two sampled feature maps through a mixing function respectively for the sampled feature map 1 and sampled feature map 2 to generate an intermediate frame, that is, feature map 3; The pyramid alignment module includes: Taking the outputs of the feature extraction module and the interpolation module in the order of feature map 1, feature map 3, and feature map 2, with every two frames as a group, as the inputs in turn, that is, input 1 and input 2; First, performing two-fold downsampling and four-fold downsampling on input 1 and input 2 respectively to obtain input 1-2, input 2-2, input 1-4, and input 2-4; Performing channel-level connection on input 1-4 and input 2-4, then performing convolution operation, calculating the offset for the result after convolution, that is, the offset 1 of input 1-4 and input 2-4, and performing variable convolution operation on offset 1 and input 2-4 to obtain a reconstructed frame 1; connecting the reconstructed frame 1 and input 2-4, and inputting the connection result into a dense residual module to obtain dense 1; Perform channel-level connection on Input 1-2 and Input 2-2, perform two convolution operations on the connected result, and perform an operation to calculate the offset on the convolution result, that is, the offset 2-1 of Input 1-2 and Input 2-2; perform 2x upsampling on Offset 1 to obtain Offset 1-2, and perform channel-level connection on Offset 2-1 and Offset 1-2 to obtain Offset 2; perform deformable convolution operation on Offset 2 and Input 2-2 to obtain Reconstructed Frame 2-1; perform 2x upsampling on Dense 1 to obtain Dense 1-2, perform channel-level connection on Dense 1-2 and Reconstructed Frame 2-1, and then input the connected result into a convolution operation to obtain Reconstructed Frame 2, perform channel-level connection on Reconstructed Frame 2 and Input 2-2, and input the connected result into a dense residual module to obtain Dense 2; Perform channel-level connection on Input 1 and Input 2, input the connected result into two convolution operations, and then perform an operation to calculate the offset on the convolved result to obtain the offset 3-1 of Input 1 and Input 2. Perform 2x upsampling on Offset 2 to obtain Offset 2-2, perform channel-level connection on Offset 2-2 and Offset 3-1 to obtain Offset 3; perform deformable convolution operation on Offset 3 and Input 2 to obtain Reconstructed Frame 3-1, perform 2x upsampling on Dense 2 to obtain Dense 2-2, perform channel-level connection on Dense 2-2 and Reconstructed Frame 3-1, and then input the connected result into a convolution operation to obtain Reconstructed Frame 3; perform channel-level connection on Reconstructed Frame 3 and Input 2, and then input the connected result into a dense residual module to obtain Dense 3; Perform channel-level connection on Input 2 and Dense 3, and then perform convolution, deformable convolution, and offset calculation operations on the connected result to obtain the offset of Input 2 and Dense 3, that is, Offset 4. Use the result of performing deformable convolution operation on Offset 4 and Dense 3 as the output result of the pyramid alignment module.
2. The high-frequency rotor visual vibration measurement method based on super-resolution reconstruction according to claim 1, characterized in that: The specific division step is as follows: The image resolution in the collected high-speed rotating rotor image dataset is A×A. Perform 4x downsampling and Gaussian blur processing on all images in the collected high-speed rotating rotor image dataset in sequence to obtain image data with a resolution of B×B, and divide the processed image data into a training dataset and a validation dataset; A > B.
3. The high-frequency rotor visual vibration measurement method based on super-resolution reconstruction according to claim 1, characterized in that: The specific mixing function in the interpolation module is as follows: F t = α * T t-1 + β * T t+1 Among them, both α and β represent learnable convolution kernels with a size of 1×1, * represents the convolution operation, and F t represents the interpolated intermediate frame, that is, feature map 3, and T t-1 represents feature map 1, and T t+1 represents feature map 2.
4. The high-frequency rotor visual vibration measurement method based on super-resolution reconstruction according to claim 1, characterized in that: The dense residual module in the pyramid alignment module takes the reconstructed frame after channel-level connection in the pyramid alignment module as the input; it includes: n residual blocks, where n is greater than or equal to 0 and is a positive integer; The residual block includes a ReLU activation function and two convolution operations; among them, the relevant parameters of the convolution operation are all the same, where the number of input and output channels is 64, the convolution kernel size is 3×3, and the convolution stride is defaulted to 1; If n is 0, the input is the output; If n is 1, the connection of the residual block is: input the input into the first convolution, ReLU activation function, and the second convolution in sequence, and the output point is Node 1; at Node 1, add the input and the output through the above operations, which is the output at Node 1, and the result is the output; If n takes a value greater than 1: For the structure of the first residual block: The input is sequentially fed into the first convolution, the ReLU activation function, and the second convolution, and the output point is node 1; At node 1, the input is added to the output obtained through the above operations, which is the output at node 1; For the structure of the second to the (n - 1)-th residual blocks, described by the i-th residual block: The output of node i - 1 is sequentially fed into the (2i - 1)-th convolution, the ReLU activation function, and the 2i-th convolution, and the output point is node i; At node i, the output result is added to the input of node i - 1, which is the output at node i; For the structure of the n residual blocks: The output of node n - 1 is sequentially fed into the (2n - 1)-th convolution, the ReLU activation function, and the 2n-th convolution, and the output point is node n; At node n, the outputs after the 2n-th convolution, the 2(n - 1)-th convolution,..., and the second convolution are added as the output of node n; The output of node n is sequentially fed into the (2n + 1)-th convolution and the ReLU activation function, and the output point is node n + 1. At node n + 1, the output is added to the original input, and the result is the output; where i = 2,...n - 1.
5. The high-frequency rotor visual vibration measurement method based on super-resolution reconstruction according to claim 1, characterized in that: The ablation experiment includes: Under the same other conditions, by changing the number of residual blocks in the dense residual module of the pyramid alignment module, multiple different prototypes of deep learning network models are obtained; Perform performance evaluations on multiple different prototypes of deep learning network models, and screen the optimal model under the current metrics according to multiple performance metrics.
6. The high-frequency rotor visual vibration measurement method based on super-resolution reconstruction according to claim 1, characterized in that: The specific extraction steps are as follows: The specific implementation steps are as follows: The newly collected high-speed rotating rotor image data is divided into training set 1 and validation set 1, and input into the deep learning network model to obtain training set 2 and validation set 2; Use annotation software to perform bounding box annotation on the rotors in the images of training set 2 and validation set 2 to generate xml files; Input training set 2, validation set 2, and the xml files generated by annotation into Fast-RCNN for training to obtain a weight parameter model; Two corresponding pkl-1 and pkl-2 files are obtained; The pkl file contains the upper left and lower right diagonal vertex coordinates of the bounding boxes generated during detection; Extract the Y-axis data of the diagonal coordinates in the pkl-1 and pkl-2 files respectively to calculate the central point Y-axis coordinates; Plot the multi-frame central point coordinates as a curve graph, which is the vibration displacement signal of the high-speed vibrating rotor.
7. A high-frequency rotor visual vibration measurement system integrating super-resolution reconstruction and video frame interpolation, characterized in that: It includes: A division unit for collecting a high-speed rotating rotor image data set and dividing it into a training data set and a validation data set; A construction unit for constructing a prototype of a deep learning network model; A determination unit for performing an ablation experiment on the prototype of the deep learning network model to determine two different deep learning network models; An acquisition unit for modifying the corresponding hyperparameters in the configuration file before formal training to obtain training parameters; A screening unit for calling the training data set and the configuration file to start training the selected deep learning network model. After the training ends, screen out the optimal candidate weights; A loading unit, configured to perform performance evaluation on the optimal candidate weights using a validation data set to quantify the performance of the weights, and load the optimal weights obtained according to the quantification result into a deep learning network model; An extraction unit, configured to use the deep learning network model loaded with the optimal weights to process a newly collected high-speed rotating rotor image data set, and extract the vibration displacement signal of the high-speed vibrating rotor; The construction step is specifically as follows: using a feature extraction module, an interpolation module, a pyramid alignment module, a ConvLSTM module, and a serial residual block to construct a prototype of a deep learning network model; The interpolation module includes: Taking the feature map 1 and the feature map 2 output by the feature extraction module as the inputs of the interpolation module, and respectively performing channel-level connection and convolution operations on the feature map 1 and the feature map 2 according to different order arrangements to obtain a convolution offset 1 and a convolution offset 2; Performing deformable convolution operations on the convolution offset 1, the feature map 1, the convolution offset 2, and the feature map 2 respectively to obtain a sampled feature map 1 and a sampled feature map 2; and generating an intermediate frame, that is, a feature map 3, by respectively giving two learnable weights to the two sampled feature maps through a mixing function; The pyramid alignment module includes: Taking the outputs of the feature extraction module and the interpolation module in the order of the feature map 1, the feature map 3, and the feature map 2, with every two frames as a group, as the inputs in sequence, that is, the input 1 and the input 2; First, performing two-fold downsampling and four-fold downsampling on the input 1 and the input 2 respectively to obtain the input 1-2, the input 2-2, the input 1-4, and the input 2-4; Performing channel-level connection on the input 1-4 and the input 2-4, then performing a convolution operation, calculating an offset for the result after convolution, that is, the offset 1 of the input 1-4 and the input 2-4, and performing a deformable convolution operation on the offset 1 and the input 2-4 to obtain a reconstructed frame 1; performing channel-level connection on the reconstructed frame 1 and the input 2-4, and inputting the connected result into a dense residual module to obtain a dense 1; Performing channel-level connection on the input 1-2 and the input 2-2, performing two convolution operations on the connected result, and performing an offset calculation operation on the convolution result, that is, the offset 2-1 of the input 1-2 and the input 2-2; performing 2-fold upsampling on the offset 1 to obtain an offset 1-2, performing channel-level connection on the offset 2-1 and the offset 1-2 to obtain an offset 2; performing a deformable convolution operation on the offset 2 and the input 2-2 to obtain a reconstructed frame 2-1; performing 2-fold upsampling on the dense 1 to obtain a dense 1-2, performing channel-level connection on the dense 1-2 and the reconstructed frame 2-1, then inputting the connected result into a convolution operation to obtain a reconstructed frame 2, performing channel-level connection on the reconstructed frame 2 and the input 2-2, and inputting the connected result into a dense residual module to obtain a dense 2; Perform channel-level connection on Input 1 and Input 2, input the connected result into two convolution operations, then perform an operation to calculate the offset on the convolved result to obtain the offset 3-1 of Input 1 and Input 2. Perform 2x upsampling on Offset 2 to obtain Offset 2-2. Perform channel-level connection on Offset 2-2 and Offset 3-1 to obtain Offset 3. Perform deformable convolution operation on Offset 3 and Input 2 to obtain the reconstructed frame 3-1. Perform 2x upsampling on Dense 2 to get Dense 2-2. Perform channel-level connection on Dense 2-2 and the reconstructed frame 3-1, and then input the connected result into a convolution operation to obtain the reconstructed frame 3. Perform channel-level connection on the reconstructed frame 3 and Input 2, and then input the connected result into the dense residual module to obtain Dense 3; Perform channel-level connection on Input 2 and Dense 3, and then perform convolution, deformable convolution, and offset calculation operations on the connected result to obtain the offset of Input 2 and Dense 3, that is, Offset 4. Use the result of performing deformable convolution operation on Offset 4 and Dense 3 as the output result of the pyramid alignment module.