Visual measurement algorithm for rotating structure vibration based on semantic segmentation network

By using a visual measurement algorithm for rotating structure vibration based on semantic segmentation networks, the accuracy and robustness issues of traditional vibration measurement methods in high-speed rotating structures are solved, enabling high-precision vibration monitoring in complex environments and improving measurement and segmentation accuracy.

CN122244125APending Publication Date: 2026-06-19ANHUI POLYTECHNIC UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI POLYTECHNIC UNIV
Filing Date
2026-02-06
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Traditional contact vibration measurement methods suffer from problems such as complex deployment, high cost, and insufficient anti-interference capability in high-speed rotating structures. Non-contact measurement technologies such as eddy current sensors and laser measurement methods have low accuracy and are affected by the environment in complex environments. Existing deep learning algorithms have problems with target detection accuracy and robustness in complex backgrounds in rotating structure measurements.

Method used

A rotational structure vibration visual measurement algorithm based on semantic segmentation network is adopted, combined with the active labeling strategy in the field of visual measurement. Through nested Unet structure, VGGModernBlock feature extraction unit with residual and depth separable convolution and dual cooperative attention module, end-to-end subpixel level displacement curve extraction is achieved, which enhances the ability to adapt to noise interference and target shape distortion in complex scenes.

Benefits of technology

It maintains robust measurement characteristics under extremely complex working conditions, improves measurement accuracy and segmentation boundary delineation accuracy, solves the performance bottleneck of traditional methods in complex environments, and realizes high-precision vibration monitoring of rotating structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244125A_ABST
    Figure CN122244125A_ABST
Patent Text Reader

Abstract

This invention relates to a visual measurement algorithm for rotating structure vibration based on a semantic segmentation network. The algorithm is based on a visual measurement system for rotating structure vibration embedded with a deep learning semantic segmentation network. The system includes an image acquisition module and a visual measurement module. The image acquisition module includes a high-speed industrial camera and a laser sensor. The visual measurement module includes a network segmentation module, a coordinate extraction module, and a signal analysis module. The network segmentation module includes a backbone encoder, a VGGModernBlock feature extraction unit, a linearly deformable convolution module, and a dual cooperative attention module. By introducing a deep learning semantic segmentation network into the field of vibration displacement measurement and combining it with an active labeling strategy in the field of visual measurement, a VGGModernBlock feature extraction unit and a spatial-channel dual cooperative attention module are constructed. This achieves an improvement of over 15% in segmentation boundary delineation and target localization accuracy compared to the benchmark model, effectively solving the performance bottleneck problem of traditional segmentation methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sensor technology, and specifically to a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks. Background Technology

[0002] Vibration measurement, as a non-invasive state sensing method, transforms the dynamic behavior of mechanical systems into quantifiable physical signals. By decoding these information-rich signals, it is possible to gain a deeper understanding of the system's dynamic characteristics, diagnose potential faults early, monitor operational status in real time, and predict and assess structural lifespan. Ultimately, this provides indispensable scientific basis and decision support for intelligent operation and maintenance, reliability design, and safe and economical operation of equipment. Vibration measurement methods are categorized into contact and non-contact methods based on whether the sensor needs to be mounted on the surface of the equipment being measured. Methods relying on contact or quasi-contact sensors such as accelerometers and strain gauges are technologically mature and can accurately measure vibration parameters of machinery under different operating conditions. However, in practical applications, they suffer from problems such as complex deployment, limited measurement points, high cost, and insufficient anti-interference capabilities. Especially for high-speed rotating structures, the sensor attachment method may alter the dynamic characteristics of the measured object, limiting its application in complex environments. To overcome the limitations of contact measurement technology, non-contact visual vibration measurement methods, with their advantages of long-distance, non-contact, and interference-free operation, are gradually gaining widespread application.

[0003] Non-contact measurement technologies include eddy current sensors, laser measurement methods, and optical measurement methods. Eddy current sensors reflect relative displacement by sensing changes in the distance between a metal conductor and the probe. To ensure the accuracy of vibration displacement signals, precise adjustment of the gap between the sensor and the target is required. Furthermore, in multi-point measurement applications, the positions and spacing of the probes must be rationally arranged to avoid coupling interference between magnetic fields. Laser measurement methods, on the other hand, are susceptible to environmental changes and interference, offering low accuracy and high cost when measuring uneven surfaces.

[0004] Due to the strong feature learning capabilities of deep learning networks, researchers have gradually introduced deep learning and convolutional neural networks into the field of visual vibration measurement. Guo et al. and Ilg et al. achieved high-precision estimation of vibration time series by combining optical flow methods with deep networks. Ren et al. used a deep learning-enhanced non-contact feature point matching method, which reduced computation time compared to traditional methods and achieved comparable accuracy to contact accelerometers and traditional model-based techniques. Bai et al. used Mask R-CNN to automatically track displacement changes of feature points on the surface of structures. Teng et al. proposed a progressive SDD method based on deep learning algorithms and digital image correlation (DIC) to effectively improve vibration detection accuracy, based on the known limitations of high-speed 3D digital image correlation vibration measurement. In addition, et al. designed a high-resolution feature learning box to learn the temporal and spatial changes of input data for feature extraction, which can better perform vibration monitoring tasks. Cheng et al. designed an optical imaging model with an optical field compensation mechanism, proving that changes in the optical field are an important factor affecting the accuracy of visual vibration measurement. Traditional methods require additional optical compensation, while deep learning segmentation methods can more naturally and robustly model optical field disturbances. Scholars have introduced bounding box-based target detection algorithms into the field of vibration measurement to achieve real-time visual measurement of structural vibrations, solving the instability of traditional tracking algorithms under conditions of changing illumination, blurring, and occlusion. Wang et al. improved the accuracy of long-distance vibration measurement by combining image super-resolution with a deep learning target detection network, demonstrating that deep models have significant advantages for long-distance scenes. Zhang et al. introduced the segmentation network Mask-RCNN into the field of structural vibration displacement monitoring and proved the effectiveness of the fitted vibration curve. Chai et al. and A. Shen et al. proposed a vibration measurement method based on a semantic segmentation network to address the existing measurement challenges of rotating structures. Later, Chai et al. further addressed the limitations of traditional rotating structure vibration displacement measurement methods by proposing a single-target vibration measurement method based on deep learning semantic segmentation, which can achieve vibration displacement measurement of rotating bodies in complex backgrounds without the need for calibration objects.

[0005] Frame-based target detection algorithms can easily detect the object being measured, effectively improving measurement efficiency. However, when the bounding box and the detected target are poorly aligned, even slight deviations in the target point displacement can significantly reduce the accuracy and precision of the vibration curve. Vibration displacement measurement methods based on deep learning semantic segmentation, utilizing precise pixel masks generated by the segmentation network, allow the target center position to be calculated to the sub-pixel level during the fitting process. Methods that segment the rotor profile holistically, however, are significantly limited in application under extreme conditions such as oil or dust coverage, where the measured object is severely occluded. Summary of the Invention

[0006] This invention addresses the shortcomings of existing technologies by providing a visual measurement algorithm for rotating structure vibration based on a semantic segmentation network. It introduces deep learning semantic segmentation networks into the field of vibration displacement measurement, combining them with active labeling strategies from the visual measurement domain. Through end-to-end sub-pixel level displacement curve extraction capabilities, it maintains robust measurement characteristics even in extremely complex working conditions. Based on a densely nested Unet structure, an improved deep learning semantic segmentation network is proposed. By introducing the concepts of residual and depthwise separable convolution, a VGGModernBlock feature extraction unit is constructed, achieving network lightweighting while maintaining basic feature representation capabilities. Furthermore, a feature decoupling mechanism enhances the model's adaptability to noise interference and target morphology distortion in complex scenes. A spatial-channel dual-cooperative attention module is constructed, combined with a deconvolution upsampling strategy to form a dynamic weight allocation system. Through fine-grained semantic information recovery in the spatial dimension and local detail feature enhancement in the channel dimension, multi-scale information in the feature space is collaboratively reconstructed.

[0007] To achieve the objective of this invention, the technical solution adopted is as follows: A visual measurement algorithm for rotating structure vibration based on a semantic segmentation network is proposed. The rotating structure vibration visual measurement algorithm is based on a rotating structure vibration visual measurement system with an embedded deep learning semantic segmentation network. The rotating structure vibration visual measurement system includes an image acquisition module and a visual measurement module.

[0008] Preferably, the image acquisition module includes a high-speed industrial camera and a laser sensor, and the vision measurement module includes a network segmentation module, a coordinate extraction module, and a signal analysis module; the network segmentation module includes a backbone encoder, a VGG ModernBlock feature extraction unit, a linear deformable convolution module, and a dual cooperative attention module.

[0009] A visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks, the specific steps of which are as follows: S1. High-speed industrial cameras synchronously capture image data of optical marks on the rotor surface; set the left and right high-speed industrial cameras at the same level, the laser of the laser sensor hits the surface of the rotating rotor shaft, the high-speed industrial cameras synchronously image the optical marks on the rotating rotor surface, continuously acquire images to form an image sequence, and store the image sequence in the form of video file to form a video sequence frame; S2. The network segmentation module outputs a mask image frame by frame. The video sequence frames captured by the high-speed industrial camera are input into the trained network segmentation module. The network segmentation module first performs step-by-step feature extraction on the video sequence frames through the backbone encoder and the VGGModernBlock feature extraction unit, outputting a multi-scale semantic feature set. ; The deep features extracted step by step are input into a linearly deformable convolutional module based on a nested upsampling strategy for adaptive resolution restoration, and the output nested upsampling features are combined with a multi-scale semantic feature set. The corresponding scale-coded features are densely spliced ​​and fused. During the fusion, a dual collaborative attention module is introduced to perform real-time weighted enhancement of the channel and spatial dimensions of the fused features, and outputs progressively restored and enhanced high-resolution decoded features. The enhanced high-resolution decoding features are mapped to pixel-level prediction results and binarized to output a binary segmentation mask image. S3. A pixel-level fitting algorithm calculates the vibration displacement offset of the target in each frame; the coordinate extraction module uses a sub-pixel fitting algorithm to extract high-precision target center coordinates from the binary segmentation mask image, and performs center positioning operations on the images acquired by the left and right high-speed industrial cameras respectively to obtain the corresponding two-dimensional image coordinates. Time-reconstruction by combining the intrinsic parameter matrix of a high-speed industrial camera and the relative pose relationship Target's three-dimensional spatial coordinates in the world coordinate system The three-dimensional vibration displacement offset is calculated using inter-frame difference. ; S4. Obtain the vibration displacement curves of the rotating body in the time and frequency domains; the signal analysis module will calculate the three-dimensional vibration displacement offset of each frame image. The original vibration displacement sequence was constructed by arranging the sequences according to the time step. For the original vibration displacement sequence Perform mean-reduction preprocessing to obtain fluctuation signals. For the preprocessed fluctuation signal Perform Discrete Fourier Transform to transform the wave signal Transform from the time domain to the frequency domain and calculate the first... The complex frequency spectrum value of each frequency point Frequency domain vibration amplitude of the frequency component ; vibration amplitude in the frequency domain Search for the frequency point of maximum vibration amplitude in the vibration amplitude spectrum to determine the main vibration frequency of the rotating rotor. The final output is based on the principal vibration frequency. Vibration characteristic data set: time-domain vibration curves and frequency domain amplitude spectrum curve And simultaneously output the main vibration frequency And its corresponding peak frequency domain vibration amplitude.

[0010] Furthermore, the specific operation of step S1 in the visual measurement algorithm for rotating structure vibration is as follows: Two high-speed industrial cameras, one on the left and one on the right, are set to the same horizontal position. The laser from the laser sensor is used to strike the surface of the rotor shaft. The two high-speed industrial cameras simultaneously image the optical marks on the rotor surface at a sampling frequency of 2000 fps. After 10 seconds of continuous acquisition, an image sequence of 20,000 frames is obtained. The image sequence is then stored as a video file to form a video sequence frame. Image frames of approximately 2 seconds are selected from this image sequence for subsequent laser point extraction and vibration displacement calculation.

[0011] Furthermore, the specific operation of step S2 in the visual measurement algorithm for rotating structure vibration is as follows: S21. Input the video sequence frames captured by the left and right high-speed industrial cameras into the trained network segmentation module. Let the size of the video sequence frame be... First, the VGGModernBlock feature extraction unit performs shallow feature mapping on the input video sequence frames to extract initial feature maps containing rich texture details. : ; in, The input video sequence frames are the original images captured by the left and right high-speed industrial cameras; 3 represents the RGB three channels; H and W are the height and width of the image in the video sequence frame, respectively. This represents a feature extraction operator composed of VGGModernBlock feature extraction units, which includes two-stage depthwise separable convolution, group normalization, GN, and GELU activation function, used to perform non-linear feature transformation while maintaining resolution. The backbone encoder consists of multiple cascaded VGGModernBlock feature extraction units and max pooling layers stacked alternately. Each layer first performs downsampling through the max pooling layer of the backbone encoder, and then inputs the downsampled features into the VGGModernBlock feature extraction unit for deep semantic abstraction. Through alternating pooling and convolution operations, it outputs a multi-scale encoded feature set with progressively decreasing resolution from shallow to deep. : ; in, The output is the i-th deep feature, which retains the same spatial resolution as the input image; For the first The hierarchical VGGModernBlock feature extraction unit has a structure similar to... Same, but different channel parameters; This is a max-pooling operator with a step size of 2; The first shallow feature output is the highest resolution feature. These are the lowest resolution semantic features, i.e., the features at the lowest level of the backbone encoder. S22. The network segmentation module uses a step-by-step decoding strategy to restore the resolution of the multi-scale encoded features output by the backbone encoder. Let the current deep encoded features be... The data is then fed into a linearly deformable convolutional module, which automatically calculates the sampling offset based on the target output size to encode deep features. Spatial resampling and resolution upsampling are performed to obtain upsampled features. : ; in, The upsampled features are those whose spatial dimensions are the same as those of deep-coded features at the same spatial scale. Consistent; It is a linearly deformable convolution upsampling operator; For the first Hierarchical deep coding features; Upsampled features are obtained through dense skip connections. Deep coding features at the same spatial scale as those in the backbone encoder Perform channel-level splicing to obtain combined features : ; in, The upsampled features are the features obtained after upsampling. In the backbone encoder and upsampled features Deep coding features at the same spatial scale; S23, Combining features Input to feature fusion unit Feature fusion unit It consists of a VGGModernBlock feature extraction unit and a dual collaborative attention module cascaded together. First, the VGGModernBlock feature extraction unit is used to process the combined features. Perform convolutional integration to suppress aliasing; Then, a dual collaborative attention module is used to perform multidimensional recalibration on the combined features after convolution: in the channel dimension, global average pooling is used to compress spatial information, and the dependencies between channels are captured through convolutional mapping to generate channel attention weights; in the spatial dimension, convolutional operations are used to focus on the spatial distribution features of the target region to generate spatial attention weights; finally, the fused features are weighted and enhanced to output the first-stage decoded features. : ; ; ; in, Channel attention weights; Spatial attention weights; This is a feature fusion function, which includes convolutional integration and attention enhancement operations; This is a global average pooling operation; This is the second 1x1 convolutional layer or fully connected layer FC2, whose function is usually to restore the number of channels from the reduced dimension back to the original number of channels; The GELU activation function is located between two fully connected layers, introducing non-linearity to help the network learn complex inter-channel dependencies. The first 1x1 convolutional layer or fully connected layer FC1 is usually used to reduce dimensionality, reduce the number of parameters, and extract core channel features. Use the Sigmoid activation function; For element-wise multiplication; It is a convolutional integration operator executed by the VGGModernBlock feature extraction unit, used to perform depthwise separable convolution operations on densely concatenated combined features to achieve cross-channel information fusion and feature extraction; S24. Repeat the above steps to construct a multi-level nested decoding path. Generate decoding features at each level in order from deep to shallow and from sparse to dense, and finally obtain the top-level enhanced decoding feature with the same size as the original image. : ; in, For level 0 (highest resolution layer) in the... The output features of each decoding stage are called "top-level enhanced decoding features," and as... With the addition of more and more semantic and detailed information, the accuracy of this feature is improved step by step; This is a feature fusion function, which includes convolutional integration and attention enhancement operations; This is a multi-channel splicing operation, representing the core idea of ​​"dense connectivity": the current node It not only receives upsampled features from the next layer, but also receives features from all nodes at the same previous layer. arrive The output characteristics; The initial shallow features output by the backbone encoder contain the most original image edges and texture details; The adaptive upsampled features from the next level, i.e., level 1, provide the deep semantic context information required by the model and have been restored to the same spatial resolution as level 0. The output feature of the previous decoding node in the same level; Indicates the index number of the decoding stage. Time generated The top-level enhancement decoding features are the final product of the decoding process; S25, Enhance the top-level decoding features of the output. The pixel-level classification head input to the end of the network segmentation module is... Convolutional layers map the number of feature channels to the number of foreground / background categories, outputting pixel-level predicted probability maps. : ; in, The top-level enhancement decoding features are the final product of the decoding process; For pixel-level classification heads ( Convolution operator); The sigmoid activation function is used to map the output to... interval; Pixel-level prediction probability map Binarization thresholding is performed to generate the final binary segmentation mask image. : ; in, Representing coordinates The probability value that the location belongs to the target region; The preset classification threshold is 0.5; 255 represents the target area and 0 represents the background area.

[0012] Furthermore, the specific operation of shallow feature mapping in step S21 of the visual measurement algorithm for rotating structure vibration is as follows: S211, Convert video sequence frames The first convolutional layer of the VGGModernBlock feature extraction unit is input to the video sequence frames. Perform depthwise convolution to independently extract local spatial structure features within each channel. And maintain local spatial structural features The spatial resolution remains unchanged; then, through point-by-point... Convolution on local spatial structure features Perform channel mapping, using a linear combination along the channel dimension, to map the number of channels from 3 to... To achieve cross-channel feature fusion and output cross-channel fused features. : ; ; in, For video sequence frames; To extract local spatial structural features, channel independence is preserved; As a cross-channel fusion feature, it integrates information from different channels; The number of input channels is the number of groups, i.e., groups=3. Depthwise convolution operator; Pointwise convolution operator; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; The number of output channels for the first level; H and W represent the height and width of the image in the video sequence frame, respectively; S212, Cross-channel fusion features of the output Applying group normalization treats all channels as a group (Group=1) for statistical standardization, eliminating internal covariate bias and obtaining normalized features. Introducing the GELU activation function to enhance normalized features. Nonlinear expressive power to obtain activation features To avoid gradient vanishing and preserve original image information, residual connections are introduced: due to the 3-channel input video sequence frames and Activation characteristics of channels The number of channels may vary and needs to be adjusted accordingly. Convolution pairs of input video sequence frames Perform channel projection alignment and obtain video sequence frame residuals. Then, the initial feature map is obtained by summing the residuals. : ; ; ; ; in, It features cross-channel fusion; Normalized features; To activate features; For a group with a group size of 1, the normalization function is used. The activation function for the Gaussian error linear unit; For use in channel alignment Convolution operator; The initial feature map generated is the output of layer 0 of the backbone encoder, and its resolution is the same as the original.Figure 1 To; S213. Map the output initial features. The input to the first-stage downsampling module of the trunk encoder is first used Max pooling reduces spatial resolution and obtains downsampled features. : ; in, Intermediate features after downsampling; This is a max-pooling operator with a step size of 2; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; The number of output channels for the first level; H and W represent the height and width of the image in the video sequence frame, respectively; Then downsample the features The data is fed into the next-level VGGModernBlock feature extraction unit for deep feature extraction to obtain first-level mid-level semantic features. : ; in, These are first-level mid-level semantic features; This is the first-level feature extraction operator, specifically the VGGModernBlock feature extraction unit; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; This represents the number of output channels for the first level. First-level mid-level semantic features The height, compared to the height of the original image. Reducing it by half indicates that the network expands its receptive field through downsampling. First-level mid-level semantic features The width of the image compared to the width of the original image. Shrink by half; The number of channels for the first-level encoded features, typically ,For example This indicates that the "thickness" of the feature has increased, and it can express richer information; S214. Repeat the above steps to progressively expand the receptive field and enhance semantic expressive power, thereby obtaining a multi-scale encoded feature set. : ; in, For the first The encoding characteristics of the output are as follows: the resolution is halved at each level, and the number of channels increases at each level. This is a max-pooling operator with a step size of 2; For the first The hierarchical VGGModernBlock feature extraction unit has a structure similar to... They are the same, but the channel parameters are different.

[0013] Furthermore, the specific operation of step S3 in the visual measurement algorithm for rotating structure vibration is as follows: S31. To eliminate noise interference in non-target areas, the output of the network segmentation module... Binary segmentation mask of the frame Introducing a region of interest constraint operator to obtain the target region mask image. : ; ; in, These are the image pixel coordinates, representing the domain of the image acquired by the high-speed industrial camera; For a predefined rectangular region of interest; This is the ROI mask function; Number the time series frames; This is a binary segmentation mask image; This is the target region mask image after being clipped by the ROI mask function; Mask image of the target region By performing a combination of morphological closing and opening operations, internal holes are filled and boundaries are smoothed to obtain an optimized mask image. : ; in, For morphological combination operators; This is an optimized mask image after morphological closing and opening operations, in which internal holes are filled and edges are smoother. This is the target region mask image after cropping using the ROI mask function, which removes interference noise from the image edges; Optimize the mask image by scanning. Extract the foreground focal length set at the position with a pixel value of 255. : ; in, Foreground point set, which is the set of pixel coordinates in the mask image that belong to the target area with a pixel value of 255; To optimize the mask image; S32, First, check the previous scenic spot. Use an edge detection algorithm to extract the sub-pixel level edge contour point set of the target. : ; in, The subpixel-level edge contour point set is the previous point set. The discrete point sequence extracted from the boundary; For sub-pixel coordinates located on the edge of the target, This represents the total number of edge points; A general quadratic curve model is used for sub-pixel level edge contour point sets. Perform least squares fitting: ; in, The parameters are the equation parameters of the ellipse to be estimated; Although it does not directly participate in the calculation of the center coordinates, it determines the size and shape of the ellipse and is an indispensable constant term in the fitting equation; When the elliptic discrimination condition is satisfied When calculating the coordinates of the geometric center of the ellipse : ; ; in, The parameters of the ellipse equation to be estimated are denoted as . Perform the above steps independently on the left and right high-speed industrial camera images respectively to acquire the images at the same time. The sub-pixel center coordinates of the target in the left view and the corresponding subpixel center coordinates in the right view : ; ; in, The subpixel-level geometric center coordinates obtained by fitting the left view; The subpixel-level geometric center coordinates obtained by fitting the right view; The sub-pixel center coordinates extracted from the left view. The sub-pixel center coordinates extracted from the right view; S33. Based on the pinhole imaging model, determine the sub-pixel center coordinates in the left and right views. , Convert them respectively to normalized imaging plane coordinate vectors : ; ; in, The coordinate vector of the target point on the normalized imaging plane eliminates the influence of camera focal length and principal point offset; These are the intrinsic parameter matrices for the left and right high-speed industrial cameras, respectively. Internal parameter matrix The inverse matrix is ​​used in the formula to transform the pixel coordinate system back to the physically normalized plane system. This represents the homogenized pixel coordinate vector; Combining the relative pose relationship between the left and right high-speed industrial cameras A line-of-sight intersection least squares model is constructed to solve for the three-dimensional spatial coordinates of the target. : ; in, The search operator for the minimum value of the independent variable means: to find a variable... The value of is such that the error function or objective function within the parentheses reaches its minimum value; Given the coordinates of a point in three-dimensional space to be estimated, during the solution process, Continuously adjust in three-dimensional space until the optimal position is found; The square of the L2 norm, which physically represents the square of the Euclidean distance, is used here to measure the geometric error or residual energy from a point in space to the line of sight. These are depth scalar factors in the line-of-sight directions of the left and right high-speed industrial cameras, which determine the specific far-field position of the point on the normalized ray. Solving the above optimization model yields the time steps. Target three-dimensional space coordinates : ; in, Three-dimensional spatial coordinates in the world coordinate system; Select reference time Three-dimensional spatial coordinates Using this as a reference position, calculate the three-dimensional vibration displacement offset between consecutive frames: ; ; ; ; in, They are time points The component displacement offset relative to the reference time. They are time points Three-dimensional spatial coordinates; They are time points Three-dimensional spatial coordinates; For a moment The amplitude of the composite vibration displacement reflects the spatial Euclidean distance of the rotating body from the reference position.

[0014] Furthermore, the specific operation of step S4 in the visual measurement algorithm for rotating structure vibration is as follows: S41, The signal analysis module calculates the three-dimensional vibration displacement offset of each frame of image. According to time step Arrange and construct the original vibration displacement sequence : ; in, This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; For the first The vibration displacement amplitude corresponding to the frame; S42. To eliminate the influence of laser sensor installation errors or static offset of the rotating body on frequency domain analysis and to provide a zero-mean stable input for subsequent spectrum calculations, the original vibration displacement sequence is... Perform mean-reduction processing to obtain fluctuation signals. : ; ; in, This is the arithmetic mean of the original signal; This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; To remove the dynamic vibration signal after static bias, i.e., the wave signal; S43. The preprocessed fluctuation signal Perform a Discrete Fourier Transform (DFT) to convert the signal from the time domain to the frequency domain, and calculate the first... The complex frequency spectrum value of each frequency point : ; in, The imaginary unit (i.e.) ), used to construct the rotation factor ; This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; It is a fluctuation signal; Calculate the next Frequency domain vibration amplitude at each frequency point : ; ; in, The sampling frequency of the left and right high-speed industrial cameras; For the first The complex frequency spectrum values ​​at each frequency point; This is the physical frequency corresponding to the kth frequency point; This represents the total number of frames sampled. This is a complex modular arithmetic operation used to obtain the magnitude of vibration energy; S44, vibration amplitude in the frequency domain Search for the frequency point with the largest amplitude in the frequency domain vibration amplitude spectrum to determine the main vibration frequency of the rotating rotor. : ; in, The operator for finding the maximum value of the independent variable represents the search for the function that makes the function... The frequency variable corresponding to the maximum value ; For the first The frequency domain vibration amplitude at each frequency point; The final output is obtained based on the principal vibration frequency. Vibration characteristic data set: time-domain vibration curves Frequency domain amplitude spectrum curve And simultaneously obtain the output main vibration frequency. and its corresponding peak frequency domain vibration amplitude; ; ; in, This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; It is a fluctuation signal; The discrete physical moment corresponding to the nth sampling frame; This is the physical frequency corresponding to the kth frequency point; For the first The frequency domain vibration amplitude at each frequency point.

[0015] This invention provides a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks, which has the following advantages: (1) Deep learning semantic segmentation network is introduced into the field of vibration displacement measurement. Combined with the active labeling strategy in the field of visual measurement, the end-to-end sub-pixel level displacement curve extraction capability can still maintain robust measurement characteristics in extremely complex working conditions. Compared with mainstream models such as traditional optical flow method and digital image correlation method, this invention shows significant superiority in measurement accuracy indicators. (2) Based on the densely nested Unet structure, an improved deep learning semantic segmentation network is proposed. By introducing the idea of ​​residual and depth-separable convolution, the VGGModernBlock feature extraction unit is constructed. This not only achieves network lightweighting while maintaining the basic feature representation ability, but also enhances the model's adaptability to noise interference and target shape distortion in complex scenes through the feature decoupling mechanism. (3) Construct a spatial-channel dual collaborative attention module, and combine it with a deconvolution upsampling strategy to form a dynamic weight allocation system. Through the restoration of fine-grained semantic information in the spatial dimension and the enhancement of local detail features in the channel dimension, the collaborative reconstruction of multi-scale information in the feature space is realized. The accuracy of segmentation boundary delineation and target localization is improved by more than 15% compared with the benchmark model, effectively solving the performance bottleneck problem of traditional segmentation methods in scenarios such as blurred edges and small target recognition. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the overall process structure of the visual measurement algorithm for rotating structure vibration based on semantic segmentation networks in this invention.

[0017] Figure 2 This is a schematic diagram of the overall network architecture of the visual measurement algorithm for rotating structure vibration based on semantic segmentation network of this invention.

[0018] Figure 3 This is a schematic diagram of the VGGModernBlock feature extraction unit network architecture of the visual measurement algorithm for rotating structure vibration based on semantic segmentation network of this invention.

[0019] Figure 4 This is a schematic diagram of the dual cooperative attention module network architecture of the rotational structure vibration visual measurement algorithm based on semantic segmentation network of this invention.

[0020] Figure 5 These are comparison images of the data augmentation effect of the visual measurement algorithm for rotating structure vibration based on semantic segmentation network of this invention; a and b are comparison images of the same data with different augmentation effects and the original image, respectively.

[0021] Figure 6 This is a comparative illustration of experimental test results for different rotational speeds of the visual measurement algorithm for rotating structure vibration based on semantic segmentation networks, as presented in this invention. Figure 1The rotational speeds from top to bottom are 5 r / s, 15 r / s, and 25 r / s, corresponding to frequencies of approximately 1 Hz, 3 Hz, and 5 Hz, respectively.

[0022] Figure 7 This is a comparative illustration of experimental test results for different rotational speeds of the visual measurement algorithm for rotating structure vibration based on semantic segmentation networks, as presented in this invention. Figure 2 The rotational speeds from left to right are 5 r / s, 15 r / s, and 25 r / s, corresponding to frequencies of approximately 1 Hz, 3 Hz, and 5 Hz, respectively.

[0023] Figure 8 This is a schematic diagram showing the comparison of vibration displacement curves obtained by different algorithms of the visual measurement algorithm for rotating structure vibration based on semantic segmentation network of this invention; the vibration frequency is 1.21 Hz.

[0024] Figure 9 This is a schematic diagram showing the superposition and comparison of vibration displacement curves obtained by different algorithms of the visual measurement algorithm for rotating structure vibration based on semantic segmentation network of this invention; the vibration frequency is 1.21 Hz. Detailed Implementation

[0025] The present invention will be further described and illustrated below with reference to specific embodiments and accompanying drawings.

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] In the description of this invention, it should be understood that the terms "upper", "lower", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0028] like Figure 1 , 2 As shown, a visual measurement algorithm for rotating structure vibration based on a semantic segmentation network is presented. The rotating structure vibration visual measurement algorithm is based on a rotating structure vibration visual measurement system with an embedded deep learning semantic segmentation network. The rotating structure vibration visual measurement system includes an image acquisition module and a visual measurement module.

[0029] like Figure 1 , 2As shown in Figures 3 and 4, a visual measurement algorithm for rotating structure vibration based on semantic segmentation network is presented. The image acquisition module includes a high-speed industrial camera and a laser sensor, and the visual measurement module includes a network segmentation module, a coordinate extraction module, and a signal analysis module. The network segmentation module includes a backbone encoder, a VGG ModernBlock feature extraction unit, a linear deformable convolution module, and a dual cooperative attention module.

[0030] like Figure 1 , 2 As shown in Figures 3 and 4, a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks is described below, with the following specific steps: S1. High-speed industrial cameras synchronously capture image data of optical marks on the rotor surface; set the left and right high-speed industrial cameras at the same level, the laser of the laser sensor hits the surface of the rotating rotor shaft, the high-speed industrial cameras synchronously image the optical marks on the rotating rotor surface, continuously acquire images to form an image sequence, and store the image sequence in the form of video file to form a video sequence frame; S2. The network segmentation module outputs a mask image frame by frame. The video sequence frames captured by the high-speed industrial camera are input into the trained network segmentation module. The network segmentation module first performs step-by-step feature extraction on the video sequence frames through the backbone encoder and the VGGModernBlock feature extraction unit, outputting a multi-scale semantic feature set. ; The deep features extracted step by step are input into a linearly deformable convolutional module based on a nested upsampling strategy for adaptive resolution restoration, and the output nested upsampling features are combined with a multi-scale semantic feature set. The corresponding scale-coded features are densely spliced ​​and fused. During the fusion, a dual collaborative attention module is introduced to perform real-time weighted enhancement of the channel and spatial dimensions of the fused features, and outputs progressively restored and enhanced high-resolution decoded features. The enhanced high-resolution decoding features are mapped to pixel-level prediction results and binarized to output a binary segmentation mask image. S3. A pixel-level fitting algorithm calculates the vibration displacement offset of the target in each frame; the coordinate extraction module uses a sub-pixel fitting algorithm to extract high-precision target center coordinates from the binary segmentation mask image, and performs center positioning operations on the images acquired by the left and right high-speed industrial cameras respectively to obtain the corresponding two-dimensional image coordinates. Time-reconstruction by combining the intrinsic parameter matrix of a high-speed industrial camera and the relative pose relationship Target's three-dimensional spatial coordinates in the world coordinate system The three-dimensional vibration displacement offset is calculated using inter-frame difference. ; S4. Obtain the vibration displacement curves of the rotating body in the time and frequency domains; the signal analysis module will calculate the three-dimensional vibration displacement offset of each frame image. The original vibration displacement sequence was constructed by arranging the sequences according to the time step. For the original vibration displacement sequence Perform mean-reduction preprocessing to obtain fluctuation signals. For the preprocessed fluctuation signal Perform Discrete Fourier Transform to transform the wave signal Transform from the time domain to the frequency domain and calculate the first... The complex frequency spectrum value of each frequency point Frequency domain vibration amplitude of the frequency component ; vibration amplitude in the frequency domain Search for the frequency point of maximum vibration amplitude in the vibration amplitude spectrum to determine the main vibration frequency of the rotating rotor. The final output is based on the principal vibration frequency. Vibration characteristic data set: time-domain vibration curves and frequency domain amplitude spectrum curve And simultaneously output the main vibration frequency And its corresponding peak frequency domain vibration amplitude.

[0031] Furthermore, such as Figure 1 , 2 As shown, a visual measurement algorithm for the vibration of a rotating structure based on a semantic segmentation network is described. The specific operation of step S1 in the visual measurement algorithm for the vibration of a rotating structure is as follows: Two high-speed industrial cameras, one on the left and one on the right, are set to the same horizontal position. The laser from the laser sensor is used to strike the surface of the rotor shaft. The two high-speed industrial cameras simultaneously image the optical marks on the rotor surface at a sampling frequency of 2000 fps. After 10 seconds of continuous acquisition, an image sequence of 20,000 frames is obtained. The image sequence is then stored as a video file to form a video sequence frame. Image frames of approximately 2 seconds are selected from this image sequence for subsequent laser point extraction and vibration displacement calculation.

[0032] Furthermore, such as Figure 1 , 2 As shown in Figures 3 and 4, a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks is described. The specific operation of step S2 in the visual measurement algorithm for the vibration of rotating structures is as follows: S21. Input the video sequence frames captured by the left and right high-speed industrial cameras into the trained network segmentation module. Let the size of the video sequence frame be... First, the VGGModernBlock feature extraction unit performs shallow feature mapping on the input video sequence frames to extract initial feature maps containing rich texture details. : ; in, The input video sequence frames are the original images captured by the left and right high-speed industrial cameras; 3 represents the RGB three channels; H and W are the height and width of the image in the video sequence frame, respectively. This represents a feature extraction operator composed of VGGModernBlock feature extraction units, which includes two-stage depthwise separable convolution, group normalization, GN, and GELU activation function, used to perform non-linear feature transformation while maintaining resolution. The backbone encoder consists of multiple cascaded VGGModernBlock feature extraction units and max pooling layers stacked alternately. Each layer first performs downsampling through the max pooling layer of the backbone encoder, and then inputs the downsampled features into the VGGModernBlock feature extraction unit for deep semantic abstraction. Through alternating pooling and convolution operations, it outputs a multi-scale encoded feature set with progressively decreasing resolution from shallow to deep. : ; in, The output is the i-th deep feature, which retains the same spatial resolution as the input image; For the first The hierarchical VGGModernBlock feature extraction unit has a structure similar to... Same, but different channel parameters; This is a max-pooling operator with a step size of 2; The first shallow feature output is the highest resolution feature. These are the lowest resolution semantic features, i.e., the features at the lowest level of the backbone encoder. S22. The network segmentation module uses a step-by-step decoding strategy to restore the resolution of the multi-scale encoded features output by the backbone encoder. Let the current deep encoded features be... The data is then fed into a linearly deformable convolutional module, which automatically calculates the sampling offset based on the target output size to encode deep features. Spatial resampling and resolution upsampling are performed to obtain upsampled features. : ; in, The upsampled features are those whose spatial dimensions are the same as those of deep-coded features at the same spatial scale. Consistent; It is a linearly deformable convolution upsampling operator; For the first Hierarchical deep coding features; Upsampled features are obtained through dense skip connections. Deep coding features at the same spatial scale as those in the backbone encoder Perform channel-level splicing to obtain combined features : ; in, The upsampled features are the features obtained after upsampling. In the backbone encoder and upsampled features Deep coding features at the same spatial scale; S23, Combining features Input to feature fusion unit Feature fusion unit It consists of a VGGModernBlock feature extraction unit and a dual collaborative attention module cascaded together. First, the VGGModernBlock feature extraction unit is used to process the combined features. Perform convolutional integration to suppress aliasing; Then, a dual collaborative attention module is used to perform multidimensional recalibration on the combined features after convolution: in the channel dimension, global average pooling is used to compress spatial information, and the dependencies between channels are captured through convolutional mapping to generate channel attention weights; in the spatial dimension, convolutional operations are used to focus on the spatial distribution features of the target region to generate spatial attention weights; finally, the fused features are weighted and enhanced to output the first-stage decoded features. : ; ; ; in, Channel attention weights; Spatial attention weights; This is a feature fusion function, which includes convolutional integration and attention enhancement operations; This is a global average pooling operation; This is the second 1x1 convolutional layer or fully connected layer FC2, whose function is usually to restore the number of channels from the reduced dimension back to the original number of channels; The GELU activation function is located between two fully connected layers, introducing non-linearity to help the network learn complex inter-channel dependencies. The first 1x1 convolutional layer or fully connected layer FC1 is usually used to reduce dimensionality, reduce the number of parameters, and extract core channel features. Use the Sigmoid activation function; For element-wise multiplication; It is a convolutional integration operator executed by the VGGModernBlock feature extraction unit, used to perform depthwise separable convolution operations on densely concatenated combined features to achieve cross-channel information fusion and feature extraction; S24. Repeat the above steps to construct a multi-level nested decoding path. Generate decoding features at each level in order from deep to shallow and from sparse to dense, and finally obtain the top-level enhanced decoding feature with the same size as the original image. : ; in, For level 0 (highest resolution layer) in the... The output features of each decoding stage are called "top-level enhanced decoding features," and as... With the addition of more and more semantic and detailed information, the accuracy of this feature is improved step by step; This is a feature fusion function, which includes convolutional integration and attention enhancement operations; This is a multi-channel splicing operation, representing the core idea of ​​"dense connectivity": the current node It not only receives upsampled features from the next layer, but also receives features from all nodes at the same previous layer. arrive The output characteristics; The initial shallow features output by the backbone encoder contain the most original image edges and texture details; The adaptive upsampled features from the next level, i.e., level 1, provide the deep semantic context information required by the model and have been restored to the same spatial resolution as level 0. The output feature of the previous decoding node in the same level; Indicates the index number of the decoding stage. Time generated The top-level enhancement decoding features are the final product of the decoding process; S25, Enhance the top-level decoding features of the output. The pixel-level classification head input to the end of the network segmentation module is... Convolutional layers map the number of feature channels to the number of foreground / background categories, outputting pixel-level predicted probability maps. : ; in, The top-level enhancement decoding features are the final product of the decoding process; For pixel-level classification heads ( Convolution operator); The sigmoid activation function is used to map the output to... interval; Pixel-level prediction probability map Binarization thresholding is performed to generate the final binary segmentation mask image. : ; in, Representing coordinates The probability value that the location belongs to the target region; The preset classification threshold is 0.5; 255 represents the target area and 0 represents the background area.

[0033] Furthermore, such as Figure 1 , 2 As shown in Figure 3, a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks is described. The specific operation of shallow feature mapping in step S21 of the visual measurement algorithm for the vibration of rotating structures is as follows: S211, Convert video sequence frames The first convolutional layer of the VGGModernBlock feature extraction unit is input to the video sequence frames. Perform depthwise convolution to independently extract local spatial structure features within each channel. And maintain local spatial structural features The spatial resolution remains unchanged; then, through point-by-point... Convolution on local spatial structure features Perform channel mapping, using a linear combination along the channel dimension, to map the number of channels from 3 to... To achieve cross-channel feature fusion and output cross-channel fused features. : ; ; in, For video sequence frames; To extract local spatial structural features, channel independence is preserved; As a cross-channel fusion feature, it integrates information from different channels; The number of input channels is the number of groups, i.e., groups=3. Depthwise convolution operator; Pointwise convolution operator; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; The number of output channels for the first level; H and W represent the height and width of the image in the video sequence frame, respectively; S212, Cross-channel fusion features of the output Applying group normalization treats all channels as a group (Group=1) for statistical standardization, eliminating internal covariate bias and obtaining normalized features. Introducing the GELU activation function to enhance normalized features. Nonlinear expressive power to obtain activation features To avoid gradient vanishing and preserve original image information, residual connections are introduced: due to the 3-channel input video sequence frames and Activation characteristics of channels The number of channels may vary and needs to be adjusted accordingly. Convolution pairs of input video sequence frames Perform channel projection alignment and obtain video sequence frame residuals. Then, the initial feature map is obtained by summing the residuals. : ; ; ; ; in, It features cross-channel fusion; Normalized features; To activate features; For a group with a group size of 1, the normalization function is used. The activation function for the Gaussian error linear unit; For use in channel alignment Convolution operator; The initial feature map generated is the output of layer 0 of the backbone encoder, and its resolution is the same as the original. Figure 1 To; S213. Map the output initial features. The input to the first-stage downsampling module of the trunk encoder is first used Max pooling reduces spatial resolution and obtains downsampled features. : ; in, Intermediate features after downsampling; This is a max-pooling operator with a step size of 2; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; The number of output channels for the first level; H and W represent the height and width of the image in the video sequence frame, respectively; Then downsample the features The data is fed into the next-level VGGModernBlock feature extraction unit for deep feature extraction to obtain first-level mid-level semantic features. : ; in, These are first-level mid-level semantic features; This is the first-level feature extraction operator, specifically the VGGModernBlock feature extraction unit; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; This represents the number of output channels for the first level. First-level mid-level semantic features The height, compared to the height of the original image. Reducing it by half indicates that the network expands its receptive field through downsampling. First-level mid-level semantic features The width of the image compared to the width of the original image. Shrink by half; The number of channels for the first-level encoded features, typically ,For example This indicates that the "thickness" of the feature has increased, and it can express richer information; S214. Repeat the above steps to progressively expand the receptive field and enhance semantic expressive power, thereby obtaining a multi-scale encoded feature set. : ; in, For the first The encoding characteristics of the output are as follows: the resolution is halved at each level, and the number of channels increases at each level. This is a max-pooling operator with a step size of 2; For the first The hierarchical VGGModernBlock feature extraction unit has a structure similar to... They are the same, but the channel parameters are different.

[0034] Furthermore, such as Figure 1 , 2 As shown in Figures 3 and 4, a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks is described. The specific operation of step S3 in the visual measurement algorithm for the vibration of rotating structures is as follows: S31. To eliminate noise interference in non-target areas, the output of the network segmentation module... Binary segmentation mask of the frame Introducing a region of interest constraint operator to obtain the target region mask image. : ; ; in, These are the image pixel coordinates, representing the domain of the image acquired by the high-speed industrial camera; For a predefined rectangular region of interest; This is the ROI mask function; Number the time series frames; This is a binary segmentation mask image; This is the target region mask image after being clipped by the ROI mask function; Mask image of the target region By performing a combination of morphological closing and opening operations, internal holes are filled and boundaries are smoothed to obtain an optimized mask image. : ; in, For morphological combination operators; This is an optimized mask image after morphological closing and opening operations, in which internal holes are filled and edges are smoother. This is the target region mask image after cropping using the ROI mask function, which removes interference noise from the image edges; Optimize the mask image by scanning. Extract the foreground focal length set at the position with a pixel value of 255. : ; in, Foreground point set, which is the set of pixel coordinates in the mask image that belong to the target area with a pixel value of 255; To optimize the mask image; S32, First, check the previous scenic spot. Use an edge detection algorithm to extract the sub-pixel level edge contour point set of the target. : ; in, The subpixel-level edge contour point set is the previous point set. The discrete point sequence extracted from the boundary; For sub-pixel coordinates located on the edge of the target, This represents the total number of edge points; A general quadratic curve model is used for sub-pixel level edge contour point sets. Perform least squares fitting: ; in, The parameters are the equation parameters of the ellipse to be estimated; Although it does not directly participate in the calculation of the center coordinates, it determines the size and shape of the ellipse and is an indispensable constant term in the fitting equation; When the elliptic discrimination condition is satisfied When calculating the coordinates of the geometric center of the ellipse : ; ; in, The parameters of the ellipse equation to be estimated are denoted as . Perform the above steps independently on the left and right high-speed industrial camera images respectively to acquire the images at the same time. The sub-pixel center coordinates of the target in the left view and the corresponding subpixel center coordinates in the right view : ; ; in, The subpixel-level geometric center coordinates obtained by fitting the left view; The subpixel-level geometric center coordinates obtained by fitting the right view; The sub-pixel center coordinates extracted from the left view. The sub-pixel center coordinates extracted from the right view; S33. Based on the pinhole imaging model, determine the sub-pixel center coordinates in the left and right views. , Convert them respectively to normalized imaging plane coordinate vectors : ; ; in, The coordinate vector of the target point on the normalized imaging plane eliminates the influence of camera focal length and principal point offset; These are the intrinsic parameter matrices for the left and right high-speed industrial cameras, respectively. Internal parameter matrix The inverse matrix is ​​used in the formula to transform the pixel coordinate system back to the physically normalized plane system. This represents the homogenized pixel coordinate vector; Combining the relative pose relationship between the left and right high-speed industrial cameras A line-of-sight intersection least squares model is constructed to solve for the three-dimensional spatial coordinates of the target. : ; in, The search operator for the minimum value of the independent variable means: to find a variable... The value of is such that the error function or objective function within the parentheses reaches its minimum value; Given the coordinates of a point in three-dimensional space to be estimated, during the solution process, Continuously adjust in three-dimensional space until the optimal position is found; The square of the L2 norm, which physically represents the square of the Euclidean distance, is used here to measure the geometric error or residual energy from a point in space to the line of sight. These are depth scalar factors in the line-of-sight directions of the left and right high-speed industrial cameras, which determine the specific far-field position of the point on the normalized ray. Solving the above optimization model yields the time steps. Target three-dimensional space coordinates : ; in, Three-dimensional spatial coordinates in the world coordinate system; Select reference time Three-dimensional spatial coordinates Using this as a reference position, calculate the three-dimensional vibration displacement offset between consecutive frames: ; ; ; ; in, They are time points The component displacement offset relative to the reference time. They are time points Three-dimensional spatial coordinates; They are time points Three-dimensional spatial coordinates; For a moment The amplitude of the composite vibration displacement reflects the spatial Euclidean distance of the rotating body from the reference position.

[0035] Furthermore, such as Figure 1 , 2 As shown in Figures 3 and 4, a visual measurement algorithm for the vibration of rotating structures based on semantic segmentation networks is described. The specific operation of step S4 in the visual measurement algorithm for the vibration of rotating structures is as follows: S41, The signal analysis module calculates the three-dimensional vibration displacement offset of each frame of image. According to time step Arrange and construct the original vibration displacement sequence : ; in, This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; For the first The vibration displacement amplitude corresponding to the frame; S42. To eliminate the influence of laser sensor installation errors or static offset of the rotating body on frequency domain analysis and to provide a zero-mean stable input for subsequent spectrum calculations, the original vibration displacement sequence is... Perform mean-reduction processing to obtain fluctuation signals. : ; ; in, This is the arithmetic mean of the original signal; This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; To remove the dynamic vibration signal after static bias, i.e., the wave signal; S43. The preprocessed fluctuation signal Perform a Discrete Fourier Transform (DFT) to convert the signal from the time domain to the frequency domain, and calculate the first... The complex frequency spectrum value of each frequency point : ; in, The imaginary unit (i.e.) ), used to construct the rotation factor ; This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; It is a fluctuation signal; Calculate the next Frequency domain vibration amplitude at each frequency point : ; ; in, The sampling frequency of the left and right high-speed industrial cameras; For the first The complex frequency spectrum values ​​at each frequency point; This is the physical frequency corresponding to the kth frequency point; This represents the total number of frames sampled. This is a complex modular arithmetic operation used to obtain the magnitude of vibration energy; S44, vibration amplitude in the frequency domain Search for the frequency point with the largest amplitude in the frequency domain vibration amplitude spectrum to determine the main vibration frequency of the rotating rotor. : ; in, The operator for finding the maximum value of the independent variable represents the search for the function that makes the function... The frequency variable corresponding to the maximum value ; For the first The frequency domain vibration amplitude at each frequency point; The final output is obtained based on the principal vibration frequency. Vibration characteristic data set: time-domain vibration curves Frequency domain amplitude spectrum curve And simultaneously obtain the output main vibration frequency. and its corresponding peak frequency domain vibration amplitude; ; ; in, This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; It is a fluctuation signal; The discrete physical moment corresponding to the nth sampling frame; This is the physical frequency corresponding to the kth frequency point; For the first The frequency domain vibration amplitude at each frequency point.

[0036] Experimental environment This invention uses a rotor with a diameter of 1 m as the experimental object. Two identical MEMERECAM-Q2m high-speed digital cameras are used as high-speed industrial cameras for image acquisition. The high-speed industrial cameras have a resolution of 1920*1080 and are equipped with 50mm magnification lenses. The two high-speed industrial cameras are set at the same horizontal level. A laser from a laser sensor is projected onto the surface of the rotating shaft and captured by the high-speed cameras. The vision system acquires 2000 images per second as the rotor operates at different speeds.

[0037] The experimental environment was based on a 64-bit Windows operating system. The specific parameters of the deep learning semantic segmentation network used in the experiment were as follows: input image size 256×256, batch size 8, initial learning rate set to 0.001, model optimizer Adaptive Moment Estimation (Adam), using BCEDiceloss as the loss function, and a maximum number of iterations of 300. The learning rate scheduling strategy was CosineAnnealingLR.

[0038] Dataset To demonstrate the detection performance of the deep learning semantic segmentation network proposed in this invention, a comparative experiment was conducted using the BUSI (Breast Ultrasound Image) and DeepCrack public datasets with the self-built dataset zd of the detection objects.

[0039] The Breast Ultrasound Images (BUSI) dataset is a widely used public benchmark dataset in the field of breast ultrasound image segmentation. Collected by the Baheya Cancer Center and compiled and released by Al-Dhabyani et al. in 2019, this dataset contains 780 breast ultrasound images of varying resolutions, all derived from actual clinical examinations. Based on lesion type, the images are divided into three categories: 437 benign, 210 malignant, and 133 normal. The tumor images (both benign and malignant) are equipped with binary segmentation masks manually annotated by experienced radiologists to accurately describe the spatial location and boundaries of the tumor region. The 647 benign and malignant images in this dataset are divided into training and validation sets at a ratio of 80% and 20%, respectively.

[0040] DeepCrack is a publicly available dataset for road crack detection and segmentation. This dataset contains high-resolution road images with detailed pixel-level annotations, covering a wide range of crack types and diverse background regions (e.g., normal road surface, crack edges). The dataset aims to reflect the complexity and diversity of crack detection in real-world road environments, providing a challenging testing ground for segmentation algorithms to comprehensively evaluate their ability to accurately detect and segment target regions under different crack morphologies.

[0041] The Zd dataset is a self-collected and labeled dataset focused on high-speed zigzag image analysis. It contains several high-quality images along the rotation axis under different operating conditions. Each image has precise region labels for detecting and segmenting target regions associated with the rotation axis.

[0042] Data Augmentation Since medical images have relatively limited data volume, this invention uses Albumentations as a data augmentation tool. Albumentations is a Python library for data augmentation, specifically designed for image processing in computer vision tasks. It provides a rich set of image augmentation techniques, featuring high speed, flexibility, and cross-platform support. It can be used for data preprocessing and augmentation to improve the robustness and generalization ability of the model. Albumentations supports many common data augmentation operations. During the data augmentation process, this invention employs various image transformation strategies provided by the Albumentations library, including: random horizontal flip (probability 0.5), random vertical flip (probability 0.2), random 90° rotation (probability 0.5), random adjustment of brightness and contrast (probability 0.3), Gaussian noise perturbation (probability 0.2), and random cropping (cropping size set to 224×224, probability 0.5). These augmentation methods effectively improve the model's robustness to input images under different orientations, brightness levels, and noise conditions. Simultaneously, the same transformation is applied to the ground truth segmentation mask corresponding to the original image. An illustration of the augmentation is shown below. Figure 5 As shown.

[0043] Evaluation indicators To objectively and realistically evaluate the performance of the algorithm proposed in this invention, Dice coefficient, Intersection over Union (IOU), Precision, Recall, and F1 score are used as evaluation metrics.

[0044] The higher the values ​​of the above five indicators, the better the segmentation result fits the label value, the higher the similarity, the better the segmentation effect, and the higher the accuracy. The calculation formulas for each indicator are as follows: ; ; ; ; In the formula: These represent true positives, false positives, and false negatives, respectively. In the experimental data of this invention, flooded areas are considered positive, and other areas are considered negative.

[0045] Comparison of different segmentation algorithms To fully verify the effectiveness and superiority of the deep learning semantic segmentation network proposed in this invention for high-precision measurement of vibrations in rotating structures, comparative experiments were conducted with various classic and improved segmentation networks, including UNet, Attention U-Net, UNet++, and DeepLabv3+. Experiments were carried out on BUSI, DeepCrack, and a self-built zd dataset. To ensure fairness and reproducibility, all models were trained under strictly standardized protocols.

[0046] Table 1 shows the performance comparison between the deep learning semantic segmentation network of this invention and mainstream segmentation models on the BUSI, DeepCrack, and self-built zd datasets. Experimental results show that the deep learning semantic segmentation network of this invention significantly outperforms UNet, Attention U-Net, UNet++, UNet+++, and Deeplabv3+ in terms of IoU, Dice, Precision, and Recall on the three datasets. Specifically, on the BUSI dataset, the IoU of the deep learning semantic segmentation network of this invention is improved by 7.78% compared to the second-best Deeplabv3+, and the Dice is improved by 3.68%; on the zd dataset, its Precision and Recall reach over 97%, respectively, which is more than 10% higher than the UNet series models, fully verifying the effectiveness of the lightweight reconstruction module, the dual collaborative attention mechanism, and the deformable convolution upsampling strategy.

[0047] Table 1. Comparison of segmentation results of different network models (%)

[0048] Compared to other models, different segmentation models exhibit progressively different performance in complex scenes. While the UNet series has a simple structure, its feature extraction capabilities are limited, and it is prone to boundary blurring in complex scenarios. Attention U-Net enhances feature focusing through an attention mechanism, but its computational cost is high. UNet+++ improves segmentation accuracy through multi-scale fusion, but its model complexity is high. Deeplabv3+ expands the receptive field using dilated convolutions, but it is prone to losing detailed information in small object segmentation scenarios. This invention's deep learning semantic segmentation network achieves lightweighting through depthwise separable convolutions, enhances feature sensitivity by combining two-dimensional attention, and accurately captures target boundaries using deformable convolutions. This reduces computational costs while maintaining high accuracy, providing a superior segmentation solution for measuring the vibration displacement of rotating objects in industrial scenarios.

[0049] Comparison of vibration displacement measurement results with different segmentation algorithms To systematically evaluate the performance of different segmentation algorithms in measuring the vibration displacement of rotating bodies, this invention conducts comparative experiments based on a self-built laser-marked rotating body dataset. Ten seconds of video data were collected at a sampling frequency of 2000 frames per second. Approximately two seconds of video were randomly selected from the 10-second video, and the segmented mask was obtained using a trained model. The performance of the algorithms was analyzed from both the time and frequency domains by performing displacement regression on the center point coordinates extracted by the segmentation algorithms. This process involved extensive research and development. Figure 6 , 7 The experimental tests at different rotational speeds are shown. For the time-domain amplitude plot, we extracted data from 0 to 3200 frames, with the vertical axis representing the distance from the test point to the center of the left camera and the horizontal axis representing the image number of each test. For the frequency domain plot, to clearly see the frequency peaks, we selected data from 0 to 200 frames, with the vertical axis representing the normalized frequency amplitude and the horizontal axis representing the frequency.

[0050] To verify the effectiveness of the deep learning semantic segmentation network of this invention for vibration measurement, this invention selected shaft data with a rotational speed of 5 r / s and a vibration frequency of approximately 1.21 Hz. This experiment compares the vibration displacement curves obtained by different algorithms to deeply analyze the performance of each algorithm in vibration signal extraction. The experiment compares the algorithm proposed in this invention with classic and advanced algorithms such as UNet, Attention U-Net, UNet++, UNet+++, and deeplabv3+, examining the effectiveness of each algorithm from both overall and local perspectives. The experimental results are shown in the figures. Figure 8 , 9 As shown.

[0051] Comparison of multiple algorithms for overall vibration signals Figure 9 It can be seen that the vibration displacement curves extracted by different algorithms have certain similarities in trend, and all can capture the general change pattern of the vibration signal. However, there are significant differences in details. Compared with other algorithms, the curve obtained by the algorithm proposed in this invention more closely follows the ideal or real vibration displacement change trajectory, and the curve fluctuation is more reasonable and stable. This indicates that the algorithm can extract the vibration signal more accurately overall and effectively reduce the influence of noise and other interference factors. In contrast, the curve fluctuation of the UNet algorithm is relatively large, indicating that it may have introduced more noise in the signal extraction process and its capture of vibration features is not accurate enough. Although Attention U-Net improves performance to some extent through the attention mechanism, the curve still has some irregular fluctuations, indicating that its ability to extract complex vibration features still has room for improvement. The curves of UNet++ and deeplabv3+ algorithms also have their own fluctuation characteristics, but their overall accuracy does not surpass that of the algorithm proposed in this invention.

[0052] Further observation of the magnified local details reveals even more significant differences. Within the 0-100 sample index range, the curve of the algorithm proposed in this invention (Ours) more delicately reflects the local changes in the vibration signal and has a stronger ability to capture minute vibration features. The numerical changes of its curve at each sample point are more consistent with the vibration law, demonstrating higher resolution and accuracy. Other algorithms, however, perform relatively worse in this local area. The UNet curve exhibits coarse fluctuations, potentially missing some crucial local vibration information; while the Attention U-Net curve shows some improvement, it still suffers from local biases; compared to the algorithm proposed in this invention, the UNet+++ and deeplabv3+ algorithms show obvious irregular fluctuations and noise in their axis-track trajectories, with trajectory shapes deviating significantly from the theoretical trajectory, indicating substantial systematic errors in displacement measurement. Experimental results demonstrate that the optimized deep learning semantic segmentation network structure effectively solves key problems such as inaccurate positioning of minute marker points and significant noise interference in the vibration measurement of rotating bodies, providing reliable technical support for high-precision non-contact vibration measurement. Its advantages in vibration displacement measurement accuracy, harmonic component capture capability, and trajectory shape fidelity make it an algorithm with great application potential in the field of rotating body vibration measurement.

[0053] ablation experiment This invention conducted a series of ablation experiments to verify the effectiveness of the three proposed key components (Depthly Separable Convolution, GELU activation function, and GroupNorm), as shown in Table 2. The baseline model (Row 1) had an IoU of 61.83% without any improved components. Experimental results show that the introduction of different components can improve model performance to varying degrees. Among them, the combination of Depthly Separable Convolution (DSC) and Group Normalization (GN) (Row 3) achieved the best segmentation results, with IoU and Dice coefficients reaching 64.28% and 76.89%, respectively, representing improvements of 2.45% and 1.98% compared to the baseline. These results fully demonstrate the effectiveness of our proposed three components in improving model segmentation performance, especially the significant impact of Depthly Separable Convolution and GroupNorm on model performance.

[0054] Table 2 Comparison of ablation experiments of key components in the VGGModernBlock feature extraction unit on the BUSI dataset.

[0055] To further analyze the data, this invention observes that the model using the GELU activation function and GroupNorm maintains high accuracy while also exhibiting relatively high recall, indicating that the deep learning semantic segmentation network model of this invention achieves a better balance in classifying positive and negative samples. Conversely, using depthwise separable convolution alone without GroupNorm results in a performance degradation, demonstrating the crucial role of GroupNorm in stabilizing training and improving the model's generalization ability. In summary, the three components proposed in this invention work together to enhance the model's segmentation accuracy and robustness.

[0056] In engineering scenarios involving large rotating equipment, long-distance photography, or obscured target parts, methods for overall rotor contour segmentation face significant limitations. This invention employs an active labeling strategy, using a high-speed industrial camera as the image acquisition medium, and introduces a deep learning-based semantic segmentation method to address the challenge of visual vibration measurement of rotating objects. To this end, a highly efficient and lightweight segmentation network—a deep learning semantic segmentation network—is constructed to robustly extract pixel-level contours of active light spots against complex backgrounds. Furthermore, this invention will investigate more efficient network architectures and algorithm optimization strategies to improve processing speed and meet real-time monitoring requirements. This will provide a more effective method for establishing complete vibration measurement systems for rotating objects in more complex industrial environments in the future.

[0057] This invention introduces a deep learning semantic segmentation network into the field of vibration displacement measurement, combined with an active labeling strategy from the field of visual measurement. Through end-to-end sub-pixel level displacement curve extraction capabilities, it maintains robust measurement characteristics even in extremely complex working conditions. Compared to mainstream models such as traditional optical flow and digital image correlation methods, this invention demonstrates significant superiority in measurement accuracy. An improved deep learning semantic segmentation network is proposed based on a densely nested Unet structure. By introducing the idea of ​​residual and depthwise separable convolution to construct the VGGModernBlock feature extraction unit, it not only achieves network lightweighting while maintaining basic feature representation capabilities, but also enhances the model's adaptability to noise interference and target morphology distortion in complex scenes through a feature decoupling mechanism. A spatial-channel dual collaborative attention module is constructed, combined with a deconvolution upsampling strategy to form a dynamic weight allocation system. Through fine-grained semantic information recovery in the spatial dimension and local detail feature enhancement in the channel dimension, it achieves collaborative reconstruction of multi-scale information in the feature space. The accuracy of segmentation boundary delineation and target localization is improved by 15% compared to the benchmark model. With a success rate of over 90%, it effectively solves the performance bottleneck problem of traditional segmentation methods in scenarios such as blurred edges and small target recognition.

[0058] The technical solutions disclosed in the embodiments of the present invention have been described in detail above. Specific embodiments have been used to illustrate the principles and implementation methods of the embodiments of the present invention. The above description of the embodiments is only for helping to understand the principles of the embodiments of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A rotation structure vibration visual measurement algorithm based on a semantic segmentation network, characterized in that, The rotating structure vibration visual measurement algorithm is based on a rotating structure vibration visual measurement system with an embedded deep learning semantic segmentation network. The rotating structure vibration visual measurement system includes an image acquisition module and a visual measurement module. The image acquisition module includes a high-speed industrial camera and a laser sensor; the vision measurement module includes a network segmentation module, a coordinate extraction module, and a signal analysis module; the network segmentation module includes a backbone encoder, a VGGModernBlock feature extraction unit, a linear deformable convolution module, and a dual cooperative attention module. The specific steps of the visual measurement algorithm for rotating structure vibration are as follows: S1. High-speed industrial cameras synchronously capture image data of optical marks on the rotor surface; set the left and right high-speed industrial cameras at the same level, the laser of the laser sensor hits the surface of the rotating rotor shaft, the high-speed industrial cameras synchronously image the optical marks on the rotating rotor surface, continuously acquire images to form an image sequence, and store the image sequence in the form of video file to form a video sequence frame; S2, the network segmentation module outputs the mask picture frame by frame; the video sequence frame collected by the high-speed industrial camera is input into the trained network segmentation module, the network segmentation module first extracts features from the video sequence frame step by step through the backbone encoder and the VGGModernBlock feature extraction unit, and outputs a multi-scale semantic feature set ; The step-by-step extracted deep features are input into a linear deformable convolution module based on a nested upsampling strategy for adaptive resolution recovery, and the output nested upsampling features are densely spliced and fused with the corresponding scale coding features in the multi-scale semantic feature set A double cooperative attention module is introduced for real-time weighting and enhancement of the channel and spatial dimensions of the fused features, and high-resolution decoding features recovered step by step and enhanced are output. The enhanced high-resolution decoding features are mapped to pixel-level prediction results and binarized to output a binary segmentation mask image. S3, the pixel-level fitting algorithm calculates the vibration displacement offset of each frame target; the coordinate extraction module extracts high-precision target center coordinates from the binary segmentation mask image by using a sub-pixel fitting algorithm, performs center positioning operations on the images collected by the left and right high-speed industrial cameras respectively, and obtains corresponding two-dimensional image coordinates , combined with the internal parameter matrix of the high-speed industrial camera and the relative pose relationship at the reconstruction moment , the three-dimensional space coordinates of the target in the world coordinate system , and the three-dimensional vibration displacement offset is calculated through inter-frame difference ; S4. Obtain the vibration displacement curves of the rotating body in the time and frequency domains; the signal analysis module will calculate the three-dimensional vibration displacement offset of each frame image. The original vibration displacement sequence was constructed by arranging the sequences according to the time step. For the original vibration displacement sequence Perform mean-reduction preprocessing to obtain fluctuation signals. For the preprocessed fluctuation signal Perform Discrete Fourier Transform to transform the wave signal Transform from the time domain to the frequency domain and calculate the first... The complex frequency spectrum value of each frequency point Frequency domain vibration amplitude of the frequency component ; vibration amplitude in the frequency domain Search for the frequency point of maximum vibration amplitude in the vibration amplitude spectrum to determine the main vibration frequency of the rotating rotor. The final output is based on the principal vibration frequency. Vibration characteristic data set: time-domain vibration curves and frequency domain amplitude spectrum curve And simultaneously output the main vibration frequency And its corresponding peak frequency domain vibration amplitude.

2. The visual measurement algorithm for vibration of rotating structures based on semantic segmentation networks according to claim 1, characterized in that, The specific operation of step S1 in the visual measurement algorithm for vibration of the rotating structure is as follows: Two high-speed industrial cameras, one on the left and one on the right, are set to the same horizontal position. The laser from the laser sensor is used to strike the surface of the rotor shaft. The two high-speed industrial cameras simultaneously image the optical marks on the rotor surface at a sampling frequency of 2000 fps. After 10 seconds of continuous acquisition, an image sequence of 20,000 frames is obtained. The image sequence is then stored as a video file to form a video sequence frame. Image frames of approximately 2 seconds are selected from this image sequence for subsequent laser point extraction and vibration displacement calculation.

3. The visual measurement algorithm for vibration of rotating structures based on semantic segmentation networks according to claim 1, characterized in that, The specific operation of step S2 in the visual measurement algorithm for vibration of the rotating structure is as follows: S21. Input the video sequence frames captured by the left and right high-speed industrial cameras into the trained network segmentation module. Let the size of the video sequence frame be... First, the VGGModernBlock feature extraction unit performs shallow feature mapping on the input video sequence frames to extract initial feature maps containing rich texture details. : ; in, The input video sequence frames are the original images captured by the left and right high-speed industrial cameras; 3 represents the RGB three channels; H and W are the height and width of the image in the video sequence frame, respectively. This represents a feature extraction operator composed of VGGModernBlock feature extraction units, which includes two-stage depthwise separable convolution, group normalization, GN, and GELU activation function, used to perform non-linear feature transformation while maintaining resolution. The backbone encoder consists of multiple cascaded VGGModernBlock feature extraction units and max pooling layers stacked alternately. Each layer first performs downsampling through the max pooling layer of the backbone encoder, and then inputs the downsampled features into the VGGModernBlock feature extraction unit for deep semantic abstraction. Through alternating pooling and convolution operations, it outputs a multi-scale encoded feature set with progressively decreasing resolution from shallow to deep. : ; in, The output is the i-th deep feature, which retains the same spatial resolution as the input image; For the first The hierarchical VGGModernBlock feature extraction unit has a structure similar to... Same, but different channel parameters; This is a max-pooling operator with a step size of 2; The first shallow feature output is the highest resolution feature. These are the lowest resolution semantic features, i.e., the features at the lowest level of the backbone encoder. S22. The network segmentation module uses a step-by-step decoding strategy to restore the resolution of the multi-scale encoded features output by the backbone encoder. Let the current deep encoded features be... The data is then fed into a linearly deformable convolutional module, which automatically calculates the sampling offset based on the target output size to encode deep features. Spatial resampling and resolution upsampling are performed to obtain upsampled features. : ; in, The upsampled features are those whose spatial dimensions are the same as those of deep-coded features at the same spatial scale. Consistent; It is a linearly deformable convolution upsampling operator; For the first Hierarchical deep coding features; Upsampled features are obtained through dense skip connections. Deep coding features at the same spatial scale as those in the backbone encoder Perform channel-level splicing to obtain combined features : ; in, The upsampled features are the features obtained after upsampling. In the backbone encoder and upsampled features Deep coding features at the same spatial scale; S23, Combining features Input to feature fusion unit Feature fusion unit It consists of a VGGModernBlock feature extraction unit and a dual collaborative attention module cascaded together. First, the VGGModernBlock feature extraction unit is used to process the combined features. Perform convolutional integration to suppress aliasing; Then, a dual collaborative attention module is used to perform multidimensional recalibration on the combined features after convolution: in the channel dimension, global average pooling is used to compress spatial information, and the dependencies between channels are captured through convolutional mapping to generate channel attention weights; in the spatial dimension, convolutional operations are used to focus on the spatial distribution features of the target region to generate spatial attention weights; finally, the fused features are weighted and enhanced to output the first-stage decoded features. : ; ; ; in, Channel attention weights; Spatial attention weights; This is a feature fusion function, which includes convolutional integration and attention enhancement operations; This is a global average pooling operation; This is the second 1x1 convolutional layer or fully connected layer FC2, whose function is usually to restore the number of channels from the reduced dimension back to the original number of channels; The GELU activation function is located between two fully connected layers, introducing non-linearity to help the network learn complex inter-channel dependencies. The first 1x1 convolutional layer or fully connected layer FC1 is usually used to reduce dimensionality, reduce the number of parameters, and extract core channel features. Use the Sigmoid activation function; For element-wise multiplication; It is a convolutional integration operator executed by the VGGModernBlock feature extraction unit, used to perform depthwise separable convolution operations on densely concatenated combined features to achieve cross-channel information fusion and feature extraction; S24. Repeat the above steps to construct a multi-level nested decoding path. Generate decoding features at each level in order from deep to shallow and from sparse to dense, and finally obtain the top-level enhanced decoding feature with the same size as the original image. : ; in, For level 0 (highest resolution layer) in the... The output features of each decoding stage are called "top-level enhanced decoding features," as... With the addition of more and more semantic and detailed information, the accuracy of this feature is improved step by step; This is a feature fusion function, which includes convolutional integration and attention enhancement operations; This is a multi-channel splicing operation, representing the core idea of ​​"dense connectivity": the current node It not only receives upsampled features from the next layer, but also receives features from all nodes at the same previous layer. arrive The output characteristics; The initial shallow features output by the backbone encoder contain the most original image edges and texture details; The adaptive upsampled features from the next level, i.e., level 1, provide the deep semantic context information required by the model and have been restored to the same spatial resolution as level 0. The output feature of the previous decoding node in the same level; Indicates the index number of the decoding stage. Time generated The top-level enhancement decoding features are the final product of the decoding process; S25, Enhance the top-level decoding features of the output. The pixel-level classification head input to the end of the network segmentation module is... Convolutional layers map the number of feature channels to the number of foreground / background categories, outputting pixel-level predicted probability maps. : ; in, The top-level enhancement decoding features are the final product of the decoding process; For pixel-level classification heads ( Convolution operator); The sigmoid activation function is used to map the output to... interval; Pixel-level prediction probability map Binarization thresholding is performed to generate the final binary segmentation mask image. : ; in, Representing coordinates The probability value that the location belongs to the target region; The preset classification threshold is 0.5; 255 represents the target area and 0 represents the background area.

4. The visual measurement algorithm for vibration of rotating structures based on semantic segmentation networks according to claim 3, characterized in that, The specific operation of shallow feature mapping in step S21 of the visual measurement algorithm for vibration of rotating structures is as follows: S211, Convert video sequence frames The first convolutional layer of the VGGModernBlock feature extraction unit is input to the video sequence frames. Perform depthwise convolution to independently extract local spatial structure features within each channel. And maintain local spatial structural features The spatial resolution remains unchanged; then, through point-by-point... Convolution on local spatial structure features Perform channel mapping, using a linear combination along the channel dimension, to map the number of channels from 3 to... To achieve cross-channel feature fusion and output cross-channel fused features. : ; ; in, For video sequence frames; To extract local spatial structural features, channel independence is preserved; As a cross-channel fusion feature, it integrates information from different channels; The number of input channels is the number of groups, i.e., groups=3. Depthwise convolution operator; Pointwise convolution operator; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; The number of output channels for the first level; H and W represent the height and width of the image in the video sequence frame, respectively; S212, Cross-channel fusion features of the output Applying group normalization treats all channels as a group (Group=1) for statistical standardization, eliminating internal covariate bias and obtaining normalized features. Introducing the GELU activation function to enhance normalized features. Nonlinear expressive power to obtain activation features To avoid gradient vanishing and preserve original image information, residual connections are introduced: due to the 3-channel input video sequence frames and Activation characteristics of channels The number of channels may vary and needs to be adjusted accordingly. Convolution pairs of input video sequence frames Perform channel projection alignment and obtain video sequence frame residuals. Then, the initial feature map is obtained by summing the residuals. : ; ; ; ; in, It features cross-channel fusion; Normalized features; To activate features; For a group with a group size of 1, the normalization function is used. The activation function for the Gaussian error linear unit; For use in channel alignment Convolution operator; The initial feature map generated is the output of layer 0 of the backbone encoder, and its resolution is the same as the original image. S213. Map the output initial features. The input to the first-stage downsampling module of the trunk encoder is first used Max pooling reduces spatial resolution and obtains downsampled features. : ; in, Intermediate features after downsampling; This is a max-pooling operator with a step size of 2; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; The number of output channels for the first level; H and W represent the height and width of the image in the video sequence frame, respectively; Then downsample the features The data is fed into the next-level VGGModernBlock feature extraction unit for deep feature extraction to obtain first-level mid-level semantic features. : ; in, These are first-level mid-level semantic features; This is the first-level feature extraction operator, specifically the VGGModernBlock feature extraction unit; It is a vector space over the real number field, used to describe the dimension and shape of data tensors; This represents the number of output channels for the first level. First-level mid-level semantic features The height, compared to the height of the original image. Reducing it by half indicates that the network expands its receptive field through downsampling. First-level mid-level semantic features The width of the image compared to the width of the original image. Shrink by half; The number of channels for the first-level encoded features, typically ,For example This indicates that the "thickness" of the feature has increased, and it can express richer information; S214. Repeat the above steps to progressively expand the receptive field and enhance semantic expressive power, thereby obtaining a multi-scale encoded feature set. : ; in, For the first The encoding characteristics of the output are as follows: the resolution is halved at each level, and the number of channels increases at each level. This is a max-pooling operator with a step size of 2; For the first The hierarchical VGGModernBlock feature extraction unit has a structure similar to... They are the same, but the channel parameters are different.

5. The visual measurement algorithm for vibration of rotating structures based on semantic segmentation networks according to claim 1, characterized in that, The specific operation of step S3 in the visual measurement algorithm for vibration of the rotating structure is as follows: S31. To eliminate noise interference in non-target areas, the output of the network segmentation module... Binary segmentation mask of the frame Introducing a region of interest constraint operator to obtain the target region mask image. : ; ; in, These are the image pixel coordinates, representing the domain of the image acquired by the high-speed industrial camera; For a predefined rectangular region of interest; This is the ROI mask function; Number the time series frames; This is a binary segmentation mask image; This is the target region mask image after being clipped by the ROI mask function; Mask image of the target region By performing a combination of morphological closing and opening operations, internal holes are filled and boundaries are smoothed to obtain an optimized mask image. : ; in, For morphological combination operators; This is an optimized mask image after morphological closing and opening operations, in which internal holes are filled and edges are smoother. This is the target region mask image after cropping using the ROI mask function, which removes interference noise from the image edges; Optimize the mask image by scanning. Extract the foreground focal length set at the position with a pixel value of 255. : ; in, Foreground point set, which is the set of pixel coordinates in the mask image that belong to the target area with a pixel value of 255; To optimize the mask image; S32, First, check the previous scenic spot. Use an edge detection algorithm to extract the sub-pixel level edge contour point set of the target. : ; in, The subpixel-level edge contour point set is the previous point set. The discrete point sequence extracted from the boundary; For sub-pixel coordinates located on the edge of the target, This represents the total number of edge points; A general quadratic curve model is used for sub-pixel level edge contour point sets. Perform least squares fitting: ; in, The parameters are the equation parameters of the ellipse to be estimated; Although it does not directly participate in the calculation of the center coordinates, it determines the size and shape of the ellipse and is an indispensable constant term in the fitting equation; When the elliptic discrimination condition is satisfied When calculating the coordinates of the geometric center of the ellipse : ; ; in, The parameters of the ellipse equation to be estimated are denoted as . Perform the above steps independently on the left and right high-speed industrial camera images respectively to acquire the images at the same time. The sub-pixel center coordinates of the target in the left view and the corresponding subpixel center coordinates in the right view : ; ; in, The subpixel-level geometric center coordinates obtained by fitting the left view; The subpixel-level geometric center coordinates obtained by fitting the right view; The sub-pixel center coordinates extracted from the left view. The sub-pixel center coordinates extracted from the right view; S33. Based on the pinhole imaging model, determine the sub-pixel center coordinates in the left and right views. , Convert them respectively to normalized imaging plane coordinate vectors : ; ; in, The coordinate vector of the target point on the normalized imaging plane eliminates the influence of camera focal length and principal point offset; These are the intrinsic parameter matrices for the left and right high-speed industrial cameras, respectively. Internal parameter matrix The inverse matrix is ​​used in the formula to transform the pixel coordinate system back to the physically normalized plane system. This represents the homogenized pixel coordinate vector; Combining the relative pose relationship between the left and right high-speed industrial cameras A line-of-sight intersection least squares model is constructed to solve for the three-dimensional spatial coordinates of the target. : ; in, The search operator for the minimum value of the independent variable means: to find a variable... The value of is such that the error function or objective function within the parentheses reaches its minimum value; Given the coordinates of a point in three-dimensional space to be estimated, during the solution process, Continuously adjust in three-dimensional space until the optimal position is found; The square of the L2 norm, which physically represents the square of the Euclidean distance, is used here to measure the geometric error or residual energy from a point in space to the line of sight. These are depth scalar factors in the line-of-sight directions of the left and right high-speed industrial cameras, which determine the specific far-field position of the point on the normalized ray. Solving the above optimization model yields the time steps. Target three-dimensional space coordinates : ; in, Three-dimensional spatial coordinates in the world coordinate system; Select reference time Three-dimensional spatial coordinates Using this as a reference position, calculate the three-dimensional vibration displacement offset between consecutive frames: ; ; ; ; in, They are time points The component displacement offset relative to the reference time. They are time points Three-dimensional spatial coordinates; They are time points Three-dimensional spatial coordinates; For a moment The amplitude of the composite vibration displacement reflects the spatial Euclidean distance of the rotating body from the reference position.

6. The visual measurement algorithm for vibration of rotating structures based on semantic segmentation networks according to claim 1, characterized in that, The specific operation of step S4 in the visual measurement algorithm for vibration of the rotating structure is as follows: S41, The signal analysis module calculates the three-dimensional vibration displacement offset of each frame of image. According to time step Arrange and construct the original vibration displacement sequence : ; in, This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; For the first The vibration displacement amplitude corresponding to the frame; S42. To eliminate the influence of laser sensor installation errors or static offset of the rotating body on frequency domain analysis and to provide a zero-mean stable input for subsequent spectrum calculations, the original vibration displacement sequence is... Perform mean-reduction processing to obtain fluctuation signals. : ; ; in, This is the arithmetic mean of the original signal; This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; To remove the dynamic vibration signal after static bias, i.e., the wave signal; S43. The preprocessed fluctuation signal Perform a Discrete Fourier Transform (DFT) to convert the signal from the time domain to the frequency domain, and calculate the first... The complex frequency spectrum value of each frequency point : ; in, The imaginary unit (i.e.) ), used to construct the rotation factor ; This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; It is a fluctuation signal; Calculate the next Frequency domain vibration amplitude at each frequency point : ; ; in, The sampling frequency of the left and right high-speed industrial cameras; For the first The complex frequency spectrum values ​​at each frequency point; This is the physical frequency corresponding to the kth frequency point; This represents the total number of frames sampled. This is a complex modular arithmetic operation used to obtain the magnitude of vibration energy; S44, vibration amplitude in the frequency domain Search for the frequency point with the largest amplitude in the frequency domain vibration amplitude spectrum to determine the main vibration frequency of the rotating rotor. : ; in, The operator for finding the maximum value of the independent variable represents the search for the function that makes the function... The frequency variable corresponding to the maximum value ; For the first The frequency domain vibration amplitude at each frequency point; The final output is obtained based on the principal vibration frequency. Vibration characteristic data set: time-domain vibration curves Frequency domain amplitude spectrum curve And simultaneously obtain the output main vibration frequency. and its corresponding peak frequency domain vibration amplitude; ; ; in, This represents the total number of frames sampled. This is a discrete time series index, with a value range of [value range missing]. arrive ; It is a fluctuation signal; The discrete physical moment corresponding to the nth sampling frame; This is the physical frequency corresponding to the kth frequency point; For the first The frequency domain vibration amplitude at each frequency point.