High-precision binocular vision rotor three-dimensional vibration measurement method and system based on segmentation constraint

By employing a high-precision binocular vision method with segmentation constraints, and utilizing a micro-displacement segmentation network and semantic segmentation-guided matching, the problem of low depth estimation accuracy in complex backgrounds of traditional binocular vision is solved, and high-precision three-dimensional vibration measurement of rotors is achieved.

CN121962223APending Publication Date: 2026-05-01KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2026-01-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional binocular vision has low depth estimation accuracy in complex backgrounds and weak texture scenes, and is highly dependent on depth annotation, making it difficult to achieve high-precision non-contact vibration measurement.

Method used

A high-precision binocular vision method based on segmentation constraints is adopted. By building a micro-displacement segmentation network model, semantic segmentation is fused to guide matching and dual-view triangulation. Segmentation information is used to constrain stereo matching, reducing dependence on external parameter matrices and reducing computational errors.

Benefits of technology

It achieves high-precision non-contact vibration measurement in complex scenarios, with a three-dimensional vibration measurement accuracy of 20.94 micrometers, providing a new approach and feasible path.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962223A_ABST
    Figure CN121962223A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraint, and the method comprises the steps: collecting a rotor vibration video at a first visual angle through a first collection device, and constructing a first vibration image data set; synchronously utilizing a second acquisition device to acquire a rotor vibration video under a second visual angle, and constructing a second vibration image data set; camera calibration is carried out on the binocular vision acquisition devices according to the calibration plate, and unit conversion coefficients under the first acquisition device and the second acquisition device are obtained; deducing the value of an included angle between two acquisition devices in the binocular vision acquisition device; extracting a first preset number of images from the first vibration image data set and the second vibration image data set, respectively marking the images, and dividing the marked images into a training set and a verification set; taking the training set and the verification set as the input of an infinitesimal displacement segmentation network model to obtain a trained infinitesimal displacement segmentation network model; inputting a continuous to-be-predicted rotor vibration image pair into the trained infinitesimal displacement segmentation network model for prediction, and obtaining a semantic segmentation mask of the image; and according to the semantic segmentation mask, using a fusion semantic segmentation guide matching and double-view triangulation method to extract a rotor vibration three-dimensional coordinate. According to the invention, the detection precision of tiny vibration displacement is high.
Need to check novelty before this filing date? Find Prior Art

Description

A High-Precision Binocular Vision Method and System for Rotor Three-Dimensional Vibration Measurement Based on Segmentation Constraints Technical Field

[0001] This invention relates to a high-precision binocular vision method and system for measuring three-dimensional vibration of a rotor based on segmentation constraints, belonging to the field of three-dimensional vibration measurement using computer vision. Background Technology

[0002] With the rapid development of modern industrial technology, higher requirements are placed on the stability, safety, and accuracy of mechanical systems, engineering structures, and precision manufacturing equipment. Vibration, as a key signal reflecting the dynamic characteristics and potential faults of a system, requires precise measurement to ensure the efficient and reliable operation of industrial systems. Vibration measurement, as a core means of structural dynamics and health monitoring, can be broadly classified into two categories: contact measurement and non-contact measurement. Contact measurement methods have advantages such as large measurement range, simple structure, and stable performance; however, this method requires direct contact with the structure, which may introduce additional mass effects into the vibration measurement. With the rapid development of digital image processing technology and high-speed industrial cameras, computer vision-based vibration displacement measurement technology, as a novel non-contact displacement measurement method, can overcome the influence of contact measurement on vibration measurement and has advantages such as high detection accuracy, low cost, and ease of setup.

[0003] Visual 3D vibration measurement is mainly divided into monocular depth estimation and binocular depth estimation. Monocular cameras use only a 2D image captured by one camera to estimate the depth distance from each pixel to the camera. However, this method still has limitations in vibration measurement in complex scenes, with significant calibration and horizontal vibration errors. Traditional techniques typically use binocular vision to estimate depth from images, i.e., using two cameras to achieve depth estimation through stereo matching and triangulation. Although machine learning-based binocular vision can achieve high measurement accuracy, traditional stereo matching methods have high requirements for detecting the surface texture of the target. In complex backgrounds and scenes with weak textures, matching errors are easily caused, leading to a decrease in depth estimation accuracy.

[0004] Therefore, how to ensure high accuracy while eliminating the dependence on depth annotation remains a problem that urgently needs to be solved. Summary of the Invention

[0005] This invention provides a high-precision binocular vision method and system for measuring the three-dimensional vibration of a rotor based on segmentation constraints. On the one hand, it proposes to build a micro-displacement segmentation network model. On the other hand, it introduces a fusion semantic segmentation-guided matching and dual-view triangulation method without relying on depth annotation. The segmentation information is used to constrain the traditional binocular matching to achieve high-precision three-dimensional matching of the rotor. At the same time, the depth displacement solution method in the fusion semantic segmentation-guided matching and dual-view triangulation methods reduces the dependence on the external parameter matrix in the calculation process and reduces the cumulative calculation error caused by multiple transformations.

[0006] The technical solution of this invention is:

[0007] According to a first aspect of the present invention, a high-precision binocular vision method for measuring three-dimensional vibration of a rotor based on segmentation constraints is proposed, comprising:

[0008] Step 1: Set up the first acquisition device and the second acquisition device as binocular vision acquisition devices. Use the first acquisition device to acquire rotor vibration video from a first perspective to construct a first vibration image dataset; simultaneously use the second acquisition device to acquire rotor vibration video from a second perspective to construct a second vibration image dataset.

[0009] Step 2: Perform camera calibration on the binocular vision acquisition device according to the calibration plate to obtain the unit conversion coefficients under the first and second acquisition devices; and derive the angle value between the two acquisition devices in the binocular vision acquisition device.

[0010] Step 3: Extract a first preset number of images from the first vibration image dataset and the second vibration image dataset, label them respectively, and divide the labeled images into a training set and a validation set;

[0011] Step 4: Using DDRNet as the framework, introduce a phase branch to capture frequency domain phase features and orientation texture features, thereby building a micro-displacement segmentation network model; use the training set and validation set as input to the micro-displacement segmentation network model to obtain a trained micro-displacement segmentation network model.

[0012] Step 5: Predict the continuous rotor vibration images to be predicted by inputting them into the trained micro-displacement segmentation network model to obtain the semantic segmentation mask of the images; wherein, the rotor vibration image pair includes the rotor vibration image from the first viewpoint and the rotor vibration image from the second viewpoint.

[0013] Step 6: Based on the semantic segmentation mask, extract the three-dimensional coordinates of rotor vibration using the fusion semantic segmentation-guided matching and dual-view triangulation method.

[0014] Furthermore, the first viewing angle is directly in front of the rotor's circumferential surface, and the optical path angle between the first and second acquisition devices and the rotor is θ°.

[0015] Furthermore, the micro-displacement segmentation network model first performs preliminary feature extraction on the rotor vibration image based on the Stem layer. Then, the extracted preliminary features are processed in parallel using an improved high-resolution branch, an improved low-resolution semantic branch, and a phase branch at different scales and emphases. The outputs of the three parallel branches are fused using a third phase feature fusion module to obtain fused features. The fused features are used as input to the segmentation head to output a semantic segmentation mask, which is then used as the output of the micro-displacement segmentation network model. The improved high-resolution branch emphasizes detail preservation, the improved low-resolution semantic branch focuses on global contextual information, and the phase branch utilizes frequency domain phase features and orientation texture features to enhance sensitivity to micro-vibrations.

[0016] Furthermore, the phase branch includes a first, second, and third frequency domain feature extraction and enhancement module and a first adaptive feature fusion module connected in sequence;

[0017] The improved low-resolution semantic branch includes a first residual block, a second residual block, a residual bottleneck block, and a second adaptive feature fusion module connected in sequence.

[0018] The improved high-resolution branch includes three windmill-filled convolutional modules that fuse Gaussian attention mechanisms. A first phase feature fusion module is provided between the first and second windmill-filled convolutional modules that fuse Gaussian attention mechanisms, and a second phase feature fusion module is provided between the second and third windmill-filled convolutional modules that fuse Gaussian attention mechanisms. The output of the first windmill-filled convolutional module that fuses Gaussian attention mechanisms is downsampled and added to the output of the first residual block in the improved low-resolution semantic branch after downsampling, serving as the input to the second residual block in the improved low-resolution semantic branch. The second windmill-filled convolutional module that fuses Gaussian attention mechanisms... The output of the convolution module is downsampled and then added to the output of the second residual block in the improved low-resolution semantic branch after downsampling, serving as the input to the residual bottleneck block in the improved low-resolution semantic branch. The first phase feature fusion module is used to fuse the output of the first windmill-shaped filling convolution module with the first fusion Gaussian attention mechanism, the output of the first frequency domain feature extraction and enhancement module in the phase branch, and the upsampled output of the first residual block. The second phase feature fusion module is used to fuse the output of the second windmill-shaped filling convolution module with the second fusion Gaussian attention mechanism, the output of the second frequency domain feature extraction and enhancement module in the phase branch, and the upsampled output of the second residual block.

[0019] The outputs of the first adaptive feature fusion module, the upsampled output of the second adaptive feature fusion module, and the output of the third windmill-filled convolution module with Gaussian attention mechanism are used as inputs to the third phase feature fusion module to obtain fused features.

[0020] Furthermore, the first, second, and third frequency domain feature extraction and enhancement modules have the same structure, including: spatial phase operation, frequency domain phase operation, learnable Gabor directional filtering, and cross-channel phase interaction attention;

[0021] The spatial phase operation and frequency domain phase operation are used to extract features from the spatial and frequency domains; then, based on the extracted spatial and frequency domain features, spatial and frequency domain fusion features are obtained.

[0022] The learnable Gabor directional filter is used to obtain directional features;

[0023] The cross-channel phase interaction attention is used to dynamically fuse spatial frequency domain fusion features and directional features through a cross attention mechanism, and then obtain phase enhancement features through pyramid pooling and SqueezeExcite channel attention mechanism.

[0024] Furthermore, the first, second, and third windmill-shaped padded convolutional modules that integrate Gaussian attention mechanisms have the same structure, which are used to divide the input features into two parts. The first part serves as the input of the windmill-shaped asymmetric padded convolution, and the output of the windmill-shaped asymmetric padded convolution, together with the output of the first and second parts, serves as the input of the Gaussian-based spatial-channel joint attention.

[0025] Furthermore, the first and second adaptive feature fusion modules have the same structure, based on the DAPPM module, and introduce the SqueezeExcite channel attention mechanism.

[0026] Furthermore, the method for fusing semantic segmentation-guided matching and dual-view triangulation includes:

[0027] The current rotor vibration image to be predicted is divided into m mask image blocks from top to bottom according to the corresponding semantic segmentation mask;

[0028] Take the center point of each masked image block as the matching point; construct a matching pair by matching the matching point of the rotor vibration image in the first view and the matching point of the corresponding number in the rotor vibration image in the second view.

[0029] Based on the unit conversion coefficient under the first acquisition device, the horizontal and vertical displacements of the rotor in the first viewpoint of the rotor vibration image to be predicted are extracted and used as the position coordinates of the rotor in the real coordinate axes x and y of the current frame.

[0030] Based on the difference in the rotor position coordinates in the x-axis direction between the matching pairs of two frames of rotor vibration images from the first and second perspectives, and the included angle between the two acquisition devices in the binocular vision acquisition device, the actual displacement change of the rotor in the z-axis direction in two adjacent frames is obtained; based on the actual displacement change, the position coordinates of the rotor in the z-axis direction of the current frame are obtained.

[0031] Furthermore, the expression for the actual displacement change of the rotor in the z-axis direction is:

[0032] ;

[0033] in, This is the difference in the rotor's position coordinates along the x-axis between the i-th mask image block in the semantic segmentation mask of the rotor vibration image from the first-view perspective in the previous frame and the i-th mask image block in the semantic segmentation mask of the rotor vibration image from the first-view perspective in the current frame. The difference in the rotor position coordinates in the x-axis direction between the i-th mask image block in the semantic segmentation mask of the rotor vibration image in the second view of the previous frame and the i-th mask image block in the semantic segmentation mask of the rotor vibration image in the second view of the current frame. It is the angle between the imaging angle of the second acquisition device and the rotor shaft.

[0034] According to a second aspect of the present invention, a high-precision binocular vision rotor three-dimensional vibration measurement system based on segmentation constraints is provided, comprising a module of the high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints described in any one of the preceding claims.

[0035] The beneficial effects of this invention are as follows: This invention takes a rotating rotor as the research object and uses two high-speed cameras as the acquisition medium. Based on traditional binocular measurement methods, a high-precision binocular matching framework based on segmentation constraints is proposed. The mask is segmented using a segmentation network model of minute vibration displacements. Then, a method fusing semantic segmentation-guided matching and dual-view triangulation is constructed. This method relies on the segmented mask as prior information to hard-constrain the stereo matching process and obtains the rotor's three-dimensional vibration displacement through a displacement solution method based on fixed-angle imaging geometry. Experimental results show that this method achieves a three-dimensional vibration measurement accuracy of 20.94 micrometers in typical vibration tests, providing a new approach and feasible path for non-contact high-precision vibration measurement in complex scenarios. Attached Figure Description

[0036] Figure 1 is a flowchart of the method of the present invention.

[0037] Figure 2 shows the data collected at the experimental site.

[0038] Figure 3 is a schematic diagram of the placement of the calibration plate during camera calibration.

[0039] Figure 4 shows the framework diagram of the micro-displacement sensing segmentation network model.

[0040] Figure 5 shows the structure of the GPAConv module.

[0041] Figure 6 is a schematic diagram of semantic segmentation-guided matching.

[0042] Figure 7 is a schematic diagram of direct world coordinate triangulation under a fixed orthogonal perspective.

[0043] Figure 8 is a schematic diagram of the principle of extracting depth information.

[0044] Figure 9 is a visual comparison of the detection effects of different networks.

[0045] Figure 10 is a comparison of rotor vibration displacement curves extracted by different networks.

[0046] Figure 11 is a comparison of the shaft center trajectory extracted by different networks for rotor vibration.

[0047] Figure 12 is a comparison chart of different depth estimation methods. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.

[0049] Example 1: As shown in Figures 1-12, according to a first aspect of the present invention, a high-precision binocular vision method for measuring three-dimensional vibration of a rotor based on segmentation constraints is proposed, comprising:

[0050] Step 1: Set up the first acquisition device and the second acquisition device as binocular vision acquisition devices. Use the first acquisition device to acquire rotor vibration video from a first perspective to construct a first vibration image dataset; simultaneously use the second acquisition device to acquire rotor vibration video from a second perspective to construct a second vibration image dataset; wherein, the first perspective is directly in front of the rotor's circumferential surface, and the optical path angle between the first and second acquisition devices and the rotor is θ° (0 < θ° < 90°); at the same time, use an eddy current sensor to acquire vibration displacement data of the rotating body in three dimensions;

[0051] For example: As shown in Figures 2 and 7, the data acquisition platform is built using two Thousand Eyes M220 high-speed industrial cameras as acquisition devices to capture rotor vibration videos (in Figure 2, the Left Camera is the first acquisition device, i.e., the first industrial camera (left camera), and the Right Camera is the second acquisition device, i.e., the second industrial camera (right camera)). The two cameras were set to a sampling frequency of 2000 Hz and a resolution of 1080×1920. The first industrial camera was positioned directly in front of the rotor's circumference (at which angle the rotor's vertical and horizontal vibrations could be clearly distinguished). The optical path angle between the first and second industrial cameras and the rotor was θ°. Simultaneously, the high-speed industrial camera was set to be at the same horizontal height as the rotor, ensuring the rotor was visible in all frames of the vibration video captured by the high-speed industrial camera. The rotor test bench was a Nanjing Dongda Z-30 rotor test bench. Eddy current sensors were arranged orthogonally in space on the rear, top, and right sides of the rotor, with a sampling range of 2000 μm, a signal input mode of SIN-DC, and a sensitivity of 5 μm / mV, used to extract the standard vibration displacement signals of the rotor in three orthogonal directions. The object of the sampling was the rotor (parameters: radius 32.75 mm, mass 590 g). A light source (Jinbei EF-200LED) was used for illumination compensation of the object being sampled. The acquisition time of the two cameras and three eddy current sensors mentioned above is set to 3 seconds (the vibration video acquired by the first industrial camera and the second industrial camera is 6000 consecutive frames).

[0052] Step 2: Perform camera calibration on the binocular vision acquisition device based on the calibration plate to obtain the unit conversion coefficients under the first and second acquisition devices; and derive the angle value between the two acquisition devices in the binocular vision acquisition device.

[0053] For example, referring to Figure 3: During calibration, a calibration plate is attached to the rotating shaft of the rotor test bench. Calibration yields the intrinsic parameter matrices of the two cameras and the extrinsic parameters (rotation matrix R and translation vector T) of the cameras relative to the world coordinate system. Based on the intrinsic parameter matrices of the two cameras and the extrinsic parameters of the cameras relative to the world coordinate system, the unit transformation coefficient from each pixel point of each camera to the real world coordinate system is calculated, thereby achieving a direct mapping from the two-dimensional pixel coordinate system to the three-dimensional world space coordinate system. Using the rotation matrix in the camera extrinsic parameters, the angular relationship between the two cameras in the binocular vision acquisition device can be derived. This lays the foundation for completing triangulation calculations in a unified world coordinate system, thereby obtaining the depth information of the target.

[0054] Step 3: Extract a first preset number of images from the first vibration image dataset and the second vibration image dataset, label them respectively, and divide the labeled images into training set and validation set.

[0055] For example, 300 images each from the left and right cameras are extracted from the videos captured from the first and second perspectives, totaling 600 images. Labels are used to manually annotate the circumferential surface of the rotor at different angles in the images (refer to Figure 6; the magenta area in the left image represents the circumferential surface of the rotor annotated from the first perspective, and the magenta area in the right image represents the circumferential surface of the rotor annotated from the second perspective). The annotated images are then combined with the corresponding original images to form an annotated dataset. The annotated images in the dataset are divided in a 9:1 ratio, with 540 images assigned to the training set and 60 images assigned to the validation set.

[0056] Step 4: Using DDRNet as the framework, introduce a phase branch to capture frequency domain phase features and orientation texture features, thereby building a micro-displacement segmentation network model; use the training set and validation set as input to the micro-displacement segmentation network model to obtain a trained micro-displacement segmentation network model.

[0057] As shown in Figure 4, the micro-displacement segmentation network model first extracts preliminary features from the rotor vibration image using the Stem layer. Then, it processes the extracted preliminary features in parallel using an improved high-resolution branch, an improved low-resolution semantic branch, and a phase branch at different scales and emphases. The outputs of the three parallel branches are fused using a third phase feature fusion module to obtain fused features. The fused features are used as input to the segmentation head to output a semantic segmentation mask, which is then used as the output of the micro-displacement segmentation network model. The improved high-resolution branch emphasizes detail preservation, the improved low-resolution semantic branch focuses on global contextual information, and the phase branch utilizes frequency domain phase features and orientation texture features to enhance sensitivity to micro-vibrations.

[0058] The phase branch includes a first, second, and third frequency domain feature extraction and enhancement module and a first adaptive feature fusion module connected in sequence.

[0059] The improved low-resolution semantic branch includes a first residual block, a second residual block, a residual bottleneck block, and a second adaptive feature fusion module connected in sequence.

[0060] The improved high-resolution branch includes three windmill-filled convolutional modules that fuse Gaussian attention mechanisms. A first phase feature fusion module is provided between the first and second windmill-filled convolutional modules that fuse Gaussian attention mechanisms, and a second phase feature fusion module is provided between the second and third windmill-filled convolutional modules that fuse Gaussian attention mechanisms. The output of the first windmill-filled convolutional module that fuses Gaussian attention mechanisms is downsampled and added to the output of the first residual block in the improved low-resolution semantic branch after downsampling, serving as the input to the second residual block in the improved low-resolution semantic branch. The second windmill-filled convolutional module that fuses Gaussian attention mechanisms... The output of the convolution module is downsampled and then added to the output of the second residual block in the improved low-resolution semantic branch after downsampling, serving as the input to the residual bottleneck block in the improved low-resolution semantic branch. The first phase feature fusion module is used to fuse the output of the first windmill-shaped filling convolution module with the first fusion Gaussian attention mechanism, the output of the first frequency domain feature extraction and enhancement module in the phase branch, and the upsampled output of the first residual block. The second phase feature fusion module is used to fuse the output of the second windmill-shaped filling convolution module with the second fusion Gaussian attention mechanism, the output of the second frequency domain feature extraction and enhancement module in the phase branch, and the upsampled output of the second residual block.

[0061] The outputs of the first adaptive feature fusion module, the upsampled output of the second adaptive feature fusion module, and the output of the third windmill-filled convolution module with Gaussian attention mechanism are used as inputs to the third phase feature fusion module to obtain fused features.

[0062] The first, second, and third frequency domain feature extraction and enhancement modules (PFA) have the same structure, including: spatial phase operation, frequency domain phase operation, learnable Gabor directional filtering, and cross-channel phase interaction attention.

[0063] The spatial phase operation and frequency domain phase operation are used to extract features from the spatial and frequency domains (using standard convolutional blocks to extract spatial features from the spatial domain); extracting phase features from the frequency domain specifically involves: processing the input features... The complex spectrum is obtained by two-dimensional Fourier transform. The phase component of the input feature is:

[0064] ;

[0065] Subsequently, the phase components are mapped to cos(P) and sin(P) to avoid the problem of periodic discontinuities in the phase. Then, sub-pixel-level displacement information is encoded in the convolutional space through concatenation and convolution operations to obtain spatial-frequency domain fusion features. In the spatial domain, this displacement information can keenly capture edge changes, ensuring that rotor boundary information is not weakened in feature extraction. From a frequency domain perspective, explicit modeling of phase information enhances the network's ability to fit minute vibration signals.

[0066] The learnable Gabor directional filter is used to obtain directional features.

[0067] Specifically, to model the orientation sensitivity of the rotor boundary, a learnable Gabor convolution kernel is introduced using a learnable Gabor orientation filter. The expression is:

[0068] ;

[0069] in, , , , For learnable parameters, The coordinates of the input feature, ( , The coordinates are those after rotation in direction Ψ, and are calculated as follows:

[0070] ;

[0071] This design enables the model to respond to edge textures in multiple directions, giving the phase branch stronger anisotropic representation capabilities. This learnable Gabor convolution combined with cross-channel phase interactive attention significantly improves the model's ability to capture orientation-sensitive features, thus balancing local detail preservation with global discriminative power.

[0072] The cross-channel phase interaction attention is used to fuse spatial frequency domain features. With directional features Phase-enhanced features are obtained through dynamic fusion using a cross-attention mechanism, followed by pyramid pooling and a SqueezeExcite channel attention mechanism.

[0073] ;

[0074] in, , Here, ⊙ represents the attention weights, and ⊙ denotes element-wise product. This approach allows the network to retain necessary details while suppressing redundant information, and enables deep integration with other branches in the subsequent PFF-Module.

[0075] Considering that the rotor vibrates in all three dimensions, it exhibits irregular elliptical motion in the two-dimensional world captured by the camera. Conventional convolution cannot effectively identify its irregular motion in all directions. Therefore, this study improves a novel windmill-shaped padding convolution module (APC) that integrates a Gaussian attention mechanism on a high-resolution backbone.

[0076] The first, second, and third windmill-filled convolutional modules (APCs) that integrate Gaussian attention mechanisms have the same structure, as shown in Figure 4. The input features are divided into two parts. The first part serves as the input to the windmill-shaped asymmetric-filled convolution. The output of the windmill-shaped asymmetric-filled convolution, combined with the output of the first and second parts, serves as the input to the Gaussian-based spatial-channel joint attention. As can be seen from the above, the design of APC mainly includes two ideas: the windmill-shaped asymmetric-filled convolution (GPAconv) and Gaussian-based spatial-channel joint attention (Gaussian modulation). As shown in Figure 5, the windmill-shaped asymmetric-filled convolution calculates responses with different emphases in parallel using four sets of convolutions with different directional fills, and then fuses them through 1×1 projection, thereby comprehensively improving the ability to identify horizontal and vertical displacements. To enhance global and local weighting, the windmill-shaped asymmetric-filled convolution (GPAconv) introduces Gaussian-based spatial-channel joint attention after concatenation. This Gaussian-based spatial-channel joint attention first obtains channel statistics through parallel global average pooling and max pooling, and then performs dimensionality reshaping through a fully connected layer to generate channel attention. Simultaneously, spatial basic attention is generated through two-channel (average pooling + max pooling) convolution. Furthermore, Gaussian-based spatial-channel joint attention also applies a learnable Gaussian modulation factor in space:

[0077] ;

[0078] in, , , These are learnable parameters that allow spatial attention to be biased toward the local regions of interest in the image; It is a very small constant (taken as 10 in the embodiments of the present invention). -5 The final attention map is composed using pointwise multiplication:

[0079] ;

[0080] The above formula enhances the ability to respond to uneven motion and changes in direction.

[0081] The first, second, and third phase feature fusion modules have the same PFF structure, including splicing operation, adaptive attention, and SqueezeExcite channel attention mechanism.

[0082] To enable the newly added phase branch to complement the high-resolution and semantic branches, this invention designs the PFF-Module (Phase-Feature Fusion Module). This module can perform deep feature concatenation and weight fusion between different branches, allowing the network to effectively preserve the fine texture and edge features provided by the phase branch while integrating semantic context information. Through this interaction of multi-source features, the overall model significantly improves both the edge segmentation accuracy of the rotor target and the sensitivity to sub-pixel vibrations. Simultaneously, to adapt to the backbone network changing from a two-branch to a three-branch system, the corresponding segmentation head and loss function have also been adaptively modified. Specifically, a CAFAFE attention fusion module is added to integrate the outputs of the three branches, enabling better utilization of phase information.

[0083] The first and second adaptive feature fusion modules (APFM) have the same structure. The first and second adaptive feature fusion modules are based on the DAPPM module and introduce the SE channel attention mechanism (i.e., the SqueezeExcite channel attention mechanism).

[0084] As can be seen from the above technical solution, the original DAPPM module is replaced by APFM (Adaptive PyramidFusion Module) in this invention. APFM introduces the SE channel attention mechanism on the multi-scale pooling convolution results to improve the stable fusion of cross-scale information and enhance semantic discrimination ability.

[0085] In summary, the newly added phase branch not only compensates for the shortcomings of traditional dual-branch networks in preserving details, but also provides structured support for high-precision analysis of rotor vibration images through the introduction of PFA-Block. Its synergistic effect with the PFF-Module lays a solid foundation for subsequent multi-scale fusion and final mask prediction.

[0086] For example, the training process specifically involves setting the parameters of the micro-displacement segmentation network model before training begins. In this experiment, the training set for all network models consisted of 540 images (270 left images and 270 right images), and the validation set consisted of 60 images (30 left images and 30 right images). The input image size was 1080×1920, the batch size was 2, and the training iterations were 120,000. The environment was a desktop computer (13th Gen Intel(R) Core(TM) i7-13700KF, 32GB RAM, NVIDIA GeForce RTX 4080 16GB GPU), and the algorithm ran on Windows 11, CUDA 12.4, PyTorch 2.4.1, and torchvision 0.19.1. After completing the parameter settings, the network model is used to train and validate the training and validation sets. After completing the set number of training iterations, the optimal weights for this training are obtained. The optimal weights are then loaded into the micro-displacement segmentation network model to obtain the trained micro-displacement segmentation network model.

[0087] Step 5: Predict the continuous rotor vibration images to be predicted using the pre-trained micro-displacement segmentation network model to obtain the semantic segmentation mask of the images; wherein, the rotor vibration image pair includes rotor vibration images from a first viewpoint and rotor vibration images from a second viewpoint. The rotor vibration images to be predicted are a continuous image sequence; for example, the first vibration image dataset and the second vibration image dataset obtained in Step 1 are used to construct a continuous rotor vibration image pair to be predicted; or, the continuous rotor vibration image pair to be predicted can also be constructed based on the rotor vibration images acquired in real time by the first and second acquisition devices.

[0088] Step 6: Use the fusion semantic segmentation-guided matching and dual-view triangulation method to extract the three-dimensional coordinates of the rotor vibration in the semantic segmentation mask.

[0089] The fusion semantic segmentation-guided matching and dual-view triangulation method includes:

[0090] The current rotor vibration image to be predicted is divided into m mask image blocks from top to bottom according to the corresponding semantic segmentation mask;

[0091] Take the center point of each masked image block as the matching point; construct a matching pair by matching the matching point of the rotor vibration image in the first view and the matching point of the corresponding number in the rotor vibration image in the second view.

[0092] Based on the unit conversion factor under the first acquisition device, the horizontal and vertical displacements of the rotor in the first viewpoint of the rotor vibration image to be predicted are extracted and used as the position coordinates of the rotor in the current frame in the true coordinate axes x and y. The expression is:

[0093]

[0094]

[0095] Where X and Y are the rotor's position coordinates along the actual coordinate axes x and y, respectively. , , , are the x and y coordinates of the center of the i-th mask image block in the semantic segmentation mask, respectively; m is the total number of generated mask image blocks; and K is the unit conversion coefficient of the first acquisition device.

[0096] When the frame number of the current frame is greater than 1, based on the difference in the rotor's position coordinates in the x-axis direction between two frames of rotor vibration images from the first and second perspectives, and the angle between the two acquisition devices in the binocular vision acquisition device, the actual displacement change of the rotor in the z-axis direction between two adjacent frames is obtained; based on the actual displacement change, the position coordinates of the rotor in the z-axis direction of the current frame are obtained; wherein, the initial frame is set as the first frame, and the actual displacement change of the rotor in the z-axis direction in the first frame is 0; for example, assuming that the present invention is used to acquire three-dimensional vibration displacement of three consecutive frames of rotor vibration images to be predicted, the calculated actual displacement change of the rotor in the z-axis direction between the second frame and the first frame is... The actual displacement change of the rotor in the z-axis direction between the third frame and the second frame is: Then, at the corresponding times of the three frames, the rotor's position coordinates in the z-direction of the true coordinate axis are 0, 0+ 0+ + .

[0097] Referring to Figure 5, the expression for the actual displacement change of the rotor in the z-axis direction is:

[0098] ;

[0099] in, This is the difference in the rotor's position coordinates along the x-axis between the i-th mask image block in the semantic segmentation mask of the rotor vibration image from the first-view perspective in the previous frame and the i-th mask image block in the semantic segmentation mask of the rotor vibration image from the first-view perspective in the current frame. The difference in the rotor position coordinates in the x-axis direction between the i-th mask image block in the semantic segmentation mask of the rotor vibration image in the second view of the previous frame and the i-th mask image block in the semantic segmentation mask of the rotor vibration image in the second view of the current frame. It is the angle between the imaging angle of the second acquisition device and the rotor shaft, that is... In Figure 8, , That is, a certain matching point pair , , .

[0100] In the above, semantic segmentation mask results are introduced as prior information before the triangulation method is introduced to obtain the true displacement change, providing hard constraints and guidance for the pixel matching process. Simultaneously, because the rotor's left-right vibration is extremely small and is the main cause of depth error, this invention optimizes the matching process: for each semantically segmented object mask block, a strategy of dividing the entire block into dozens of small blocks from top to bottom is adopted to reduce false detections in the left-right direction of some segmentation masks, thus reducing their impact on the overall matching. Then, the pixel coordinate system is converted into displacement changes in the world coordinate system using a unit transformation coefficient between the pixel coordinate system and the world coordinate system. Simultaneously, triangulation is used to calculate the rotor's pixel displacement changes between the current frame and the previous frame, thereby realizing the calculation of the rotor's three-dimensional vibration displacement in the world coordinate system.

[0101] Step 7: Based on the three-dimensional coordinates of the rotor vibration, regress the rotor vibration displacement curve.

[0102] Step 8: The experiment uses the global pixel classification accuracy (aAcc), average accuracy per class (mAcc), and mean intersection over union (mIoU) as metrics to evaluate the segmentation performance of the algorithms. The root mean square error (RMSE) is used to evaluate the accuracy of different algorithms in fitting the vibration curve. The formulas for calculating aAcc, mAcc, and mIoU are as follows:

[0103] ;

[0104] ;

[0105] ;

[0106] In the above formula, k is the total number of categories (in this invention, it is taken as 2, divided into rotor and background categories). For a true example, it represents the number of pixels that are actually of class i and are also predicted to be of class i. A false counterexample represents the number of pixels that are actually of class i but are predicted as other classes. A false positive is a pixel that is actually not of class i but is predicted to be of class i. The formula for calculating RMSE is as follows:

[0107] ;

[0108] Where M represents the number of sampling points, and y represents the vibration displacement collected by the eddy current sensor. This represents the vibration displacement generated by the algorithm.

[0109] Furthermore, to verify the feasibility of the algorithm, ablation experiments were conducted to compare the performance of different versions of the proposed model. The original network DDRNet, which incorporates algorithm improvements, uses an improved adaptive feature fusion module APFM in its backbone extraction network, adds a phase branch + PFF, and introduces an APC model. All of these models use the fusion semantic segmentation-guided matching and dual-view triangulation method of this invention for 3D coordinate extraction, and are compared using the RMSE metric. Table 1 shows the algorithm proposed in this invention, which incorporates all the above innovations. As shown in Table 1, the APFM module, when integrated alone, exhibits a clear functional orientation. While it doesn't alter the global modeling capability of the semantic branch, it improves regression accuracy by optimizing cross-scale feature weights to reduce the impact of background interference on displacement estimation, demonstrating its effectiveness in semantic branch optimization. The performance of the DDRNet+phase+PFF model, when the phase branch is integrated alone, degrades across the board. This is because the phase branch, when used alone, lacks APFM's branch fusion adaptive noise suppression and APC's spatial positioning guidance, causing ambient light interference in the original phase information to directly disrupt the feature distribution. However, when the phase branch collaborates with other modules, performance significantly improves—the RMSE of the DDRNet+APFM+phase+PFF model drops to 24.57 μm, and the DDRNet+APC+phase+PFF model achieves a maximum mIoU of 99.55% and a low RMSE of 22.49 μm. This precisely verifies the core idea in the method design: "Only through the synergy of the phase branch, APFM's noise suppression, and APC's directional response can vibration-related phase features be effectively extracted." It is worth noting that the rationality of the three-module collaboration depends on the fusion mechanism of the PFF-Module: the DDRNet+APFM+APC model suffers from significant performance degradation due to the lack of a phase branch (aAcc 98.82%, mAcc 99.39%, mIoU 99.11%, RMSE 88.46μm), indicating that the feature processing of APFM and APC is redundant when the three-branch feature fusion lacks a phase branch. The model proposed in this invention achieves "three-branch dimension alignment - dynamic weight allocation - residual fusion" through the PFF-Module, preserving the detail advantages of the high-resolution branch APC (maintaining segmentation performance of aAcc 99.90% and mAcc 99.76%) while integrating the semantic noise suppression of APFM and the sub-pixel sensitive features of the phase branch, ultimately achieving a minimum RMSE of 20.94μm, perfectly matching the design goal of "three-branch complementarity to solve the unified measurement of large displacement and sub-pixel displacement" in the method section. In other words, as can be seen from the above, the performance of using a single module has limitations, while the model of this invention performs optimally across all metrics.It can effectively improve the edge accuracy of rotor mask extraction and the sub-pixel sensitivity of vibration displacement estimation, providing more discriminative feature support for subsequent tasks.

[0110] Table 1 Comparison of ablation experiments (bold indicates optimal index)

[0111]

[0112] Furthermore, to verify the accuracy of rotor vibration displacement measurement of the proposed algorithm, its overall performance was compared with that of semantic segmentation networks such as KNet, DeepLabV3+, Segmenter, HRNet, SegFormer, and DDRNet. As shown in Table 2, the proposed algorithm is among the best in all metrics. In terms of segmentation performance, DDRNet and the proposed method are tied for the highest aAcc and mAcc, at 99.90% and 99.76%, respectively. Regarding mIoU, DDRNet slightly outperforms the proposed algorithm at 99.51% with 99.52%. The proposed method benefits from its three-branch backbone design, particularly the sensitivity of the phase branch to minute vibrations, achieving top-level segmentation performance similar to DDRNet while maintaining the accuracy of boundary details. On the core metric of vibration detection, RMSE, the proposed method significantly outperforms other models with an error of 20.94 μm, compared to KNet's 23.75 μm, SegFormer's 31.07 μm, DDRNet's 26.86 μm, HRNet's 34.52 μm, and DeepLabV3+'s 44.57 μm, demonstrating superior vibration quantification accuracy.

[0113] Table 2. Comparison of overall performance of different networks (bold indicates best performance indicators)

[0114]

[0115] As shown in Figure 9, the effectiveness of the proposed algorithm is first demonstrated by comparing the detection performance of different algorithm models. The proposed algorithm (ours) is compared experimentally with DeepLabV3+, SegFormer, Segmenter, DDRNet, HRNet, and KNet. Visual comparison of the detection results from different networks shows that all models can identify the target rotor, but there are differences in boundary refinement. The DeepLabV3+ model performs well in segmenting the corners of the rotor in the front view, but exhibits obvious unevenness at the left and right boundaries, resulting in a coarse overall segmentation. Conversely, in the side view, the segmentation of the left and right boundaries is relatively smooth, but performance in the edge and corner areas is poor. The Segmenter model has the worst segmentation performance, failing to accurately fit the rotor edges in both the front and side views. In contrast, the other models, including the network proposed in this invention, show little difference in visual effect and can all accurately fit the target boundaries. This indicates that the proposed algorithm can generate high-quality segmentation masks for rotor images from both perspectives, demonstrating good segmentation performance for the rotor.

[0116] The three-dimensional vibration displacement curve of the rotor obtained by the algorithm of this invention is compared with the rotor vibration displacement signal synchronously acquired by the eddy current sensor as the standard vibration displacement offset of the rotor. Comparisons are made between the rotor vibration displacement curves extracted by different algorithms in the X-axis, Y-axis, and depth (Z-axis) directions, and between the rotor block axis center trajectory extracted by different networks. Figure 10 shows a comparison between the rotor vibration displacement curves extracted by different algorithms and the standard signal acquired by the eddy current sensor. The black curve represents the vibration displacement curve (GT) acquired by the eddy current sensor, and the other colored curves represent the vibration displacement curves generated by different algorithms (red corresponds to the X-axis, blue to the Y-axis, and yellow to the Z-axis). Algorithms such as DDRNet, DeepLabv3+, HRNet, KNet, and SegFormer can all show a general oscillation trend, but their performance in terms of smooth transition and matching with the real curve is poor. In contrast, Ours' vibration displacement curves in the X, Y, and Depth directions all highly coincide with the real curves: the peak height, trough depth, and curve transition slope in the X direction, the oscillation period and amplitude changes in the Y direction, and the oscillation rhythm and local detail fluctuations in the Depth direction all accurately match the real curves. It exhibits excellent temporal continuity and morphological reproduction, fully demonstrating its ability to accurately capture the "global oscillation trend + local detail features" of the vibration displacement curve. This shows that the algorithm of this invention has high detection accuracy for minute vibration displacements.

[0117] To further illustrate the effectiveness and feasibility of the proposed method, we compared the extraction of rotor vibration axis center trajectories using different algorithms. As shown in Figure 11, the black curve represents the vibration displacement curve collected by the eddy current sensor (because the plotted data is a continuous data segment randomly selected from 6000 frames, the ground truth data for comparison differs between different networks), and the red curve represents the vibration displacement curve extracted by the algorithm. The network proposed in this study extracts a red axis center trajectory that highly overlaps with the real black trajectory in the 3D overall view, YZ plane projection, and XZ plane projection. Not only does it perfectly match the overall closed shape and the direction of the long axis extension, but the density of local trajectory points and the smoothness of inter-segment connections are also consistent with the real trajectory, demonstrating a more accurate learning ability for the spatial morphology and detailed texture of the axis center trajectory.

[0118] Furthermore, considering that comparing segmentation algorithms within the framework proposed in this invention cannot fully demonstrate the superiority of the proposed method, a comparison was made between the method of this invention and other binocular vision algorithms in the depth direction. The algorithms compared included deep learning-based binocular vision algorithms—DEFOM, IGEV++, and MonSter++, a deep learning-based monocular vision algorithm—DNet, and traditional binocular vision matching algorithms—SGBM and Graph Cut, as shown in Figure 12. As can be seen from Figure 12, except for this invention, the depth vibration curves obtained by other algorithms are basically unable to fit the standard ground truth (GT), thus demonstrating the superiority of the high-precision binocular matching framework proposed in this invention in vibration measurement.

[0119] According to a second aspect of the present invention, a high-precision binocular vision rotor three-dimensional vibration measurement system based on segmentation constraints is provided, comprising modules of the high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints described in any one of the preceding embodiments. Each module in the above-described high-precision binocular vision rotor three-dimensional vibration measurement system based on segmentation constraints can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0120] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A high-precision binocular vision method for measuring three-dimensional vibration of a rotor based on segmentation constraints, characterized in that, include: Step 1: Set up the first acquisition device and the second acquisition device as binocular vision acquisition devices. Use the first acquisition device to acquire rotor vibration video from the first perspective and construct the first vibration image dataset. The process involves: 1) Simultaneously acquiring rotor vibration video from a second perspective using a second acquisition device to construct a second vibration image dataset; 2) Calibrating the binocular vision acquisition device using a calibration board to obtain unit conversion coefficients for the first and second acquisition devices, and deriving the angle between the two acquisition devices in the binocular vision acquisition device; 3) Extracting a first preset number of images from the first and second vibration image datasets and labeling them, then dividing the labeled images into training and validation sets; 4) Using DDRNet as a framework, introducing a phase branch to capture frequency domain phase features and directional texture features, thereby constructing a micro-displacement segmentation network model; using the training and validation sets as inputs to the micro-displacement segmentation network model to obtain a trained micro-displacement segmentation network model; 5) Predicting the continuous rotor vibration images to be predicted using the input trained micro-displacement segmentation network model to obtain the semantic segmentation mask of the image; wherein, the rotor vibration image pair includes rotor vibration images from the first perspective and rotor vibration images from the second perspective; 6) Based on the semantic segmentation mask, using a fusion semantic segmentation guided matching and dual-view triangulation method to extract the three-dimensional coordinates of the rotor vibration.

2. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 1, characterized in that, The first viewing angle is directly in front of the rotor's circumferential surface, and the optical path angle between the first and second acquisition devices and the rotor is θ°.

3. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 1, characterized in that, The micro-displacement segmentation network model first extracts preliminary features from the rotor vibration image using the STEM layer. Then, it processes the extracted preliminary features in parallel using an improved high-resolution branch, an improved low-resolution semantic branch, and a phase branch at different scales and emphases. The outputs of the three parallel branches are fused using a third phase feature fusion module to obtain fused features. The fused features are used as input to the segmentation head to output a semantic segmentation mask, which is then used as the output of the micro-displacement segmentation network model. The improved high-resolution branch emphasizes detail preservation, the improved low-resolution semantic branch focuses on global contextual information, and the phase branch utilizes frequency domain phase features and orientation texture features to enhance sensitivity to micro-vibrations.

4. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 3, characterized in that, The phase branch includes a first, second, and third frequency domain feature extraction and enhancement module and a first adaptive feature fusion module connected in sequence. The improved low-resolution semantic branch includes a first residual block, a second residual block, a residual bottleneck block, and a second adaptive feature fusion module connected in sequence. The improved high-resolution branch includes three windmill-shaped filled convolutional modules that fuse Gaussian attention mechanisms. A first phase feature fusion module is provided between the first and second windmill-shaped filled convolutional modules that fuse Gaussian attention mechanisms, and a second phase feature fusion module is provided between the second and third windmill-shaped filled convolutional modules that fuse Gaussian attention mechanisms. The output of the first windmill-shaped filled convolutional module that fuses Gaussian attention mechanisms is downsampled and added to the output of the first residual block in the improved low-resolution semantic branch after downsampling, serving as the input of the second residual block in the improved low-resolution semantic branch. The output of the second windmill-shaped filled convolutional module that fuses Gaussian attention mechanisms is downsampled and added to the output of the second residual block in the improved low-resolution semantic branch after downsampling, serving as the input of the residual bottleneck block in the improved low-resolution semantic branch. The first phase feature fusion module is used to fuse the output of the first fused Gaussian attention mechanism windmill-filled convolution module, the output of the first frequency domain feature extraction and enhancement module in the phase branch, and the upsampled output of the first residual block. The second phase feature fusion module is used to fuse the output of the second fused Gaussian attention mechanism windmill-filled convolution module, the output of the second frequency domain feature extraction and enhancement module in the phase branch, and the upsampled output of the second residual block. The output of the first adaptive feature fusion module, the upsampled output of the second adaptive feature fusion module, and the output of the third fused Gaussian attention mechanism windmill-filled convolution module are used as inputs to the third phase feature fusion module to obtain fused features.

5. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 4, characterized in that, The first, second, and third frequency domain feature extraction and enhancement modules have the same structure, including: spatial phase operation, frequency domain phase operation, learnable Gabor directional filtering, and cross-channel phase interaction attention. The spatial phase operation and frequency domain phase operation are used to extract features from the spatial and frequency domains. Then, based on the extracted spatial and frequency domain features, spatial-frequency domain fusion features are obtained. The learnable Gabor directional filtering is used to obtain directional features. The cross-channel phase interaction attention is used to dynamically fuse the spatial-frequency domain fusion features and directional features through a cross-attention mechanism, and then obtain phase enhancement features through pyramid pooling and SqueezeExcite channel attention mechanisms.

6. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 4, characterized in that, The first, second, and third windmill-shaped padded convolutional modules that integrate Gaussian attention mechanisms have the same structure. They are used to divide the input features into two parts. The first part serves as the input of the windmill-shaped asymmetric padded convolution. The output of the windmill-shaped asymmetric padded convolution, together with the output of the first and second parts, serves as the input of Gaussian-based spatial-channel joint attention.

7. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 4, characterized in that, The first and second adaptive feature fusion modules have the same structure, based on the DAPPM module, and introduce the SqueezeExcite channel attention mechanism.

8. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 1, characterized in that, The fusion semantic segmentation-guided matching and dual-view triangulation method includes: dividing the semantic segmentation mask corresponding to the current rotor vibration image pair to be predicted into m mask image blocks from top to bottom; taking the center point of each mask image block as a matching point; constructing a matching pair by matching the matching point of the rotor vibration image in the first view and the matching point of the corresponding sequence number of the rotor vibration image in the second view; extracting the horizontal and vertical displacements of the rotor in the first view of the current rotor vibration image pair to be predicted according to the unit conversion coefficient under the first acquisition device, as the position coordinates of the rotor in the real coordinate axes x and y of the current frame; obtaining the real displacement change of the rotor in the z-axis direction between two adjacent frames of the rotor vibration image pair in the first and second view and the included angle between the two acquisition devices in the binocular vision acquisition device according to the difference of the rotor position coordinates in the x-axis direction between the matching pairs of the rotor vibration images in the first and second view and the actual displacement change of the rotor in the z-axis direction between the two frames of the rotor vibration image pair; and obtaining the position coordinates of the rotor in the real coordinate axis z of the current frame according to the real displacement change.

9. The high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints according to claim 8, characterized in that, The expression for the actual displacement change of the rotor in the z-axis direction is: ;in, This is the difference in the rotor's position coordinates along the x-axis between the i-th mask image block in the semantic segmentation mask of the rotor vibration image from the first-view perspective in the previous frame and the i-th mask image block in the semantic segmentation mask of the rotor vibration image from the first-view perspective in the current frame. The difference in the rotor position coordinates in the x-axis direction between the i-th mask image block in the semantic segmentation mask of the rotor vibration image in the second view of the previous frame and the i-th mask image block in the semantic segmentation mask of the rotor vibration image in the second view of the current frame. It is the angle between the imaging angle of the second acquisition device and the rotor shaft.

10. A high-precision binocular vision-based three-dimensional vibration measurement system for rotors based on segmentation constraints, characterized in that, The module includes the high-precision binocular vision rotor three-dimensional vibration measurement method based on segmentation constraints as described in any one of claims 1-9.