Target-free structure vibration signal acquisition method based on unmanned aerial vehicle platform
By adopting the target-free structural vibration signal acquisition method on the drone platform, using optical flow neural network and optical flow compensation algorithm for motion compensation and optical flow estimation, the scene limitation of the fixed camera method in the measurement of vibration signal of large infrastructure structures and the error problems caused by drone jitter are solved, and high-precision structural vibration signal acquisition and modal analysis are achieved.
Patent Information
- Application Number
- CN202510248153.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In the prior art, computer vision methods based on fixed cameras have scene limitations when measuring structure vibration signals of large infrastructure, especially when the test area is far away from the camera, it is difficult to extract high-precision vibration signals, and the jitter of the drone during hovering shooting will lead to huge errors or errors in the estimation of structural vibration signals.
The target-free structure vibration signal acquisition method based on the drone platform is adopted to collect and pre-process video data through the drone approaching the infrastructure structure, and compensate the drone motion using optical flow neural network and optical flow compensation algorithm to eliminate the large displacement of the whole pixel and small displacement noise of the sub-pixel, and then perform optical flow estimation and post-processing to obtain the structural vibration signal.
It realizes that without additional targets, the vibration information of multiple measurement points is extracted from the structural vibration video captured by the drone, and the modal information of the overall structure is obtained, which reduces the cost, overcomes the long-distance resolution limitation of the fixed camera, and can quickly process the target vibration image and output high-precision vibration information.
Smart Images

Figure CN120213192A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of infrastructure structural health monitoring, and in particular to a method for collecting non-target structural vibration signals based on an unmanned aerial vehicle (UAV) platform. Background Art
[0002] Structural vibration analysis is one of the important methods for structural health monitoring (SHM) and equipment fault diagnosis. Traditional vibration signal measurement methods include contact acceleration sensors and non-contact optical sensors, etc. The acceleration sensor needs to be attached to the structure to achieve vibration measurement, which is relatively expensive, cumbersome to operate, and prone to quality problems in the collected signals due to line failures. The most widely used non-contact optical sensor is the Doppler laser vibrometer, which requires a complex interference optical path auxiliary device, has too high a cost and a small measurement range, and it is difficult to achieve large-scale multi-point measurement of infrastructure structures.
[0003] In recent years, with the rise of computer technology, the structural vibration measurement method based on computer vision has attracted more and more attention from researchers in the field of structural health monitoring. The vision-based method can achieve structural vibration measurement without additional loads, over a large range, and at low cost.
[0004] In the existing patented technologies, for example, CN118706244A proposes a real-time monitoring technology and method for bridge vibration based on vision enhancement. It uses a fixed camera to capture real-time vibration images of the bridge structure, selects a specific area for image cropping, and performs lossless and lossy data compression processing. The template matching combined with the LK sub-pixel optical flow technology is used to extract the vibration signals in the lossless compressed image, and the fundamental frequency of the bridge is determined through frequency domain analysis. The lossy compressed image is subjected to video synthesis processing, and the multi-layer filtering technology is used to visually display the bridge structure vibration and amplify the micro-movement, effectively improving the accuracy of vibration signal extraction and realizing the visual amplification of vibration through multi-layer filtering.
[0005] However, the computer vision method based on the fixed camera method has certain application scenario limitations. In actual engineering applications, if the test area is far from the camera and exceeds the resolution ability of the camera, it is difficult to extract high-precision vibration signals. By using a movable acquisition platform such as a UAV for close-range shooting, the shooting distance problem in the signal acquisition of large real structures can be effectively solved. However, when the UAV hovers for shooting, it will generate jitter, causing camera movement, which will bring huge errors or mistakes to the estimation of structural vibration signals.
[0006] Therefore, a method for collecting non-target structural vibration signals based on a UAV platform is provided to solve the above problems. Summary of the Invention
[0007] The object of the present invention is to provide a method for collecting vibration signals of a targetless structure based on a drone platform, which non-contact measures the vibration signals of infrastructure structures based on a drone camera platform, effectively solving the problem of scene limitation in the existing vibration signal measurement method using a fixed camera.
[0008] To achieve the above object, the present invention provides a method for collecting vibration signals of a targetless structure based on a drone platform, including the following steps:
[0009] S1: Fly a drone close to the infrastructure structure to collect and preprocess video data, obtaining a video image v1 of the vibration process and a video image v2 of irregular motion.
[0010] S2: Compensate for the drone motion based on the video image v2 to eliminate the noise of large displacement of whole pixels, obtaining a processed video image v3.
[0011] S3: Obtain a video image v4 without the structure to be measured based on the video image v3, compensate for the drone motion based on the video image v4 to eliminate the noise of small displacement of sub-pixels, obtaining a video image v5 after precise motion compensation.
[0012] S4: Perform optical flow estimation and post-processing based on the video image v5.
[0013] Preferably, step S1 specifically includes the following steps:
[0014] S11: Fly a drone to take a video of the area to be measured, obtaining a video image v1 of the vibration process.
[0015] S12: Obtain the static background of the video image v1. When the drone hovers in the static background, the camera generates irregular motion, obtaining a video image v2 of the corresponding irregular motion of the static background.
[0016] S13: Input the video image v2 into the optical flow neural network FlowFormer frame by frame as a time series.
[0017] Preferably, step S2 specifically includes the following steps:
[0018] S21: Train the optical flow neural network FlowFormer with the public dataset FlyingChairs, and calculate the optical flow between frames of the image through the optical flow neural network FlowFormer.
[0019] S22: Calculate the target whole-pixel displacement data through the optical flow data.
[0020] S23: Based on the integer pixel displacement data, perform heuristic image translation processing on each frame of the video image v1 through a heuristic algorithm, limit the noise caused by camera movement within a unit pixel, and generate the processed video image v3.
[0021] Preferably, in step S21, the input signal of the optical flow neural network FlowFormer is an image pair composed of a video reference frame and a current frame, and the output signal of the optical flow neural network FlowFormer is the full-field optical flow data corresponding to the image pair.
[0022] Preferably, step S21 specifically includes the following steps:
[0023] Step 1: Input an image of H I ×W I ×3RGB, extract a feature map of H×W×D f through the backbone vision network, extract the source image of the previous frame and the target image of the current frame based on the feature map, calculate the dot product similarity corresponding to all pixel points in the source image and the target image, and construct a 4D cost volume of H×W×H×W, where (H, W) = (H I / 8, W I / 8);
[0024] Step 2: Identify the position corresponding to the source pixel in the target image through the source-target visual similarity encoded in the 4D cost volume, and eliminate the repeated patterns and non-distinguishing regions in the two frames of images through the cost encoder;
[0025] Step 3: Through the cost decoder, predict the optical flow residual Δf(x) in a recursive manner. The optical flow residual Δf(x) is expressed as:
[0026] Δf(x) = ConvGRU(Concat(c x , q x , t x , f(x)))
[0027] where t x represents the context feature of the source image, gradually optimizes the optical flow prediction in an iterative manner, and in the iterative process, regresses the optical flow residual Δf(x) through the gated recurrent unit ConvGRU.
[0028] Preferably, step 2 specifically includes the following steps:
[0029] Step 1: Block the cost map M x of each source pixel point x through strided convolution. The cost map M x is expressed as:
[0030]
[0031] Obtain multiple cost blocks q x Embedded cost graph M x , and convert multiple cost blocks q x Embedded cost graph M x into a feature map F through the activation function ReLU x , and the feature map F x is expressed as:
[0032]
[0033] wherein, the cost block feature F x in the feature map F x corresponds one-to-one with the 8×8 cost block q x in the cost graph;
[0034] Step 2: Randomly initialize the latent codeword C, and summarize each cost block feature F through K latent codewords x , query each cost block feature F x through the dot product attention mechanism, and summarize the cost graph into K latent vectors x with dimension D through the latent representation T Determine the embedding position of each cost block feature F x through the COTR method to obtain a 4D cost volume T, and the 4D cost volume T is expressed as:
[0035]
[0036] Update the latent codeword C through backpropagation and share the latent codeword C among different source pixel points x;
[0037] Step 3: Group the 4D cost volume T through the alternating grouped Transformer layer AGT, and generate corresponding cost queries Q x based on the grouping result through the feed-forward neural network FFN x , and the cost query Q
[0038] Q x = FFN(FFN(q x ) + PE(p))
[0039] wherein, q x represents a 9×9 cost block extracted from the cost graph M x centered at the position p;
[0040] Step 4: Predict the optical flow of the source pixel through the cost query, and calculate the corresponding position p of each source pixel point x in the target image, and the position p is expressed as:
[0041] p = x + f(x)
[0042] Among them, f(x) represents the current optical flow estimation;
[0043] Step Five: Generate the cost memory information c of the aggregated source pixel point x through the cross-attention mechanism x , the cost memory information c x is expressed as:
[0044] c x = Attention(Q x , K x , V x ).
[0045] Preferably, in Step Two, the latent representation T x is expressed as:
[0046] K x = Conv 1×1 (Concat(F x , PE))
[0047] V x = Conv 1×1 (Concat(F x , PE))
[0048] T x = Attention(C, K x , V x ).
[0049] Among them, the key K x and the value V x both represent the projection of the cost block feature F x , the key K x and the value V x are respectively expressed as:
[0050] K x = FFN(T x ).
[0051] V x = FFN(T x ),
[0052] Before projecting the cost block feature F x into the key K x and the value V x , the cost block feature F x and the position embedding sequence PE are concatenated, and the position embedding sequence PE is expressed as:
[0053]
[0054] Among them, D p represents the coding length.
[0055] Preferably, in step three, the AGT layer specifically includes the following two grouping methods:
[0056] Method 1: Combine the latent representations of each source pixel point x and set them as a group. Perform spatial-separated self-attention calculation on all latent representations within the group to obtain the updated latent representation T x . The updated latent representation T x is expressed as:
[0057]
[0058] where T x (i) represents the i-th latent representation of the source pixel point x. Process the latent representation T x after self-attention calculation through the feed-forward network FFN, and reorganize it into the 4D cost volume T to form the cost map M x intra-self-attention;
[0059] Method 2: Based on the latent representation T x , group all 4D cost volumes T according to, with each group containing H×W cost tokens T. Perform spatial-separated self-attention calculation on all cost tokens T within the group to obtain the updated cost token T i . The updated cost token T i is expressed as:
[0060] T i = FFN(SS-SelfAttention(T i ))), i = 1, 2,..., K.
[0061] Preferably, step S3 specifically includes the following steps:
[0062] S31: Based on the video image v3, obtain the background without the structure to be measured, and set the reference region video image without the structure to be measured as the video image v4. The coordinate positions of the video image v4 and the video image v2 in the video image v1 are the same;
[0063] S32: Combine the KLT optical flow compensation algorithm of the variational pyramid, extract the feature points in the video image v4 frame by frame, perform optical flow tracking on the feature points, and represent the sub-pixel transformation relationship between different frames caused by camera movement through the affine transformation matrix;
[0064] S33: Affinely transform each frame of the video image v4 through the affine transformation matrix to eliminate the sub-pixel small displacement noise in the video image v3, and obtain the video image v5 after precise motion compensation.
[0065] Preferably, step S4 specifically includes the following steps:
[0066] S41: Based on the video image v5, set the area to be measured as the ROI area, and extract the local phase information of the image in the ROI area through the complex steerable pyramid filter;
[0067] S42: Track the ROI area through the optical flow compensation algorithm KLT, calculate the displacement time history data at the sub-pixel level of the feature points in the area to be measured, multiply the pixel displacement time history data by the scale factor, and convert the pixel displacement time history data into the true displacement;
[0068] S43: Based on the true displacement data of multiple measurement points, obtain the frequency-domain power spectrum PSD through the fast Fourier transform FFT, and perform modal analysis on the data of the frequency-domain power spectrum PSD through the frequency decomposition method FDD to obtain the modal information of the structure to be measured.
[0069] Therefore, the present invention adopts the above-mentioned method for collecting structure vibration signals without a target based on a drone platform, and has the following beneficial effects:
[0070] (1) The present invention does not require additional targets, can extract the vibration information of multiple measurement points from the structure vibration video captured by the drone, and further obtain the overall modal information of the structure through the vibration information;
[0071] (2) The present invention effectively reduces costs, realizes high-precision shooting of local vibration information, overcomes the long-distance resolution limitation of fixed cameras at the same time, and can quickly process the target vibration image and output the vibration information.
[0072] Next, through the drawings and embodiments, the method scheme of the present invention will be further described in detail. Description of the Drawings
[0073] Figure 1 is the flowchart of a method for collecting structure vibration signals without a target based on a drone platform of the present invention;
[0074] Figure 2 is the architecture diagram of the optical flow neural network FlowFormer of the present invention. Detailed Embodiments
[0075] The following further illustrates the method scheme of the present invention through the drawings and embodiments.
[0076] Unless otherwise defined, the method terms or scientific terms used in this invention shall have the ordinary meanings as understood by those with ordinary skills in the field to which this invention pertains.
[0077] The words such as "including" or "comprising" used in this invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation to this invention. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly. In this invention, unless otherwise clearly specified and limited, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this invention can be understood according to specific circumstances.
[0078] Embodiment
[0079] As Figure 1 and Figure 2 shown, this invention provides a method for collecting vibration signals of a non-target structure based on a drone platform, including the following steps:
[0080] S1: Fly a drone close to the infrastructure structure to collect and preprocess video data, obtaining a video image v1 of the vibration process and a video image v2 of irregular motion;
[0081] Step S1 specifically includes the following steps:
[0082] S11: Use a drone to take a video of the area to be measured, obtaining a video image v1 of the vibration process;
[0083] S12: Obtain the static background of the video image v1. When the drone hovers in the static background, the camera generates irregular motion, obtaining a video image v2 of the irregular motion corresponding to the static background;
[0084] S13: Input the video image v2 into the optical flow neural network FlowFormer frame by frame as a time series.
[0085] S2: Compensate for the drone motion based on the video image v2, eliminate the noise of large displacement of whole pixels, and obtain a processed video image v3;
[0086] Step S2 specifically includes the following steps:
[0087] S21: Train the optical flow neural network FlowFormer using the publicly available dataset FlyingChairs, and calculate the optical flow between images frame by frame through the optical flow neural network FlowFormer;
[0088] In step S21, the input signal of the optical flow neural network FlowFormer is an image pair composed of a video reference frame and a current frame, and the output signal of the optical flow neural network FlowFormer is the full-field optical flow data corresponding to the image pair.
[0089] Step S21 specifically includes the following steps:
[0090] Step 1: Input an image of H I ×W I ×3 RGB, extract a feature map of H×W×D f through the backbone vision network, extract the source image of the previous frame and the target image of the current frame based on the feature map, calculate the dot product similarity corresponding to all pixel points in the source image and the target image, and construct a 4D cost volume of H×W×H×W, where (H,W) = (H I / 8, W I / 8);
[0091] Step 2: Identify the position corresponding to the source pixel in the target image through the source-target visual similarity encoded in the 4D cost volume, and eliminate the repeated patterns and non-distinguishing regions in the two frames of images through the cost encoder;
[0092] Step 2 specifically includes the following steps:
[0093] Step one: Block the cost map M x for each source pixel point x through strided convolution. The cost map M x is expressed as:
[0094]
[0095] to obtain multiple cost blocks q x The embedded cost map M x , and convert the multiple cost blocks q x The embedded cost map M x into a feature map F x through the activation function ReLU. The feature map F x is expressed as:
[0096]
[0097] where the cost block feature F x in the feature map F x corresponds to the 8×8 cost block q in the cost mapx One-to-one correspondence;
[0098] Step 2: Randomly initialize the potential codeword C, and through K potential codewords For each cost block feature F x Perform a summary, query each cost block feature F through the dot product attention mechanism x , through the latent representation T x Summarize the cost map into K latent vectors of dimension D Determine the embedding position of each cost block feature F through the COTR method x to obtain a 4D cost volume T, and the 4D cost volume T is expressed as:
[0099]
[0100] The above steps convert the original 4D cost volume into a compact latent 4D cost volume T.
[0101] Update the potential codeword C through backpropagation and share the potential codeword C among different source pixel points x;
[0102] In Step 2, the latent representation T x is expressed as:
[0103] K x = Conv 1×1 (Concat(F x , PE))
[0104] V x = Conv 1×1 (Concat(F x , PE))
[0105] T x = Attention(C, K x , V x )
[0106] where the key K x and the value V x both represent the projection of the cost block feature F x , and the key K x and the value V x are respectively expressed as:
[0107] K x = FFN(T x )
[0108] V x = FFN(T x ),
[0109] The cost block feature Fx The projection is the key K x and the value V x Before that, for the cost block feature F x and the position embedding sequence PE are concatenated. The position embedding sequence PE is expressed as:
[0110]
[0111] where D p represents the encoding length.
[0112] Step 3: To reduce the geometric complexity growth caused by the increase in the number of marker vectors and lower the computational cost, the 4D cost volume T is grouped by the alternating group Transformer layer AGT. Based on the grouping result, the corresponding number of cost queries Q is generated through the feed-forward neural network FFN x , the cost query Q x is expressed as:
[0113] Q x = FFN(FFN(q x ) + PE(p))
[0114] where q x represents a 9×9 cost block centered at position p extracted from the cost map M x ;
[0115] In step 3, the AGT layer specifically includes the following two grouping methods:
[0116] Method 1: The latent representations of each source pixel point x are combined and set as a group. Spatial separated self-attention calculation is performed on all the latent representations within the group to obtain the updated latent representation T x , the updated latent representation T x is expressed as:
[0117]
[0118] where T x (i) represents the i-th latent representation of the source pixel point x. The latent representation T x after self-attention calculation is processed through the feed-forward network FFN and reorganized into the 4D cost volume T to form the self-attention within the cost map M x ;
[0119] Method 2: Based on the latent representation T x, all 4D cost volumes \(T\) are grouped, with each group containing \(H\times W\) cost tokens \(T\). Spatial-separated self-attention calculation is performed on all cost tokens \(T\) within the group to obtain the updated cost tokens \(T\). i , the updated cost tokens \(T\) i are expressed as:
[0120] \(T\) i = FFN(SS - SelfAttention(\(T\) i ), \(i = 1, 2, \ldots, K\).
[0121] Through the alternating operations of these two methods, AGT can effectively exchange information between source pixels and latent representations, and finally convert the 4D cost volume into a more compact cost memory for the decoding of optical flow estimation.
[0122] Step 4: Predict the optical flow of the source pixels through cost queries, and calculate the corresponding position \(p\) of each source pixel point \(x\) in the target image. The position \(p\) is expressed as:
[0123] \(p=x + f(x)\)
[0124] where \(f(x)\) represents the current optical flow estimation;
[0125] Step 5: Generate the cost memory information \(c\) of the aggregated source pixel points \(x\) through the cross-attention mechanism x , the cost memory information \(c\) x is expressed as:
[0126] \(c\) x = Attention(Q x , K x , V x ).
[0127] Step 3: Through the cost decoder, predict the optical flow residual \(\Delta f(x)\) in a recursive manner. The optical flow residual \(\Delta f(x)\) is expressed as:
[0128] \(\Delta f(x)=\text{ConvGRU}(\text{Concat}(c x , q x , t x , f(x)))\)
[0129] where \(t\) x represents the context features of the source image, and the optical flow prediction is gradually optimized in an iterative manner. During the iteration, the optical flow residual \(\Delta f(x)\) is regressed through the gated recurrent unit ConvGRU.
[0130] S22: Calculate the target integer-pixel displacement data through the optical flow data;
[0131] S23: Based on the whole-pixel displacement data, perform heuristic image translation processing on the video image v1 frame by frame through a heuristic algorithm, limit the noise caused by camera movement within a unit pixel, and generate the processed video image v3.
[0132] S3: Obtain the video image v4 without the structure to be measured based on the video image v3, compensate for the movement of the drone based on the video image v4, eliminate the sub-pixel small displacement noise, and obtain the video image v5 after precise motion compensation;
[0133] Step S3 specifically includes the following steps:
[0134] S31: Based on the video image v3, obtain the background without the structure to be measured, set the reference region video image without the structure to be measured as the video image v4, and the coordinate positions of the video image v4 and the video image v2 in the video image v1 are the same;
[0135] During the video stabilization process, first, the user needs to manually select a static reference region in the region to be measured as the benchmark for subsequent optical flow calculation. This operation is implemented through mouse interaction. The user selects a static region in the first frame, uses the OpenCV library callback function cv2.setMouseCallback to register the mouse event to capture the user's input. This function listens for mouse events and calls the select_roi callback function to determine the region according to mouse clicks and drags. After the user finishes selecting, extract the coordinates of the selected region for subsequent processing.
[0136] Combine the optical flow compensation algorithm KLT of the variational pyramid, extract the feature points in the video image v4 frame by frame, use the OpenCV library function cv2.goodFeaturesToTrack() to detect the corner points that are easy to track in the image. Through detection, obtain 100 feature points for tracking. These feature points will be tracked by the optical flow method in subsequent frames. The processing of each frame is based on the optical flow method. Estimate the movement of the camera or object by calculating the movement amount of the feature points between the current frame and the previous frame. The optical flow method is implemented through the OpenCV library function cv2.calcOpticalFlowPyrLK();
[0137] To eliminate the sub-pixel noise of camera vibration, estimate the translational amount between frames by tracking the movement of these feature points. For each frame, calculate the average displacement of the feature points to obtain the translational vectors in the horizontal and vertical directions. Obtain the displacement of the feature points by calculating the difference between the coordinates of the feature points in the current frame and the previous frame, and then take the average to obtain the overall translation;
[0138] After obtaining the translation between frames, represent the sub-pixel transformation relationship between different frames caused by camera movement through an affine transformation matrix;
[0139] S33: Perform an affine transformation on the video image v4 frame by frame using an affine transformation matrix, and adjust the current frame back to its original position using a translation transformation matrix to compensate for sub-pixel jitter.
[0140] Use the OpenCV library function cv2.warpAffine() to perform an image translation matrix transformation. After processing the current frame, use the current frame as the previous frame for the next frame, and repeat the above optical flow calculation and correction process. To maintain the consistency of feature points, update the tracked feature points for the next iteration.
[0141] Thereby eliminating the sub-pixel small displacement noise in the video image v3 and obtaining the video image v5 after precise motion compensation.
[0142] First, perform integer pixel correction on the video image through a heuristic algorithm, restricting the movement of the camera within the range of one pixel, thereby effectively eliminating the vibration noise caused by large movements during the hovering of the drone, creating the necessary conditions for the next step of using the optical flow frame-by-frame motion compensation algorithm, and effectively solving the problem of the optical flow tracking algorithm failing due to the movement of the drone.
[0143] S4: Perform optical flow estimation and post-processing based on the video image v5, implemented using the Python language and the OpenCV module.
[0144] Step S4 specifically includes the following steps:
[0145] S41: Based on the video image v5, set the area to be measured as the ROI area, and extract the local phase information of the image in the ROI area through a complex steerable pyramid filter.
[0146] S42: Track the ROI area through the optical flow compensation algorithm KLT, calculate the displacement time history data at the sub-pixel level of the feature points in the area to be measured, multiply the pixel displacement time history data by a scale factor, and convert the pixel displacement time history data into real displacements.
[0147] S43: Based on the real displacement data of multiple measurement points, obtain the frequency domain power spectrum PSD through the fast Fourier transform FFT, and perform modal analysis on the data of the frequency domain power spectrum PSD through the frequency decomposition method FDD to obtain the modal information of the structure to be measured.
[0148] Therefore, the present invention adopts the above-mentioned method for collecting structure vibration signals without a target based on a drone platform, performs integer pixel large displacement denoising and sub-pixel small displacement denoising based on computer vision methods, and performs displacement tracking and modal analysis, realizing high-precision shooting of local vibration information, overcoming the long-distance resolution limitation of a fixed camera, and being able to quickly process target vibration images and output vibration information.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the method of the present invention, and these modifications or equivalent replacements cannot make the modified method deviate from the spirit and scope of the method of the present invention.
Claims
1. A target-free structural vibration signal acquisition method based on an unmanned aerial vehicle platform, characterized in that: The following steps are involved: S1: The drone approaches the infrastructure structure to collect and preprocess video data, and obtains the video image v1 of the vibration process and the video image v2 of the irregular motion; S2: Compensate the drone motion based on the video image v2, eliminate the integer pixel large displacement noise, and obtain the processed video image v3; S3: Based on the video image v3, a video image v4 without the structure to be measured is obtained. Based on the video image v4, the motion of the UAV is compensated to eliminate the sub-pixel small displacement noise, and the video image v5 after accurate motion compensation is obtained; S4: Optical flow estimation and post-processing based on video image v5.
2. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 1 is characterized in that: Step S1 specifically includes the following steps: S11: Use a drone to shoot a video of the area to be tested to obtain a video image v1 of the vibration process; S12: Acquire a static background of the video image v1, and the camera generates irregular motion when the drone hovers in the static background, so as to obtain a video image v2 with irregular motion corresponding to the static background; S13: Input the video image v2 as a time series frame by frame into the optical flow neural network FlowFormer.
3. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 1 is characterized in that: Step S2 specifically includes the following steps: S21: Train the optical flow neural network FlowFormer through the public dataset FlyingChairs, and calculate the optical flow of images between frames through the optical flow neural network FlowFormer; S22: Calculate the target integer pixel displacement data through the optical flow data; S23: Based on the integer pixel displacement data, a heuristic image translation process is performed on the video image v1 frame by frame through a heuristic algorithm, the noise caused by the camera motion is limited within a unit pixel, and a processed video image v3 is generated.
4. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 3 is characterized in that: In step S21, the input signal of the optical flow neural network FlowFormer is an image pair consisting of a video reference frame and a current frame, and the output signal of the optical flow neural network FlowFormer is the full-field optical flow data corresponding to the image pair.
5. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 3 is characterized in that: Step S21 specifically includes the following steps: Step 1: Type H I ×W I ×3RGB image, extract H×W×D through the backbone visual network f Based on the feature map, the source image of the previous frame and the target image of the current frame are extracted, the dot product similarity of all pixels in the source image and the target image is calculated, and a 4D cost volume of H×W×H×W is constructed, where (H,W)=(H I / 8,W I / 8); Step 2: Identify the corresponding positions of source pixels in the target image through the source-eye visual similarity encoded in the 4D cost volume, and eliminate repeated patterns and non-distinguishing areas in the two frames through the cost encoder; Step 3: Through the cost decoder, the optical flow residual Δf(x) is predicted recursively. The optical flow residual Δf(x) is expressed as: Δf(x)=ConvGRU(Concat(c x ,q x ,t x ,f(x))) Among them, t x It represents the contextual features of the source image and gradually optimizes the optical flow prediction in an iterative manner. During the iteration, the optical flow residual Δf(x) is regressed through the gated recurrent unit ConvGRU.
6. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 5 is characterized in that: Step 2 specifically includes the following steps: Step 1: Cost map M for each source pixel x through strided convolution x Block processing, cost map M x It is expressed as: Get multiple cost blocks q x Embedded cost map M x , multiple cost blocks q are transformed through the activation function ReLU x Embedded cost map M x Convert to feature map F x , feature map F x It is expressed as: Among them, the feature map F x The cost block feature F in x With the 8×8 cost block q in the cost map x One to one correspondence; Step 2: Randomly initialize the potential codeword C, and use K potential codewords For each cost block feature F x Perform a summary and query each cost block feature F through the dot product attention mechanism x , through the potential representation T x The cost map is summarized into K latent vectors of dimension D. Determine the feature F of each cost block by COTR method x The embedded position of , we get the 4D cost volume T, which is expressed as: Update the latent codeword C through back-propagation and share the latent codeword C between different source pixels x; Step 3: The 4D cost volume T is grouped through the alternating grouping Transformer layer AGT. Based on the grouping results, the corresponding number of cost queries Q are generated through the feed-forward neural network FFN x , cost query Q x It is expressed as: Q x =FFN(FFN(q x )+PE(p)) Among them, q x Represents the cost graph M x Extract a 9×9 cost block centered at position p; Step 4: Predict the optical flow of the source pixel through cost query and calculate the corresponding position p of each source pixel point x in the target image. The position p is expressed as: p=x+f(x) Where f(x) represents the current optical flow estimate; Step 5: Generate the cost memory information c of the aggregated source pixel x through the cross attention mechanism x , cost memory information c x It is expressed as: c x =Attention(Q x ,K x ,V x )。 7. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 6 is characterized in that: In step 2, the latent representation T x It is expressed as: K x =Conv 1×1 (Concat(F x ,ON)) V x =Conv 1×1 (Concat(F x ,ON)) T x =Attention(C,K x ,V x ) Among them, key K x Sum value V x Both represent the cost block feature F x Projection of key K x Sum value V x Respectively expressed as: K x =FFN(T x ) V x =FFN(T x ), The cost block feature F x Projection is key K x Sum value V x Before, for the cost block feature F x And the position embedding sequence PE is concatenated, and the position embedding sequence PE is expressed as: Among them, D p Indicates the encoding length.
8. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 6 is characterized in that: In step 3, the AGT layer specifically includes the following two grouping methods: Method 1: The potential representation of each source pixel x Collect and set as a group, and all potential representations in the group Perform spatial separation self-attention calculation to obtain the updated potential representation T x , the updated potential representation T x It is expressed as: Among them, T x (i) represents the i-th potential representation of the source pixel x, and the potential representation T calculated by the self-attention through the feed-forward network FFN x Process and reorganize into 4D cost volume T to form cost map M x internal self-attention; Method 2: Based on potential representation T x , group all 4D cost volumes T according to , each group contains H×W cost markers T, perform spatial separation self-attention calculation on all cost markers T in the group, and obtain the updated cost marker T i , the updated cost mark T i It is expressed as: T i =FFN(SS-SelfAttention(T i )),i=1,2,…,K。 9. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31: based on the video image v3, a background without the structure to be measured is obtained, and a video image of a reference area without the structure to be measured is set as a video image v4, where the coordinate positions of the video image v4 and the video image v2 in the video image v1 are the same; S32: Combined with the variational pyramid optical flow compensation algorithm KLT, feature points in the video image v4 are extracted frame by frame, and optical flow tracking is performed on the feature points. The sub-pixel transformation relationship between different frames caused by camera motion is represented by the affine transformation matrix; S33: performing affine transformation on the video image v4 frame by frame through an affine transformation matrix, eliminating sub-pixel small displacement noise in the video image v3, and obtaining a video image v5 after precise motion compensation.
10. The method for collecting target-free structural vibration signals based on an unmanned aerial vehicle platform according to claim 1, characterized in that: Step S4 specifically includes the following steps: S41: Based on the video image v5, the area to be measured is set as the ROI area, and the local phase information of the image in the ROI area is extracted by a complex steerable pyramid filter; S42: Tracking the ROI area through the optical flow compensation algorithm KLT, calculating the sub-pixel displacement time-history data of the feature points in the measured area, multiplying the pixel displacement time-history data by the scale factor, and converting the pixel displacement time-history data into real displacement; S43: Based on the real displacement data of multiple measuring points, the frequency domain power spectrum PSD is obtained by fast Fourier transform FFT, and the modal analysis of the frequency domain power spectrum PSD data is performed by frequency decomposition method FDD to obtain the modal information of the structure to be measured.
Citation Information
Patent Citations
High-resolution aerial video moving object detection method based on depth neural network
CN109063549A
Structural vibration monitoring and correcting method and device based on unmanned aerial vehicle and storage medium
CN117935096A
Bridge vibration real-time monitoring method, system and device based on visual enhancement
CN118706244A
Structural vibration displacement identification method based on diamond search and improved optical flow method
CN118762054A
Computer vision measurement method for unmarked structure vibration
CN118918099A