A computer vision structural vibration identification method and device
By using spatial decomposition, resampling and reconstruction, and reweighted residual sparsity processing, the labeling requirements of traditional computer vision measurement methods and the noise problem of phase motion amplification technology are solved, thus achieving high-precision structural vibration recognition.
Patent Information
- Application Number
- CN202411709583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Traditional computer vision measurement methods require high-contrast markings to be placed on the surface of the structure being measured, which makes it difficult to handle large or hard-to-reach areas. Furthermore, structural vibration identification methods based on phase motion amplification technology are susceptible to noise and phase step phenomena.
By employing spatial decomposition, resampling and reconstruction, and reweighted residual sparsity processing, structural vibration videos are converted into image sequences, enhancing temporal correlation, resisting noise, and reducing phase step problems.
It achieves high-precision structural vibration analysis, effectively resists noise, reduces phase steps, and provides clear structural vibration identification results.
Smart Images

Figure CN119741631B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of video image processing technology, and in particular to a computer vision structural vibration identification method and device. BACKGROUND
[0002] The digital camera-based optical measurement non-contact motion identification method is a technology for detecting and analyzing object motion by using a digital camera and image processing technology. The camera-based optical measurement method is non-contact, and does not need to attach a traditional sensor to the measured object, so it does not introduce any additional mass and does not change the natural motion state of the object. When measuring with a digital camera, the image or video of the entire scene is captured by a single or multiple synchronized cameras at the same time. Since all motion data comes from the same video stream or image sequence captured simultaneously, the image frames are naturally synchronized, avoiding the problem of time difference between multiple sensors. That is, the digital camera-based optical measurement non-contact motion identification method can effectively solve the problems of additional mass caused by traditional sensors and different synchronization of wireless sensors.
[0003] However, although the digital camera-based optical measurement non-contact motion identification method has significant advantages in many aspects, the traditional computer vision measurement method needs to arrange a large number of high-contrast markers or speckle patterns on the surface of the measured structure, which is relatively time-consuming. When the to-be-measured area is large or difficult to access, the limitations of this method are obvious. In addition, for the traditional optical measurement system, using only a small amount of energy to excite the to-be-measured structure may result in the inability to capture subtle vibration responses.
[0004] The structural vibration identification method based on phase motion amplification technology can cope with the above challenges, however, the structural vibration identification method based on phase motion amplification technology is based on a basic assumption that small motion can be converted into phase shift of pixel color. In fact, the spatial displacement encoded in the phase information is limited, which may cause phase step phenomenon. In addition, phase estimation is an ill-posed problem in which observation noise is amplified and propagated into the estimation of phase and amplitude, which may cause image blurring. Therefore, there is an urgent need for a structural vibration identification method that can efficiently resist noise and solve the phase step problem. SUMMARY
[0005] The present application provides a computer vision structural vibration identification method and device, which can efficiently resist noise and solve the phase step problem.
[0006] The present application provides a computer vision structural vibration identification method, comprising:
[0007] Converting the original structural vibration video into a first image sequence;
[0008] spatially decompose the first image sequence frame by frame to obtain a first spatial signal, and calculate initial phase difference information of a second image sequence after spatial decomposition;
[0009] resample the first spatial signal and reconstruct to obtain a third image sequence of the second spatial signal;
[0010] process the third image sequence using the reweighted residual sparsity to enhance the temporal correlation of the third image sequence;
[0011] based on the phase difference information in the third image sequence, use conversion calculation to obtain structural vibration information.
[0012] In an exemplary instance, the original structural vibration video is converted into the first image sequence, including:
[0013] the original structural vibration video is converted into the first image sequence;
[0014] the first image sequence is converted from an RGB color space to a YIQ color space.
[0015] In an exemplary instance, the first image sequence is converted from an RGB color space to a YIQ color space according to the following formula:
[0016]
[0017] wherein Y represents a luminance signal, I represents a color difference signal sensitive to human eyes, Q represents a color difference signal insensitive to human eyes, R represents a red channel image signal, G represents a green channel image signal, and B represents a blue channel image signal.
[0018] In an exemplary instance, the first image sequence is spatially decomposed frame by frame to obtain a first spatial signal, and initial phase difference information of a second image sequence after spatial decomposition, including:
[0019] a complex-valued steerable pyramid model is used to spatially decompose the first image sequence frame by frame, and a second image sequence corresponding to each frequency is calculated to obtain the first spatial signal;
[0020] for each frequency feature of the second image sequence, the phase information of the current image and the first frame image are subtracted respectively to calculate the initial phase difference information Δψ of each frame image corresponding to the first frame image;
[0021] the initial phase difference information of all frequency components is obtained to obtain the initial phase difference information of the second image sequence.
[0022] In an exemplary instance, the complex-valued steerable pyramid model is as follows:
[0023]
[0024] where A denotes the amplitude, ψ denotes the initial phase, and θ denotes the frequency, denotes the convolution operation, denotes the real part of the convolution kernel with frequency θ, denotes the imaginary part of the convolution kernel with frequency θ, and M is the image;
[0025] The initial phase difference information Δψ is calculated as: Δψ θ (t) = ψ θ (t) - ψ θ (1), where t denotes time.
[0026] In an exemplary instance, the resampling and reconstruction of the first spatial domain signal to obtain the second spatial domain signal of the third image sequence comprises:
[0027] According to the initial phase difference information of the second image sequence, the first spatial domain information of the video image sequence is resampled and reconstructed using the compressed sensing sparse theorem to obtain the third image sequence of the second spatial domain signal.
[0028] In an exemplary instance, the reconstruction is as follows:
[0029]
[0030] where i denotes the image block index; P denotes the image block size; Φ P denotes the image-to-block conversion matrix; M denotes the video frame, which is used to represent a frame of image in the video; M P (i, t) is the resampling mode of M(i, t); is the hypothesis of M(i, t); H(i, t) is a single hypothesis, is the weight vector to be solved, which is used to represent the linear combination of the hypothesis H(i, t); the constraint condition is
[0031] In an exemplary instance, it further comprises:
[0032] The uncertainty estimation of the third image sequence is calculated using sparse coding to obtain the noise-free block domain matrix M s (t).
[0033] In an exemplary instance, the block domain uncertainty estimation is as follows:
[0034]
[0035] wherein denotes the optimized sparse coefficient matrix; C denotes the sparse coded coefficient matrix; denotes the noisy image block matrix at time t; ‖·‖ F denotes the Frobenius norm; ‖·‖1 denotes the L1 norm; D denotes the data dictionary, and γ denotes the regularization parameter;
[0036] the non-noisy block domain matrix M s (t) is:
[0037] In an exemplary instance, the third image sequence is processed by the following formula to enhance the time correlation of the third image sequence:
[0038]
[0039] s.t.T=M(t),
[0040] wherein T denotes an intermediate variable of the frame M(t); M(t) denotes the original vibration video frame at time t; denotes the reconstructed video frame at time t; denotes the non-noisy block domain matrix of the key video frame; key denotes the selected video key frame; Φ denotes the image-to-block conversion matrix; β is the regularization parameter, nB denotes the total sum of the number of blocks in the frame, F is the weighted residual error; the constraint condition is T=M(t).
[0041] In an exemplary instance, the structural vibration information is obtained using conversion calculation based on the phase difference information in the third image sequence, comprising:
[0042] obtaining the structural motion response in the video based on the phase difference information in the third image sequence;
[0043] According to the actual length of the structure, the obtained structural motion response in the video is converted into the structural vibration information.
[0044] In an exemplary instance, the structural motion response in the video is calculated by the following formula based on the phase difference information in the third image sequence:
[0045]
[0046] wherein u denotes the structural motion response in the x direction, v denotes the structural motion response in the y direction, ψ0 denotes the phase of the structure in the x direction, ψ π / 2 denotes the phase of the structure in the y direction.
[0047] In an exemplary instance, the conversion of the obtained structural motion response in the video into the structural vibration information comprises:
[0048] calculating a ratio between a length of the known structure and a number of pixels the structure spans on the image;
[0049] calculating a product of the structure motion response and the obtained ratio in the video to obtain the structure vibration information.
[0050] The embodiment of the present application further provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are used for executing the computer vision structure vibration identification method.
[0051] The embodiment of the present application further provides a computer device, which comprises a memory and a processor, wherein the memory stores instructions executable by the processor, and the instructions are used for executing the steps of the computer vision structure vibration identification method.
[0052] The embodiment of the present application further provides a computer vision structure vibration identification device, which comprises a conversion module, a decomposition processing module, a reconstruction module, an enhancement processing module and an acquisition module.
[0053] The conversion module is used for converting an original structure vibration video into a first image sequence.
[0054] The decomposition processing module is used for performing spatial domain decomposition on the first image sequence frame by frame to obtain a first spatial domain signal, and calculating initial phase difference information of a second image sequence after spatial domain decomposition.
[0055] The reconstruction module is used for resampling the first spatial domain signal and reconstructing to obtain a third image sequence of a second spatial domain signal.
[0056] The enhancement processing module is used for processing the third image sequence of the reconstructed second spatial domain signal using a reweighted residual sparsity to enhance time correlation of the third image sequence of the reconstructed second spatial domain signal.
[0057] The acquisition module is used for obtaining structure vibration information using conversion calculation based on phase difference information in the processed third image sequence.
[0058] The computer vision structure vibration identification method provided by the embodiment of the present application comprises spatial domain decomposition, resampling and reconstruction, reweighted residual sparsity processing and time correlation enhancement, realizes processing and identification from a video to vibration information, effectively resists noise, reduces phase step problems and provides high-precision structure vibration analysis results.
[0059] Additional features and advantages of the application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the application. The objectives and other advantages of the application will be realized and attained by the structure particularly pointed out in the description and claims. BRIEF DESCRIPTION OF DRAWINGS
[0060] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application.
[0061] Figure 1 A flow chart of the structural vibration identification method of computer vision in the embodiments of the application is shown in the figure.
[0062] Figure 2 A schematic diagram of the composition structure of the structural vibration identification device of computer vision in the embodiments of the application is shown in the figure.
[0063] Figure 3(a) is a schematic diagram of the motion of the selected three-layer frame and the motion shown in the slice thereof in the embodiments of the application. Figure 1 ;
[0064] Figure 3(b) is a schematic diagram of the motion of the selected three-layer frame and the motion shown in the slice thereof in the embodiments of the application.
[0065] Figure 3(c) is a schematic diagram of the motion of the selected three-layer frame and the motion shown in the slice thereof in the embodiments of the application.
[0066] Figure 3(d) is a schematic diagram of the motion of the selected three-layer frame and the motion shown in the slice thereof in the embodiments of the application.
[0067] Figure 4(a) is a frame similarity quantitative index curve graph showing the change of noise variance from 0.01 to 0.145 using the peak signal-to-noise ratio in the embodiments of the application.
[0068] Figure 4(b) is a frame similarity quantitative index curve graph showing the change of noise variance from 0.01 to 0.145 using the structural similarity in the embodiments of the application.
[0069] Figure 5(a) is a structural vibration table plane layout diagram of a stone curtain wall in the embodiments of the application.
[0070] Figure 5(b) is a top view of a stone curtain wall structural vibration table test device in the embodiments of the application.
[0071] Figure 6(a) is a stone curtain wall structural vibration curve of measuring point 1 in the embodiments of the application. Figure 1 ;
[0072] Fig. 6(b) is a schematic diagram of the vibration curve of the stone curtain wall structure at measuring point 1 in the embodiment of the present application; Figure 2 ;
[0073] Fig. 6(c) is a schematic diagram of the vibration curve of the stone curtain wall structure at measuring point 1 in the embodiment of the present application;
[0074] Fig. 6(d) is a schematic diagram of the vibration curve of the stone curtain wall structure at measuring point 2 in the embodiment of the present application Figure 1 ;
[0075] Fig. 6(e) is a schematic diagram of the vibration curve of the stone curtain wall structure at measuring point 2 in the embodiment of the present application Figure 2 ;
[0076] Fig. 6(f) is a schematic diagram of the vibration curve of the stone curtain wall structure at measuring point 2 in the embodiment of the present application;
[0077] Fig. 7(a) is a vibration error diagram of the stone curtain wall structure at measuring point 1 in the embodiment of the present application;
[0078] Fig. 7(b) is a vibration error diagram of the stone curtain wall structure at measuring point 2 in the embodiment of the present application. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be described in detail below with reference to the drawings. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other as long as there is no conflict.
[0080] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings. The embodiments of the present application are shown in the drawings. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0081] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0082] It can be understood that the terms "first", "second" used in the present application are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited.
[0083] It can be understood that the "connection" in the following embodiments should be understood as "electrical connection", "communication connection" and the like if the circuits, modules, units and the like connected by the connection have the transmission of electrical signals or data between each other.
[0084] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", or the like, when used in this specification, specify the presence of stated features, integers, steps, operations, components, parts, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, parts, or combinations thereof.
[0085] Figure 1 A flowchart of a structural vibration identification method of computer vision in the embodiments of the present application is shown in FIG. 1, which can include the following steps: Figure 1
[0086] Step 100: converting the original structural vibration video into a first image sequence.
[0087] In an exemplary instance, step 100 can include:
[0088] converting the original structural vibration video into a first image sequence;
[0089] converting the obtained first image sequence from RGB color space to YIQ color space.
[0090] In an exemplary instance, the continuous vibration video is decomposed into a first image sequence composed of frame-by-frame images, so as to analyze each frame independently. In an embodiment, a video processing library (such as OpenCV) can be used to decompose the video frame by frame and save it as a separate image file.
[0091] In an exemplary instance, each frame of image in the obtained first image sequence is usually represented in RGB color space. In order to further process, especially when processing color information or doing some specific analysis (such as phase information extraction), it is necessary to convert the image from RGB color space to YIQ or other color space. YIQ color space helps to separate color information and brightness information. In an embodiment, formula (1) can be used to convert the first image sequence from RGB color space to YIQ color space:
[0092]
[0093] In formula (1), Y represents a luminance signal, I represents a color difference signal sensitive to the human eye, Q represents a color difference signal insensitive to the human eye, R represents a red channel image signal, G represents a green channel image signal, and B represents a blue channel image signal.
[0094] Step 101: performing spatial domain decomposition on the first image sequence frame by frame to obtain a first spatial domain signal, and calculating initial phase difference information of the second image sequence after spatial domain decomposition.
[0095] In an exemplary example, the image sequence can be decomposed frame by frame in the spatial domain by using a complex steerable pyramid model, a wavelet transform, a Laplacian pyramid, or other phase extraction methods, and initial phase difference information can be calculated and obtained.
[0096] In an exemplary example, when the complex steerable pyramid model is used to perform spatial domain decomposition on the first image sequence frame by frame and calculate initial phase difference information of the second image sequence after spatial domain decomposition, step 101 can include:
[0097] The first image sequence in the YIQ color space is decomposed frame by frame in the spatial domain by using the complex steerable pyramid model, and a second image sequence corresponding to each frequency is calculated, thereby obtaining the first spatial domain signal.
[0098] For each frequency feature of the second image sequence, the phase information of the current image and the phase information of the first frame image are subtracted respectively, and initial phase difference information Δψ of each frame image relative to the first frame image is calculated.
[0099] In step 101, first, each frame image in the YIQ color space is decomposed in the spatial domain, and the complex steerable pyramid model is used to decompose the image into multiple frequency components (i.e., features of different scales and directions); then, for each frame image, the phase information is calculated on each frequency component, and the corresponding frequency image sequence is generated. These image sequences represent the features of each frame image in different scales and directions. Then, for each frame image, the phase information is calculated on each frequency component, and the corresponding frequency image sequence is generated. These frequency image sequences represent the features of each frame image in different scales and directions. Then, for each frequency feature, the phase information of the current frame and the phase information of the first frame image are subtracted, and the initial phase difference information Δψ of each frame image relative to the first frame image is calculated. Finally, the phase difference information of all frequency components is integrated to obtain the complete initial phase difference data of the second image sequence. The initial phase difference information represents the subtle changes between images in different frames and is the key data for identifying vibrations.
[0100] In an embodiment, the complex steerable pyramid model can be represented as shown in formula (2):
[0101]
[0102] In formula (2), A represents amplitude, ψ represents initial phase, and θ represents frequency, represents convolution operation, represents real part of the convolution kernel with frequency θ, represents imaginary part of the convolution kernel with frequency θ, and M is an image.
[0103] In an embodiment, the phase difference Δψ θ (t) is calculated as shown in formula (3):
[0104] Δψ θ (t) = ψ θ (t) - ψ θ (1) (3)
[0105] In formula (3), t represents time.
[0106] Step 102: Resampling and reconstructing the first spatial domain signal to obtain a third image sequence of the second spatial domain signal.
[0107] In an exemplary instance, the first spatial domain signal can be resampled and the third image sequence of the second spatial domain signal can be reconstructed using methods such as multi-hypothesis prediction method, interpolation technique, Laplace reconstruction, wavelet reconstruction, deep learning method, and frequency domain reconstruction.
[0108] In an embodiment, step 102 can include:
[0109] According to the calculated initial phase difference information of the second image sequence, the first spatial domain information of the video image sequence is resampled and reconstructed using the compressed sensing sparse theorem to obtain the third image sequence of the second spatial domain signal.
[0110] In an embodiment, the spatial domain reconstruction is shown in formula (4):
[0111]
[0112] In formula (4), i represents image block label, the entire image can be divided into multiple small blocks, each small block has a unique label; P represents image block size; Φ P represents image-to-block conversion matrix, that is, matrix operation for converting the entire image into each image block; M represents video frame, used to represent a frame of image in the video; M P (i, t) is a resampling mode of M(i, t), that is, a resampling result of a specific image block i extracted from the video frame M at time t; is a hypothesis of M(i, t). H(i, t) is a single hypothesis, is the weight vector to be solved, which is used to represent the linear combination of the hypotheses H(i,t).
[0113] By formula (4), the minimization of M P (i,t) and Φ P The two-norm error between H(i,t)ω is achieved. The two-norm ‖·‖2 represents the Euclidean distance, and the goal of the optimization problem is to find the optimal weight vector ω, so that the signal Φ P H(i,t)ω obtained by the linear combination of the hypotheses H(i,t) is closest to the resampled signal M P (i,t). The constraint condition is means that the reconstructed image block must be equal to the signal obtained by the combination of the optimal weight vector . This processing ensures that the final image block is consistent with the linear combination of the hypothesis model.
[0114] In an embodiment, the structural changes in each frame of image are estimated based on the multi-hypothesis prediction method, and the resampled signal is more accurate or has better spatiotemporal characteristics than the original signal, and the accuracy of the signal is enhanced by resampling. After resampling, spatial signal reconstruction is also performed, so that the spatial characteristics of the image sequence are reconstructed, ensuring that the vibration information can be accurately represented at different scales.
[0115] In an exemplary instance, step 102 can further include:
[0116] The uncertain area in the third image sequence is identified and processed by sparse coding based on the weight matrix. In an embodiment, it can include: using sparse coding to estimate the uncertainty of the third image sequence, so as to calculate the noise-free block domain matrix M s (t).
[0117] In an embodiment, the block domain uncertainty estimation is as shown in formula (5):
[0118]
[0119] In formula (5), denotes the optimized sparse coefficient matrix, which is obtained by minimizing the objective function; C denotes the coefficient matrix of sparse coding; denotes the noisy image block matrix at time t; ‖·‖ F denotes the Frobenius norm, which is usually used to measure the difference between matrices; ‖·‖1 denotes the L1 norm, which is usually used to realize the regularization representation of sparsity; D denotes the data dictionary, and γ denotes the regularization parameter. It should be noted that,
[0120] solved The process can be implemented using the Alternating Direction Method of Multipliers (ADMM). The specific implementation is not intended to limit the scope of protection of this application and will not be elaborated here.
[0121] In one embodiment, the noise-free block domain matrix M s (t) is the result of reconstruction after sparse coding, as shown in formula (6):
[0122]
[0123] During phase motion estimation, phase information exceeding the phase limit is compressed due to the phase constraint, leading to inaccurate phase estimation results. Furthermore, phase estimation is also affected by ill-conditioned estimation, significantly impacting its accuracy. Therefore, step 102 resamples and reconstructs the structural phase motion information, avoiding phase step and ill-conditioned estimation problems.
[0124] Step 103: Use reweighted residual sparsity to process the third image sequence of the reconstructed second spatial domain signal to enhance the temporal correlation of the third image sequence of the reconstructed second spatial domain signal.
[0125] In one embodiment, the reconstructed second spatial signal's third image sequence can be processed using formula (7) to enhance the temporal correlation of the third image sequence:
[0126]
[0127] stT=M(t)
[0128] In formula (7), T represents the intermediate variable of frame M(t); M(t) represents the original video frame at time t. The reconstructed video frame at time t represents the objective of the optimization problem. The matrix represents the noise-free block domain of the key video frame. Key frames are typically frames with higher sampling rates or greater importance, used to guide the reconstruction of other frames. `key` represents the selected video key frame, i.e., one with a higher sampling rate. `Φ` represents the image-to-block transformation matrix. `β` is the regularization parameter, `nB` represents the total number of blocks within the frame, and `F` is the weighted residual. The constraint `T = M(t)` means that the intermediate variable `T` is actually a representation of the original frame `M(t)`, ensuring that the final solution uses the correct frame. It is a reconstructed version based on the original frame.
[0129] The objective function of formula (7) consists of two parts: Keyframe The error between the intermediate variable T and the target is to ensure that the intermediate variable T is as close as possible to the key frame after conversion, thereby maintaining consistency between frames; β∑ 1≤i≤nB ‖F(M s The sum of the weighted and sparse residual errors of each image block in the video frame encourages the sparsity of the residual error, meaning that the reconstruction error of the image block is concentrated on as few blocks as possible to maintain the temporal correlation between frames. Formula (7) solves the And The reconstructed frame can maintain high consistency with the key frame (enhanced temporal correlation) and sparsity at the block level, thereby reducing noise and uncertainty.
[0130] Since spatial domain and block domain resampling and reconstruction will lose video information in the time domain, the reweighted residual sparsity is used in the embodiments of the present application to enhance the temporal correlation of the image sequence. Step 103 realizes the enhancement of the reweighted sparsity of the residual error between the blocks of the key video frame and the linear combination of the multi-hypothesis prediction of the non-key video frame. The temporal correlation of the third image sequence of the second spatial signal obtained after step 103 enhances the reconstructed video quality, making the transition between frames smoother, thereby reducing the artifacts or inconsistencies introduced by reconstruction.
[0131] Step 104: Based on the phase difference information in the processed third image sequence, the structural vibration information is obtained using conversion calculation.
[0132] In an exemplary example, the optimized phase difference information contains structural motion information in the video, and the related motion information also needs to be converted into actual structural vibration information of the structure. Step 104 can include:
[0133] Based on the phase difference information in the processed third image sequence, the structural motion response in the video is obtained;
[0134] According to the actual length of the structure, the obtained structural motion response in the video is converted into structural vibration information.
[0135] In an exemplary example, based on the optimized phase difference information in the third image sequence, the structural motion response in the video can be calculated by formula (8) and formula (9):
[0136]
[0137] In formula (8) and formula (9), u represents the structural motion response in the x direction, v represents the structural motion response in the y direction, ψ0 represents the phase of the structure in the x direction, and ψ π / 2 represents the phase of the structure in the y direction.
[0138] In an exemplary instance, converting the obtained structural motion response in a video into structural vibration information without considering image deformation can include:
[0139] calculating a ratio between a length of a known structure and a number of pixels of the structure spanning on an image;
[0140] calculating a product of the structural motion response in the video and the obtained ratio to obtain the structural vibration information.
[0141] The computer vision structural vibration identification method provided by the embodiments of the present application includes spatial domain decomposition, resampling and reconstruction, re-weighted residual sparsity processing, and time correlation enhancement, realizes processing and identification from a video to vibration information, effectively resists noise, reduces phase step problems, and provides high-precision structural vibration analysis results.
[0142] The computer vision structural vibration identification method provided by the embodiments of the present application can also be applied to computer vision micro-vibration identification based on sparse enhancement compressed sensing.
[0143] The present application also provides a computer readable storage medium storing computer executable instructions for executing the computer vision structural vibration identification method of any one of the above.
[0144] The present application further provides a computer device including a memory and a processor, wherein the memory stores instructions executable by the processor, for performing the steps of the computer vision structural vibration identification method of any one of the above.
[0145] The embodiments of the present application also provide a computer vision structural vibration identification device, which can include a conversion module, a decomposition processing module, a reconstruction module, an enhancement processing module, and an obtaining module; wherein,
[0146] The conversion module is configured to convert an original structural vibration video into a first image sequence;
[0147] The decomposition processing module is configured to perform spatial domain decomposition on the first image sequence frame by frame to obtain a first spatial domain signal, and calculate initial phase difference information of a second image sequence after spatial domain decomposition;
[0148] The reconstruction module is configured to resample and reconstruct the first spatial domain signal to obtain a third image sequence of a second spatial domain signal;
[0149] The enhancement processing module is configured to use re-weighted residual sparsity to process the third image sequence of the reconstructed second spatial domain signal to enhance the time correlation of the third image sequence of the reconstructed second spatial domain signal;
[0150] The acquisition module is used to obtain structural vibration information based on the phase difference information in the processed third image sequence and by using conversion calculation.
[0151] The computer vision-based structural vibration recognition device provided in this application realizes the processing and recognition of vibration information from video through spatial decomposition, resampling and reconstruction, reweighted residual sparsity processing, and temporal correlation enhancement. It effectively resists noise, reduces phase step problems, and provides high-precision structural vibration analysis results.
[0152] The computer vision-based structural vibration recognition method provided in this application can be evaluated from two aspects: motion amplification and vibration recognition. In the motion amplification analysis, this embodiment selects vibration videos of a three-layer frame. To ensure the stability of the algorithm, the method of this embodiment is compared with traditional inference-based and phase-based methods. To increase the difficulty of the analysis, Gaussian noise with a variance range of 0.01 to 0.145 is pre-added to all analyzed video signals.
[0153] In the evaluation of the magnified results of the frame structure, the selected horizontal slice line recording the structural motion is parallel to the floor level and located on the 2nd floor, see [link / reference]. Figures 3(a)-3(d) The upper half of the image. In the unmagnified original video shown in Figure 3(a), it can be seen that the original vibration information is basically imperceptible. When the minute motion is magnified 40 times, as shown in Figures 3(b) and 3(c), it can be seen that the phase-based and reasoning-based methods can roughly reveal the motion of the structure. However, the original noise is amplified, and the motion revealed by the shape of the frame structure and the horizontal tangent is basically severely blurred, as shown globally in Figure 3(b) and in the white area in Figure 3(c). In contrast, in Figure 3(d), the video motion processed by the computer vision structural vibration recognition method provided in this application embodiment (i.e., the proposed method in the figure) is clearer and has almost no noise. Quantitative comparisons of video motion amplification are shown in Figures 4(a) and 4(b). For each method, the similarity between the amplified video and the original video is calculated frame-by-frame using Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM). In Figures 4(a) and 4(b), the dashed line represents the phase-based method, the dashed line represents the inference-based method, and the solid line represents the computer vision-based structural vibration recognition method provided in this application embodiment. It can be observed that the computer vision-based structural vibration recognition method provided in this application embodiment significantly outperforms the amplified motion of minute structures.
[0154] The shaking table test used to verify the effectiveness of the computer vision-based structural vibration recognition process provided in this application embodiment can be shown in Figures 5(a) and 5(b). Figure 5(a) is a plan view of the shaking table layout for the stone curtain wall structure in this application embodiment. The structure has a planar dimension of 3000mm (x-direction) × 3000mm (y-direction) and a total height of 4260mm. The camera sampling frequency is 50Hz, and the resolution of the video image sequence is 1920×1080. Figure 5(b) is a top view of the shaking table test device for the stone curtain wall structure in this application embodiment, where measuring point 1 (shown as a circular black dot) is located in the middle of the layer, and measuring point 2 (shown as a square black dot) is located on the top of the layer. The vibration recognition results are compared as follows: Figures 6(a)-6(f) As shown in Figures 7(a) and 7(b), the error comparison is illustrated in Figures 6(a), 6(b), 6(c), and 7(a), respectively. The results for measurement point 1 are shown in Figures 6(d), 6(e), 6(f), and 7(b), respectively. Although phase-based and inference-based methods can accurately estimate some structural vibrations, significant deviations occur when large vibrations occur. Furthermore, the outliers obtained by phase-based and inference-based methods are far more numerous than those obtained by the computer vision-based structural vibration recognition method (i.e., the method proposed in the figures) provided in this application embodiment, as specifically shown in Figures 7(a) and 7(b). In summary, the computer vision-based structural vibration recognition method provided in this application embodiment has better accuracy.
[0155] Although the embodiments disclosed in this application are as described above, the content described is merely for the purpose of understanding this application and is not intended to limit this application. Any person skilled in the art to which this application pertains may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in this application; however, the scope of patent protection of this application shall still be determined by the scope defined in the appended claims.
Claims
1. A computer vision-based method for structural vibration recognition, characterized in that, include: The original structural vibration video was converted into a first image sequence; The first image sequence is decomposed frame by frame to obtain the first spatial domain signal, and the initial phase difference information of the second image sequence after spatial domain decomposition is calculated. The third image sequence of the second spatial domain signal is obtained by resampling and reconstructing the first spatial domain signal; The third image sequence is processed using reweighted residual sparsity to enhance its temporal correlation. Based on the phase difference information in the third image sequence, structural vibration information is obtained using conversion calculation; The step of performing spatial domain decomposition frame by frame on the first image sequence to obtain the first spatial domain signal, and calculating the initial phase difference information of the second image sequence after spatial domain decomposition, includes: The first image sequence is decomposed frame by frame using a complex-valued manipulable pyramid model to calculate the second image sequence corresponding to each frequency, thereby obtaining the first spatial domain signal. For each frequency feature of the second image sequence, the phase information of the current image is subtracted from that of the first frame image to calculate the initial phase difference information of each frame image corresponding to the first frame image. ; The initial phase difference information of the second image sequence is obtained based on the initial phase difference information of all frequency components; The complex-valued manipulable pyramid model is shown in the following equation: , in, Indicates amplitude, Indicates the initial phase. Indicates frequency, This represents the convolution operation. Indicates frequency as The real part of the convolution kernel, Indicates frequency as The imaginary part of the convolution kernel, It is an image; The initial phase difference information The calculation is as follows: ,in, Indicates time; The third image sequence obtained by resampling and reconstructing the first spatial domain signal to obtain the second spatial domain signal includes: Based on the initial phase difference information of the second image sequence, the first spatial domain signal of the video image sequence is resampled and reconstructed using the compressed sensing sparsity theorem to obtain the third image sequence of the second spatial domain signal. The reconstruction is shown in the following equation: , s.t. , in, Indicates the image block number; Indicates the size of the image block; This represents the transformation matrix from image to block; Represents a video frame, used to represent a single image frame in a video; yes The resampling mode, that is, from video frames Extract a specific image patch In time Resampling results at time step; yes The assumption; It is a single hypothesis. It is the weight vector that needs to be solved, used to represent the hypothesis. A linear combination; the constraint condition is ; The temporal correlation of the third image sequence is enhanced by processing it using the following formula: , s.t. , in, Representing a frame intermediate variables; Indicates time The original vibration video frame at that moment; Indicates time Reconstructing video frames in real time; The noise-free block domain matrix representing the key video frame; This indicates the selected video keyframe; This represents the transformation matrix from image to block; It is a regularization parameter. This represents the total number of blocks within the frame. It is the weighted residual; constraints ; The method of obtaining structural vibration information based on phase difference information in the third image sequence using transformation calculation includes: Based on the phase difference information in the third image sequence, the structural motion response in the video is obtained; Based on the actual length of the structure, the structural motion response in the obtained video is converted into the structural vibration information; It also includes: using sparse coding to estimate the uncertainty of the third image sequence, and calculating a noise-free block matrix. .
2. The structural vibration identification method according to claim 1, wherein, The process of converting the original structural vibration video into a first image sequence includes: The original structural vibration video is converted into the first image sequence; The first image sequence is converted from the RGB color space to the YIQ color space.
3. The structural vibration identification method according to claim 2, wherein, The first image sequence is converted from the RGB color space to the YIQ color space according to the following formula: , in, Indicates the brightness signal. The color difference signal that the human eye is sensitive to. Color difference signals that indicate insensitivity of the human eye Represents the red channel image signal. Indicates the green channel image signal. This represents the blue channel image signal.
4. The structural vibration identification method according to claim 1, wherein, The block uncertainty estimation is shown in the following equation: , in, This represents the optimized sparse coefficient matrix; The coefficient matrix represents the sparse coding. Indicates time A matrix of noisy image patches at each time step; Denotes the Frobenius norm; Represents the L1 norm; Represents a data dictionary. Represents the regularization parameter; The noise-free block matrix for: .
5. The structural vibration identification method according to claim 1, wherein, Based on the phase difference information in the third image sequence, the structural motion response in the video is calculated using the following formula: , , in, Indicates in Structural motion response in the direction, Indicates in Structural motion response in the direction, Indicates in Phase of the structure in the direction, Indicates in The phase of the structure in the direction.
6. The structural vibration identification method according to claim 5, wherein, The process of converting the structural motion response in the obtained video into structural vibration information includes: Calculate the ratio between the length of a known structure and the number of pixels that the structure spans on the image; The structural vibration information is obtained by calculating the product of the structural motion response in the video and the obtained ratio.
7. A computer-readable storage medium storing computer-executable instructions for performing the computer vision-based structural vibration recognition method according to any one of claims 1 to 6.
8. A computer device comprising a memory and a processor, wherein, The memory stores the following instructions that can be executed by a processor: for performing the steps of the computer vision structural vibration recognition method according to any one of claims 1 to 6.
9. A computer vision-based structural vibration recognition device, characterized in that, The structural vibration recognition method based on computer vision according to any one of claims 1-6 is used, comprising: a conversion module, a decomposition processing module, a reconstruction module, an enhancement processing module, and an acquisition module; wherein, A conversion module is used to convert the original structural vibration video into a first image sequence; The decomposition processing module is used to perform spatial domain decomposition on the first image sequence frame by frame to obtain the first spatial domain signal, and to calculate the initial phase difference information of the second image sequence after spatial domain decomposition. The reconstruction module is used to resample and reconstruct the first spatial domain signal to obtain a third image sequence of the second spatial domain signal; An enhancement processing module is used to process the third image sequence of the reconstructed second spatial domain signal using reweighted residual sparsity, so as to enhance the temporal correlation of the third image sequence of the reconstructed second spatial domain signal. The acquisition module is used to obtain structural vibration information based on the phase difference information in the processed third image sequence and by using conversion calculation.
Citation Information
Patent Citations
Structure tiny vibration measurement method and system based on broadband phase motion amplification
CN114993452A
Method and device for measuring displacement of machine vision structure
CN116576781A