Structure high-frequency dynamic characteristic identification method based on deep learning and video processing
By using DFC-3DNet network for self-supervised training and constructing high frame rate self-supervised learning pairs using low frame rate videos, the problem of insufficient sampling rate in engineering monitoring videos is solved, and high-precision high-frequency dynamic characteristic recognition and reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, the sampling rate of engineering monitoring videos is insufficient, resulting in a spatiotemporal sampling rate lower than the Nyquist rate. This makes it impossible to fully capture the high-frequency vibration characteristics of rigid components such as bridge towers, thus affecting the accuracy of dynamic characteristic recognition.
A method based on deep learning and video processing is adopted. The DFC-3DNet network is used for self-supervised training. Low frame rate-high frame rate self-supervised learning pairs are constructed using low frame rate videos to generate high frame rate videos and realize high frequency dynamic feature recognition.
It improves the peak signal-to-noise ratio of video reconstruction by about 6dB, significantly enhances the structural similarity index, accurately identifies higher-order mode frequencies that exceed the original sampling frequency limit, solves the mode aliasing problem, and ensures reliable extraction of high-frequency dynamic characteristics.
Smart Images

Figure CN121904547A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, structural health monitoring and artificial intelligence, and in particular to a method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing. Background Technology
[0002] With the increasing service life of large-scale engineering structures such as bridges and buildings, structural health monitoring (SHM) plays a crucial role in ensuring structural safety and extending their lifespan. Compared to contact sensors, non-contact measurement methods based on computer vision offer advantages such as ease of deployment and the ability to acquire information across the entire field, thus finding widespread application in the identification and monitoring of structural dynamic characteristics.
[0003] However, due to limitations in data storage, transmission bandwidth, and computing power, engineering monitoring videos often suffer from insufficient sampling rates. Their spatiotemporal sampling rates are lower than the Nyquist rate, making it impossible to fully capture the high-frequency vibration characteristics of rigid components such as bridge towers and piers. This can easily lead to signal distortion and mode mixing, affecting the accuracy of dynamic characteristic recognition.
[0004] The information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] This invention provides a method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing, which can solve the technical problem that related technologies cannot guarantee the accuracy of dynamic characteristic identification.
[0006] According to a first aspect of the present invention, a method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing is provided, comprising:
[0007] Acquire the raw low frame rate video and the video to be processed;
[0008] Based on the original low frame rate video, obtain training sample data;
[0009] Based on the original low frame rate video, construct a low frame rate-high frame rate self-supervised learning pair;
[0010] Based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, the DFC-3DNet network is self-supervised trained to obtain the trained DFC-3DNet network.
[0011] The trained DFC-3DNet network is subjected to performance testing, and the video to be processed is processed using the trained DFC-3DNet network to obtain the reconstruction result.
[0012] According to the present invention, obtaining training sample data based on the original low frame rate video includes:
[0013] Based on the original low frame rate video, obtain spatiotemporal video blocks;
[0014] New video slices were determined by swapping the spatial temperature and time dimensions;
[0015] The spatiotemporal video block is processed to generate the first training sample data;
[0016] Augmentation operations are performed on the first sample data to obtain training sample data.
[0017] According to the present invention, a DFC-3DNet network is self-supervised trained based on the low frame rate-high frame rate self-supervised learning pair and the training sample data to obtain a trained DFC-3DNet network, comprising:
[0018] Determine the network architecture of the DFC-3DNet network;
[0019] Determine the activation functions for the fully connected layers of the DFC-3DNet network;
[0020] Determine the forward propagation formula for the 3D convolutional layers of the DFC-3DNet network;
[0021] Determine the residual learning strategy;
[0022] Determine the combined loss function for the DFC-3DNet network;
[0023] The DFC-3DNet network is trained according to the combined loss function to obtain the trained DFC-3DNet network.
[0024] According to the present invention, determining the activation function of the fully connected layer of the DFC-3DNet network includes: according to the formula:
[0025]
[0026] Determine the activation functions for the fully connected layers of the DFC-3DNet network, where, y is the output of the first hidden layer, and y is the two-dimensional input after video compression-sensory encoding. The weight matrix consists of linear filters. For bias terms, It is a non-linear activation function.
[0027] According to the present invention, determining the forward propagation formula for the three-dimensional convolutional layer of the DFC-3DNet network includes: according to the formula:
[0028]
[0029] Determine the forward propagation formula for the 3D convolutional layers of the DFC-3DNet network, where, For a 3D convolution kernel, k=2,…,8.
[0030] According to the present invention, determining a residual learning strategy includes:
[0031] Use the interpolated video as the baseline input;
[0032] The DFC-3DNet network learns the residual between the benchmark input and the real high-resolution video.
[0033] According to the present invention, determining the combined loss function of the DFC-3DNet network includes: according to the formula:
[0034]
[0035] Determine the combined loss function of the DFC-3DNet network, where, , It includes all the weights and biases of the model, weights =0.55、 =0.45.
[0036] According to a second aspect of the present invention, a structural high-frequency dynamic characteristic recognition system based on deep learning and video processing is provided, comprising:
[0037] The video acquisition module is used to acquire the raw low frame rate video and the video to be processed;
[0038] The sample generation module is used to obtain training sample data based on the original low frame rate video.
[0039] The learning generation module is used to construct low frame rate-high frame rate self-supervised learning pairs based on the original low frame rate video.
[0040] The network training module is used to perform self-supervised training on the DFC-3DNet network based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, so as to obtain the trained DFC-3DNet network.
[0041] The video reconstruction module is used to perform performance testing on the trained DFC-3DNet network and process the video to be processed using the trained DFC-3DNet network to obtain the reconstruction result.
[0042] Technical Results: According to the present invention, test results based on physical graphical model simulation data show that the video sequences generated by the present invention maintain stable reconstruction results even under complex interference environments. Even with low input quality, high-fidelity video reconstruction and more accurate detail recovery can be achieved. Specifically, the peak signal-to-noise ratio is improved by up to approximately 6 dB compared to the contrast model FC-7 (trained using a seven-layer fully connected video compression and sensing network DFCNet), the structural similarity index is also significantly improved, and low frame rate videos can be effectively reconstructed into high frame rate videos. Furthermore, it accurately identifies higher-order modal frequencies exceeding the original sampling frequency limit, thereby effectively solving the modal aliasing problem and ensuring reliable extraction of high-frequency dynamic characteristics.
[0043] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Other features and aspects of the invention will become clearer from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0045] Figure 1 An exemplary flowchart of a structural high-frequency dynamic characteristic recognition method based on deep learning and video processing according to an embodiment of the present invention is shown.
[0046] Figure 2 An exemplary diagram illustrates the overall algorithm framework of a high-frequency dynamic characteristic recognition method for structures based on deep learning and video processing according to an embodiment of the present invention;
[0047] Figure 3 An exemplary schematic diagram of cross-dimensional recursion of small video spatiotemporal blocks according to an embodiment of the present invention is shown;
[0048] Figure 4 An exemplary diagram illustrating the generation process of a training example according to an embodiment of the present invention is shown;
[0049] Figure 5 An exemplary flowchart of a PBGM-based numerical simulation test according to an embodiment of the present invention is shown;
[0050] Figures 6(a), (b) and (c) exemplarily illustrate actual bridge test diagrams according to embodiments of the present invention;
[0051] Figures 7(a), (b), (c) and (d) exemplarily illustrate qualitative evaluation comparison diagrams of reconstruction results under different motion fuzzing conditions according to embodiments of the present invention;
[0052] Figures 8(a), (b) and (c) exemplarily illustrate quantitative evaluation comparisons of reconstruction results under different motion fuzzing conditions according to embodiments of the present invention;
[0053] Figures 9(a), (b), (c) and (d) exemplarily illustrate qualitative evaluation comparisons of reconstruction results under different noise conditions according to embodiments of the present invention;
[0054] Figures 10(a), (b), (c) and (d) exemplarily illustrate comparison diagrams of reconstructed video dynamic characteristic recognition based on full-field visual measurement according to embodiments of the present invention;
[0055] Figures 11(a), (b), (c), (d), and (e) exemplarily illustrate comparisons of actual bridge test results according to embodiments of the present invention;
[0056] Figure 12 An exemplary block diagram of a structural high-frequency dynamic characteristic recognition method based on deep learning and video processing according to an embodiment of the present invention is shown;
[0057] Figure 13 An exemplary comparison chart of the average running time of the model under different operating conditions according to embodiments of the present invention is shown;
[0058] Figure 14 An exemplary comparison chart of modal frequency identification results according to an embodiment of the present invention is shown. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The technical solution of the present invention will be described in detail below with reference to specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0061] Figure 1 An exemplary flowchart illustrates a method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to an embodiment of the present invention. The method includes:
[0062] Step S1: Obtain the original low frame rate video and the video to be processed;
[0063] Step S2: Obtain training sample data based on the original low frame rate video;
[0064] Step S3: Construct a low frame rate-high frame rate self-supervised learning pair based on the original low frame rate video;
[0065] Step S4: Based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, perform self-supervised training on the DFC-3DNet network to obtain the trained DFC-3DNet network.
[0066] Step S5: Perform performance testing on the trained DFC-3DNet network, and process the video to be processed using the trained DFC-3DNet network to obtain the reconstruction result.
[0067] According to embodiments of the present invention, the high-frequency dynamic characteristic recognition method based on deep learning and video processing demonstrates, based on physical graphical model simulation data, that the video sequences generated by the present invention maintain stable reconstruction performance even under complex interference environments. Even with low input quality, it achieves high-fidelity video reconstruction and more accurate detail recovery. Specifically, it improves the peak signal-to-noise ratio by approximately 6 dB compared to the contrast model FC-7 (trained using a seven-layer fully connected video compression sensing network DFCNet), significantly improves the structural similarity index, effectively reconstructs low-frame-rate videos into high-frame-rate videos, and accurately identifies higher-order modal frequencies exceeding the original sampling frequency limit, thereby effectively solving the modal aliasing problem and ensuring reliable extraction of high-frequency dynamic characteristics.
[0068] According to one embodiment of the present invention, in step S1, the original low frame rate video and the video to be processed are acquired.
[0069] For example, a low-frame-rate, low-resolution structural surveillance video can be used as input data (i.e., the original low-frame-rate video) without any external annotation data.
[0070] According to one embodiment of the present invention, in step S2, training sample data is obtained based on the original low frame rate video.
[0071] According to an embodiment of the present invention, step S2 includes:
[0072] Step S21: Obtain spatiotemporal video blocks based on the original low frame rate video;
[0073] Step S22: Determine a new video slice by exchanging the spatial temperature and time dimensions;
[0074] Step S23: Process the spatiotemporal video block to generate the first training sample data;
[0075] Step S24: Perform augmentation operations on the first sample data to obtain training sample data.
[0076] For example, a sliding window approach is used to densely extract spatiotemporal video blocks of size 8×8×8 (width×height×frame) from the input video for subsequent training sample construction. By swapping the spatial dimension (x, y) and the temporal dimension (t), new video slices are constructed, utilizing the recursiveness of fast and slow moving objects in the xy, xt, yt planes to improve the model's adaptability to different motion patterns. Spatial downsampling (e.g., bicubic downsampling), temporal downsampling (e.g., frame averaging to simulate motion blur), and uniform scaling are performed on the video blocks to generate samples with multiple resolutions and blur levels to improve the model's robustness. Further enhancement operations such as mirror flipping, rotation (90°, 180°, 270°), and temporal flipping are performed to increase sample diversity and avoid model overfitting. The enhanced video blocks are used as low-quality inputs, and the corresponding original video blocks are used as high-quality outputs to form self-supervised training sample pairs (low resolution-high resolution, low frame rate-high frame rate, low blur-high blur) for subsequent network training.
[0077] According to an embodiment of the present invention, in step S3, a low frame rate-high frame rate self-supervised learning pair is constructed based on the original low frame rate video.
[0078] For example, small-scale spatiotemporal video blocks can be extracted from input low frame rate videos and expanded through cross-dimensional transformations (e.g., spatial-temporal axis interchange) and various data augmentation strategies (e.g., downsampling, rotation, flipping, etc.) to construct a large number of self-supervised training sample pairs for specific videos, including low-resolution-high-resolution sample pairs and low-motion-blur-high-motion-blur sample pairs generated by dimensional transformation, thereby achieving training data generation without relying on external datasets.
[0079] According to an embodiment of the present invention, in step S4, the DFC-3DNet network is self-supervised trained based on the low frame rate-high frame rate self-supervised learning pair and the training sample data to obtain the trained DFC-3DNet network.
[0080] According to an embodiment of the present invention, step S4 includes:
[0081] Step S41: Determine the network architecture of the DFC-3DNet network;
[0082] Step S42: Determine the activation functions of the fully connected layers of the DFC-3DNet network;
[0083] Step S43: Determine the forward propagation formula for the 3D convolutional layers of the DFC-3DNet network;
[0084] Step S44: Determine the residual learning strategy;
[0085] Step S45: Determine the combined loss function of the DFC-3DNet network;
[0086] Step S46: Train the DFC-3DNet network according to the combined loss function to obtain the trained DFC-3DNet network.
[0087] For example, DFC-3DN uses the interpolated video as the baseline input. The network learns the residual between the interpolated video and the real high-resolution video, reducing the learning difficulty and improving the convergence speed. It uses a 3D tensor for subsequent convolutional processing, with ReLU as the activation function. The 3D convolutional layers use 3×3×3 and 1×3×3 kernels with a stride of 1, 128 channels per layer, for a total of eight layers, used for depth extraction and spatiotemporal feature reconstruction. This loss function can balance the tasks of compressed sensing reconstruction and temporal super-resolution reconstruction, adjusting the model's weights and bias parameters. Backpropagation updates are performed to optimize the network in both spatial and temporal dimensions. A DFC-3DNet model is built using the deep learning framework (PyTorch). An internal training set is generated from the input video using the method described above. The optimizer is set to Stochastic Gradient Descent (SGD) with a momentum of 0.9 and an initial learning rate of 0.01. The optimization rate is increased every 3×... The next iteration reduces the learning rate by a factor of 10 and the batch size to 200. A combined loss function is used. =0.55, Training was performed with a value of 0.45, and the total number of iterations was approximately 4× Second-rate.
[0088] According to an embodiment of the present invention, step S42 includes: determining the activation function of the fully connected layer of the DFC-3DNet network according to formula (1).
[0089] (1)
[0090] in, y is the output of the first hidden layer, and y is the two-dimensional input after video compression-sensory encoding. The weight matrix consists of linear filters. For bias terms, It is a non-linear activation function.
[0091] According to one embodiment of the present invention, It is a nonlinear activation function, specifically using a rectified linear unit (ReLU), which is defined as follows: .
[0092] According to an embodiment of the present invention, step S43 includes: determining the forward propagation formula of the three-dimensional convolutional layer of the DFC-3DNet network according to formula (2).
[0093] (2)
[0094] in, For a 3D convolution kernel, k=2,…,8.
[0095] According to one embodiment of the present invention, It is a three-dimensional convolutional kernel used to extract features simultaneously in the spatial and temporal dimensions.
[0096] According to an embodiment of the present invention, step S44 includes:
[0097] Step S441: Use the interpolated video as the reference input;
[0098] Step S442, the DFC-3DNet network learns the residual between the benchmark input and the real high-resolution video.
[0099] (3)
[0100] According to an embodiment of the present invention, step S45 includes: determining the combined loss function of the DFC-3DNet network according to formula (3),
[0101] in, , It includes all the weights and biases of the model, weights =0.55、 =0.45.
[0102] According to one embodiment of the present invention, a weighted mean square error (MSE) loss function is adopted because MSE is directly related to the peak signal-to-noise ratio (PSNR), a video quality evaluation metric. Minimizing MSE can effectively improve the PSNR of video reconstruction, thereby ensuring the detail accuracy and overall quality of the reconstructed video.
[0103] According to an embodiment of the present invention, in step S5, the performance of the trained DFC-3DNet network is tested, and the video to be processed is processed by the trained DFC-3DNet network to obtain the reconstruction result.
[0104] For example, the first part of the performance testing was a simulation test. In this simulation test, to verify the effectiveness of the proposed DFC-3DNet, a bridge tower numerical simulation environment was constructed based on the Physical-based Graphics Model (PBGM), generating videos at different resolutions and frame rates to reproduce the high-frequency vibration process. Based on this, the model's reconstruction performance was comprehensively evaluated, mainly including image reconstruction quality, modal aliasing suppression, and the accuracy of dynamic characteristic reconstruction and recognition. At the metrics level, image reconstruction quality was quantified using Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM), while modal aliasing and dynamic characteristic recognition accuracy were evaluated using existing research methods. In terms of comparison, DFC-3DNet was only compared with a deep fully connected network for video compressed sensing (DFCNet). This is because DFC-3DNet integrates the advantages of video compressed sensing and temporal super-resolution on top of DFCNet, enabling higher-precision reconstruction in both spatial and temporal dimensions. To obtain reliable comparative verification data, a highly realistic bridge tower simulation environment was constructed based on the Physical-based Graphics Model (PBGM), such as... Figure 5As shown. PBGM calculates structural deformation through finite element analysis and renders it using the Blender graphics engine, thereby generating a high-quality video consistent with the actual engineering scene. The experiment uses the Xiaoxihu pedestrian overpass as the background and uses synthetic cameras to achieve global video acquisition, which is used as input data for DFC-3DNet to evaluate the network's reconstruction performance and vibration measurement effect. In the experimental design, the generation and model training of PBGM fully consider the super-resolution requirements of both time and space: (1) Time dimension: Vibration data obtained by finite element step-by-step calculation is extracted through the Python interface of Abaqus, and reference videos of 25FPS, 50FPS and 100FPS are generated as comparison data at different frame rates. The finite element analysis frequency is limited to below 50 Hz, and the 100FPS reference video can completely cover the selected mode without aliasing. (2) Spatial dimension: In order to simulate the differences of different acquisition devices (such as network cameras, industrial cameras and SLR cameras), multiple resolution videos such as 720P, 1080P, 2K and 4K are rendered using the synthetic environment. The second part of the performance test is the real bridge test. In the real bridge test, in order to further verify the performance of DFC-3DNet in the real environment, a real bridge test was carried out on the pedestrian bridge in Xiaoxihu Park, Lanzhou. (1) Overview of the bridge background: The bridge is a single-tower cable-stayed bridge (without backstays), with a total length of 44.1 m. The main beam adopts a box section with a height of 0.5 m and an inclined web. The bridge tower is a rectangular section with a height of 11.5 m (above the bridge deck). The overall shape is like a "swan spreading its wings". The cable stays are UU-shaped stainless steel cables with a diameter of 50 mm. The actual scene and structural dimensions are shown in Figure 6(a) and (b). (2) Test plan: ① Video acquisition: Canon EOS 5D Mark IV camera, sampling rate 25Hz, resolution 4K, global shooting of the bridge tower; the camera is arranged at a distance of 16.5m and the horizontal angle is about 20° (Figure 6(c)). ②Acceleration test: A contact-type high-frequency accelerometer (CT1050LCIEPE) with a sampling rate of 100Hz, sensitivity of 500mV / g, and range of 10g is installed in the middle of the bridge tower to obtain accurate acceleration response.
[0105] Furthermore, to fully verify the performance advantages of the proposed DFC-3DNet network, a multi-dimensional comparative experiment was conducted on the video generated by the Physical Graphics Model (PBGM), covering aspects such as motion blur suppression, noise robustness, inference efficiency, dynamic feature reconstruction and modal aliasing suppression, and compared with the traditional FC-7 model (trained by the DFCNet network). (1) Motion blur suppression capability, 1) Experimental settings: The blur intensity was controlled by setting different shutter times (0.2s, 0.5s, 1.0s) in Blender, and the reconstruction results were compared. 2) Qualitative evaluation (Figure 7): DFC-3DNet outperformed FC-7 under different blur conditions, proving its significant anti-blur capability. 3) Quantitative evaluation (Figure 8): Under mild blur (0.2s), DFC-3DNet is basically unaffected, while FC-7 performance drops significantly; under moderate blur (0.5s), DFC-3DNet performance improves with resolution, while FC-7 is more severely damaged; under severe blur (1.0s), DFC-3DNet is affected but still maintains an improving trend, while FC-7 almost fails; (2) Noise robustness verification, 1) Experimental setup: Gaussian noise with different signal-to-noise ratios (SNR=5, 10, 20) is superimposed on the video, and the qualitative evaluation of the reconstruction results is shown in Figure 9. 2) Result comparison: FC-7 is limited by the fully connected structure and has difficulty effectively suppressing noise; DFC-3DNet relies on the spatiotemporal feature extraction capability of 3D convolution to effectively alleviate the impact of noise; 3) Conclusion: The overall reconstruction quality is better under noisy conditions than under blurred conditions, indicating that the information loss caused by blur is more serious than that caused by noise. At the same time, the performance gap between the two models narrows, reflecting the universal robustness of DFC-3DNet. (3) Reasoning efficiency evaluation 1) Results ( Figure 13 ): The inference time of both models increases with increasing resolution, with DFC-3DNet being slightly higher than FC-7, but in return for a significant quality improvement. 2) Conclusion: DFC-3DNet achieves a reasonable balance between computational cost and reconstruction accuracy. (4) Dynamic feature reconstruction and modal aliasing suppression, 1) Nyquist constraint: It can identify modal frequencies below 50Hz at 100Hz sampling. Obvious high-order modal aliasing appears in the 25Hz and 50Hz videos (Fig. 10(a), (b)). 2) Super-resolution results: DFC-3DNet successfully reconstructs the 25Hz video into a 50Hz video, and its frequency and energy distribution are highly consistent with the 100Hz video (Fig. 10(c), (d)), which verifies the effective mitigation of modal aliasing.
[0106] Furthermore, for the actual bridge test results, the original 25Hz video was first segmented (see Figure 11(a)), and then full-field visual measurement was performed to plot the maximum displacement cloud map (see Figure 11(b)). Since the 25Hz video exhibits modal aliasing and energy distribution distortion in frequency analysis (see Figure 11(c)), a spatiotemporal super-resolution model based on DFC-3DNet was trained with parameters t=2 and t=4 respectively to reconstruct the video, obtaining two enhanced results at 50Hz and 100Hz. The time-frequency recognition results are shown in Figure 11(d), demonstrating that the super-resolution method effectively alleviates the modal aliasing problem and maintains good consistency with the inertial accelerometer recognition results (Figure 11(e)) in terms of high-order modal frequency recovery. The relative error is shown in Figure 11(d). Figure 14 .
[0107] Furthermore, after generating training samples and completing self-supervised network training and performance verification, the system can directly output reconstruction results for the input video, obtaining the corresponding high frame rate, high resolution video sequence. This reconstruction result not only improves the spatiotemporal resolution of the video but can also serve as input for subsequent full-field visual measurements and signal analysis, used to identify the dynamic characteristics of structures and extract high-frequency modal information.
[0108] According to embodiments of the present invention, the high-frequency dynamic characteristic recognition method based on deep learning and video processing demonstrates, based on physical graphical model simulation data, that the video sequences generated by the present invention maintain stable reconstruction performance even under complex interference environments. Even with low input quality, it achieves high-fidelity video reconstruction and more accurate detail recovery. Specifically, it improves the peak signal-to-noise ratio by approximately 6 dB compared to the contrast model FC-7 (trained using a seven-layer fully connected video compression sensing network DFCNet), significantly improves the structural similarity index, effectively reconstructs low-frame-rate videos into high-frame-rate videos, and accurately identifies higher-order modal frequencies exceeding the original sampling frequency limit, thereby effectively solving the modal aliasing problem and ensuring reliable extraction of high-frequency dynamic characteristics.
[0109] Figure 12 An exemplary block diagram of a high-frequency dynamic characteristic recognition system for structures based on deep learning and video processing according to an embodiment of the present invention is shown, the system comprising:
[0110] The video acquisition module is used to acquire the raw low frame rate video and the video to be processed;
[0111] The sample generation module is used to obtain training sample data based on the original low frame rate video.
[0112] The learning generation module is used to construct low frame rate-high frame rate self-supervised learning pairs based on the original low frame rate video.
[0113] The network training module is used to perform self-supervised training on the DFC-3DNet network based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, so as to obtain the trained DFC-3DNet network.
[0114] The video reconstruction module is used to perform performance testing on the trained DFC-3DNet network and process the video to be processed using the trained DFC-3DNet network to obtain the reconstruction result.
[0115] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0116] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and explained in the embodiments, and any variations or modifications may be made to the implementation of the present invention without departing from the stated principles.
Claims
1. A method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing, characterized in that, include: Acquire the raw low frame rate video and the video to be processed; Based on the original low frame rate video, obtain training sample data; Based on the original low frame rate video, construct a low frame rate-high frame rate self-supervised learning pair; Based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, the DFC-3DNet network is self-supervised trained to obtain the trained DFC-3DNet network. The trained DFC-3DNet network is subjected to performance testing, and the video to be processed is then processed using the trained DFC-3DNet network to obtain the reconstruction result.
2. The method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to claim 1, characterized in that, Based on the original low frame rate video, training sample data is obtained, including: Based on the original low frame rate video, obtain spatiotemporal video blocks; New video slices were determined by swapping the spatial temperature and time dimensions; The spatiotemporal video block is processed to generate the first training sample data; Augmentation operations are performed on the first sample data to obtain training sample data.
3. The method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to claim 1, characterized in that, Based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, the DFC-3DNet network is self-supervised trained to obtain a trained DFC-3DNet network, including: Determine the network architecture of the DFC-3DNet network; Determine the activation functions for the fully connected layers of the DFC-3DNet network; Determine the forward propagation formula for the 3D convolutional layers of the DFC-3DNet network; Determine the residual learning strategy; Determine the combined loss function for the DFC-3DNet network; The DFC-3DNet network is trained according to the combined loss function to obtain the trained DFC-3DNet network.
4. The method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to claim 3, characterized in that, Determine the activation functions of the fully connected layers in the DFC-3DNet network, including: according to the formula: Determine the activation functions for the fully connected layers of the DFC-3DNet network, where, y is the output of the first hidden layer, and y is the two-dimensional input after video compression-sensory encoding. The weight matrix consists of linear filters. For bias terms, It is a non-linear activation function.
5. The method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to claim 3, characterized in that, Determine the forward propagation formula for the 3D convolutional layers of the DFC-3DNet network, including: based on the formula: Determine the forward propagation formula for the 3D convolutional layers of the DFC-3DNet network, where, For a 3D convolution kernel, k=2,…,8.
6. The method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to claim 3, characterized in that, Determine the residual learning strategy, including: Use the interpolated video as the baseline input; The DFC-3DNet network learns the residual between the benchmark input and the real high-resolution video.
7. The method for identifying high-frequency dynamic characteristics of structures based on deep learning and video processing according to claim 3, characterized in that, Determine the combined loss function of the DFC-3DNet network, including: according to the formula: Determine the combined loss function of the DFC-3DNet network, where, , It includes all the weights and biases of the model, weights =0.55、 =0.
45.
8. A structural high-frequency dynamic characteristic recognition system based on deep learning and video processing, characterized in that, include: The video acquisition module is used to acquire the raw low frame rate video and the video to be processed; The sample generation module is used to obtain training sample data based on the original low frame rate video. The learning generation module is used to construct low frame rate-high frame rate self-supervised learning pairs based on the original low frame rate video. The network training module is used to perform self-supervised training on the DFC-3DNet network based on the low frame rate-high frame rate self-supervised learning pair and the training sample data, so as to obtain the trained DFC-3DNet network. The video reconstruction module is used to perform performance testing on the trained DFC-3DNet network and process the video to be processed using the trained DFC-3DNet network to obtain the reconstruction result.