Computer vision-based intelligent non-contact structural displacement detection method and system

By improving the YOLOv11 and StrongSORT algorithms, and combining the frequency domain-time domain fusion module and the inverse perspective transformation module, the shortcomings of existing contact and non-contact displacement detection technologies have been solved, realizing high-precision and stable long-distance multi-target structural displacement detection, especially efficient displacement monitoring in complex and dynamic environments.

CN121095167BActive Publication Date: 2026-04-28TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2025-08-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, contact displacement detection methods interfere with the structure itself and lack durability, while non-contact detection methods have low accuracy and low sampling frequency, making it difficult to achieve high-precision long-distance multi-target displacement detection, especially in complex environments and dynamic conditions.

Method used

An improved YOLOv11 algorithm and StrongSORT algorithm are adopted, combined with a frequency domain-time domain fusion module and an inverse perspective transformation module to construct a structural displacement analysis algorithm. Through multi-scale attention feature extraction and dynamic upsampling module, high-precision multi-target displacement detection is achieved, and image distortion caused by camera vibration and angle changes is corrected.

Benefits of technology

It significantly improves the detection accuracy and system stability of small targets, and enables long-distance real-time detection of multi-target structural displacement in complex backgrounds and dynamic environments, thereby improving detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095167B_ABST
    Figure CN121095167B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent non-contact structural displacement detection method and system based on computer vision, belong to the field of deep learning.The method comprises: obtaining the first video image containing displacement target, establishing the first detection model based on YOLOv11 algorithm, using the first detection model to analyze and process the first video image, obtain the detection result of displacement target;Based on StrongSORT algorithm, construct information network according to the appearance features of multiple displacement targets, input the detection result into the information network, form the motion trajectory of multiple displacement targets;Based on the structure displacement analysis algorithm constructed in combination with frequency domain-time domain fusion module and inverse perspective transformation module, the actual structure displacement of displacement target is calculated based on the mapping relationship between target actual size and pixel coordinate;Based on the structure displacement analysis algorithm, the displacement change of target structure is detected in real time by the displacement calculation of continuous frames.The application realizes non-contact long-distance multi-target structure displacement real-time detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning, specifically relating to an intelligent non-contact structural displacement detection method and system based on computer vision. Background Technology

[0002] Structural displacement is a crucial indicator for assessing structural mechanical properties and damage evolution in structural health monitoring. Current structural displacement detection methods can be categorized into contact and non-contact methods. Contact monitoring primarily relies on devices such as wire-type displacement sensors, capacitive displacement sensors, and accelerometers. These devices suffer from drawbacks, including excessive interference with the structure during installation and sensor susceptibility to environmental corrosion leading to insufficient durability. Non-contact detection employs technologies such as laser vibrometers, total stations, and microwave radar. However, these methods are hampered by high cost, low accuracy, and low sampling frequency, making it difficult to achieve high-frequency continuous detection of multiple targets over long distances.

[0003] Computer vision technology has provided a new paradigm for structural displacement monitoring. Subpixel-level displacement analysis is achieved through template matching based on feature point detection. Deep learning technology significantly improves measurement accuracy and scene adaptability through feature self-extraction, 3D displacement reconstruction, and end-to-end displacement prediction. Patent CN202411443228.8 discloses an intelligent non-contact structural displacement detection method based on computer vision, including the following steps: acquiring a first image containing a displacement target; establishing a first recognition model based on the YOLOv7 algorithm; processing the first image using the first recognition model to obtain the recognition result of the displacement target; constructing a tracking network for the displacement target based on the DeepSORT algorithm; inputting the recognition result into the tracking network to form the motion trajectory of the displacement target; constructing a structural displacement analysis network based on the IPM-ED method; analyzing the motion trajectory to generate the actual structural displacement of the displacement target; forming a first time history curve of the displacement target based on the actual structural displacement; and comparing the first time history curve with a first measurement curve to obtain the structural displacement detection effect.

[0004] However, the above method has the following problems:

[0005] First, the YOLOv7 and DeepSORT algorithms used in the above methods are relatively outdated and cannot meet the requirements for efficient and high-precision target detection and tracking. They perform poorly, especially in complex environments and with small targets, easily resulting in missed detections and poor overall tracking performance. Second, the method has limited processing capabilities for low-resolution images, making it difficult to effectively improve the detection accuracy and detail of distant targets. Furthermore, in dynamic environments, especially when load vibrations cause high-frequency camera vibrations, it fails to effectively address the problem of inaccurate displacement detection caused by camera vibrations, leading to deviations in the target's displacement trajectory in the image and affecting the overall detection and tracking accuracy. Therefore, the above methods suffer from high missed detection rates and low positioning accuracy in practical applications, making it difficult to meet the requirements for precise displacement detection. Summary of the Invention

[0006] To address the aforementioned problems, this invention provides a non-contact, long-distance, multi-target structural displacement detection method based on computer vision, thereby resolving the issues in the prior art.

[0007] To achieve the aforementioned objectives, this invention proposes a non-contact, long-range, multi-target structural displacement detection method based on computer vision, comprising:

[0008] A first video image containing a displacement target is acquired, a first detection model is established based on the improved YOLOv11 algorithm, and the first detection model is used to process the first video image to obtain the detection result of the displacement target.

[0009] Based on the StrongSORT algorithm, an information network is constructed according to the appearance features of various displacement targets. The detection results are input into the information network to form the motion trajectory of various displacement targets.

[0010] A structural displacement analysis algorithm is constructed by combining a frequency domain-time domain fusion module and an inverse perspective transformation module. The structural displacement analysis algorithm calculates the actual structural displacement of the target based on the mapping relationship between the actual size of the target and the pixel coordinates.

[0011] Based on the structural displacement analysis algorithm, displacement changes of the target structure are detected in real time through displacement calculation of consecutive frames. Furthermore, the first detection model based on the improved YOLOv11 algorithm includes the following steps:

[0012] The first detection model includes a backbone layer, a neck layer, and a head layer. Based on the spatial-depth convolution module in the backbone layer, fine-grained first features are extracted from the first video image. The first features are coupled across stages through a feature extraction module. The first features are aggregated at multiple scales by replacing the bottleneck structure with an embedded multi-scale attention feature extraction module. The dynamic upsampling module in the neck layer performs multi-scale fusion on the aggregated first feature information to obtain a first fused feature. Then, a detection head added to the head layer is used to construct a four-layer cascaded feature pyramid to process the first fused feature and obtain the detection result of the displacement target.

[0013] Furthermore, constructing the tracking network includes the following steps:

[0014] Based on the detection results, the StrongSORT algorithm obtains the first feature vectors of the various displacement targets, corrects the inter-frame offset of the first video image through a motion compensation mechanism that maximizes the correlation coefficient, predicts the motion state of the displacement targets in the first video image through a gating mechanism based on the NSA Kalman algorithm, assigns IDs to the various displacement targets, and the StrongSORT algorithm re-associates short-term occlusion trajectories based on an appearance-free linking model. It also completes the missing frame positions of the short-term occlusions using a Gaussian smoothing interpolator and outputs continuous motion trajectories.

[0015] Furthermore, the structural displacement analysis algorithm constructed based on the combination of the frequency-time domain fusion module and the inverse perspective transformation module includes the following steps:

[0016] The virtual position information of the displacement target is obtained based on the motion trajectory. The virtual position information of the displacement target is corrected based on the combination of the frequency domain-time domain fusion module and the inverse perspective transformation module. The virtual position information is converted into the actual structural position information based on the scaling factor through the mapping relationship between the actual size of the target and the pixel coordinates. The actual structural displacement of the displacement target is calculated using Euclidean distance.

[0017] Furthermore, the correction based on the combination of the frequency-time domain fusion module and the inverse perspective transformation module includes the following steps:

[0018] The first video image features are extracted based on the frequency-time domain fusion module, the features are mapped to the virtual location information, the virtual location information is unified using the inverse perspective transformation module, and the correction is completed based on the virtual location information and the mapping relationship.

[0019] Furthermore, the structural displacement analysis algorithm calculates structural displacement based on the mapping relationship between the actual size of the target and the pixel coordinates. This mapping relationship is obtained by calibrating the physical size of the target and the mapping of pixels in the image.

[0020] Furthermore, displacement changes of the target structure are calculated and detected through continuous frames, wherein the continuous frames include at least three frames in the first video image, and the monitoring generates real-time displacement data by analyzing the motion trajectory of the displacement target in each frame and combining it with the actual structural displacement of the displacement target.

[0021] Furthermore, the real-time displacement curve is generated based on the displacement calculation in the continuous frames and is updated in real time. The real-time update is achieved by tracking the motion trajectory of the displacement target in each frame image to generate the displacement change trend curve of the target structure.

[0022] Furthermore, the detection results include category information and confidence scores.

[0023] Secondly, the present invention also provides an intelligent non-contact structural displacement detection system based on computer vision, comprising:

[0024] The first module is configured to acquire a first video image containing a displacement target, establish a first detection model based on the YOLOv11 algorithm, and use the first detection model to analyze and process the first video image to obtain the detection result of the displacement target.

[0025] The second module is configured to construct an information network based on the StrongSORT algorithm according to the appearance features of various displacement targets, input the detection results into the information network, and form the motion trajectory of various displacement targets.

[0026] The third module is configured to construct a structural displacement analysis algorithm based on the combination of the frequency domain-time domain fusion module and the inverse perspective transformation module. The structural displacement analysis algorithm calculates the actual structural displacement of the displacement target based on the mapping relationship between the actual size of the target and the pixel coordinates.

[0027] The fourth module is configured to detect the displacement changes of the target structure in real time by calculating the displacement of continuous frames based on a structural displacement analysis algorithm.

[0028] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0029] First, this invention employs improved YOLOv11 and StrongSORT algorithms. The former achieves better performance in target detection accuracy and speed by introducing the SPD-Conv module to replace the convolution module of YOLOv11, introducing the MAB module to replace the bottleneck structure of the C3K2 module in the YOLOv11 algorithm, and introducing the Dy_Sample module to replace the upsampling module in the YOLOv11 algorithm. The latter provides higher stability and accuracy in target tracking, especially in the detection of complex backgrounds and small targets, and has significant advantages over YOLOv7 and DeepSORT. Secondly, this invention addresses the issues of recognition errors and missed detections in structural displacement target detection under complex environments, as well as the target edge blurring caused by low-resolution imaging, by replacing the bottleneck structure of the C3K2 module in the YOLOv11 algorithm with an embedded multi-scale attention feature extraction module. In the embedded multi-scale attention feature extraction module (C3K2_MAB), the input features are first processed by layer normalization, and then long-distance and local information are captured by the large kernel convolution decomposition of the multi-scale large kernel attention (MLKA) module. The aggregation of spatial information is further optimized by the gate spatial attention unit (GSAU), thereby effectively reducing the number of parameters and improving the model's performance in super-resolution tasks.

[0030] To address image distortion caused by camera vibration and changes in shooting angle, this invention proposes an innovative structural displacement analysis method that combines a frequency domain and time domain fusion module with an inverse perspective module, improving the accuracy and robustness of displacement detection. This invention effectively avoids the accuracy loss caused by vibration in the original method, improving the system's adaptability and stability in dynamic environments, especially in high-vibration and complex backgrounds, accurately capturing the target's displacement trajectory. The improvements not only enhance the accuracy of small target detection but also optimize the overall system's stability and anti-interference capabilities, significantly improving the accuracy of displacement tracking.

[0031] This invention addresses the problems of excessive interference with the structure itself and insufficient durability in contact displacement detection methods, as well as the low accuracy, low sampling frequency, and sensitivity to distance changes in non-contact displacement detection methods. It proposes a non-contact, long-range, multi-target structural displacement detection method based on computer vision. First, high-precision long-range detection of multiple displacement targets is achieved using the YOLOv11 algorithm. Then, a target information network is constructed based on the StrongSORT algorithm to improve the problem of decreased confidence in motion model predictions and missed detection of small targets due to small-scale movement. Next, a structural displacement analysis algorithm is constructed by combining a frequency-time domain fusion module and an inverse perspective transformation module to calculate the actual structural displacement of the displacement targets. The proposed algorithm model can significantly improve the detection accuracy of small targets. It corrects target displacement information for image distortion caused by high-frequency camera vibration and shooting angle, achieving long-range real-time detection of multi-target structural displacement. Furthermore, this detection method has a relatively simple algorithm, high accuracy, high sampling frequency, and is sensitive to distance changes. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating the steps of a non-contact, long-distance, multi-target structural displacement detection method based on computer vision according to the present invention.

[0033] Figure 2 This is a structural diagram of the model of the present invention;

[0034] Figure 3 This is a diagram of the improved YOLOv11 algorithm network structure of the present invention;

[0035] Figure 4 Here is a structural diagram of the feature extraction module;

[0036] Figure 5 This is a structural diagram of an embedded multi-scale attention feature acquisition module;

[0037] Figure 6 This is a diagram of the network structure of the StrongSORT algorithm of this invention;

[0038] Figure 7 This is an analytical diagram of the structural displacement based on the combination of the frequency-time domain fusion module and the inverse perspective transformation module of the present invention; Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0040] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various elements, but unless otherwise stated, these elements are not limited by these terms. These terms are used only to distinguish one element from another.

[0041] As described in the background section, existing non-contact structural displacement detection methods involve multiple complex algorithms, resulting in high computational load, sensitivity to distance changes, and unfavorable conditions for real-time structural displacement detection. Furthermore, the reliability of relying on structural surface texture matching is insufficient. Therefore, a simple, flexible, and stable long-distance non-contact multi-point displacement detection method and system are needed. This embodiment addresses the problems of excessive interference with the structural body and insufficient durability of contact displacement detection methods, as well as the low accuracy, low sampling frequency, and sensitivity to distance changes of non-contact displacement detection methods. It proposes a non-contact long-distance multi-target structural displacement detection method based on computer vision.

[0042] like Figure 1 As shown, this embodiment provides a non-contact, long-range multi-target structural displacement detection method based on computer vision. First, high-precision long-range detection of multiple displacement targets is achieved using the YOLOv11 algorithm. Then, a target information network is constructed based on the StrongSORT algorithm to improve the problem of decreased confidence in motion model predictions and missed detection of small targets caused by small-range movement. Next, a structural displacement analysis algorithm is constructed based on a combination of a frequency-time domain fusion module and an inverse perspective transformation module to calculate the actual structural displacement of the displacement targets. The algorithm model proposed in this invention can significantly improve the detection accuracy of small target targets. It corrects target displacement information for image deformation caused by high-frequency camera vibration and shooting angle, achieving long-range real-time detection of multi-target structural displacement. Specifically, the method includes:

[0043] Step 1: Obtain the first video image containing the displacement target; establish the first detection model based on the improved YOLOv11 algorithm; use the first detection model to analyze and process the first video image to obtain the detection result of the displacement target.

[0044] Step 2: Based on the StrongSORT algorithm, an information network is constructed according to the appearance features of various displacement targets. The detection results are input into the information network to form the motion trajectories of various displacement targets.

[0045] Step 3: Based on the combination of frequency domain-time domain fusion module and inverse perspective transformation module, construct a structural displacement analysis algorithm. The structural displacement analysis algorithm calculates the actual structural displacement of the target based on the mapping relationship between the actual size of the target and the pixel coordinates.

[0046] Step 4: Based on the structural displacement analysis algorithm, the displacement changes of the target structure are detected in real time through displacement calculation of consecutive frames.

[0047] Specifically, in this embodiment, step 1 involves acquiring a first video image with a displaced target attached at the test site. To improve the robustness of the training results, the first video image is augmented to form a dataset. The displacement detection of the target in the first video image mainly consists of three parts: target recognition, target tracking, and displacement analysis, such as... Figure 2 The diagram shown is a model structure diagram of the present invention. The improved YOLOv11 algorithm is used to identify and detect the displacement target in the first video image and obtain the detection results, which include category information and confidence values.

[0048] The specific process of establishing the first detection model based on the improved YOLOv11 algorithm is as follows: The first detection model includes a backbone layer, a neck layer, and a head layer. Based on the spatial-depth convolution module in the backbone layer, fine-grained first features are extracted from the first video image. These first features are then coupled across stages using a feature extraction module. An embedded multi-scale attention feature extraction module replaces the bottleneck structure to achieve multi-scale information aggregation of the first features. The dynamic upsampling module in the neck layer performs multi-scale fusion of the aggregated first feature information to obtain the first fused feature. Then, a detection head added to the head layer is used to construct a four-layer cascaded feature pyramid to process the first fused feature, obtaining the detection result of the displacement target. The following section combines... Figure 3 The entire process will be explained below:

[0049] The first detection model includes a backbone network layer, a neck network layer, and a head network layer. The first video image containing the displacement target is input to the first detection model. Fine-grained first features are extracted from the first video image using a first convolutional module and a first spatial-depth convolutional module in the backbone network layer. Multi-scale information aggregation of these first features is achieved through a first embedding multi-scale attention feature extraction module. This first embedding multi-scale attention feature extraction module aggregates the first feature multi-scale information by concatenating and layering information processed by the convolutional module and two multi-scale attention modules. The multi-scale attention module then achieves the aggregation of the first feature multi-scale information. The video image information details are repaired; then, the input information is processed by the second spatial-depth convolution module, the second embedded multi-scale attention feature extraction module, and the third spatial-depth convolution module, and then input into the first feature extraction module for first feature cross-stage feature coupling. The first feature extraction module performs first feature cross-stage feature coupling by splicing and layering the information that has passed through the convolution module and the two feature extraction units. Then, the coupled information is processed by the fourth spatial-depth convolution module and the second feature extraction module, and then, after passing through spatial pyramid pooling and the cross-stage attention module, the information is output from the backbone network layer and input into the neck network layer.

[0050] The first dynamic upsampling module in the neck network layer performs multi-scale fusion of the information. The fused information is then spliced ​​together with the information output from the first feature extraction module via a first splicing module. The spliced ​​information is then processed by a third embedded multi-scale attention feature extraction module and a second dynamic upsampling module. The spliced ​​information is then spliced ​​together with the information output from the second embedded multi-scale attention feature extraction module and the second dynamic upsampling module via a second splicing module. The spliced ​​information is then processed by a fourth embedded multi-scale attention feature extraction module and a third dynamic upsampling module. The information is then spliced ​​together with the information output from the first embedded multi-scale attention feature extraction module and the third dynamic upsampling module via a third splicing module. The spliced ​​information is then output from the neck network via a third feature extraction module and fed into the first detection head of the head network.

[0051] Meanwhile, the information extracted by the third feature extraction module is processed by the fifth spatial-depth convolution module, and then spliced ​​with the information output by the fourth embedded multi-scale attention feature extraction module through the fourth splicing module. The spliced ​​information is then output from the neck network through the fifth embedded multi-scale attention feature extraction module and input into the second detection head of the head network.

[0052] Simultaneously, the information extracted by the fifth embedded multi-scale attention feature extraction module enters the sixth spatial-depth convolution module; the information processed by the sixth spatial-depth convolution module is then concatenated with the information transmitted from the third embedded multi-scale attention feature extraction module through the fifth splicing module; the spliced ​​information enters the sixth embedded multi-scale attention feature extraction module; the neck network is transmitted out through the sixth embedded multi-scale attention feature extraction module and the additional detection head is transmitted into the head network;

[0053] Simultaneously, the information from the sixth embedded multi-scale attention feature extraction module is fed into the seventh spatial-depth convolution module; the information processed by the seventh spatial-depth convolution module is concatenated with the information from the cross-stage attention module through the sixth splicing module, and the spliced ​​information is fed into the third detection head of the head network through the fourth feature extraction module.

[0054] Finally, the information processed by the first detection head, the second detection head, the additional detection head, and the third detection head is transmitted through the output terminal to obtain the detection result of the displacement target.

[0055] Furthermore, the structures of the first, second, third, and fourth feature extraction modules mentioned above are identical, as follows: Figure 4 As shown, it includes, in sequence, a convolution module, a layering module, a first feature extraction unit, a second feature extraction unit, a concatenation module, and another convolution module; the specific information processing procedure is as follows:

[0056] The information input to the feature extraction module first passes through the first convolution module, then through the layering module for layering. After layering, one path of information sequentially enters the first feature extraction unit and the second feature extraction unit for feature extraction. The second feature extraction unit outputs one piece of information, which then enters the splicing module. The other path of information after layering directly enters the splicing module, which splices the two paths of information. Then, the spliced ​​information is processed by the second convolution module and output to the feature extraction module.

[0057] Furthermore, the structures of the first, second, third, fourth, fifth, and sixth embedded multi-scale attention feature extraction modules described above are identical, as follows: Figure 5 As shown, it sequentially includes a first convolutional module, a layered module, a first multi-scale attention module, a second multi-scale attention module, a concatenation module, and a second convolutional module; the specific information processing procedure is as follows:

[0058] The information input to the embedded multi-scale attention feature extraction module first passes through the first convolution module, then through the layering module for layering. The layered information then sequentially enters the first multi-scale attention module and the second multi-scale attention module. The second multi-scale attention module outputs one piece of information, which then enters the splicing module. The other layered signal directly enters the splicing module, which splices the two pieces of information. Finally, the spliced ​​information is processed by the second convolution module and then output to the feature extraction module.

[0059] Furthermore, the convolutional module in this invention corresponds to the Conv module; the spatial-depth convolutional module corresponds to the SPD-Conv module. This invention replaces the convolutional module of YOLOv11 with the SPD-Conv module to solve the problems of high computational complexity and fixed receptive field characteristics of standard convolutional modules, which limit model efficiency and insufficient accuracy in small target detection. The SPD-Conv module consists of a spatial-to-depth transformation layer and a non-stretch convolutional layer. By mapping the spatial blocks of the feature map to the channel dimension, it preserves fine-grained information and optimizes feature representation using non-downsampling convolution. This effectively solves the problem of small target detail loss caused by stretch operations and pooling in traditional convolution, and significantly improves the detection robustness in low-resolution scenes. The embedded multi-scale attention feature extraction module corresponds to the C3K2_MAB module. This invention introduces the C3K2_MAB module to replace the bottleneck structure of the C3K2 module in the YOLOv11 algorithm to solve the problems of recognition errors and false negatives in the detection of structural displacement targets in complex environments, as well as the problem of target edge blurring caused by low-resolution imaging. In this module, the input features are first processed through layer normalization, then long-range and local information are captured by the large kernel convolution decomposition of the MLKA component, and the aggregation of spatial information is further optimized by introducing a gating mechanism and a spatial attention GSAU component, thereby effectively reducing the number of parameters and improving the model's performance in super-resolution tasks; spatial pyramid pooling corresponds to the SPPE module; cross-stage attention module corresponds to the C2PSA module; dynamic upsampling module corresponds to the Dy_Sample module. This invention introduces the Dy_Sample module to replace the upsampling module in the YOLOv11 algorithm. The Dy_Sample module uses a point resampling method, which not only improves resource efficiency but also reduces a lot of computational burden and time; the concatenation module corresponds to the Concat module; and the layering module corresponds to the Split module.

[0060] Specifically, in this embodiment, step 2 uses the StrongSORT algorithm to match and track multiple displacement targets. First, a high-dimensional feature extraction method is used to extract the detection results from the improved YOLOv11 algorithm. Then, a trajectory prediction algorithm is used to predict the target's trajectory. Next, a target association algorithm is used to match the detection results with the trajectory, and the trajectory prediction is updated based on the matching results. Finally, a new trajectory is created to track the new detection results output by the improved YOLOv11 algorithm, and trajectories that have not matched the detection results for a long time are deleted, thereby achieving matching and tracking of multiple displacement targets. See [link to details] for more information. Figure 6 ;

[0061] In step 3, a structural displacement analysis algorithm is constructed based on a combination of a frequency-time domain fusion module and an inverse perspective transformation module to calculate the actual structural displacement of the displacement target. Specifically, the structural displacement analysis algorithm based on the combination of a frequency-time domain fusion module and an inverse perspective transformation module includes the following steps: extracting the displacement target image within the bounding box of the time series based on the tracking displacement detection module; then correcting the displacement target within the bounding box using the inverse perspective transformation module to obtain a corrected displacement target; then obtaining a transformation matrix using the ratio of the real-world displacement target size to the corrected displacement target size in the video image and the effect of the frequency-time domain fusion module; and generating real-time displacement data using the transformation matrix and the trajectory of the displacement target, such as... Figure 7 As shown.

[0062] Furthermore, the correction based on the combination of the frequency-time domain fusion module and the inverse perspective transformation module includes the following steps: extracting the first video image features based on the frequency-time domain fusion module, mapping the features to the virtual position information, unifying the virtual position information using the inverse perspective transformation module, and completing the correction based on the virtual position information and the mapping relationship.

[0063] Furthermore, the structural displacement analysis algorithm calculates structural displacement based on the mapping relationship between the actual size of the target and the pixel coordinates. This mapping relationship is obtained by calibrating the physical size of the target and the mapping of pixels in the image.

[0064] In step 4, based on a structural displacement analysis algorithm, the displacement changes of the target structure are detected in real time through displacement calculation of consecutive frames. Specifically, in this embodiment, the consecutive frames include at least three frames in the first video image. The real-time detection generates a real-time displacement curve by analyzing the motion trajectory of the displacement target in each frame and combining it with the actual structural displacement of the displacement target. The real-time displacement curve is generated based on the displacement calculation in the consecutive frames and is updated in real time. The real-time update generates a displacement change trend curve of the target structure by tracking the motion trajectory of the displacement target in each frame.

[0065] The above method achieves high-precision long-distance detection of multiple displacement targets using the YOLOv11 algorithm; then, it constructs a target information network based on the StrongSORT algorithm to improve the problem of decreased confidence in motion model prediction and missed detection of small targets caused by small-range movement; next, it constructs a structural displacement analysis algorithm based on a combination of frequency-time domain fusion module and inverse perspective transformation module to calculate the actual structural displacement of the displacement target. The algorithm model proposed in this invention can significantly improve the detection accuracy of small target targets, correct the target displacement information for image deformation caused by high-frequency vibration of the camera and shooting angle, and realize long-distance real-time detection of multi-target structural displacement.

[0066] Furthermore, this embodiment also provides a computer vision-based intelligent non-contact structural displacement detection system. This computer vision-based intelligent non-contact structural displacement detection system is used to implement a computer vision-based intelligent non-contact structural displacement detection method. The system includes four modules: a first module, a second module, a third module, and a fourth module.

[0067] The first module is configured to acquire a first video image containing a displacement target, establish a first detection model based on the YOLOv11 algorithm, and use the first detection model to analyze and process the first video image to obtain the detection result of the displacement target.

[0068] The second module is configured to construct an information network based on the StrongSORT algorithm according to the appearance features of various displacement targets, input the detection results into the information network, and form the motion trajectory of various displacement targets.

[0069] The third module is configured to construct a structural displacement analysis algorithm based on the combination of the frequency domain-time domain fusion module and the inverse perspective transformation module. The structural displacement analysis algorithm calculates the actual structural displacement of the displacement target based on the mapping relationship between the actual size of the target and the pixel coordinates.

[0070] The fourth module is configured to detect the displacement changes of the target structure in real time by calculating the displacement of continuous frames based on a structural displacement analysis algorithm.

[0071] The specific steps for each module will not be elaborated here; please refer to the individual steps of the computer vision-based intelligent non-contact structural displacement detection method for details.

[0072] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0073] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0074] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0075] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.

[0076] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A computer vision-based intelligent non-contact structural displacement detection method, characterized in that, The method includes the following steps: Step 1: Obtain the first video image containing the displacement target; establish the first detection model based on the improved YOLOv11 algorithm; use the first detection model to analyze and process the first video image to obtain the detection result of the displacement target. The steps for establishing the first detection model based on the improved YOLOv11 algorithm are as follows: The first detection model includes a backbone network layer, a neck network layer, and a head-neck network layer. Based on the spatial-depth convolution module in the backbone network layer, fine-grained first features are extracted from the first video image. The first features are coupled across stages through a feature extraction module. The first features are aggregated at multiple scales by replacing the bottleneck structure with an embedded multi-scale attention feature extraction module. The dynamic upsampling module in the neck network layer performs multi-scale fusion of the first feature's multi-scale information to obtain a first fused feature. Then, a four-layer cascaded feature pyramid is constructed using a detection head added to the head-neck network layer to process the first fused feature and obtain the detection result of the displacement target. The embedded multi-scale attention feature extraction module corresponds to the C3K2_MAB module. In the C3K2_MAB module, the input features are first processed by layer normalization, and then long-distance and local information are captured by the large kernel convolution decomposition of the MLKA component. The aggregation of spatial information is further optimized by introducing a gating mechanism and the GSAU component with spatial attention. Step 2: Based on the StrongSORT algorithm, an information network is constructed according to the appearance features of various displacement targets. The detection results are input into the information network to form the motion trajectories of various displacement targets. Step 3: Based on the combination of frequency domain-time domain fusion module and inverse perspective transformation module, construct a structural displacement analysis algorithm. The structural displacement analysis algorithm calculates the actual structural displacement of the target based on the mapping relationship between the actual size of the target and pixel coordinates. Step 4: Based on the structural displacement analysis algorithm, the displacement changes of the target structure are detected in real time through displacement calculation of consecutive frames.

2. The intelligent non-contact structural displacement detection method based on computer vision according to claim 1, characterized in that, Constructing the information network based on the appearance features of various displacement targets includes the following steps: Based on the detection results, the StrongSORT algorithm obtains the first feature vectors of the various displacement targets, corrects the inter-frame offset of the first video image through a motion compensation mechanism that maximizes the correlation coefficient, predicts the motion state of the displacement targets in the first video image through a gating mechanism based on the NSA Kalman algorithm, assigns IDs to the various displacement targets, and the StrongSORT algorithm re-associates short-term occlusion trajectories based on an appearance-free linking model. It also completes the missing frame positions of the short-term occlusions using a Gaussian smoothing interpolator and outputs continuous motion trajectories.

3. The intelligent non-contact structural displacement detection method based on computer vision according to claim 2, characterized in that, The structural displacement analysis algorithm based on the combination of a frequency-time domain fusion module and an inverse perspective transformation module includes the following steps: The virtual position information of the displacement target is obtained based on the motion trajectory. The virtual position information of the displacement target is corrected based on the combination of the frequency domain-time domain fusion module and the inverse perspective transformation module. The virtual position information is converted into the actual structural position information based on the scaling factor through the mapping relationship between the actual size of the target and the pixel coordinates. The actual structural displacement of the displacement target is calculated using Euclidean distance.

4. The intelligent non-contact structural displacement detection method based on computer vision according to claim 3, characterized in that, The correction based on the combination of frequency-time domain fusion module and inverse perspective transformation module includes the following steps: The first video image features are extracted based on the frequency-time domain fusion module, the features are mapped to the virtual location information, the virtual location information is unified using the inverse perspective transformation module, and the correction is completed based on the virtual location information and the mapping relationship.

5. The intelligent non-contact structural displacement detection method based on computer vision according to claim 3, characterized in that, The structural displacement analysis algorithm calculates structural displacement based on the mapping relationship between the actual size of the target and the pixel coordinates. The mapping relationship is obtained by calibrating the physical size of the target and the mapping of pixels in the image.

6. The intelligent non-contact structural displacement detection method based on computer vision according to claim 2, characterized in that, In step 4, the continuous frames include at least three frames in the first video image, and the real-time detection generates a real-time displacement curve by analyzing the motion trajectory of the displacement target in each frame and combining it with the actual structural displacement of the displacement target.

7. The intelligent non-contact structural displacement detection method based on computer vision according to claim 6, characterized in that, The real-time displacement curve is generated based on the displacement calculation in the continuous frames and is updated in real time. The real-time update is generated by tracking the motion trajectory of the displacement target in each frame image to generate the displacement change trend curve of the target structure.

8. The intelligent non-contact structural displacement detection method based on computer vision according to claim 1, characterized in that, The detection results include category information and confidence scores.

9. A computer vision-based intelligent non-contact structural displacement detection system, characterized in that, include: The first module is configured to acquire a first video image containing a displacement target, establish a first detection model based on the improved YOLOv11 algorithm, and use the first detection model to analyze and process the first video image to obtain the detection result of the displacement target. Establishing the first detection model based on the improved YOLOv11 algorithm includes the following steps: The first detection model includes a backbone network layer, a neck network layer, and a head-neck network layer. Based on the spatial-depth convolution module in the backbone network layer, fine-grained first features are extracted from the first video image. The first features are coupled across stages through a feature extraction module. The first features are aggregated at multiple scales by replacing the bottleneck structure with an embedded multi-scale attention feature extraction module. The dynamic upsampling module in the neck network layer performs multi-scale fusion of the first feature's multi-scale information to obtain a first fused feature. Then, a four-layer cascaded feature pyramid is constructed using a detection head added to the head-neck network layer to process the first fused feature and obtain the detection result of the displacement target. The embedded multi-scale attention feature extraction module corresponds to the C3K2_MAB module. In the C3K2_MAB module, the input features are first processed by layer normalization, and then long-distance and local information are captured by the large kernel convolution decomposition of the MLKA component. The aggregation of spatial information is further optimized by introducing a gating mechanism and the GSAU component with spatial attention. The second module is configured to construct an information network based on the StrongSORT algorithm according to the appearance features of various displacement targets, input the detection results into the information network, and form the motion trajectory of various displacement targets. The third module is configured to construct a structural displacement analysis algorithm based on the combination of the frequency domain-time domain fusion module and the inverse perspective transformation module. The structural displacement analysis algorithm calculates the actual structural displacement of the displacement target based on the mapping relationship between the actual size of the target and the pixel coordinates. The fourth module is configured to detect the displacement changes of the target structure in real time by calculating the displacement of continuous frames based on a structural displacement analysis algorithm.

Citation Information

Patent Citations

  • Intelligent non-contact structure displacement detection method based on computer vision

    CN119399273A

  • Shaving board surface defect detection method based on deep learning

    CN120147248A