Feature matching-based visual vibration measurement method for mark-free low-texture structure
Through the deep learning visual vibration measurement model, combined with multi-scale feature extraction and adaptive attention enhancement, the problems of low matching accuracy and poor environmental robustness on low texture structures are solved, and the vibration measurement accuracy and calculation efficiency at the subpixel level are achieved.
Patent Information
- Application Number
- CN202510624834.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
AI Technical Summary
The existing visual vibration measurement methods lack significant characteristics in low texture structures, resulting in low matching accuracy and poor environmental robustness, making it difficult to meet the engineering needs of high precision and full-field coverage.
The visual vibration measurement model based on deep learning is adopted, combined with multi-scale feature extraction, adaptive attention enhancement and hierarchical matching optimization technologies, and vibration measurement of markless low-texture structures is achieved through the dual-branch feature extraction module, feature enhancement module and dual-stage matching optimization module.
The vibration characteristics are stably extracted in low-texture areas, improving matching accuracy and environmental robustness, and achieving subpixel-level measurement accuracy and computing efficiency.
Smart Images

Figure CN120489320A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual vibration measurement, and in particular to a visual vibration measurement method of a markerless low-texture structure based on feature matching. Background Art
[0002] Structural vibration monitoring is a core technology for ensuring the service safety of civil and mechanical engineering structures. Its measurement accuracy directly determines the reliability of modal parameter identification, damage location, and health assessment. Traditional contact measurement technologies (such as accelerometers and strain gauges) are prone to altering the inertial properties of lightweight structures due to the added mass effect. Furthermore, they suffer from inherent drawbacks such as complex wiring and susceptibility to electromagnetic interference when monitoring large-scale infrastructure. Furthermore, sensor network deployment is limited by accessibility, making full coverage difficult. These challenges are driving the innovative development of non-contact measurement technologies. Among the current mainstream non-contact devices, laser Doppler vibrometers and microwave interferometer radars avoid mass interference, but their single-point scanning mode cannot meet the spatial resolution requirements of modal analysis. GPS technology is limited by the inherent contradiction between millimeter-level positioning accuracy and high-frequency sampling. Against this backdrop, vision-based vibration measurement methods have become a research hotspot due to their non-contact, full-field measurement, and low cost. The widespread adoption of consumer imaging devices (digital cameras, smartphones, and drones) has significantly expanded the application scenarios of these technologies in engineering.
[0003] Currently, the mainstream computer vision methods used in vibration measurement include template matching, feature matching, and dense optical flow calculations. Template matching uses preset high-contrast markers to track displacement. While the algorithm is simple and efficient, it has serious application limitations. Research has shown that when light intensity changes by more than 50% or template occlusion reaches 30%, measurement errors increase dramatically to the millimeter level, and the durability of artificial markers directly affects monitoring continuity. Feature matching, which analyzes motion based on surface texture, eliminates reliance on artificial markers but faces a conflict between the distribution density of natural feature points and tracking stability. Research has shown that auxiliary markers are still required on low-texture surfaces, effectively violating the principle of non-contact measurement. Dense optical flow calculations achieve full-field motion estimation using pixel-level displacement vector fields. While offering sub-pixel accuracy, they are sensitive to lighting and prone to motion vector divergence, limiting their reliability on engineering sites.
[0004] In recent years, deep learning technology has broken through the limitations of traditional methods. In the field of rotating machinery diagnosis, target detection algorithms have achieved 15%-20% higher accuracy than traditional methods. In civil engineering monitoring, the integration of deep neural networks and optical flow estimation has achieved stable submillimeter-level measurements, marking an acceleration in the technology's practical application. However, existing methods still have core limitations in their applicability to low-texture areas: traditional network architectures struggle to capture the essential characteristics of weak vibration signals, and model generalization in dynamic environments faces significant challenges. Summary of the Invention
[0005] In response to the above-mentioned shortcomings of the existing technology, the present invention provides a visual vibration measurement method for unmarked low-texture structures based on feature matching. By integrating multi-scale feature extraction, adaptive attention enhancement and hierarchical matching optimization technology of deep learning, it solves the technical problems of low matching accuracy and poor environmental robustness in vibration measurement of unmarked low-texture structures due to the lack of significant features.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A visual vibration measurement method for unmarked low-texture structures based on feature matching is used to obtain a vibration video sequence of a target structure to be measured, input the obtained vibration video sequence of the target structure to be measured into a pre-trained visual vibration measurement model, and output a vibration measurement result of the vibration video sequence of the target structure to be measured;
[0008] The visual vibration measurement model includes a dual-branch feature extraction module, a feature enhancement module, a dual-stage matching optimization module, and a displacement conversion module. The visual vibration measurement model transmits the input vibration video sequence of the target structure to the dual-branch feature extraction module to extract an enhanced multi-scale feature map. The feature enhancement module then performs global spatial position modeling and cross-frame alignment on the enhanced multi-scale feature map to output an enhanced low-resolution feature map. The dual-stage matching optimization module then performs bidirectional softmax and mutual nearest neighbor filtering operations on the enhanced low-resolution feature map to generate coarse matching feature point pairs, and determines sub-pixel coordinates through differentiable expected value regression to obtain matching results at this resolution. Finally, the displacement conversion module converts the sub-pixel coordinates into physical displacements through a preset scaling factor and constructs a time series vibration signal to output the vibration measurement results of the vibration video sequence of the target structure.
[0009] As a preferred solution, the dual-branch feature extraction module combines the information perception diversion unit and the multi-scale feature fusion mechanism. The processing process of the dual-branch feature extraction module includes:
[0010] For a vibration video sequence of a target structure as input to a visual vibration measurement model, the first frame image of the vibration video sequence of the target structure to be measured is obtained as a reference frame image and a current frame image and used as the input of the dual-branch feature extraction module, and a low-resolution feature map and a high-resolution feature map are respectively obtained through a multi-scale feature pyramid; and the first-frame reference image and the current frame image are multiplied by weights of different information amounts using the information perception diversion unit through a dynamic weight allocation strategy to obtain a high-information feature map and a low-information feature map; CSPNet and bidirectional orthogonal convolution are used to extract features from the high-information feature map and the low-information feature map, respectively, and the fusion output is used to obtain an enhanced multi-scale feature map.
[0011] As a preferred solution, the information perception and diversion unit uses a dynamic weight allocation strategy to multiply the first frame reference image and the current frame image by weights of different information amounts to obtain a high-information feature map and a low-information feature map. The processing process is as follows:
[0012] For the input reference frame image I0 and the current frame image I t After the initial convolution extraction, the feature map X is obtained, and the group normalization calculation is performed on the feature map X to obtain the group normalized feature map X', and the scaling factor γ corresponding to each channel in the group normalization calculation is calculated. i Normalized to the corresponding weight factor ω i , by the weight factor ω of each channel i Constitute the initial weight vector W γ ; Then, the group normalized feature map X' is combined with the initial weight vector W γ After multiplication, the probabilistic weight matrix is obtained by mapping it to the range (0,1) through the Sigmoid function. Then, the probabilistic weight matrix is converted into Perform binary conversion processing and convert the probabilistic weight matrix The elements greater than or equal to the threshold τ are all set to 1, and the high information weight matrix W is obtained. high , the probabilistic weight matrix The elements greater than or equal to the threshold τ are all set to 0, which means the low-quality information weight matrix W is converted low ; Finally, the feature map X is respectively combined with the high information weight matrix W high and the low-quality information weight matrix W low Multiply to get the high-information feature map F high and low-information feature map F low .
[0013] As a preferred solution, the formula for calculating the feature map X using group normalization is:
[0014]
[0015] Where, X' i Represents the i-th channel data of the group normalized feature map X', X i Represents the i-th channel data of the input feature map X; μ i and σ i Represents X i The mean and standard deviation of γ i , β i They represent the scaling factor and offset factor corresponding to the i-th channel in the group normalization calculation respectively; ε represents a preset constant;
[0016] The initial weight vector W γ The calculation formula is:
[0017] W γ ={ω1,ω2,…,ω i ,…,ω C};
[0018]
[0019] Where C represents the total number of channels in the feature map, and i represents the i-th channel of the feature map X;
[0020] The probabilistic weight matrix The calculation formula is:
[0021]
[0022] Where X' represents the group normalized feature map, and Sigmiod(·) represents the Sigmoid function operation.
[0023] As a preferred solution, the process of extracting features from the high-information feature map and the low-information feature map using CSPNet and bidirectional orthogonal convolution respectively, and fusing the output to obtain an enhanced multi-scale feature map includes:
[0024] For high-information feature maps, CSPNet is used to first split the high-information feature map into two feature maps according to the channel and then input them into the direct branch and the dense calculation branch respectively; the feature map input to the direct branch is convolved with a 3×3 convolution kernel to obtain direct features; the feature map input to the dense calculation branch is first extracted by the CBR unit, and the output of the CBR unit is spliced with the feature map input to the dense calculation branch, and then the high-order features are obtained by the compression conversion unit. Finally, the direct features output by the direct branch are spliced with the high-order features output by the dense calculation branch to obtain the output feature map of CSPNet; wherein, the CBR unit includes a cascaded 3×3 convolution layer, a BN batch normalization layer and a ReLU activation function, and the compression conversion unit includes a cascaded 1×1 convolution layer, a BN batch normalization layer, an LReLU activation function and a maximum pooling layer;
[0025] For low-information feature maps, bidirectional orthogonal convolution is used to convolve the low-information feature maps with 1×3 convolution kernels and 3×1 convolution kernels respectively, and then channel fusion operation is performed to obtain the output feature map of bidirectional orthogonal convolution;
[0026] Finally, the output feature map of CSPNet is added element-wise to the output feature map of bidirectional orthogonal convolution to obtain the enhanced multi-scale feature map.
[0027] As a preferred solution, the processing process of the Transformer-based feature enhancement module includes:
[0028] According to the results of the dual-branch feature extraction module as the input of the position-aware encoding, each feature map is divided into multiple local blocks and flattened into a feature sequence. Position encoding is added to the flattened feature sequence blocks through position encoding to generate a learnable position encoding matrix. The position encoding matrix is added to the feature sequence element by element to generate a dual-channel encoding feature sequence; the encoding feature sequence of one channel is passed through the self-attention mechanism, and a linear transformation is used to generate a query, key, and value triple, and the attention weight of the global spatial position is calculated to obtain an enhanced low-resolution feature map The encoded feature sequence of another channel is passed through the cross attention mechanism with the reference frame image feature as the query and the current frame image feature as the key value to calculate the cross-frame attention weight and generate the aligned enhanced low-resolution feature map
[0029] As a preferred solution, the two-stage matching optimization module includes a coarse matching unit and a fine matching unit, and its processing process includes:
[0030] In the coarse matching unit, first, the enhanced low-resolution feature map of the input and The cosine similarity of all pixels in the image is calculated and normalized to generate a score matrix S. A bidirectional Softmax normalization operation is then performed on the score matrix S in both the row and column directions to convert the similarity scores into probability distributions, representing the matching probabilities of the feature points in the reference frame with all feature points in the current frame, and the matching probabilities of the feature points in the current frame with all feature points in the reference frame, respectively, to obtain a confidence matrix. Finally, a mutual nearest neighbor screening strategy is used to retain only the coarse matching feature point pairs and their confidence scores that simultaneously meet the maximum values of the row and column probability distributions.
[0031] In the fine matching unit, first, the coarse matching result is mapped to the high-resolution feature space by upsampling, and the high-resolution feature map is positioned as input by bidirectional interpolation. A local window is intercepted on the high-resolution feature map with the mapped coordinates as the center to obtain the local feature blocks of the reference frame and the current frame respectively; then, the intercepted local feature block is optimized by the Transformer of the feature enhancement module to generate the enhanced feature. and Next, the similarity between the central feature of the reference frame window and each position in the current frame window is calculated to generate a heat map; finally, the sub-pixel coordinates are determined by differentiable expected value regression, and the matching result M at this resolution is finally output. f .
[0032] As a preferred solution, the processing of the displacement conversion module includes:
[0033] According to the results of the two-stage matching optimization module, the two-dimensional pixel displacement vector Δd between the reference frame and the current frame is calculated by the Euclidean distance based on the coordinates of the matching point pairs; then, the displacement-time series D(t) = {Δd1, Δd2, ..., Δd n}, forming a complete time-course signal; finally, the pixel displacement is converted into physical displacement through a preset scaling factor, and the time-domain displacement data with physical units and spectrum analysis results are finally output.
[0034] As a preferred solution, the visual vibration measurement model is trained by optimizing and updating the parameters of a two-stage matching optimization module with the goal of minimizing a total loss function composed of a coarse matching loss function and a fine matching loss function, thereby optimizing the parameters of the visual vibration measurement model;
[0035] As a preferred solution, the total loss function is:
[0036]
[0037] Where L represents the total loss function, L c represents the coarse matching feature, L f represents fine matching features, represents the true mutual nearest neighbor matching result in the coarse matching unit, The confidence matrix representing the probability of matching feature points between the reference frame and the current frame in the coarse matching unit, Represents the reference frame query point in the fine matching unit The variance value of the similarity heat map with the current frame, Represents the query point of the reference frame obtained after calculation by the fine matching unit The corresponding feature points, Represents the query point in the reference frame The real feature points in the current frame are obtained after transformation calculation; ||·||2 represents the L2 norm operation.
[0038] As a preferred solution, the processing of the displacement conversion module includes:
[0039] According to the results of the two-stage matching optimization module, the two-dimensional pixel displacement vector Δd between the reference frame and the current frame is calculated by the Euclidean distance based on the coordinates of the matching point pairs; then, the displacement-time series D(t) = {Δd1, Δd2, ..., Δd n}, forming a complete time-course signal; finally, the pixel displacement is converted into physical displacement through a preset scaling factor, and the time-domain displacement data with physical units and spectrum analysis results are finally output.
[0040] Compared with the prior art, the present invention has the following technical effects:
[0041] (1) The dual-branch feature extraction module of the present invention adopts a dual-branch differentiation architecture, which combines an information-aware diversion unit and a multi-scale feature fusion mechanism. The information-aware diversion unit adaptively enhances the feature expression ability of low-texture areas through a dynamic weight allocation strategy, so that the model can stably extract the vibration characteristics of weak-texture surfaces without relying on manual labeling. Specifically, the high / low information feature maps are distinguished based on information entropy, and the CSP structure is used to improve the CSPNet architecture of DenseNet for the high-information branch, so as to extract multi-level, high-discriminative features in texture-rich areas, and bidirectional orthogonal convolution is used for the low-information branch, so as to effectively extract directional features in low-texture areas. This design optimizes the allocation of computing resources and improves the generalization ability of the model in low-texture areas; the multi-scale feature fusion mechanism hierarchically fuses the different-scale features output by the dual branches through a feature pyramid network, and combines the cross-stage feature reuse mechanism to enhance the correlation between shallow details and high-level semantics.
[0042] (2) The feature enhancement module (FAM) in the present invention is designed based on the Transformer architecture and combined with the Transformer's self-attention mechanism. The self-attention mechanism enhances the global context modeling capability, solves the problem of limited local receptive field of traditional CNN, and combines with the cross-attention mechanism to establish cross-view feature associations, significantly improving the robustness of feature matching in complex scenarios such as target occlusion and perspective offset, ensuring the accuracy of vibration displacement measurement. At the same time, through position-aware coding, the position-related vector is embedded to display the modeling space order, solving the problem of position information loss in convolution operation.
[0043] (3) The dual-stage matching optimization module of the present invention uses the mutual nearest neighbor matching algorithm to construct a probability matching matrix on the low-resolution feature map after downsampling through the coarse matching unit, and generates a confidence matrix through row-column bidirectional Softmax normalization to quickly lock the potential matching area, thereby greatly reducing the computational complexity; the fine matching unit uses differentiable interpolation to achieve sub-pixel displacement optimization, so that the final matching accuracy reaches the sub-pixel level, thereby ensuring high precision while significantly improving the computational efficiency. Specifically, through local window refinement, the local neighborhood window of the high-resolution feature map is intercepted within the candidate area determined by the coarse matching, and the feature map is upsampled by differentiable bilinear interpolation to restore detail information. The feature enhancement module (FAM) is used to perform global context modeling on the features in the window, optimize the semantic coherence of the local features, and based on the similarity heat map, differentiable expected value regression is used to accurately calculate the sub-pixel coordinate offset of the feature point. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to make the purpose, technical solutions and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0045] Figure 1 is a schematic diagram of a high-speed imaging system of the present invention;
[0046] Figure 2 Schematic diagram of the visual vibration measurement model of the present invention;
[0047] Figure 3 Schematic diagram of the process of obtaining high-information feature maps and low-information feature maps in the solution of the present invention;
[0048] Figure 4 Schematic diagram of the processing process of the enhanced multi-scale feature map after fusion in the solution of the present invention;
[0049] Figure 5 A comparison chart of matching results of different methods on dataset A according to an embodiment of the present invention;
[0050] Figure 6 This is a comparison chart of the time domain results of displacement signals using different methods on dataset A according to an embodiment of the present invention;
[0051] Figure 7 This is a comparison chart of the displacement signal frequency domain results of different methods on dataset A according to an embodiment of the present invention;
[0052] Figure 8 This is a visual comparison diagram of dataset B after different degrees of degradation processing according to an embodiment of the present invention;
[0053] Figure 9 This is a visualization comparison diagram of the data set C after being degraded to different degrees according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0055] The present invention will be described in further detail below with reference to the accompanying drawings.
[0056] Existing traditional contact sensors (such as accelerometers and strain gauges) have an added mass effect in structural vibration monitoring, which may change the inertial characteristics of lightweight structures. In addition, the wiring is complex and susceptible to electromagnetic interference, making it difficult to achieve full-area coverage. At the same time, existing visual methods (such as template matching, feature matching, and dense optical flow calculation) rely on manual markers or natural feature points on low-texture surfaces. However, natural feature points are sparsely distributed and have poor tracking stability, making it difficult to meet high-precision measurement requirements, resulting in insufficient representation of low-texture surface features. Traditional methods are sensitive to dynamic environmental interference such as lighting changes, target occlusion, and perspective offset, resulting in reduced measurement reliability and weak generalization of dynamic scenes. Although methods such as dense optical flow can achieve sub-pixel accuracy, they are sensitive to lighting and have high computational complexity, limiting their engineering practicality.
[0057] To address the above problems and shortcomings, the present invention proposes a visual vibration measurement method for unlabeled low-texture structures based on feature matching, and constructs a visual vibration measurement model based on a multi-scale feature pyramid network and a differentiable feature enhancement module. The model first uses a dual-branch feature extraction module with a cascaded dual-frame input-multi-scale output architecture to distinguish high / low information content feature maps based on information entropy, and adopts CSPNet (high information content) and bidirectional orthogonal convolution (low information content) to extract features, respectively, to optimize computational efficiency and accuracy; however, a feature enhancement module (FMA) is constructed, and the self-attention mechanism is used based on the Transformer to construct long-range dependencies, thereby improving feature robustness under occlusion and perspective changes; then, a two-stage matching optimization module is designed, in which a coarse matching unit uses a mutual nearest neighbor matching algorithm on a low-resolution feature map to quickly locate the candidate region, thereby reducing computational complexity; then, a fine matching unit uses a differentiable bilinear interpolation to map to a high-resolution feature space, and combines feature similarity measurement to achieve sub-pixel displacement field reconstruction; finally, a displacement conversion module is used to obtain a parameterized motion description by calculating the coordinate difference between the corresponding feature points of the two images.
[0058] Specifically, the present invention proposes a method for visual vibration measurement of a markerless low-texture structure based on feature matching. The method obtains a vibration video sequence of a target structure to be measured, inputs the obtained vibration video sequence of the target structure to be measured into a pre-trained visual vibration measurement model, and outputs a vibration measurement result of the vibration video sequence of the target structure to be measured.
[0059] The visual vibration measurement model includes a dual-branch feature extraction module, a feature enhancement module, a dual-stage matching optimization module, and a displacement conversion module. The visual vibration measurement model transmits the input vibration video sequence of the target structure to the dual-branch feature extraction module to extract an enhanced multi-scale feature map. The feature enhancement module then performs global spatial position modeling and cross-frame alignment on the enhanced multi-scale feature map to output an enhanced low-resolution feature map. The dual-stage matching optimization module then performs bidirectional softmax and mutual nearest neighbor filtering operations on the enhanced low-resolution feature map to generate coarse matching feature point pairs, and determines sub-pixel coordinates through differentiable expected value regression to obtain matching results at this resolution. Finally, the displacement conversion module converts the sub-pixel coordinates into physical displacements through a preset scaling factor and constructs a time series vibration signal to output the vibration measurement results of the vibration video sequence of the target structure.
[0060] To address key technical bottlenecks in existing vibration monitoring technologies, such as the added mass effect introduced by contact sensors, insufficient representation of low-texture surface features by traditional visual methods, and weak system generalization in dynamic scenarios, this paper proposes a deep learning-based visual vibration measurement model, VibraNet. This model, through the collaborative design of a multi-scale feature pyramid network and a differentiable feature enhancement module, adopts a dynamic weight allocation strategy for adaptive feature extraction in low-texture areas, breaking through the limitations of traditional methods that rely on explicit surface texture.
[0061] To solve the problem of feature mismatch under complex working conditions, the present invention introduces a Transformer-based feature enhancement module (FAM). By constructing a self-attention-driven long-range dependency modeling network in the spatial dimension and combining it with the cross-view feature correlation matrix established by the cross-attention mechanism, the feature robustness in interference scenarios such as target occlusion and perspective offset is significantly improved. To further ensure measurement accuracy, this model designs a two-stage matching optimization structure: in the first stage, the candidate feature point area is quickly locked through the mutual nearest neighbor matching algorithm on the low-resolution feature map, and the hierarchical abstract characteristics of the feature pyramid are used to reduce the computational complexity; in the second stage, the coarse matching results are upsampled and mapped to the high-resolution feature space, and the sub-pixel displacement field is reconstructed through the differentiable bilinear interpolation operator. It is then iteratively optimized in combination with the feature similarity measurement function to finally achieve sub-pixel matching accuracy.
[0062] The visual vibration measurement method for a markerless low-texture structure based on feature matching of the present invention is described in more detail below.
[0063] 1. Obtain vibration video sequence
[0064] Use high-speed cameras and laser displacement sensors to build a complete acquisition system, such as Figure 1 As shown, a high-speed imaging system is used to obtain a vibration video sequence of the target structure and convert it into a video file; the high-speed camera is used to collect motion video data of the target structure, thereby capturing the full-field displacement through a global image, which complements the laser single-point data to ensure the consistency of local measurement results with the overall motion; the laser displacement sensor is used as a standard signal to verify the collected motion video data of the target structure, thereby calibrating the scale factor of the visual measurement and verifying the absolute accuracy of the visual algorithm, eliminating systematic deviations such as lens distortion and perspective error.
[0065] 2. Visual vibration measurement model
[0066] The present invention proposes a deep learning-based visual vibration measurement model, VibraNet, which uses a feature extraction network with a cascaded dual-frame input and multi-scale output architecture, including a dual-branch feature extraction module, a Transformer-based feature enhancement module, a dual-stage matching optimization module, and a displacement conversion module. The visual vibration measurement model transmits the input vibration video sequence of the target structure to the dual-branch feature extraction module to extract an enhanced multi-scale feature map. The feature enhancement module then performs global spatial position modeling and cross-frame alignment on the enhanced multi-scale feature map, outputting an enhanced low-resolution feature map. The dual-stage matching optimization module then performs bidirectional Softmax and mutual nearest neighbor filtering operations on the enhanced low-resolution feature map to generate coarse matching feature point pairs, and determines sub-pixel coordinates through differentiable expected value regression to obtain matching results at this resolution. Finally, the displacement conversion module converts the sub-pixel coordinates into physical displacements using a preset scaling factor, and constructs a time series vibration signal to output the vibration measurement results of the vibration video sequence of the target structure.
[0067] like Figure 2 As shown, the first frame reference image I0 and the current frame I t As input, it is decomposed into high- and low-information feature maps. High-information features are reused through the CSPNet cross-stage mechanism to enhance feature diversity, while low-information features are captured through bidirectional orthogonal convolution kernels. To further address the problem of insufficient representation capability caused by local receptive fields and high-level semantic associations, we use FAM to enhance feature maps at different granularities, embed global structural constraints, and form highly discriminative and robust feature representations. Next, we use a dual-cascade matching strategy to achieve sub-pixel accuracy. First, a probabilistic matching matrix is constructed based on the mutual nearest neighbor principle to complete coarse matching, and then fine matching optimization is implemented through upsampling and reconstruction of the feature space. Finally, the displacement conversion module obtains a parameterized motion description by calculating the coordinate difference between the corresponding feature points in the two images.
[0068] The dual-branch feature extraction module, feature enhancement module, dual-stage matching optimization module and displacement conversion module are introduced in detail below.
[0069] 2.1 Dual-branch feature extraction module
[0070] For the vibration video sequence of the target structure used as input of the visual vibration measurement model, the first frame image of the vibration video sequence of the target structure to be measured is obtained as the reference frame image and the current frame image and used as the input of the dual-branch feature extraction module, and a low-resolution feature map and a high-resolution feature map are respectively obtained through a multi-scale feature pyramid.
[0071] In order to better distinguish the information features of targetless and low-texture areas, an information-entropy-based information perception diversion unit is used to multiply the first-frame reference image and the current frame image by weights of different information amounts through a dynamic weight allocation strategy to obtain high-information feature maps and low-information feature maps; the high-information feature maps and low-information feature maps are respectively extracted using CSPNet and bidirectional orthogonal convolution, and the fusion output is used to obtain an enhanced multi-scale feature map.
[0072] Specifically, the information perception diversion unit multiplies the first frame reference image and the current frame image by weights of different information amounts through a dynamic weight allocation strategy to obtain the overall processing process of high-information feature maps and low-information feature maps as follows: Figure 3 As shown, specifically including:
[0073] For the input reference frame image I0 and the current frame image I t After the initial convolution extraction, the feature map X is obtained, and the group normalization calculation is performed on the feature map X to obtain the group normalized feature map X', and the scaling factor γ corresponding to each channel in the group normalization calculation is calculated. i Normalized to the corresponding weight factor ω i , by the weight factor ω of each channel i Constitute the initial weight vector W γ , the process is as follows Figure 3 As shown in part a of FIG; Then, the group normalized feature map X' is combined with the initial weight vector W γ After multiplication, the probabilistic weight matrix is obtained by mapping it to the range (0,1) through the Sigmoid function. Then, the probabilistic weight matrix is converted into Perform binary conversion processing and convert the probabilistic weight matrix The elements greater than or equal to the threshold τ are all set to 1, and the high information weight matrix W is obtained. high , the probabilistic weight matrix The elements greater than or equal to the threshold τ are all set to 0, which means the low-quality information weight matrix W is converted low , the process is as follows Figure 3 As shown in part b of the figure; Finally, the feature map X is respectively combined with the high information weight matrix W high and the low-quality information weight matrix W low Multiply to get the high-information feature map F high and low-information feature map F low , the process is as follows Figure 3 As shown in part c.
[0074] Among them, the formula for calculating the feature map X using group normalization can be expressed as:
[0075]
[0076] Where, X' i Represents the i-th channel data of the group normalized feature map X', X i Represents the i-th channel data of the input feature map X; μ i and σ i Represents X i The mean and standard deviation of γ i , β i They represent the scaling factor and offset factor corresponding to the i-th channel in the group normalization calculation respectively; ε represents a preset constant;
[0077] Initial weight vector W γ The calculation formula can be expressed as:
[0078] W γ ={ω1,ω2,…,ω i ,…,ω C};
[0079]
[0080] Where C represents the total number of channels in the feature map, and i represents the i-th channel of the feature map X;
[0081] Probabilistic weight matrix The calculation formula can be expressed as:
[0082]
[0083] Where X' represents the group normalized feature map, and Sigmiod(·) represents the Sigmoid function operation.
[0084] In this embodiment, the scaling factor in group normalization means that a certain feature channel is more critical to the task, and the response of the channel needs to be amplified during the model training process. Conversely, it means that the information of the channel may be suppressed or redundant, thereby distinguishing the amount of information contained in the feature map in this way.
[0085] The high-information feature map and the low-information feature map are respectively subjected to feature extraction using CSPNet and bidirectional orthogonal convolution, and the overall processing process of the fused output to obtain the enhanced multi-scale feature map is as follows: Figure 4 As shown, specifically including:
[0086] For high-information feature maps, CSPNet first splits the high-information feature map into two feature maps according to the channel and then inputs them into the direct branch and dense calculation branch respectively. The direct branch is used to retain the original information flow to extract the direct features, and the dense calculation branch obtains high-order features through the compression-activation mechanism. The direct features and the high-order features are spliced along the channel dimension to output a high-information feature map with enhanced texture structure. The processing process of CSPNet is as follows Figure 4 As shown in part a of the figure, the specific processing process is as follows:
[0087] The high-information feature map F high After being split into two feature maps by channel, they are input into the direct branch and the dense calculation branch respectively; the feature map input to the direct branch is convolved with a 3×3 convolution kernel to obtain the direct feature; the feature map input to the dense calculation branch is first extracted by the CBR unit, and the output of the CBR unit is spliced with the feature map input to the dense calculation branch, and then the high-order feature is obtained by the compression conversion unit. Finally, the direct feature output of the direct branch is spliced with the high-order feature output of the dense calculation branch to obtain the output feature map of CSPNet; wherein, the CBR unit includes a cascaded 3×3 convolution layer, a BN batch normalization layer and a ReLU activation function, and the compression conversion unit includes a cascaded 1×1 convolution layer, a BN batch normalization layer, an LReLU activation function and a maximum pooling layer. The mathematical expression of the CSPNet processing process is:
[0088] Split(F high )=F pass ,F dense ;
[0089] F pass,out =Conv 3×3 (F pass );
[0090] F CBR,out =concat(F dense ,ReLU(BN(Conv 3×3 (F dense ))));
[0091] F dense,out =F Transition,out =MPool(LReLU(BN(Conv 1×1 (F CBR,out ))));
[0092] F out1 =concat(F pass,out ,F dense,out ).
[0093] Where, F out1 represents the output feature map of CSPNet, Split(·) represents splitting by channel, F pass represents the input feature map for the direct branch after splitting, F dense Represents the input feature map for dense computation branch after splitting, F pass,out Represents the direct feature map output by the direct branch, Conv 3×3 (·) represents a 3×3 convolution kernel, F CBR,out represents the output feature map of the CBR module, concat(·) represents concatenation along the channel dimension, ReLU(·) represents the ReLU activation function, BN(·) represents batch normalization, and F dense,out Represents the high-order feature map output by the dense computation branch, F Transition,out Represents the output feature map of the compression conversion unit, MPool(·) represents the maximum pooling operation, LReLU(·) represents the LReLU activation function, Conv 1×1 (·) represents a 1×1 convolution kernel.
[0094] For low-information feature maps, bidirectional orthogonal convolution is used to construct a spatial orthogonal basis using 1×3 and 3×1 convolution kernels, and the output is a low-information feature map containing geometric structure. The process is as follows Figure 4 As shown in part b of the figure, the specific processing process is to transform the low information feature map F low After convolution processing with 1×3 convolution kernel and 3×1 convolution kernel respectively, channel fusion operation is performed to obtain the output feature map of bidirectional orthogonal convolution, whose mathematical expression is:
[0095]
[0096] Where, F out2 Represents the output feature map of bidirectional orthogonal convolution, c 1×3 (·) and c 3×1 (·) represents 1×3 convolution kernel and 3×1 convolution kernel respectively, Represents a channel fusion operation.
[0097] Finally, the output feature map F of CSPNet is out1 The output feature map F of the bidirectional orthogonal convolution out2 Perform element-by-element addition to obtain the enhanced multi-scale feature map.
[0098] In this embodiment, the high-information feature map after feature diversion contains rich detail features. Taking into account the computational complexity of feature extraction and the expression ability of the model, this embodiment uses the CSP structure to improve DenseNet to achieve feature extraction of high-information feature maps. For feature maps with low information content, due to the sparse features, the traditional 3×3 convolution is prone to introduce invalid calculations. This embodiment proposes to use anisotropic decomposition (1×3 and 3×1 convolution) to construct orthogonal basis functions, which equivalently cover the 3×3 receptive field. The high and low information content branches achieve "divide and conquer" through differentiated convolution strategies, avoiding the suboptimal processing of heterogeneous information by a single structure, and achieving the complementarity of coarse and fine granularity features through channel fusion.
[0099] 2.2 Transformer-based feature enhancement module
[0100] The processing of the Transformer-based feature enhancement module includes:
[0101] According to the results of the dual-branch feature extraction module as the input of the position-aware encoding, each feature map is divided into multiple local blocks and flattened into a feature sequence. Position encoding is added to the flattened feature sequence blocks through position encoding to generate a learnable position encoding matrix. The position encoding matrix is added to the feature sequence element by element to generate a dual-channel encoding feature sequence; the encoding feature sequence of one channel is passed through the self-attention mechanism, and a linear transformation is used to generate a query, key, and value triple, and the attention weight of the global spatial position is calculated to obtain an enhanced low-resolution feature map The encoded feature sequence of another channel is passed through the cross attention mechanism with the reference frame image feature as the query and the current frame image feature as the key value to calculate the cross-frame attention weight and generate the aligned enhanced low-resolution feature map
[0102] While the multi-level feature maps generated by the CNN module have a spatial pyramid structure, they suffer from two key bottlenecks: locality limitations and semantic fragmentation. This embodiment introduces an attention structure and implements feature association through a self-attention mechanism, breaking the local receptive field limitations of CNN and achieving global context. Furthermore, a cross-attention mechanism is used to establish semantic-level correspondences between image pairs, enabling feature interaction and addressing the perspective limitations of single-image modeling.
[0103] 2.3. Two-stage matching optimization module
[0104] The two-stage matching optimization module includes a coarse matching unit and a fine matching unit, and its processing process includes:
[0105] In the coarse matching unit, first, the enhanced low-resolution feature map of the input and The cosine similarity of all pixels in the image is calculated and normalized to generate a score matrix S. A bidirectional Softmax normalization operation is then performed on the score matrix S in both the row and column directions to convert the similarity scores into probability distributions, representing the matching probabilities of the feature points in the reference frame with all feature points in the current frame, and the matching probabilities of the feature points in the current frame with all feature points in the reference frame, respectively, to obtain a confidence matrix. Finally, a mutual nearest neighbor screening strategy is used to retain only the coarse matching feature point pairs and their confidence scores that simultaneously meet the maximum values of the row and column probability distributions.
[0106] In the fine matching unit, first, the coarse matching result is mapped to the high-resolution feature space by upsampling, and the high-resolution feature map is positioned as input by bidirectional interpolation. A 5x5 local window is intercepted on the high-resolution feature map with the mapped coordinates as the center to obtain the local feature blocks of the reference frame and the current frame respectively; then, the intercepted local feature block is optimized by the Transformer of the feature enhancement module to generate the enhanced feature. and Next, the similarity between the central feature of the reference frame window and each position in the current frame window is calculated to generate a heat map; finally, the sub-pixel coordinates are determined by differentiable expected value regression, and the matching result M at this resolution is finally output. f .
[0107] Directly performing dense matching on full-resolution feature maps results in a quadratic increase in the computational complexity of the similarity matrix due to the large feature size. Therefore, this embodiment adopts a coarse-to-fine hierarchical matching strategy to quickly locate potential matching areas on downsampled feature maps, narrowing the search space. Matching coordinates are refined within a local window to achieve sub-pixel accuracy.
[0108] 2.4 Displacement conversion module
[0109] The displacement conversion module calculates the two-dimensional pixel displacement vector Δd between the reference frame and the current frame based on the results of the two-stage matching optimization module and the coordinates of the matching point pair through the Euclidean distance, where; then, the displacement-time series D(t) = {Δd1, Δd2, ..., Δd n}, forming a complete time-course signal; finally, the pixel displacement is converted into physical displacement through a preset scaling factor, and the time-domain displacement data with physical units and spectrum analysis results are finally output.
[0110] The calculation formula of Δd is:
[0111]
[0112] Where (x0, y0) represents the reference frame coordinates extracted from the matching point pair, (x t ,yt ) represents the current frame coordinates extracted from the matching point pair.
[0113] 3. Training of visual vibration measurement model
[0114] The visual vibration measurement model is trained by optimizing and updating the parameters of the two-stage matching optimization module with the goal of minimizing the total loss function composed of the coarse matching loss function and the fine matching loss function, thereby optimizing the parameters of the visual vibration measurement model.
[0115] The total loss function used in training is:
[0116] L=L c +L f
[0117]
[0118] Where L represents the total loss function, L c represents the coarse matching feature, L f represents fine matching features, represents the true mutual nearest neighbor matching result in the coarse matching unit, The confidence matrix representing the probability of matching feature points between the reference frame and the current frame in the coarse matching unit, Represents the reference frame query point in the fine matching unit The variance value of the similarity heat map with the current frame, Represents the query point of the reference frame obtained after calculation by the fine matching unit The corresponding feature points, Represents the query point in the reference frame The real feature points in the current frame are obtained after transformation calculation; ||·||2 represents the L2 norm operation.
[0119] 4. Examples
[0120] In order to further verify the effectiveness of the visual vibration measurement model proposed in this embodiment, Figure 5 As shown in the figure, a high-speed imaging system is used to obtain a vibration video A of the target structure. Video A contains a high-definition vibration image of the target structure. In this embodiment, a 400×100 area is selected as the vibration measurement object. A comparative experiment was conducted on various methods based on this data, and the time domain and frequency domain results of the structural displacement signal were compared. Figure 6 and Figure 7 The numerical results are shown in Table 1. The measurement accuracy of the method in this embodiment surpasses the traditional visual vibration measurement method and the existing deep learning-based method in the unmarked low-texture area. The feature matching results are shown in Table 1. Figure 5As shown, the method of this embodiment detects a higher number of feature points and a higher visual matching accuracy in the target-free and low-texture area, which is also confirmed by the numerical results in Table 1.
[0121] Table 1 Measurement results of VibraNet and other visual vibration measurement methods in high-definition video images
[0122]
[0123] In order to further evaluate the stability of our method in the face of poor image visual quality caused by environmental interference or equipment limitations, we processed the original collected data to form Figure 8 As shown in Video B and Figure 9 Video C shown, where Figure 8 (a) is the original image without processing and blur intensity s=0. Figure 8 (b) is a slightly blurred image with blur strength s=10, Figure 8 (c) is a moderately blurred image with blur strength s=20. Figure 8 (d) is a heavily blurred image with blur strength s = 30; Figure 9 (a) is a low-brightness image with illumination intensity α=0.5, Figure 9 (b) is the original image without processing, with illumination intensity α=1. Figure 9 (c) is a moderate brightness image with illumination intensity α=1.5, Figure 9 (d) is a high-brightness image with an illumination intensity of α = 2.0. Video B uses a linear motion blur model to degrade Video A, simulating motion blur caused by a mismatch between the camera sampling rate and the video vibration frequency. Video C uses a linear photometric transformation to simulate brightness perturbations based on Video A, adjusting parameters to simulate low-light and overexposure environments.
[0124] At the same time, this embodiment uses RMSE and R 2 Two indicators are used to measure performance. RMSE (root mean square error) can directly reflect the average deviation between the predicted value and the true value. 2 (Pearson correlation coefficient) evaluates the goodness of fit of the model by comparing the difference between the model prediction results and the original data, and can quantify the model's ability to explain the signal variability. In VibraNet, RMSE and R 2 These two indicators can fully evaluate the accuracy and precision of vibration displacement.
[0125] The calculation formula of the RMSE is:
[0126]
[0127] Where,
[0128] The R 2 The calculation formula is:
[0129]
[0130] Where y i and represent the true value and the predicted value respectively, and Indicates the average of the true and measured values.
[0131] Table 2 shows the RMSE and R of each method when facing motion blur. 2 Results comparison. Table 3 shows the RMSE and R under different light intensities. 2 Results comparison: As can be seen from the table, our method has higher robustness in the face of environmental changes and can adapt to changes in various environments.
[0132] Table 2 RMSE / R of VibraNet under different blur intensities 2 result.
[0133]
[0134] Table 3 RMSE / R of VibraNet under different light intensities 2 result.
[0135]
[0136] 5. Overview
[0137] In summary, the method of the present invention has the following technical advantages:
[0138] Compared to other methods, VibraNet effectively improves its ability to extract image details and its ability to withstand environmental changes, particularly when dealing with motion blur and illumination changes. VibraNet utilizes enhanced feature extraction and attention mechanisms, effectively improving its ability to extract image details and its ability to withstand interference from environmental changes. Specifically, it first optimizes the feature learning process using an information-entropy-based information-aware splitting mechanism. This mechanism allows the network to focus more on areas containing more variation and detail, thereby improving performance when processing complex or low-texture images. High-information feature maps separated by this mechanism are used for feature extraction, while low-information feature maps are subjected to a bidirectional orthogonal convolutional architecture to retrieve valuable information. Through these two steps, the model extracts effective feature representations. Traditional methods assign equal weight to all dimensions of a feature vector, but in practical tasks, certain local features are more important for matching. To address this issue, this embodiment introduces a multi-head attention mechanism that, through efficient query-key interaction, can more accurately learn and capture scene features. These contributions enable VibraNet to achieve high matching accuracy even for unlabeled, low-texture objects, further laying the foundation for the development and application of VibraNet in visual vibration measurement.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described with reference to the preferred embodiments of the present invention, it should be understood by those skilled in the art that various changes can be made in form and details without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A method for visual vibration measurement of low-texture structures without markers based on feature matching, characterized in that: Obtaining a vibration video sequence of a target structure to be measured, inputting the obtained vibration video sequence of the target structure to be measured into a pre-trained visual vibration measurement model, and outputting a vibration measurement result of the vibration video sequence of the target structure to be measured; The visual vibration measurement model includes a dual-branch feature extraction module, a feature enhancement module, a dual-stage matching optimization module, and a displacement conversion module. The visual vibration measurement model transmits the input vibration video sequence of the target structure to the dual-branch feature extraction module to extract an enhanced multi-scale feature map. The feature enhancement module then performs global spatial position modeling and cross-frame alignment on the enhanced multi-scale feature map to output an enhanced low-resolution feature map. The dual-stage matching optimization module then performs bidirectional softmax and mutual nearest neighbor filtering operations on the enhanced low-resolution feature map to generate coarse matching feature point pairs, and determines sub-pixel coordinates through differentiable expected value regression to obtain matching results at this resolution. Finally, the displacement conversion module converts the sub-pixel coordinates into physical displacements through a preset scaling factor and constructs a time series vibration signal to output the vibration measurement results of the vibration video sequence of the target structure.
2. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 1, characterized in that: The dual-branch feature extraction module combines the information perception diversion unit and the multi-scale feature fusion mechanism. The processing process of the dual-branch feature extraction module includes: For a vibration video sequence of a target structure as input to a visual vibration measurement model, the first frame image of the vibration video sequence of the target structure to be measured is obtained as a reference frame image and a current frame image and used as the input of the dual-branch feature extraction module, and a low-resolution feature map and a high-resolution feature map are respectively obtained through a multi-scale feature pyramid; and the first-frame reference image and the current frame image are multiplied by weights of different information amounts using the information perception diversion unit through a dynamic weight allocation strategy to obtain a high-information feature map and a low-information feature map; CSPNet and bidirectional orthogonal convolution are used to extract features from the high-information feature map and the low-information feature map, respectively, and the fusion output is used to obtain an enhanced multi-scale feature map.
3. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 2, characterized in that: The processing process of using the information perception and diversion unit to multiply the first frame reference image and the current frame image by weights of different information amounts through a dynamic weight allocation strategy to obtain a high-information feature map and a low-information feature map is as follows: For the input reference frame image I0 and the current frame image I t After the initial convolution extraction, the feature map X is obtained, and the group normalization calculation is performed on the feature map X to obtain the group normalized feature map X', and the scaling factor γ corresponding to each channel in the group normalization calculation is calculated. i Normalized to the corresponding weight factor ω i , by the weight factor ω of each channel i Constitute the initial weight vector W γ ; Then, the group normalized feature map X' is combined with the initial weight vector W γ After multiplication, the probabilistic weight matrix is obtained by mapping it to the range (0,1) through the Sigmoid function. Then, the probabilistic weight matrix is converted into Perform binary conversion processing and convert the probabilistic weight matrix The elements greater than or equal to the threshold τ are all set to 1, and the high information weight matrix W is obtained. high , the probabilistic weight matrix The elements greater than or equal to the threshold τ are all set to 0, which means the low-quality information weight matrix W is converted low ; Finally, the feature map X is respectively combined with the high information weight matrix W high and the low-quality information weight matrix W low Multiply to get the high-information feature map F high and low-information feature map F low .
4. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 3, characterized in that: The formula for calculating the feature map X using group normalization is: Where X′ i Represents the i-th channel data of the group normalized feature map X', X i Represents the i-th channel data of the input feature map X; μ i and σ i Represents X i The mean and standard deviation of γ i , β i They represent the scaling factor and offset factor corresponding to the i-th channel in the group normalization calculation respectively; ε represents a preset constant; The initial weight vector W γ The calculation formula is: W γ ={ω1,ω2,…,ω i ,…,oh C }; Where C represents the total number of channels in the feature map, and i represents the i-th channel of the feature map X; The probabilistic weight matrix The calculation formula is: Where X' represents the group normalized feature map, and Sigmiod(·) represents the Sigmoid function operation.
5. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 2, characterized in that: The process of extracting features from the high-information feature map and the low-information feature map using CSPNet and bidirectional orthogonal convolution respectively, and fusing the output to obtain an enhanced multi-scale feature map includes: For high-information feature maps, CSPNet is used to first split the high-information feature map into two feature maps according to the channel and then input them into the direct branch and the dense calculation branch respectively; the feature map input to the direct branch is convolved with a 3×3 convolution kernel to obtain direct features; the feature map input to the dense calculation branch is first extracted by the CBR unit, and the output of the CBR unit is spliced with the feature map input to the dense calculation branch, and then the high-order features are obtained by the compression conversion unit. Finally, the direct features output by the direct branch are spliced with the high-order features output by the dense calculation branch to obtain the output feature map of CSPNet; wherein, the CBR unit includes a cascaded 3×3 convolution layer, a BN batch normalization layer and a ReLU activation function, and the compression conversion unit includes a cascaded 1×1 convolution layer, a BN batch normalization layer, an LReLU activation function and a maximum pooling layer; For low-information feature maps, bidirectional orthogonal convolution is used to convolve the low-information feature maps with 1×3 convolution kernels and 3×1 convolution kernels respectively, and then channel fusion operation is performed to obtain the output feature map of bidirectional orthogonal convolution; Finally, the output feature map of CSPNet is added element-wise to the output feature map of bidirectional orthogonal convolution to obtain the enhanced multi-scale feature map.
6. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 1, characterized in that: The processing process of the Transformer-based feature enhancement module includes: According to the results of the dual-branch feature extraction module as the input of the position-aware encoding, each feature map is divided into multiple local blocks and flattened into a feature sequence. Position encoding is added to the flattened feature sequence blocks through position encoding to generate a learnable position encoding matrix. The position encoding matrix is added to the feature sequence element by element to generate a dual-channel encoding feature sequence; the encoding feature sequence of one channel is passed through the self-attention mechanism, and a linear transformation is used to generate a query, key, and value triple, and the attention weight of the global spatial position is calculated to obtain an enhanced low-resolution feature map The encoded feature sequence of another channel is passed through the cross attention mechanism with the reference frame image feature as the query and the current frame image feature as the key value to calculate the cross-frame attention weight and generate the aligned enhanced low-resolution feature map 7. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 1, characterized in that: The two-stage matching optimization module includes a coarse matching unit and a fine matching unit, and its processing process includes: In the coarse matching unit, first, the enhanced low-resolution feature map of the input and The cosine similarity of all pixels in the image is calculated and normalized to generate a score matrix S. A bidirectional Softmax normalization operation is then performed on the score matrix S in both the row and column directions to convert the similarity scores into probability distributions, representing the matching probabilities of the feature points in the reference frame with all feature points in the current frame, and the matching probabilities of the feature points in the current frame with all feature points in the reference frame, respectively, to obtain a confidence matrix. Finally, a mutual nearest neighbor screening strategy is used to retain only the coarse matching feature point pairs and their confidence scores that simultaneously meet the maximum values of the row and column probability distributions. In the fine matching unit, first, the coarse matching result is mapped to the high-resolution feature space by upsampling, and the high-resolution feature map is positioned as input by bidirectional interpolation. A local window is intercepted on the high-resolution feature map with the mapped coordinates as the center to obtain the local feature blocks of the reference frame and the current frame respectively; then, the intercepted local feature block is optimized by the Transformer of the feature enhancement module to generate the enhanced feature. and Next, the similarity between the central feature of the reference frame window and each position in the current frame window is calculated to generate a heat map; finally, the sub-pixel coordinates are determined by differentiable expected value regression, and the matching result M at this resolution is finally output. f .
8. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 1, wherein: The processing process of the displacement conversion module includes: According to the results of the two-stage matching optimization module, the two-dimensional pixel displacement vector Δd between the reference frame and the current frame is calculated by the Euclidean distance based on the coordinates of the matching point pairs; then, the displacement-time series D(t) = {Δd1, Δd2, ..., Δd n }, forming a complete time-course signal; finally, the pixel displacement is converted into physical displacement through a preset scaling factor, and the time-domain displacement data with physical units and spectrum analysis results are finally output.
9. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 1, characterized in that: The visual vibration measurement model is trained by optimizing and updating the parameters of a two-stage matching optimization module with the goal of minimizing a total loss function composed of a coarse matching loss function and a fine matching loss function, thereby optimizing the parameters of the visual vibration measurement model.
10. The method for visual vibration measurement of a markerless low-texture structure based on feature matching according to claim 9, characterized in that: The total loss function is: Where L represents the total loss function, L c represents the coarse matching feature, L f represents fine matching features, represents the true mutual nearest neighbor matching result in the coarse matching unit, The confidence matrix representing the probability of matching feature points between the reference frame and the current frame in the coarse matching unit, Represents the reference frame query point in the fine matching unit The variance value of the similarity heat map with the current frame, Represents the query point of the reference frame obtained after calculation by the fine matching unit The corresponding feature points, Represents the query point in the reference frame The real feature points in the current frame are obtained after transformation calculation; ||·||2 represents the L2 norm operation.
Citation Information
Cited By
Miniature intelligent terminal intestinal image processing method and device, storage medium and computer equipment
CN121353289A
Micro intelligent terminal intestinal tract image processing method, processing device, storage medium and computer equipment
CN121353289B
Wafer image matching method and device and computer storage medium
CN121353703A