Steel-wood combined beam column joint damage identification method and system applying machine vision
By applying machine vision technology, combined with adaptive filtering and denoising, multi-scale feature fusion and deep learning network, the problems of high damage identification cost of steel-wood composite beam and column nodes in the existing technology and difficulty in achieving global continuous monitoring are solved, and the precise identification and positioning of the damage of steel-wood composite beam and column nodes is achieved.
Patent Information
- Application Number
- CN202411979253.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing steel-wood composite beam-column node damage recognition method relies on sensor measurement, which increases detection cost and makes it difficult to achieve continuous monitoring of global damage.
Machine vision technology is adopted, combining adaptive filtering and denoising, multi-scale feature fusion and deep learning networks, and multi-modal image sequences are obtained, so as to achieve accurate identification and positioning of node damage of steel-wood combined beam columns.
The precise identification and positioning of the damage of the steel-wood composite beam and column nodes is achieved, which reduces the detection cost and can realize continuous monitoring of global damage.
Smart Images

Figure CN120014284A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of steel-wood composite beam-column node damage, and in particular to a steel-wood composite beam-column node damage identification method and system using machine vision. Background Art
[0002] With the rapid development of building industrialization, steel-wood composite structures have been widely used in the field of construction engineering due to their advantages such as environmental protection, high efficiency and economy. Especially in new building forms such as prefabricated buildings, green buildings and high-rise wooden structures, the steel-wood composite beam-column nodes are key load-bearing components, and their performance directly affects the safety, durability and service life of the overall structure. As an important load-bearing component in the building structure, the damage state of the steel-wood composite beam-column nodes directly affects the safety performance of the overall structure. At present, the structural damage identification methods mainly include traditional methods based on physical quantities such as vibration, strain and displacement.
[0003] Traditional structural damage identification methods mainly rely on the analysis of physical quantities measured by sensors. For example, the invention patent with application number 201910800633.3 discloses a beam structure damage identification method based on support reaction and strain. This method measures the strain curve and support reaction of the damaged structure, calculates the stiffness at each position, and identifies the damage location and degree through stiffness mutation. This type of method requires a large number of sensors to be arranged on the structure, which increases the detection cost. At the same time, the choice of sensor installation position has a great influence on the identification result, and only damage information at discrete measuring points can be obtained, making it difficult to achieve continuous monitoring of global damage. Summary of the invention
[0004] In view of this, the present invention provides a method and system for identifying damage at steel-wood composite beam-column nodes using machine vision. The purpose is to provide a method and system for identifying damage at steel-wood composite beam-column nodes using machine vision, combining adaptive filtering denoising, multi-scale feature fusion and deep learning network to achieve accurate identification and positioning of damage at steel-wood composite beam-column nodes.
[0005] To achieve the above object, the present invention provides a method for identifying damage at a steel-wood composite beam-column node using machine vision, comprising the following steps:
[0006] S1: Obtain visible light, near infrared and thermal imaging multimodal image sequences of steel-wood composite beam-column joints, perform denoising processing through adaptive anisotropic diffusion filtering, obtain denoised multimodal image sequences, and construct denoised multimodal image sequence tensors based on the denoised multimodal image sequences;
[0007] S2: Perform Gaussian pyramid decomposition and Laplacian pyramid feature fusion on the denoised multimodal image sequence tensor to obtain a fused feature map;
[0008] S3: Perform adaptive threshold segmentation and morphological processing on the fused feature map to obtain the final binary image;
[0009] S4: The deep learning network is trained to obtain a trained deep learning network;
[0010] S5: Use the trained deep learning network to identify damage, obtain segmentation probability maps and annotate segmentation results to visualize the damaged area, including:
[0011] Use the trained deep learning network to extract features from the denoised multimodal image sequence tensor, and perform spatial attention guidance on the extracted features and the final binary image to obtain guided features;
[0012] The guided features and the final binary image are concatenated in the feature dimension and input into the Transformer module. The features are fused through the self-attention mechanism of the Transformer module to obtain the fused features.
[0013] The fused features are input into the trained deep learning network to perform damaged pixel segmentation and output a segmentation probability map;
[0014] Annotate the segmentation results on the multimodal image sequences to visualize the damaged areas.
[0015] Optionally, the S1 includes:
[0016] S11: Acquiring multimodal image sequences using sensor arrays of three different wavelengths I total , specifically:
[0017] I total = {I vis , I nir , I thm};
[0018] Among them, I vis ,I nir and I thm These are modal images collected in the visible light band, near infrared band, and thermal imaging band;
[0019] S12: Using iterative optimization strategy to optimize multimodal image sequence I total Adaptive anisotropic diffusion filtering is performed to denoise the multimodal image sequence I after denoising. den , specifically:
[0020]
[0021] in, and are the denoised pixel values at the pixel point (x, y) of the mth modality image in the denoised multimodal image sequence at the tth and t+1th iterations respectively; m = 1, 2, 3, are the modality numbers, and the first modality image, the second modality image and the third modality image are I vis ,I nir and I thm ; When t = 0, is the pixel value of the mth modality image at the pixel point (x, y) in the multimodal image sequence; η is the iteration step parameter; N is the set of four directions: up, down, left, and right; is the gradient value of the mth modality image at the pixel point (x, y) along the direction d; c d (x, y) is the diffusion coefficient of the pixel point (x, y) in direction d, specifically:
[0022]
[0023] Where K(x, y) is the adaptive local threshold at the pixel point (x, y):
[0024] K(x,y)=μ(x,y)+λ·σ(x,y);
[0025] Where μ(x,y) is The mean value of the gradient amplitude in a 7×7 local window centered at the pixel point (x, y); σ(x, y) is The standard deviation of the gradient amplitude in a 7×7 local window centered at the pixel point (x, y); λ is the adjustment parameter;
[0026] The iteration termination condition is:
[0027]
[0028] Among them, ε is the convergence threshold; is the L2 norm of the t-th iteration result of the m-th modality image in the denoised multimodal image sequence; is the L2 norm of the difference between the t+1th and tth iteration results of the mth modal image; after the iteration is completed, the mth modal image in the denoised multimodal image sequence is obtained
[0029] S13: Construct the denoised multimodal image sequence tensor T:
[0030]
[0031] Optionally, the S2 includes:
[0032] S21: The denoised multimodal image sequence tensor is decomposed into a Gaussian pyramid, specifically:
[0033]
[0034] in, is the pixel value of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) in the layer-th Gaussian pyramid decomposition result image; layer = 1, 2, 3, is the number of Gaussian pyramid layers; G sigma is a Gaussian kernel function with a standard deviation of 1.6; * is a convolution operation; It is the Gaussian pyramid decomposition result image of the mth modality image in the denoised multimodal image sequence tensor at the layer-1 level. When layer=1 ↓2 is a 2x downsampling operation in the horizontal and vertical directions;
[0035] S22: Construct a Laplace pyramid, specifically:
[0036]
[0037] in, is the result of the Laplacian pyramid of the mth modality image in the denoised multimodal image sequence tensor at layer-1; Up is the bilinear interpolation upsampling operation, which expands the input image by 2 times in the horizontal and vertical directions; It is the Gaussian pyramid decomposition result image of the mth modality image in the denoised multimodal image sequence tensor at the layer;
[0038] S23: Calculate the fusion weight, specifically:
[0039]
[0040] in, is the fusion weight of the mth modality image at the pixel (x, y) in the layer-1 layer and the denoised multimodal image sequence tensor; is the saliency map of the kth modality image at the pixel (x, y) in the denoised multimodal image sequence tensor, layer-1, k = 1, 2, 3; (x, y) is the layer-1 layer, and the saliency map of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) is:
[0041]
[0042] Where Ω is the pixel point contained in the 7×7 neighborhood window centered on the pixel point (x, y), and (i, j) is the element in Ω; for The mean value of the pixel value in the 7×7 local window centered at the pixel point (x, y);
[0043] S24: Laplacian pyramid feature fusion and reconstruction to obtain a fused feature map.
[0044] Optionally, the S24 includes:
[0045] S241: Perform weighted fusion on the Laplacian pyramid to obtain the fused Laplacian pyramid:
[0046]
[0047] in, is the pixel value after fusion at the pixel point (x, y) of layer-1; is the pixel value of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) in the result image of the layer-1 Laplacian pyramid;
[0048] S242: Reconstruct the fused Laplacian pyramid to obtain the fused feature map F:
[0049]
[0050] Among them, Up layer-1 Indicates a layer-1 2x upsampling operation; F(x,y) is the pixel value of the fused feature map at the pixel point (x, y).
[0051] Optionally, in step S3, adaptive threshold segmentation and morphological processing are performed on the fused feature map to obtain a final binary map, including:
[0052] S31: Calculate the adaptive local threshold of the fused feature map, specifically:
[0053]
[0054] Where Th(x,y) is the adaptive local threshold at the pixel point (x,y); is the mean of the fused feature map F in the 15×15 local window centered at the pixel point (x, y); is the standard deviation of the fused feature map F in the 15×15 local window centered at the pixel point (x, y); α is the threshold adjustment coefficient;
[0055] S32: Adaptively perform local threshold segmentation on the fused feature map, and compare the size relationship between the fused feature map F and the adaptive threshold Th at each pixel point (x, y) to obtain the initial binary map M0, specifically:
[0056] When F(x,y)>Th(x,y), M0(x,y)=1;
[0057] When F(x,y)≤Th(x,y), M0(x,y)=0;
[0058] Among them, M0(x,y) is the pixel value of the initial binary image M0 at the pixel point (x,y);
[0059] S33: Perform morphological processing on the initial binary image to obtain a final binary image.
[0060] Optionally, the S33 includes:
[0061] S331: Perform opening operation:
[0062] An opening operation is performed on the initial binary image M0. The opening operation uses a 3×3 rectangular structure element to first perform an erosion operation on the initial binary image M0, and then perform an expansion operation to obtain an image M1 after the opening operation.
[0063] S332: Perform closing operation:
[0064] Perform a closing operation on M1. The closing operation also uses a 3×3 rectangular structure element to first perform an expansion operation on M1, and then perform an erosion operation to obtain an image M2 after the closing operation.
[0065] S333: Perform area filtering:
[0066] The area filtering process is performed on M2. The area filtering process is to calculate the pixel area of each connected area in the image M2, and set the pixel values of the pixels contained in the connected areas with an area less than 20 to 0, so as to obtain the final binary image M2. final .
[0067] Optionally, the S4 includes:
[0068] S41: Calculate segmentation loss Loss seg and consistency constraint loss Loss consist , get the total loss Loss total :
[0069] Loss seg =CE(Prob seg , Label);
[0070] Loss consist=MSE(Prob seg , Label);
[0071] Loss total =Loss seg +Loss consist ;
[0072] Among them, CE is the cross entropy loss; MSE is the mean square error; Label is the true segmentation label; Prob seg is the segmentation probability map;
[0073] S42: Use the stochastic gradient descent algorithm to train the parameters in the deep learning network to reduce the total loss; after reaching the set number of iterations, a trained deep learning network is obtained.
[0074] Optionally, the S5 includes:
[0075] The trained deep learning networks include ResNet50 network and UNet network;
[0076] The ResNet50 network is used to extract features from the denoised multimodal image sequence tensor to obtain the denoised multimodal image sequence tensor features;
[0077] The denoised multimodal image sequence tensor features are combined with the final binary image M final Perform spatial attention guidance to obtain guided multimodal image sequence tensor features;
[0078] The guided multimodal image sequence tensor features and the final binary map M final After splicing in the feature dimension, it is input into the Transformer module for feature fusion to obtain the fused multimodal image sequence tensor feature Fea fuse ;
[0079] The fused multimodal image sequence tensor features are input into the UNet network for damaged pixel segmentation to obtain a segmentation probability map, where each pixel value of the segmentation probability map represents the probability that the position belongs to the damaged area;
[0080] The segmentation probability map is binarized to obtain the final segmentation mask, and the segmentation results are annotated on the multimodal image sequence to achieve visualization of the damaged area.
[0081] The present invention also discloses a steel-wood composite beam-column node damage identification system using machine vision, comprising:
[0082] Denoising module: obtains visible light, near infrared and thermal imaging multimodal image sequences of steel-wood composite beam-column joints, performs denoising processing through adaptive anisotropic diffusion filtering, obtains denoised multimodal image sequences, and constructs denoised multimodal image sequence tensors based on the denoised multimodal image sequences;
[0083] Fusion module: Perform Gaussian pyramid decomposition and Laplacian pyramid feature fusion on the denoised multimodal image sequence tensor to obtain a fused feature map;
[0084] Binarization module: performs adaptive threshold segmentation and morphological processing on the fused feature map to obtain the final binary map;
[0085] Recognition model training module: Recognition model training module: train the deep learning network to obtain a trained deep learning network;
[0086] Identification model application module: Use the trained deep learning network to perform damage identification, obtain the segmentation probability map and annotate the segmentation results to realize the visualization of the damaged area.
[0087] Compared with the prior art, the present invention has at least the following beneficial effects:
[0088] The present invention uses image sequences of three different bands, namely visible light, near infrared and thermal imaging, for damage identification, making full use of the complementary information of different modal images. The multimodal images are denoised by an adaptive anisotropic diffusion filter algorithm, which can adaptively adjust the diffusion coefficient according to local image features, effectively suppressing noise while maintaining edge and detail information. At the same time, iterative optimization strategies and adaptive termination conditions are used to ensure the stability and reliability of the denoising effect, laying the foundation for subsequent feature extraction and fusion.
[0089] The present invention decomposes the image into features of different scales through Gaussian pyramid decomposition, and realizes multi-scale expression of features in combination with Laplacian pyramid reconstruction. An adaptive weight fusion strategy based on local saliency is innovatively introduced, so that features of different modalities and scales can be optimally combined according to their importance. In addition, the adaptive threshold segmentation and morphological processing methods are used to effectively extract the preliminary outline of the damaged area, providing valuable prior information for the deep learning network.
[0090] The deep learning network architecture designed in the present invention has strong feature extraction and fusion capabilities. Feature extraction is performed through ResNet50, combined with the spatial attention mechanism guided by binary images, to highlight the feature expression of damage-related areas. The innovative introduction of the Transformer module realizes the deep fusion of multimodal features and fully explores the correlation between different modal features. The UNet network is used for damage pixel segmentation, and an optimization target combining cross entropy loss and consistency constraints is designed to improve the accuracy and robustness of damage identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 The present invention is a flowchart of a method for identifying damage at a steel-wood composite beam-column joint using machine vision according to an embodiment of the present invention. DETAILED DESCRIPTION
[0092] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.
[0093] Example 1: A method for identifying damage at a steel-wood composite beam-column joint using machine vision, such as Figure 1 As shown, the following steps are included:
[0094] S1: Obtain visible light, near infrared and thermal imaging multimodal image sequences of steel-wood composite beam-column joints, and perform denoising processing through adaptive anisotropic diffusion filtering to obtain denoised multimodal image sequences, and construct denoised multimodal image sequence tensors based on the denoised multimodal image sequences:
[0095] S11: Acquiring multimodal image sequences using sensor arrays of three different wavelengths I total , specifically:
[0096] I total = {I vis , I nir , I thm};
[0097] Among them, I vis ,I nir and I thm These are modal images collected in the visible light band, near infrared band, and thermal imaging band;
[0098] S12: Using iterative optimization strategy to optimize multimodal image sequence I total Adaptive anisotropic diffusion filtering is performed to denoise the multimodal image sequence I after denoising. den , specifically:
[0099]
[0100] in, and are the pixel values at the pixel point (x, y) of the mth modality image in the denoised multimodal image sequence at the tth and t+1th iterations; m = 1, 2, 3, are the modality numbers, and the first modality image, the second modality image, and the third modality image are I vis ,I nir and I thm ; When t = 0, is the pixel value of the mth modality image at the pixel point (x, y) in the multimodal image sequence; η is the iteration step parameter, and its value range is (0, 1]; N is the set of four directions: up, down, left, and right; is the gradient value of the mth modality image at the pixel point (x, y) along the direction d; c d (x, y) is the diffusion coefficient of the pixel point (x, y) in direction d, specifically:
[0101]
[0102] Where K(x,y) is the adaptive local threshold at the pixel point (x,y):
[0103] K(x,y)=μ(x,y)+λ·σ(x,y);
[0104] Where μ(x, y) is The mean value of the gradient amplitude in a 7×7 local window centered at the pixel point (x, y); σ(x, y) is The standard deviation of the gradient amplitude in a 7×7 local window centered at the pixel point (x, y); λ is a tuning parameter with a value range of [0.8, 1.2];
[0105] The iteration termination condition is:
[0106]
[0107] Among them, ε is the convergence threshold, which is set to 0.00l; is the L2 norm of the t-th iteration result of the m-th modality image in the denoised multimodal image sequence; is the L2 norm of the difference between the t+1th and tth iteration results of the mth modal image; after the iteration is completed, the mth modal image in the denoised multimodal image sequence is obtained
[0108] S13: Construct the denoised multimodal image sequence tensor T:
[0109]
[0110] This step collects image data in three different bands: visible light, near infrared, and thermal imaging, making full use of the complementary information characteristics of different bands. Visible light images can provide clear surface details and texture features, near infrared images have the ability to penetrate the surface to display the internal structure, and thermal imaging images can reflect the internal stress distribution and thermal anomaly areas of the material. This multimodal fusion significantly improves the comprehensiveness and reliability of damage detection.
[0111] The adaptive anisotropic diffusion filter used in this step has unique advantages in image processing. Through iterative optimization strategy, this method can automatically adjust the filter strength according to the local features of the image, strengthen denoising in smooth areas while protecting the detail information of edge areas. The diffusion coefficient design in the filtering process takes into account the local gradient information, and the introduction of adaptive local threshold enables the algorithm to intelligently adapt to the image characteristics of different areas. By setting a reasonable iterative termination condition, the convergence of the algorithm is ensured and the loss of details caused by excessive filtering is avoided.
[0112] S2: Perform Gaussian pyramid decomposition and Laplacian pyramid feature fusion on the denoised multimodal image sequence tensor to obtain a fused feature map:
[0113] S21: Perform Gaussian pyramid decomposition on the denoised multimodal image sequence tensor, specifically:
[0114]
[0115] in, is the pixel value of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) in the layer-th Gaussian pyramid decomposition result image; layer = 1, 2, 3, is the number of Gaussian pyramid layers; G sigma is a Gaussian kernel function with a standard deviation of 1.6; * is a convolution operation; It is the Gaussian pyramid decomposition result image of the mth modality image in the denoised multimodal image sequence tensor at the layer-1 level. When layer=1 ↓2 is a 2x downsampling operation in the horizontal and vertical directions;
[0116] S22: Construct a Laplace pyramid, specifically:
[0117]
[0118] in, is the result of the Laplacian pyramid of the mth modality image in the denoised multimodal image sequence tensor at layer-1; Up is the bilinear interpolation upsampling operation, which expands the input image by 2 times in the horizontal and vertical directions; It is the Gaussian pyramid decomposition result image of the mth modality image in the denoised multimodal image sequence tensor at the layer;
[0119] S23: Calculate the fusion weight, specifically:
[0120]
[0121] in, is the fusion weight of the mth modality image at the pixel (x, y) in the layer-1 layer and the denoised multimodal image sequence tensor; is the saliency map of the kth modality at the pixel (x, y) in the denoised multimodal image sequence tensor at layer-1, k = 1, 2, 3; is the layer-1 layer, and the saliency value of the mth modality in the denoised multimodal image sequence tensor at the pixel point (x, y) is:
[0122]
[0123] Where Ω is the pixel point contained in the 7×7 neighborhood window centered on the pixel point (x, y), and (i, j) is the element in Ω; for The mean value of the pixel value in the 7×7 local window centered at the pixel point (x, y);
[0124] S24: Feature fusion and reconstruction to obtain a fused feature map; specifically:
[0125] S241: performing weighted fusion on the Laplacian pyramid to obtain a fused Laplacian pyramid; wherein the weighted fusion operation is to multiply the pixel values at the corresponding positions of each modality image in the denoised multimodal image sequence tensor by the corresponding weights for each layer of the Laplacian pyramid, and then add the weighted pixel values, specifically:
[0126]
[0127] in, is the pixel value after fusion at the pixel point (x, y) of layer-1; is the pixel value of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) in the result image of the layer-1 Laplacian pyramid;
[0128] S242: Reconstruct the fused Laplacian pyramid to obtain the fused feature map F:
[0129]
[0130] Among them, Up layer-1 Indicates a layer-1 2x upsampling operation; F(x, y) is the pixel value of the final fused feature map at the pixel point (x, y).
[0131] This step achieves effective feature extraction and fusion of multimodal images through Gaussian pyramid decomposition and fusion strategy. First, Gaussian pyramid decomposition is used to decompose the image of each modality into representations of different scales. This hierarchical structure can capture the multi-scale feature information of the image from details to the global. The introduction of Gaussian kernel function ensures the smoothness and stability of the feature extraction process, while the pyramid downsampling operation effectively reduces the computational complexity.
[0132] The Laplacian pyramid constructed on this basis highlights the edge and detail features of the image by calculating the differences between adjacent scale levels. This bandpass filtering feature enables the Laplacian pyramid to effectively separate and retain image information of different frequency components, providing an ideal feature representation for subsequent feature fusion. At the same time, bilinear interpolation is used for upsampling operations to ensure the smoothness of the image reconstruction process.
[0133] S3: Perform adaptive threshold segmentation and morphological processing on the fused feature map to obtain the final binary image:
[0134] S31: Calculate the adaptive local threshold of the fused feature map, specifically:
[0135]
[0136] Where Th(x,y) is the adaptive local threshold at the pixel point (x,y); is the mean of the fused feature map F in the 15×15 local window centered at the pixel point (x, y); is the standard deviation of the fused feature map F in a 15×15 local window centered at the pixel point (x, y); α is the threshold adjustment coefficient, which is 0.1 in this embodiment;
[0137] S32: Adaptively perform local threshold segmentation on the fused feature map, and obtain the initial binary map M0 by comparing the size relationship between the fused feature map F and the adaptive threshold Th at each pixel point (x, y), specifically:
[0138] When F(x,y)>Th(x,y), M0(x,y)=1;
[0139] When F(x,y)≤Th(x,y), M0(x,y)=0;
[0140] Wherein, M0(x, y) is the pixel value of the initial binary image M0 at the pixel point (x, y);
[0141] S33: Perform morphological processing on the initial binary image to obtain a final binary image; specifically:
[0142] S331: Perform opening operation:
[0143] An opening operation is performed on the initial binary image M0. The opening operation uses a 3×3 rectangular structure element to first perform an erosion operation on M0, and then perform an expansion operation to obtain an image M1 after the opening operation.
[0144] S332: Perform closing operation:
[0145] Perform a closing operation on M1. The closing operation also uses a 3×3 rectangular structure element to first perform an expansion operation on M1, and then perform an erosion operation to obtain an image M2 after the closing operation.
[0146] S333: Perform area filtering:
[0147] The area filtering process is performed on M2. The area filtering process is to calculate the pixel area of each connected area in the image M2, and set the pixel values of the pixels contained in the connected area with an area less than 20 to 0, so as to obtain the final binary image M2. final .
[0148] This step uses the method of adaptive threshold segmentation combined with morphological processing to achieve efficient conversion from fusion feature map to final binary map. By introducing the calculation mechanism of adaptive local threshold, the statistical characteristics of local areas of the image are fully considered, so that the segmentation process can better adapt to the brightness and contrast changes of different areas of the image. The larger local window size ensures the reliability of statistical features, and the reasonable threshold adjustment coefficient effectively controls the false detection rate while maintaining the detection sensitivity.
[0149] The binarization process in this step uses a simple and intuitive threshold comparison method, but the introduction of adaptive local threshold overcomes the limitations of the traditional fixed threshold method. This adaptive mechanism is particularly suitable for processing images with uneven illumination or local contrast differences, and can improve the robustness of the algorithm while maintaining detection accuracy.
[0150] S4: training the deep learning network to obtain a trained deep learning network;
[0151] S41: Calculate segmentation loss Loss seg and consistency constraint loss Lossconsist , get the total loss Loss total :
[0152] Loss seg =CE(Prob seg , Label);
[0153] Loss consist =MSE(Prob seg , Label);
[0154] Loss total =Loss seg +Loss consist ;
[0155] Among them, CE is the cross entropy loss; MSE is the mean square error; Label is the true segmentation label; Prob seg is the segmentation probability map;
[0156] S42: Use the stochastic gradient descent algorithm to train the parameters in the deep learning network to reduce the total loss; after reaching the set number of iterations, a trained deep learning network is obtained.
[0157] S5: Use the trained deep learning network to identify damage, obtain segmentation probability maps and annotate segmentation results to visualize the damaged area, including:
[0158] S51: Feature extraction and binary image guidance:
[0159] S511: Use the ResNet50 network to extract features from the denoised multimodal image sequence tensor, specifically:
[0160]
[0161] in, is the mth modality image in the denoised multimodal image sequence tensor Features extracted by ResNet50 network;
[0162] S512: The final binary image M final The spatial attention is guided by the tensor features of the denoised multimodal image sequence, specifically:
[0163]
[0164] in, is the mth modality image in the denoised multimodal image sequence tensor After M final The guided features;° is element-wise multiplication;
[0165] S52: And the final binary image M final After splicing in the feature dimension, the data is input into the Transformer module, and feature fusion is performed through the self-attention mechanism of the Transformer module. Specifically:
[0166]
[0167] Fea fuse = Transformer(Fea input );
[0168] Among them, Fea input is the input of the Transformer module layer; Fea fus e is the fused feature;
[0169] S53: Using UNet network for damaged pixel segmentation:
[0170] The fused features are input into the UNet network for damaged pixel segmentation, and the segmentation probability map is output through the Softmax function;
[0171] Prob seg =Softmax(UNet(Fea fuse ));
[0172] Among them, Prob seg is the segmentation probability map, and each pixel value in the segmentation probability map is the probability value of damage; Softmax is the normalized exponential function.
[0173] S54: Binarize the segmentation probability map to obtain a final segmentation mask, and annotate the segmentation result on the multimodal image sequence to achieve visualization of the damaged area.
[0174] This step achieves efficient feature extraction and fusion segmentation of multimodal images through deep learning networks. First, the classic deep convolutional network ResNet50 is used to extract features from each modality image, making full use of the network's advantages in image feature extraction. ResNet50 is a deep residual network. The design of residual connections not only enables the network to extract deeper feature information, but also effectively alleviates the gradient vanishing problem in deep network training, ensuring the effectiveness and stability of feature extraction.
[0175] This step innovatively introduces the binary image as the guiding information of the spatial attention mechanism, and achieves the selective enhancement of features by element-by-element multiplication. This guiding mechanism enables the network to focus on potential damage areas, effectively suppresses the interference of background areas, and improves the pertinence of feature extraction.
[0176] Embodiment 2: The present invention also discloses a steel-wood composite beam-column node damage identification system using machine vision, comprising the following five modules:
[0177] Denoising module: obtains visible light, near infrared and thermal imaging multimodal image sequences of steel-wood composite beam-column joints, performs denoising processing through adaptive anisotropic diffusion filtering, obtains denoised multimodal image sequences, and constructs denoised multimodal image sequence tensors based on the denoised multimodal image sequences;
[0178] Fusion module: Perform Gaussian pyramid decomposition and Laplacian pyramid feature fusion on the denoised multimodal image sequence tensor to obtain a fused feature map;
[0179] Binarization module: performs adaptive threshold segmentation and morphological processing on the fused feature map to obtain the final binary map;
[0180] Recognition model training module: train the deep learning network to obtain a trained deep learning network;
[0181] Identification model application module: Use the trained deep learning network to perform damage identification, obtain the segmentation probability map and annotate the segmentation results to realize the visualization of the damaged area.
[0182] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0183] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0184] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for identifying damage at steel-wood composite beam-column joints using machine vision, characterized in that: The following steps are involved: S1: Obtain visible light, near infrared and thermal imaging multimodal image sequences of steel-wood composite beam-column joints, perform denoising processing through adaptive anisotropic diffusion filtering, obtain denoised multimodal image sequences, and construct denoised multimodal image sequence tensors based on the denoised multimodal image sequences; S2: Perform Gaussian pyramid decomposition and Laplacian pyramid feature fusion on the denoised multimodal image sequence tensor to obtain a fused feature map; S3: Perform adaptive threshold segmentation and morphological processing on the fused feature map to obtain the final binary image; S4: training the deep learning network to obtain a trained deep learning network; S5: Use the trained deep learning network to identify damage, obtain segmentation probability maps and annotate segmentation results to visualize the damaged area, including: Use the trained deep learning network to extract features from the denoised multimodal image sequence tensor, and perform spatial attention guidance on the extracted features and the final binary image to obtain guided features; The guided features and the final binary image are concatenated in the feature dimension and input into the Transformer module. The features are fused through the self-attention mechanism of the Transformer module to obtain the fused features. The fused features are input into the trained deep learning network to perform damaged pixel segmentation and output a segmentation probability map; Annotate the segmentation results on the multimodal image sequences to visualize the damaged areas.
2. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 1, characterized in that: The S1 comprises the following steps: S11: Acquiring multimodal image sequences using sensor arrays of three different wavelengths I total , specifically: I total ={I vis ,I nir ,I thm }; Among them, I vis ,I nir and I thm These are modal images collected in the visible light band, near infrared band, and thermal imaging band; S12: Using iterative optimization strategy to optimize multimodal image sequence I total Adaptive anisotropic diffusion filtering is performed to denoise the multimodal image sequence I after denoising. den , specifically: in, and are the denoised pixel values at the pixel point (x, y) of the mth modality image in the denoised multimodal image sequence at the tth and t+1th iterations respectively; m = 1, 2, 3, are the modality numbers, and the first modality image, the second modality image and the third modality image are I vis ,I nir and I thm ; When t = 0, is the pixel value of the mth modality image at the pixel point (x, y) in the multimodal image sequence; η is the iteration step parameter; N is the set of four directions: up, down, left, and right; is the gradient value of the mth modality image at the pixel point (x, y) along the direction d; c d (x,y) is the diffusion coefficient of the pixel point (x,y) in direction d, specifically: Where K(x,y) is the adaptive local threshold at the pixel point (x,y): K(x,y)=μ(x,y)+λ·σ(x,y); Where μ(x,y) is The mean value of the gradient amplitude in a 7×7 local window centered at the pixel point (x, y); σ(x, y) is The standard deviation of the gradient amplitude in a 7×7 local window centered at the pixel point (x, y); λ is the adjustment parameter; The iteration termination condition is: Among them, ε is the convergence threshold; is the L2 norm of the t-th iteration result of the m-th modality image in the denoised multimodal image sequence; is the L2 norm of the difference between the t+1th and tth iteration results of the mth modal image; after the iteration is completed, the mth modal image in the denoised multimodal image sequence is obtained S13: Construct the denoised multimodal image sequence tensor T:
3. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 2 is characterized in that: The S2 comprises the following steps: S21: Perform Gaussian pyramid decomposition on the denoised multimodal image sequence tensor, specifically: in, is the pixel value of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) in the layer-th Gaussian pyramid decomposition result image; layer = 1, 2, 3, is the number of Gaussian pyramid layers; G sigma is a Gaussian kernel function with a standard deviation of 1.6; * is a convolution operation; It is the Gaussian pyramid decomposition result image of the mth modality image in the denoised multimodal image sequence tensor at the layer-1 level. When layer=1 ↓2 is a 2x downsampling operation in the horizontal and vertical directions; S22: Construct a Laplace pyramid, specifically: in, is the result of the Laplacian pyramid of the mth modality image in the denoised multimodal image sequence tensor at layer-1; Up is the bilinear interpolation upsampling operation, which expands the input image by 2 times in the horizontal and vertical directions; It is the Gaussian pyramid decomposition result image of the mth modality image in the denoised multimodal image sequence tensor at the layer; S23: Calculate the fusion weight, specifically: in, is the fusion weight of the mth modality image at the pixel (x, y) in the layer-1 layer and the denoised multimodal image sequence tensor; is the saliency map of the kth modality image at the pixel (x, y) in the denoised multimodal image sequence tensor at layer-1, k = 1, 2, 3; is the layer-1 layer, and the saliency map of the mth modality image in the denoised multimodal image sequence tensor at the pixel (x, y) is: Where Ω is the pixel point contained in the 7×7 neighborhood window centered on the pixel point (x, y), and (i, j) is the element in Ω; for The mean value of the pixel value in the 7×7 local window centered at the pixel point (x, y); S24: Laplacian pyramid feature fusion and reconstruction to obtain a fused feature map.
4. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 3 is characterized in that: The S24 includes the following steps: S241: Perform weighted fusion on the Laplacian pyramid to obtain the fused Laplacian pyramid: in, is the pixel value after fusion at the pixel point (x, y) of layer-1; is the pixel value of the mth modality image in the denoised multimodal image sequence tensor at the pixel point (x, y) in the result image of the layer-1 Laplacian pyramid; S242: Reconstruct the fused Laplacian pyramid to obtain the fused feature map F: Among them, Up layer-1 Indicates a layer-1 2x upsampling operation; F(x,y) is the pixel value of the fused feature map at the pixel point (x,y).
5. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 3 is characterized in that: The S3 comprises the following steps: S31: Calculate the adaptive local threshold of the fused feature map, specifically: Among them, Th(x,y) is the adaptive local threshold at the pixel point (x,y) of the fusion feature map; is the mean of the fused feature map F in the 15×15 local window centered at the pixel point (x, y); is the standard deviation of the fused feature map F in the 15×15 local window centered at the pixel point (x, y); α is the threshold adjustment coefficient; S32: Adaptively threshold segment the fused feature map, and compare the size relationship between the fused feature map F and the adaptive local threshold Th at each pixel point (x, y) to obtain the initial binary map M0, specifically: When F(x,y)>Th(x,y), M0(x,y)=1; When F(x,y)≤Th(x,y), M0(x,y)=0; Among them, M0(x,y) is the pixel value of the initial binary image M0 at the pixel point (x,y); S33: Perform morphological processing on the initial binary image to obtain a final binary image.
6. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 5 is characterized in that: The S33 includes the following steps: S331: Perform opening operation: An opening operation is performed on the initial binary image M0. The opening operation uses a 3×3 rectangular structure element to first perform an erosion operation on the initial binary image M0, and then perform an expansion operation to obtain an image M1 after the opening operation. S332: Perform closing operation: Perform a closing operation on M1. The closing operation also uses a 3×3 rectangular structure element to first perform an expansion operation on M1, and then perform an erosion operation to obtain an image M2 after the closing operation. S333: Perform area filtering: The area filtering process is performed on M2. The area filtering process is to calculate the pixel area of each connected area in the image M2, and set the pixel values of the pixels contained in the connected areas with an area less than 20 to 0, so as to obtain the final binary image M2. final .
7. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 6, characterized in that: The S4 comprises the following steps: S41: Calculate segmentation loss Loss seg and consistency constraint loss Loss consist , get the total loss Loss total : Loss seg =CE(Prob seg ,Label); Loss consist =MSE(Prob seg ,Label); Loss total =Loss seg +Loss consist ; Among them, CE is the cross entropy loss; MSE is the mean square error; Label is the true segmentation label; Prob seg is the segmentation probability map; S42: Use the stochastic gradient descent algorithm to train the parameters in the deep learning network to reduce the total loss; after reaching the set number of iterations, a trained deep learning network is obtained.
8. The method for identifying damage at the joints of steel-wood composite beams and columns using machine vision according to claim 7, characterized in that: The S5 comprises the following steps: The trained deep learning networks include ResNet50 network and UNet network; The ResNet50 network is used to extract features from the denoised multimodal image sequence tensor to obtain the denoised multimodal image sequence tensor features; The denoised multimodal image sequence tensor features are combined with the final binary image M final Perform spatial attention guidance to obtain guided multimodal image sequence tensor features; The guided multimodal image sequence tensor features and the final binary map M final After splicing in the feature dimension, it is input into the Transformer module for feature fusion to obtain the fused multimodal image sequence tensor feature Fea fuse ; The fused multimodal image sequence tensor features are input into the UNet network for damaged pixel segmentation to obtain the segmentation probability map Prob seg , where each pixel value of the segmentation probability map represents the probability that the position belongs to the damaged area; The segmentation probability map is binarized to obtain the final segmentation mask, and the segmentation results are annotated on the multimodal image sequence to achieve visualization of the damaged area.
9. A machine vision-based steel-wood composite beam-column joint damage identification system, characterized in that: include: Denoising module: obtains visible light, near infrared and thermal imaging multimodal image sequences of steel-wood composite beam-column joints, performs denoising processing through adaptive anisotropic diffusion filtering, obtains denoised multimodal image sequences, and constructs denoised multimodal image sequence tensors based on the denoised multimodal image sequences; Fusion module: Perform Gaussian pyramid decomposition and Laplacian pyramid feature fusion on the denoised multimodal image sequence tensor to obtain a fused feature map; Binarization module: performs adaptive threshold segmentation and morphological processing on the fused feature map to obtain the final binary map; Recognition model training module: train the deep learning network to obtain a trained deep learning network; Identification model application module: Use the trained deep learning network to identify damage, obtain segmentation probability maps and annotate segmentation results to visualize the damaged area; To realize the steel-wood composite beam-column node damage identification method using machine vision as described in any one of claims 1-8.
Citation Information
Patent Citations
Beam structure damage identification method based on support reaction force and strain
CN110487578A
Transformer substation safe operation monitoring method and system based on OpenCV
CN118485973A
Small target detection method and system based on multispectral image fusion
CN118675057A
Multi-modal medical image fusion and disease prediction method, computer program and terminal
CN118967480A
Image processing method and apparatus, computer device and storage medium
WO2024021134A1
Cited By
Aircraft surface damage area segmentation method based on adaptive multi-modal fusion
CN120635119A
An aircraft surface damage region segmentation method based on adaptive multi-modal fusion
CN120635119B
Real-time damage detection method and system for support hanger based on image processing
CN120852420A
A real-time damage detection method and system for support hangers based on image processing
CN120852420B