Steel pipe joint defect identification method based on cross-modal fusion technology

By using cross-modal fusion technology, the collaborative processing of image and ultrasonic signals was achieved, which solved the problem of limited recognition accuracy of single image data, improved the ability to identify multiple types of defects in steel pipe joints, especially the recognition effect of hidden defects, and balanced system load and stability.

CN120995269APending Publication Date: 2025-11-21GUANGZHOU MAYER CORP LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511087771.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing steel pipe joint defect identification technologies rely on single image data, making it difficult to accurately identify multiple types of defects, especially hidden defects, in complex environments. Furthermore, the lack of a collaborative processing mechanism for heterogeneous sensing data leads to high rates of missed and false detections, making it difficult to balance identification speed and accuracy.

Method used

By employing cross-modal fusion technology, image data and ultrasound signals are spatiotemporally aligned and feature fused. A deep feature extraction network and a cross-modal hybrid attention mechanism are used to adaptively evaluate signal reliability. Mutual information constraint loss and a category-specific untangling error correction structure are introduced to optimize features for defect identification.

Benefits of technology

It significantly improves the accuracy of identifying various types of defects, especially hidden defects, balances system load and algorithm stability, and ensures both recognition speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995269A_ABST
    Figure CN120995269A_ABST
Patent Text Reader

Abstract

The invention discloses a steel pipe joint defect identification method based on a cross-modal fusion technology, and the method comprises the steps: collecting image data and corresponding ultrasonic signals of the surface of a steel pipe joint, adding a time sequence label to construct original multi-modal data, and carrying out the rough alignment through a signal registration algorithm in combination with space-time correction; respectively extracting multilevel characteristic spectrums of the image and the ultrasonic mode by using an exclusive depth characteristic extraction network; a cross-modal mixed attention mechanism is introduced, the reliability weight of each modal signal is adaptively evaluated, the weight of an abnormal signal is automatically reduced or dynamic correction is triggered, and feature recombination is guided based on category-level prior; a mutual information constraint loss and category-specific unwrapping error correction structure is introduced to optimize recombination features; and performing defect identification based on the optimization features to obtain defect types and distribution information of the steel pipe joints. According to the method, through multi-modal data fusion and a cross-modal attention mechanism, the capability of distinguishing hidden defects is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial defect detection, in particular to a steel pipe joint defect recognition method based on cross-modal fusion technology. BACKGROUND

[0002] As the core component connecting pipelines in various pipeline systems, the structural integrity of the steel pipe joint is directly related to the operation safety of the entire pipeline system. In particular, in the fields of petroleum, chemical industry, municipal engineering, etc., if the defects (such as cracks, pores, and oxide scale adhesion) of the steel pipe joint cannot be accurately identified in time, it may cause serious safety accidents such as leakage, pressure imbalance, and even explosion. Therefore, efficient and accurate detection of steel pipe joint defects is a key link to ensure the reliability of the pipeline system.

[0003] At present, the steel pipe joint defect recognition technology mainly relies on single image data acquisition and analysis. This method acquires joint surface images through visual sensors and then uses image processing algorithms for defect detection. However, in practical applications, there are significant limitations: when facing complex processing environments (such as uneven lighting, surface oil covering, and processing texture interference), the recognition accuracy and category distinction of multiple types of defects such as cracks, pores, and oxide scale are obviously limited. Especially when multiple defects overlap or the defects are in hidden positions (such as the inside of the joint and the deep part of the weld seam), single image data cannot capture enough defect features, resulting in high rates of missed detection and false detection.

[0004] Meanwhile, in industrial detection scenarios, the system needs to balance recognition speed and multi-type defect detection accuracy. On the one hand, the continuity of the production line requires the detection process to be efficient and fast to avoid affecting the production rhythm. On the other hand, the subtle differences between multiple types of defects require the algorithm to have high feature discrimination ability, which makes the system face the outstanding contradiction between the load pressure caused by the complexity of the algorithm and the stability of the detection: complex algorithms can improve accuracy but easily lead to system response delay and high resource occupation, while simplified algorithms may sacrifice detection accuracy and fail to meet engineering requirements.

[0005] In addition, there is a lack of effective collaborative processing mechanism for heterogeneous perception data in existing technologies. Image data can intuitively reflect the surface defect morphology, but has limited representation ability for internal or deep defects. Other perception modalities (such as ultrasonic signals) can capture internal structural information, but it is difficult to accurately associate with surface features. Therefore, how to realize the efficient collaborative processing and information complementary of heterogeneous perception data and improve the discrimination ability of hidden defects has become a key problem to be solved in the field of steel pipe joint defect recognition. SUMMARY

[0006] The technical solution of the present application is as follows: a steel pipe joint defect recognition method based on cross-modal fusion technology, comprising:

[0007] S1, collect image data and corresponding ultrasonic signals of the surface of the steel pipe joint, add time sequence labels to the collected data, construct original multi-modal data, perform coarse alignment on the original multi-modal data stream by using a signal registration algorithm combined with space-time correction, and obtain coarse alignment multi-modal data;

[0008] S2, based on the coarse alignment multi-modal data, respectively use corresponding deep feature extraction networks to extract features of different modal pairs, and obtain comprehensive multi-level feature spectrum;

[0009] S3, introduce a cross-modal mixed attention mechanism to adaptively evaluate the reliability weight of each modal signal in the comprehensive multi-level feature spectrum, automatically weight or trigger dynamic correction of abnormal signals, and guide feature reorganization based on class-level priori, and obtain reorganized features;

[0010] S4, introduce mutual information constraint loss and class-specific unwinding error correction structure to optimize the reorganized features, and obtain optimized features;

[0011] S5, based on the optimized features, perform defect recognition to obtain the defect type and distribution information of the steel pipe joint.

[0012] Further, in step S1, collect image data and corresponding ultrasonic signals of the surface of the steel pipe joint, add time sequence labels to the collected data, construct original multi-modal data, and perform coarse alignment on the original multi-modal data stream by using a signal registration algorithm combined with space-time correction, and obtain coarse alignment multi-modal data;

[0013] Further, in step S1, the specific steps are:

[0014] Synchronously start the visual sensor and the acoustic sensor, and collect image data of a specific region of the surface of the steel pipe joint and an ultrasonic signal region corresponding to the specific region of the surface in space, construct original multi-modal data;

[0015] For the original multi-modal data, add time sequence labels including collection time and device coordinates to each frame of image and corresponding ultrasonic signal segment, and perform label association to obtain multi-modal data with time sequence labels;

[0016] Based on the multi-modal data with time sequence labels, use a cross-correlation coefficient registration algorithm to compare time sequences, calculate phase difference to eliminate time axis misalignment, and obtain time calibration multi-modal data;

[0017] Combine the preset space calibration parameters to perform space-time correction on the time calibration multi-modal data, map the image and ultrasonic data to the same coordinate system through a coordinate conversion matrix, correct the spatial misalignment, and obtain coarse alignment multi-modal data.

[0018] Further, in step S2, based on the coarse alignment multi-modal data, the multi-level visual feature spectrum of the image modality and the multi-level acoustic feature spectrum of the ultrasound modality are extracted respectively by using a dedicated deep feature extraction network, to obtain a comprehensive multi-level feature spectrum;

[0019] Further, in step S2, the specific steps are:

[0020] For the image modality in the coarse alignment multi-modal data, an improved residual network is used as a deep feature extraction network, and features are extracted hierarchically through multi-scale convolution kernels, and the image multi-level feature spectrum from low order to high order is output;

[0021] For the ultrasound signal modality in the coarse alignment multi-modal data, a deep feature extraction network based on a time convolution network is used, the time domain ultrasound signal is preprocessed and input into the time convolution network, and multi-scale time domain features are extracted by using causal convolution, dilated convolution and residual connection, and the ultrasound multi-level feature spectrum is output;

[0022] The image multi-level feature spectrum and the ultrasound multi-level feature spectrum are dimensionally adapted, mapped to the same feature dimension space through an adaptive dimension conversion layer, and the features at each level are spliced by weight through a channel fusion operation to form a comprehensive multi-level feature spectrum.

[0023] Further, in step S3, a cross-modal mixed attention mechanism is introduced to adaptively evaluate the reliability weight of each modality signal in the comprehensive multi-level feature spectrum, automatically reduce the weight of abnormal signals or trigger dynamic correction, and guide feature reorganization based on class-level priori to obtain reorganized features;

[0024] Further, in step S3, the specific steps are:

[0025] For the comprehensive multi-level feature spectrum, a cross-modal mixed attention mechanism is constructed, a spatial correlation matrix within the image features is calculated through a self-attention module, a modality correlation matrix between the image and ultrasound features is established through a cross-attention module, and an attention map with a fusion of spatial-modality two-dimensional dependence is formed;

[0026] Based on the attention map, a reliability weighting mechanism is designed to automatically reduce the weight of non-defect echo features in the ultrasound signal and noise regions in the image, while enhancing the weight of features in suspected defect regions, to form an adaptive feature importance distribution;

[0027] For the adaptive feature importance distribution, a dynamic correction mechanism is introduced, when the reliability of a certain modality feature is detected to be lower than a preset threshold, a feature compensation algorithm based on historical data is triggered to correct the current feature using the historical feature pattern of the region, and abnormal signal interference is suppressed;

[0028] The feature reorganization function is constructed based on the category level prior knowledge, the typical feature templates of different defect types are matched with the current features to calculate the matching degree, the features are reorganized and transformed according to the matching result, the feature components related to the specific defect type are highlighted, and the reorganized features for defect classification are formed.

[0029] Further, in step S4, the mutual information constraint loss and the category-specific unwinding error correction structure are introduced to optimize the reorganized features, and the optimized features are obtained.

[0030] Further, in step S4, the specific steps are as follows:

[0031] The mutual information value of the image modal feature and the ultrasound modal feature in the reorganized feature is calculated, a mutual information constraint loss function is constructed, and the mutual information constraint loss function is applied to the cross-modal feature.

[0032] For different defect categories, a category-specific unwinding error correction structure is constructed, the reorganized features are decomposed into category-related feature components and interference feature components, and the core features strongly related to the defect type are separated through the unwinding network.

[0033] An unwinding error correction loss term is designed to impose a penalty constraint on the separated interference feature components, and a bias correction is performed on the category-related feature components based on the pre-set defect category feature template.

[0034] The mutual information constraint loss and the unwinding error correction loss are fused to construct a joint optimization objective function, the parameters of the deep feature extraction network and the cross-modal attention mechanism are iteratively updated through back propagation, the reorganized features are iteratively optimized, and the optimized features are obtained.

[0035] Further, in step S5, based on the optimized features, the defect recognition is performed to obtain the defect type and distribution information of the steel pipe joint.

[0036] Further, in step S5, the specific steps are as follows:

[0037] The optimized features are input into the defect classifier, the probability distribution of each defect type is output through a fully connected network combined with a softmax activation function, the class with the highest probability is selected as the defect type, and the corresponding confidence is output.

[0038] Based on the spatial position coding in the optimized features, the features are mapped back to the original physical coordinate system of the steel pipe joint, the spatial coverage range of the defect features is determined through threshold segmentation, and the specific position coordinates of the defects on the steel pipe joint are located.

[0039] The intensity distribution of the optimized features is combined with the pre-set defect quantification rule to calculate the geometric parameters of the defect area.

[0040] Fusion of defect type, confidence, position coordinates and geometric parameters, generate identification result report, complete the comprehensive identification of steel pipe joint defect.

[0041] The steel pipe joint defect identification method based on the cross-modal fusion technology has the advantages that:

[0042] The application discloses a steel pipe joint defect identification method based on a cross-modal fusion technology. The method can fully utilize the complementary advantages of different modal data, and can significantly improve the identification accuracy of multiple types of defects such as cracks, pores and oxide skins in a complex processing environment. Especially for the superimposed or hidden defects that are difficult to identify by traditional methods, the application can effectively extract key features and suppress interference signals through the cross-modal hybrid attention mechanism and the category-specific unwinding error correction structure, and greatly improve the identification effect. Meanwhile, the application balances the contradiction between system load pressure and algorithm stability through the adaptive evaluation of the reliability weight of each modal signal and the dynamic correction mechanism, balances the identification speed and the detection accuracy of multiple types of defects, and guarantees the stable operation of the system. BRIEF DESCRIPTION OF DRAWINGS

[0043] Fig. 1 The application discloses a steel pipe joint defect identification method based on a cross-modal fusion technology.

[0044] Fig. 2 The application discloses a steel pipe joint defect identification method based on a cross-modal fusion technology. DETAILED DESCRIPTION

[0045] In order to make the purpose and advantages of the application clearer and more apparent, the application will be further described below with reference to the embodiments; it should be understood that the specific embodiments described herein are only used to explain the application, and do not limit the application.

[0046] The preferred implementation methods of the application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation methods are only used to explain the technical principles of the application, and do not limit the protection scope of the application.

[0047] Please refer to Figs. 1-2 The steel pipe joint defect identification method based on the cross-modal fusion technology comprises the following steps.

[0048] S1, collect image data and corresponding ultrasonic signals on the surface of the steel pipe joint, add time sequence labels to the collected data, construct original multi-modal data, perform coarse alignment on the original multi-modal data stream by using a signal registration algorithm combined with space-time correction, and obtain coarse alignment multi-modal data;

[0049] S2, based on the coarse alignment multi-modal data, respectively extract features of different modal pairs by using corresponding deep feature extraction networks, and obtain comprehensive multi-level feature spectrum;

[0050] S3, introduce a cross-modal mixed attention mechanism to adaptively evaluate the reliability weight of each modal signal in the comprehensive multi-level feature spectrum, automatically weight or trigger dynamic correction of abnormal signals, and guide feature reorganization based on class-level priori to obtain reorganized features;

[0051] S4, introduce mutual information constraint loss and class-specific unwinding error correction structure to optimize the reorganized features, and obtain optimized features;

[0052] S5, based on the optimized features, perform defect recognition to obtain the defect type and distribution information of the steel pipe joint.

[0053] As shown in Fig. 2 In step S1, image data and corresponding ultrasonic signals on the surface of the steel pipe joint are collected, time sequence labels are added to the collected data, original multi-modal data is constructed, and coarse alignment multi-modal data is obtained by using a signal registration algorithm combined with space-time correction.

[0054] Specifically, in this embodiment, the visual sensor and the acoustic sensor are simultaneously started for data collection, and the synchronization accuracy is controlled within ±2 microseconds. The visual sensor used is an industrial area camera with a resolution of 2560x1920, equipped with an 8mm fixed-focus lens, installed on an adjustable bracket, with the lens optical axis perpendicular to the surface of the steel pipe joint. The shooting range covers the joint welding area and a range of 60mm on both sides, and the frame rate is set to 30 frames / second to ensure clear capture of surface details. The acoustic sensor uses an 8-channel linear array ultrasonic probe with a center frequency of 5MHz. The probe array length matches the width range of the camera shooting, and is tightly attached to the surface of the steel pipe joint through a coupling agent. Each channel corresponds to a specific lateral area in the camera field of view, ensuring that the ultrasonic signal area collected and the specific area on the surface in the image form a one-to-one correspondence in spatial position. The two sensors collect data in real time, of which the image data is stored in a lossless format and the ultrasonic signal is saved in the form of original time domain waveform, together constituting the original multi-modal data.

[0055] The acquired data is labeled with time sequence tags and associated. For each image frame, a tag with a unique identifier is generated. The tag content includes the acquisition time accurate to milliseconds, the three-dimensional coordinates (X, Y, Z) of the camera in the preset coordinate system, and the storage path of the image frame. For the ultrasound signal segment corresponding to the image, each segment contains 2048 sampling points. Similarly, a tag containing the same acquisition time, the three-dimensional coordinates of the ultrasound probe, and the storage path of the signal segment is generated. Using the acquisition time as the association basis, the image tag is bound to the corresponding ultrasound signal segment tag, establishing a one-to-one mapping relationship and forming multimodal data with time sequence tags, which facilitates subsequent data retrieval and processing.

[0056] Time calibration is performed on time-labeled multimodal data. The cross-correlation coefficient registration algorithm is used to compare and analyze the timestamp information of image frame sequences and ultrasound signal segment sequences. The cross-correlation coefficient between the two sequences is calculated. When the coefficient value is greater than 0.85, it is determined that there is a strong time correlation. The phase difference between the two sequences is determined based on the cross-correlation result. The time axis is adjusted by linear interpolation to eliminate the time axis misalignment caused by sensor start-up delay or data transmission time difference, so that the image and the corresponding ultrasound signal are synchronized in the time dimension, and time-calibrated multimodal data is obtained.

[0057] Then, spatiotemporal correction is performed by combining the preset spatial calibration parameters. The spatial calibration parameters are obtained through previous calibration experiments. Specifically, a standard calibration board is used to obtain the camera's intrinsic parameter matrix, distortion coefficients, and spatial position parameters of each channel of the ultrasonic probe. The transformation relationship between image pixel coordinates and the physical coordinates of ultrasonic detection points is established. Based on these parameters, a coordinate transformation matrix is ​​constructed to map the time-calibrated image data and ultrasonic data to the same world coordinate system. Matrix operations are used to correct spatial misalignment caused by sensor installation position deviation or the surface curvature of the steel pipe joint, so that the spatial deviation between each pixel in the image and the corresponding ultrasonic detection area is controlled within ±0.5mm. Finally, coarsely aligned multimodal data is obtained, laying the foundation for subsequent feature extraction work.

[0058] like Fig. 2 As shown, in step S2, based on the coarsely aligned multimodal data, features are extracted from different modal pairs using the corresponding deep feature extraction networks to obtain a comprehensive multi-level feature spectrum.

[0059] Specifically, in this embodiment, for the image modality in the coarse alignment multi-modal data, a modified residual network (ResNet-50) is used as the deep feature extraction network. The network is optimized on the basis of the traditional residual network, and a multi-scale convolution kernel structure is introduced. In the initial stage of the network, three different size convolution kernels of 1x1, 3x3 and 7x7 are set. The small size convolution kernel is responsible for capturing the details and textures in the image, and the large size convolution kernel is used to extract the global contour information. Through a hierarchical feature extraction process, low-order to high-order image features are output from the shallow layer to the deep layer of the network: the low-order features mainly include the edge, color change and other basic visual information of the surface of the steel pipe joint, which can reflect the subtle defects on the surface; the middle-order features are a preliminary integration of the low-order features, which can identify the shape features of the local area; the high-order features further integrate the global information, which can express the overall distribution trend of the defects. Finally, these different levels of features jointly constitute the multi-level feature spectrum of the image, which fully retains the key information in the image modality;

[0060] For the ultrasonic signal modality in the coarse alignment multi-modal data, a deep feature extraction network based on a time convolution network (TCN) is used. First, the time domain ultrasonic signal is preprocessed, which includes two steps: first, the low-frequency (50Hz / 60Hz) power frequency interference introduced by the sensor power supply is suppressed through adaptive band-stop filtering (for signal baseline drift), and then the wavelet threshold denoising method is used to reduce the high-frequency random noise (1-10MHz frequency band), to ensure the purity of the signal. The preprocessed ultrasonic signal is input into the TCN. The network uses causal convolution to ensure the time sequence of signal processing and avoid interference of future information on current feature extraction. Through dilated convolution, the receptive field is expanded, which can capture signal features of different time scales. At the same time, residual connection is introduced to alleviate the gradient vanishing problem of deep network and enhance the transmission efficiency of features. Through the synergistic effect of these structures, the network extracts features of the ultrasonic signal from low-order to high-order. The low-order features reflect the instantaneous fluctuations of the signal, and the high-order features reflect the periodic changes and abnormal patterns of the signal. Finally, the multi-level feature spectrum of the ultrasonic signal is formed.

[0061] After obtaining the image multi-level feature spectrum and the ultrasonic multi-level feature spectrum, dimension adaptation and fusion are performed. Dimension adaptation is realized through an adaptive dimension conversion layer, which contains multiple fully connected layers and BatchNorm layers. It can automatically adjust parameters according to the dimensions of the two modal features. For the image feature spectrum, the conversion layer compresses or expands the dimensions of the features at different levels to meet the unified dimension requirement. For the ultrasonic feature spectrum, the conversion layer is also used for processing to ensure that the features of the two modalities are consistent in dimension.

[0062] After dimensional adaptation is completed, channel fusion is performed. During fusion, weights are assigned based on the information entropy of each level of features. Features with higher information entropy contain more effective information and are assigned greater weights. By concatenating features at each level according to weights, the features of image and ultrasound modalities are organically combined to form a comprehensive multi-level feature spectrum. This feature spectrum not only retains the unique information of the two modalities but also achieves information complementarity, providing high-quality input for subsequent cross-modal attention mechanisms.

[0063] like Fig. 2 As shown, in step S3, a cross-modal hybrid attention mechanism is introduced to adaptively evaluate the reliability weights of each modal signal in the integrated multi-level feature spectrum, automatically reduce the weights of abnormal signals or trigger dynamic correction, and obtain recombined features based on category-level prior guidance.

[0064] Specifically, in this embodiment, a cross-modal hybrid attention mechanism is constructed. For the comprehensive multi-level feature spectrum, a structure combining multi-head self-attention and cross-attention is adopted: the self-attention module has 8 attention heads, which divide the image features into 8 groups according to the channels. The spatial similarity of each group is calculated through the Query, Key, and Value matrix. The calculation adopts the scaling dot product attention formula to generate an 8×(H×W)×(H×W) spatial correlation matrix (H and W are the spatial dimensions of the image features), capturing the local and global dependencies between pixels, such as the continuous edge correlation of cracks;

[0065] The cross-attention module also has 8 heads, with image features as queries and ultrasound features as keys and values. It calculates an 8×1024×1024 modal correlation matrix to quantify the complementarity of the two modal features, such as the matching degree between surface image defects and internal ultrasound echoes. The spatial correlation matrix and the modal correlation matrix are concatenated by channel and normalized by Softmax to form a fused spatial-modal dual-dimensional attention map with dimensions of 8×(H×W+1024)×(H×W+1024).

[0066] A reliability weighting mechanism is designed, and the feature reliability is evaluated based on the weight mean of the attention map. The average attention weight of each feature channel is calculated, and the reliability threshold is set to 0.6 (determined by 500 sample tests, and features with a value lower than the threshold are determined as low reliability features). For non-defect echoes in the ultrasonic signal that meet the signal-to-noise ratio SNR<6dB and the duration >3ms, such as the reflection signal at the steel pipe joint interface, the weight of the corresponding feature channel is multiplied by a decay coefficient of 0.2. For noise regions in the image with a gray scale standard deviation <20 and an area >80 pixels, such as the surface oil stain coverage area, the weight of the corresponding feature channel is multiplied by a decay coefficient of 0.3. For suspected defect edge regions in the image with a gradient change rate >0.5 (calculated by the Sobel operator), the feature channel weight is multiplied by an enhancement coefficient of 1.3. For abnormal segments in the ultrasonic signal with an amplitude mutation rate >30% (the ratio of the amplitude difference to the average value of adjacent sampling points), such as defect reflection pulse signals, the feature channel weight is multiplied by an enhancement coefficient of 1.4. Through the above processing, an adaptive feature importance distribution is formed, highlighting the contribution of high-value features.

[0067] A dynamic correction mechanism is introduced. When the reliability mean of a certain modal feature is lower than the threshold value 0.6 for 3 consecutive frames, the feature compensation process is triggered: a preset historical feature database (containing 1000+ normal and defect features of the same type of steel pipe joint, stored according to diameter and wall thickness) is called, the matching degree of the current feature and the historical feature is calculated by cosine similarity, the threshold is set to 0.8, the 3 historical features with the highest similarity are selected as references, and the current feature is corrected using weighted average method, where the current feature weight accounts for 0.6, and the historical features are weighted and summed according to similarity (weight accounts for 0.4), to suppress abnormal signals caused by temporary sensor failure or environmental interference, such as ultrasonic signal distortion caused by poor probe coupling;

[0068] A feature reorganization function is constructed based on category-level prior knowledge. First, 5 typical defect (crack, porosity, incomplete fusion, slag inclusion, and scale) templates are generated by training 10000+ labeled samples, each template is a 1024-dimensional vector representing the average feature of the defect, and a 1024x1024 projection matrix corresponding to each defect is obtained by training, which is used to decompose the comprehensive multi-level feature spectrum into class-related components.

[0069] For the current comprehensive multi-level feature spectrum, feature decomposition is performed by the projection matrix to obtain defect-related feature components. The Euclidean distance between the feature components and the corresponding category typical feature templates is calculated, and the matching degree (range 0-1) is obtained by normalization. For example, the matching degree between the crack-related component in the current feature and the crack template is 0.85, and the matching degree between the porosity-related component and the porosity template is 0.3. The feature component weight is allocated by the reorganization function according to the matching degree, and the reorganization formula is:

[0070]

[0071] wherein F C is the output feature vector after recombination, m i is a match degree representing the current feature and the typical feature template of the i-th defect class, F i is a feature component in the current comprehensive feature related to the i-th defect class;

[0072] The component weight of the defect type with the highest match degree is increased to 1.2 times the original weight, and the remaining components are scaled according to the match degree. Finally, a 1024-dimensional recombination feature is output, highlighting the feature components strongly related to the target defect.

[0073] As Fig. 2 shown, in step S4, the recombination feature is optimized by introducing a mutual information constraint loss and a class-specific unwinding error correction structure to obtain an optimized feature.

[0074] Specifically, in the present embodiment, the mutual information value of the image modal and ultrasound modal features in the recombination feature is calculated and a constraint loss is constructed. Mutual information is used to quantify the dependence of the two modal features. The higher the value, the closer the correlation between the features. In the calculation, the marginal probability distribution of the image features and the ultrasound features is first constructed by kernel density estimation. For the estimation of the joint probability distribution, considering the high-dimensional characteristics of the recombination feature, a neural network-based mutual information estimator (such as the MINE algorithm) is used. A light-weight discriminator network (containing 3 fully connected layers with hidden dimensions of 512, 256, and 1) takes the concatenation of the image features and the ultrasound features as input and outputs the probability of the two features coming from the joint distribution. By minimizing the discriminator loss, the joint probability distribution is approximated. Based on the estimation results of the marginal probability distribution and the joint probability distribution, the mutual information value is calculated. When the mutual information value is low, the mutual information constraint loss function value increases, forcing the network to improve the correlation between the cross-modal features through backpropagation, ensuring that the observed defect features in the image and the internal features reflected by the ultrasound signal form consistent correlations, such as strong matching between the visual features of surface cracks and the abnormal signals of ultrasound echoes.

[0075] A specific unwinding error correction structure is constructed for different defect classes. An independent unwinding network is designed for each defect class (such as cracks, pores, incomplete fusion, slag inclusion, and scale). The network contains 3 fully connected layers (with hidden dimensions of 512, 256, and 128) and BatchNorm layers. Through nonlinear transformation, the recombination feature is decomposed into two types of components: class-related feature components, which are strongly related to the current defect type and account for about 70%, and interference feature components, which are composed of noise, background texture, and other irrelevant information, accounting for about 30%. Taking the crack defect as an example, the unwinding network will mainly separate the features related to the crack direction and depth, filter out irrelevant interference such as scale, and ensure the purity of the core features.

[0076] For the separated interference feature components, L2 regularization is used to impose a penalty constraint on them. By reducing the overall energy of the component, the influence of irrelevant information on subsequent defect identification is suppressed. For the core feature components related to the category, they are compared with the preset defect category feature template. The template is a typical pattern obtained by statistical averaging of the core features of a large number of similar samples in the training set (such as the crack template containing linear edge features in the image and continuous echo features in the ultrasonic signal). By minimizing the difference between the current feature and the template, the current feature is forced to move closer to the typical pattern of this type of defect, thereby improving the feature's ability to distinguish different defect types.

[0077] A joint optimization objective function is constructed and iteratively trained. The mutual information constraint loss, interference feature penalty loss, and core feature correction loss are weighted and summed according to a certain ratio to form a comprehensive optimization objective. The weighted sum is:

[0078] L=α·L1+β·L2+γ·L3

[0079] Where L is the total loss function, L1 is the mutual information constraint loss, L2 is the interference feature penalty loss, L3 is the core feature correction loss, α is the mutual information constraint loss weight, β is the interference feature penalty loss weight, and γ is the core feature correction loss weight.

[0080] The weights of each loss term were determined as follows: 50 labeled samples covering various types of defects were selected, and a 5-fold cross-validation method was used. The weight combination was adjusted in increments of 0.1 (range 0.5-1.5) for each validation. The defect identification accuracy (target ≥95%) and cross-modal consistency (target mutual information value ≥0.7) on the validation set were used as optimization indicators.

[0081] The Adam optimizer is used to iteratively optimize the objective function. An initial learning rate and decay strategy are set, and the parameters of the deep feature extraction network and cross-modal attention mechanism are continuously adjusted through backpropagation. This allows the features to retain key information about defects while gradually reducing interference and enhancing cross-modal consistency. The iterative process continues until the loss value stabilizes and converges (the fluctuation is less than the set threshold for multiple consecutive rounds). The final optimized features are then tested and verified.

[0082] like Fig. 2 As shown, in step S5, defect identification is performed based on the optimized features to obtain the defect type and distribution information of the steel pipe joint.

[0083] Specifically, in the present embodiment, the defect type and confidence are determined by the defect classifier, and the optimized features obtained in step S4 are input into the preset defect classifier. The classifier includes a 3-layer fully connected network (the hidden layer dimensions are 256, 128, and 64, respectively), and a Dropout layer (dropout rate 0.3) is arranged after each layer to prevent overfitting. After the features are processed by the fully connected network, the probability distribution of 5 defect types (crack, pore, incomplete fusion, slag inclusion, and scale) is converted by a softmax activation function, the sum of the probability values is 1, the class with the highest probability is selected as the final defect type, and the highest probability value is taken as the recognition confidence. For example, if the output probability is: crack 0.92, pore 0.05, and others 0.03, the defect type is determined to be crack, and the confidence is 0.92. To ensure the reliability of recognition, when the confidence is lower than 0.75, it is marked as a suspected defect and manual review is prompted.

[0084] The spatial encoding is used to realize accurate positioning of the defect position. The optimized features include the spatial position encoding consistent with the spatial correction in step S1. The encoding records the physical coordinate information corresponding to the features. By analyzing the encoding information, the features are mapped back to the original physical coordinate system of the steel pipe joint, and the coordinate accuracy is retained to 0.1 mm. An adaptive threshold segmentation method is used to automatically determine the segmentation threshold based on the feature intensity distribution. The area with a feature intensity higher than the threshold is determined as the defect coverage range. The coordinates of the minimum bounding rectangle of the defect (such as the upper left corner (X1, Y1, Z1) and the lower right corner (X2, Y2, Z2)) are output, and the center coordinates are calculated as the representative position of the defect. For example, the center coordinates of a certain crack defect are (152.3 mm, 45.6 mm, 2.1 mm) after mapping, and the coverage range is a rectangular area with a length of 8.5 mm and a width of 1.2 mm.

[0085] The defect geometric parameters are calculated based on the feature intensity distribution. The preset defect quantification rules are as follows: for image features, the actual size is converted from the pixel range of the intensity distribution (1 pixel corresponds to 0.05 mm), and the length (the longest axis distance) and area (pixel number x unit area) of the defect are calculated; for ultrasonic features, the defect depth is estimated by an intensity attenuation model (the higher the intensity, the smaller the depth, and the calibration coefficient is determined based on 50 standard defect samples). Taking a pore defect as an example, its diameter is calculated to be 2.3 mm based on the image features, and the depth is estimated to be 1.5 mm based on the ultrasonic features. The final output geometric parameters include key indicators such as length, width, area, and depth, with an error controlled within ±0.2 mm.

[0086] The fusion multi-dimensional information generates an identification result report, which contains three core contents: first, a defect type label, which clearly labels the name and confidence of the identified defect, such as a crack with a confidence of 0.92; second, a spatial distribution heat map, which uses different colors (red for high confidence and yellow for medium confidence) to label the specific location and coverage of the defect on the steel pipe joint based on the world coordinate system in step S1, and the heat map resolution is consistent with the original image (0.05 mm / pixel); third, a quantitative parameter table, which lists the geometric parameters (length, width, area, depth) of the defect, detection time, and equipment number, etc. information; the report supports two output forms: visual interface real-time display (for detection personnel to check immediately) and PDF file archiving (contains original data link for traceability), and finally completes the comprehensive identification of the steel pipe joint defects.

[0087] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to related technical features without departing from the principles of the present application, and the technical solutions after such changes or replacements will fall within the protection scope of the present application.

[0088] The above description is only the preferred embodiments of the present application and is not intended to limit the present application; for those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A steel pipe joint defect recognition method based on a cross-modal fusion technology, characterized by, The method comprises the following steps: S1, collecting image data and corresponding ultrasonic signals of the surface of the steel pipe joint, adding time sequence labels to the collected data, constructing original multi-modal data, using a signal registration algorithm combined with space-time correction to coarsely align the original multi-modal data stream to obtain coarsely aligned multi-modal data; S2, based on the coarsely aligned multi-modal data, respectively using corresponding deep feature extraction networks to extract features of different modal pairs to obtain comprehensive multi-level feature spectrum; S3, introducing a cross-modal mixed attention mechanism to adaptively evaluate the reliability weight of each modal signal in the comprehensive multi-level feature spectrum, automatically weighting or triggering dynamic correction of abnormal signals, and based on the category level priori to guide feature reorganization to obtain reorganized features; S4, introducing mutual information constraint loss and class-specific unwinding error correction structure to optimize the reorganized features to obtain optimized features; S5, based on the optimized features, defect recognition is performed to obtain the defect type and distribution information of the steel pipe joint.

2. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 1, characterized in that: The S1 step specifically comprises: Collecting image data of a specific area of the surface of the steel pipe joint and an ultrasonic signal region corresponding to the specific area of the surface in space, and constructing original multi-modal data; For each frame of image and corresponding ultrasonic signal segment, a time sequence label including acquisition time and device coordinates is added, and label association is performed to obtain multi-modal data with time sequence labels; Based on the multi-modal data with time sequence labels, a cross-correlation coefficient registration algorithm is used to compare time sequences, calculate phase difference to eliminate time axis misalignment, and obtain time-calibrated multi-modal data; The time-calibrated multi-modal data is combined with a preset space calibration parameter for space-time correction, the image and ultrasonic data are mapped to the same coordinate system through a coordinate conversion matrix, the spatial misalignment is corrected, and coarsely aligned multi-modal data is obtained.

3. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 1, characterized in that: The S2 step specifically comprises: For the image modality in the coarsely aligned multi-modal data, an improved residual network is used as a deep feature extraction network, multi-scale convolution kernel hierarchical feature extraction is performed, and low-order to high-order image multi-level feature spectrum is outputted; For the ultrasonic signal modality in the coarsely aligned multi-modal data, a deep feature extraction network based on a time convolution network is used, the time domain ultrasonic signal is preprocessed and inputted into the time convolution network, multi-scale time domain features are extracted by using causal convolution, dilated convolution and residual connection, and ultrasonic multi-level feature spectrum is outputted; The image multi-level feature spectrum and the ultrasonic multi-level feature spectrum are dimensionally adapted, mapped to the same feature dimension space through an adaptive dimension conversion layer, and each level of feature is spliced by weight through a channel fusion operation to form a comprehensive multi-level feature spectrum.

4. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 3, characterized in that: The improved residual network is ResNet-50 as a deep feature extraction network, which hierarchically extracts low-order to high-order image multi-level feature spectrum through 1×1, 3×3 and 7×7 convolution kernel.

5. The method of claim 1, wherein the method is based on a cross-modal fusion technique. The S3 step specifically comprises: A cross-modal mixed attention mechanism is constructed for the comprehensive multi-level feature spectrum, a spatial correlation matrix within the image features is calculated through self-attention, a modal correlation matrix between the image and ultrasound features is established through cross-attention, and an attention atlas that fuses spatial-modal two-dimensional dependence is formed; Based on the attention atlas, a reliability weighting mechanism is designed to automatically reduce the weight of non-defect echo features in the ultrasound signal and noise regions in the image, while enhancing the weight of features in suspected defect regions, forming an adaptive feature importance distribution; A dynamic correction mechanism is introduced for the adaptive feature importance distribution. When the reliability of a certain modal feature is detected to be lower than a preset threshold, a feature compensation algorithm based on historical data is triggered to correct the current feature using the historical feature pattern of the region; Based on the class-level prior knowledge, a feature reorganization function is constructed to calculate the matching degree between the typical feature templates of different defect types and the current features, reorganize and transform the features according to the matching results, highlight the feature components related to specific defect types, and form reorganized features for defect classification.

6. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 5, characterized in that: The typical feature templates of different defect types include typical feature templates of cracks, pores, incomplete fusion, slag inclusion, and scale.

7. The method of claim 1, wherein the method is based on a cross-modal fusion technique. The S4 step specifically includes: Calculate the mutual information value of the image modal feature and the ultrasound modal feature in the reorganized feature, construct a mutual information constraint loss function, and apply the mutual information constraint loss function to the cross-modal feature; For different defect categories, a category-specific unwinding error correction structure is constructed to decompose the reorganized feature into category-related feature components and interference feature components, and separate the core features strongly related to the defect type through an unwinding network; An unwinding error correction loss term is designed to impose a penalty constraint on the separated interference feature components, and based on the preset defect category feature templates, a bias correction is performed on the category-related feature components; The mutual information constraint loss and the unwinding error correction loss are fused to construct a joint optimization objective function, the parameters of the deep feature extraction network and the cross-modal attention mechanism are updated iteratively through back propagation, the reorganized feature is iteratively optimized, and the optimized feature is obtained.

8. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 7, characterized in that: The joint optimization objective function is composed of the mutual information constraint loss, the interference feature penalty loss, and the core feature correction loss.

9. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 1, characterized in that: The S5 step specifically includes: The optimized feature is input into a defect classifier, the probability distribution of each defect category is output through a fully connected network combined with a softmax activation function, the category with the highest probability is selected as the defect type, and the corresponding confidence is output; Based on the spatial position coding in the optimized feature, the feature is mapped back to the original physical coordinate system of the steel pipe joint, the spatial coverage range of the defect feature is determined through threshold segmentation, and the specific position coordinates of the defect on the steel pipe joint are located; The intensity distribution of the optimized feature is combined with the preset defect quantification rules to calculate the geometric parameters of the defect region; The defect type, confidence, position coordinates, and geometric parameters are fused to generate an identification result report.

10. The steel pipe joint defect identification method based on cross-modal fusion technology according to claim 9, characterized in that: The defect identification result report contains a defect type label, a spatial distribution heat map, a quantitative parameter table, and supports visual interface display and PDF file archiving.

Citation Information

Cited By

  • Pavement slab defect nondestructive testing method fusing ultrasonic pulse echoes and visual information

    CN121476234A

  • Intelligent voice interaction system and method

    CN121583247A

  • Pile slot sediment thickness detection method, system and equipment based on magnetostriction signal inversion

    CN122107920A