Multi-mode-based steel box girder complex weld defect detection system

By working in concert with multimodal data acquisition, processing and intelligent recognition modules, the problems of blind spots and low automation in the inspection of complex welds in steel box girders have been solved, and high-precision, full-coverage inspection and digital management of complex welds have been achieved.

CN121499657APending Publication Date: 2026-02-10JSTI GRP CO LTD +2

Patent Information

Application Number
CN202610031255.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies for inspecting complex welds in steel box girders suffer from problems such as blind spots in single inspection modes, difficulty in effectively integrating multi-source data, and low levels of automation and digitization in the inspection process.

Method used

A multimodal data acquisition module is used to simultaneously acquire ultrasonic, X-ray and visual image data. The data processing and intelligent recognition module performs spatiotemporal registration and feature-level fusion. A deep learning model is used to identify, locate and quantify weld defects. Combined with a control and integration platform, automated scanning and integrated output of results are achieved.

Benefits of technology

It has achieved full coverage inspection of complex welds in steel box girders, improved the identification rate and positioning accuracy of hidden defects and mixed defects, reduced the rate of missed and false detections, and realized the full automation of the inspection process and the digital management of inspection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121499657A_ABST
    Figure CN121499657A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of nondestructive testing, in particular to a multi-modal-based steel box girder complex weld defect detection system which comprises a multi-modal data acquisition module, a data processing and intelligent identification module and a control and integration platform. The multi-modal data acquisition module comprises an ultrasonic phased array probe, an X-ray imaging sensor, a visual imaging subunit and a positioning unit, and is used for synchronously acquiring ultrasonic, ray, visual images and spatial position information of a welding seam; the data processing and intelligent identification module is used for carrying out space-time registration and feature level fusion on the multi-modal data and realizing defect identification, positioning and quantification by utilizing a deep learning model; and the control and integration platform drives the movement mechanism to realize automatic scanning and provides result visualization and digital twin interfaces. According to the invention, the detection precision, efficiency and reliability of the internal and surface defects of the complex structure of the steel box girder can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nondestructive testing technology, and in particular to a multimodal steel box girder complex weld defect detection system. Background Technology

[0002] In the early stages, the assessment of weld quality in steel box girders mainly relied on single inspection methods such as manual visual inspection, radiographic testing, and conventional ultrasonic testing. Among these, manual visual inspection can only detect obvious surface defects and is inefficient and highly subjective; while radiographic testing can provide intuitive internal images, it poses radiation safety risks, is costly, and lacks sufficient sensitivity in detecting surface defects such as cracks; conventional ultrasonic testing is highly dependent on the operator's experience, has low scanning coverage, is difficult to effectively inspect complex structures such as U-rib fillet welds, and the test results are not intuitive and difficult to digitally record.

[0003] With the advancement of nondestructive testing (NDT) technology, advanced methods such as ultrasonic phased array, digital X-ray imaging, and machine vision have been gradually applied to bridge inspection. Phased array technology improves inspection flexibility and efficiency by electronically controlling the deflection and focusing of the sound beam; digital X-ray improves imaging speed; and machine vision enables automated identification of surface defects. However, most of these methods use only one technology for independent inspection and analysis. In practical applications, the defect morphology of complex welds in steel box girders is diverse and their spatial location is complex. Single inspection methods have inherent limitations: ultrasound is not sensitive to surface-opening micro-cracks, vision cannot detect internal defects, and X-ray imaging has strict requirements for defect orientation and requires bulky equipment. As a result, the inspection results are not comprehensive, and the risk of missed or false detections remains high.

[0004] Currently, although some studies have attempted to combine various detection technologies in a simple way, most of them remain at the level of separate detection and result comparison, failing to achieve deep integration and collaborative analysis at the data level. The key problems are: the lack of an effective spatiotemporal registration mechanism for multi-source heterogeneous data; the lack of a fusion strategy that can adaptively weight different modal features and give full play to the complementary advantages of information; insufficient automation and intelligence in the detection process, still relying on manual interpretation and equipment operation; and the disconnect between the detection results and the bridge's digital management model, failing to form structured data that can be used for subsequent operation and maintenance decisions.

[0005] Chinese invention patent CN118961875A discloses a defect detection device and method for steel beam box girder construction sites. This invention solves the problems of cumbersome centering and difficulty in batch detection in traditional blind hole methods by using device positioning, visual alignment, drilling tool calibration, step-by-step blind hole detection, and real-time position adjustment. However, this invention does not involve the core objective of weld defect identification and the corresponding multimodal nondestructive testing technology. It also lacks intelligent processing modules for multimodal data spatiotemporal registration, feature-level fusion, and hybrid deep learning models. Furthermore, it lacks the linkage function between digital twin data interfaces and BIM models. The technical approach focuses on residual stress measurement rather than defect identification.

[0006] Therefore, this invention discloses a multimodal steel box girder complex weld defect detection system. Summary of the Invention

[0007] The purpose of this invention is to address the technical shortcomings of existing technologies in the inspection of complex welds in steel box girders, such as blind spots in single detection modes, difficulty in effectively integrating multi-source data, and low levels of automation and digitization in the inspection process. In response, this invention proposes a multimodal steel box girder complex weld defect detection system.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: a multimodal steel box girder complex weld defect detection system, comprising:

[0009] The multimodal data acquisition module is used to simultaneously acquire ultrasonic data, radiographic images, and visual images of the U-rib fillet welds or butt welds of the transverse diaphragms of the steel box girder, and simultaneously record the spatial pose information during acquisition.

[0010] The data processing and intelligent recognition module is used to perform spatiotemporal registration and feature-level fusion on the collected data, and to use deep learning models to identify, locate and quantify weld defects.

[0011] The control and integration platform, based on the identification, location and quantification of weld defects, automatically controls the multimodal data acquisition module to complete the automated scanning of the weld and integrates and outputs the identification results.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] This invention constructs a multimodal data acquisition module that integrates ultrasonic, X-ray, visual, and positioning units, and enables it to operate under unified synchronous triggering. This allows for the simultaneous acquisition of comprehensive physical information of U-rib fillet welds or butt welds of transverse diaphragms in steel box girders, thereby achieving full coverage of volumetric, area-based, and surface defects at the data source and fundamentally eliminating the detection blind spots of single sensors.

[0014] This invention, by setting up a data processing and intelligent recognition module, performs spatiotemporal registration and feature-level fusion on synchronously collected multi-source data. It can accurately align data from different physical sensors in a unified three-dimensional coordinate system and generate high-dimensional fusion features that are more sensitive to defects based on their complementarity. This significantly improves the recognition rate and positioning accuracy of hidden defects and mixed defects, and effectively reduces missed detections and false detections caused by incomplete information.

[0015] This invention designs a control and integration platform that automatically plans the scanning path and controls the execution mechanism based on the digital three-dimensional model of the weld to be inspected. This enables the multimodal data acquisition module to autonomously complete full-coverage, high-precision scanning of complex welds on steel box girders, achieving full automation of the inspection process. This frees operators from heavy, repetitive, and experience-dependent labor, significantly improving inspection efficiency and operational safety.

[0016] Through the collaborative work of the above modules, this invention ultimately integrates and outputs the identification results, delivering structured information such as defect type, precise three-dimensional coordinates, and quantified dimensions in various forms such as visual reports and digital interfaces. This provides a directly usable and accurate data foundation for the quality acceptance, health monitoring, and digital operation and maintenance of steel box girders, achieving a deep integration of the inspection process and the digitalization of engineering management. Attached Figure Description

[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the system configuration provided in an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram illustrating the principle of ultrasonic phased array detection of U-rib welds provided in an embodiment of the present invention.

[0020] Figure 3 Comparison of trench imaging using probes of different frequencies provided in this embodiment of the invention;

[0021] Figure 4 This is a schematic diagram of the TOFD detection principle provided in an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of feature fusion provided for an embodiment of the present invention. Detailed Implementation

[0023] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a multimodal steel box girder complex weld defect detection system proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0025] The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] The following description, in conjunction with the accompanying drawings, details the specific scheme of a multimodal steel box girder complex weld defect detection system provided by the present invention.

[0027] Example

[0028] Please see Figure 1 The diagram illustrates a system configuration of a multimodal steel box girder complex weld defect detection system according to an embodiment of the present invention, including:

[0029] The multimodal data acquisition module is used to simultaneously acquire ultrasonic data, radiographic images, and visual images of the U-rib fillet welds or butt welds of the transverse diaphragms of the steel box girder, and simultaneously record the spatial pose information during acquisition.

[0030] The data processing and intelligent recognition module is used to perform spatiotemporal registration and feature-level fusion on the collected data, and to use deep learning models to identify, locate and quantify weld defects.

[0031] The control and integration platform, based on the identification, location and quantification of weld defects, automatically controls the multimodal data acquisition module to complete the automated scanning of the weld and integrates and outputs the identification results.

[0032] It should be noted that ultrasonic data refers to digital data generated by emitting a deflectable and focused ultrasonic beam into the weld using a linear ultrasonic phased array probe, receiving the signals reflected or diffracted within the weld, and converting them into digital data. This can be achieved using a 7.5MHz frequency, 16-element probe configuration, with the beam scanning angle electronically adjusted between 30° and 80°. Its primary purpose is to capture the acoustic response characteristics of defects such as incomplete penetration, porosity, and lack of fusion within the weld, providing data support for internal defect detection.

[0033] X-ray imaging refers to the process of emitting X-rays into the weld area using a low-dose X-ray imaging sensor, capturing the X-ray signal after it penetrates the weld using a flat panel detector, and converting it into a grayscale image. It can be implemented using parameters such as tube voltage from 50 to 300 kV and tube current from 0.5 to 5 mA, with a lead equivalent of ≥2 mm for safety. Its main purpose is to visually present the morphology and distribution of internal weld defects, supplementing the blind spots of ultrasonic testing in identifying some volumetric defects.

[0034] A visual image refers to a surface image captured by an image processor and corrected for distortion after the weld surface is illuminated by an active ring-shaped LED light source from the visual imaging subunit. It can be achieved using an adjustable illumination intensity of 300 to 1000 lux, combined with Zhang's calibration method to correct geometric distortion. Its main purpose is to obtain the visual characteristics of defects such as cracks and undercuts on the weld surface, enabling rapid identification of surface defects.

[0035] Spatial pose information refers to the three-dimensional spatial coordinates and three-dimensional attitude angles of the data acquisition device relative to the weld at the instant of data acquisition. This can be achieved by combining a laser tracker and an inertial measurement unit, synchronously recording with millisecond-level precision timestamps. Its main purpose is to establish a precise correspondence between the acquired data and the physical position of the weld, providing a foundation for subsequent spatiotemporal registration.

[0036] Spatiotemporal registration refers to the process of using the spatial pose information provided by the positioning unit to map the ultrasonic beam path, X-ray image pixels, and visual image pixels to the same global coordinate system by solving the rigid body transformation matrix. This can be achieved by using coordinate transformation algorithms to align data from different sensors to a coordinate system with the feature points of the steel box girder as the origin. Its main purpose is to eliminate spatial misalignment between different modal data and ensure accurate correlation of multi-source data at the same physical location.

[0037] Feature-level fusion refers to the process of dynamically adjusting the weight ratios of features from different modalities based on the physical characteristics of defects using a lightweight attention network and then fusing them in a weighted manner. This can be achieved using an attention network composed of fully connected layers and a softmax function, combined with end-to-end training. Its main purpose is to fully leverage the complementary advantages of different modalities, enhance the discriminative power of defect features, and improve subsequent recognition accuracy.

[0038] Deep learning models refer to hybrid neural networks employing a dual-branch parallel architecture, specifically designed for processing ultrasonic time-series data, X-ray data, and visual image data. This particular model can be trained using a dataset augmented with 462 TOFD defect images. Its primary purpose is to automatically identify weld defect types, locate the defect's 3D coordinates, and quantify its dimensions, reducing reliance on manual intervention.

[0039] Automated scanning refers to the process where a control and integration platform plans the optimal path based on a digital model of the weld, drives a multi-degree-of-freedom motion actuator to move the acquisition module along the path, and completes data acquisition. This can be achieved by importing BIM or point cloud models to automatically identify weld features and generate a comprehensive tracking trajectory. Its main purpose is to replace manual operation, achieving full coverage and high-efficiency inspection of the weld.

[0040] Integrated output refers to the process by which the control and integration platform receives defect identification results, displays them through a visual interface, automatically generates inspection reports, and maps the defect information to the BIM model. This can be achieved through simultaneous overlay display of multimodal data, combined with data push in the IFC standard format. Its main purpose is to provide users with intuitive and structured inspection results, supporting subsequent quality assessments and operational decisions.

[0041] In one specific implementation, the multimodal data acquisition module integrates a lightweight bracket with a linear ultrasonic phased array probe, a low-dose X-ray imaging sensor, a visual imaging subunit, and a positioning unit. The ultrasonic phased array probe is set to a frequency of 7.5MHz, with 16 elements, an acoustic beam scanning angle adjusted to 50°, and a focusing depth set to 1.5 times the weld thickness, suitable for U-rib fillet weld inspection. The X-ray imaging sensor uses a 150kV tube voltage and a 3mA tube current, with a flat panel detector acquiring images at a frame rate of 5 frames / s. The annular LED light source of the visual imaging subunit is installed at a 45° angle to the shooting direction, with the illumination intensity adjusted to 600 lux. After calibration using the Zhang calibration method, the geometric error is controlled within 0.05mm / pixel. The positioning unit uses a combination of a laser tracker and an inertial measurement unit to synchronously record three-dimensional coordinates and attitude angles with millisecond-level timestamps.

[0042] The control and integration platform imports the steel box girder BIM model, automatically identifies the features of the U-rib fillet welds, and generates an optimal scanning path along the weld centerline. The path spacing is set according to the optimal working distance of the ultrasonic phased array probe, which is 8mm. A multi-degree-of-freedom motion actuator drives the acquisition module to move along the path at a speed of 5mm / s, while maintaining the angle between the ultrasonic phased array probe and the weld surface normal within ±5°.

[0043] During the acquisition process, ultrasound data, X-ray images, visual images, and spatial pose information are bound together by the same timestamp and transmitted in real time to the data processing and intelligent recognition module. This module first performs wavelet packet transform noise reduction on the ultrasound A-scan signal, performs median filtering on the X-ray and visual images, and then completes spatiotemporal registration through a rigid body transformation matrix, with the registration error controlled within 0.5 mm.

[0044] Subsequently, the attention network dynamically assigns weights based on defect characteristics. Surface defects are assigned a weight of 0.65 to visual features, while internal defects are assigned a combined weight of 0.75 to ultrasonic and X-ray features. The weighted and fused data is then input into the hybrid neural network model. This model, trained on a TOFD dataset containing five types of defects including cracks and incomplete penetration, achieves an accuracy of 98.05% and outputs the defect type, 3D coordinates, and size parameters.

[0045] The control and integration platform's visual interface synchronously overlays multimodal data and defect annotations, automatically generating standardized reports containing weld information, inspection parameters, and defect maps. Simultaneously, defect information is encapsulated into IFC standard data packets via the MQTT protocol and pushed to the digital twin platform, where 3D visual annotations are performed at the corresponding locations in the BIM model. This enables synchronized updates of inspection results and the digital management platform, with a data transmission latency of <1 second.

[0046] I. Multimodal Data Acquisition Module:

[0047] The multimodal data acquisition module includes a data acquisition device unit and a positioning unit;

[0048] Please see Figure 2 and Figure 3 This is a schematic diagram of the ultrasonic phased array detection principle of U-rib weld seam and a comparison diagram of imaging by probes of different frequencies provided in the embodiments of the present invention.

[0049] The data acquisition unit includes a linear ultrasonic phased array probe, an X-ray imaging sensor consisting of a low-dose X-ray source and a flat panel detector, and a visual imaging subunit. It is used to synchronously acquire multimodal data of the U-rib fillet weld or the butt weld of the diaphragm of the steel box girder with millisecond-level timestamps under the synchronous trigger signal control of the control and integration platform. The multimodal data includes ultrasonic data, X-ray images and visual images.

[0050] The positioning unit is used to synchronously record the three-dimensional spatial coordinates and three-dimensional attitude angles of the data acquisition device unit relative to the weld seam at each data acquisition moment. The positioning data and the acquisition data are bound together by the same millisecond-level precision timestamp under the control of the synchronous trigger signal to achieve spatiotemporal correlation.

[0051] Furthermore, the linear ultrasonic phased array probe is configured to inspect the U-rib fillet weld of the steel box girder. Its electronic deflection range for the acoustic beam scanning angle covers 30° to 80°, and its focusing depth is dynamically set according to the actual plate thickness of the weld. The wedge angle of the probe is adapted to the angle between the U-rib and the top plate to ensure that the incident acoustic beam can cover and focus on the weld root, weld toe, and the set penetration depth requirement area.

[0052] Furthermore, the visual imaging subunit includes:

[0053] An active ring-shaped LED light source is installed at an angle of 30° to 60° to the shooting direction, and the light intensity is adjustable to eliminate reflections on the weld surface;

[0054] An image processor is used to perform geometric distortion correction based on Zhang's calibration method. It acquires multi-view images of a checkerboard calibration plate placed next to the weld, calculates camera intrinsic parameters and distortion coefficients, and uses these coefficients to correct the acquired weld images in real time, so that the geometric error of the corrected image is controlled within a set threshold.

[0055] It should be noted that a linear ultrasonic phased array probe refers to an ultrasonic testing probe that uses a multi-element array design and can achieve sound beam deflection and focusing through electronic control. Its center frequency can be selected within the range of 2 to 10 MHz; for example, 8 to 10 MHz can be selected for thin plate welds, and 2 to 5 MHz for thick plate welds. The number of array elements can be selected between 16 and 128, with an element spacing between 0.3 and 1.0 mm. The probe wedge is made of plexiglass, with a longitudinal wave velocity of approximately 2700 to 2800 m / s. The wedge angle can be finely adjusted according to the weld inclination angle. To adapt to U-rib fillet welds with inclination angles between 60° and 70°, the wedge angle is selected between 15° and 30° to ensure that the sound beam effectively covers critical areas such as the weld root and weld toe.

[0056] A low-dose X-ray source refers to a radiation source whose output dose meets ionizing radiation protection standards and can emit an X-ray beam that can penetrate welds. It can be a compact X-ray source with a tube voltage of 50 to 300 kV and a tube current of 0.5 to 5 mA, paired with an industrial-grade amorphous silicon flat panel detector with an effective detection area of ​​not less than 200 mm × 300 mm and a pixel size ≤ 100 μm. Its main purpose is to obtain direct images of internal weld defects, supplementing the blind spots of ultrasonic testing in identifying volumetric defects.

[0057] The visual imaging subunit, consisting of an illumination source and an image processor, is a device used to acquire images of the weld surface and perform distortion correction. It can utilize a combination of a ring-shaped LED light source and an embedded image processor. The light source is adjustable in intensity from 300 to 1000 lux, and the image processor is an industrial-grade processor supporting real-time processing of 8 megapixels. Its primary purpose is to clearly capture the visual characteristics of defects such as cracks and undercuts on the weld surface, eliminate surface reflection interference, and improve the accuracy of surface defect measurement.

[0058] A positioning unit is a measurement unit capable of acquiring the spatial position and attitude of a data acquisition device in real time. It can be a combination of a laser tracker and an inertial measurement unit (IMU). The laser tracker is an industrial-grade device with a positioning accuracy of ±0.1mm, and the IMU is a compact unit supporting millisecond-level data output. Its main purpose is to establish a precise correlation between the acquired data and the physical location of the weld, providing fundamental data for subsequent spatiotemporal registration.

[0059] Synchronization trigger signal refers to the synchronization command issued by the control and integration platform to control the coordinated operation of the data acquisition device unit and the positioning unit. It can be transmitted via industrial Ethernet and triggered using a pulse signal. Its main purpose is to ensure that ultrasonic data, X-ray images, visual images, and positioning data are acquired at the same time, guaranteeing the spatiotemporal consistency of the data.

[0060] Millisecond-level timestamps refer to time stamps with millisecond-level precision added to each set of collected data. This can be achieved through the clock synchronization function of the positioning unit, ensuring a unified time base for all data. Its main purpose is to achieve precise binding between collected data and positioning data, avoiding misalignment of position and data during subsequent data processing.

[0061] Electronic beam scanning angle deflection refers to the electronically controlled deflection of the sound beam propagation direction by controlling the excitation timing of each element of the ultrasonic phased array probe. This can be achieved through dedicated control software in the phased array equipment, with a deflection range covering 30° to 80°. Its primary purpose is to adapt to the inclined structure of U-rib fillet welds, enabling the sound beam to cover critical areas such as the weld root and weld toe, eliminating blind spots in detection.

[0062] Dynamic focusing depth setting refers to adjusting the focusing position of the ultrasonic phased array probe's sound beam according to the actual weld thickness. This can be achieved through the parameter configuration interface of the phased array equipment, with the focusing depth set to 1.0 to 2.0 times the weld thickness. Its main purpose is to ensure the sound beam is focused at the depth of the defect, thereby increasing the amplitude and resolution of the defect signal and improving defect detection accuracy.

[0063] Wedge angle adaptation refers to selecting the corresponding ultrasonic phased array probe wedge angle based on the angle between the U-rib and the top plate. Wedges ranging from 15° to 30° can be used to adapt to U-rib to 70° angles between the probe and the top plate. This is primarily to ensure that the incident sound beam enters the weld area at a suitable angle, effectively covering the weld root, weld toe, and areas requiring sufficient penetration.

[0064] Zhang's calibration method refers to a calibration method that uses multi-view images of a checkerboard calibration board to solve for camera intrinsic parameters and distortion coefficients. It can be implemented using a dedicated algorithm in an image processor to correct distortion in the visual imaging sub-units. Its main purpose is to eliminate image distortion caused by shooting angle and lens distortion, ensuring the geometric accuracy of surface defect measurements.

[0065] Geometric distortion correction refers to the process of correcting acquired weld images using camera intrinsic parameters and distortion coefficients obtained through Zhang's calibration method. This can be achieved in real-time using an image processor, with the geometric error of the corrected image controlled within 0.05 mm / pixel. Its main purpose is to ensure that the visual image accurately reflects the geometric morphology of the weld surface, providing a guarantee for the precise measurement of surface defects.

[0066] In one specific implementation, the data acquisition unit uses the following combination of equipment: the linear ultrasonic phased array probe is a GE Phasor XS model with 16 elements and a 7.5MHz linear array; the low-dose X-ray source is an industrial-grade X-ray source with a tube voltage of 150kV and a tube current of 3mA; the matching flat panel detector is an amorphous silicon flat panel detector with an effective detection area of ​​200mm×300mm and a pixel size of 75μm; the visual imaging subunit uses a ring LED light source with an adjustable illumination intensity of 300 to 1000 lux and an embedded image processor that supports real-time processing of 8 million pixels; and the positioning unit uses a laser tracker with a positioning accuracy of ±0.1mm and an inertial measurement unit with millisecond-level output.

[0067] The control and integration platform sends a synchronous trigger signal via industrial Ethernet, with the trigger frequency set to 10Hz, ensuring that the data acquisition unit and the positioning unit synchronously acquire data once every 100 milliseconds. The electronic deflection range of the beam scanning angle of the linear ultrasonic phased array probe is set to 30° to 80°. For an 8mm thick U-rib plate, the dynamic focusing depth is set to 12mm. A 25° wedge is selected to match the 65° angle between the U-rib and the top plate, ensuring that the sound beam covers the weld root, weld toe, and areas requiring penetration depth.

[0068] The ring LED light source of the visual imaging subunit is installed at a 45° angle to the shooting direction, and the light intensity is adjusted to 600 lux to eliminate the reflection on the weld surface. The image processor executes the Zhang calibration method to acquire 10 images from different perspectives of a 10mm×10mm checkerboard calibration plate placed next to the weld. The camera intrinsic parameters and distortion coefficients are calculated, and the acquired weld images are corrected in real time. The geometric error of the corrected image is controlled within 0.05mm / pixel.

[0069] The positioning unit outputs the three-dimensional spatial coordinates and three-dimensional attitude angles of the data acquisition device in real time. Each set of acquired data is timestamped at the millisecond level, and is precisely bound to the ultrasound data, X-ray images, and visual images through timestamps to ensure the spatiotemporal consistency of all data. During the acquisition process, all data is transmitted in real time to the cache module of the control and integration platform via high-speed Ethernet, providing complete and accurate raw data for subsequent data processing and intelligent recognition.

[0070] II. Data Processing and Intelligent Recognition Module:

[0071] Please see Figure 4 and Figure 5 This is a schematic diagram of the TOFD detection principle and a schematic diagram of feature fusion provided in the embodiments of the present invention.

[0072] The data processing and intelligent recognition module includes a data preprocessing unit and a feature fusion and recognition unit;

[0073] The data preprocessing unit is used to perform spatiotemporal registration on multimodal data. The spatiotemporal registration uses positioning data and solves the rigid body transformation matrix to uniformly map the ultrasonic beam path, X-ray image pixels and visual image pixels to a global coordinate system with the feature points of the steel box girder as the origin.

[0074] The feature fusion and recognition unit adopts an adaptive feature-level fusion strategy based on an attention mechanism. The registered multimodal data is input into a hybrid neural network model. The adaptive strategy uses a lightweight attention network trained end-to-end with the recognition network to dynamically calculate the weight coefficients of ultrasonic, X-ray and visual features based on the physical characteristics of the defect. The hybrid neural network model outputs the type, three-dimensional spatial coordinates and quantized size of the weld defect.

[0075] Furthermore, in the adaptive feature-level fusion strategy based on the attention mechanism, the lightweight attention network is a network composed of a fully connected layer and a Softmax function; the physical characteristics of the defect are characterized by the defect depth, orientation, and acoustic / optical imaging features; the process of dynamically allocating weights is to receive the feature vectors of each modality through the attention network, output the corresponding weight coefficients, and obtain the fused feature vector through weighted summation.

[0076] It should be noted that spatiotemporal registration refers to the process of mapping different modal data to the same global coordinate system by solving the rigid body transformation matrix using spatial pose data acquired by the positioning unit. This can be achieved using coordinate transformation algorithms and relying on an industrial-grade data processing server, specifically an industrial control computer equipped with an Intel Core i9 processor and 64GB of memory. Its main purpose is to eliminate spatial misalignment of ultrasonic, X-ray, and visual data, ensuring accurate correlation of multi-source data at the same physical location.

[0077] A rigid body transformation matrix is ​​a mathematical matrix that describes the spatial position and attitude transformations of a data acquisition device, used to achieve coordinate transformations from data from different sensors. It can be obtained by solving for the 3D coordinates and attitude angles output by the positioning unit, with the assistance of the matrix operation toolkit in MATLAB software. Its main purpose is to provide the mathematical foundation for spatiotemporal registration and ensure the accuracy of coordinate transformations.

[0078] Attention mechanisms refer to algorithms that dynamically allocate weights to different modal features based on the physical characteristics of defects. They can be implemented using a lightweight network consisting of fully connected layers and a Softmax function, deployed within the PyTorch deep learning framework. Their primary purpose is to highlight modal features that are more effective for defect identification and improve the discriminative power of fused features.

[0079] Lightweight attention networks are neural networks with simple structures and low computational cost, used for dynamically assigning feature weights. They can employ a network structure containing two fully connected layers and one Softmax function, and are trained and inferred using an NVIDIA Tesla V100 graphics card. Their main purpose is to improve data processing efficiency while maintaining the accuracy of weight assignment, thus meeting real-time detection requirements.

[0080] Defect physical characteristic characterization refers to the key features that can distinguish different defect types, including defect depth, orientation, and acoustic / optical imaging features. It can be obtained by extracting information such as amplitude, texture, and edges from multimodal data, relying on feature extraction algorithms. Its main purpose is to provide a basis for weight allocation in attention networks, enabling targeted feature enhancement for different types of defects.

[0081] Hybrid neural network models are deep learning models that employ a dual-branch parallel architecture, capable of processing both time-series and image data simultaneously. They can use YOLOv5s as the base network architecture, combined with LSTM units to construct a dual-branch structure, and trained using the PyTorch framework. Their primary purpose is to adapt to the heterogeneous characteristics of ultrasound time-series data and X-ray and visual image data, enabling efficient defect identification, localization, and quantification.

[0082] Dynamic weight allocation refers to the process by which the attention network adjusts the weight ratios of ultrasound, X-ray, and visual features in real time based on the physical characteristics of the defect. This is achieved by inputting feature vectors from each modality into the attention network, outputting weight coefficients, and then performing a weighted sum. Its main purpose is to fully leverage the complementary advantages of data from different modalities, optimize feature combinations for different defect types, and improve recognition accuracy.

[0083] In one specific implementation, the data preprocessing unit uses an industrial control computer equipped with an Intel Core i9 processor and 64GB of memory, and solves the rigid body transformation matrix using the matrix operation toolkit of MATLAB software. First, it receives multimodal data and the three-dimensional coordinates and attitude angles output by the positioning unit. Wavelet packet transform is used to reduce noise in the ultrasonic A-scan signal, and median filtering is applied to the X-ray image and visual image. Then, the rigid body transformation matrix is ​​used to uniformly map the ultrasonic beam path, X-ray image pixels, and visual image pixels to a global coordinate system with the corner point of the steel box girder top plate as the origin, controlling the registration error to within 0.5mm.

[0084] The lightweight attention network for the feature fusion and recognition unit adopts a structure containing two fully connected layers and one Softmax function, deployed on the PyTorch framework, and runs on an NVIDIA Tesla V100 graphics card. The hybrid neural network model is based on the YOLOv5s architecture, with an added LSTM temporal processing branch, forming a dual-branch parallel structure. The training dataset uses a dataset containing 462 TOFD defect images, which, after data augmentation processing such as Mosaic stitching and rotation, is divided into training and validation sets in a 7:3 ratio.

[0085] During inspection, the registered ultrasonic A-scan signal is input into the LSTM branch to extract temporal features; the X-ray and visual images are input into the CNN branch to extract image features; the two types of features are input into a lightweight attention network, which dynamically assigns weights based on the physical characteristics of the defect: surface defects such as surface cracks are assigned a weight of 0.65 to visual features, while internal defects such as incomplete penetration are assigned a combined weight of 0.75 to ultrasonic and X-ray features. The weighted and fused features are then input into a classification and regression network, which outputs the defect type (crack, incomplete penetration, porosity, inclusion, lack of fusion), three-dimensional spatial coordinates, and quantized dimensions.

[0086] Tests showed that the module achieved a defect identification accuracy of 98.05%, with classification confidence levels all above 60%, reaching a maximum of 97%. It can accurately process multimodal detection data of U-rib fillet welds and butt welds of transverse diaphragms in steel box girders, meeting the accuracy and efficiency requirements of engineering inspection.

[0087] III. Control and Integration Platform:

[0088] The control and integration platform includes a path planning and motion control unit and a result integration and output unit;

[0089] The path planning and motion control unit is used to import the building information model or 3D laser scanning point cloud model of the steel box girder, automatically identify the weld features, and automatically generate the optimal scanning path without omissions or repetitions based on the working distance and viewing angle of each sensor in the multimodal data acquisition module. It also drives the multi-degree-of-freedom motion actuator to move the multimodal data acquisition module along the path.

[0090] The results integration and output unit receives the type, three-dimensional spatial coordinates, and quantified dimensions of weld defects, provides visualization and automated report generation functions, and maps defect information to the building information model in real time through a digital twin data interface.

[0091] Furthermore, in the path planning and motion control unit:

[0092] Automatically generating the optimal scanning path includes: calculating the weld centerline based on the building information model or point cloud model, and generating a probe tracking trajectory along the centerline and according to the optimal working distance of the ultrasonic phased array probe.

[0093] The multi-degree-of-freedom motion actuator includes calculating the tracking trajectory into motion commands for each joint and controlling the actuator to drive the data acquisition device unit to move along the trajectory, while keeping the angle between the ultrasonic phased array probe and the normal of the weld surface within a preset range.

[0094] Furthermore, the real-time mapping function of the digital twin data interface is specifically implemented as follows:

[0095] The system receives structured defect information output by the feature fusion and recognition unit. The structured defect information includes defect type encoding, three-dimensional spatial coordinates and size, and confidence level.

[0096] Encapsulate structured defect information into data packets that conform to the Industry Foundation Classes standard;

[0097] Data packets are pushed to the target digital twin platform via HTTP RESTful API or MQTT protocol, triggering the platform to perform 3D visualization annotation and information binding at the corresponding spatial location in the building information model.

[0098] Furthermore, in the results integration and output unit:

[0099] The visualization display function provides an integrated interface that can simultaneously overlay, compare and analyze ultrasound B / C scan images, X-ray images, visual optical images and three-dimensional spatial positioning information in the same time axis and spatial coordinate system.

[0100] The automated report generation function can automatically fill in weld information, inspection parameters, defect maps, data lists and conclusions based on preset templates to generate inspection report documents that conform to industry standards.

[0101] It should be noted that a Building Information Model (BIM) is a 3D digital model containing information such as the geometric dimensions, weld locations, and material parameters of the steel box girder, used to provide the basic data for path planning. It can be built using Revit 2023 or Bentley OpenBridge Modeler software, and the model format supports the IFC 4.0 standard. Its main purpose is to provide accurate spatial location information of welds for path planning, ensuring that the scanning path covers the entire inspection area.

[0102] A 3D laser scanning point cloud model refers to the point cloud data of a steel box girder surface acquired through a 3D laser scanner, which can be used to reconstruct the actual shape of the weld. It can be acquired using a FARO Focus S70 3D laser scanner, with a point cloud density ≥100 points / square millimeter. Its main purpose is to adapt to the deviation between the actual weld location on site and the design model, thereby improving the accuracy of path planning.

[0103] The optimal scanning path refers to the motion trajectory that covers the entire weld area, is non-repetitive, and meets the sensor's working distance requirements. It can be generated using a path planning algorithm and runs on an industrial-grade control computer equipped with an Intel Core i7 processor and 32GB of RAM. Its main purpose is to ensure that the multimodal data acquisition module can efficiently and comprehensively acquire weld data, avoiding blind spots in detection.

[0104] A multi-degree-of-freedom motion actuator is a mechanical device capable of moving a multi-modal data acquisition module along a planned path, enabling multi-directional attitude adjustment. It can be a six-axis industrial robot, such as the ABB IRB 1200 or KUKAKR C4, with a repeatability of ±0.05mm. Its primary function is to accurately execute the scanning path and maintain the optimal relative position and angle between the sensor and the weld seam.

[0105] Structured defect information refers to standardized defect data that has been classified and quantified, including key parameters such as type, coordinates, size, and confidence level. It can be output by the data processing and intelligent recognition module and stored in JSON format. This is primarily for facilitating subsequent visualization, report generation, and data reception and processing by the digital twin platform.

[0106] A digital twin data interface refers to the communication interface that enables data interaction between the inspection system and the digital twin platform, supporting standardized data transmission. It can use an industrial Ethernet interface and supports HTTP RESTful API or MQTT protocol. Its main purpose is to push defect information to the digital twin platform in real time, enabling synchronized updates between inspection results and the 3D model.

[0107] Visualization refers to the function of synchronously displaying multimodal inspection data and defect annotation information on the same interface, supporting multi-dimensional analysis. It can be developed using the Qt framework and runs on industrial-grade touchscreen displays. Its main purpose is to provide operators with intuitive inspection results, facilitating comprehensive interpretation and verification of defects.

[0108] Automated report generation refers to the function of automatically filling in test data and generating test reports that conform to industry standards based on pre-set templates. It can be implemented using Python scripts combined with the Jinja2 template engine, and supports PDF and Word format output. Its main purpose is to reduce the workload of manual report compilation and ensure the standardization and accuracy of reports.

[0109] In one specific implementation, the control and integration platform uses an industrial control computer equipped with an Intel Core i7 processor, 32GB of memory, and a 1TB solid-state drive, running the Windows 10 IoT operating system. The path planning algorithm and visualization software are both developed based on Python 3.8. The path planning and motion control unit imports a steel box girder building information model (IFC 4.0 format) built using Revit 2023, automatically identifies the characteristics of the U-rib fillet welds and the butt welds of the diaphragms, calculates the weld centerline, and generates a spiral probe tracking trajectory with a spacing of 5mm based on the optimal working distance of 8mm for the ultrasonic phased array probe. The multi-degree-of-freedom motion actuator uses an ABB IRB 1200 six-axis industrial robot, which calculates the tracking trajectory into joint motion commands, controlling the robot to move the multimodal data acquisition module along the trajectory at a speed of 5mm / s. Simultaneously, the robot's attitude adjustment function maintains the angle between the ultrasonic phased array probe and the normal to the weld surface within ±5°. The results integration and output unit receives structured defect information from the data processing and intelligent recognition module, including defect type codes (0-crack, 1-lack fusion, 2-incomplete penetration, 3-porosity, 4-slag inclusion), three-dimensional spatial coordinates, dimensions, and confidence levels. The visualization software, developed using the Qt framework, provides an integrated interface that synchronously overlays ultrasonic B / C scan images, X-ray images, visual optical images, and three-dimensional spatial positioning information on the same time axis and spatial coordinate system, supporting defect location backtracking and multi-modal data linkage comparison. The automated report generation function, based on the Jinja2 template engine, pre-sets report templates conforming to the "Technical Specifications for Highway Bridge and Culvert Construction," automatically fills in weld information (location, plate thickness, penetration requirements), detection parameters (sensor model, scanning speed, probe parameters), defect atlas (defect images under each modality), data list (defect type, coordinates, dimensions, confidence level), and detection conclusions, generating a PDF format inspection report with automatic timestamps and project numbers. The digital twin data interface uses the MQTT protocol to encapsulate structured defect information into data packets conforming to the IFC 4.0 standard. These packets are then pushed to the digital twin platform equipped with the Bentley iTwin Platform via industrial Ethernet. This triggers the platform to perform 3D visualization annotations on the corresponding spatial locations of the steel box girder building information model. Different colored icons are used to distinguish defect types (red - crack, yellow - lack of fusion, blue - incomplete penetration, green - porosity, orange - slag inclusion). A pop-up window displays detailed defect parameters and confidence levels, enabling synchronous updates of detection results and the digital twin platform with a data transmission latency of <1 second.

[0110] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A multimodal defect detection system for complex welds in steel box girders, characterized in that, include: The multimodal data acquisition module is used to simultaneously acquire ultrasonic data, radiographic images, and visual images of the U-rib fillet welds or butt welds of the transverse diaphragms of the steel box girder, and simultaneously record the spatial pose information during acquisition. The data processing and intelligent recognition module is used to perform spatiotemporal registration and feature-level fusion on the collected data, and to use deep learning models to identify, locate and quantify weld defects. The control and integration platform, based on the identification, location and quantification of weld defects, automatically controls the multimodal data acquisition module to complete the automated scanning of the weld and integrates and outputs the identification results.

2. The multimodal steel box girder complex weld defect detection system according to claim 1, characterized in that, The multimodal data acquisition module includes: The data acquisition unit includes a linear ultrasonic phased array probe, an X-ray imaging sensor consisting of a low-dose X-ray source and a flat panel detector, and a visual imaging subunit. It is used to synchronously acquire multimodal data of the U-rib fillet weld or the butt weld of the transverse diaphragm of the steel box girder with millisecond-level timestamps under the synchronous trigger signal control of the control and integration platform. The multimodal data includes ultrasonic data, X-ray images and visual images. The positioning unit is used to synchronously record the three-dimensional spatial coordinates and three-dimensional attitude angle of the data acquisition device unit relative to the weld at each data acquisition moment. The positioning data and the acquisition data are bound together by the same millisecond-level precision timestamp under the control of the synchronization trigger signal to achieve spatiotemporal correlation.

3. The multimodal steel box girder complex weld defect detection system according to claim 2, characterized in that, The linear ultrasonic phased array probe is configured to detect the U-rib fillet weld of the steel box girder. Its electronic deflection range for the acoustic beam scanning angle covers 30° to 80°, and its focusing depth is dynamically set according to the actual plate thickness of the weld. The wedge angle of the probe is adapted to the angle between the U-rib and the top plate to ensure that the incident acoustic beam can cover and focus on the weld root, weld toe, and the set penetration depth requirement area.

4. The multimodal steel box girder complex weld defect detection system according to claim 2, characterized in that, The visual imaging subunit includes: An active ring-shaped LED light source is installed at an angle of 30° to 60° to the shooting direction, and the light intensity is adjustable to eliminate reflections on the weld surface; An image processor is used to perform geometric distortion correction based on Zhang's calibration method. It acquires multi-view images of a checkerboard calibration plate placed next to the weld, calculates camera intrinsic parameters and distortion coefficients, and uses these coefficients to correct the acquired weld images in real time, so that the geometric error of the corrected image is controlled within a set threshold.

5. The multimodal steel box girder complex weld defect detection system according to claim 1, characterized in that, The data processing and intelligent recognition module includes: The data preprocessing unit is used to perform spatiotemporal registration on multimodal data. The spatiotemporal registration uses positioning data and solves the rigid body transformation matrix to uniformly map the ultrasonic beam path, X-ray image pixels and visual image pixels to a global coordinate system with the feature points of the steel box girder as the origin. The feature fusion and recognition unit adopts an adaptive feature-level fusion strategy based on an attention mechanism. The registered multimodal data is input into a hybrid neural network model. The adaptive strategy uses a lightweight attention network trained end-to-end with the recognition network to dynamically calculate the weight coefficients of ultrasonic, X-ray and visual features based on the physical characteristics of the defect. The hybrid neural network model outputs the type, three-dimensional spatial coordinates and quantized size of the weld defect.

6. The multimodal steel box girder complex weld defect detection system according to claim 5, characterized in that, In the adaptive feature-level fusion strategy based on the attention mechanism, the lightweight attention network is a network composed of a fully connected layer and a Softmax function; The physical characteristics of the defect include defect depth, orientation, and acoustic / optical imaging features; the dynamic weight allocation process involves receiving feature vectors of each modality through an attention network, outputting corresponding weight coefficients, and obtaining a fused feature vector through weighted summation.

7. The multimodal steel box girder complex weld defect detection system according to claim 1, characterized in that, The control and integration platform includes: The path planning and motion control unit is used to import the building information model or 3D laser scanning point cloud model of the steel box girder, automatically identify the weld features, and automatically generate the optimal scanning path without omissions or repetitions based on the working distance and viewing angle of each sensor in the multimodal data acquisition module. It also drives the multi-degree-of-freedom motion actuator to move the multimodal data acquisition module along the path. The results integration and output unit receives the type, three-dimensional spatial coordinates, and quantified dimensions of weld defects, provides visualization and automated report generation functions, and maps defect information to the building information model in real time through a digital twin data interface.

8. The multimodal steel box girder complex weld defect detection system according to claim 7, characterized in that, In the path planning and motion control unit: The automatic generation of the optimal scanning path includes: calculating the weld centerline based on the building information model or point cloud model, and generating a probe tracking trajectory along the centerline and according to the optimal working distance of the ultrasonic phased array probe. The multi-degree-of-freedom motion actuator includes calculating the tracking trajectory into motion commands for each joint, controlling the actuator to drive the data acquisition device unit to move along the trajectory, while keeping the angle between the ultrasonic phased array probe and the normal of the weld surface within a preset range.

9. The multimodal steel box girder complex weld defect detection system according to claim 7, characterized in that, The real-time mapping function of the digital twin data interface is specifically implemented as follows: The system receives structured defect information output by the feature fusion and recognition unit, the structured defect information including defect type encoding, three-dimensional spatial coordinates and size, and confidence level; Encapsulate structured defect information into data packets that conform to the Industry Foundation Classes standard; Data packets are pushed to the target digital twin platform via HTTP RESTful API or MQTT protocol, triggering the platform to perform 3D visualization annotation and information binding at the corresponding spatial location in the building information model.

10. The multimodal steel box girder complex weld defect detection system according to claim 7, characterized in that, In the result integration and output unit: The visualization display function specifically provides an integrated interface that can synchronously overlay, compare and analyze ultrasound B / C scan images, X-ray images, visual optical images and three-dimensional spatial positioning information under the same time axis and spatial coordinate system. The automated report generation function can automatically fill in weld information, inspection parameters, defect maps, data lists and conclusions based on preset templates to generate inspection report documents that conform to industry standards.

Citation Information

Patent Citations

  • Steel beam box construction site defect detection device and method

    CN118961875A

  • Ultrasonic detection fast adjustment and calibration method through diffraction time difference process

    CN102253126A

  • Ultrasonic detecting device and detecting method for interface corrugation of explosive welding composite material

    CN102608209A

  • Welding seam defect detection and quick extraction system based on ultrasonic phased array and method

    CN107576729A

  • Real-time object three-dimensional reconstruction method based on depth camera

    CN107833270A

Cited By

  • Weld joint center line real-time tracking method and device, computer equipment and storage medium

    CN121829324A

  • Chip defect diagnosis method based on computer vision and detection system thereof

    CN121978138A