Intelligent visual defect detection method and system for steel structure

CN122651733APending Publication Date: 2026-08-28SICHUAN YIHUA ZHIYUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610787833.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

本发明的目的在于提供一种钢结构智能视觉缺陷检测方法及系统,能够有效解决现有技术中单纯依赖视觉检测手段难以发现闭合微裂纹等隐性缺陷的技术问题

Benefits of technology

通过声振动激励与视觉捕捉相结合的方式,能够有效激活闭合微裂纹等隐性缺陷,使其产生可识别的微幅振动,解决了单纯视觉检测无法发现隐蔽缺陷的技术难题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122651733A_ABST
    Figure CN122651733A_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field and discloses a steel structure intelligent visual defect detection method and system, aiming to solve the technical problem that it is difficult to find hidden defects such as closed micro-cracks by simply relying on visual detection means in the prior art. The method comprises the following steps: performing sweep excitation of a specific frequency range on a steel structure to make the defect area produce a micro-vibration response; using a high-frame-rate industrial camera combined with Lagrange video amplification technology to capture the sub-pixel level vibration amplitude; and outputting a defect detection report based on a support vector machine classifier. The system comprises an acoustic vibration excitation module, a high-frame-rate visual capture module, a cross-modal fusion analysis module and a defect detection report output module. By combining acoustic vibration excitation and visual capture, the application can effectively activate and detect hidden defects, realize non-contact, rapid, high-sensitivity intelligent defect identification, and significantly improve the detection accuracy and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer technology, specifically relating to a method and system for intelligent visual defect detection of steel structures. Background Technology

[0002] As the core load-bearing components of modern industrial and civil buildings, the safety and reliability of steel structures directly affect the overall stability of major projects. With the large-scale application of steel structures in bridges, high-rise buildings, offshore platforms, and special equipment, various defects inevitably arise during manufacturing, installation, and long-term service. Among these, latent defects such as closed microcracks, hidden welding defects, and stress concentration areas, which are difficult to detect using conventional testing methods, have become key hidden dangers threatening the safe operation of steel structures. Traditional steel structure inspection methods mainly include manual visual inspection, penetrant testing, magnetic particle testing, ultrasonic testing, and radiographic testing. However, these methods are no longer sufficient to meet the quality control requirements of modern large-scale steel structure projects in terms of testing efficiency, coverage, and level of automation.

[0003] With the rapid development of machine vision and artificial intelligence technologies, vision-based steel structure defect detection technology has gradually become a research hotspot in the industry. Among them, visual recognition methods based on convolutional neural networks can achieve automatic classification and localization of surface defects to a certain extent. However, this technical approach has inherent limitations: for microcracks in a closed state, embedded defects, or latent defects covered by anti-corrosion coatings, traditional optical imaging is difficult to effectively capture because they do not form significant visual features on the surface, often leading to missed or false detections. In addition, pure visual inspection is highly dependent on lighting conditions, camera resolution, and imaging angle, and the stability of detection results is insufficient in complex industrial environments.

[0004] To improve the ability of visual inspection to capture latent defects, researchers have explored various auxiliary enhancement techniques. A common strategy involves using phased array ultrasound or laser ultrasound to excite defects, generating a recognizable ultrasonic response in the defect area, followed by signal acquisition via visual or ultrasonic probes. While this method can activate latent defects to some extent, it suffers from practical problems such as coupling agent dependence, bulky equipment, and limited detection efficiency. Another approach utilizes infrared thermal imaging technology, visualizing defects through the difference in thermal conductivity between defective and intact areas. However, this method is highly selective for defect depth and type, and lacks sufficient sensitivity for detecting shallow closed cracks. Some researchers have proposed a detection approach based on acoustic vibration, using acoustic waves to excite micro-vibrations in the defect area, followed by measurement using contact sensors or laser Doppler vibrometers. However, this method typically requires point-by-point scanning, resulting in low detection efficiency, and its adaptability to complex structural systems needs improvement. Overall, existing steel structure defect detection technologies are still unable to effectively capture various latent defects in a non-contact, rapid, and highly sensitive manner, failing to meet the application requirements of intelligent and automated inspection. Summary of the Invention

[0005] Latent defects make it difficult to meet the application requirements of intelligent and automated detection. (Invention Content) The purpose of this invention is to provide an intelligent visual defect detection method and system for steel structures, which can effectively solve the technical problem that existing technologies relying solely on visual inspection methods are unable to detect latent defects such as closed microcracks. By deeply integrating acoustic vibration excitation technology with machine vision technology, this invention achieves effective activation and high-precision detection of latent defects inside steel structures, significantly improving the accuracy and reliability of defect identification.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: a method and system for intelligent visual defect detection of steel structures, characterized in that it includes an acoustic vibration excitation module, a high frame rate visual capture module, a cross-modal fusion analysis module, and a defect detection report output module. The modules work together through a preset synchronization mechanism and communication protocol to jointly complete the all-round detection of potential defects on the surface and inside of the steel structure.

[0007] The specific testing method includes the following steps: S1: The steel structure is subjected to frequency sweep excitation within a specific frequency range using an acoustic vibration excitation device, causing micro-amplitude vibration responses in the steel structure surface and potential defect areas. The acoustic vibration excitation device includes a power amplifier, a signal generator, and an ultrasonic transducer. The power amplifier has a preset maximum output power range, and the signal generator supports both linear and logarithmic frequency sweep modes. The frequency sweep excitation adopts a sinusoidal frequency sweep method, and the sweep period is continuously adjustable within a preset time period. Under the action of acoustic vibration excitation, the defect areas and healthy areas on the steel structure surface exhibit different vibration response characteristics. Due to the structural discontinuity, the defect areas will generate unique vibration modes, which carry information such as the location, type, and severity of the defects.

[0008] S2: A high frame rate industrial camera combined with a video magnification algorithm is used to perform real-time imaging of the steel structure surface, capturing the minute vibration amplitudes generated by acoustic excitation in defect areas. The high frame rate industrial camera has preset frame rate and pixel resolution requirements. The camera is equipped with an infrared filter to eliminate interference from ambient light sources on the detection of minute vibrations. Simultaneously, the camera is equipped with a high-precision external trigger synchronization interface to achieve precise synchronization with the acoustic vibration excitation signal, ensuring the time consistency between visual capture and acoustic excitation. The video magnification algorithm uses Lagrange video magnification technology, performing spatial pyramid decomposition on the video sequence and then magnifying it within a selected frequency band. The magnification factor is continuously adjustable within a preset range. The magnified vibration amplitude is accurately measured using a sub-pixel registration algorithm, achieving effective extraction of sub-pixel level vibration information. The video magnification algorithm performs temporal domain decomposition and frequency domain filtering on the original video to extract vibration information within a preset amplitude range. This vibration information reflects the degree of minute deformation on the steel structure surface.

[0009] S3: Cross-modal fusion analysis is performed between the visually captured vibration spectrum and the feedback signal collected by the acoustic sensor. Based on the correlation calculation of vibration characteristics, defect location and type identification are achieved. The cross-modal fusion uses a time-series alignment algorithm to ensure the time synchronization of the acoustic vibration signal and the visual vibration, and the synchronization error is controlled within a preset time accuracy range. The cross-modal fusion analysis specifically includes the following steps: First, the visual vibration spectrum is segmented into regions to extract the vibration time-series signal of the suspected defect region; second, the acoustic feedback signal is subjected to spectrum analysis to extract the main vibration frequency and energy distribution characteristics; then, the correlation coefficient between the visual vibration signal and the acoustic feedback signal is calculated. When the correlation coefficient is greater than a preset correlation coefficient threshold, the region is determined to have a defect. Through the cross-modal fusion analysis, environmental noise interference can be effectively eliminated, improving the accuracy and reliability of defect detection.

[0010] S4: Output a defect detection report based on the fusion analysis results, including defect location coordinates, defect type classification, and defect severity rating; the defect type classification uses a support vector machine classifier, with input features being a fusion feature vector of vibration amplitude, vibration frequency, vibration phase, and acoustic energy density, and classification categories including closed cracks, welding defects, corrosion areas, and stress concentration areas; the kernel function of the support vector machine classifier uses a radial basis function kernel function, the penalty parameter is set to a preset value, the feature dimension is a multi-dimensional fusion feature, and the classification accuracy remains at a high level.

[0011] Furthermore, the present invention also includes a defect severity rating module, which calculates the signal-to-noise ratio (SNR) based on the ratio of vibration amplitude to background noise, and rates the defect severity based on the preset range in which the SNR falls; when the SNR is greater than a first preset threshold, it is rated as a severe defect; when the SNR is between the first and second preset thresholds, it is rated as a moderate defect; and when the SNR is between the second and third preset thresholds, it is rated as a minor defect.

[0012] The technical principle of this invention is as follows: A sweep excitation signal within a specific frequency range generated by the acoustic vibration excitation device is applied to the surface of a steel structure through an ultrasonic transducer. The excitation signal propagates inside the steel structure and interacts with the defect area. Due to the structural discontinuity of the defect area, the excitation signal will produce phenomena such as reflection, refraction, and mode conversion at the defect, resulting in the defect area exhibiting different vibration response characteristics than the healthy area. A high frame rate industrial camera amplifies these minute vibration responses to the visible range through a video magnification algorithm, thereby effectively capturing latent defects. An acoustic sensor synchronously collects the response signal of the steel structure to the excitation signal, which contains information such as the type, location, and severity of the defect. A cross-modal fusion analysis module deeply fuses the visual vibration spectrum with the acoustic feedback signal, and determines the specific location and type of the defect through correlation analysis.

[0013] The interaction logic between the components is as follows: the sweep frequency excitation signal generated by the signal generator is amplified by the power amplifier and drives the ultrasonic transducer to emit acoustic vibration excitation; at the same time, the signal generator sends a synchronization trigger signal to the high frame rate industrial camera to ensure that the camera starts image acquisition at the same time as the excitation signal is applied; the acoustic sensor collects the vibration response signal of the steel structure in real time and transmits it together with the vibration spectrum captured by vision to the cross-modal fusion analysis module; the fusion analysis module performs time alignment, feature extraction and correlation calculation on the signals of the two modes, and finally outputs the defect detection results to the report output module.

[0014] The applications of this invention include, but are not limited to: non-destructive testing of bridge steel structures, weld inspection of pressure vessels, corrosion monitoring of building steel structures, inspection of key components in the aerospace field, and inspection of power transmission line towers. In these applications, this invention can effectively detect hidden defects such as closed microcracks, incomplete penetration, and slag inclusions that are difficult to detect using traditional visual inspection methods, providing reliable technical assurance for the safe operation of steel structures.

[0015] Compared with the prior art, the present invention has the following beneficial effects: By combining acoustic vibration excitation with visual capture, latent defects such as closed microcracks can be effectively activated, causing them to produce identifiable micro-amplitude vibrations, thus solving the technical problem that simple visual inspection cannot detect hidden defects. By combining a high frame rate industrial camera with a video magnification algorithm, sub-pixel level vibration measurement is achieved, significantly improving detection sensitivity; By combining visual vibration spectrum with acoustic feedback signal through cross-modal fusion analysis, environmental noise interference is effectively reduced and the reliability of detection results is improved. The entire detection process is non-contact, fast, and highly automated, meeting the application requirements of intelligent detection. The defect severity rating module classifies defects based on the signal-to-noise ratio, providing a quantitative reference for subsequent maintenance decisions. Support vector machine classifiers use multi-dimensional fusion features to identify defect types, and have high classification accuracy and generalization ability. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the overall technical solution architecture proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the cross-modal fusion analysis of acoustic vibration excitation and high frame rate visual capture in this invention; Figure 3 This is a flowchart illustrating the logic of the acoustic vibration excitation module in this invention performing frequency sweep excitation on a steel structure. Figure 4 This is a flowchart illustrating the logic of capturing micro-vibrations using a high frame rate industrial camera combined with a video magnification algorithm in this invention. Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow of cross-modal fusion analysis and defect detection report output in this invention. Detailed Implementation

[0017] Please refer to Figures 1 to 5 To further illustrate the technical means and effects of the present invention in order to achieve the intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0018] Example 1 I. Static System Construction of the Acoustic Vibration Excitation Module: The acoustic vibration excitation module is the core excitation unit of the detection system of this invention. Its hardware architecture consists of four main parts: a signal generator, a power amplifier, an ultrasonic transducer, and a synchronous control interface. The components are connected by carefully designed electrical connections and communication protocols to form a complete excitation signal generation and distribution system.

[0019] The signal generator employs programmable digital synthesis technology, integrating a direct digital frequency synthesizer chip to output a wide frequency range covering 20Hz to 200kHz. In this embodiment, for the inspection needs of large welded components such as bridge steel structures, the signal generator is configured to support a sweep frequency operation mode from 20Hz to 500kHz. The low-frequency band (20Hz to 10kHz) is used to excite vibration responses near the structure's natural frequencies, while the high-frequency band (100kHz to 500kHz) is used to excite local resonance in areas of microscopic defects. The signal generator has an independent digital-to-analog conversion channel with a resolution of 16 bits and a sampling rate of 100MHz, ensuring the frequency continuity and phase stability of the sweep signal. The signal generator panel features a touchscreen human-machine interface, allowing operators to set key parameters such as the sweep start frequency, end frequency, sweep period, sweep mode (linear or logarithmic), and sweep signal amplitude. The signal generator also integrates a clock synchronization module, receiving a high-precision clock signal from the main control computer via an external 10MHz reference clock input interface to ensure timing consistency among multiple devices.

[0020] The power amplifier adopts a Class D digital power amplifier architecture design, with a rated output power of 500W and a frequency response range covering DC to 1MHz. Internally, the power amplifier incorporates overcurrent and overtemperature protection circuits. When the output current exceeds a preset safety threshold (typically set to 120% of the rated current), the protection circuit will cut off the output within microseconds to prevent equipment damage caused by transducer short circuits or abnormal loads. The power amplifier's gain adjustment range is 0dB to 60dB, continuously adjustable via a potentiometer or digital control interface. Its input impedance is set to a standard value of 50Ω, matching the output impedance of the signal generator. The power amplifier also features an LED indicator array to indicate power status, signal input status, amplification operation status, and protection circuit trigger status, facilitating real-time monitoring of equipment operation.

[0021] The ultrasonic transducer employs a piezoelectric ceramic array design, with its core sensing element being a lead zirconate titanate piezoelectric ceramic sheet. A multi-layer stacked structure is used to improve electroacoustic conversion efficiency. The center frequency of the ultrasonic transducer is set to 40kHz, with a -6dB bandwidth covering 30kHz to 50kHz within this frequency range. The transducer's electroacoustic conversion efficiency can reach over 80%. The ultrasonic transducer is equipped with a stainless steel housing protection structure, achieving an IP67 protection rating, suitable for the complex environment of bridge field testing. The ultrasonic transducer is connected to the power amplifier via a BNC coaxial cable with a characteristic impedance of 50Ω. The cable length is selected between 5m and 20m depending on the field testing distance to ensure that the attenuation of the excitation signal during transmission is controlled within an acceptable range. A special acoustic coupling agent is used for coupling between the ultrasonic transducer and the steel structure surface. This coupling agent has low viscosity and high acoustic impedance matching characteristics, effectively eliminating air gaps between the transducer and the tested component, and improving the efficiency of sound wave energy transmission.

[0022] The synchronization control interface is a crucial bridge between the acoustic vibration excitation module and the high frame rate vision capture module. Its function is to ensure precise time synchronization between the application of the excitation signal and the start of image acquisition. The synchronization control interface uses TTL level triggering, with an output trigger signal level of 0V to 5V and a trigger pulse width that is continuously adjustable from 1μs to 100μs. After receiving control commands from the main control computer, the synchronization control interface first sends a frequency sweep start command to the signal generator, and then sends a frame synchronization trigger signal to the high frame rate industrial camera after a preset delay time (typically adjustable from 10ms to 500ms), ensuring that the camera begins image acquisition at the optimal moment after the steel structure begins to vibrate.

[0023] II. Static System Construction of High Frame Rate Visual Capture Module: The core device of the high frame rate visual capture module is a high-speed industrial camera. This camera uses a global shutter CMOS image sensor with a total of 2 million pixels, specifically the Gpixel GMAX200 Sensor, with a photosensitive area of ​​1 inch and a pixel size of 5.5μm × 5.5μm. The camera's maximum frame rate can reach 2000fps in full-resolution mode. In this embodiment, considering the sensitivity of defect detection and the economy of data storage, it is usually set to 1000fps, at which point the resolution can be reduced to 1280×1024 to improve the frame rate. The camera has a dynamic range of 72dB and a quantum efficiency of over 70% in the visible light band, which can meet the stringent image quality requirements of micro-vibration detection. The camera is equipped with a Camera Link high-speed data interface with a data transmission bandwidth of 6.8Gbps, ensuring that high frame rate image data can be transmitted to the image processing computer in real time.

[0024] The high frame rate industrial camera is equipped with an infrared filter. This filter uses multi-layer dielectric film coating technology to cut off the wavelength to 650nm, effectively filtering out infrared radiation with wavelengths greater than 650nm. This eliminates interference from artificial light sources such as high-pressure sodium lamps and metal halide lamps commonly found in on-site inspection environments. The filter's light transmittance reaches over 95% in the visible light band (400nm to 650nm), with negligible attenuation to image brightness. The camera lens adopts a telecentric optical design, with a magnification ratio of 1:1, an adjustable working distance of 200mm to 500mm, and a depth of field range of ±10mm, capable of covering the inspection field of typical bridge steel structure weld areas.

[0025] A high-precision external trigger synchronization interface is the key hardware for achieving time consistency between visual capture and acoustic excitation. The interface adopts the industry-standard GenICam trigger protocol, supporting both hardware and software trigger modes. In this embodiment, the hardware trigger mode is preferred, receiving TTL trigger pulses from the acoustic vibration excitation module's synchronization control interface via an SMA coaxial cable. The rising edge of the trigger signal triggers the camera to begin acquiring one frame of image. The delay time from the rising edge of the trigger pulse to the start of exposure by the image sensor is strictly calibrated to less than 100ns, ensuring that the time synchronization error between visual capture and acoustic excitation is controlled at the sub-millisecond level. An internal trigger de-shake filter is incorporated into the camera to eliminate false triggers caused by electromagnetic interference on the trigger signal line.

[0026] The video upscaling algorithm unit is deployed on a dedicated GPU computing server equipped with an NVIDIA Tesla series graphics processor and 24GB of video memory, used for real-time computation of the Lagrange video upscaling algorithm. The core idea of ​​the Lagrange video upscaling algorithm is to decompose the input video sequence into a multi-scale spatial pyramid representation, then perform temporal domain upscaling on the video frames of each pyramid layer within a specific frequency band, and finally reassemble the upscaled pyramid layers to output the video. In this embodiment, the number of spatial pyramid decomposition layers is set to 4, and a bandpass filter is used for temporal filtering, with a passband frequency range of 0.5Hz to 10Hz, corresponding to the frequency range of micro-amplitude vibrations in the steel structure defect area. The video upscaling factor is continuously adjustable from 20x to 100x, manually set by the operator or automatically optimized by the system based on on-site lighting conditions and defect vibration amplitude. The measurement of the upscaled vibration amplitude uses a sub-pixel registration algorithm, an improved version based on the Lucas-Kanade optical flow method, capable of achieving displacement measurement accuracy at the 1 / 10 pixel level.

[0027] III. Static System Construction of the Cross-Modal Fusion Analysis Module: The hardware platform for the cross-modal fusion analysis module adopts an industrial-grade server architecture, equipped with a multi-core CPU processor and a massively parallel computing GPU accelerator card. The CPU processor is an Intel Xeon series processor with a clock speed of 3.0GHz and 8 cores, used for computational tasks such as time-series data preprocessing, feature extraction, and classifier inference. The GPU accelerator card is an NVIDIA Quadro series, containing 2048 CUDA cores, used for massively parallel computation of video upscaling algorithms and accelerated training and inference of support vector machine classifiers. The system memory capacity is configured with 64GB of DDR4 ECC register memory for caching high frame rate video data and intermediate computation results. The storage system uses an NVMe solid-state drive array with a total capacity of 2TB and a read / write speed of 3000MB / s, meeting the real-time writing requirements of high frame rate video data.

[0028] The acoustic sensor array is a crucial component of the cross-modal fusion analysis module, used to acquire the response signals of the steel structure to acoustic vibration excitation. The acoustic sensors are piezoelectric accelerometers with a sensitivity of 100mV / g, a measurement range of ±50g, and a frequency response range covering 0.5Hz to 10kHz. The acoustic sensors are connected to the data acquisition module via BNC coaxial cables. The data acquisition module has 8 channels of simultaneous sampling capability, a sampling rate of 102.4kHz per channel, and a resolution of 24 bits. The data acquisition module integrates an anti-aliasing filter with a cutoff frequency set to 50kHz, effectively preventing spectral aliasing during sampling. The acoustic sensor array employs a multi-point distributed layout, with 4 to 16 sensor points strategically placed on the surface of the steel structure under test to achieve comprehensive acquisition of the excitation response signals and maximize spatial coverage.

[0029] IV. Dynamic Evolution of the Detection Method Flow: Based on the above system architecture, the defect detection method in this embodiment is executed in sequence according to the following steps. Each step is tightly coupled through data flow and control signals to form a complete detection closed loop.

[0030] Step S1-1: Pre-configuration stage of detection parameters: Before the formal testing begins, operators pre-configure the testing system parameters through the human-machine interface of the main control computer. First, the sweep excitation parameters are set: the sweep start frequency is set to 30Hz, the end frequency to 200Hz, the sweep period to 10 seconds, the sweep mode to linear sweep, and the excitation signal amplitude to 70% of the power amplifier's rated output power. Then, the visual capture parameters are set: the camera frame rate is set to 1000fps, the resolution to 1280×1024, the exposure time to 0.5ms, and the gain to 6dB. Finally, the video magnification parameters are set: the spatial pyramid decomposition layer number is set to 4, the temporal bandpass filter passband is set to 0.5Hz to 10Hz, and the video magnification factor is set to 50x. After the above parameters are configured, the system automatically performs a self-test to verify the communication connections between modules. Once confirmed to be correct, it enters the testing state.

[0031] Step S1-2: Acoustic vibration excitation application stage: After receiving the start command from the main control computer, the signal generator generates an excitation signal according to the preset frequency sweep parameters. The excitation signal is first amplified by a power amplifier, and the amplified electrical signal drives the ultrasonic transducer to convert the electrical signal into acoustic mechanical vibration. The sound waves emitted by the ultrasonic transducer are transmitted to the steel structure surface through an acoustic coupling agent, generating longitudinal and transverse waves within the steel structure and propagating in all directions. When the sound waves propagate to the defect area, due to the structural discontinuities in the defect area (closed cracks, incomplete penetration, slag inclusions, etc.), the sound waves undergo reflection, refraction, and mode conversion at the defect interface, resulting in a significant difference in the vibration response characteristics of the defect area compared to the surrounding healthy area. The vibration response of the defect area contains rich defect information, such as the location, depth, type, and severity of the defect, which is manifested on the steel structure surface in the form of micro-amplitude vibrations.

[0032] Simultaneously with the application of the excitation signal, the synchronization control interface sends a frame synchronization trigger signal to the high frame rate industrial camera according to a preset delay time, ensuring that the camera begins image acquisition at the optimal moment after the steel structure begins to vibrate. The delay time of the synchronization trigger signal is calibrated based on the propagation characteristics of the excitation signal and the response characteristics of the camera, and is typically set to 50ms to 100ms.

[0033] Step S2-1: High frame rate image acquisition stage: Upon receiving the synchronization trigger signal, the high-frame-rate industrial camera immediately begins continuous image acquisition at the set frame rate and resolution. The image sensor converts the acquired light signals into electrical signals and transmits them in real time to the GPU computing server via the Camera Link interface. The image acquisition process lasts for the entire frequency sweep cycle (10 seconds), acquiring a total of 10,000 frames to form a complete time-series image sequence. During transmission, the image data is compressed using a lossless compression algorithm to reduce storage space usage and data transmission bandwidth requirements. The GPU computing server receives and caches the image data in real time, preparing it for subsequent video upscaling processing.

[0034] Step S2-2: Video magnification and vibration extraction stage: Upon receiving the time-series image sequence, the GPU computing server immediately initiates the Lagrange video upscaling algorithm to process the image sequence. First, spatial pyramid decomposition is performed on the input video, dividing each frame into four pyramid representations at different resolution scales. Correlation between the pyramid layers is established through Gaussian downsampling. Then, bandpass filtering is applied to each pyramid layer in the temporal domain to extract vibration components within the 0.5Hz to 10Hz frequency band, corresponding to the characteristic frequency range of micro-amplitude vibrations in the steel structure defect area. Next, the extracted vibration components are upscaled in the temporal domain at a magnification of 50 times, amplifying the micro-amplitude vibrations to within the visible range. Finally, the upscaled pyramid layers are synthesized to reconstruct the output video.

[0035] Subpixel registration algorithms perform precise positioning measurements on magnified video. The algorithm first selects several feature points (corner points, edge points, etc.) in the video sequence, and then tracks the displacement of these feature points between consecutive frames using the Lucas-Kanade optical flow method. The displacement data output by the algorithm represents the vibration amplitude at each point on the steel structure surface, forming a visualized vibration spectrum. The spatial resolution of the vibration spectrum is determined by both the camera resolution and the accuracy of the registration algorithm, achieving subpixel levels.

[0036] Step S3-1: Acoustic feedback signal acquisition stage: The acoustic sensor array begins acquiring vibration response signals from the steel structure simultaneously with the application of acoustic vibration excitation. Eight piezoelectric accelerometers sample synchronously at a sampling rate of 102.4 kHz, with a sampling duration equal to the sweep frequency period (10 seconds). Each sensor collects 1,024,000 data points. The acquired raw signals are first preprocessed by an anti-aliasing filter and amplifier within the data acquisition module, and then transmitted to the main control computer via a USB 3.0 interface for storage and subsequent analysis. The response signals acquired by the acoustic sensors contain information such as the type, location, and severity of defects, which can be compared and verified with visual vibration maps.

[0037] Step S3-2: Cross-modal fusion analysis stage: The cross-modal fusion analysis module receives two data streams: a visual vibration spectrum and an acoustic feedback signal. It first performs time alignment processing. Since visual image acquisition and acoustic signal acquisition are performed by different hardware devices, there is a certain time deviation. The time alignment algorithm calculates the time offset between the two signals through cross-correlation. Then, based on the offset, the visual vibration spectrum and acoustic feedback signal are resampled and aligned to ensure that the visual vibration information and acoustic vibration information at the same moment correspond precisely. The accuracy requirement for time alignment is controlled within 1ms.

[0038] After time alignment is completed, cross-modal fusion analysis enters the feature extraction stage. The visual vibration spectrum is segmented into several sub-regions, and vibration time-series signals are extracted from each sub-region. Characteristic parameters such as the mean, variance, maximum value, and peak frequency of the vibration amplitude are statistically analyzed. The acoustic feedback signal undergoes spectral analysis, and the frequency components of the signal are calculated using Fast Fourier Transform to extract characteristic parameters such as the principal vibration frequency, energy distribution characteristics, and harmonic components.

[0039] After feature extraction is complete, the correlation analysis stage begins. The correlation coefficient between the visual vibration signal and the acoustic feedback signal is calculated, and the presence of a defect is determined using the following formula:

[0040] in, The first visual vibration signal One sampling point, The first part representing the acoustic feedback signal Each sampling point, and These are the corresponding means. Total number of sampling points. Correlation coefficient. The value range is from -1 to 1, when If the correlation coefficient is greater than a preset threshold (usually set to 0.7), the area is considered defective. If the correlation coefficient is less than the threshold, it indicates a lack of correlation between the visual vibration and acoustic response in the area, which may be caused by environmental noise interference, and is therefore considered defect-free.

[0041] Step S4-1: Defect Type Classification Stage: The cross-modal fusion analysis module classifies regions with identified defects. The classifier employs a support vector machine (SVM) algorithm, with input features consisting of a fused feature vector of vibration amplitude, frequency, phase, and acoustic energy density, resulting in a 4-dimensional feature set. The SVM classifier uses a radial basis function (RBF) kernel with a penalty parameter of 10. Classification categories include closed cracks, welding defects, corrosion areas, and stress concentration zones. Before deployment, the classifier underwent extensive training on numerous samples containing standard feature patterns of various defects, maintaining a classification accuracy above 90%. Classification results are labeled on the vibration spectrum, enabling visualization of defect types.

[0042] Step S4-2: Defect Severity Rating Stage The defect severity rating module calculates the signal-to-noise ratio (SNR) based on the ratio of vibration amplitude to background noise, and then rates the defect severity according to the preset range within which the SNR falls. The formula for calculating the SNR is:

[0043] in, is the average vibration amplitude of the defect area, and is the average vibration amplitude of the background noise. The rating thresholds are set as follows: a signal-to-noise ratio (SNR) greater than 40 dB is rated as a severe defect; a SNR between 25 dB and 40 dB is rated as a moderate defect; and a SNR between 10 dB and 25 dB is rated as a minor defect. The rating results are color-coded on the defect distribution map: red indicates a severe defect, orange indicates a moderate defect, and yellow indicates a minor defect.

[0044] Step S4-3: Test Report Output Stage The defect detection report output module integrates the results of defect location, type classification, and severity rating to generate the final detection report. The report includes basic information about the inspected object (name, specifications, inspection date, etc.), inspection parameter settings, a general defect distribution map, detailed parameters for each defect (location coordinates, type, rating, confidence level, etc.), and inspection conclusions and recommendations. The report is output in PDF format and supports electronic signatures and print archiving.

[0045] V. System Fault Tolerance and Exception Handling Mechanism: This embodiment is designed with a robust system fault tolerance and anomaly handling mechanism to ensure the reliability of the detection process and the integrity of the data in complex industrial environments.

[0046] When the power amplifier of the acoustic vibration excitation module detects an output abnormality (such as overcurrent or overtemperature), the protection circuit immediately cuts off the output and sends a fault alarm signal to the main control computer. Upon receiving the alarm signal, the main control computer automatically stops the frequency sweep excitation and saves the collected data, while simultaneously displaying the fault type and troubleshooting suggestions on the human-machine interface. After troubleshooting, the operator can choose to continue testing from the interrupted point or restart the testing process.

[0047] When a high-frame-rate industrial camera experiences a communication interruption or image acquisition malfunction, the camera will automatically trigger a watchdog reset mechanism, attempting to re-establish the connection within a preset number of retries. If the retry fails, the camera will save the acquired image data and send an error report. Upon receiving the camera error report, the main control computer will prompt the operator to check the camera connection status. After troubleshooting, the operator can choose to re-acquire data.

[0048] When data synchronization anomalies (such as discontinuous timestamps or missing data) are detected during cross-modal fusion analysis, the system will automatically activate data interpolation and completion algorithms to restore data integrity to the greatest extent possible. For data segments that cannot be recovered, the system will mark the data anomaly in the analysis report to ensure the reliability of the detection results.

[0049] Example 2 I. Hardware adaptation for high temperature and high pressure environments: In pressure vessel weld inspection scenarios, the tested objects are typically high-temperature and high-pressure vessels in operation, and the inspection environment is characterized by high temperature, high humidity, and the presence of potentially hazardous media. To address these specific environmental factors, this embodiment provides targeted hardware adaptation for the system.

[0050] The ultrasonic transducer of the acoustic vibration excitation module features a high-temperature resistant design, with the housing material replaced by Hastelloy, extending its temperature range to 300℃. The coupling between the transducer and the steel structure surface utilizes a high-temperature acoustic coupling agent with an operating temperature range of -40℃ to 250℃, capable of withstanding the surface temperatures during pressure vessel operation. The power amplifier employs a water-cooled design with a cooling water flow rate of 5L / min, ensuring stable operation over extended periods in high-temperature environments. The signal generator's housing is explosion-proof, achieving an Exd II BT4 protection rating, suitable for industrial environments containing flammable gases.

[0051] The high frame rate industrial camera features a high-temperature resistant protective housing with an integrated semiconductor cooling module to keep the camera's operating temperature within acceptable limits. The lens is made of quartz glass, capable of withstanding higher ambient temperatures. The camera is equipped with a high-pressure air curtain protection device that continuously purges protective gas during testing to prevent high-temperature steam and dust from contaminating the optical components. The infrared filter is made of high-temperature resistant glass, ensuring stable filtering performance even at high temperatures.

[0052] The acoustic sensor features a corrosion-resistant design, with the sensing element encased in a titanium alloy shell, achieving an IP68 protection rating, making it suitable for detection environments containing corrosive media. The connecting cable between the sensor and the data acquisition module has a stainless steel sheath, capable of withstanding high temperatures and mechanical damage.

[0053] II. Optimized parameter settings for weld inspection: The main types of defects in pressure vessel welds include incomplete penetration, slag inclusions, porosity, and cracks. The response characteristics of these defects under acoustic vibration excitation differ from those of bridge steel structures. Considering the unique characteristics of weld inspection, this embodiment optimizes and adjusts the inspection parameters.

[0054] The frequency sweep excitation parameters were optimized as follows: the sweep start frequency was set to 50Hz, the sweep end frequency to 100kHz, and the sweep period to 5 seconds, using a segmented sweep method. The first low-frequency band (50Hz to 1kHz) was used to detect large defects, the second mid-frequency band (10kHz to 50kHz) was used to detect medium-sized defects, and the third high-frequency band (50kHz to 100kHz) was used to detect small defects. The excitation signal amplitude was set to 80% of the rated output power of the power amplifier to elicit a more pronounced vibration response in the weld area.

[0055] The visual capture parameters were optimized as follows: the camera frame rate was set to 2000fps and the resolution to 640×512 to improve the temporal resolution of micro-vibrations in the weld area. The bandpass filter passband of the video magnification algorithm was set to 1Hz to 20Hz, and the magnification factor was set to 80x to more clearly display the vibration characteristics of weld defects.

[0056] The cross-modal fusion analysis algorithm was optimized as follows: the sub-region size for region segmentation was set to 5mm × 5mm, which is more suitable for the fine structure of welds. The correlation coefficient threshold was adjusted to 0.75 to reduce false positive detections. The support vector machine classifier was optimized by adding two categories, porosity and incomplete penetration, to target common defect types in pressure vessel welds.

[0057] III. Dynamic Adjustment of the Testing Process: In view of the special characteristics of pressure vessel weld inspection, the inspection process in this embodiment has been adjusted accordingly in terms of execution details.

[0058] During the pre-test preparation phase, additional checks on high-temperature protection measures should be conducted to confirm that the high-temperature resistant components of the acoustic vibration excitation module and the visual capture module are properly installed, and that the refrigeration system and water cooling system are operating normally. Operators should maintain a safe distance from the container being tested, and a warning area should be set up to prevent accidental contact.

[0059] The detection execution phase employs a segmented frequency sweep method. After each frequency sweep is completed, the system automatically performs data preprocessing and preliminary analysis, identifies suspected defect areas, and then focuses on inspecting these areas in the next frequency sweep. This segmented detection method improves detection efficiency while allowing for more detailed analysis of key areas.

[0060] In the inspection report output phase, to meet the safety assessment requirements of pressure vessels, additional sections for defect hazard level assessment and maintenance recommendations have been added. The defect hazard level assessment is based on a comprehensive calculation of defect type, size, location, and operating parameters (pressure, temperature, medium, etc.), with output levels categorized into three levels: immediate shutdown, time-limited repair, and monitoring operation. The maintenance recommendations section provides specific repair method suggestions for different types of defects, such as welding repair, grinding, and replacement.

[0061] Example 3 I. Hardware upgrades for ultra-high precision testing requirements: Critical components in the aerospace field require extremely high sensitivity for defect detection. The width of closed microcracks may be less than 10 μm, making them difficult to detect effectively using traditional methods. To address this ultra-high precision requirement, this embodiment features a comprehensive upgrade to the system hardware.

[0062] The high frame rate industrial camera has been upgraded to an ultra-high resolution model, employing a 100-megapixel back-illuminated CMOS image sensor with a pixel size of 3.76μm × 3.76μm. The camera achieves a frame rate of 500fps in full-resolution mode, and in this embodiment, a 2048 × 2048 resolution can reach 1500fps. The camera's dynamic range has been improved to 80dB, and its quantum efficiency exceeds 85%, enabling it to capture sub-pixel-level vibration displacements caused by microcracks. The lens features an apochromatic design, with maximum optical distortion of less than 0.1%, ensuring high geometric accuracy even in ultra-wide field-of-view inspections.

[0063] The video upscaling algorithm has been upgraded to an improved version of the Euler video upscaling technique—the coherent video upscaling algorithm. This algorithm introduces a coherence detection mechanism based on the Lagrange video upscaling algorithm, enabling more accurate differentiation between defect vibration signals and environmental noise interference. The number of spatial pyramid decomposition layers has been increased to 6, and the time-domain filter adopts an adaptive bandwidth design, dynamically adjusting filtering parameters according to the characteristics of the input signal. The magnification has been increased to 200x, and combined with the sub-pixel registration algorithm, displacement measurement accuracy at the 1 / 50 pixel level can be achieved.

[0064] The acoustic sensor has been upgraded to a capacitive microelectromechanical system (MEMS) sensor, with sensitivity increased to 1000mV / g, measurement range expanded to ±500g, and frequency response range covering 0.1Hz to 100kHz. The sensor size has been reduced to 5mm×5mm×3mm, allowing it to be placed closer to the defect area and reducing signal attenuation along the transmission path. The data acquisition module's sampling rate has been increased to 204.8kHz with a 32-bit resolution, enabling complete capture of vibration response signals across an ultra-wide frequency band.

[0065] II. Improved accuracy of cross-modal fusion analysis: To address the extremely high requirements for defect location accuracy in aerospace component inspection, the cross-modal fusion analysis module has undergone algorithm upgrades.

[0066] The temporal alignment algorithm employs a precise alignment method based on mutual information. It determines the optimal alignment parameters by maximizing the mutual information between the visual vibration spectrum and the acoustic feedback signal, improving the temporal alignment accuracy to 0.1 ms. The feature extraction algorithm uses a time-frequency analysis method combining wavelet transform and short-time Fourier transform, simultaneously obtaining fine features in both the time and frequency domains of the vibration signal. The correlation coefficient calculation uses a sliding window approach with a window length of 100 ms and a sliding step size of 10 ms, yielding the time evolution curve of the vibration response in the defect area.

[0067] Defect localization employs a multi-sensor triangulation method, combining the response signals from multiple acoustic sensors to determine the spatial location of the defect. The localization algorithm is based on the time-difference positioning principle, calculating the three-dimensional coordinates of the defect by determining the time difference between the sound wave's propagation to each sensor. The localization accuracy can reach ±0.5mm within a detection range of 100mm × 100mm × 50mm.

[0068] III. Implementation of Intelligent Detection Process: This embodiment realizes the fully automated execution of the testing process, which greatly reduces the workload and professional skill requirements of operators.

[0069] The system integrates an automatic target recognition module, which utilizes a deep learning convolutional neural network to automatically identify and partition the structure of the tested component. The convolutional neural network employs a ResNet50 backbone network and has been trained on a large number of aerospace component images, enabling it to accurately identify different types of component structures and key detection areas. The output of the automatic target recognition module guides the adaptive configuration of the sweep frequency excitation parameters and the automatic adjustment of the camera's field of view.

[0070] The system also integrates an adaptive parameter optimization module, which analyzes signal quality in real time during detection and automatically adjusts excitation parameters and amplification algorithm parameters based on the signal-to-noise ratio. When a signal quality degradation is detected, the system automatically increases the excitation power or amplification factor; when signal saturation is detected, the system automatically decreases the parameter settings. This adaptive mechanism ensures optimal detection results are always achieved in complex and ever-changing detection environments.

[0071] The inspection report output module integrates 3D visualization capabilities, presenting defect inspection results as a 3D model that can be rotated, zoomed, and panned. The 3D model is labeled with the location, type, and rating information of each defect; clicking on a defect marker displays detailed inspection data and analysis curves. The report also supports augmented reality (AR) format output, allowing inspectors wearing AR glasses to directly overlay defect information on-site, achieving a precise correspondence between inspection results and actual components.

[0072] In summary, this invention achieves comprehensive detection of potential defects on the surface and inside of steel structures through the collaborative operation of an acoustic vibration excitation module, a high frame rate visual capture module, a cross-modal fusion analysis module, and a defect detection report output module. The precise synchronization and deep fusion between these modules enable the detection system to effectively discover latent defects such as closed microcracks that are difficult to capture using traditional visual inspection methods, significantly improving the accuracy and reliability of defect identification and providing a reliable technical guarantee for the safe operation of steel structures.

Claims

1. A method for intelligent visual defect detection in steel structures, characterized in that, include: A sweeping excitation signal within a specific frequency range is applied to the steel structure through an acoustic vibration excitation device to stimulate micro-amplitude vibration responses in the steel structure's surface and internal potential defect areas. A high frame rate industrial camera is used to synchronously acquire time-series image sequences of the steel structure surface under excitation, and the time-series image sequences are processed in combination with video magnification algorithms to extract sub-pixel level micro-amplitude vibration information and generate vibration spectrum. The acoustic feedback signal of the steel structure to the sweep frequency excitation signal is synchronously acquired by an acoustic sensor; The vibration spectrum and the acoustic feedback signal are subjected to cross-modal fusion analysis. Based on the correlation between the two in the time and frequency domains, the existence of defects is determined, the location of defects is located, and the type of defects is identified. Based on the fusion analysis results, an inspection report is generated that includes the defect location coordinates, defect type classification, and defect severity rating.

2. The intelligent visual defect detection method for steel structures according to claim 1, characterized in that, Processing the time-series image sequence to extract sub-pixel level micro-amplitude vibration information and generate vibration spectra includes: Spatial pyramid decomposition is performed on the temporal image sequence to obtain multi-scale image representation; Time-domain filtering and amplification are performed on image sequences of various scales within a preset frequency band. The magnified multi-scale image sequence is reconstructed into an enhanced video, and the displacement time-series data of each pixel is calculated using a sub-pixel registration algorithm to form a vibration spectrum.

3. The intelligent visual defect detection method for steel structures according to claim 2, characterized in that, The vibration spectrum and the acoustic feedback signal are fused across modes for analysis, including: The vibration spectrum is segmented into regions to extract the vibration time-series signals of each sub-region; The acoustic feedback signal is subjected to spectral analysis to extract the dominant frequency components and energy distribution characteristics; The correlation coefficient between the vibration time-series signal and the acoustic feedback signal of each sub-region is calculated, and the defect area is determined based on the preset correlation threshold.

4. The intelligent visual defect detection method for steel structures according to claim 3, characterized in that, Based on the results of the fusion analysis, defect types are identified, including: Construct a fused feature vector, which includes vibration amplitude, vibration frequency, vibration phase, and acoustic energy density; The fused feature vectors are input into a pre-trained support vector machine classifier, which outputs a defect type classification result. The defect types include closed cracks, welding defects, corrosion areas, and stress concentration areas.

5. The intelligent visual defect detection method for steel structures according to claim 4, characterized in that, The vibration spectrum is segmented into regions to extract the vibration time-series signals of each sub-region, including: Adaptive mesh generation of vibration spectrum based on image gradient or structural prior information; The pixel vibration data within each grid area are statistically aggregated to generate a time-series signal representing the overall vibration characteristics of that area.

6. The intelligent visual defect detection method for steel structures according to claim 5, characterized in that, Calculate the correlation coefficient between the vibration time-series signal and the acoustic feedback signal in each sub-region, including: Timing alignment is performed on the two signals to eliminate time discrepancies between the acquisition devices; The local correlation coefficient is calculated segment by segment on the aligned signal using a sliding window method. The maximum or mean of the local correlation coefficient is used as the final correlation measure for that sub-region.

7. The intelligent visual defect detection method for steel structures according to claim 6, characterized in that, The degree of defect is rated based on the results of the fusion analysis, including: The signal-to-noise ratio is obtained by calculating the ratio of the vibration amplitude in the defect area to the vibration amplitude in the background noise area. The signal-to-noise ratio is compared with multiple preset rating threshold ranges to determine the defect severity level, which includes minor defects, moderate defects, and severe defects.

8. The intelligent visual defect detection method for steel structures according to claim 7, characterized in that, The spatial pyramid decomposition of the temporal image sequence is performed to obtain a multi-scale image representation, including: Each frame of the image is downsampled using either a Gaussian pyramid or a Laplacian pyramid. Subsequent temporal filtering and magnification operations are performed independently on each pyramid image to preserve vibration details at different spatial scales.

9. A steel structure intelligent visual defect detection system, characterized in that, include: The acoustic vibration excitation module is used to apply a sweep frequency excitation signal within a specific frequency range to the steel structure to excite the defect area to generate a micro-amplitude vibration response. A high frame rate visual capture module is used to synchronously acquire time-series image sequences of the steel structure surface and generate vibration spectra through video magnification algorithms; The acoustic sensing module is used to synchronously acquire the acoustic feedback signal of the steel structure to the excitation signal; The cross-modal fusion analysis module is used to fuse and analyze the vibration spectrum and the acoustic feedback signal to achieve defect localization, type identification and severity rating; The defect detection report output module is used to generate a detection report containing defect location, type, and rating information based on the analysis results.

10. The intelligent visual defect detection system for steel structures according to claim 9, characterized in that, The cross-modal fusion analysis module is used for: The vibration spectrum is segmented into regions to extract the vibration time-series signals of each sub-region; The acoustic feedback signal is subjected to spectral analysis to extract the dominant frequency components and energy distribution characteristics; The correlation coefficient between the vibration time-series signal and the acoustic feedback signal of each sub-region is calculated, and the defect area is determined based on the preset correlation threshold.