A multi-modal industrial data three-dimensional fusion and visualization detection method based on brain-like perception
By using brain-like perception-driven multimodal data fusion and tensor decomposition, the problem of synchronous high-precision three-dimensional visualization of internal and external defects in weld inspection is solved, realizing high precision and adaptive capability of weld inspection, which is suitable for industrial weld inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUILIN UNIV OF ELECTRONIC TECH
- Filing Date
- 2026-03-09
- Publication Date
- 2026-05-29
AI Technical Summary
Existing weld inspection technologies suffer from problems such as low cross-modal fusion accuracy, insufficient visualization of reconstruction results, high rate of missed detection of small and dense defects, and poor adaptability to dynamic working conditions in multimodal data fusion, making it difficult to achieve synchronous high-precision three-dimensional visualization of internal and external defects.
A multimodal data fusion method based on brain-like perception is adopted. Data is collected synchronously by a binocular camera and an ultrasonic array sensor. Adaptive filtering and noise reduction, temporal and spatial alignment are performed. The spatiotemporal feature extraction and fusion are simulated by the human brain processing mechanism. Combined with tensor decomposition, three-dimensional reconstruction is performed to achieve high-precision visualization of internal and external defects in welds.
It achieves deep fusion of multimodal data, improves the accuracy and robustness of weld inspection, and can adaptively adjust under dynamic working conditions, significantly improving the detection effect of small defects.
Smart Images

Figure CN122115986A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of nondestructive testing technology, specifically relating to a method and system for synchronous detection of internal and external defects in welds based on multimodal data fusion and brain-like perception in three-dimensional visualization. Background Technology
[0002] Welding, as a critical structural manufacturing process, presents a core challenge in ensuring the safety and manufacturing quality of major engineering projects through the simultaneous, accurate, and visualized detection of internal and surface defects. Traditional nondestructive testing methods (such as ultrasonic, radiographic, eddy current, and visual methods) are mostly based on single-modal data, making it difficult to achieve collaborative perception and three-dimensional visualization of internal and external defects. Although multimodal data fusion technology has been developed, in industrial weld inspection scenarios, there are still fundamental differences in spatiotemporal references, physical mechanisms, and feature dimensions between multi-source heterogeneous data (such as optical images and ultrasonic images). This leads to problems such as low cross-modal fusion accuracy, insufficient visualization of reconstruction results, high missed detection rate of small and dense defects, and poor adaptability to dynamic working conditions.
[0003] In the prior art, patent CN116773547B discloses a weld defect detection method based on multimodal data fusion, which uses visual and ultrasonic related data fusion, but feature fusion relies on manual design and has limited adaptability. Patent CN114863018A proposes a three-dimensional reconstruction system based on deep learning multi-sensor data, which achieves point cloud generation, but does not simultaneously reconstruct internal and external weld defects, and does not consider unified spatiotemporal feature modeling. Patent CN117115538A relates to a biomimetic perception-driven industrial vision inspection method that simulates the visual cortex processing mechanism, but it is limited to surface defect detection and does not fuse internal ultrasonic data. Patent CN120491605A discloses a multimodal data fusion method based on tensor decomposition for structural health monitoring, but it lacks spatiotemporal consistency modeling and inverse problem solving mechanisms. In addition, existing methods are mostly limited to a single physical mode or simple multimodal stitching, lacking deep cross-modal fusion guided by biomimetic perception mechanisms. They are difficult to achieve integrated, high-fidelity 3D visualization reconstruction of internal and external defects, and do not consider the adaptive evolution requirements of detection models in dynamic industrial scenarios.
[0004] In summary, existing weld inspection technologies still have significant shortcomings in terms of the accuracy of multimodal data fusion, the completeness of 3D visualization reconstruction, and the system's adaptability in dynamic scenarios. Therefore, there is an urgent need for a new intelligent inspection method that can deeply integrate multi-source heterogeneous data, achieve synchronous high-precision 3D visualization of internal and external defects, and possess continuous learning and adaptation capabilities. Summary of the Invention
[0005] To address the problem that existing industrial weld nondestructive testing technologies struggle to achieve simultaneous, high-precision detection and three-dimensional visualization of internal and external defects using single-modal or simple multimodal fusion methods, this invention provides a brain-like perception-based multimodal industrial data three-dimensional fusion and visualization detection method. This method aims to simulate the multimodal information processing mechanism of the human brain, deeply integrating optical and ultrasonic heterogeneous data to achieve integrated, high-precision, visualized, simultaneous detection and quantitative analysis of internal and external weld defects.
[0006] To achieve the above objectives, the present invention provides the following solution:
[0007] A method for 3D fusion and visualization detection of multimodal industrial data based on brain-like perception includes the following steps:
[0008] S1. Synchronous acquisition and preprocessing of multimodal data
[0009] Optical images of the weld surface and ultrasonic data of the interior are acquired simultaneously by a binocular camera and an ultrasonic array sensor. The two types of data are then subjected to adaptive filtering and noise reduction processing to provide high-quality input for subsequent steps.
[0010] S2, Multi-source data temporal and spatial alignment
[0011] The denoised optical and ultrasonic data are subjected to temporal alignment and spatial registration to unify them into the same spatiotemporal coordinate system, thereby eliminating data inconsistencies caused by differences in sensor sampling and spatial location.
[0012] S3, Spatiotemporal Feature Fusion Driven by Brain-like Perception
[0013] This step is the core of the invention. By simulating the mechanism by which the human brain processes multimodal information, a complementary spatiotemporal feature extraction and fusion framework is constructed.
[0014] S3.1 Temporal Feature Extraction
[0015] Based on aligned optical and ultrasound image data, a time-series inverse model is established using the approach of solving the inverse problem of electroencephalogram (EEG) signals, and dynamic features with high temporal resolution are deduced from the observation data.
[0016] S3.2 Spatial Feature Extraction
[0017] Using the same data, a spatially sparse inverse model is established by solving the inverse problem of functional magnetic resonance imaging (fMRI) to inversely deduce the distribution characteristics with high spatial resolution from the observation data.
[0018] S3.3 Adaptive Fusion Based on Composite Spatiotemporal Metric
[0019] Adaptive Feature Fusion: Introducing the concept of spatiotemporal metric, a composite spatiotemporal fusion model is constructed. Based on the confidence and uncertainty of each feature component, the fusion weight is dynamically calculated to achieve adaptive weighted fusion of temporal and spatial features, forming a unified spatiotemporal feature representation.
[0020] S4. Three-dimensional reconstruction based on tensor decomposition
[0021] The fused spatiotemporal features are organized into high-order tensors, and the feature tensors are reduced in dimensionality and their structure is extracted using tensor decomposition. Then, through reconstruction and optimization, high-precision three-dimensional visualization images of internal and external defects in the weld are generated.
[0022] S5, Intelligent Defect Detection and Quantitative Analysis
[0023] By inputting 3D visualized images into a lightweight multi-task detection network, automatic location, classification, and dimensional quantification of defects can be achieved. The network is trained by introducing physical constraints to improve the physical rationality and reliability of the detection results.
[0024] The present invention also provides a multimodal three-dimensional visualization detection system for internal and external defects in welds for implementing the above method, characterized in that it comprises:
[0025] A multimodal image acquisition unit, including a binocular structured light camera and an ultrasonic array sensor, is used to simultaneously acquire data on the surface and interior of the weld.
[0026] The data fusion processing unit is used to perform preprocessing, feature extraction and fusion, three-dimensional reconstruction and defect detection algorithms for the multimodal data;
[0027] The visualization and output unit is used to display the 3D reconstruction results and output a structured inspection report;
[0028] The synchronization control unit is used to coordinate the synchronous acquisition of data from each sensor.
[0029] An embedded computing unit is used to deploy and run the algorithm on an edge device, supporting real-time or near-real-time processing of the system.
[0030] Beneficial effects of the present invention
[0031] 1. By introducing a brain-like perception mechanism, a complete framework from temporal-spatial feature complementarity extraction to adaptive fusion was constructed, realizing deep fusion of optical and ultrasonic data at the mechanistic level, and effectively solving the problem of spatiotemporal mismatch of multimodal data.
[0032] 2. The proposed composite spatiotemporal metric and adaptive weighting mechanism enable the system to dynamically adjust the fusion strategy based on data quality, thereby improving the robustness and adaptability of the method in complex industrial environments.
[0033] 3. Using tensor decomposition for 3D reconstruction can fully explore the intrinsic structure of high-dimensional features, significantly improve the accuracy and integrity of the reconstructed image, and is especially beneficial for revealing small and internal defects.
[0034] 4. The entire methodology is clear and the steps are logically connected, forming a calculable and optimizable technical closed loop. This not only improves the detection performance but also enhances the interpretability and engineering feasibility of the method, providing a new technical path for industrial intelligent non-destructive testing. Attached Figure Description
[0035] Figure 1 This is the overall method flowchart.
[0036] Figure 2 This is a schematic diagram of feature extraction and fusion of data.
[0037] Figure 3 A schematic diagram of 3D reconstruction using tensor decomposition and joint inverse optimization.
[0038] Figure 4 This is a flowchart for intelligent defect detection and quantitative analysis.
[0039] Figure 5 This is a schematic diagram of the structural components of the detection system. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of protection of this invention.
[0041] This embodiment provides a detailed implementation process for a multimodal industrial data 3D fusion and visualization detection method based on brain-like perception. The method flow is as follows: Figure 1 As shown.
[0042] S1. Synchronous acquisition and preprocessing of multimodal data
[0043] In this embodiment, a Hikvision MV-CA020-20GM binocular structured light camera and an Olympus OmniScan X3 ultrasonic phased array sensor are used for data acquisition. Both receive the same high-precision clock signal via an external synchronization trigger to ensure synchronization of the acquisition operations. When scanning the weld area, the binocular camera acquires a sequence of surface optical images, while the ultrasonic array sensor transmits and receives a sequence of ultrasonic pulse echoes. The acquired raw optical image data is denoted as a matrix. Ultrasonic echo data are denoted as a matrix. In the preprocessing stage, optical image data is processed. With ultrasound data Adaptive filtering and denoising are performed separately to suppress environmental noise and sensor noise, resulting in denoised data. and , which serves as input for subsequent processing.
[0044] S2, Multi-source data temporal and spatial alignment
[0045] First, time synchronization is performed. Since the two sensors have different sampling frequencies, the data needs to be mapped to a unified time axis. This implementation uses a time alignment method based on a dynamic time warping algorithm. First, the grayscale change curve of the weld area in the optical image sequence is extracted as a reference time series, and the ultrasonic signal envelope is extracted as the target time series. Then, the optimal bending path is calculated using DTW, and the ultrasonic signal is nonlinearly stretched or compressed in the time dimension to achieve time alignment with the optical image, ultimately controlling the time alignment error between the two to within ±5ms. Second, spatial registration is performed. A high-precision calibration board is used beforehand to perform joint extrinsic parameter calibration on the binocular camera and ultrasonic sensor, obtaining the coordinate transformation matrix from the ultrasonic sensor coordinate system to the camera coordinate system (with the left camera as the reference). The three-dimensional coordinates of each defect point detected by ultrasound. Mapping to the camera coordinate system via homogeneous coordinate transformation:
[0046]
[0047] This step unifies the optical images and ultrasound data to the same reference in both time and space.
[0048] S3, Spatiotemporal Feature Fusion Driven by Brain-like Perception
[0049] This step is the core of the invention, and its detailed implementation steps are as follows: Figure 2 As shown.
[0050] S3.1 Time Feature Extraction Based on EEG Inverse Problem Solving
[0051] This implementation simulates the high temporal resolution mechanism of the EEG inverse problem, extracting high dynamic temporal features from aligned time-series data, and constructing the following optimization problem:
[0052]
[0053] in: and This is the known forward model matrix constructed based on the sensor physical model. Let be the first-order difference operator matrix in the time domain. The regularization parameter controls the strength of the time smoothing constraint. An iterative optimization algorithm is used to solve the problem, yielding the optimal time source activity matrix. and Vectorize and concatenate them to form a unified time feature vector. .
[0054] S3.2 Spatial Feature Extraction Based on fMRI Inverse Problem Solving
[0055] This implementation simulates the high spatial resolution mechanism of the fMRI inverse problem to extract the spatial distribution features of the data. The following optimization problem is constructed:
[0056]
[0057] in: This indicates a convolution operation. and The point spread function is known. It is the L1 norm, used to introduce spatial sparsity priors. This is the regularization parameter.
[0058] The spatial source distribution is obtained by using a sparse optimization algorithm. and Vectorize and concatenate them to form a unified spatial feature vector. .
[0059] S3.3 Adaptive Fusion Based on Composite Spatiotemporal Metric
[0060] To achieve the organic integration of spatiotemporal characteristics, this implementation introduces the concept of relativistic spacetime transformation to construct a composite spacetime metric. A four-dimensional spacetime metric tensor is defined. as follows:
[0061]
[0062] in, and These are time weights and spatial weights, respectively, related to the uncertainty of time characteristics. and spatial characteristics uncertainty It is inversely proportional, thus achieving adaptive adjustment. Using this metric, the square of the spatiotemporal distance is calculated. :
[0063]
[0064] in This is a spatiotemporal interval vector. Finally, the features are fused. Calculated by weighted sum:
[0065]
[0066] in, and To base on feature local confidence Dynamically calculated diagonal weight matrix.
[0067] S4. Three-dimensional reconstruction based on tensor decomposition
[0068] This step aims to reconstruct the three-dimensional structure of internal and external defects in the weld from the fusion features, as illustrated in the diagram below. Figure 3 As shown. The fused features The tensor is reorganized into a higher-order tensor to characterize the spatiotemporal feature field of the weld region. For efficient and robust 3D reconstruction, Tucker decomposition is used to reduce the dimensionality and extract the structure of this tensor.
[0069]
[0070] in: For the core tensor; This corresponds to a high-dimensional factor matrix; Represent the tensor and the matrix of the first Modular multiplication of factors. The above decomposition is solved using an optimized algorithm, and tensor reconstruction is performed using the core tensor and factor matrices to generate a high-precision 3D visualization image of internal and external weld defects. .
[0071] S5, Intelligent Defect Detection and Quantitative Analysis
[0072] This step involves reconstructing the 3D visualization image. The process for performing automated defect analysis is as follows: Figure 4 As shown. A 3D visualization image. Input a lightweight multi-task detection network. This network is an improvement on the YOLO architecture, and its forward propagation process can be described as follows:
[0073]
[0074] in, For having trainable parameters The network function outputs a set of defect bounding boxes B, a category probability vector P, and a quantization size parameter d. The network is trained using a loss function incorporating physical constraints to achieve accurate defect localization, identification, and quantization analysis.
[0075] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the invention. Adjustments to hardware selection, network structure, and parameters can be made by those skilled in the art without departing from the core principles of the invention, and such adjustments should be considered within the scope of protection of the present invention.
Claims
1. A method for three-dimensional fusion and visualization detection of multimodal industrial data based on brain-like perception, characterized in that, Includes the following steps: S1. Optical image data of the weld surface and ultrasonic image data of the interior are acquired simultaneously through a binocular camera and an ultrasonic array sensor, and adaptive filtering and noise reduction processing are performed on the optical image data and ultrasonic image data respectively. S2. Based on the denoised optical and ultrasonic image data, time alignment processing is performed through dynamic time warping algorithm to achieve time synchronization of multi-source data. Then, geometric mapping processing is performed through coordinate transformation matrix to complete the spatial registration of multi-source data. S3. Perform temporal feature extraction based on EEG inverse problem solving and spatial feature extraction based on functional magnetic resonance imaging (fMRI) inverse problem solving on the aligned optical image data and ultrasound image data respectively. Then, perform weighted fusion of the extracted temporal and spatial features and construct a unified composite spatiotemporal metric data. S4. Based on the composite spatiotemporal metric data, a three-dimensional reconstruction is performed using tensor decomposition and joint inverse solution methods to obtain a three-dimensional visualization image of internal and external defects in the weld. S5. By inputting the three-dimensional visualization image into the lightweight detection network, the location, identification, and quantitative analysis of defects are achieved.
2. The method according to claim 1, characterized in that, In S2: By using a time-series mapping and dynamic time warping algorithm based on a synchronous trigger signal, time synchronization of optical and ultrasonic image data is achieved; spatial registration of image data of different modalities is achieved by using a coordinate transformation matrix calculated based on the sensor extrinsic calibration results.
3. The method according to claim 1, characterized in that, In S3: The specific method for constructing a composite spacetime metric is to introduce the concept of relativistic spacetime transformation to construct a four-dimensional spacetime metric tensor. Its form is:
4. Among them, The weights are set based on the high temporal resolution characteristics of EEG signals. Weights set based on the high spatial resolution characteristics of fMRI signals; Calculate the square of the spatiotemporal distance using this metric. This allows for the quantification and fusion of spatiotemporal differences in cross-modal image data.
5. The method according to claim 1, characterized in that, In S4: By organizing the fused features into high-dimensional tensors The tensor is then decomposed and reconstructed using the Tucker decomposition method to generate the final 3D visualization image.
6. A multimodal three-dimensional visualization detection system for weld internal and external defects for implementing the method of any one of claims 1 to 4, characterized in that, include: The multimodal image acquisition unit acquires optical images of the weld surface through a binocular structured light camera and acquires ultrasonic images of the weld interior through an ultrasonic array sensor. The data fusion processing unit is used to execute the data preprocessing, cross-modal feature fusion and three-dimensional reconstruction algorithm described in claim 1, so as to realize the conversion from multimodal image data to three-dimensional visualization results; The visualization and output unit is used to display 3D reconstructed images and output defect detection reports. The synchronization control unit coordinates the acquisition actions of each sensor in the multimodal image acquisition unit through trigger signals to achieve synchronous image data acquisition; The embedded computing unit enables the system to operate in real time or near real time by deploying the algorithms in the data fusion processing unit.