Welding structure performance intelligent evaluation method and system based on multi-modal large model
By using multimodal large-scale model data acquisition, preprocessing, fusion, and mechanism-constrained reasoning, the problem of instability and uncertainty in multimodal data fusion in welded structure evaluation is solved, achieving high-precision and interpretable evaluation results and reducing decision-making risks. It is applicable to the life management of welded structures such as bridges, ships, engineering machinery, and automobiles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI UNIV
- Filing Date
- 2026-03-23
- Publication Date
- 2026-05-05
AI Technical Summary
Existing welded structure performance evaluation technologies suffer from problems such as unstable multimodal data fusion, insufficient uncertainty quantification, and high decision-making risks, making it difficult to achieve accurate evaluation under complex working conditions.
A multimodal large model is used for data acquisition and access, preprocessing and calibration alignment, multimodal large model layer, mechanism constraint and state inference, interpretability and decision support, and closed-loop learning, to achieve high-precision spatiotemporal alignment, quality-aware fusion, uncertainty quantification, and mechanism constraint inference of multi-source heterogeneous data.
It achieves high-precision spatiotemporal alignment and traceability of multimodal data, improves the stability and interpretability of evaluation results, reduces the risk of false alarms and missed alarms, and provides reliable decision-making basis and long-term adaptability.
Smart Images

Figure CN121980522A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of in-service health monitoring, non-destructive testing and life assessment of welded structures, and specifically relates to an intelligent evaluation method and system for the performance of welded structures based on a multimodal large model. Background Technology
[0002] Welded structures are widely used in various industrial equipment. Under complex load spectra, corrosive environments, and temperature fluctuations, their weld toes, weld roots, and heat-affected zones are prone to fatigue cracking, defect propagation, or localized failure, directly impacting equipment operational safety. Existing performance evaluation technologies have significant shortcomings:
[0003] 1. Relying on single-modal data, health monitoring based solely on signals such as strain / vibration / acoustic emission, or defect interpretation based solely on non-destructive testing images such as ultrasound / radiation, is prone to false alarms and false negatives when noise increases, operating conditions change, or testing conditions change, and lacks cross-modal evidence fusion and interpretability;
[0004] 2. Although multimodal fusion technology has been applied, the field data has problems such as inconsistent sampling frequency, timestamp drift, difficulty in unifying detection pose and weld coordinates, data loss and large quality fluctuations, which makes it difficult to stably reproduce the fusion results.
[0005] 3. Traditional models lack a linkage mechanism for uncertainty quantification, out-of-distribution identification, and re-examination strategies, resulting in a lack of confidence and risk basis for engineering decisions, making it difficult to meet the needs of practical applications.
[0006] Therefore, there is an urgent need for an integrated evaluation method that can achieve multimodal data alignment, quality-aware fusion, mechanism-constrained reasoning, and uncertainty quantification, in order to address the core pain points of poor stability, insufficient interpretability, and high decision-making risk in existing technologies. Summary of the Invention
[0007] To address the problems existing in the prior art, this invention provides a method and system for intelligent evaluation of welded structure performance based on a multimodal large model. The purpose is to achieve deep fusion and intelligent processing of multi-source heterogeneous data, including online monitoring signals, non-destructive testing images, visual or infrared images, and working condition text, in order to solve the problem of insufficient evaluation accuracy of the prior art under conditions of multi-source heterogeneity, quality fluctuation, and working condition transfer.
[0008] To achieve the above objectives, the specific solution of the present invention is as follows:
[0009] A smart performance evaluation system for welded structures based on a multimodal large model, comprising:
[0010] The data acquisition and access layer is used to acquire online monitoring signals, non-destructive testing images, visual images or infrared images, as well as working condition text or historical text, and output raw multimodal data packets and their metadata. A unified traceability identifier is written into the metadata. The unified traceability identifier serves as the primary key for cross-layer data association and includes at least a structure identifier, a weld identifier, and an acquisition timestamp.
[0011] The preprocessing and calibration alignment layer, connected to the data acquisition and access layer, is used to receive the original multimodal data packets and their metadata. It performs denoising, segmentation, outlier processing, and missing value completion on the online monitoring signals, non-destructive testing images, visual images, or infrared images in the original multimodal data packets, and conducts quality assessment, calculating modal quality indicators such as signal-to-noise ratio, sharpness, and missing rate. It is also used for time synchronization calibration and spatial positioning calibration, ROI mapping, mapping signal segments to the weld seam region of interest or weld toe region of interest, and the detection pose to the structural coordinate system, outputting aligned multimodal samples. The aligned multimodal samples carry a unified traceability identifier.
[0012] The multimodal large model layer, connected to the preprocessing and calibration alignment layer, is used to receive aligned multimodal samples, encode them using an encoder group to obtain a unified embedding representation; it is also used to adaptively adjust the contribution weights of each modality through quality-aware gating fusion based on the calculated modality quality index, and perform cross-modal attention fusion to obtain a fusion representation; it performs initial state prediction based on the fusion representation, and performs uncertainty quantification and out-of-distribution detection on the fusion representation, outputting the initial state prediction result, uncertainty index, and out-of-distribution detection result;
[0013] The mechanism constraint and state inference layer, connected to the multimodal large model layer, receives fused representations, uncertainty indices, and out-of-distribution detection results. It retrieves relevant knowledge fragments from the mechanism model library, standard threshold library, and historical case library. Combining the uncertainty indices and out-of-distribution detection results, it constrains and calibrates the inference output of the fused representation, outputting constraint calibration results. Based on the constraint calibration results, it estimates state quantities and predicts remaining lifetime intervals, outputting crack presence probability, crack size or damage index, health index, remaining lifetime interval, and development trend. It also performs condition attribution and trend analysis, outputting condition attribution results and a key evidence index.
[0014] The interpretable and decision support layer, connected to the mechanism constraint and state inference layer, is used to receive the crack size or damage index, health index, remaining life range and development trend, the operating condition attribution results and key evidence index, as well as the uncertainty index and out-of-distribution detection results; output confidence level, range and out-of-distribution judgment; and trigger re-inspection or manual review strategy based on the out-of-distribution judgment; generate evidence heat map or key fragments, and form reports and audit records; it is also used to integrate the crack size or damage index, health index, remaining life range and development trend, the operating condition attribution results, and the confidence level and out-of-distribution judgment to calculate risk score and map risk level, and output maintenance suggestions or re-inspection suggestions and audit report;
[0015] The closed-loop learning and digital ledger layer, connected to the interpretability and decision support layer, receives the risk level, maintenance or re-inspection recommendations, and audit reports. It obtains the maintenance verification results or re-inspection results corresponding to the maintenance or re-inspection recommendations output by the interpretability and decision support layer as monitoring signals, performs incremental updates, and records the model version, data version, and threshold version. The updated model version is then sent back to the multimodal large model layer, and the updated threshold version is sent back to the mechanism constraint and state inference layer for subsequent evaluation. A digital ledger is constructed by combining the version management and computer maintenance management system interface or the enterprise asset management interface.
[0016] Furthermore, a first data interface is provided between the data acquisition and access layer and the preprocessing and calibration alignment layer. The first data interface is used to transmit the original multimodal data packets and their metadata. The metadata includes at least the sampling rate, timestamp, pose or coordinates, device parameters and operating condition text, and contains a unified traceability identifier.
[0017] A second data interface is provided between the preprocessing and calibration alignment layer and the multimodal large model layer. The second data interface is used to transmit the aligned multimodal samples. The aligned multimodal samples include at least time alignment parameters, spatial calibration matrix or region of interest coordinates, missing completion mask and quality index vector, and contain a unified traceability identifier.
[0018] A third data interface is provided between the multimodal large model layer and the mechanism constraint and state inference layer. The third data interface is used to transmit fusion representation, initial state prediction and uncertainty and out-of-distribution detection results, and contains the unified traceability identifier.
[0019] A fourth data interface is provided between the mechanism constraint and state inference layer and the interpretability and decision support layer. The fourth data interface is used to transmit crack state quantity or damage state quantity, health index, remaining life interval, trend and working condition attribution, and key evidence index, and contains a unified traceability identifier. The key evidence index is used to indicate the corresponding time segment, image frame, ROI coordinate and the matched standard or mechanism entry number, so that the interpretability and decision support layer can generate evidence heat map, key segment and risk classification report.
[0020] A fifth data interface is provided between the explainable and decision support layer and the closed-loop learning and digital ledger layer. The fifth data interface is used to transmit risk level, maintenance or re-inspection suggestions and audit records. The audit records include at least a unified traceability identifier, data version, model version, threshold version, quality indicators and uncertainty description.
[0021] Furthermore, the time alignment parameters include at least a time drift correction amount, which is a scalar, a segmented scalar sequence, or a function parameter that varies with time, and is stored in association with a unified traceability identifier. When generating an evaluation report, the alignment parameters, including the time drift correction amount, are output.
[0022] Furthermore, the spatial calibration matrix or region of interest coordinates transmitted by the second data interface are used to establish the correspondence between signal segments, image frames and weld positions. This correspondence is used for spatial reprojection when the evidence heatmap is generated by the interpretable and decision support layer, as well as for audit tracing.
[0023] Furthermore, the unified traceability identifier also includes detection pose, acquisition channel number and / or operating condition label; the multimodal large model layer includes an inference calibration module, and a feedback interface is provided between the inference calibration module and the mechanism constraint and state inference layer. The feedback interface is used to transmit the constraint verification results and deviation of the mechanism constraint and state inference layer to the inference calibration module; the encoder group includes a signal encoder, an image encoder and a text encoder.
[0024] Furthermore, the closed-loop learning and digital ledger layer sends the updated model version back to the multimodal large model layer, and sends the updated threshold version back to the mechanism constraint and state inference layer.
[0025] A method for intelligent performance evaluation of welded structures based on a multimodal large model includes the following steps:
[0026] S1, Multi-source data acquisition and traceability: Acquire online monitoring signals, non-destructive testing images, visual images or infrared images, and working condition text or historical text to obtain the original multimodal data package and its metadata. Write a unified traceability identifier into the metadata. The unified traceability identifier serves as the primary key for cross-layer data association and includes at least the structural identifier, weld identifier, and acquisition timestamp.
[0027] S2, Preprocessing and Spatiotemporal Calibration Alignment: The online monitoring signals, non-destructive testing images, visual images or infrared images in the original multimodal data packets described in step S1 are denoised, segmented, outlier processed and missing data filled in. The modal quality indicators of signal-to-noise ratio, sharpness and missing rate are calculated, and time synchronization calibration and spatial positioning calibration are performed. The signal segments are mapped to the region of interest of the weld or the region of interest of the weld toe, the detection pose and the structural coordinate system to obtain aligned multimodal samples. The aligned multimodal samples carry a unified traceability identifier.
[0028] S3, Multimodal Large Model Inference: The aligned multimodal samples described in step S2 are encoded using a signal encoder, an image encoder, and a text encoder respectively to obtain a unified embedding representation; based on the calculated modal quality index, the contribution weights of each modality are adaptively adjusted through quality-aware gating fusion, and cross-modal attention fusion is performed to obtain a fused representation; based on the fused representation, initial state prediction is performed, and uncertainty quantification and out-of-distribution detection are performed on the fused representation, outputting the initial state prediction result, uncertainty index, and out-of-distribution detection result;
[0029] S4, Mechanism Constraints and State Inference: Relevant knowledge fragments are retrieved from the mechanism model library, standard threshold library, and historical case library. Combined with the uncertainty index and out-of-distribution detection results described in step S3, the inference output of the fusion representation described in step S3 is constrained and calibrated, and the constraint calibration results are output. Based on the constraint calibration results, state quantity estimation and remaining life interval prediction are performed, and the crack existence probability, crack size or damage index, health index, remaining life interval, and development trend are output. Working condition attribution and trend analysis are performed, and the working condition attribution results and key evidence index are output.
[0030] S5, Explanatory and Decision Support: Outputs confidence level, range, and out-of-distribution judgment, and triggers re-inspection or manual review strategies based on the out-of-distribution judgment; generates evidence heatmaps or key fragments, and forms reports and audit records; integrates the crack size or damage index, health index, remaining life range, development trend, working condition attribution results, and output confidence level and out-of-distribution judgment as described in step S4, calculates risk score and maps risk level, and outputs maintenance suggestions or re-inspection suggestions and audit reports;
[0031] S6, Closed-loop learning and digital ledger: Obtain the maintenance verification results or re-inspection results corresponding to the maintenance suggestions or re-inspection suggestions output in step S5 as a supervision signal, perform incremental updates, and record the model version, data version and threshold version; send back the updated model version and threshold version, and build a digital ledger by combining version management with the computer maintenance management system interface or enterprise asset management interface;
[0032] The steps S1 to S6 above are executed sequentially, with the output of the previous step serving as the input data, parameters, or constraints for the next step.
[0033] Furthermore, in step S2, when the unified traceability identifier is incomplete, the corresponding original multimodal data packet is marked as needing to be reviewed or needing to be re-collected.
[0034] Furthermore, the conditions for triggering the re-inspection or manual review strategy in step S5 include at least one of the following: confidence level is lower than the first threshold, out-of-distribution is determined to be out-of-distribution, risk level is not lower than the second risk level, key modality missing rate exceeds the second threshold, timestamp drift exceeds the preset threshold, and mechanism constraint verification fails.
[0035] Furthermore, the first threshold is 0.75, the second risk level is R2, and the second threshold is 20%.
[0036] Advantages of the present invention
[0037] First, this invention establishes a unified traceability identifier in step S1 and performs time synchronization calibration and spatial positioning calibration in step S2, which solves the technical problems of inconsistent sampling frequencies, timestamp drift, and inconsistent spatial coordinates of multi-source heterogeneous data. It achieves high-precision spatiotemporal alignment and full-process traceability of multimodal data, enabling cross-modal evidence to be consistently mapped and reproduced, and significantly improving the auditability and traceability of the evaluation process.
[0038] Secondly, this invention calculates modal quality indicators such as signal-to-noise ratio, clarity, and missing rate through step S3, and adaptively adjusts the contribution weight of each modality through quality-aware gating fusion based on the quality indicators. This solves the technical problem of unstable evaluation results caused by data quality fluctuations and missing data. Even when data is missing, there is noise interference, or the detection conditions change, the evaluation conclusion can still be output stably. This significantly reduces the risk of false alarms and missed alarms compared to traditional single-modality or fixed-weight fusion schemes.
[0039] Third, this invention retrieves relevant knowledge fragments from the mechanism model library, standard threshold library, and historical case library in step S4, and performs physical boundary verification, threshold consistency verification, and mechanism parameter correction on the inference output. This solves the technical problem of the lack of engineering consistency and interpretability of pure data-driven models, making the evaluation results conform to engineering physical laws. At the same time, it outputs key evidence indexes, providing traceable basis for crack existence probability, health index, and remaining life range, significantly improving the interpretability and engineering credibility of the evaluation results.
[0040] Fourth, this invention quantifies uncertainty and detects out-of-distribution distributions in the fusion representation through step S3, and automatically triggers a re-inspection or manual review strategy in step S5 based on multi-dimensional conditions such as confidence level below the threshold, out-of-distribution determination, excessive key mode missing rate, and failure of mechanism constraint verification. This solves the technical problem of decision-making risks caused by the lack of confidence representation in the evaluation results and false alarms and omissions. By outputting confidence level, interval, and risk level, it provides a quantitative risk basis for maintenance and re-inspection decisions, effectively reducing decision-making risks.
[0041] Fifth, this invention uses step S6 to feed back the maintenance verification results or re-inspection results as a monitoring signal, performs incremental updates and records the model version, data version and threshold version, and sends the updated version back for subsequent evaluation. This solves the technical problem that the evaluation model cannot adapt to the dynamic changes of the welded structure during long-term service, and forms a closed-loop mechanism of "evaluation-decision-verification-optimization". This allows the model to be continuously calibrated and optimized as the distribution of service data changes, and has good long-term adaptability.
[0042] The aforementioned beneficial effects have been verified by field tests on welded joints of pressure vessels and can be extended to scenarios such as vehicle welded structures. This invention, through the synergistic effect of core technologies such as spatiotemporal alignment, quality perception fusion, mechanism constraint reasoning, and uncertainty quantification, ultimately provides accurate crack / damage status, health index, and remaining life range for the performance evaluation of welded structures of equipment such as bridges, ships, pressure vessels, engineering machinery, and automobiles. It also simultaneously provides uncertainty, out-of-distribution judgments, and interpretable evidence. Through a closed-loop update mechanism, it continuously optimizes and provides full life-cycle technical support for maintenance decisions and life management. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the intelligent performance evaluation system for welded structures based on a multimodal large model, as described in this invention.
[0044] Figure 2 This is a flowchart of the intelligent performance evaluation method for welded structures based on a multimodal large model, as described in this invention.
[0045] Figure 3 This is a flowchart illustrating the multimodal large model reasoning and mechanism constraint state inference in this invention (corresponding to...). Figure 2 Steps S3 and S4 in the process.
[0046] Figure 4 This is a schematic diagram of the process of explaining the digital ledger of decision support and closed-loop learning in this invention (corresponding to...). Figure 2 Steps S5 and S6 in the process. Detailed Implementation
[0047] The present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. It should be noted that the specific embodiments are not intended to limit the scope of the present invention.
[0048] like Figure 1 As shown in the figure, the intelligent performance evaluation system for welded structures based on a multimodal large model provided in this specific embodiment includes a data acquisition and access layer, a preprocessing and calibration alignment layer, a multimodal large model layer, and a mechanism constraint and state inference layer.
[0049] The system comprises an interpretable and decision support layer, a closed-loop learning and digital ledger layer, and layers that work together to achieve end-to-end processing and intelligent evaluation of multimodal data.
[0050] The data acquisition and access layer is used to acquire online monitoring signals, non-destructive testing images, visual images or infrared images, and operating condition text or historical text, and outputs raw multimodal data packets and their metadata. A unified traceability identifier is written into the metadata. To ensure the traceability and reproducibility of multimodal evidence and support closed-loop updates, the unified traceability identifier is used as the primary key for cross-layer data association. The unified traceability identifier includes a structure identifier (referred to as: structure ID), a weld identifier (referred to as: weld ID), and an acquisition timestamp; it also includes the detection pose, acquisition channel number, and / or operating condition label. The online monitoring signals include at least one or more of strain, acoustic emission (AE), and temperature signals; the non-destructive testing images include at least one or more of offline acquired ultrasonic testing (UT) and offline acquired radiographic testing (RT) images.
[0051] The system establishes a logical correspondence between data objects based on a unified traceability identifier between its layers. Combined with the alignment parameters output by the alignment layer, it achieves accurate registration of multi-source data in spatial and temporal dimensions, thereby ensuring the continuous transmission, associated storage, and closed-loop feedback of raw data, feature vectors, and evaluation results between layers.
[0052] A first data interface is provided between the data acquisition and access layer and the preprocessing and calibration alignment layer. The first data interface is used to transmit the original multimodal data packets and their metadata. The metadata includes at least the sampling rate, timestamp, pose or coordinates, device parameters and operating condition text, and contains a unified traceability identifier.
[0053] The preprocessing and calibration alignment layer, connected to the data acquisition and access layer, is used to receive the original multimodal data packets and their metadata. It performs denoising, segmentation, outlier processing, and missing value completion on the online monitoring signals, non-destructive testing images, visual images, or infrared images in the original multimodal data packets, and conducts quality assessment, calculating modal quality indicators such as signal-to-noise ratio, sharpness, and missing rate. It is also used for time synchronization calibration and spatial positioning calibration, ROI mapping, mapping signal segments to the weld seam region of interest (ROI) or weld toe region of interest (ROI), detection pose, and structural coordinate system, outputting aligned multimodal samples. The aligned multimodal samples carry a unified traceability identifier.
[0054] A multimodal large model layer, connected to the preprocessing and calibration alignment layer, is used to receive aligned multimodal samples. For example... Figure 3 As shown, after receiving aligned multimodal samples, the multimodal large model layer encodes them using an encoder group, which includes a signal encoder, an image encoder, and a text encoder, to obtain a unified embedding representation. Based on the modal quality index, the contribution weights of each modality are adaptively adjusted through quality-aware gating fusion, and cross-modal attention fusion is performed to obtain a fused representation. Further, initial state prediction is performed based on the fused representation, and uncertainty quantification and out-of-distribution detection are performed on the fused representation, outputting the initial state prediction result, uncertainty index, and out-of-distribution detection result. Step S4 receives the fused representation, uncertainty index, and out-of-distribution detection result, retrieves relevant knowledge fragments from the mechanism model library, standard threshold library, and historical case library, i.e., performs knowledge retrieval, performs physical boundary verification, threshold consistency verification, and mechanism parameter correction on the inference output, and then completes state quantity estimation, remaining lifetime interval prediction, and operating condition attribution and trend analysis.
[0055] A second data interface is provided between the preprocessing and calibration alignment layer and the multimodal large model layer. This second data interface is used to transmit aligned multimodal samples. The aligned multimodal samples include at least temporal alignment parameters, a spatial calibration matrix or region of interest (ROI) coordinates, a missing completion mask, and a quality index vector, and contain a unified traceability identifier. The temporal alignment parameters include at least a temporal drift correction, which can be a scalar, a segmented scalar sequence, or a time-varying function parameter, and is stored in association with the unified traceability identifier. When generating the evaluation report, the alignment parameters, including the temporal drift correction, are output. The spatial calibration matrix or ROI coordinates transmitted by the second data interface are used to establish the correspondence between signal segments, image frames, and weld positions. This correspondence is used for spatial re-projection and audit traceability when the interpretable and decision support layer generates evidence heatmaps. In quality-aware gated fusion, the multimodal large model layer uses the quality index vector carried by the aligned multimodal samples to calculate the modal weights.
[0056] The mechanism constraint and state inference layer, connected to the multimodal large model layer, receives fused representations, uncertainty indices, and out-of-distribution detection results. It retrieves relevant knowledge fragments from the mechanism model library, standard threshold library, and historical case library (i.e., performs knowledge retrieval). Combining the uncertainty indices and out-of-distribution detection results, it constrains and calibrates the inference output of the fused representation. This constraint and calibration includes physical boundary verification, threshold consistency verification, and mechanism parameter correction, thus achieving fatigue / fracture mechanism constraints and outputting constraint calibration results. Based on the constraint calibration results, it estimates state quantities and predicts remaining life intervals, outputting crack presence probability, crack size or damage index, health index, remaining life interval, and development trend. It also performs condition attribution and trend analysis, outputting condition attribution results and a key evidence index.
[0057] A third data interface is provided between the multimodal large model layer and the mechanism constraint and state inference layer. The third data interface is used to transmit fusion representation, initial state prediction results and uncertainty and out-of-distribution detection results, and contains the unified traceability identifier.
[0058] The multimodal large model layer includes an inference calibration module. A feedback interface is provided between the inference calibration module and the mechanism constraint and state inference layer. The feedback interface is used to transmit the constraint verification results and deviation of the mechanism constraint and state inference layer to the inference calibration module for consistency correction.
[0059] The interpretable and decision support layer, connected to the mechanism constraint and state inference layer, is used to receive the crack size or damage index, health index, remaining life range and development trend, the working condition attribution results and key evidence index, as well as the uncertainty index and out-of-distribution detection results. It outputs confidence level, range and out-of-distribution judgment, and triggers re-inspection or manual review strategy based on the out-of-distribution judgment; generates evidence heat map or key fragments, and forms reports and audit records; it is also used to integrate the crack size or damage index, health index, remaining life range and development trend, the working condition attribution results, and the confidence level and out-of-distribution judgment to calculate risk score and map risk level, that is, to perform risk classification and threshold judgment, and output maintenance suggestions or re-inspection suggestions and audit reports.
[0060] A fourth data interface is provided between the mechanism constraint and state inference layer and the interpretability and decision support layer. The fourth data interface is used to transmit crack state quantity or damage state quantity, health index, remaining life interval, trend and working condition attribution, and key evidence index, and contains a unified traceability identifier. The key evidence index is used to indicate the corresponding time segment, image frame, ROI coordinate and the matched standard or mechanism entry number, so that the interpretability and decision support layer can generate evidence heat map, key segment and risk classification report.
[0061] The closed-loop learning and digital ledger layer, connected to the interpretability and decision support layer, receives the risk level, maintenance or re-inspection recommendations, and audit reports. It obtains the maintenance verification results or re-inspection results corresponding to the maintenance or re-inspection recommendations output by the interpretability and decision support layer as monitoring signals, performs incremental updates and incremental learning, and records model versions, data versions, and threshold versions to achieve model / data / threshold version management. The closed-loop learning and digital ledger layer transmits the updated model version to the multimodal large model layer and transmits the threshold version back to the mechanism constraint and state inference layer for subsequent evaluation. This layer also has external interfaces, including computer maintenance management system interfaces or enterprise asset management interfaces, i.e., CMMS interfaces or EAM interfaces, which, combined with version management, construct a digital ledger.
[0062] A fifth data interface is provided between the explainable and decision support layer and the closed-loop learning and digital ledger layer. The fifth data interface is used to transmit risk level, maintenance or re-inspection suggestions and audit records. The audit records include at least a unified traceability identifier, data version, model version, threshold version, quality indicators and uncertainty descriptions, and are written into the digital ledger. They can also be synchronized to the computer maintenance management system interface or the enterprise asset management interface.
[0063] The closed-loop learning and digital ledger layer transmits the updated model version to the multimodal large model layer, and transmits the updated threshold version to the mechanism constraint and state inference layer.
[0064] The data acquisition and access layer, preprocessing and calibration alignment layer, multimodal large model layer, mechanism constraint and state inference layer, interpretability and decision support layer, and closed-loop learning and digital ledger layer are connected sequentially to form a data processing serial closed loop. The implementation methods, parameter settings, and model structures of the above layers are adjusted according to different welding structure types and application scenarios. As long as multimodal spatiotemporal alignment, quality perception fusion, mechanism constraint inference, and uncertainty quantification are achieved, they are all within the protection scope of this invention.
[0065] like Figure 2 As shown, a method for intelligent performance evaluation of welded structures based on a multimodal large model using the above system includes the following steps:
[0066] S1, Multi-source data acquisition and traceability: Acquire online monitoring signals, non-destructive testing images, visual images or infrared images, and working condition text or historical text to obtain the original multimodal data package and its metadata. Write a unified traceability identifier into the metadata. The unified traceability identifier serves as the primary key for cross-layer data association and includes at least the structural identifier, weld identifier, and acquisition timestamp.
[0067] S2, Preprocessing and Spatiotemporal Calibration Alignment: Receive the unified traceability identifier and metadata mentioned in step S1 to segment, calculate quality indicators and perform missing statistics on multimodal data; when the unified traceability identifier is incomplete, mark the corresponding original multimodal data packet as needing to be reviewed or supplemented.
[0068] The online monitoring signals, non-destructive testing images, visual images, or infrared images in the original multimodal data packets described in step S1 are denoised, segmented, outlier-handled, and missing data filled in. The modal quality indices of signal-to-noise ratio, sharpness, and missing rate are calculated, and the quality index vector and missing mask are output.
[0069] Time synchronization calibration and spatial positioning calibration are performed. Time synchronization calibration is used to handle situations where the sampling frequency of multi-source sensor data is inconsistent or timestamp drift exists. Based on a unified time reference, the data of each channel is resampled and aligned, and time alignment parameters are output. The time alignment parameters include at least the timestamp drift correction amount Δt; where Δt can be a scalar, a segmented scalar sequence, or a function parameter that varies with time, and is stored in association with a unified traceability identifier; the alignment parameters, including Δt, are output when generating the evaluation report to support result reproducibility.
[0070] Spatial positioning and calibration is used to map the detection pose or weld coordinates to the structural coordinate system, establish the correspondence between signal segments, image frames and the region of interest (ROI) of the weld or weld toe, and between the detection pose and the structural coordinate system, and output the ROI coordinates or weld mileage coordinates. When alignment fails, it is marked as requiring review.
[0071] The aligned multimodal samples are obtained, which carry a unified traceability identifier and include time alignment parameters, spatial calibration matrix or region of interest coordinates, missing completion mask and quality index vector.
[0072] Time synchronization parameters and spatial positioning calibration or ROI mapping results are used to determine the one-to-one correspondence between signal segments, image frames and weld positions during subsequent encoding; the correspondence is also used for spatial reprojection when generating evidence heatmaps in the interpretable and decision support layer and for audit tracing.
[0073] The quality index vector and the missing mask are used as the basis for calculating modal weights in subsequent gating fusion. When the quality of a certain modality is substandard or missing, the weight of that modality is reduced and the priority of re-inspection is increased in subsequent steps.
[0074] S3, Multimodal Large Model Inference: Receive the aligned multimodal sample passed from step S2. The aligned multimodal sample carries a unified traceability identifier and contains a quality index vector. Encode the aligned multimodal sample from step S2 using a signal encoder, an image encoder, and a text encoder respectively to obtain a unified embedding representation.
[0075] Based on the modal quality indicators such as signal-to-noise ratio, clarity, and missing rate in the quality indicator vector, the contribution weights of each modality are adaptively adjusted through quality-aware gating fusion, and cross-modal attention fusion is performed to obtain the fusion representation; the quality indicators serve as the basis for generating fusion weights and are retained in the audit records.
[0076] Initial state prediction is performed based on fusion representation, and uncertainty quantification and out-of-distribution detection are performed on the fusion representation. The output includes initial state prediction results, uncertainty index and out-of-distribution detection results, including confidence level, interval and out-of-distribution determination.
[0077] In step S3, the modal embedding representations and fusion representations are output; the modal embedding representations, as inputs to cross-modal attention fusion, have already been completed in step 3; the fusion representations serve as the common inputs to retrieval enhancement and mechanistic constraint reasoning in mechanistic constraint and state inference in step S4.
[0078] S4, Mechanism Constraints and State Inference: Receive the fusion representation, uncertainty index, and out-of-distribution detection results output from step S3. Retrieve relevant knowledge fragments from the mechanism model library, standard threshold library, and historical case library. Combine the uncertainty index and out-of-distribution detection results described in step S3 to constrain and calibrate the inference output of the fusion representation described in step S3. The constraint and calibration process includes at least physical boundary verification, threshold consistency verification, and mechanism parameter correction. Output the constraint calibration results, which are used to correct the state variables and remaining lifetime range. When the constraint verification fails, the reason is written as evidence into the audit report, triggering a re-inspection or manual review.
[0079] Based on the constraint calibration results, state variables are estimated and remaining life intervals are predicted. The output includes crack existence probability, crack size or damage index, health index, remaining life interval and development trend. The working condition attribution and trend analysis are performed, and the working condition attribution results and key evidence index are output.
[0080] S5, Explanatory and Decision Support: Receives the uncertainty index and out-of-distribution detection results output from step S3, including confidence level, interval and out-of-distribution judgment, and triggers a re-inspection or manual review strategy based on the out-of-distribution judgment.
[0081] The system receives the crack presence probability, crack size, health index, remaining lifespan range, development trend, operating condition attribution results, and key evidence index output from step S4. The crack presence probability, crack size, health index, remaining lifespan range, and development trend output from step S4, along with the uncertainty and out-of-distribution detection results output from step S3, are used together for risk classification and strategy output.
[0082] The risk score is calculated by comprehensively considering severity, trend, operating conditions, and uncertainty, and then mapped to a risk level. The output includes a risk level, maintenance or re-inspection recommendations, and an audit report. Uncertainty is used to adjust the risk score and determine the re-inspection or manual review path.
[0083] Simultaneously, the uncertainty index and out-of-distribution detection results output from step S3, along with the constraint calibration results output from step S4, are received to comprehensively determine whether to trigger a re-inspection or manual review strategy. Triggering conditions include any of the following: confidence level below the first threshold; out-of-distribution determination as out-of-distribution; risk level not lower than the second risk level; key modality missing rate exceeding the second threshold; timestamp drift exceeding a preset threshold; mechanism constraint verification failing. The first threshold is 0.75, the second risk level is R², and the second threshold is 20%.
[0084] Generate evidence heatmaps or key fragments, and form reports and audit records. The evidence heatmaps are spatially retrospectively analyzed based on the spatial correspondence established in step S2.
[0085] like Figure 4 As shown, step S5 receives the uncertainty index and out-of-distribution detection results output by step S3, and the crack presence probability, crack size or damage index, health index, remaining life range, development trend, operating condition attribution results, and key evidence index output by step S4. Step S5 calculates the risk score and maps the risk level, outputting a risk score S based on severity, trend, operating condition, and uncertainty, and generates maintenance or re-inspection recommendations. Maintenance recommendations include welding repair, reinforcement, or replacement, while re-inspection recommendations include NDT type or cycle. Step S5 also generates an evidence chain or heat map and outputs the risk level, maintenance or re-inspection recommendations, and audit records to step S6. Step S6 receives the maintenance verification results or re-inspection results feedback, performs incremental updates, records the model version, data version, and threshold version, and sends the updated model version back to step S3 and the updated threshold version back to step S4 to achieve continuous optimization of the model and thresholds and full-process audit traceability.
[0086] The output of step S5 includes at least the risk level, maintenance or re-inspection recommendations, audit log, chain of evidence or heat map, and the probability of crack presence, crack size or damage index, health index, remaining life range, development trend, operating condition attribution results, key evidence index, confidence level and out-of-distribution determination used to form the audit log.
[0087] S6, Closed-loop learning and digital ledger: Obtain the maintenance verification results or re-inspection results corresponding to the maintenance suggestions or re-inspection suggestions output in step S5 as a supervision signal, perform incremental updates, and record the model version, data version, and threshold version; In the maintenance suggestions or re-inspection suggestions output in step S5 and the audit records, the audit records and the subsequently obtained maintenance verification results or re-inspection results together serve as the supervision signal for step S6.
[0088] The updated model version is sent to step S3, and the threshold version is sent back to step S4, so that the next round of evaluation can be repeated and compared under the same traceability identification system.
[0089] Digital ledgers can be built by combining version management with computer maintenance management system interfaces or enterprise asset management interfaces.
[0090] The steps S1 to S6 are executed sequentially, forming a closed-loop process. The output of the previous step serves as the input data, parameters, or constraints for the next step, making the evaluation process closed at the data level and traceable.
[0091] The encoder type in this embodiment can be replaced according to the data characteristics. For example, a signal encoder can use GRU, and an image encoder can use ResNet. As long as effective features can be extracted and a unified embedding representation can be generated, it falls within the protection scope of this invention. The content of the knowledge base can be expanded according to the application scenario, adding industry standards, new mechanism models, etc., without affecting the core technical solution. The weight coefficients and thresholds of risk classification can be flexibly adjusted according to the importance of the equipment and the usage scenario to adapt to different security requirements. The update frequency of closed-loop learning and the sample size threshold of incremental learning can be adjusted according to the actual data accumulation, with the core requirement of continuously optimizing the model. Those skilled in the art can split or merge the module functions without deviating from the core technology of this invention. All modifications based on the technical solution of this invention fall within the protection scope of this invention.
[0092] This invention is applicable to welded structures in bridges, ships, engineering machinery, and automobiles, achieving consistent technical results. Below is an application example of using the above system and method for intelligent assessment of the quality status of welded joints in pressure vessels. This application example is merely illustrative and does not constitute a limitation. The specific example is as follows:
[0093] 1. Multi-source acquisition and traceability: This involves acquiring online monitoring signals from the welded joints of the pressure vessel, ultrasonic non-destructive testing images, infrared thermal imaging images, video recordings of the processing, and operational condition text. The online monitoring signals include strain signals acquired by strain sensors, temperature signals acquired by temperature sensors, and acoustic emission signals acquired by acoustic emission sensors. The operational condition text includes operating pressure, medium type, operating time, and historical maintenance records. A unified traceability identifier is established for the acquired data, for example: “RQ-001-Weld-05-Pose-3-CH-2-202406151430-Load-1.2MPa”, used to bind all associated data corresponding to this inspection.
[0094] 2. Preprocessing and quality assessment: Wavelet denoising was applied to the acoustic emission signal, detrending was applied to the strain signal, and grayscale enhancement and denoising were performed on the ultrasonic image. The calculated signal-to-noise ratio was 18dB, the image sharpness was 0.85, and the text missing rate was 8%. It was determined that all quality indicators of this group met the fusion quality requirements.
[0095] 3. Spatiotemporal calibration and alignment: Multimodal data time synchronization is achieved through timestamp calibration. Based on the three-dimensional model of the pressure vessel, the mapping of ultrasonic images, infrared images and physical coordinate system is completed to locate the ROI area of the weld toe of the weld joint.
[0096] 4. Representation learning encoding: The signal encoder uses TCN to extract frequency domain features, the image encoder uses ViT to extract defect texture features, and the text encoder uses BERT to extract working condition semantic features, generating a 384-dimensional unified embedding vector.
[0097] 5. Quality-aware gated fusion: The gated network assigns weights according to quality indicators, with signal weight of 0.4, image weight of 0.4, and text weight of 0.2. A unified fusion representation is obtained through cross-modal attention fusion.
[0098] 6. Retrieve relevant knowledge fragments and perform mechanism constraint reasoning: Retrieve the crack propagation model and standard threshold for this type of welded joint. The crack propagation model is the Paris equation, and the standard threshold is that the crack length is ≥3mm and repair is required. Apply monotonicity constraints and calibrate the reasoning results.
[0099] 7. Condition and lifespan estimation: The probability of output crack presence is 92%, the crack length is 3.2mm, the health index is 0.35, the remaining lifespan range is 12-18 months, and the damage shows a slow growth trend.
[0100] 8. Judgment of uncertainty and out-of-distribution results: The confidence level is 0.82, which indicates that the sample is within the distribution. No emergency re-examination is required. It is recommended to repeat the ultrasound image after 3 months.
[0101] 9. Risk Classification and Decision Output: The overall risk score is calculated to be 75 points, which is mapped to medium risk. The output is the suggestion to "develop a maintenance plan within 1 month and shorten the monitoring cycle to 15 days / time". An audit report containing evidence heatmap, standard references and uncertainty analysis is generated.
[0102] 10. Closed-loop learning and updating: Repair is completed after 1 month, the crack length is confirmed to be 3.3mm, the repair result is fed back as a supervision signal, the system performs incremental updates, and the model version is recorded as V1.1.
[0103] 11. Actual test results show that the evaluation results of this embodiment deviate from the actual damage state by ≤5%, the risk classification is accurate, the maintenance suggestions are operable, the efficiency is improved by 60% compared with the traditional evaluation method, and the risk of misjudgment is reduced by 70%.
[0104] The method and system of this invention are not limited to the application of the aforementioned pressure vessel welded joints, but are also applicable to the performance evaluation of welded structures in bridges, ships, engineering machinery, and automobiles. Depending on the characteristics of different application scenarios, the quality index thresholds, risk level classification standards, and maintenance recommendation types can be adaptively adjusted.
Claims
1. A smart performance evaluation system for welded structures based on a multimodal large model, characterized in that, include: The data acquisition and access layer is used to acquire online monitoring signals, non-destructive testing images, visual images or infrared images, as well as working condition text or historical text, and output raw multimodal data packets and their metadata. A unified traceability identifier is written into the metadata. The unified traceability identifier serves as the primary key for cross-layer data association and includes at least a structure identifier, a weld identifier, and an acquisition timestamp. The preprocessing and calibration alignment layer, connected to the data acquisition and access layer, is used to receive the original multimodal data packets and their metadata. It performs denoising, segmentation, outlier processing, and missing value completion on the online monitoring signals, non-destructive testing images, visual images, or infrared images in the original multimodal data packets, and conducts quality assessment, calculating modal quality indicators such as signal-to-noise ratio, sharpness, and missing rate. It is also used for time synchronization calibration and spatial positioning calibration, ROI mapping, mapping signal segments to the weld seam region of interest or weld toe region of interest, and the detection pose to the structural coordinate system, outputting aligned multimodal samples. The aligned multimodal samples carry a unified traceability identifier. The multimodal large model layer, connected to the preprocessing and calibration alignment layer, is used to receive aligned multimodal samples, which are then encoded using an encoder group to obtain a unified embedded representation. It is also used to adaptively adjust the contribution weights of each modality through quality-aware gating fusion based on the calculated modal quality index, and to perform cross-modal attention fusion to obtain a fusion representation; Initial state prediction is performed based on fusion representation, and uncertainty quantification and out-of-distribution detection are performed on the fusion representation. The initial state prediction results, uncertainty index and out-of-distribution detection results are output. The mechanism constraint and state inference layer, connected to the multimodal large model layer, receives fused representations, uncertainty indices, and out-of-distribution detection results. It retrieves relevant knowledge fragments from the mechanism model library, standard threshold library, and historical case library. Combining the uncertainty indices and out-of-distribution detection results, it constrains and calibrates the inference output of the fused representation, outputting constraint calibration results. Based on the constraint calibration results, it estimates state quantities and predicts remaining lifetime intervals, outputting crack presence probability, crack size or damage index, health index, remaining lifetime interval, and development trend. It also performs condition attribution and trend analysis, outputting condition attribution results and a key evidence index. The interpretable and decision support layer, connected to the mechanism constraint and state inference layer, is used to receive the crack size or damage index, health index, remaining life range and development trend, the operating condition attribution results and key evidence index, as well as the uncertainty index and out-of-distribution detection results; output confidence level, range and out-of-distribution judgment; and trigger re-inspection or manual review strategy based on the out-of-distribution judgment; generate evidence heat map or key fragments, and form reports and audit records; it is also used to integrate the crack size or damage index, health index, remaining life range and development trend, the operating condition attribution results, and the confidence level and out-of-distribution judgment to calculate risk score and map risk level, and output maintenance suggestions or re-inspection suggestions and audit report; The closed-loop learning and digital ledger layer, connected to the interpretability and decision support layer, receives the risk level, maintenance or re-inspection recommendations, and audit reports. It obtains the maintenance verification results or re-inspection results corresponding to the maintenance or re-inspection recommendations output by the interpretability and decision support layer as monitoring signals, performs incremental updates, and records the model version, data version, and threshold version. The updated model version is then sent back to the multimodal large model layer, and the updated threshold version is sent back to the mechanism constraint and state inference layer for subsequent evaluation. A digital ledger is constructed by combining the version management and computer maintenance management system interface or the enterprise asset management interface.
2. The system according to claim 1, characterized in that, A first data interface is provided between the data acquisition and access layer and the preprocessing and calibration alignment layer. The first data interface is used to transmit the original multimodal data packets and their metadata. The metadata includes at least the sampling rate, timestamp, pose or coordinates, device parameters and operating condition text, and contains a unified traceability identifier. A second data interface is provided between the preprocessing and calibration alignment layer and the multimodal large model layer. The second data interface is used to transmit the aligned multimodal samples. The aligned multimodal samples include at least time alignment parameters, spatial calibration matrix or region of interest coordinates, missing completion mask and quality index vector, and contain a unified traceability identifier. A third data interface is provided between the multimodal large model layer and the mechanism constraint and state inference layer. The third data interface is used to transmit fusion representation, initial state prediction results and uncertainty and out-of-distribution detection results, and contains the unified traceability identifier. A fourth data interface is provided between the mechanism constraint and state inference layer and the interpretability and decision support layer. The fourth data interface is used to transmit crack state quantity or damage state quantity, health index, remaining life interval, trend and working condition attribution, and key evidence index, and contains a unified traceability identifier. The key evidence index is used to indicate the corresponding time segment, image frame, ROI coordinate and the matched standard or mechanism entry number, so that the interpretability and decision support layer can generate evidence heat map, key segment and risk classification report. A fifth data interface is provided between the explainable and decision support layer and the closed-loop learning and digital ledger layer. The fifth data interface is used to transmit risk level, maintenance or re-inspection suggestions and audit records. The audit records include at least a unified traceability identifier, data version, model version, threshold version, quality indicators and uncertainty description.
3. The system according to claim 2, characterized in that, The time alignment parameters include at least a time drift correction amount, which can be a scalar, a segmented scalar sequence, or a function parameter that varies with time. It is stored in association with a unified traceability identifier, and the alignment parameters, including the time drift correction amount, are output when the evaluation report is generated.
4. The system according to claim 2, characterized in that, The spatial calibration matrix or region of interest coordinates transmitted by the second data interface are used to establish the correspondence between signal segments, image frames and weld positions. The correspondence is used for spatial reprojection when the evidence heatmap is generated by the interpretable and decision support layer, as well as for audit tracing.
5. The system according to claim 1, characterized in that, The unified traceability identifier also includes detection pose, acquisition channel number and / or operating condition label; the multimodal large model layer includes an inference calibration module, and a feedback interface is provided between the inference calibration module and the mechanism constraint and state inference layer. The feedback interface is used to transmit the constraint verification results and deviation of the mechanism constraint and state inference layer to the inference calibration module; the encoder group includes a signal encoder, an image encoder and a text encoder.
6. The system according to claim 1, characterized in that, The closed-loop learning and digital ledger layer transmits the updated model version to the multimodal large model layer, and transmits the updated threshold version to the mechanism constraint and state inference layer.
7. A method for intelligent performance evaluation of welded structures based on a multimodal large model, characterized in that, Includes the following steps: S1, Multi-source data acquisition and traceability: Acquire online monitoring signals, non-destructive testing images, visual images or infrared images, and working condition text or historical text to obtain the original multimodal data package and its metadata. Write a unified traceability identifier into the metadata. The unified traceability identifier serves as the primary key for cross-layer data association and includes at least the structural identifier, weld identifier, and acquisition timestamp. S2, Preprocessing and Spatiotemporal Calibration Alignment: The online monitoring signals, non-destructive testing images, visual images or infrared images in the original multimodal data packets described in step S1 are denoised, segmented, outlier processed and missing data filled in. The modal quality indicators of signal-to-noise ratio, sharpness and missing rate are calculated, and time synchronization calibration and spatial positioning calibration are performed. The signal segments are mapped to the region of interest of the weld or the region of interest of the weld toe, the detection pose and the structural coordinate system to obtain aligned multimodal samples. The aligned multimodal samples carry a unified traceability identifier. S3, Multimodal large model inference: The aligned multimodal samples described in step S2 are encoded using a signal encoder, an image encoder, and a text encoder respectively to obtain a unified embedding representation; Based on the calculated modal quality index, the contribution weights of each modality are adaptively adjusted through quality-aware gating fusion, and cross-modal attention fusion is performed to obtain a fused representation; Initial state prediction is performed based on fusion representation, uncertainty quantification and out-of-distribution detection are performed on the fusion representation, and the initial state prediction results, uncertainty index and out-of-distribution detection results are output. S4, Mechanism Constraints and State Inference: Relevant knowledge fragments are retrieved from the mechanism model library, standard threshold library, and historical case library. Combined with the uncertainty index and out-of-distribution detection results described in step S3, the inference output of the fusion representation described in step S3 is constrained and calibrated, and the constraint calibration results are output. Based on the constraint calibration results, state quantity estimation and remaining life interval prediction are performed, and the crack existence probability, crack size or damage index, health index, remaining life interval, and development trend are output. Working condition attribution and trend analysis are performed, and the working condition attribution results and key evidence index are output. S5, Explanatory and Decision Support: Outputs confidence level, range, and out-of-distribution judgment, and triggers re-inspection or manual review strategies based on the out-of-distribution judgment; generates evidence heatmaps or key fragments, and forms reports and audit records; integrates the crack size or damage index, health index, remaining life range, development trend, working condition attribution results, and output confidence level and out-of-distribution judgment as described in step S4, calculates risk score and maps risk level, and outputs maintenance suggestions or re-inspection suggestions and audit reports; S6, Closed-loop learning and digital ledger: Obtain the maintenance verification results or re-inspection results corresponding to the maintenance suggestions or re-inspection suggestions output in step S5 as a supervision signal, perform incremental updates, and record the model version, data version, and threshold version; send the updated model version back to the multimodal large model layer corresponding to step S3, send the updated threshold version back to the mechanism constraint and state inference layer corresponding to step S4, and build a digital ledger by combining version management and computer maintenance management system interface or enterprise asset management interface; The steps S1 to S6 above are executed sequentially, with the output of the previous step serving as the input data, parameters, or constraints for the next step.
8. The method according to claim 7, characterized in that, In step S2, when the unified traceability identifier is incomplete, the corresponding original multimodal data packet is marked as needing to be reviewed or needing to be supplemented.
9. The method according to claim 7, characterized in that, The conditions for triggering the re-inspection or manual review strategy in step S5 include at least one of the following: confidence level is lower than the first threshold, out-of-distribution is determined to be out-of-distribution, risk level is not lower than the second risk level, key modality missing rate exceeds the second threshold, timestamp drift exceeds the preset threshold, and mechanism constraint verification fails.
10. The method according to claim 9, characterized in that, The first threshold is 0.75, the second risk level is R2, and the second threshold is 20%.
Citation Information
Cited By
Intelligent inspection method and system for hoisting machinery based on multi-modal large model
CN122133529A
An automobile welding part position recognition method based on deep learning
CN122223022A