Intelligent inspection method and system for hoisting machinery based on multi-modal large model

By combining a multimodal large model with a numerical simulation program for fatigue damage evolution in fracture mechanics, the problem of accuracy in identifying latent fatigue damage under varying working conditions of lifting machinery was solved, and effective identification and early warning of early latent fatigue damage were achieved.

CN122133529AActive Publication Date: 2026-06-02ANHUI SPECIAL EQUIP INSPECTION INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI SPECIAL EQUIP INSPECTION INST
Filing Date
2026-05-07
Publication Date
2026-06-02

Smart Images

  • Figure CN122133529A_ABST
    Figure CN122133529A_ABST
Patent Text Reader

Abstract

This invention relates to the field of crane machinery safety inspection technology, and discloses an intelligent inspection method and system for crane machinery based on a multimodal large model. The method simultaneously collects multi-source data including acoustic emission, 3D vision, and strain, preprocesses and desensitizes the data at the edge, and then uploads it to the cloud. The cloud-based large model exhibits spatiotemporal decoupling features and extracts damage-sensitive features. Combined with fracture mechanics simulation, it generates damage state and life curves, and optimizes model weights through causal reasoning to form a closed loop. The model is then incrementally updated through federated aggregation, and finally, a judgment result and graded early warning are generated by comparing strength thresholds. The system integrates data and mechanism-driven approaches to solve the problem of missed detection of hidden defects under varying working conditions, improves inspection accuracy and stability, and provides reliable support for the safe operation and maintenance of crane machinery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety inspection technology for lifting machinery, and discloses an intelligent inspection method and system for lifting machinery based on a multimodal large model. Background Technology

[0002] Currently, the inspection of lifting machinery often employs a combination of multi-source heterogeneous sensing and centralized data processing. Specifically, acoustic emission sensors, vision cameras, and strain sensors are deployed on the metal structure and transmission components of the lifting machinery to simultaneously collect inspection data, which is then directly aggregated to a cloud server. The cloud server uses a multimodal deep learning model to extract features from the multi-source data. Typically, data from different modalities are concatenated or weighted at the feature layer to generate a comprehensive feature vector. This comprehensive feature vector is input into a classifier or inspection head, outputting the identification results of surface defects such as cracks and deformations. A manual inspection report is then generated based on these identification results. Simultaneously, the cloud server uses the aggregated raw data from all equipment to centrally train and update the model, directly distributing the updated model parameters to each piece of equipment.

[0003] The existing technical solutions described above have limitations in handling early latent fatigue damage in cranes operating under varying conditions. Because multimodal deep learning models only stitch features together from multiple data sources without incorporating fracture mechanics damage evolution mechanisms for constraint, the models cannot distinguish between feature changes caused by actual fatigue damage and pseudo-feature changes caused by operating condition disturbances and environmental noise. When cranes are operating under varying conditions, pseudo-features generated by operating condition disturbances dominate the comprehensive feature vector, causing the model's output to only reflect obvious surface defects and failing to extract feature vectors sensitive to early latent fatigue damage, resulting in missed detection of latent defects and delayed inspection results. Summary of the Invention

[0004] The purpose of this invention is to provide a solution to the problems described in the background section.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: Intelligent inspection methods for lifting machinery based on multimodal large models include: Simultaneously acquire acoustic emission, 3D vision, and distributed fiber optic strain multi-source heterogeneous data of key stress sections and transmission and braking components of lifting machinery; At the edge, multi-source heterogeneous data is preprocessed and preliminary feature extraction is performed. The extracted feature vectors are then de-identified and uploaded to the cloud. The cloud-based multimodal large model performs spatiotemporal feature decoupling and correlation mapping on feature vectors, removes pseudo-features, and extracts damage-sensitive feature vectors; Input the damage-sensitive feature vector into the fracture mechanics fatigue damage evolution numerical simulation program to generate the current damage state and remaining life evolution curve; The evolution curve is input in reverse into the causal inference engine of the multimodal large model to optimize the feature extraction weight parameters to enhance the extraction of early latent damage features, forming a closed-loop coupling driven by both data and mechanism. The model is incrementally updated based on the feature vectors of multiple lifting machines, and the updated model weights are sent to the edge. The damage state and evolution trend are compared with the structural strength verification threshold under rated load conditions to generate compliance judgment results and risk warning signals.

[0006] Preferably, the spatiotemporal feature decoupling and correlation mapping includes: Simultaneous compression transformation and wavelet packet decomposition are performed on acoustic emission and vibration data to construct time-frequency domain spectra. The time-frequency domain spectra are mapped to the node attributes in the graph structure data, along with the 3D visual features and strain features. By using graph neural networks to capture the topological relationships of nodes of different modalities in time series, cross-modal neighbor node information is aggregated through graph convolution operations, and isolated subgraphs corresponding to environmental noise and operating condition disturbances are decoupled. Damage-sensitive feature vectors that retain the attributes of core nodes in connected subgraphs are extracted.

[0007] Preferably, the causal reasoning engine employs a counterfactual reasoning mechanism, including: By introducing intervention variables into the feature vector space, a counterfactual feature generator is constructed. The counterfact feature generator performs mask replacement on the working condition covariates in the core node attributes of the connected subgraph to generate counterfact feature vectors without working condition disturbances. Calculate the difference in causal effects between the true feature vector and the counterfactual feature vector on the causal graph, eliminate false causal paths whose difference value is greater than a preset perturbation threshold and have no physical connection with fatigue damage, and output the purified damage-sensitive feature vector.

[0008] Preferably, the optimized feature extraction weight parameters include: The partial differential equations of the numerical simulation program for fracture mechanics fatigue damage evolution are transformed into computational graph constraints of a physical information neural network. Using the purified damage-sensitive feature vector as input, the residual of the partial differential equation at the current time step is added as a penalty term to the loss function of the multimodal large model; A model prediction control strategy is adopted. Within a set time prediction window, the weight parameters of the graph convolution operation and the counterfactual feature generator in the multimodal large model are updated on a rolling basis by backpropagating through the loss function that minimizes the physical residual penalty term.

[0009] Preferably, the de-identification process and uploading to the cloud includes: At the edge, a non-orthogonal multiple access power allocation strategy is adopted, and the transmission power is allocated according to the contribution of each mode data in the feature vector to fatigue damage prediction. A Laplace noise mechanism is superimposed on the core feature components allocated high transmission power, and dimension reduction and compression are performed on the redundant feature components allocated low transmission power. The feature vectors with superimposed noise and dimensionality reduction are multiplexed and uploaded to the cloud via a non-orthogonal multiple access channel. The cloud then separates and restores the feature vectors using a continuous interference cancellation algorithm.

[0010] Preferably, incremental model updates include: After obtaining the feature vectors uploaded by multiple lifting machinery edge devices in the cloud, a ensemble Kalman filter strategy is used for federated aggregation. The weights of the multimodal large model updated locally at each edge are used as state variables, and the consistency between the local model weights and the global damage evolution trend is evaluated by the covariance matrix. Based on the consistency assessment results, dynamic confidence Kalman gain is assigned to the weights of each local model to suppress global model drift caused by local abnormal operating conditions, and updated model weights are generated.

[0011] Preferably, the generation of compliance determination results and risk warning signals includes: After the updated model weights are distributed to the edge, Monte Carlo simulation is used to randomly sample the remaining lifetime evolution curve within the future time window to generate a probability distribution sequence of damage evolution. A conditional value at risk model is introduced to calculate the expected tail risk of a damage evolution probability distribution sequence exceeding the structural strength verification threshold at a set confidence level. When the expected value of tail risk is greater than zero, a dynamic graded early warning signal characterizing the risk of latent fatigue damage is generated, and the compliance judgment result is output in combination with the current damage status.

[0012] Preferably, the multimodal large model employs a speculative decoding strategy when processing feature vectors: Deploy draft models with small parameters and large multimodal models with large parameters in the cloud for parallel operation; The draft model quickly generates a draft sequence of damage feature evolution paths based on feature vectors uploaded from the edge. The large-parameter multimodal model verifies the damage features of each time step in the draft sequence in parallel during a single forward propagation, retains the draft sequence segments that pass the verification, and only performs recalculation of the causal inference engine for the segments that fail the verification.

[0013] Preferably, deep reinforcement learning dynamic scheduling is used when preprocessing and initially extracting features from multi-source heterogeneous data at the edge: Construct an action space that includes a candidate set of sampling frequencies and preprocessing precision for multi-source heterogeneous data, and a state space that uses the test miss rate and edge computing resource consumption as a joint penalty function; The edge agent outputs an action strategy based on the current working condition of the lifting machinery using a deep deterministic policy gradient algorithm. By adjusting the sampling frequency and preprocessing accuracy of the next time step based on the action strategy, the extraction gain of early latent damage features is maximized under the condition of limited computing resources.

[0014] Preferably, a multimodal large-scale intelligent inspection system for lifting machinery is used to implement the method, comprising: The sensing and acquisition device is deployed on the key stress-bearing sections and transmission and braking components of the lifting machinery to simultaneously acquire multi-source heterogeneous data of acoustic emission, three-dimensional vision and distributed fiber strain. Edge computing nodes are connected to sensing and acquisition devices to preprocess and extract preliminary features from multi-source heterogeneous data, and to desensitize and upload the extracted feature vectors. The cloud-based causal reasoning and simulation server communicates with edge computing nodes and has a built-in multimodal large model and fracture mechanics fatigue damage evolution numerical simulation program. It is used to perform spatiotemporal decoupling and causal reasoning on feature vectors and forms a closed-loop coupling with the numerical simulation program driven by both data and mechanism. It completes incremental model updates based on the feature vectors of multiple cranes and distributes the updated model weights to the edge computing nodes. The judgment and early warning terminal communicates with the cloud-based causal reasoning and simulation server to compare the damage state and evolution trend with the structural strength verification threshold under rated load conditions, and generate compliance judgment results and risk warning signals.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention solves the problem of missed detection of latent defects caused by pseudo-feature interference under varying working conditions in existing technologies by constructing a closed-loop coupled architecture of a multimodal large-scale model and a numerical simulation program for fracture mechanics fatigue damage evolution. The multimodal large-scale model performs spatiotemporal decoupling and correlation mapping on feature vectors, and combines graph neural networks and counterfactual reasoning mechanisms to eliminate isolated subgraphs and false causal paths corresponding to environmental noise and working condition disturbances, thereby extracting damage-sensitive feature vectors. The numerical simulation program transforms partial differential equations into computational graph constraints of a physical information neural network, using physical residuals as penalty terms to back-optimize the weight parameters of feature extraction in the large-scale model. This forces the model to follow the physical laws of fatigue damage evolution during the feature extraction stage, eliminating interference from working condition covariates, and ensuring that the extracted feature vectors directly correspond to the early latent fatigue damage state of the structure.

[0016] 2. In the edge data transmission and model update stages, the edge end allocates non-orthogonal multiple access transmission power based on feature contribution and adds Laplace noise to core feature components, reducing transmission resource consumption while ensuring data privacy. In the cloud, an ensemble Kalman filter strategy is used for federated aggregation. The covariance matrix is ​​used to evaluate the consistency between local weights and global trends and to allocate Kalman gains, suppressing global model drift caused by local anomalies. In the cloud inference stage, a speculative decoding strategy is employed. Draft sequences are quickly generated using draft models and verified in parallel by large models, reducing redundant computations in causal inference and lowering the latency of test result generation. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall process of intelligent inspection of lifting machinery according to the present invention. Figure 2 This is a flowchart of the spatiotemporal dimension feature decoupling and correlation mapping of the present invention; Figure 3 This is a flowchart of the causal reasoning and model weight optimization process of the present invention; Figure 4 This is a flowchart of the edge feature desensitization and uploading process of the present invention; Figure 5 This is a flowchart of the model federation incremental update process of the present invention; Figure 6 This is a flowchart of the risk warning and compliance determination process for the present invention. Detailed Implementation

[0018] The intelligent inspection system for lifting machinery based on a multimodal large model disclosed in this invention includes a sensor acquisition device, an edge computing node, a cloud-based causal inference and simulation server, and a judgment and early warning terminal. The sensor acquisition device is deployed on the key stress sections and transmission and braking components of the lifting machinery and is coupled to the edge computing node; the edge computing node is coupled to the cloud-based causal inference and simulation server via a wireless communication link; and the cloud-based causal inference and simulation server is coupled to the judgment and early warning terminal.

[0019] The sensing and acquisition device includes an acoustic emission sensing unit, a three-dimensional vision acquisition unit, and a distributed fiber optic strain sensing unit. The acoustic emission sensing unit employs a broadband resonant acoustic emission sensor with a resonant frequency range of 100kHz-1MHz. It is deployed at the mid-span section of the main beam, the end beam support connection section, the stress-bearing section of the hook pulley block, and transmission and braking components such as the reducer gearbox, brake wheel, and drum drive shaft. This unit is used to collect acoustic emission stress wave signals generated by material plastic deformation and crack propagation during the lifting machinery's loading process. The three-dimensional vision acquisition unit uses a binocular structured light camera with a baseline distance of 100mm-300mm and a measurement accuracy of ±0.1mm. It is deployed below the crane operator's cab and at the ends of the outriggers, covering the upper and lower flanges of the main beam, the web welding area, and the moving contact surfaces of the transmission and braking components. This unit is used to collect three-dimensional point cloud data and surface deformation images of key stress-bearing sections. The distributed fiber optic strain sensing unit uses a single-mode sensing fiber based on Brillouin optical time-domain analysis, continuously deployed along the longitudinal and transverse web welding seams of the main beam. It has a spatial resolution of 0.5m and a strain measurement range of ±15000. The temperature measurement range is -20℃ to 80℃, and it is used to collect distributed strain and temperature data of the entire range of lifting machinery.

[0020] The edge computing node adopts an industrial-grade edge computing gateway, equipped with a multi-core ARM processor and FPGA hardware acceleration unit, and has built-in preprocessing module, preliminary feature extraction module, desensitization processing module and local inference module. The preprocessing module is coupled to the output end of each sensing unit of the sensing acquisition device, the preliminary feature extraction module is coupled to the output end of the preprocessing module, the desensitization processing module is coupled to the output end of the preliminary feature extraction module, and the local inference module is coupled to the input end of the model weights issued by the cloud causal inference and simulation server.

[0021] The multi-core ARM processor is responsible for running the actor and critic networks of the deep reinforcement learning dynamic scheduling algorithm, the feature desensitization processing algorithm, and the model inference task of the local inference module. The FPGA hardware acceleration unit is divided into a parallel preprocessing acceleration channel, a feature extraction acceleration channel, and a data transmission channel. The parallel preprocessing acceleration channel corresponds to three parallel preprocessing logics for acoustic emission signals, 3D visual point cloud data, and distributed fiber strain data. Each channel is equipped with a configurable filter IP core and a feature statistical calculation IP core, which are coupled with the action policy signal output by the deep reinforcement learning dynamic scheduling algorithm. The filter order, sampling frequency division coefficient, and feature extraction dimension are adjusted in real time according to the action policy. The feature extraction acceleration channel is equipped with a parallel multiply-accumulate operation array to realize the parallel calculation of time-domain and frequency-domain features in the initial feature extraction process. The data transmission channel is equipped with a non-orthogonal multiple access signal encoding IP core, which is coupled with the output of the feature desensitization processing module. After completing the channel encoding and power allocation of the feature vector, it is uploaded to the cloud server through the 5G industrial module. The local inference module loads the updated model weights from the cloud and implements lightweight damage feature inference in the ARM processor. The inference results can be directly output to the local display screen to achieve real-time damage assessment and on-site early warning at the edge. At the same time, the inference results and feature vectors are uploaded to the cloud synchronously.

[0022] The cloud-based causal inference and simulation server employs a GPU cluster server, integrating a multimodal large model, a causal inference engine, a fracture mechanics fatigue damage evolution numerical simulation program, and a federated incremental update module. The input of the multimodal large model is coupled to the feature vector output uploaded by the edge computing nodes; the causal inference engine is coupled to the feature output of the multimodal large model; the input of the fracture mechanics fatigue damage evolution numerical simulation program is coupled to the damage-sensitive feature vector output by the multimodal large model; and the output of the numerical simulation program is coupled to the weight optimization input of the causal inference engine, forming a closed-loop coupled architecture driven by both data and mechanism. The federated incremental update module is coupled to the feature vectors and local model weights uploaded by multiple lifting machinery edge computing nodes; and the output of the federated incremental update module is coupled to the weight update input of the multimodal large model.

[0023] The GPU cluster server adopts a multi-node distributed architecture, with each node configured with no fewer than 8 NVIDIA A100 80GB TensorCore GPUs. Nodes communicate with each other via an InfiniBand HDR200G high-speed interconnect network. The cluster uses Kubernetes for resource scheduling and containerized deployment, allocating independent GPU resource pools and CPU scheduling resources to the multimodal large model service, causal inference engine service, numerical simulation service, and federated aggregation service. Specifically, the multimodal large model service employs a distributed deployment method combining tensor parallelism and pipelined parallelism, deploying the Transformer encoder layer, graph neural network module, and causal inference engine of the large-parameter multimodal large model on different GPU nodes to achieve parallel computation of feature decoupling, counterfactual inference, and weight optimization. The small-parameter draft model is deployed on a single GPU node and interacts with the large model service through shared memory to complete the parallel inference process of the inference decoding strategy. The numerical simulation program for fracture mechanics fatigue damage evolution is accelerated using CUDA and deployed on dedicated GPU computing nodes. It decomposes the partial differential equation solution process into multi-threaded parallel computing tasks, with a solution latency of no more than 200ms for a single remaining life evolution curve. It can simultaneously support numerical simulation requests from at least 1000 edge nodes of lifting machinery. The federated incremental update module is deployed on the cluster's management node, receiving feature vectors and local model weights uploaded by each edge node via a message queue. The federated aggregation algorithm based on an ensemble Kalman filter strategy employs multi-threaded parallel computing, enabling minute-level global model incremental updates and weight distribution.

[0024] The judgment and early warning terminal adopts an industrial-grade touch display terminal, which has a built-in threshold comparison module, early warning generation module and inspection report output module. The input end of the threshold comparison module is coupled with the damage status and remaining life evolution curve output by the cloud server. The threshold comparison module has a built-in structural strength verification threshold under rated load conditions. The early warning generation module is coupled with the output end of the threshold comparison module. The inspection report output module is coupled with the output ends of the early warning generation module and the threshold comparison module.

[0025] This terminal adopts an explosion-proof industrial design, making it suitable for flammable, explosive, high-dust, and high-vibration environments at crane operation sites. It features a 10.1-inch high-brightness touchscreen display, a built-in quad-core industrial-grade processor, and more than 2GB of RAM, and supports offline storage of at least one year's worth of inspection data and early warning records. The terminal communicates with the cloud server via a wireless LAN, receiving damage status, remaining life evolution curves, and tail risk expected values ​​data from the cloud in real time. The threshold comparison module verifies thresholds based on built-in GB / T3811-2008 and GB / T6067.1-2010 standards, performing secondary verification of the compliance judgment logic locally. The early warning generation module drives the terminal's built-in audible and visual alarm module to provide corresponding on-site audible and visual warnings based on the early warning level corresponding to the tail risk expected value. The inspection report output module has built-in report templates conforming to the periodic inspection specifications for lifting machinery. It can automatically fill in damage status data, compliance judgment results, early warning information, and remaining life curves into the templates to generate editable and printable standardized inspection reports. It also supports uploading reports to the special equipment safety supervision platform, enabling full-process traceability of inspection data. The terminal is also equipped with a manual inspection trigger interface, which can connect to handheld flaw detectors, thickness gauges, and other equipment used by on-site inspectors, merging manual inspection data with intelligent inspection results to form a complete safety inspection file for lifting machinery.

[0026] Example 1: Please refer to the appendix Figure 1 The sensing and acquisition device synchronously collects multi-source heterogeneous data on acoustic emission, 3D vision, and distributed fiber optic strain from key stress sections and transmission and braking components of the lifting machinery. Specifically, a BeiDou time synchronization module provides a synchronous clock signal to each sensing unit of the sensing and acquisition device, with a synchronization accuracy better than 1μs, ensuring the timestamp alignment of acoustic emission signals, 3D vision point cloud data, and distributed fiber optic strain data. During the full-condition cycle of lifting, running, luffing, and slewing of the lifting machinery under rated load, the acoustic emission sensing unit continuously collects acoustic emission time-domain signals at a sampling frequency of 2MHz, the 3D vision acquisition unit acquires the 3D point cloud sequence of key sections at a sampling frequency of 20Hz, and the distributed fiber optic strain sensing unit collects full-range distributed strain data at a sampling frequency of 5Hz, forming a timestamp-aligned multi-source heterogeneous raw dataset.

[0027] Each sensing unit of the sensing acquisition device is connected to the FPGA hardware acceleration unit of the edge computing node via an industrial Ethernet bus. The 1PPS pulse signal and the synchronization clock signal output by the Beidou time synchronization module are respectively connected to the synchronization trigger terminal of each sensing unit and the clock synchronization terminal of the edge computing node to achieve sub-microsecond time synchronization of the entire sensing link. The analog signal output by the acoustic emission sensing unit is converted into a digital signal by the built-in 16-bit high-speed ADC chip and then transmitted to the preprocessing acceleration channel of the FPGA through the LVDS interface. The 3D point cloud data output by the 3D vision acquisition unit is transmitted to the ARM processor of the edge computing node through the gigabit Ethernet interface. The demodulator of the distributed fiber optic strain sensing unit communicates with the edge computing node through the RS485 bus to upload strain and temperature data to the preprocessing module in real time. The action strategy signal output by the deep reinforcement learning dynamic scheduling algorithm is connected to the sampling control terminal of each sensing unit through the GPIO interface of the edge computing node. When the algorithm outputs a high sampling frequency strategy, it directly triggers the sampling frequency of the acoustic emission sensing unit to be increased to 5MHz, the sampling frequency of the three-dimensional vision acquisition unit to be increased to 50Hz, and the sampling frequency of the distributed fiber optic strain sensing unit to be increased to 20Hz. This achieves hardware-level synchronous adjustment of the sampling parameters of the entire sensing link, ensuring the complete acquisition of damage characteristics under high-risk conditions.

[0028] Edge computing nodes preprocess and perform preliminary feature extraction on multi-source heterogeneous data, de-identify the extracted feature vectors, and upload them to the cloud. For acoustic emission time-domain signals, preprocessing includes bandpass filtering, amplitude thresholding denoising, ringing counting, and energy parameter extraction; for 3D visual point cloud data, preprocessing includes point cloud denoising, registration, point cloud segmentation, and deformation displacement calculation; for distributed fiber optic strain data, preprocessing includes temperature compensation, outlier removal, and moving average smoothing. In the initial feature extraction process, for the preprocessed acoustic emission signal, time-domain features such as peak value, rise time, duration, energy, ring count, and root mean square value are extracted, along with frequency-domain features such as spectral peak value, center frequency, and frequency band energy distribution extracted through Fast Fourier Transform. For the 3D visual point cloud data, 3D coordinate displacement, deformation, and curvature change features of key points are extracted. For the distributed fiber optic strain data, strain peak value, strain gradient, and strain accumulation features of each monitoring section are extracted. The features extracted from each mode are aligned according to the timestamp and concatenated to form an initial feature vector of dimension N×D, where N is the time step length and D is the sum of the dimensions of each mode feature. After anonymizing the initial feature vector, it is uploaded to the cloud-based causal inference and simulation server via a 5G wireless communication link.

[0029] The cloud-based multimodal large-scale model performs spatiotemporal feature decoupling and correlation mapping on feature vectors, eliminating pseudo-features and extracting damage-sensitive feature vectors. This multimodal large-scale model adopts a Transformer architecture, embedding a graph neural network module and a causal inference engine in the encoder layer. The input is the desensitized feature vector uploaded from the edge, and the output is the damage-sensitive feature vector. Specifically, the multimodal large-scale model decouples the input feature vectors in both temporal and spatial dimensions. In the temporal dimension, it captures the temporal evolution correlation of feature sequences through a self-attention mechanism. In the spatial dimension, it captures the spatial topological correlation of different modal features and features from different monitoring locations through a graph neural network. This decoupling and elimination of pseudo-features corresponding to environmental noise and operational disturbances retains feature components directly related to fatigue damage evolution, generating damage-sensitive feature vectors.

[0030] The damage-sensitive feature vector is input into a numerical simulation program for fatigue damage evolution in fracture mechanics to generate the current damage state and remaining life evolution curves. This numerical simulation program is based on linear elastic fracture mechanics and continuous damage mechanics, employing the Paris formula for fatigue crack propagation and partial differential equations for damage evolution. The inputs are the current stress state, initial crack size, and material fracture toughness parameters corresponding to the damage-sensitive feature vector. The outputs are the current damage variable, crack propagation length, remaining load cycle count, and remaining life evolution curves. Specifically, the strain peak and strain gradient extracted from the damage-sensitive feature vector are transformed into the stress field distribution of key structural sections. Combined with the material property parameters of the structural steel used in the lifting machinery, including elastic modulus, Poisson's ratio, fracture toughness KIC, fatigue crack propagation parameters C and m, the fatigue damage evolution equations are solved through the numerical simulation program to calculate the damage variable D at the current time step and the damage evolution path under future load cycles, generating the remaining life evolution curve as a function of the number of load cycles.

[0031] The evolution curve is input inversely into the causal inference engine of the multimodal large model to optimize the feature extraction weight parameters, thereby enhancing the extraction of early latent damage features and forming a closed-loop coupling driven by both data and mechanism. Specifically, the physical laws of damage evolution corresponding to the remaining life evolution curve output by the numerical simulation program are transformed into physical constraints of the multimodal large model. Based on these physical constraints, the causal inference engine calculates the residual between the current feature extraction result of the multimodal large model and the physical constraints. With the goal of minimizing the residual, the weight parameters of the feature extraction layer in the multimodal large model are updated through the backpropagation algorithm. This forces the model's feature extraction process to follow the physical laws of fatigue damage evolution, enhancing the ability to extract weak features corresponding to early latent fatigue damage, and forming a closed-loop coupling between the data-driven multimodal large model and the mechanism-driven fatigue damage numerical simulation.

[0032] Incremental model updates are performed based on feature vectors from multiple cranes, and the updated model weights are then distributed to the edge computing nodes. The cloud server aggregates feature vectors and local model update parameters uploaded from the edge computing nodes of multiple cranes of the same model and operating conditions. A federated learning strategy is used for incremental updates of the global model. Without acquiring the original data from each edge computing node, global optimization of the multimodal large model is achieved. The updated global model weights are then distributed to each participating edge computing node via wireless communication links. The local inference modules of the edge computing nodes load the updated model weights and perform local feature extraction and preliminary damage identification.

[0033] The damage state and evolution trend are compared with the structural strength verification thresholds under rated load conditions to generate compliance judgment results and risk warning signals. The structural strength verification thresholds under rated load conditions include allowable stress threshold, allowable damage variable threshold, critical crack size threshold, and minimum remaining life threshold. The damage variables, crack size, remaining life, and damage evolution trend of the current damage state are compared one by one with the corresponding verification thresholds. When the damage parameters exceed the allowable thresholds, a corresponding risk warning signal is generated. Simultaneously, based on the comparison results, compliance judgment results for the structural safety and transmission braking performance of the lifting machinery are generated and output to the judgment and warning terminal.

[0034] As a preferred embodiment, refer to the appendix. Figure 2 When performing spatiotemporal feature decoupling and correlation mapping on feature vectors in a multimodal large model, synchronous compression transformation and wavelet packet decomposition are first performed on acoustic emission and strain data to construct a time-frequency domain spectrum. For acoustic emission time-domain signals... With distributed fiber strain time-domain signal High-resolution time-frequency distribution is obtained by synchronous compression transform. The mathematical expression is:

[0035] in, For the short-time Fourier transform of the signal, The value of the window function at zero. For a signal at time t and frequency instantaneous frequency at that point This is the Dirac function.

[0036] A three-level wavelet packet decomposition is performed on the time-frequency signal after synchronous compression transformation. The db4 wavelet basis function is used to decompose the time-frequency signal into 8 frequency bands. The wavelet packet energy of each frequency band is calculated, and a time-frequency domain spectrum with dimension T×F is constructed, where T is the time step length and F is the number of frequency bands.

[0037] The time-frequency domain spectra, 3D visual features, and strain features are mapped to node attributes in the graph structure data, respectively. The core components of the heterogeneous graph structure are divided into three parts: a node set, an edge set, and an adjacency matrix. The node set includes three types of nodes: acoustic emission feature nodes, 3D visual feature nodes, and strain feature nodes. Each type of node corresponds to a feature sequence of a monitoring location. The edge set E consists of topological connections between nodes, including spatial connections between nodes of different modes at the same monitoring location and temporal connections between nodes of the same mode at different time steps. The node attribute matrix A consists of the feature vectors of each node. The attributes of acoustic emission feature nodes are the time-frequency domain spectra feature vectors of the corresponding time step, the attributes of 3D visual feature nodes are the deformation and curvature feature vectors of the corresponding time step, and the attributes of strain feature nodes are the strain peak value and strain gradient feature vectors of the corresponding time step.

[0038] A graph neural network is used to capture the topological relationships between nodes of different modalities over time. Cross-modal neighbor node information is aggregated through graph convolution operations to decouple isolated subgraphs corresponding to environmental noise and operational disturbances. Damage-sensitive feature vectors that retain the core node attributes of the connected subgraphs are then extracted. The graph neural network employs a graph attention convolutional network, and the mathematical expression for the graph convolution operation is:

[0039] in, For the first The output feature vector of node i in the layer network. It is a non-linear activation function. For nodes The set of neighboring nodes, Let be the attention weight between node i and its neighbor node j. Let be the trainable weight matrix of the l-th layer network. For the first The input feature vector of neighbor node j in the layer network.

[0040] By employing multi-layer graph convolution operations, neighbor node information across time and modality is aggregated to generate embedded feature vectors for all nodes in the graph. Subgraph segmentation is performed based on these node embedded feature vectors. A spectral clustering algorithm divides the heterogeneous graph into multiple connected subgraphs. The temporal stability and damage correlation of each connected subgraph are calculated. Subgraphs with temporal fluctuations exceeding a preset fluctuation threshold and unrelated to the physical laws of fatigue damage are identified as isolated subgraphs corresponding to environmental noise and operational disturbances and are discarded. Connected subgraphs whose temporal evolution conforms to the laws of fatigue damage accumulation are retained, and the attribute feature vectors of the core nodes of these connected subgraphs are extracted and output as damage-sensitive feature vectors.

[0041] As a preferred embodiment, refer to the appendix. Figure 3The causal reasoning engine employs a counterfactual reasoning mechanism, and its specific implementation process is as follows: An intervention variable is introduced into the feature vector space to construct a counterfactual feature generator. A damage feature causal map is constructed based on a structural causal model. This causal map includes the outcome variable Y (fatigue damage state), the treatment variable T (damage-sensitive feature vector), the operating condition covariate X (operating condition disturbance characteristics such as lifting load, operating speed, and ambient temperature), and the confounding variable U (unobserved environmental disturbance variables). An intervention variable is then introduced. ,in To establish a baseline covariate value without operational disturbance, a counterfactual feature generator based on a generative adversarial network is constructed. The input of the counterfactual feature generator is the true feature vector and the operational covariate, and the output is the counterfactual feature vector after intervention.

[0042] A counterfactual feature generator is used to mask and replace the load condition covariates in the core node attributes of the connected subgraph, generating counterfactual feature vectors without load condition perturbations. Specifically, the load condition covariate components in the core node attribute feature vector of the connected subgraph are masked and replaced with... Input the counterfactual feature generator to generate counterfactual feature vectors under baseline operating conditions without operating condition disturbances. The optimization objective of the counterfactual feature generator is to minimize the distribution difference between the generated features and the true features under no operational disturbances, and its loss function is... The expression is:

[0043] in, For the true feature vector, For the generated counterfactual feature vector, The coefficient of the gradient penalty term, is the gradient operator, and D is the discriminator used to distinguish between true features and counterfactual features.

[0044] The difference in causal effects between the true feature vector and the counterfactual feature vector on the causal graph is calculated. False causal paths with a difference value greater than a preset perturbation threshold and no physical connection to fatigue damage are eliminated, and the purified damage-sensitive feature vector is output. The difference in causal effects is the average processing effect. Its mathematical expression is:

[0045] in, The expected damage state corresponding to the true feature vector. This represents the expected damage state corresponding to the counterfactual feature vector.

[0046] For each causal path in the causal graph, its corresponding ATE value is calculated. When the ATE value is greater than the preset perturbation threshold and the feature change corresponding to the path cannot be explained by the physical laws of fatigue damage evolution, the path is determined to be a false causal path and is removed. Causal paths whose ATE values ​​conform to physical laws and are directly related to fatigue damage are retained, and the feature vectors of the corresponding paths are extracted as the purified damage-sensitive feature vectors.

[0047] The specific implementation process for optimizing feature extraction weight parameters is as follows: The partial differential equations of the numerical simulation program for fatigue damage evolution in fracture mechanics are transformed into computational graph constraints of a physical information neural network. The partial differential equations for fatigue damage evolution are constructed based on continuous damage mechanics, and their expressions are as follows:

[0048] in, As a damage variable, For the number of load cycles, For stress amplitude, The material hardening index. Where is Poisson's ratio, and E is the elastic modulus. This is the threshold value for fatigue crack propagation. The Paris formula exponent is used to define fatigue crack propagation.

[0049] The aforementioned partial differential equations are used as physical constraints and embedded into the computational graph of a physical information neural network to construct a system based on damage variables. The neural network, with stress amplitude, material parameters, and load cycle number as inputs, transforms the solution of partial differential equations into a neural network optimization problem.

[0050] Using the purified damage-sensitive feature vector as input, the residual of the partial differential equation at the current time step is added as a penalty term to the loss function of the multimodal large model. The total loss function of the multimodal large model... The expression is:

[0051] in, Let cross-entropy be the loss function for the damage recognition task. The coefficient of the physical residual penalty term. The residuals of the partial differential equation for fatigue damage evolution are... The counterfactual loss coefficient, This is the loss function for the counterfactual feature generator.

[0052] The physical residual The expression is:

[0053] in, The number of samples at each time step. Let be the damage variable output by the neural network at the i-th time step. Let be the stress amplitude transformed from the damage-sensitive eigenvector at the i-th time step.

[0054] A model predictive control strategy is adopted. Within a set time prediction window, the weight parameters of the graph convolution operation and the counterfactual feature generator in the multimodal large model are updated on a rolling basis by backpropagation that minimizes the loss function including the physical residual penalty term. Specifically, a time prediction window of length H is set. At each control step, based on the damage-sensitive feature vector at the current time step, the damage evolution curve for the next H time steps is predicted through numerical simulation. The total loss function within the prediction window is then calculated. Backpropagation is performed through the adaptive moment estimation optimizer to calculate the gradient of the loss function with respect to the weight matrix of the graph convolution operation and the network weights of the counterfactual feature generator, and update the corresponding weight parameters. In the next control step, the damage-sensitive feature vector is re-extracted based on the updated model, and the above optimization process is repeated to achieve rolling optimization of the model weights and enhance the ability to extract early latent damage features.

[0055] As a preferred embodiment, refer to the appendix. Figure 4 At the edge, a non-orthogonal multiple access (MOA) power allocation strategy is adopted, allocating transmission power based on the contribution of each modal data in the feature vector to fatigue damage prediction. First, the correlation between each modal feature component in the feature vector and the fatigue damage state is calculated using the Pearson correlation coefficient. This correlation is the contribution of that feature component to fatigue damage prediction, and its mathematical expression is:

[0056] in, For the first The contribution of each feature component For the first each feature component With damage variables covariance, For the first The variance of each characteristic component, The variance of the damage variable.

[0057] Based on the contribution of each characteristic component, a non-orthogonal multiple access power allocation algorithm is used to allocate transmission power, with the total transmission power constrained as follows: The transmission power allocated to the kth characteristic component satisfy:

[0058] in, This represents the total dimension of the feature vector. Feature components with a contribution greater than a preset contribution threshold are classified as core feature components and assigned high transmission power; feature components with a contribution less than or equal to the preset contribution threshold are classified as redundant feature components and assigned low transmission power.

[0059] A Laplace noise mechanism is superimposed on the core feature components allocated high transmission power, while dimensionality reduction and compression are performed on redundant feature components allocated low transmission power. For the core feature components, a differential privacy Laplace noise mechanism is employed, and the superimposed Laplace noise satisfies the scale parameter. ,in For the global sensitivity of the feature components, For differential privacy budgeting, the expression for the core feature components after adding noise is:

[0060] in, These are the core feature components after adding noise. To comply with scale parameters Random noise with a Laplace distribution.

[0061] For redundant feature components, principal component analysis algorithm is used to perform dimensionality reduction and compression, retaining principal component components with a cumulative contribution rate greater than 95%, and mapping high-dimensional redundant feature components into low-dimensional compressed feature vectors to reduce the amount of data transmitted.

[0062] The noise-superimposed and dimension-reduced feature vectors are multiplexed and uploaded to the cloud via a non-orthogonal multiple access channel. The cloud then uses a continuous interference cancellation algorithm to separate and reconstruct the feature vectors. At the edge, the processed feature vectors are non-orthogonally superimposed on the same time-frequency resource block and uploaded to the cloud server via a 5G wireless channel. After receiving the superimposed signal, the cloud server uses a serial interference cancellation algorithm to decode the core feature components and compressed redundant feature components sequentially according to the transmission power from high to low, eliminating signal interference between multiple users, separating and restoring each feature component, and reconstructing the complete feature vector.

[0063] As a preferred embodiment, refer to the appendix. Figure 5 After obtaining feature vectors uploaded from multiple edge computing nodes of lifting machinery in the cloud, a ensemble Kalman filter strategy is used for federated aggregation. The number of edge computing nodes participating in the federated aggregation is set to M, and each edge node corresponds to a local multimodal large model. The local model weight vector of the i-th edge node is... Where d is the total dimension of the model weights, The set of real numbers represents the global model weight vector. .

[0064] Construct the state equation and observation equation for the ensemble Kalman filter. The state equation describes the evolution of the global model weights, and its expression is:

[0065] in, For the first The global model weights for each update step. For the first The global model weights for each update step. The noise is the process noise, which has a mean of 0 and a covariance matrix of... The Gaussian distribution.

[0066] The observation equation describes the mapping relationship between the local model weights and the global model weights, and its expression is:

[0067] in, For the first The first update step Local model weights of each edge node For the observation matrix, The observed noise follows a mean of 0 and a covariance matrix of... The Gaussian distribution.

[0068] The weights of the multimodal large model, updated locally at each edge, are used as state variables. The consistency between the local model weights and the global damage evolution trend is evaluated through the covariance matrix. First, a set of samples of global model weights is generated. ,in Given the number of samples in the set, calculate the mean and covariance matrix of the set:

[0069]

[0070] in, The mean of the global model weight set. The covariance matrix of the global model weights. for The first time in the global model weight set at time step Each sample weight vector.

[0071] The consistency coefficient between each local model weight and the global damage evolution trend is calculated based on the covariance matrix. This consistency coefficient is the reciprocal of the Mahalanobis distance between the local and global model weights, expressed as:

[0072] in, For the first The consistency coefficient of the local model at each edge node is determined by the Mahalanobis distance. The larger the Mahalanobis distance, the smaller the consistency coefficient, indicating that the local model is less consistent with the global trend.

[0073] Based on the consistency assessment results, dynamic confidence Kalman gains are assigned to the weights of each local model to suppress global model drift caused by local abnormal operating conditions, generating updated model weights. The expression for the Kalman gain matrix is:

[0074] in, Here is the Kalman gain matrix. This is the weighted average observation noise covariance matrix, with weights equal to the consistency coefficients of each local model. The expression is:

[0075] This is the total number of edge nodes participating in this global model weight aggregation.

[0076] Update the global model weights and covariance matrix based on the Kalman gain matrix:

[0077]

[0078] in, It is an identity matrix.

[0079] Through the above process, higher weights are assigned to local models with high consistency, and lower weights are assigned to local models with low consistency. This suppresses global model drift caused by local abnormal operating conditions, generates updated global model weights, and completes the incremental update of the multimodal large model.

[0080] As a preferred embodiment, refer to the appendix. Figure 6 After distributing the updated model weights to the edge, Monte Carlo simulation is used to randomly sample the remaining life evolution curve within the future time window, generating a probability distribution sequence of damage evolution. The future prediction time window length is set to T. Load amplitude, material performance parameters, and environmental parameters in the fracture mechanics fatigue damage evolution model are used as random variables. Based on the historical statistical distribution of each parameter, Monte Carlo random sampling is performed, with the number of samplings being [number missing]. For each set of sampling parameters, the damage variable evolution sequence for the next T time steps is solved using a fatigue damage evolution numerical simulation program, generating... The damage evolution path is analyzed, and the probability distribution of damage variables at each time step is statistically analyzed to form a probability distribution sequence of damage evolution.

[0081] A conditional value at risk (VAT) model is introduced to calculate the expected tail risk of a damage evolution probability distribution sequence exceeding a structural strength verification threshold at a set confidence level. This conditional value at risk... The mathematical expression is:

[0082] in, For the conditional expectation operator, The set confidence level has a value range of 0.95-0.999. Confidence level Value at risk (VaR), i.e., the damage variable at a confidence level. The quantiles below satisfy the probability condition , For damage variables.

[0083] Based on the probability distribution sequence of damage evolution, calculate the probability distribution at each time step. and The allowable damage variable threshold in the structural strength verification threshold is denoted as... Calculate the expected value of tail risk:

[0084] in, For the expected value of tail risk, when A value greater than 0 indicates that, at the set confidence level, the tail expectation of the damage variable exceeds the allowable threshold, indicating a risk of latent fatigue damage.

[0085] When the expected tail risk value is greater than zero, a dynamic graded early warning signal characterizing the risk of latent fatigue damage is generated, and a conformity judgment result is output in conjunction with the current damage status. Based on the magnitude of the expected tail risk value, the early warning signal is divided into four levels: Level 1 early warning corresponds to… >0.5 The structure indicates a serious risk of fatigue damage, and continued operation is prohibited; a level two warning corresponds to 0.2. < ≤0.5 The characterization structure indicates a significant risk of fatigue damage, requiring immediate shutdown and inspection; Level 3 warning corresponds to 0 < ≤0.2 The characterization structure indicates a risk of early-stage latent fatigue damage, necessitating shorter testing cycles and enhanced monitoring; Level IV early warning corresponds to... ≤0 indicates that the structural damage is within the acceptable range and there is no risk.

[0086] At the same time, the damage variables, crack size, remaining life and structural strength verification threshold of the current damage state are compared one by one to generate three types of compliance judgment results: qualified, unqualified and required to be rectified within a time limit. These results, along with the early warning signal, are output to the judgment and early warning terminal.

[0087] As a preferred embodiment, the multimodal large model employs a speculative decoding strategy when processing feature vectors. The specific implementation process is as follows: A draft model with small parameters and a large multimodal model with large parameters are deployed in the cloud and run in parallel. The draft model is a lightweight Transformer model with 1 / 10 to 1 / 5 of the parameters of the large multimodal model. It retains only the encoder layer and feature mapping layer of the large multimodal model, and removes the causal inference engine and physical constraint module. The inference speed is 5 to 10 times that of the large parameter model. The large multimodal model with large parameters is a complete Transformer model with embedded physical constraints and a causal inference engine. It has complete feature decoupling, causal inference and damage recognition capabilities.

[0088] The draft model rapidly generates a draft sequence of damage feature evolution paths based on feature vectors uploaded from the edge. The draft model receives feature vectors uploaded from the edge, extracts features through a shallow encoder, and quickly generates a draft sequence of damage feature evolution paths of length K. This draft sequence contains predicted damage feature values ​​and damage state predictions for the next K time steps.

[0089] The large-parameter multimodal model verifies the impairment features of each time step in the draft sequence in parallel during a single forward propagation. It retains the successfully verified draft sequence segments and only recalculates the causal inference engine for segments that fail verification. Specifically, the large-parameter multimodal model inputs the features of K time steps of the draft sequence into the network in parallel. During a single forward propagation, it calculates the confidence score of each time step's draft feature. When the confidence score of a draft feature is greater than a preset confidence threshold, it is considered to have passed verification, and the draft sequence segment is retained. When the confidence score of a draft feature is less than or equal to the preset confidence threshold, it is considered to have failed verification, and the counterfactual inference and physical constraint optimization of the causal inference engine are re-executed for the features of that time step to generate a corrected feature sequence. Through this process, the amount of redundant computation in the causal inference process is reduced, and the latency in generating the verification results is lowered.

[0090] When preprocessing and performing preliminary feature extraction on multi-source heterogeneous data at the edge, deep reinforcement learning dynamic scheduling is used. The specific implementation process is as follows: An action space is constructed, comprising candidate sets of sampling frequencies and preprocessing accuracies for multi-source heterogeneous data, and a state space with a joint penalty function of the false negative rate and edge computing resource consumption. The state space S includes the current operating state of the lifting machinery, including lifting load, operating speed, number of operation cycles, edge computing resource occupancy, and the damage identification confidence score of the current feature extraction. The action space A includes candidate sets of sampling frequencies for acoustic emission sensing units, 3D vision acquisition units, distributed fiber optic strain sensing units, preprocessing filter orders, and feature extraction dimensions.

[0091] Construct a reward function, using the false negative rate and edge computing resource consumption as a joint penalty term, with the expression as follows:

[0092] in, for state of time Next action Instant rewards This is the penalty coefficient for the missed detection rate. The damage miss rate under the current action strategy. To calculate the resource consumption penalty coefficient, Calculate resource utilization for the edge device under the current action strategy.

[0093] The edge agent outputs an action policy based on the current working state of the lifting machinery using a deep deterministic policy gradient algorithm. This algorithm comprises an actor network and a critic network. The actor network outputs a deterministic action policy based on the current state, while the critic network evaluates the value of the action policy and updates the actor network's parameters. The optimization objective of the actor network is to maximize the cumulative reward, while the optimization objective of the critic network is to minimize the temporal difference error. The expression for the temporal difference error is:

[0094] in, For timing difference error, For immediate returns, As a discount factor, For the target critic network, For the target actor network, For the target actor's network parameters, Here, Q represents the target critic network parameters, and Q represents the current critic network. For the current commentator network parameters, for The state at any given moment.

[0095] By adjusting the sampling frequency and preprocessing accuracy of the next time step based on the action strategy, the extraction gain of early latent damage features is maximized under the condition of limited computing resources. When the lifting machinery is under high-load, high-cycle-number high-risk conditions, the agent outputs an action strategy with high sampling frequency and high preprocessing accuracy to improve feature extraction accuracy and enhance the ability to extract early latent damage features. When the lifting machinery is under low-load, low-cycle-number low-risk conditions, the agent outputs an action strategy with low sampling frequency and low preprocessing accuracy to reduce the consumption of computing resources and data transmission volume at the edge, and realize dynamic adaptive scheduling of sampling frequency and preprocessing accuracy.

Claims

1. A method for intelligent inspection of lifting machinery based on a multimodal large model, characterized in that, include: Simultaneously acquire acoustic emission, 3D vision, and distributed fiber optic strain multi-source heterogeneous data of key stress sections and transmission and braking components of lifting machinery; At the edge, multi-source heterogeneous data is preprocessed and preliminary feature extraction is performed. The extracted feature vectors are then de-identified and uploaded to the cloud. The cloud-based multimodal large model performs spatiotemporal feature decoupling and correlation mapping on feature vectors, removes pseudo-features, and extracts damage-sensitive feature vectors; Input the damage-sensitive feature vector into the fracture mechanics fatigue damage evolution numerical simulation program to generate the current damage state and remaining life evolution curve; The evolution curve is input in reverse into the causal inference engine of the multimodal large model to optimize the feature extraction weight parameters to enhance the extraction of early latent damage features, forming a closed-loop coupling driven by both data and mechanism. The model is incrementally updated based on the feature vectors of multiple lifting machines, and the updated model weights are sent to the edge. The damage state and evolution trend are compared with the structural strength verification threshold under rated load conditions to generate compliance judgment results and risk warning signals.

2. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 1, characterized in that, Spatiotemporal feature decoupling and correlation mapping include: Simultaneous compression transformation and wavelet packet decomposition are performed on acoustic emission and vibration data to construct time-frequency domain spectra. The time-frequency domain spectra are mapped to the node attributes in the graph structure data, along with the 3D visual features and strain features. By using graph neural networks to capture the topological relationships of nodes of different modalities in time series, cross-modal neighbor node information is aggregated through graph convolution operations, and isolated subgraphs corresponding to environmental noise and operating condition disturbances are decoupled. Damage-sensitive feature vectors that retain the attributes of core nodes in connected subgraphs are extracted.

3. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 2, characterized in that, The causal reasoning engine employs a counterfactual reasoning mechanism, including: By introducing intervention variables into the feature vector space, a counterfactual feature generator is constructed. The counterfact feature generator performs mask replacement on the working condition covariates in the core node attributes of the connected subgraph to generate counterfact feature vectors without working condition disturbances. Calculate the difference in causal effects between the true feature vector and the counterfactual feature vector on the causal graph, eliminate false causal paths whose difference value is greater than a preset perturbation threshold and have no physical connection with fatigue damage, and output the purified damage-sensitive feature vector.

4. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 3, characterized in that, Optimizing feature extraction weight parameters includes: The partial differential equations of the numerical simulation program for fracture mechanics fatigue damage evolution are transformed into computational graph constraints of a physical information neural network. Using the purified damage-sensitive feature vector as input, the residual of the partial differential equation at the current time step is added as a penalty term to the loss function of the multimodal large model; A model prediction control strategy is adopted. Within a set time prediction window, the weight parameters of the graph convolution operation and the counterfactual feature generator in the multimodal large model are updated on a rolling basis by backpropagating through the loss function that minimizes the physical residual penalty term.

5. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 4, characterized in that, De-identification and uploading to the cloud include: At the edge, a non-orthogonal multiple access power allocation strategy is adopted, and the transmission power is allocated according to the contribution of each mode data in the feature vector to fatigue damage prediction. A Laplace noise mechanism is superimposed on the core feature components allocated high transmission power, and dimension reduction and compression are performed on the redundant feature components allocated low transmission power. The feature vectors with superimposed noise and dimensionality reduction are multiplexed and uploaded to the cloud via a non-orthogonal multiple access channel. The cloud then separates and restores the feature vectors using a continuous interference cancellation algorithm.

6. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 5, characterized in that, Completing incremental model updates includes: After obtaining the feature vectors uploaded by multiple lifting machinery edge devices in the cloud, a ensemble Kalman filter strategy is used for federated aggregation. The weights of the multimodal large model updated locally at each edge are used as state variables, and the consistency between the local model weights and the global damage evolution trend is evaluated by the covariance matrix. Based on the consistency assessment results, dynamic confidence Kalman gain is assigned to the weights of each local model to suppress global model drift caused by local abnormal operating conditions, and updated model weights are generated.

7. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 6, characterized in that, The generation of compliance assessment results and risk warning signals includes: After the updated model weights are distributed to the edge, Monte Carlo simulation is used to randomly sample the remaining lifetime evolution curve within the future time window to generate a probability distribution sequence of damage evolution. A conditional value at risk model is introduced to calculate the expected tail risk of a damage evolution probability distribution sequence exceeding the structural strength verification threshold at a set confidence level. When the expected value of tail risk is greater than zero, a dynamic graded early warning signal characterizing the risk of latent fatigue damage is generated, and the compliance judgment result is output in combination with the current damage status.

8. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 7, characterized in that, Multimodal large models employ a speculative decoding strategy when processing feature vectors. Deploy draft models with small parameters and large multimodal models with large parameters in the cloud for parallel operation; The draft model quickly generates a draft sequence of damage feature evolution paths based on feature vectors uploaded from the edge. The large-parameter multimodal model verifies the damage features of each time step in the draft sequence in parallel during a single forward propagation, retains the draft sequence segments that pass the verification, and only performs recalculation of the causal inference engine for the segments that fail the verification.

9. The intelligent inspection method for lifting machinery based on a multimodal large model according to claim 8, characterized in that, Deep reinforcement learning dynamic scheduling is used for preprocessing and preliminary feature extraction of multi-source heterogeneous data at the edge: Construct an action space that includes a candidate set of sampling frequencies and preprocessing precision for multi-source heterogeneous data, and a state space that uses the test miss rate and edge computing resource consumption as a joint penalty function; The edge agent outputs an action strategy based on the current working condition of the lifting machinery using a deep deterministic policy gradient algorithm. By adjusting the sampling frequency and preprocessing accuracy of the next time step based on the action strategy, the extraction gain of early latent damage features is maximized under the condition of limited computing resources.

10. A lifting machinery intelligent inspection system based on a multimodal large model, characterized in that, To implement the method according to any one of claims 1 to 9, comprising: The sensing and acquisition device is deployed on the key stress-bearing sections and transmission and braking components of the lifting machinery to simultaneously acquire multi-source heterogeneous data of acoustic emission, three-dimensional vision and distributed fiber strain. Edge computing nodes are connected to sensing and acquisition devices to preprocess and extract preliminary features from multi-source heterogeneous data, and to desensitize and upload the extracted feature vectors. The cloud-based causal reasoning and simulation server communicates with edge computing nodes and has a built-in multimodal large model and fracture mechanics fatigue damage evolution numerical simulation program. It is used to perform spatiotemporal decoupling and causal reasoning on feature vectors and forms a closed-loop coupling with the numerical simulation program driven by both data and mechanism. It completes incremental model updates based on the feature vectors of multiple cranes and distributes the updated model weights to the edge computing nodes. The judgment and early warning terminal communicates with the cloud-based causal reasoning and simulation server to compare the damage state and evolution trend with the structural strength verification threshold under rated load conditions, and generate compliance judgment results and risk warning signals.