Defect detection method and system based on joint distribution optimization and structural knowledge guidance
By constructing the device physical topology graph and graph neural network to generate structured feature vectors, combining attention model and joint distribution optimization, the problems of inaccurate feature alignment and poor interpretability in the existing technology are solved, and defect detection and root cause diagnosis with higher accuracy are achieved.
Patent Information
- Application Number
- CN202510900787.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The lack of physical topological constraints in the existing industrial multimodal detection methods leads to inaccurate feature alignment and poor interpretability, making it difficult to effectively mine the intrinsic connections between multimodal data, especially when facing complex industrial scenarios, misjudgment and missed detection are prone to occur.
By constructing the physical topology diagram of the device, using the graph neural network to generate structured feature vectors, and guiding the alignment of perceptual feature vectors in the attention model, performing cross-modal feature fusion, combining joint distribution optimization and Lipschitz stability constraints, ensuring that features are stable aligned in the shared space and conform to physical laws.
Improve the accuracy and interpretability of defect detection, generate a root cause diagnostic report with physical interpretability, significantly improve detection accuracy and reduce missed detection rate.
Smart Images

Figure CN120411083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to defect detection, in particular to a defect detection method and system based on joint distribution optimization and structural knowledge guidance. Background Art
[0002] Industrial vision defect detection is the core technical cornerstone for ensuring the yield of intelligent manufacturing products and production safety. The detection accuracy and intelligent level directly determine the upper limit of quality control in the high-end equipment manufacturing industry. In precision industrial fields such as semiconductor manufacturing, aeroengines, and new energy, equipment defects often present in complex forms spanning multiple physical dimensions. For example, micron-level cracks in circuits may be accompanied by subtle local temperature rises and abnormal fluctuations in operating parameters. Therefore, it is of great significance to develop a defect detection method that can deeply integrate multi-source information, has physical interpretability, and can give early and accurate warnings.
[0003] Currently, research on industrial defect detection mainly focuses on the combined application of deep learning and various sensing technologies. The mainstream solution is first to analyze single-modal visible light images. By introducing object detection frameworks such as YOLO, Faster R-CNN, or image segmentation networks such as U-Net, defects in the images are identified and located. To break through the limitations of single information, some research has begun to explore fusion strategies for multi-modal data. Initial attempts included simply weighted superposition or channel merging of infrared thermal maps and visible light images at the pixel level, in order to supplement temperature information in visual features. Subsequently, methods based on the attention mechanism were proposed. By constructing a feature interaction module, infrared and visible light features are enhanced with each other in the deep network. In addition, some work has tried to use optical character recognition (OCR) technology to extract text reports on equipment maintenance and use them as auxiliary information to classify or corroborate the defect categories analyzed from the images. In recent years, although cross-modal pre-training models represented by CLIP have made significant progress in image-text association, they are mainly oriented to general fields and are not specifically designed for the structured characteristics of specific industrial equipment.
[0004] However, when dealing with complex industrial scenarios, the existing technologies still face some problems, which fundamentally stem from their failure to effectively establish the internal connection between heterogeneous data and the physical entities of devices. For example, there are problems of pseudo-alignment and shallow association in multi-modal features, and there is a lack of a unified and physically interpretable metric criterion in the cross-modal feature space. Specifically, at the level of modal collaboration, traditional feature alignment methods (such as CCA and adversarial training) only focus on statistical distribution matching and ignore the constraint effect of the physical structure of devices on multi-modal association, resulting in the failure to effectively mine the essential associations such as heat maps and key circuit nodes, and text descriptions and three-dimensional model regions; at the level of knowledge fusion, existing methods simplify domain knowledge into rule bases or label extensions and fail to use structured knowledge such as device topology maps as the spatial constraint conditions for feature learning, resulting in the deviation of defect location from physical reality. The root cause lies in the modeling gap between the heterogeneity of industrial multi-modal data (such as the locality of images and the globality of texts) and the graph structure characteristics of domain knowledge, and it is difficult for existing deep learning architectures to uniformly handle the contradiction between non-Euclidean space knowledge representation and Euclidean space feature extraction. Summary of the Invention
[0005] The object of the invention is to provide a defect detection method and system based on joint distribution optimization and structural knowledge guidance to solve the problems of inaccurate feature alignment and poor interpretability caused by the lack of physical topology constraints in existing industrial multi-modal detection methods.
[0006] Technical solution: A defect detection method based on joint distribution optimization and structural knowledge guidance includes: Analyze the device structured document and construct a device physical topology map; Use a graph neural network to process the device physical topology map and generate a structured feature vector; Collect multi-modal perception data from industrial devices and extract perception feature vectors from the multi-modal perception data; In an attention model, use the structured feature vector to guide the alignment of the perception feature vector, perform cross-modal feature fusion, and obtain a fused feature representation; Generate a defect detection result based on the fused feature representation.
[0007] Preferably, in the attention model, the step of using the structured feature vector to guide the alignment of the perception feature vector, perform cross-modal feature fusion, and obtain a fused feature representation includes: Map the structured feature vector into a set of key vectors and a set of value vectors respectively, and the set of key vectors and value vectors are used to describe the structural information of the device physical topology; Transform the perception feature vector into a query vector, and the query vector is used to characterize the real-time state collected from the device; Probe the key vector with the query vector to determine the correlation degree between the perceived feature vector and the device physical topology, and dynamically weight the value vector based on the correlation degree to aggregate and form a fused feature representation.
[0008] Preferably, the optimization steps of the attention model include: Inside the attention model, generate a corresponding modal feature distribution for each of multiple modalities, thereby obtaining a set of modal feature distributions; Calculate the arithmetic mean of all distributions in the set of modal feature distributions to construct a mixed central distribution; Measure the Kullback-Leibler divergence of each modal feature distribution in the set of modal feature distributions tending to the mixed central distribution, and integrate all the measured divergences into a joint distribution divergence loss; Optimize the attention model based on the joint distribution divergence loss to minimize the topological differences between the modal feature distributions.
[0009] Preferably, the optimization steps further include: Obtain the gradient norm of the feature transformation function in the attention model, and based on the gradient norm and a preset Lipschitz constant, obtain a stability constraint loss; Perform a weighted combination of the stability constraint loss and the joint distribution divergence loss to form a combined optimization objective; Finally optimize the attention model according to the combined optimization objective to ensure that the feature transformation function meets the requirements of smoothness and continuity of physical laws.
[0010] Preferably, the step of probing the key vector with the query vector to determine the correlation degree further includes: Generate a spatial constraint matrix according to the device physical topology graph, and the spatial constraint matrix is used to calibrate the physical adjacency relationship allowing associations between device nodes; When calculating the correlation degree between the query vector and the key vector, apply the spatial constraint matrix to mask the invalid associations between physically non-adjacent nodes to obtain the correlation degree.
[0011] Preferably, generating a defect detection result based on the fused feature representation includes: Decode the fused feature representation to generate a defect localization mask indicating the spatial position of the defect; Perform a correlation inference on the features corresponding to the defect positions in the fused feature representation and the node attributes in the device physical topology graph to trace the physical root cause of the defect, and generate a root cause diagnosis report including the confidence of the faulty component; Integrate the defect localization mask and the root cause diagnosis report to form the defect detection result.
[0012] According to another aspect of the present application, there is also provided a defect detection system based on joint distribution optimization and structural knowledge guidance, including: A knowledge parsing module, configured to parse the device structured document to construct a device physical topology diagram; A data acquisition module, configured to collect multi-modal perception data from industrial devices; A feature processing module, communicatively connected to the knowledge parsing module and the data acquisition module, A result generation module, communicatively connected to the feature processing module, for generating a defect detection result based on the fused feature representation; Wherein the feature processing module includes: A graph embedding unit, configured to process the device physical topology diagram to generate a structured feature vector; A perception feature extraction unit, configured to extract a perception feature vector from the multi-modal perception data; A feature fusion unit, built-in with an attention model, configured to use the structured feature vector to guide the alignment of the perception feature vector, perform cross-modal feature fusion, and obtain a fused feature representation.
[0013] Preferably, the feature fusion unit is further configured to: Map the structured feature vector received from the graph embedding unit into a set of key vectors and a set of value vectors respectively, for describing the structural information of the device physical topology; Transform the perception feature vector received from the perception feature extraction unit into a set of query vectors, for characterizing the real-time state collected from the device; Wherein, the built-in attention model is configured to: drive the query vector to explore the key vector to determine the correlation degree between the perception feature vector and the device physical topology, and dynamically weight the value vector based on the correlation degree, and aggregate to form a fused feature representation.
[0014] Preferably, the feature processing module is further configured to be trained through an optimization objective, and the optimization objective includes: For each of multiple modalities, generate a corresponding modality feature distribution inside the feature processing module, so as to obtain a set of modality feature distributions; Construct a mixed central distribution, which is the arithmetic mean of all distributions in the set of modality feature distributions; Integrate the Kullback-Leibler divergence of each modality feature distribution tending to the mixed central distribution to form a joint distribution divergence loss, and optimize based on this loss.
[0015] Preferably, the optimization objective further includes a stability constraint loss, wherein the feature processing module is further configured to: Obtain the gradient norm from the internal feature transformation function, and based on the gradient norm and a preset Lipschitz constant, obtain the stability constraint loss; Perform weighted combination of the stability constraint loss and the joint distribution divergence loss to form a combined optimization objective, and complete the training according to the combined optimization objective.
[0016] Preferably, the feature fusion unit is further configured to: Generate a spatial constraint matrix according to the device physical topology map received from the knowledge parsing module, which is used to calibrate the physical adjacency relationship allowing associations between device nodes; Among them, when its built-in attention model determines the degree of association, the spatial constraint matrix is applied to mask the invalid associations between physically non-adjacent nodes to obtain the degree of association.
[0017] Preferably, the result generation module is further configured to: Decode the fused feature representation received from the feature processing module to generate a defect localization mask indicating the spatial position of the defect; Perform the association inference between the fused feature representation and the device physical topology map to trace the physical root cause of the defect, and generate a root cause diagnosis report including the confidence of the faulty component; Integrate the defect localization mask and the root cause diagnosis report to obtain the defect detection result.
[0018] Beneficial effects: Using the device physical topology map as a knowledge prior, by constructing an attention mechanism with structured features as keys (Key) and values (Value) and perceptual features as queries (Query), the pseudo-alignment problem is solved. Semantic-level alignment is performed under physical constraints to associate, for example, thermal anomaly points with specific circuit nodes, improving the accuracy of defect localization. By introducing a joint distribution optimization objective based on Jensen-Shannon divergence and Lipschitz stability constraints, the problems of metric unification and physical authenticity in the feature space are solved. This combined optimization objective ensures that heterogeneous modal features can be stably and smoothly mapped and aligned in the shared space, making the model get rid of the black box characteristics and enhancing the robustness. It can significantly improve the accuracy of defect detection and generate a root cause diagnosis report with physical interpretability. Description of the Drawings
[0019] Figure 1 is the flowchart of the present invention.
[0020] Figure 2 is the flowchart of the present invention for obtaining the fused feature representation.
[0021] Figure 3 is the flowchart of the present invention for optimizing the attention model.
[0022] Figure 4 It is a flowchart of another embodiment of the optimized attention model of the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following will combine Figures 1 to 4 to describe the present invention in further detail. It should be noted that the specific embodiments herein are only used to explain the present invention, rather than limiting the protection scope of the present invention. The following embodiments and the technical features therein can be combined with each other without conflict.
[0024] The applicant has found that existing fusion methods, whether based on statistical correlation (such as CCA) or distribution alignment based on adversarial training, essentially stay at the distribution matching at the data level, while ignoring the hard constraints of the inherent physical topology of the device on the association of multimodal information. For example, the physical root cause of an abnormal high-temperature area on an infrared thermal map must correspond to one or several specific electronic components on the circuit topology diagram. Due to the lack of guidance of such topological knowledge, existing methods may wrongly associate the thermal anomaly with physically unrelated adjacent areas, resulting in a superficial combination of feature fusion. This alignment method cannot penetrate into the physical semantic level, causing the understanding and positioning of defects to deviate from physical reality. Especially when facing tiny defects or coupled faults, it is extremely easy to produce misjudgments and missed detections.
[0025] In addition, industrial multimodal data (such as the local pixel grid of an image, the global abstract semantics of text, and the non-Euclidean graph structure of device topology) has natural heterogeneity. When they are forcibly projected into the same feature space, traditional metric methods such as Euclidean distance often fail. More critically, existing deep learning fusion models are like black boxes, and the internal feature transformation process lacks the constraints of physical laws. For example, the model may learn a non-smooth mapping, such that a small and continuous change in the input temperature causes a drastic and discontinuous jump in the feature space, which completely violates the physical reality of heat conduction. This feature representation lacking physical stability and interpretability not only reduces the generalization ability and robustness of the model, but also makes its diagnostic results difficult to be trusted and adopted by domain experts, limiting its practical application in safety-critical industrial scenarios. Therefore, the following method is constructed.
[0026] Embodiment 1. Describe the construction and operation process of a defect detection system guided by structural knowledge. Taking a general industrial device as an object, the processing flow of the present invention is elaborated. Specifically, it includes the following steps: Step S100, parse the device structured document to construct a device physical topology diagram; and use a graph neural network to process the device physical topology diagram to generate a structured feature vector.
[0027] In this embodiment, the step is used to convert the prior knowledge representing the physical structure of the device into digital features that can be processed by a deep learning model.
[0028] In this embodiment, the device structured document refers to a standardized engineering document that can describe the device components and their interrelationships, such as computer-aided design models, circuit schematics, bills of materials (BOMs), etc.
[0029] The physical topology diagram (i.e., the device physical topology diagram) is a mathematical graph structure, where the nodes represent the physical components of the device, and the edges represent the physical or functional connections between the components.
[0030] The structured feature vector is a low-dimensional dense vector generated by a graph neural network and including the topological graph structure information.
[0031] In some embodiments, the process is as follows: Obtain the structured documents of the device to be detected from the device manufacturer or the operation and maintenance party. Automatically or semi-automatically parse these documents through software tools, extract the key components as nodes, extract the connection relationships between the components as edges, and construct the device physical topology diagram.
[0032] Adopt a graph neural network (GNN), such as a graph convolutional network (GCN) or a graph attention network (GAT), to process the constructed physical topology diagram. The GNN generates a structured feature vector for each node in the graph through layer-by-layer aggregation of the information of neighboring nodes, which can characterize its own attributes and its position in the entire topology.
[0033] Since the physical structure of the device is an objectively existing and extremely valuable prior knowledge, which determines the physical laws of fault occurrence and conduction. Converting this structured knowledge in non-Euclidean space into vector form through the GNN is a prerequisite for enabling the subsequent neural network to understand and utilize these physical laws.
[0034] Step S200: Collect multi-modal perception data from industrial devices and extract perception feature vectors from the multi-modal perception data.
[0035] In this embodiment, the purpose of this step is to obtain the raw data representing the real-time operating state of the device from the external world and extract its high-level semantic features. The multi-modal perception data refers to data with different sources or types but jointly describing the same detection object. For example, visible light images, infrared thermal maps, operation log texts, etc. The perception feature vector refers to a high-dimensional abstract vector obtained after the raw perception data is processed by a feature extraction network and can represent the core content of the data.
[0036] In this embodiment, data is synchronously collected through various sensors (such as industrial cameras, thermal imagers) deployed at the industrial site, and text data is obtained through technologies such as optical character recognition (OCR).
[0037] For different types of data, corresponding deep learning feature extractors are adopted. For example, a pre-trained convolutional neural network (CNN), such as ResNet, is used to extract visual perception features of images and heatmaps; a pre-trained language model, such as BERT, is used to extract semantic perception features of text records.
[0038] In some embodiments, different feature extraction networks can be selected according to specific detection tasks and data characteristics. For example, EfficientNet for high-resolution images or long short-term memory network (LSTM) for sequence data.
[0039] In step S300, in the attention model, the structured feature vector is used to guide the alignment of the perception feature vector for cross-modal feature fusion, thereby producing a fused feature representation.
[0040] In this embodiment, this step is the core bridge connecting physical knowledge and perception information, which realizes a knowledge-guided feature fusion mechanism.
[0041] The attention model can dynamically and selectively focus on the most relevant parts of the input information. In the present invention, it specifically refers to a fusion model with the Transformer architecture as the core.
[0042] The fused feature representation is used to effectively combine the feature vectors from multiple modalities and the structured feature vector to produce a unified and more informative feature vector.
[0043] The structured feature vector generated in step S100 and the perception feature vector generated in step S200 are jointly input into the attention model based on Transformer.
[0044] Inside this model, through its core cross-attention mechanism, the structured feature vector is used as a knowledge reference system providing context background, and the perception feature vector is weighted and aggregated under the guidance of this reference system. In this way, when the model fuses perception features, it will preferentially focus on the relevance that conforms to the physical structure.
[0045] This step aims to solve the blindness of traditional fusion methods. Instead of simply splicing or averaging features from different sources, it establishes a fusion mode that follows a logical pattern. Perception information (for example, an abnormal hot spot) must find its most reasonable attribution under the guidance of physical topology knowledge, thus ensuring the logic and physical interpretability of the fusion process.
[0046] In step S400, based on the fused feature representation, a defect detection result is generated.
[0047] This step in this embodiment is used to convert the internal and abstract fusion feature representations of the model into specific and actionable detection conclusions for end-users.
[0048] The defect detection result refers to the final judgment on whether there are defects in the device, where the defects are located, and what types they belong to. Its form can be classification labels, bounding boxes, segmentation masks, etc.
[0049] Feed the fusion feature representation output by step S300 into one or more TaskHead networks.
[0050] For example, a classification head (usually composed of fully connected layers) can be used to judge the category of the defect; a segmentation head (usually composed of upsampling and convolutional layers) can be used to generate a pixel-level defect localization mask on the original image.
[0051] Since the input fusion feature representation already contains physical topology knowledge and multi-modal perception information, the defect detection results decoded from it have improved accuracy and reliability compared to single-modal or traditional fusion methods.
[0052] Embodiment 2: In this embodiment, the electro-luminescence (EL) image defect detection in the production process of photovoltaic (PV) modules is used as a specific application scenario to detail the core fusion architecture. Based on the overall framework disclosed in the above Embodiment 1 (i.e., the best embodiment), this embodiment focuses on how to use the circuit topology knowledge of photovoltaic modules to guide the deep fusion of EL image and infrared thermal image features to achieve accurate identification of defects such as dark spots and virtual soldering.
[0053] Refer to Figure 1 , the method disclosed in this embodiment specifically includes the following steps: Step S100, parse the device structured document to construct a device physical topology graph; and use a graph neural network to process the device physical topology graph to generate a structured feature vector.
[0054] The main purpose of this step is to convert the electrical connection relationships inside the photovoltaic module into a graph data structure and extract its topological features. In this embodiment, the device physical topology graph is the circuit connection graph of the photovoltaic module. Among them, each solar cell is regarded as a node, and the series or parallel relationships between the cells are represented by edges. The attributes of the nodes can include their coordinates, nominal power, material batch, etc. in the module. The attribute of the edge can be the connection resistance value. The knowledge source is the layout design file LayoutFile and the electrical wiring diagram of this model of photovoltaic module. The process of graph construction is to automatically identify the positions of each cell and create nodes by parsing the layout design file. By parsing the electrical wiring diagram, determine the connection relationships between the cells and between the cells and the busbars, and create the corresponding edges, thereby constructing the physical topology graph G representing the electrical topology of the entire module. The process of knowledge embedding is to process the graph G using the graph attention network GAT to generate a 512-dimensional structured feature vector for each cell node.
[0055] Step S200, collect multi-modal perception data from industrial devices and extract perception feature vectors from the multi-modal perception data.
[0056] In this embodiment, the multi-modal perception data is mainly electroluminescence EL images and infrared IR thermograms.
[0057] On the production line, apply a forward bias voltage to the photovoltaic module and use a dedicated near-infrared camera to capture its EL image IEL, which can intuitively reflect the internal defects (such as dark spots, black hearts, cracks) of the cells. At the same time, use an infrared thermal imager to capture the temperature distribution map TIR of the module in the working state.
[0058] Extract EL image features. Considering that the texture features of the EL image are relatively simple, a lightweight convolutional neural network CNN, such as MobileNetV3 or a customized shallow CNN, can be used to process the EL image and extract its perception feature vector.
[0059] Extract infrared image features. Use a lightweight CNN to process the infrared thermogram and extract its perception feature vector related to the temperature distribution.
[0060] Step S300, in an attention model, use the structured feature vector to guide the alignment of the perception feature vector and perform cross-modal feature fusion, thereby producing a fused feature representation.
[0061] Through a novel attention mechanism configuration, let the visual perception information (EL and IR images) be fused under the guidance of the circuit topology knowledge.
[0062] Construct a Transformer fusion unit. Its core cross-attention mechanism is configured as: Use the photovoltaic module structured feature vector generated in step S100 (representing the position and connection relationship of each cell in the circuit) as the key (Key, K) and value (Value, V) of the attention mechanism.
[0063] After concatenating or adding the EL image perception features and IR image perception features extracted in step S200, use them as the query (Query, Q).
[0064] The motivation for this configuration is that the dark spots (defects) on the EL image are closely related to their positions in the circuit in terms of the causes and severity of their generation. For example, the failure of a cell at the end of a branch and the failure of a cell in the middle of a branch have different impacts on the performance of the entire module.
[0065] By having the Query representing the image features query the Key representing the circuit topology, the model can calculate the highest degree of association between the features of a certain region in the image and which specific cell node in the circuit. For example, when the Query comes from the dark spot region in the image, the attention mechanism will be guided to the Key of the cell node corresponding to this region and obtain the Value of this node (i.e., its structured feature vector).
[0066] In this way, the obtained fused feature representation contains both information about what the defect looks like (from Q) and where the defect is in the circuit (from K and V), achieving semantic-level alignment rather than simple pixel superposition.
[0067] Step S400, generate a defect detection result based on the fused feature representation. In this embodiment, the final defect detection result not only marks the defect location but also can associate its circuit attributes.
[0068] Input the fused feature representation generated in step S300 into a fully connected classifier and regressor.
[0069] The classifier outputs the defect type (e.g., cell fragmentation, solder joint looseness, bypass diode failure).
[0070] The regressor outputs the precise bounding box of the defect.
[0071] When the system detects that a certain cell has an electroluminescence dark spot, it first locates the circuit node corresponding to this region in the three-dimensional model of the module (for example, belonging to series branch 3, the 5th cell), synchronously analyzes the thermal imaging data (no obvious abnormality in temperature) and EL image features (presence of an irregular dark area) at this point. The knowledge graph automatically associates the circuit characteristics at this position, and after the Transformer architecture fuses this information, it not only marks the defect region but also outputs a diagnostic conclusion: there is a microcrack in cell C3-5, with a confidence level of 96%.
[0072] Through the above steps, this embodiment can effectively distinguish whether it is a problem of a single cell or a series problem caused by solder tapes or busbars, greatly improving the accuracy and automation level of photovoltaic module defect detection.
[0073] Embodiment 3: Taking the multi-modal non-destructive testing of aero-engine turbine blades as the background, a complete optimization process for training a detection model is described in detail.
[0074] Steps S100 - S200, knowledge graph construction and multi-modal feature extraction. In this embodiment, first, steps similar to those in the foregoing embodiments are executed.
[0075] Construct a knowledge graph: Based on the finite element analysis (FEA) model and 3D CAD file of the blade, construct a topology graph. Among them, the nodes represent the finite element mesh units on the blade, and the weight of the edge can represent the stress conduction relationship between the units. The node attributes can include the design stress and temperature field distribution at this position, etc.
[0076] Extract multi-modal data and features: Synchronously collect the X-ray imaging, visible light surface coating image of the blade, and the data of the main shaft vibration sensor during operation. Use corresponding feature extractors (such as CNN for processing images and time series networks for processing vibration signals) to extract their respective perceptual feature vectors.
[0077] Step S300, (training stage) Optimize the attention model using a combined optimization objective.
[0078] In this embodiment, the detailed systematic optimization process is the key for the model to learn a feature representation that is both unified and conforms to physical laws. The combined optimization objective refers to a total optimization objective formed by the weighted combination of multiple loss functions with different purposes. In the present invention, it drives the model parameters to be updated in the direction of small reconstruction error, consistent distribution among modalities, and stable transformation.
[0079] In each training iteration, the forward propagation process of the model is similar to that in Embodiment 2 to obtain the feature representations of each modality. Subsequently, through the combined optimization objective min θ L recon + λ1L JSD + λ2L Lip to calculate the total loss and perform backpropagation.
[0080] Calculate the joint distribution divergence loss LJSD: First, construct the feature distributions F X , F V , F Vib of the three modalities of X-ray, visible light, and vibration signals into a set of modal feature distributions inside the model.
[0081] Then, calculate the arithmetic mean of all the distributions within the set to construct the mixed central distribution F* = (1 / 3)(F X + F V + F Vib ).
[0082] Finally, measure the KL divergence of each modal feature distribution tending towards this mixed central distribution, and integrate (e.g., average) all the KL divergence values to obtain the joint distribution divergence loss LJSD. The role of this loss term is, at the information theory level, to pull heterogeneous data from different physical sources towards a unified and shared semantic space, such that, for example, a tiny crack on an X-ray image can correspond to a specific harmonic in a vibration signal at the feature level.
[0083] Calculate the stability constraint loss L Lip : Identify the key feature transformation function in the recognition model, such as the mapping function f(•) from the original X-ray image to its feature vector.
[0084] During the training process, through techniques such as spectral normalization, constrain the gradient norm ||gradxf(x)|| 2 of this function f(•) not to exceed a preset Lipschitz constant L (e.g., L = 1.5).
[0085] Derive the stability constraint loss L Lip based on the difference between this gradient norm and L. The motivation for this loss term is that the crack on the blade is physically continuous, and a tiny movement of its position on the X-ray image should not cause a drastic and irrelevant jump in its feature representation. Imposing the Lipschitz constraint injects this prior of physical continuity into the model's learning process, ensuring the stability of the model and its robustness to tiny changes.
[0086] Weightedly combine the above two core loss terms with the standard reconstruction loss term L recon to form the final combined optimization objective, and update the parameters θ of the entire network using an optimizer such as Adam according to this objective.
[0087] Through the above training steps, the model obtained in this embodiment can not only accurately fuse multi-source data to locate the micro-cracks on the blade, but more importantly, its internal feature representation is highly ordered and conforms to physical intuition, thus enabling a more reliable prediction of the future expansion path of the crack, reducing the missed detection rate to below 0.3%, and meeting the stringent requirements of aero-engine predictive maintenance.
[0088] In this embodiment, the steps to optimize the attention model specifically include: Inside the attention model, a corresponding modal feature distribution is generated for each of multiple modalities. Preset an ordered circular alignment path among multiple modal feature distributions. Along the circular alignment path, sequentially measure the Kullback-Leibler divergence of the latter modal feature distribution relative to the former modal feature distribution, and accumulate all the measured divergences to form a chain alignment loss. Optimize the attention model based on the chain alignment loss to drive the local and global consistency of each modal feature distribution on the path.
[0089] Specifically, the ordered circular alignment path refers to a one-way, closed-loop alignment sequence artificially predefined among modal feature distributions. For example, for three modalities of vision V, thermal H, and text T, a valid alignment path can be V→H→T→V. The chain alignment loss refers to the total loss obtained by accumulating the KL divergence values between all adjacent modalities on the circular alignment path.
[0090] Before the start of model training, preset an alignment path according to domain knowledge or experiments. For example, in circuit breaker diagnosis, it is generally considered that visible light anomalies are usually the surface causes of thermal anomalies, and text records are the summaries of the phenomena. Therefore, an alignment path of V→H→T→V can be defined. This means that engineers expect: the thermal feature distribution FH should be aligned with the visual feature distribution FV; the text feature distribution FT should be aligned with the thermal feature distribution FH; at the same time, to ensure the global closed loop, the visual feature distribution FV should also be aligned with the text feature distribution FT.
[0091] In each training iteration, the loss function is calculated as: L chain =KL(F H ||F V )+KL(F T ||F H )+KL(F V ||F T ); Take this chain alignment loss L chain (which can be combined with other loss terms such as reconstruction loss, Lipschitz loss, etc.) as the total optimization objective, and perform backpropagation and update on the parameters of the entire attention model.
[0092] This method establishes an asymmetric and ordered dependence structure for the alignment relationship between modalities, which may be more in line with physical reality in some industrial scenarios. For example, the evolution of physical events itself has a sequence.
[0093] Through this one - by - one chained constraint, not only is the local feature consistency between adjacent modalities guaranteed, but also through the final closed - loop constraint, it is ensured that the entire feature space will not have information drift due to the overly long chain, achieving global topological consistency.
[0094] Compared with the center - radiating alignment, this ordered circular alignment method provides different inductive biases for model optimization and may obtain better convergence characteristics and performance on specific tasks.
[0095] In some embodiments, the alignment path can be non - circular. For example, in a scenario where the dominant modality is very clear, a linear alignment path with a start and an end can be defined (such as V→H→T). In this case, the last term in the loss function that connects the start and the end will not be included.
[0096] The order of the path can be adjusted and optimized as a hyperparameter. For example, two different path orders, V→H→T→V and V→T→H→V, may bring different model performances, and the better one can be selected according to the experimental results.
[0097] Embodiment 4: Taking the multi - modal intelligent diagnosis of high - voltage circuit breakers in substations in the power industry as a specific application scenario, the complete implementation process of the defect detection method and its system based on joint distribution optimization and structural knowledge guidance disclosed by the present invention is elaborated in detail. The aim is to achieve accurate positioning and interpretable root - cause tracing of early faults of the circuit breaker (such as internal insulation deterioration, minor deformation of the mechanical structure, etc.). The specific steps are as follows: [[ID=?]] [[ID=?]]
[0098] In this embodiment, this step aims to transform the domain knowledge describing the physical structure and electrical connection relationship of the circuit breaker into a graph - structured data that can be understood and processed by a computer, and use a graph neural network (GNN) to extract its deep - structure features as the knowledge skeleton for the subsequent fusion process.
[0099] In this embodiment, the device physical topology graph is a graph G=(V,E) composed of nodes (V) and edges (E), which is used to mathematically represent the physical composition of the device and the connection relationship between components. Among them, the nodes represent device components (such as main contacts, insulators), and the edges represent the physical or electrical connections between them.
[0100] It should be noted that there seem to be some incorrect or incomplete tag numbers in the original text (e.g., the "?" in the tag numbers in the middle part). This translation is based on the content that can be clearly understood.Knowledge extraction is performed based on the 3D CAD assembly model (such as a STEP format file) of this type of high-voltage circuit breaker and the primary and secondary electrical wiring schematic diagrams (such as DWG format files). Specifically, by parsing the CAD model and using the octree space partitioning algorithm, the circuit breaker is disassembled into hierarchical components, with each component serving as a node, and its attributes including 3D spatial coordinates, materials, etc. At the same time, by parsing the electrical wiring diagram, electrical components are extracted as nodes, and the connection relationships are used as edges, with their attributes including rated voltage, current, etc. Finally, the nodes and edges from the two sources are fused to construct a unified physical topology graph G.
[0101] A 4-layer graph attention network (GAT) is used to process the constructed physical topology graph G. At the l-th layer, the feature h_i^{(l + 1)} of node i is aggregated by weighting the features of its neighbor nodes N_i. Preferably, the dimension of the output structured feature vector is set to 512 dimensions to match the processing requirements of the subsequent attention model.
[0102] Traditional deep learning models cannot directly process the topological knowledge of devices in non-Euclidean spaces. Constructing these valuable prior knowledge into a knowledge graph and using GAT to embed it into the vector space is the key bridge for injecting physical constraints and domain knowledge into the model. Through the attention mechanism, GAT can adaptively learn the importance of different neighbor nodes, making the generated structured feature vector not only contain the attributes of the node itself, but also aggregate its context information in the entire device topology structure.
[0103] In step S402, multi-modal perception data is collected from industrial devices, and perception feature vectors are extracted from the multi-modal perception data. Synchronously obtain multi-source and heterogeneous status monitoring data of the circuit breaker through a sensor array, perform standardized preprocessing, and use a deep neural network designed according to the characteristics of different modal data to extract high-level semantic features.
[0104] In this embodiment, the multi-modal perception data specifically includes: visible light inspection images, infrared thermal imaging videos, device operation logs, and historical maintenance record texts. [[ID=IS=14]]
[0105] During operation, visible light images I_raw and thermal imaging videos T_raw of the key components of the circuit breaker are synchronously collected through a high-resolution industrial camera and a high-precision infrared thermal imager. Parse and associate the corresponding text documents through an OCR engine.
[0106] The collected raw data is corrected, calibrated, and spatio-temporally aligned to eliminate sensor noise and environmental interference, forming a regular data triple {I_t, T_t, Text_t}, which is used as the input to the feature extractor.
[0107] Extract visual features: Use an improved ResNet-50 network to process the visible light image Icorrected. To enhance the ability to capture irregular-shaped defects such as surface micro-cracks, a deformable convolutional layer is inserted after the third residual block of ResNet-50, and its calculation process is y(p)=∑ k=1 K w k •x(p+p k +Δp k ).
[0108] y(p) is the output feature at position p. Wk is the weight of the k-th sampling point in the deformable convolution. P k is the predefined sampling offset. Δp k is the additional offset learned by the network.
[0109] Extract thermal features, including: Using the discrete wavelet transform DWT, selecting the db4 wavelet basis, decomposing the calibrated thermal map Tcalib into coefficients of different frequency bands to separate the low-frequency components representing temperature steady state and the high-frequency components representing transient anomalies, which together constitute the thermal perception features.
[0110] Extract text features, specifically: Use the RoBERTa model to process the aligned text records, and extract text perception features containing information such as fault descriptions and operation records.
[0111] In some embodiments, other advanced network architectures such as EfficientNet and SwinTransformer can also be selected for the visual feature extractor. The text feature extractor can also be replaced with BERT or other variants thereof. More data augmentation strategies such as random rotation and color jitter can also be introduced in the preprocessing stage to improve the generalization ability of the model.
[0112] Step S403, in the attention model, use the structured feature vector to guide the alignment of the perception feature vector, perform cross-modal feature fusion, and thus produce a fused feature representation.
[0113] In this step, a knowledge-enhanced attention architecture is constructed, and the extracted structured knowledge is used to guide the deep fusion of the multi-modal perception features extracted in step S402.
[0114] Knowledge-enhanced cross-attention architecture: Construct a multi-head (preferably, the number of heads h = 8) Transformer architecture as the feature fusion unit. The core cross-attention calculation process is configured as follows: Use the structured feature vector generated in step S401 as the key (Key, K) and value (Value, V) of the attention mechanism.
[0115] Use the perceptual feature vector generated in step S402 as the query (Q).
[0116] The attention calculation formula is CrossAttention(Q, K, V) = softmax(QK T / sqrt(d k ) + M mask )V.
[0117] d k is the dimension of the key (K) vector; M mask is the spatial constraint matrix (attention mask) generated based on the physical topology.
[0118] The spatial constraint matrix Mmask is generated according to the physical topology graph in step S401. If nodes i and j in the topology graph are not physically adjacent or there is no direct electrical connection, the value at the corresponding position in the matrix is set to a very large negative number. This can force the model to mask out physically unreasonable associations when calculating attention.
[0119] Optimization objective for model training: The combined optimization objective function for training this attention model is min θ L recon + λ1L JSD + λ2L Lip . λ1 and λ2 are weight hyperparameters used to balance different loss terms. θ is all the trainable parameters of the entire model. L is a preset Lipschitz constant.
[0120] Among them, the joint distribution divergence loss LJSD is obtained by calculating the KL divergence between the feature distribution of each modality and the mixed central distribution of all modalities, aiming to drive the features of different modalities to be semantically aligned.
[0121] The stability constraint loss LLip is obtained by imposing a Lipschitz constraint on the feature transformation function in the model, aiming to ensure that the computational process inside the model conforms to the smoothness and continuity of the physical world.
[0122] Preferably, the weight coefficients are set as λ1 = 0.7 and λ2 = 0.3.
[0123] To solve the two major problems of pseudo-alignment and lack of unified physical metrics in the existing technology. K / V / Q directly materializes the prior knowledge that perceptual information should query the physical structure into the network architecture, achieving in-depth guidance of knowledge. The joint optimization objective provides a solid mathematical theory guarantee for fusion from the perspectives of information theory and differential geometry.
[0124] Step S404, generate a defect detection result based on the fused feature representation.
[0125] In this embodiment, this step is used to convert the fused feature representation output by the model, which contains rich information, into a diagnostic result that can be directly understood and used by users, including specific locations and possible causes.
[0126] Design a differentiable clustering head at the end of the model to decode the final fused feature representation into a pixel-level defect probability map Mdefect, highlighting potential defect areas to form a defect localization mask.
[0127] Extract the fused feature vector hdj corresponding to the defect area.
[0128] Calculate the cosine similarity between this defect feature vector and the embedding vectors hek of all possible faulty components in the knowledge graph.
[0129] Convert the similarity score into a posterior probability distribution P(ei∣dj) through the Softmax function. This probability represents the likelihood that the root cause is component ei under the condition that defect dj is detected.
[0130] Finally, output a structured root cause diagnosis report.
[0131] For example, in the application scenario of a substation circuit breaker, when an abnormality is detected in the bushing of phase A, the defect detection results output by the system may be as follows: Defect localization mask: Generate a red highlighted area at the lower skirt position of the bushing of phase A in the circuit breaker image.
[0132] Root cause diagnosis report: Discharge traces are detected in the bushing of phase A, and the associated heat map shows that the temperature in this area has abnormally increased by 1.2°C. Root cause tracing result: Internal insulation deterioration leads to partial discharge, with a confidence level of 92%; Associated historical text records: The acetylene content in the oil chromatography of this equipment has exceeded the standard.
[0133] Since the feature representation of the present invention is deeply bound to the physical topology map from the beginning of learning, it contains physical semantics, making it possible to trace the root cause by translating abstract features back to specific physical components. <X
[0134] Through the above steps, the solution disclosed in this embodiment has successfully achieved multi-modal and interpretable fault diagnosis for complex industrial equipment. Test data shows that compared with the method that only uses visual information, the present invention has improved the comprehensive detection F1-score by about 27.6%. At the same time, due to the effective guidance of knowledge, the number of training samples required for the model to converge has been reduced by about 83%. Its diagnostic results are more likely to be adopted by domain experts because of their physical interpretability, demonstrating great industrial application value.
[0135] In another embodiment of the present application, the nodes of the device physical topology map include at least one of electrical parameters and material properties; The generation process of the structured feature vector includes: using a graph attention network to process the physical topology graph, where the attention coefficient between nodes depends on both the node feature similarity and the physical distance between nodes.
[0136] In another embodiment of the present application, the step of performing cross-modal feature fusion further includes: Generating a spatial constraint matrix based on the device physical topology graph; when calculating the correlation degree between the query and the key, applying the spatial constraint matrix to mask the invalid correlations between physically non-adjacent nodes.
[0137] In another embodiment of the present application, the process of calculating the correlation degree is specifically: generating a spatial constraint matrix based on the device physical topology graph, and the spatial constraint matrix is used to calibrate the physical adjacency relationship that allows correlations to occur between device nodes; When calculating the correlation degree between the query vector and the key vector, applying the spatial constraint matrix to mask the invalid correlations between physically non-adjacent nodes to obtain the correlation degree.
[0138] In another embodiment of the present application, the step of driving the query vector to explore the key vector to determine the correlation degree further includes: Generating a physical attenuation attention matrix based on the physical attributes between nodes in the device physical topology graph, where each element value in the physical attenuation attention matrix is a continuous value between 0 and 1, and is used to characterize the attenuation degree of the mutual influence between the corresponding node pairs; After calculating the original correlation degree score between the query vector and the key vector, modulating the score by multiplying the original correlation degree score element-wise with the physical attenuation attention matrix to obtain the final correlation degree.
[0139] In this embodiment, a soft-constrained physical attenuation attention mechanism is adopted to more finely incorporate physical prior knowledge into the attention calculation.
[0140] The physical attenuation attention matrix is a matrix with the same dimension as the attention score matrix, and its element (Matten)ij is a continuous value in the range of [0,1], and this value is inversely proportional to the attenuation degree of the physical influence between nodes i and j. The closer the value is to 1, the smaller the influence attenuation; the closer the value is to 0, the greater the influence attenuation.
[0141] Element-wise multiplication refers to multiplying the corresponding elements of two matrices with the same dimension to obtain a new matrix with the same dimension, that is, the Hadamard product.
[0142] The generation process of the attenuation matrix is as follows: Extract the physical attributes between all node pairs (i,j) from the physical topology graph. In this embodiment, it is preferably to use the spatial distance dij of the nodes in the 3D CAD model.
[0143] Calculate each element value of the physical attenuation attention matrix through a preset Gaussian attenuation function: (Matten)ij = exp(-d ij 2 / (2σ 2 )); where σ is a hyperparameter that controls the attenuation range. When two nodes physically coincide (dij = 0), the attenuation value is 1; as the distance increases, the attenuation value smoothly approaches 0.
[0144] The attention modulation process is as follows: First, according to the standard attention calculation process, obtain the original, unconstrained correlation score matrix S = QK T / sqrt(d k ). Then, multiply the original score matrix S element-wise with the physical attenuation attention matrix Matten we generated to obtain the modulated score matrix S′ = M atten OS. O represents the Hadamard product. Finally, send the modulated score matrix S′ into the Softmax function to calculate the final attention weights.
[0145] This method is more in line with physical reality. Since the influence changes continuously, allowing the model to pay attention to nodes that are not directly adjacent but still have weak physical connections to a certain extent is crucial for capturing some complex cascading failures that require considering the global stress field or temperature field.
[0146] By changing the constraint method from adding a very large negative number to multiplying by an attenuation coefficient between 0 and 1, the attention modulation process becomes smoother and differentiable, which may be more conducive to the gradient propagation and stable convergence of the model.
[0147] In some embodiments for electrical equipment, the attenuation degree can be calculated based on the equivalent resistance or impedance between nodes. For example, (M atten ) ij = 1 / (1 + R ij ), where R ij is the equivalent resistance between nodes i and j.
[0148] The form of the attenuation function can also be other functions, such as an exponential attenuation function or an inverse proportional function, as long as it can reflect the trend of the influence decaying with the change of physical properties (such as distance, resistance).
[0149] Example Five: Describe the data processing process of another embodiment.
[0150] Step 1: Multimodal data collection and structured preprocessing This step realizes the synchronous acquisition and spatio-temporal alignment of multi-source data through an industrial-grade sensor array. The visible light vision channel uses a high-resolution camera equipped with a telecentric lens, which acquires surface images at a rate of 15 frames per second under the control of a trigger signal, and the pixel resolution needs to meet Δx ≤ 5μm / pixel. The thermodynamics channel uses a FLIR A8580 infrared thermal imager, and the acquisition frequency is strictly synchronized with the visible light frame rate, and the temperature measurement accuracy needs to reach ±0.3°C. The text modal data uses an OCR engine to parse the device operation log in real time and extract the maintenance record keyword fields containing timestamps. In the data preprocessing stage, first, perform dark current correction and non-linear illumination compensation on the RGB image: I corrected =(I raw -B) / GOM illu , where B is the dark field reference, G is the gain matrix, M illu is the light intensity distribution mask, and O is the Hadamard product. The thermal data is unified to the same spatial resolution as the visual data through bicubic interpolation, and temperature drift compensation is performed: T calib =T raw +α(t)•Δt, where α(t) is the environmental temperature change rate. After the text data is extracted with word vectors by the BERT model, a mapping relationship is established with the image sequence through the time alignment module. The multi-modal tensor output by this step needs to meet the normalization constraint of |I_corrected|_L2 ∈ [0.2, 0.8].
[0151] Step 2: Device Topological Knowledge Graph Construction and Embedding Based on the device CAD model and process documents, construct a knowledge graph containing physical connection relationships and functional semantics. Use the AutoCAD Electrical plugin to parse the electrical schematic diagram and generate a topological graph G=(V,E) whose node attributes include position coordinates, rated current, and material conductivity, where the edge weight w ij ∈E represents the connector impedance value. The three-dimensional mechanical structure is imported through a STEP file, and an octree space segmentation algorithm is used to establish a hierarchical assembly relationship. The knowledge embedding module uses a two-layer graph attention network: h i l+1 =σ(∑ j∈N(i)) α ij W l h j l ), and the attention coefficient α ij is jointly determined by the node spacing d ij and the functional similarity: α ij =(exp(LeakyReLU(a T [Wh i ||Wh j ))) / (∑ k∈N(i)exp(LeakyReLU(a T [Wh i ||Wh k )))•e^(-βd ij ), the output dimension is set to 512 dimensions to match the subsequent Transformer architecture, and the best topological coverage is achieved when the number of layers L = 4. The knowledge graph update cycle is synchronized with the equipment maintenance records to ensure that the node attributes reflect the latest status.
[0152] Step 3: Deep extraction and representation learning of multimodal features In this step, a heterogeneous network is used to implement the feature encoding of each modality. The visual branch uses an improved ResNet-50 architecture, and a deformable convolutional layer is inserted after the third residual block to adapt to small defect deformations: y(p)=∑ k=1 K w k •x(p + p k +Δp k ), Δp k is predicted by the offset network. The thermal feature extraction adopts a frequency-domain decomposition method, and the steady-state and transient components are separated by wavelet transform: T feature =DWT(T calib ,ψ db4 )•W thermal . After the text modality extracts word vectors using the RoBERTa model, the key descriptive information is aggregated through temporal convolution. The output dimension of all features is unified to d = 256, and the distribution is standardized through LayerNorm. The feature extractor adopts a contrastive learning strategy in the pre-training stage, constructing positive sample pairs (x i ,x j +) and negative sample pairs (x i ,x j -), and the optimization objective is: L cont =-log[(exp(s(x i ,x j + ) / τ)) / (∑ k=1 N exp(s(x i ,x k - ) / τ))], and the best modality discrimination is achieved when the temperature coefficient τ = 0.07.
[0153] Step 4: Knowledge-guided cross-modal attention alignment Construct a Transformer-based multimodal interaction architecture, the core of which is a knowledge-enhanced cross-attention mechanism. The knowledge features H output by the GCN kUsing the keys and values, and each modal feature F_m as the query, calculate the hierarchical attention weight CrossAttention(Q,F,K,H)=softmax((Q(F)K(H) T ) / sqrt(d k )+M mask )V(H), where M mask is the spatial constraint matrix generated according to the device topology, which prevents incorrect associations between non-adjacent nodes. Design a two-way mutual information maximization module: within the modality, the Jensen-Shannon divergence is used to constrain the consistency of the feature distribution, and between modalities, the InfoNCE loss is used to enhance semantic alignment: L JSD =1 / 3∑ m=1 3 KL(F m ||1 / 3∑F m ), L info NCE=-E[log[(exp(f(x) T f(x + ))) / (exp(f(x) T f(x + ))+∑exp(f(x) T f(x - ))]], and the hyperparameters in this stage include the number of attention heads h = 8, the hidden layer dimension d k = 512, and the dropout rate p = 0.1.
[0154] Step 5: Joint Distribution Optimization and Stability Control Achieve manifold alignment of multi-modal features through differential equation constraints. Construct a joint optimization objective function: min θ L recon +λ1L jSD +λ2L Lip . The reconstruction loss L recon includes the auto-encoding errors of each modality, and the Lipschitz constraint term is: L Lip =max(|gradxf(x)|2 - L, 0), and spectral normalization technology is used to ensure that the network meets the stability requirement of L ≤ 1.5. The optimization process adopts an alternating training strategy: fix the parameters of the knowledge encoder and update the feature extractor; then jointly optimize the entire network. Ablation experiments on the semiconductor defect dataset show that when λ1 = 0.7 and λ2 = 0.3, the model achieves an optimal balance between precision and robustness.
[0155] Step 6: Interpretable Defect Localization and Root Cause Tracing In the final stage, generate a pixel-level defect heat map and associate it with the physical cause. Design a differentiable clustering head: Mdefect = ∑ c=1 C π c N(μ c , Σ c ), mixing coefficient π c is dynamically generated by knowledge features. The visual attention area is extracted through the Grad-CAM++ algorithm and spatially intersected with the thermal anomaly area R final = GeLU(W g •(A vis OA thermal ))), and the root cause tracing module performs probabilistic reasoning based on the knowledge graph: P(e i | d j ) = (exp(sim(h_e i , h_(d j )))) / (∑exp(sim(h_e k , h_d j ))), and the output result includes the localization mask, the confidence of the defect type, and the failure probability of the associated component, meeting the interpretability requirements of industrial inspection.
[0156] d j represents the specific defect detected by the model, the j-th defect example. e i represents the i-th component, the physical entity that may cause the defect. h_d j represents the feature vector corresponding to the region where the defect d j is located. h_e i represents the embedding vector generated by the i-th component e i .
[0157] In the intelligent operation and maintenance of power equipment, the present invention can solve the problem of multi-source monitoring data fusion of substation equipment. Defects in high-voltage circuit breakers and transformers often manifest as surface cracks in visible light images, local overheating in infrared spectra, and mechanical vibration anomalies in ultrasonic detection. By constructing a knowledge graph from the device's 3D CAD model and electrical wiring diagram, and using the Transformer architecture to synchronously analyze visible light inspection images, infrared thermal video, and acoustic signal, cross-modal causal reasoning of mechanical structure anomalies and electrical parameter fluctuations is achieved. For example, when a discharge trace on the bushing surface is detected, the system can automatically associate the phenomenon of excessive acetylene content in the oil chromatographic analysis data and generate an interpretable diagnostic report of internal insulation deterioration causing partial discharge, providing a basis for condition-based maintenance decisions and avoiding misjudgments caused by traditional single-sensor detection.
[0158] In terms of engine health management, the present invention can achieve multi-dimensional collaborative diagnosis of blade damage. Blade defects involve multi-modal characterizations such as microscopic cracks (X-ray imaging), coating spalling (visible light images), and aero-performance degradation (pressure sensor data). By using the blade finite element analysis model as domain knowledge constraints, the theory of maximizing mutual information is adopted to fuse multi-source detection data, and a mapping relationship between the material stress concentration area and the abnormal surface temperature field is established. When a vibration signal of a specific frequency is detected, the system can automatically associate the micro-crack image features at the corresponding position, predict the crack propagation path, and evaluate the remaining life. Compared with the single-modal detection method, the missed detection rate is reduced to less than 0.3%, meeting the stringent requirements for engine predictive maintenance.
[0159] The breakthrough direction of the present invention lies in constructing a cross-modal joint optimization framework under domain knowledge constraints: 1) Convert the device structure topology diagram into the adjacency matrix of the graph convolutional network to establish an explicit mapping between multi-modal features and the physical structure; 2) Design a two-stream Transformer architecture based on maximizing mutual information, and achieve semantic-level alignment of text descriptions, thermal anomaly areas, and 3D model components through Jensen-Shannon divergence optimization; 3) Introduce the Lipschitz continuity condition to constrain the feature projection process to ensure that the multi-modal association conforms to the physical laws of the actual device conditions. This method theoretically solves the mathematical unity problem of heterogeneous data and structured knowledge collaborative representation, and practically realizes the synchronous improvement of defect localization accuracy and annotation efficiency, providing a new path for constructing a self-explanatory industrial vision detection system.
[0160] In complex industrial scenarios, device vision defect detection faces the core technical problems of scarce labeled data and fragmented cross-modal information. Traditional single-modal image analysis methods rely on manually labeled defect samples, which have problems such as high labeling costs and weak guidance. Especially for tiny defects or anomalies in non-visible light bands (such as thermodynamic distribution anomalies), a single vision modality is difficult to provide sufficient characterization. Although existing cross-modal methods attempt to fuse multi-source data, they lack the embedding of prior knowledge of the device physical structure, resulting in insufficient semantic consistency during feature alignment and false or missed detections. A deeper problem is that the heterogeneity of multi-modal data makes traditional Euclidean space metric criteria ineffective, and the requirement for the interpretability of defect localization in the industrial field restricts the direct application of black-box models. The present invention proposes a knowledge-guided collaborative representation framework for the above three major technical bottlenecks of multi-modal representation mismatch, insufficient utilization of domain knowledge, and lack of physical interpretability.
[0161] The technical solution of the present invention constructs a three-stage processing flow of knowledge embedding-cross-modal alignment-joint optimization. First, through the domain knowledge graph construction module, structured knowledge such as device circuit diagrams and 3D CAD models is transformed into topological graphs, and graph convolutional networks (GCNs) are used to extract hierarchical features to form node embedding vectors with physical meanings. Then, a multi-modal Transformer architecture is established, and its core is an improved mutual information maximization module: let the visual modal feature be V ∈ R d×m , the text modal feature be T ∈ R d×n , the heat map modal be H ∈ R d×p , and the joint distribution constraint JSD(V, T, H) = 1 / 3[KL(V||M) + KL(T||M) + KL(H||M)] is constructed through Jensen-Shannon divergence, where M = (V + T + H) / 3 is the mixed distribution and KL is the Kullback-Leibler divergence. This mathematical model forces each modality to maintain topological consistency in the shared subspace, and at the same time controls the stability of feature transformation through the Lipschitz continuity condition ‖f(x) - f(y)‖ ≤ L‖x - y‖. The key parameters include the number of graph convolutional layers K (typical value 3 - 5), the number of Transformer heads h (recommended 8 - 12), and the Lipschitz constant L (constrained in the interval of 1.2 - 1.8), and the optimal configuration of these parameters is determined through domain knowledge verification experiments. The algorithm alternately optimizes the intra-modal reconstruction loss and the cross-modal alignment loss during the training stage, and finally outputs a defect region probability map with physical interpretability.
[0162] The overall technical solution process of the present invention is as follows: After the system starts, it first loads the device CAD model and process documents, and generates a topological graph containing node attributes (such as electrical parameters, material properties) and edge relationships (such as connection methods, spatial orientations) through the knowledge parsing engine; then the multi-modal acquisition module synchronously obtains visible light images, infrared heat maps, and maintenance record texts, and extracts the original features using ResNet, a specific band decomposition algorithm, and the BERT model respectively; in the core processing stage, the graph convolutional network hierarchically aggregates the topological knowledge, and its output is used as the Key-Value pair of the Transformer, and performs attention interaction with the Query of the multi-modal features. During this process, the mutual information maximization module dynamically adjusts the contribution weights of each modality; finally, a defect heat map is generated through differentiable clustering, and the root cause of the defect is traced in combination with the knowledge graph. The entire process adopts end-to-end training, but retains the editability of the knowledge graph to adapt to different device models.
[0163] Taking the EL detection of photovoltaic panels as an example, the implementation process is specifically described as follows: When the system detects an electroluminescence dark spot in a certain cell, first locate the circuit node corresponding to this area from the three-dimensional model of the module (such as belonging to series branch 3), and synchronously analyze the thermal imaging data at this point (the temperature rises abnormally by 0.8 °C) and the maintenance record text (the historical record shows that the solder tape has fallen off); the knowledge graph automatically associates the current path characteristics upstream and downstream of this position. After the Transformer architecture fuses this information, it not only marks the defective area but also outputs a probabilistic diagnosis conclusion (the virtual connection of the solder tape leads to an increase in local resistance, with a confidence level of 92%). During this process, the number of GCN layers is set to 4 to cover the three-level topology of cell - series solder tape - bus bar, and the Lipschitz constant is constrained to 1.5 to ensure the stable mapping of thermodynamic characteristics and electrical characteristics.
[0164] The improvements of this solution include: 1) The method of integrating the equipment structure topology diagram as a trainable parameter into the multi-modal learning framework; 2) The multi-modal joint distribution optimization algorithm based on Jensen-Shannon divergence and its Lipschitz constraint implementation; 3) The dynamic mapping mechanism between the node attributes of the knowledge graph and the deep learning feature vectors; 4) The end-to-end interpretable feature alignment system architecture for industrial defect detection; 5) The cross-modal attention weight allocation strategy integrating physical priors.
[0165] In this embodiment, in terms of detection performance, verified by a certain semiconductor equipment dataset, compared with the pure vision method, multi-modal collaborative representation improves the F1-score by 27.6%, and especially the recall rate increases by 41.3% in the detection of micron-level cracks; at the level of engineering implementation, due to the introduction of structured knowledge guidance, the number of training samples required for model convergence is reduced by 83%, and the GPU video memory occupancy during the inference stage is reduced by 62%, meeting the real-time requirements of the industrial site (the single-frame processing time < 50 ms); in terms of technical scalability, the constructed knowledge anchor points support the automatic reasoning of defect attributes (such as the overheating of transistor Q2 leads to the virtual connection of the solder joint), providing semantic annotations for the subsequent data closed-loop. These advantages directly stem from the newly added domain knowledge embedding step and the joint distribution optimization module. The former encodes the physical constraints of the equipment through a graph structure, and the latter uses information theory criteria to solve the problem of metric unification for heterogeneous modalities. The two work together to break through the dependence on manual annotation of traditional methods.
[0166] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A defect detection method based on joint distribution optimization and structural knowledge guidance, characterized in that Including: Analyze the structured document of the device and construct the physical topology diagram of the device; Use a graph neural network to process the physical topology diagram of the device and generate a structured feature vector; Collect multi-modal perception data from industrial devices and extract perception feature vectors from the multi-modal perception data; In the attention model, use the structured feature vector to guide the alignment of the perception feature vector, perform cross-modal feature fusion, and obtain a fused feature representation; Generate a defect detection result based on the fused feature representation.
2. The method according to claim 1, characterized in that In the attention model, the steps of using the structured feature vector to guide the alignment of the perception feature vector, performing cross-modal feature fusion, and obtaining a fused feature representation include: Map the structured feature vector into a set of key vectors and a set of value vectors respectively, and this set of key vectors and value vectors are used to describe the structural information of the physical topology of the device; Transform the perception feature vector into a query vector, and the query vector is used to characterize the real-time state collected from the device; Explore the key vector through the query vector to determine the correlation degree between the perception feature vector and the physical topology of the device, and dynamically weight the value vector based on the correlation degree, and aggregate to form a fused feature representation.
3. The method according to claim 1, wherein The optimization steps of the attention model include: Inside the attention model, for each of multiple modalities, generate a corresponding modality feature distribution respectively, so as to obtain a set of modality feature distributions; Calculate the arithmetic mean of all distributions in the set of modality feature distributions to construct a mixed central distribution; Measure the Kullback-Leibler divergence of each modality feature distribution in the set of modality feature distributions tending to the mixed central distribution, and integrate all the measured divergences into a joint distribution divergence loss; Optimize the attention model based on the joint distribution divergence loss to minimize the topological difference between modality feature distributions.
4. The method according to claim 3, characterized in that, The optimization steps also include: Obtain the gradient norm of the feature transformation function in the attention model, and based on this gradient norm and a preset Lipschitz constant, obtain a stability constraint loss; Perform a weighted combination of the stability constraint loss and the joint distribution divergence loss to form a combined optimization objective; Finally optimize the attention model according to the combined optimization objective to make the feature transformation function meet the requirements of smoothness and continuity of physical laws.
5. The method according to claim 2, wherein The step of exploring the key vector through the query vector to determine the correlation degree further includes: Generate a spatial constraint matrix according to the physical topology diagram of the device, and the spatial constraint matrix is used to calibrate the physical adjacency relationship allowing associations between device nodes; When calculating the correlation degree between the query vector and the key vector, apply the spatial constraint matrix to shield the invalid associations between physically non-adjacent nodes to obtain the correlation degree; Generating a defect detection result based on the fused feature representation includes: Decode the fused feature representation to generate a defect localization mask indicating the spatial position of the defect; Perform a correlation inference on the features corresponding to the defect position in the fused feature representation and the node attributes in the physical topology diagram of the device, trace the physical root cause of the defect, and generate a root cause diagnosis report including the confidence of the faulty component; Integrate the defect localization mask and the root cause diagnosis report to form the defect detection result.
6. A defect detection system based on joint distribution optimization and structural knowledge guidance, characterized in that, Including: A knowledge parsing module, configured to parse the device structured document to construct a device physical topology graph; A data acquisition module, configured to acquire multi-modal perception data from industrial devices; A feature processing module, communicatively connected to the knowledge parsing module and the data acquisition module, A result generation module, communicatively connected to the feature processing module, for generating a defect detection result based on the fused feature representation; wherein the feature processing module includes: A graph embedding unit, configured to process the device physical topology graph to generate a structured feature vector; A perception feature extraction unit, configured to extract a perception feature vector from the multi-modal perception data; A feature fusion unit, built-in with an attention model, configured to use the structured feature vector to guide the alignment of the perception feature vector, perform cross-modal feature fusion, and obtain a fused feature representation.
7. The system according to claim 6, wherein The feature fusion unit is further configured to: Map the structured feature vector received from the graph embedding unit into a set of key vectors and a set of value vectors respectively, for describing the structural information of the device physical topology; Transform the perception feature vector received from the perception feature extraction unit into a set of query vectors, for characterizing the real-time state collected from the device; wherein, the built-in attention model is configured to: drive the query vector to explore the key vector to determine the correlation degree between the perception feature vector and the device physical topology, and dynamically weight the value vector based on the correlation degree, and aggregate to form a fused feature representation.
8. The system according to claim 6, wherein The feature processing module is further configured to be trained through an optimization objective, and the optimization objective includes: For each of multiple modalities, generate a corresponding modality feature distribution inside the feature processing module respectively, so as to obtain a set of modality feature distributions; Construct a mixed central distribution, which is the arithmetic mean of all distributions in the set of modality feature distributions; Integrate the Kullback-Leibler divergence of each modality feature distribution tending to the mixed central distribution to form a joint distribution divergence loss, and optimize based on this loss.
9. The system according to claim 8, wherein The optimization objective further includes a stability constraint loss, wherein the feature processing module is further configured to: Obtain the gradient norm from the internal feature transformation function thereof, and obtain the stability constraint loss based on the gradient norm and a preset Lipschitz constant; Perform weighted combination of the stability constraint loss and the joint distribution divergence loss to form a combined optimization objective, and complete the training according to this combined optimization objective.
10. The system according to claim 9, characterized in that, The feature fusion unit is further configured to: Generate a spatial constraint matrix according to the device physical topology graph received from the knowledge parsing module, for calibrating the physical adjacency relationship allowing associations between device nodes; wherein, when the built-in attention model determines the correlation degree, it applies the spatial constraint matrix to mask the invalid associations between physically non-adjacent nodes to obtain the correlation degree; The result generation module is further configured to: Decode the fused feature representation received from the feature processing module to generate a defect localization mask indicating the spatial position of the defect; Perform the correlation inference between the fused feature representation and the device physical topology graph to trace the physical root cause of the defect, and generate a root cause diagnosis report including the confidence of the faulty component; Integrate the defect localization mask and the root cause diagnosis report to obtain the defect detection result.
Citation Information
Patent Citations
Data search method and system based on artificial intelligence
CN119227013A
Method for detecting defects of power equipment and related products
CN119557716A
EOSIO smart contract vulnerability detection method of adaptive multi-channel graph convolutional network
CN119583145A
Industrial Internet of Things equipment detection method and system based on federal map neural network
CN119676094A
Multi-modal data prediction method based on causal markov model
US20240143999A1
Cited By
AI visual inspection method for battery manufacturing defects and self-adaptive correction system
CN120778741A
Artificial intelligence visual inspection method for battery manufacturing defects and adaptive correction system
CN120778741B
Cross-modal semantic alignment driven power grid equipment fault diagnosis method and system
CN121071667A
Power distribution station equipment defect discrimination method, equipment and medium
CN121095255A
Alignment method based on natural language and machine vision
CN121117953A