Defect generation root cause tracing method and system
Patent Information
- Application Number
- CN202611007229.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请提供一种缺陷生成根因追溯方法和系统,利用物理参数化缺陷生成模型根据设定的物理参数在正常的产品图像上合成逼真的缺陷,由缺陷生成反演引擎通过优化算法,从真实的缺陷图像中逆向推断出最可能产生该缺陷的物理参数组合,以该物理参数组合为桥梁与设备数据进行因果匹配,实现了缺陷的物理级可解释诊断,实现了从缺陷表现感知到缺陷物理根因认知的根本性跨越,克服现有工业视觉检测技术仅能感知缺陷表象,而无法自动、准确、可解释地追溯缺陷产生根源的局限性
[0015]结合第一方面,在第一方面的某些实现方式中,利用缺陷生成反演引擎根据真实缺陷图像和映射关系从物理参数空间中反演出第一物理参数组合,具体包括:获取与真实缺陷图像对应的正常的产品图像。利用缺陷生成模型在物理参数空间内进行逆向优化搜索第一物理参数组合,以使得基于正常的产品图像和第一物理参数组合生成的模拟缺陷图像为与真实缺陷图像最相似的图像。
Smart Images

Figure CN122819474A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial intelligent detection and fault diagnosis technology, and more specifically, to a method and system for tracing the root causes of defects. Background Technology
[0002] Currently, deep learning-based machine vision technology is widely used in industrial product quality inspection, automatically identifying defects such as scratches, dents, and stains on product surfaces, greatly improving the efficiency and consistency of industrial quality inspection. However, current mainstream methods, such as object detection based on You Look Only Once (YOLO), Faster R-CNN (Faster Region-Convolutional Neural Network), or unsupervised anomaly detection based on Anomaly Detection, only output defect bounding boxes, category labels, or anomaly scores. Therefore, current detection and diagnosis are disconnected; they can only perceive the surface of defects but cannot trace their root causes, and problem localization still relies on time-consuming manual investigation by experienced engineers.
[0003] Therefore, how to trace the root cause of defects is an urgent problem to be solved. Summary of the Invention
[0004] This application provides a method and system for tracing the root causes of defects. It utilizes a physically parameterized defect generation model to synthesize realistic defects on normal product images based on set physical parameters. A defect generation inversion engine, through optimization algorithms, inversely infers the most likely combination of physical parameters that caused the defect from the real defect image. This combination of physical parameters is then used as a bridge for causal matching with equipment data, achieving physically interpretable diagnosis of defects. This represents a fundamental leap from perceiving defect manifestations to recognizing the physical root causes of defects, overcoming the limitations of existing industrial vision inspection technologies that can only perceive the surface of defects but cannot automatically, accurately, and interpretably trace the root causes. It breaks through the technical bottleneck of traditional machine vision, which can only answer "what is the defect and where is it" but cannot locate the source equipment, significantly shortening manual troubleshooting time and significantly reducing unplanned downtime losses.
[0005] Firstly, a method for tracing the root causes of defects is provided. The method includes acquiring a defect generation model, which is used to synthesize simulated defect images based on normal product images and physical parameters in a defined physical parameter space, thereby constructing a mapping relationship between the simulated defect images and the physical parameters. Next, a real defect image to be diagnosed is acquired, and a first combination of physical parameters is inverted from the physical parameter space based on the real defect image and the mapping relationship using a defect generation inversion engine. Finally, a causal reasoning engine performs joint reasoning based on the associated equipment data corresponding to the real defect image and the first combination of physical parameters, outputting a root cause diagnosis report. The root cause diagnosis report includes the faulty production equipment and fault mode that caused the defect.
[0006] In this embodiment, a defect mapping is formed by learning normal samples and physical laws through a forward simulation model. During the inference stage, generation parameters that conform to physical laws are inferred from a single real defect image through reverse optimization. Finally, the combination of these physical parameters is used as a bridge to perform causal matching with equipment data, achieving a physically interpretable diagnosis of defects. This transforms the difficult-to-understand low-level pixel information into clear and interpretable physical parameters, achieving a fundamental leap from perceiving defect manifestations to recognizing the physical root causes of defects. It overcomes the limitation of existing industrial vision inspection technologies, which can only perceive the appearance of defects but cannot automatically, accurately, and interpretably trace the root cause of defects. It breaks through the technical bottleneck of traditional machine vision, which can only answer "what is the defect and where is it" but cannot locate the source equipment. It effectively solves the problem of multi-source heterogeneous data in the production site being unable to be analyzed collaboratively due to the island effect, significantly shortening the time for manual fault diagnosis and significantly reducing unplanned downtime losses.
[0007] In conjunction with the first aspect, in some implementations of the first aspect, the defect generation model is constructed based on a diffusion model. The defect generation model includes a physical parameter encoder, a conditional fusion module, and a multi-level conditional injection network. Specifically: the physical parameter encoder converts physical parameters into physical embedding vectors; the conditional fusion module obtains text feature embeddings, fuses the text feature embeddings with the physical embedding vectors to generate fused conditional embeddings; and the multi-level conditional injection network performs conditional injection based on the fused conditional embeddings from a normal product image to generate a simulated defect image.
[0008] In this embodiment, a diffusion model is used as the basic generative skeleton. A dedicated physical parameter encoder decouples the original physical parameter space and performs high-dimensional semantic representation. A conditional fusion module performs cross-modal alignment and fusion of discrete textual process descriptions and precise continuous physical variables to form a unified control condition representation. Finally, a multi-level conditional injection network implements multi-dimensional, multi-layered pervasive control on the denoising trajectory of the diffusion model. The excellent latent space modeling capability of the diffusion model ensures the high realism and richness of detail in the generated defects. This allows conventional image generation models to effectively accept continuous parameters with clear physical meaning as constraints, establishing a precise transmission channel between abstract physical parameters and the generated image.
[0009] In conjunction with the first aspect, in some implementations of the first aspect, the conditional fusion module is used to fuse the physical embedding vector with the text feature embedding to generate a fused conditional embedding. Specifically, this includes: using a first attention stream, performing attention calculation with the text feature embedding as the query vector and the physical embedding vector as the key and value vectors to extract the constraint features of physical conditions on text positions; using a second attention stream, performing self-attention calculation with the text feature embedding simultaneously as the query vector, key vector, and value vector to extract semantic dependency features between text features; and fusing the constraint features and semantic dependency features through gating weight calculation, then performing residual connection and layer normalization processing with the text feature embedding to output the fused conditional embedding.
[0010] In this embodiment, a dual-stream parallel attention interaction pipeline is employed. The first attention stream is a cross-attention mechanism used to force text tokens to retrieve and correspond to precise physical scale information. The second attention stream is a self-attention mechanism used to maintain the contextual logic of the text itself. Finally, an adaptive gating weight matrix dynamically adjusts the fusion depth of the two streams, and residual connections and layer normalization are used to ensure gradient stability. This effectively overcomes the shortcomings of conventional splicing projection schemes, which suffer from low requirements for physical condition independence, leading to the dilution or loss of some key physical information in deep networks. Through adaptive hierarchical attention fusion, the model can deeply inject precise constraints of physical conditions while preserving the rich semantic information of the original text conditions. Moreover, even in scenarios without text descriptions, learnable embeddings can still maintain structural consistency, resulting in stronger input generalization. The introduction of layer normalization and residual connections prevents gradient vanishing and significantly improves the stability of model training during multimodal condition fusion.
[0011] In conjunction with the first aspect, in some implementations of the first aspect, the multi-level conditional injection network includes a cross-attention layer and a feature affine modulation layer. The multi-level conditional injection network is used to perform conditional injection based on the fusion conditional embedding of a normal product image. Specifically, during the denoising process, the latent representation of the normal product image is used as the basis. The cross-attention layer uses the fusion conditional embedding to perform high-level semantic guidance, and the feature affine modulation layer uses the fusion conditional embedding to perform spatial feature modulation to complete the conditional injection.
[0012] In this embodiment, the internal architecture of the conditional UNet is improved. In each layer of the diffusion denoising network, the fused conditional embedding is simultaneously fed into the cross-attention layer and the feature affine modulation layer (FiLM layer), achieving synergistic effects of precise control and semantic guidance. A multi-level conditional injection mechanism is constructed to ensure that the conditional information can provide directional guidance at the high-level semantic dimension and perform precise pixel-level modulation at the low-level spatial feature dimension. This achieves synergy between "precise control" and "semantic guidance" in the defect generation process, ensuring that the generated defective image is strictly constrained by physical parameters at both the micro and macro levels, including position, scale, and texture.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the defect generation model is trained using a training loss function, which includes: a pixel-level reconstruction loss for constraining the pixel differences between the synthesized simulated defect image and the training ground value; a perceptual loss for constraining the distance between the simulated defect image and the training ground value in a high-level feature space; and a physical constraint loss for penalizing combinations of physical parameters that do not conform to physical laws.
[0014] In this embodiment, a weighted training loss function is constructed. The pixel-level reconstruction loss is responsible for constraining the absolute errors of color and shape at the microscopic level; the perceptual loss is responsible for constraining the high-level visual fit at the macroscopic level; and the physical constraint loss serves as a physical regularization term, based on industrial manufacturing mechanisms and boundary rules, directly blocking and penalizing the network for outputs that do not conform to real physical common sense. This ensures that the constructed defect generation model possesses both high visual realism and rigorous physical rationality, enabling the network to reflect the physical mechanism of parameter-defect mapping, and significantly improving the model's numerical stability and inversion accuracy under extreme parameter boundaries.
[0015] In conjunction with the first aspect, in some implementations of the first aspect, a defect generation inversion engine is used to invert the first combination of physical parameters from the physical parameter space based on the real defect image and the mapping relationship. Specifically, this includes: obtaining a normal product image corresponding to the real defect image; and using a defect generation model to perform inverse optimization search for the first combination of physical parameters in the physical parameter space, so that the simulated defect image generated based on the normal product image and the first combination of physical parameters is the image most similar to the real defect image.
[0016] In this embodiment, a defect generation inversion mechanism is used to reverse-engineer the capabilities of the forward mapping network. When searching for the optimal combination of physical parameters in the parameter space, images of defect-free normal products of the same model or corresponding region are retrieved as background benchmarks for defect simulation and synthesis comparison. This minimizes the distance gradient between the forward-generated simulated image and the currently observed real defect image to be diagnosed, thereby obtaining the globally optimal physical solution. This overcomes the severe limitation of conventional AI diagnostic technologies, which rely on massive, costly, manually annotated databases of real defect samples. Since the training data for the forward model can be automatically synthesized from normal samples through physical simulation or procedural rendering methods, the physical causes can be reverse-engineered from a single defect image during the diagnostic phase. This provides a practical intelligent diagnostic method for industrial production lines with high yield rates, small sample sizes, or even zero unknown samples.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: extracting the production path of the defective product from the manufacturing execution system to determine a causal time window based on the production identifier and timestamp of the defective product corresponding to the actual defect image; and obtaining the associated equipment data corresponding to the actual defect image, which includes the timing data and process logs of the relevant equipment in the production path within the causal time window.
[0018] This application provides a precise alignment scheme for cross-modal data association. By querying the historical production history of defective products based on their serial numbers, the time interval of specific process steps is identified, and then the high-frequency sensor time-series data of the equipment within that time period is accurately extracted. This breaks down the barriers between visual data, equipment time-series data, and process data, enabling cross-modal joint reasoning and improving the comprehensiveness and accuracy of diagnosis.
[0019] In conjunction with the first aspect, in certain implementations of the first aspect, a causal reasoning engine is used to perform joint reasoning based on the associated equipment data corresponding to the real defect image and the combination of first physical parameters. Specifically, this includes: constructing a dynamic causal graph, which includes a physical causal layer, an intermediate phenomenon layer, and a defect manifestation layer. The physical causal layer contains equipment state nodes and / or process parameter nodes representing observable or controllable entities on the production line; the intermediate phenomenon layer contains phenomenon nodes representing invisible physical or chemical changes that occur in the product during processing; and the defect manifestation layer contains defect nodes representing the final observable defect characteristics of the product. Associated equipment data is input to the nodes corresponding to the physical causal layer, and the combination of first physical parameters is input to the defect nodes corresponding to the defect manifestation layer. A probabilistic reasoning algorithm is run on the dynamic causal graph to solve for the posterior probability of the faulty production equipment and the fault mode, and a root cause diagnosis report is output.
[0020] In this embodiment, expert knowledge and probabilistic graphical models are deeply integrated to construct a three-layer progressive Bayesian network. The physical causal layer directly receives directly controllable process and temporal characteristics; the intermediate phenomenon layer maps the invisible intermediate physical / chemical state transitions; and the defect manifestation layer receives the inversely derived physical parameter manifestations of defects. Through bidirectional probability propagation and the Bayesian evidence chain fusion formula, rigorous mathematical reasoning from result plus evidence to the posterior probability of the fault source is achieved. The originally unknowable and unreliable AI black box decision-making process is completely transformed into a transparent causal graph with extremely high physical interpretability. Strictly following the essential chain of physical evolution of industrial defects, it not only has high diagnostic accuracy but also provides production line engineers with reasonable counterfactual reasoning basis. The output results can be directly converted into physical maintenance actions, forming an intelligent decision-making closed loop.
[0021] In conjunction with the first aspect, some implementations of the first aspect further include collecting production data from the production line in real time and dynamically updating the conditional probability table of each node in the dynamic causal graph using a Bayesian update algorithm. A causal discovery algorithm is then run using production data accumulated over a preset historical time period to automatically add or delete nodes in the dynamic causal graph and / or adjust the causal relationships between nodes, thereby achieving dynamic updates to the dynamic causal graph.
[0022] In this embodiment, industrial production lines often experience initial causal graph failures due to gradual wear and tear of physical components, secondary improvements to process windows, or environmental drift. The probability weights are continuously adjusted using new samples generated by the production line, and causal discovery algorithms are periodically used to uncover hidden new nodes or correct causal connections, giving the root cause tracing system an adaptive lifecycle and evolutionary capability. This allows the root cause tracing system to adapt to dynamic changes in industrial production lines, such as long-term equipment wear and aging, and process parameter upgrades; it avoids the causal knowledge base becoming invalid over time, greatly reduces reliance on manual rule maintenance by domain experts, and maintains the system's diagnostic accuracy throughout its entire lifecycle.
[0023] In conjunction with the first aspect, in some implementations of the first aspect, a root cause diagnosis report is output, specifically including outputting a real defect image in the interactive interface, a simulated defect image obtained by the defect generation model based on a combination of first physical parameters, a dynamic causal graph used to characterize the decision evidence chain of the causal reasoning engine, and maintenance suggestions.
[0024] In this embodiment, a multimodal simultaneous screen presentation mechanism is designed at the system's application and interaction layer. On the engineer's terminal interface, real defect images, inverted physical-level simulation diagrams, dynamic evidence causal trees with probability strength connections, and expert action suggestions directly linked to the spare parts maintenance database are displayed simultaneously. This provides on-site engineers with an extremely intuitive, operable, and highly trustworthy central platform for troubleshooting. The visual consistency between the real and simulated images allows engineers to directly verify the accuracy of the inversion engine; the transparent display of the decision evidence chain completely eliminates engineers' concerns about AI false positives, reducing the traditional hours of manual cross-departmental data retrieval and game-theoretic troubleshooting to minutes of precise closed-loop maintenance, maximizing the commercial value of data-driven zero-defect manufacturing.
[0025] Secondly, a defect generation root cause tracing system is provided, which is a module or unit for performing the methods described in the first aspect and its various implementations.
[0026] Thirdly, an apparatus is provided, comprising a processor and a memory, wherein the processor and the memory are connected together, wherein the memory is used to store program code, and the processor is used to invoke the program code to execute the methods described in the first aspect and its various implementations.
[0027] Fourthly, a computer-readable storage medium is provided storing a computer program that is executed by a processor to implement the method in any possible implementation of the method design of the first aspect.
[0028] Fifthly, a computer program product is provided, including instructions that, when executed by a processor, cause a computer to perform any possible implementation of the method design of the first aspect described above.
[0029] Other beneficial effects can be found in the description of the first aspect, and will not be repeated here. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of a defect generation root cause tracing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the network architecture of a defect generation model provided in an embodiment of this application; Figure 3 This is a schematic diagram of a physical parameter encoder network structure provided in an embodiment of this application; Figure 4 This is a schematic diagram of an adaptive hierarchical attention fusion scheme provided in an embodiment of this application; Figure 5 This is a schematic diagram of an improved conditional Unet internal architecture provided in an embodiment of this application; Figure 6 This is a pseudocode diagram of a defect generation and inversion engine provided in an embodiment of this application; Figure 7 This is a schematic diagram of a defect generation model training process provided in an embodiment of this application; Figure 8 This is a schematic diagram of a defect generation and inversion reasoning process provided in an embodiment of this application; Figure 9 This is a schematic diagram of the specific structure of a dynamic evidence chain visualization causal graph provided in an embodiment of this application; Figure 10 This is a schematic diagram of a defect generation root cause tracing system architecture provided in an embodiment of this application; Figure 11 This is a schematic diagram of an optimal combination of physical parameters provided in an embodiment of this application; Figure 12 This is a schematic diagram of a root cause diagnosis report provided in an embodiment of this application; Figure 13 This is a schematic diagram of a defect generation root cause tracing system provided in an embodiment of this application. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0032] The "range" disclosed in this application is defined by a lower limit and an upper limit. A given range is defined by selecting a lower limit and an upper limit, which define the boundaries of a specific range. Ranges defined in this way include endpoint values and can be arbitrarily combined; that is, any lower limit can be combined with any upper limit to form a range. For example, if ranges of 60-120 and 80-110 are listed for a specific parameter, it is also expected that ranges of 60-110 and 80-120 are also included. Furthermore, if the minimum range values are listed as 1 and 2, and the maximum range values are listed as 3, 4, and 5, then the following ranges are all expected: 1-3, 1-4, 1-5, 2-3, 2-4, and 2-5. In this application, unless otherwise stated, the numerical range "ab" represents a shortened representation of any combination of real numbers between a and b, where a and b are real numbers. For example, the numerical range "0-5" means that all real numbers between "0-5" have been listed herein; "0-5" is simply a shortened representation of these numerical combinations. Furthermore, when a parameter is described as an integer greater than or equal to 2, it is equivalent to disclosing that the parameter is, for example, an integer such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, etc.
[0033] Unless otherwise specified, all embodiments and optional embodiments of this application can be combined with each other to form new technical solutions.
[0034] Unless otherwise specified, all technical features and optional technical features of this application may be combined to form new technical solutions.
[0035] The terminology used in the following embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, “at least one” and “one or more” refer to one, two, or more than two. The term “and / or” is used to describe the relationship between related objects, indicating that three relationships may exist; for example, A and / or B can indicate: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.
[0036] References to "one embodiment," "some embodiments," "one example," or "some examples" used in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0037] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different descriptive objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0038] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. For clarity, the various parts in the drawings are not drawn to scale. Furthermore, some well-known parts may not be shown in the drawings.
[0039] Currently, deep learning-based machine vision is widely used in industrial product quality inspection, automatically identifying defects such as scratches, dents, and stains on product surfaces, greatly improving inspection efficiency and consistency. However, the output of current mainstream methods is limited to defect bounding boxes, category labels, or anomaly scores. This result can only answer "whether there is a defect" and "where the defect is," failing to address more pressing questions on the production floor, such as "how the defect occurred" and "which piece of equipment or process parameter caused the anomaly." This leads to a disconnect between inspection and diagnosis, and problem localization still relies on time-consuming manual investigation by experienced engineers.
[0040] To address the aforementioned issues, this application provides a method and system for tracing the root causes of defects. It utilizes a physically parameterized defect generation model to synthesize realistic defects on normal product images based on set physical parameters. A defect generation inversion engine, through optimization algorithms, inversely infers the most likely combination of physical parameters that caused the defect from the real defect image. This combination of physical parameters serves as a bridge for causal matching with equipment data, achieving physically interpretable diagnosis of defects. This represents a fundamental leap from perceiving defect manifestations to recognizing the physical root causes of defects, overcoming the limitation of existing industrial vision inspection technologies that can only perceive the surface of defects but cannot automatically, accurately, and interpretably trace the root causes. It breaks through the technical bottleneck of traditional machine vision, which can only answer "what is the defect and where is it" but cannot locate the source equipment, significantly shortening manual troubleshooting time and significantly reducing unplanned downtime losses.
[0041] like Figure 1 As shown, the defect generation root cause tracing method 100 provided in this application embodiment includes at least the following steps: S110, Obtain the defect generation model.
[0042] To overcome the limitations of existing industrial vision inspection technologies, which can only perceive the appearance of defects but cannot trace the root cause of defects, i.e., specific faulty equipment and abnormal process parameters, it is necessary to establish an effective correlation and causal reasoning model between the physical characteristics of product defects and multi-source equipment data of the production line.
[0043] Traditional supervised inspection models require a large number of labeled defect samples for training. However, in actual industrial production lines, especially high-yield production lines, defect samples are scarce and their shapes are diverse, making model training difficult and making it hard to cover all unknown defect types.
[0044] Figure 2 The diagram below shows the network architecture of the defect generation model provided in this application. The defect generation model is used to synthesize simulated defect images based on normal product images and physical parameters in a set physical parameter space, in order to construct a mapping relationship between simulated defect images and physical parameters.
[0045] First, a physical parameter space Θ needs to be defined. For the target defect type, such as stamping scratches, a set of quantifiable physical parameters is defined in collaboration with domain experts. For example: Θ_scratch = {length L, average depth D, orientation angle α, width change rate β, edge buildup height H}. Each parameter has a reasonable physical range.
[0046] Then, a conditional generation network is constructed.
[0047] In one embodiment, the defect generation model is constructed based on a diffusion model. The defect generation model includes a physical parameter encoder, a conditional fusion module, and a multi-level conditional injection network. The physical parameter encoder converts physical parameters into physical embedding vectors. The conditional fusion module acquires text feature embeddings, fuses the text feature embeddings with the physical embedding vectors, and generates fused conditional embeddings. The multi-level conditional injection network performs conditional injection based on the fused conditional embeddings from a normal product image to generate a simulated defect image.
[0048] Using the diffusion model as the framework, a defect image generation model G(I_norm, θ) → I_defect is constructed. Here, I_norm is the input normal product image, and θ∈Θ is the condition vector.
[0049] Compared to the original StableDiffusion diffusion model architecture, the model architecture has been further improved to incorporate the physical parameter space as a conditional input. The overall conditional generation model architecture is as follows: Figure 2 As shown in the diagram, a physical parameter encoder, a conditional fusion module, and a multi-level conditional injection network are introduced into the original framework. The multi-level conditional injection network improves upon the original UNet structure. The physical parameter encoder network structure is shown below. Figure 3 As shown, this module converts physical space parameters into physical embedding vectors. The input is the original physical parameters such as [scratch length, depth, angle, ...], and the output is two 768-dimensional vectors representing the "content" and "style" of the physical condition. For example, token1 represents the defect type and severity (content), and token2 represents the texture and appearance details of the defect (style).
[0050] The physical parameter encoder employs a multilayer perceptron (MLP) topology. From input to output, the encoder consists of a first fully connected layer (FC), a first activation layer (ReLU), a second fully connected layer (FC), a second activation layer (ReLU), and a third fully connected layer (FC). Specifically, the input of the first fully connected layer serves as the input interface for the encoder, receiving the raw physical parameter sequence extracted or defined from the physical parameter space. For example, the raw physical parameter sequence includes, but is not limited to, one or more combinations of length L, average depth D, orientation angle α, width change rate β, and edge stacking height H. The first fully connected layer performs a linear combination transformation on the input raw physical parameter sequence, and its output is connected to the input of the first activation layer. The first activation layer uses a modified linear unit activation function for nonlinear mapping to introduce nonlinear feature representation capabilities, and its output is connected to the input of the second fully connected layer. The second fully connected layer further performs spatial dimension transformation and feature aggregation on the previous features, and its output is connected to the input of the second activation layer. The second activation layer also employs a modified linear unit activation function to perform nonlinear transformations, enhancing the feature fitting accuracy of the network. Its output is connected to the input of the third fully connected layer. The third fully connected layer, as the final output transformation layer, projects the aggregated high-order nonlinear features and ultimately outputs two independent 768-dimensional vectors. These are defined as the first physical embedding vector (token1) representing the physical condition "content" and the second physical embedding vector (token2) representing the physical condition "style." The first physical embedding vector characterizes the type and severity of the defect, while the second physical embedding vector characterizes the microscopic texture and appearance details of the defect. This achieves a decoupled representation of the original physical parameters in the latent semantic space, encompassing both content and style.
[0051] Then, in the conditional fusion module, the text embedding and the physical embedding are fused to form the conditional embedding cond.
[0052] Optionally, the conditional fusion strategy adopts a conventional splicing projection scheme, splicing the physical condition embedding [B, 77, 768] and the text embedding [B, 77, 768] into [B, 77, 768]. This scheme is suitable for scenarios with relatively low requirements for the independence of physical conditions, but may lose some physical information.
[0053] In one embodiment, the conditional fusion module is used to fuse the physical embedding vector and the text feature embedding to generate a fused conditional embedding. Specifically, this includes: using a first attention flow, performing attention calculation with the text feature embedding as the query vector and the physical embedding vector as the key and value vectors to extract the constraint features of physical conditions on text positions; using a second attention flow, performing self-attention calculation with the text feature embedding simultaneously as the query vector, key vector, and value vector to extract semantic dependency features between text features; and fusing the constraint features and semantic dependency features through gating weight calculation, then performing residual connection and layer normalization processing with the text feature embedding to output the fused conditional embedding.
[0054] Preferably, considering that physical conditions are more precise and task-relevant than textual conditions in this scenario, a simple concatenation or weighted averaging approach is not used. Instead, an adaptive hierarchical attention fusion scheme is employed. The relevance between physical conditions and each text token is automatically learned through an attention mechanism. Residual connections preserve the semantic information of the original textual conditions while injecting precise constraints from the physical conditions. When no textual description is provided, a set of learnable embeddings is used as the textual condition embeddings to maintain model structural consistency. The adaptive hierarchical attention fusion scheme is as follows: Figure 4 As shown.
[0055] The adaptive hierarchical attention fusion, from input to output, comprises: an input layer, a physical information enhancement layer, a two-stream attention mechanism layer, a gated fusion layer, a residual connection and normalization layer, and an output layer. The input layer serves as the data receiving interface for the entire fusion scheme, simultaneously acquiring conditional feature embeddings from multiple modalities. In this embodiment, the features received by the input layer include a physical embedding vector, whose tensor dimension is [B, 2, 768], where B is the batch size, 2 represents the number of decoupled physical conditions (i.e., two feature tokens representing physical condition content and physical condition style respectively), and 768 is the feature hiding layer dimension. The text feature embedding tensor dimension is [B, 77, 768], where 77 represents the maximum length of the text sequence. Due to the asymmetry between the physical embedding vector and the text feature embedding in the sequence length dimension, the physical embedding vector [B, 2, 768] output by the input layer is directly fed into the physical information enhancement layer. In the physical information augmentation layer, cross-modal scaling of the physical embedding vectors is performed through linear projection operations, specifically including copy expansion and linear transformation. The physical embedding vectors are expanded to a sequence length of 77 via copy expansion, and then projected into the feature space using a learnable linear transformation matrix, ultimately outputting a physically augmented embedding vector with tensor dimensions aligned to [B, 77, 768]. The physical embeddings are then expanded to [B, 77, 768] dimensions via linear projection, aligning with the text embeddings. The aligned text feature embeddings and the physically augmented embedding vectors are simultaneously input into the two-stream attention mechanism layer. The two-stream attention mechanism layer internally constructs a two-stream parallel computation pipeline to simultaneously extract cross-modal constraints and unimodal semantic dependencies. The first attention stream is a physics-driven attention stream, employing a cross-attention mechanism. It uses text feature embeddings as the query vector (Q) and physics-enhanced embeddings as both key and value vectors (V) for scaled dot product attention calculation. This extracts the cross-modal precise constraint features of physics conditions on each text position, ultimately outputting the first feature matrix A1 with tensor dimensions [B, 77, 768]. This essentially allows the text token to query physics information, determining which physics features should be focused on at each text position. The second attention stream is a text self-attention stream, employing a self-attention mechanism. It uses text feature embeddings as the query vector (Q), key vector (K), and value vector (V) for self-attention calculation. This extracts long-distance contextual semantic dependencies within text features to capture semantic dependencies between text tokens, ultimately outputting the second feature matrix B1 with tensor dimensions [B, 77, 768].The gated fusion layer, connected downstream of the two-stream attention mechanism layer, receives the first feature matrix A1 and the second feature matrix B1 and implements a non-linear adaptive weighting. The first feature matrix A1 and the second feature matrix B1 are concatenated along their feature dimensions to obtain the concatenated feature Concat(A1, B1). This concatenated feature is then fed into a learnable linear mapping layer (Linear), and finally, the dynamic total gate weight matrix gate is calculated using the Sigmoid activation function. Its calculation formula is: gate = σ(linear(Concat(A1, B1))). After obtaining the total gate weight matrix gate, a split operation is performed, decoupling it into a first gate weight matrix gate_physics for physical features and a second gate weight matrix gate_text for text features. The first and second gate weight matrices are then multiplied element-wise with their respective feature matrices and summed to obtain the fused intermediate feature matrix fused, calculated as: gate_physics ⊙ A1 + gate_text ⊙ B1. This allows the model to adaptively determine whether each position should rely more on physical or textual information.
[0056] To prevent gradient vanishing and stabilize training, the intermediate fused feature matrix output from the gated fusion layer is input to the residual connection and normalization layer via residual connections and normalization. To prevent gradient vanishing or numerical fluctuations during backpropagation in the deep neural network, the system performs element-wise residual addition of the intermediate fused feature matrix with the original text feature embedding from the input layer. The result is then fed into the layer normalization module for standardization and scaling. The final output feature vector is calculated as: output = LayerNorm(fused + TextEmbedding). The feature vector output from the residual connection and normalization layer is finally input to the output layer, outputting the final fused conditional embedding vector (fused conditional embedding), whose tensor dimension remains stably [B, 77, 768]. This fused embedding contains both physical and textual semantic information.
[0057] In one embodiment, the multi-level conditional injection network includes a cross-attention layer and a feature affine modulation layer. The multi-level conditional injection network is used to perform conditional injection based on fusion conditional embedding of a normal product image. Specifically, during the denoising process, the latent representation of the normal product image is used as the basis. The cross-attention layer uses fusion conditional embedding to perform high-level semantic guidance, and the feature affine modulation layer uses fusion conditional embedding to perform spatial feature modulation to complete the conditional injection.
[0058] like Figure 5As shown, the latent feature vector obtained from the image encoder is input together with the conditional embedding into the improved conditional UNet. During the diffusion process, the fused conditional embedding is used as a condition, and conditional information is injected simultaneously through the cross-attention layer and the FiLM layer to ensure that the conditional information is modulated by FiLM in terms of spatial features and guided by cross-attention in terms of high-level semantics.
[0059] The UNet main body presents a symmetrical encoder-decoder structure, including the encoder side, the bottleneck layer, and the decoder side. The encoder side comprises Encoder Block 1, Encoder Block 2, and Encoder Block 3, used for progressive downsampling and high-order feature extraction of the initial latent features x. The bottleneck layer, connecting the encoder and decoder sides, consists of several residual blocks and is used to aggregate global latent semantics at the lowest resolution scale. The decoder side comprises Decoder Block 3, Decoder Block 2, and Decoder Block 1, used for progressive upsampling and feature reconstruction of the features output from the bottleneck layer. To simultaneously embed precise physical geometric constraints and macroscopic contextual semantics into the denoising trajectory, a multi-level conditional injection network, consisting of spatial feature modulation and high-level semantic guidance, is embedded in parallel along the feature flow path of the UNet main body. Spatial feature modulation is achieved through a multi-layer feature affine modulation layer (FiLM layer). The physical embedding vector with tensor dimensions [B, 2, 768] output by the physical parameter encoder is input to the FiLM layer. Inside the FiLM layer, the input physical embedding vector (physics_emb) is subjected to affine parameter prediction using a multilayer perceptron, and the multidimensional scaling factor γ and the bias translation factor β are calculated, expressed as γ = MLP(physics_emb); β = MLP(physics_emb). Subsequently, the FiLM layer uses the calculated γ and β to perform channel-by-channel affine transformation modulation on the intermediate feature matrix (feature) output by the previous residual block, outputting the spatial modulation feature matrix modulated_feature, expressed as: modulated_feature = γ feature+β.
[0060] High-level semantic guidance is achieved through multiple cross-attention control layers (cross-attention layers). The fused conditional embedding vector with tensor dimensions [B, 77, 768] output by the conditional fusion module is used as the control condition input to the cross-attention layer. During computation, the cross-attention layer uses the intermediate feature matrix Feature as the query tensor (Query, Q) and the input fused conditional embedding vector fused_condition as both the key tensor (Key, K) and value tensor (Value, V) to perform cross-modal scaled dot product attention mapping, ultimately outputting attention-weighted features with macroscopic process semantic guidance. The multidimensional feature stream, after fine-grained spatial scale constraints through spatial feature modulation and macroscopic semantic alignment through high-level semantic guidance, is aggregated and calculated at the decoding end of the UNet main body, ultimately outputting a denoised latent representation. This denoised latent representation is provided to the downstream variational autoencoder (VAE) decoder to reconstruct and generate the final simulated defect image.
[0061] During training, noise is added to the latent representation z_norm of the normal image, and then UNet is trained to predict the noise based on condition θ. During inference, starting with random noise, the image is gradually denoised using condition θ to obtain a latent representation with defects, and finally the defective image is obtained through a VAE decoder.
[0062] Before model training, training data needs to be generated. A large number of defect-free product images collected under normal production conditions are acquired and defined as normal images I_norm. For these normal images I_norm, parameters θ are randomly sampled or combined according to physical laws, and matched defect images I_defect_sim are generated using physical simulation software or procedural rendering methods as training ground values. For example, finite element analysis (FEA) can be used to simulate the scratch generation process and render the images.
[0063] In one embodiment, the defect generation model is trained using a training loss function, which includes: a pixel-level reconstruction loss for constraining the pixel differences between the synthesized simulated defect image and the training ground value; a perceptual loss for constraining the distance between the simulated defect image and the training ground value in a high-level feature space; and a physical constraint loss for penalizing combinations of physical parameters that do not conform to physical laws.
[0064] To ensure that the generated simulated defect images possess microscopic pixel accuracy, macroscopic structural realism, and physical consistency of the underlying mechanism, a three-dimensional collaborative composite objective loss function is designed. The overall loss function of the defect generation model G includes pixel-level reconstruction loss, perceptual loss, and physical constraint loss. During simulation training, the defect generation model G is trained with (I_norm, θ) as input and I_defect_gen as output target. The loss function includes pixel-level reconstruction loss (L1 / L2), perceptual loss (VGG feature distance), and physical constraint loss (penalizing combinations of θ that do not conform to physical laws). The pixel-level reconstruction loss is used to strongly constrain the absolute spatial difference between the simulated defect image and the target ground truth at the low-level pixel level, ensuring the absolute accuracy of the defect's edge morphology and basic color difference. The pixel-level reconstruction loss can be either the L1 loss function or the L2 loss function. Its specific expression is: loss_pixel = L1_loss(I_defect_real, I_defect_gen).
[0065] Since pixel-level loss functions tend to result in smooth and blurry network outputs, a perceptual loss based on latent space feature distances from pre-trained deep networks is introduced to capture higher-order structural textures and visual semantic consistency. I_defect_real and I_defect_gen are simultaneously input into a pre-defined VGG feature extraction network to extract their distances in a specific feature map space. The specific calculation formula is: loss_feature = perceptual_loss(I_defect_real, I_defect_gen).
[0066] To prevent deep generative networks from getting bogged down in purely statistical pixel-based calculations and producing images that violate real-world industrial manufacturing principles, an explicit physical regularization term, known as physical constraint loss, is introduced. Its specific calculation formula is: loss_physics = physics_regularizer(θ).
[0067] Finally, to balance the aforementioned multimodal and multidimensional loss terms, the individual loss terms are weighted and linearly combined to construct the final overall composite loss function, total_loss. Its mathematical formula is: total_loss = loss_pixel + λ1 * loss_feature + λ2 * loss_physics, where λ1 and λ2 are preset adjustable hyperparameters, used as weighting balancing factors for perceptual loss and physical constraint loss, respectively, to dynamically adjust the contribution weights of each loss term to gradient updates during training.
[0068] According to the scheme in this application, physical-level interpretable diagnosis of defects is achieved. By generating inversion, the pixel information of defect images that is difficult to understand is transformed into a series of interpretable physical parameters (such as force direction, energy magnitude, and material response), making the diagnostic process transparent. It also overcomes the challenge of diagnosing defects with small samples. The generating inversion model mainly relies on normal samples and physical rules for training, and also has diagnostic potential for rare or novel defects, reducing the dependence on massive defect samples.
[0069] S120: Obtain the real defect image to be diagnosed, and use the defect generation inversion engine to invert the first physical parameter combination from the physical parameter space based on the real defect image and mapping relationship.
[0070] In one embodiment, a defect generation and inversion engine is used to invert a first combination of physical parameters from a physical parameter space based on a real defect image and a mapping relationship. Specifically, this includes: obtaining a normal product image corresponding to the real defect image; and using a defect generation model to perform inverse optimization search for the first combination of physical parameters in the physical parameter space, so that the simulated defect image generated based on the normal product image and the first combination of physical parameters is the image most similar to the real defect image.
[0071] First, obtain the actual defective product image I_defect_real to be diagnosed, and the corresponding normal product image template or reference area I_norm.
[0072] To infer the physically-compliant generation process parameters that cause a defect from a single defect image, the defect generation model is used to perform inverse optimization in the physical parameter space to search for the optimal best_θ, i.e., the first combination of physical parameters, such that I_defect_real is most similar to the generated I_defect_gen. This yields the final first combination of physical parameters, best_θ, and the corresponding I_defect_gen_best. A pseudocode example is shown below. Figure 6 As shown in the figure. Then, the first physical parameter estimate best_θ, the parameter uncertainty measure, and the simulated image I_defect_gen that is "most like" the real defect generated based on best_θ are output.
[0073] Traditional approaches directly regress defect physical parameters by training large multimodal models. However, this method has a fundamental limitation: its training process relies on a massive number of defect image samples with precisely labeled physical parameters. In the industrial manufacturing field, defect samples are extremely scarce, and labeling defects with physical parameters requires the intervention of domain experts, which is very costly and inconsistent. Furthermore, such models have poor generalization ability for unseen defect types, and the generation process is uninterpretable, making it difficult to implement this method in actual production.
[0074] According to the scheme in this application, the physical parameters obtained from the inversion of the defect generation model can be automatically synthesized from normal samples through physical simulation or programmed methods, without the need for real defect images. Furthermore, during the inference stage, physical parameters are solved from a single defect image through inversion optimization, eliminating the need for a defect sample library. This fundamentally overcomes the bottleneck of scarce industrial defect data and difficult annotation, enabling root cause diagnosis of defects with small or even zero samples.
[0075] The overall process of defect generation model training and defect generation inversion inference is as follows: Figure 7 and Figure 8 As shown.
[0076] like Figure 7 As shown, the system receives a normal image I_norm and physical parameters θ sampled from the physical parameter space. The normal image I_norm and physical parameters θ are combined and input into the conditional generation model. The conditional generation model uses its internally constructed mapping network G to perform forward feature transformation and defect synthesis, outputting a corresponding simulated defect-generated image, denoted as I_defect_gen = G(I_norm, θ). The target ground truth establishment branch obtains a simulated defect image I_defect_sim pre-generated by physical simulation software or procedural rendering methods and directly uses it as the supervision target image in the training phase, feeding it into the subsequent loss calculation module. The composite loss function calculation module receives the simulated defect-generated image I_defect_gen and the simulated defect image I_defect_sim, performs dynamic comparison and numerical calculation between the two, constructs a comprehensive loss function composed of MSE loss (mean squared error loss, used as pixel-level reconstruction loss), perceptual loss, and physical constraint terms, and uses this function to calculate the total loss gradient and perform backpropagation until the conditional generation model converges, thus completing the training evolution of the forward mapping network.
[0077] like Figure 8As shown, in the inference and diagnosis phase after model training, a defect generation and inversion engine is used to solve the problem inversely from real images. The defect generation and inversion engine acquires the real defect image I_defect_real acquired online on the production line and the normal product image I_norm corresponding to the real defect image from the standard template library, and inputs them into the inversion optimization module. The inversion optimization module uses the normal product image I_norm as the reconstruction basis and starts inverse iterative optimization calculation in the continuous physical parameter space. Its specific optimization criterion is expressed as: min L(G(I_norm, θ), I_defect_real) + R(θ), where L(G(I_norm, θ), I_defect_real) is used to characterize the multi-scale difference loss between the simulated image reconstructed by the conditional generation model based on the current candidate parameters θ and the real defect image, and R(θ) is used to characterize the physical rationality constraint regularization term applied to the physical parameter space. The inversion optimization module performs a directional search along the gradient direction in the parameter space by continuously minimizing this objective function. When the iteration meets the preset termination convergence condition, the inversion optimization module completes the calculation and outputs the final result. The optimal parameter best_θ is then fed back into the model to reconstruct and generate an optimal generated defect image I_defect_gen_best that is visually and stylistically highly consistent with the current real defect for on-screen verification and interactive presentation.
[0078] S130 utilizes a causal reasoning engine to perform joint reasoning based on the associated device data corresponding to the real defect image and the combination of the first physical parameters.
[0079] Production sites contain multi-source heterogeneous data, such as visual images, time-series data from equipment sensors, process logs from the Manufacturing Execution System (MES), and material batch information from Enterprise Resource Planning (ERP). Currently, there is a lack of effective technology to deeply integrate and causally analyze this data with specific defect instances, making it impossible to form a closed-loop traceability chain of "defect-equipment-process".
[0080] In one embodiment, the production path of the defective product is extracted from the manufacturing execution system to determine a causal time window based on the production identifier and timestamp of the defective product corresponding to the actual defect image. Related equipment data corresponding to the actual defect image is obtained, including time-series data and process logs of relevant equipment in the production path within the causal time window.
[0081] In actual industrial production, equipment generates massive amounts of time-series data, including a significant amount of noisy data unrelated to the current defect. To achieve accurate root cause tracing, the causal inference engine first performs a rigorous spatiotemporal data alignment process. When the system receives an image of a real defect to be diagnosed, it first obtains the unique production identifier of the defective product (e.g., production batch number or specific panel serial number, Panel_ID). Based on the production identifier, the system extracts the complete production path of the product from the manufacturing execution system. The production path includes not only all production lines, processing equipment, and specific process chambers through which the product flows, but also precisely records the absolute timestamps of the product entering and leaving each piece of equipment and process step. Since a specific type of defect is usually caused by a specific process, it is not necessary to trace back the entire production cycle; instead, the "causal time window" that may lead to this type of defect is determined. For example, for a common stamping scratch defect, the system will focus on tracing back the stamping process and the time periods of the preceding and following handling processes. For example, in the manufacturing of organic light-emitting diode (OLED) panels, if a low-frequency, blurred-border murmur defect is detected, the knowledge base can determine that the defect is most likely caused by a uniformity problem in the thin film encapsulation (TFE) or organic light-emitting layer evaporation process. The analysis then focuses on these two processes and adjacent handling and cleaning steps, establishing the corresponding causal time window. After identifying the causal time window and related production equipment, the system further extracts high-frequency time-series data and process logs from the factory's underlying data platforms, such as supervisory control and data acquisition (SCADA), statistical process control (SPC), or equipment automation programs (EAP), within that specific time window.
[0082] To achieve accurate feature matching and causal probability calculation, a fault-defect causal knowledge base was pre-built and maintained. This knowledge base uses a directed graph (knowledge graph) data structure to structurally store the causal relationship network accumulated from domain expert experience and mined from massive historical production data through causal discovery algorithms. For example, for a typical machining scenario, the causal relationship path stored in the graph structure is defined as: (Fault mode node: 'Die wear') → [Cause] → (Defect physical feature node: {Scratch direction fixed, depth progressively increasing}) → [Associated equipment] → (Equipment: 'Pressing machine A - Upper die').
[0083] After acquiring associated device data and constructing a graph-structured knowledge base, the causal inference engine executes probability matching and inference steps. First, feature matching is performed, calculating and matching the optimal combination of physical parameters output by the defect generation inversion engine (i.e., the first combination of physical parameters, such as the parameter combination best_θ with the characteristics of "deep and fixed direction") with the typical defect physical features corresponding to various candidate fault modes in the fault-defect causal knowledge base, thereby solving for the preliminary probability P(Fault|best_θ) of each candidate fault mode. Next, data evidence fusion is performed, further checking for anomalies in the sensor time-series data of each candidate associated device within the aforementioned causal time window. For example, if feature matching initially suspects the fault mode as "mold wear," the system will focus on retrieving and checking the vibration signal spectrum of the stamping machine for wear-related characteristic frequencies and whether their energy is abnormally elevated. The system quantifies the degree of anomaly in the extracted device data and converts it into the corresponding data evidence strength P(Data|Fault).
[0084] Finally, a Bayesian update is performed, combining the prior probability (e.g., the basic failure probability P(Fault) calculated based on the device's accumulated runtime or historical failure rate), the feature matching probability P(best_θ|Fault), and the aforementioned data evidence strength P(Data|Fault), to calculate the final posterior probability of each candidate failure mode using the Bayesian update algorithm. The specific posterior probability derivation is proportional to the following formula: P(Fault|best_θ, Data) ∝ P(best_θ|Fault) × P(Data|Fault) × P(Fault).
[0085] Existing models are "black boxes," lacking physical interpretability, and their decision-making processes are difficult to understand. When models misdetect or miss, engineers cannot understand the reasons from a physical perspective, and it is difficult to trust the model's judgment, hindering the in-depth application of the technology.
[0086] In one embodiment, a causal reasoning engine is used to perform joint reasoning based on the associated equipment data corresponding to the real defect image and the combination of first physical parameters. Specifically, this includes: constructing a dynamic causal graph, which includes a physical causal layer, an intermediate phenomenon layer, and a defect manifestation layer. The physical causal layer contains equipment state nodes and / or process parameter nodes representing observable or controllable entities on the production line; the intermediate phenomenon layer contains phenomenon nodes representing invisible physical or chemical changes occurring in the product during processing; and the defect manifestation layer contains defect nodes representing the final observable defect characteristics of the product. Associated equipment data is input to the nodes corresponding to the physical causal layer, and the combination of first physical parameters is input to the defect nodes corresponding to the defect manifestation layer. A probabilistic reasoning algorithm is run on the dynamic causal graph to solve for the posterior probability of the faulty production equipment and the fault mode, and a root cause diagnosis report is output.
[0087] To provide engineers with a visually intuitive and highly interpretable reasoning process, a dynamic evidence chain visualization causal graph is constructed based on the defect cause evidence chain determined in the aforementioned probabilistic reasoning stage. This dynamic causal graph is designed with a bottom-up, three-layer progressive structure, including: a physical causal layer, an intermediate phenomenon layer, and a defect manifestation layer.
[0088] The first layer is the physical causal layer, which includes physical entities on the production line that can be directly observed or controlled, such as equipment operating status, process parameter settings, and environmental factors. This layer design strictly follows the objective physical causal chain of industrial defects, giving the entire causal diagram strong physical interpretability.
[0089] This provides physically reasonable operational control points for counterfactual reasoning and human intervention analysis, enabling the system not only to diagnose the past but also to answer practical production simulation questions such as "how will the defect morphology change if specific process parameters of a certain piece of equipment are changed?" It greatly enhances the robustness and generalization ability of the cause-effect graph structure because physical objective laws are relatively stable; even if the appearance of specific defects at the upper level changes, it will not easily shake the underlying structural foundation of the cause-effect graph. It achieves a fundamental technological leap from simple "data association mining" to "rigorous causal diagnosis." The root cause results output by the system can directly guide on-site physical operations (such as adjusting specific equipment parameters or replacing worn parts), thus forming a complete intelligent decision-making closed loop in the industrial field.
[0090] The second layer is the intermediate phenomenon layer, which mainly describes the invisible physical or chemical changes that occur during the product's processing and manufacturing. It is a necessary bridge connecting the underlying causal entities and the top-level defect manifestations, including intermediate phenomenon nodes such as "uneven distribution of organic film thickness" and "plastic deformation of material surface".
[0091] The third layer is the defect representation layer, which corresponds to the physical features of defects on the surface of the final product that can be directly observed. The nodes in this layer also serve as the top input interface for the interaction between the visual inspection system and the inversion engine.
[0092] refer to Figure 9 , Figure 9 This embodiment illustrates the specific structure and associated paths of the dynamic evidence chain visualization causal graph. In the dynamic causal graph, nodes represent specific production variables, and directed edges represent causal relationships between variables. Each node is associated with a conditional probability table (CPT) to quantify the influence of its parent node on the node's state.
[0093] In order to correspond with the aforementioned three-layer progressive network structure, Figure 9 The layer type of each node is distinguished by different attributes: the root node represents the physical entity and device state in the physical cause-and-effect layer; process nodes represent process states or parameter nodes (such as film thickness, packaging quality, etc.) in the intermediate phenomenon layer; and defect nodes represent specific defect features (such as murmurs, scratches, etc.) in the defect manifestation layer. Figure 9 The following are three typical causal transmission paths as examples. Path 1 (Evaporator Temperature → Film Thickness → Mura Defect): When the bottom-level device status node "Evaporator Temperature" becomes abnormal and persists for 30 minutes, it will cause uneven distribution of the middle-level node "Film Thickness" (corresponding conditional probability / correlation strength is 0.8); and after "uneven film thickness" persists for 2 hours, it is very likely to cause the top-level node to exhibit "Mura" defects (corresponding correlation strength is 0.9). Path 2 (Encapsulation Chamber Pressure → Encapsulation Quality → Mura Defect): After the bottom-level node "Encapsulation Chamber Pressure" deviates for 1 hour, it will cause the middle-level node "Encapsulation Quality" to decrease (correlation strength is 0.6); after the decrease in encapsulation quality persists for 3 hours, it will cause the top-level node to exhibit "Mura" defects (correlation strength is 0.5). Path 3 (Handling Robot Vibration → Scratch Defect): After the bottom-level node "Robot Arm Vibration" becomes abnormal and persists for 10 minutes, it can directly cross layers and cause the top-level node to produce "Scratch" defects (correlation strength is 0.7).
[0094] In one embodiment, production data from the production line is collected in real time, and the conditional probability table of each node in the dynamic causal graph is dynamically updated using a Bayesian update algorithm. A causal discovery algorithm is run using production data accumulated over a preset historical time period to automatically add or delete nodes in the dynamic causal graph and / or adjust the causal relationships between nodes, thereby achieving dynamic updates to the dynamic causal graph.
[0095] Considering that the actual physical causal relationships in industrial production processes may drift or change over time (e.g., due to equipment component wear and aging, or production process optimization and improvement), the graph structure in this embodiment is designed as a dynamically adjustable network, allowing for self-evolution and updates as the system acquires new evidence. Specifically, based on the above structure, the specific construction and operation process of the visualized causal graph mainly consists of the following three key stages: Phase 1: Constructing the initial cause-effect graph.
[0096] Step 11 (Define Nodes): Based on the knowledge of industry experts and the results of historical production data mining, identify the key influencing factors (including equipment status, process parameters, environmental factors, etc.) related to the target defect, and abstract them into random variables (i.e., graph nodes) in the network.
[0097] Step 12 (Determine Node State): Define a discrete or continuous operating state for each node. In a preferred embodiment, discrete states are used for quantization (e.g., divided into: normal / abnormal, or high / medium / low state levels).
[0098] Step 13 (Determine causal relationships): Based on the guidance of domain experts or by running a causal discovery algorithm using historical data, preliminarily determine the causal transmission direction and relationship between each node, forming directed edges in the graph structure.
[0099] Step 14 (Learning the Conditional Probability Table): Using massive amounts of historical production data, the conditional probability table for each node is trained and learned through statistical methods or machine learning models, thus completing the construction of the initial causal graph network.
[0100] Phase 2: Dynamically update the cause-effect graph.
[0101] Step 21 (Collecting New Evidence Data): The system obtains the latest operational data in real time from the data monitoring and acquisition system on the production line, including the latest equipment sensor timing data, process parameter execution records, and visual inspection results.
[0102] Step 22 (Update Probability Table): Using the latest collected evidence data, dynamically update the probability weight values in the conditional probability table of each node through a Bayesian update algorithm or an online learning algorithm.
[0103] Step 23 (Structure Adaptive Update, Optional): Based on a preset maintenance cycle (e.g., weekly or monthly), rerun the causal discovery algorithm using production data accumulated over a recent period to check whether the production line has evolved new causal relationship paths or whether old causal connections have become invalid, and adaptively adjust the topology of the causal graph accordingly (e.g., add or delete nodes, adjust connections).
[0104] Phase 3: Reasoning for the root cause of the failure.
[0105] Step 31 (Evidence State Setting): Using the current real defect image features to be diagnosed (or the first combination of physical parameters obtained by inversion) and the abnormal device data extracted from the previous spatiotemporal alignment step as input evidence, set the exact observation state of the corresponding node in the dynamic causal graph.
[0106] Step 32 (Peripheral Probability Calculation): Run the Bayesian network inference algorithm on the dynamic causal graph after setting the evidence, and use probability propagation to calculate the final posterior probability of each possible potential root cause node (i.e., the underlying device state node) causing the defect.
[0107] Step 33 (Sorting and Output): Sort the candidate root cause nodes in descending order according to the calculated posterior probability, and finally lock in the list of most likely root causes of the failure and assign corresponding diagnostic confidence.
[0108] This solution breaks down the barriers between visual data and equipment time-series data and process data, enabling cross-modal joint reasoning and improving the comprehensiveness and accuracy of diagnosis.
[0109] S140, output the root cause diagnosis report.
[0110] In one embodiment, the interactive interface outputs a real defect image, a simulated defect image obtained by the defect generation model based on a combination of first physical parameters, a dynamic causal graph used to characterize the decision evidence chain of the causal reasoning engine, and maintenance suggestions.
[0111] After completing the core reasoning process described above, the system executes a comprehensive root cause diagnosis report output to support closed-loop maintenance decisions at the factory site. Specific output content and display format include: a list of potential faulty equipment and patterns sorted by confidence level. For example, the system output format is [('Pressing press 03 - Upper die wear', confidence level 87%), ('Conveyor belt 05 - Position offset', confidence level 12%)]. Furthermore, key relevant data evidence points supporting the reasoning conclusion are attached to this list (e.g., highlighted: "Data evidence: Vibration peak value of pressing press 03 exceeded the standard by 150% at frequency X").
[0112] The root cause diagnosis report also includes a chain of evidence visualization diagram, which is a dynamic causal diagram that visually displays fault-highlighted paths and probability propagation strengths on the interface.
[0113] The root cause diagnosis report also includes specific repair recommendations, which are automatically generated from a knowledge base to provide standard operating procedures (SOPs) or repair guidance for specific failure modes.
[0114] When displayed on the front-end user interface, a multimodal simultaneous view mode is adopted, comprehensively displaying the following elements: the actual defect image to be diagnosed, the simulated defect image obtained by forward rendering based on the combination of the first physical parameters derived from the defect generation model (used to intuitively verify the accuracy of the inversion parameters), a visualized association diagram of the evidence chain to characterize the system's decision-making path, and a structured maintenance suggestion work order. Through the collaborative display of the above interface information, the pain point of traditional black-box diagnostic models being difficult to interpret is completely solved, providing on-site maintenance engineers and process engineers with panoramic, traceable, and highly reliable intelligent decision support. The diagnostic results not only provide information on potentially faulty equipment but also include confidence levels, related evidence (such as abnormal sensor readings), and specific maintenance suggestions, directly supporting on-site actions.
[0115] The following describes the complete application process for root cause tracing of "Mura" (display non-uniformity) defects in the high-precision display panel manufacturing industry. Panel manufacturing is a typical scenario with extremely complex processes and "zero tolerance" for defects. Any slight fluctuation in process parameters can lead to costly scrapping of finished products. This embodiment will demonstrate in detail how the system integrates visual data with underlying equipment data to perform intelligent closed-loop diagnosis.
[0116] The product in question is an OLED display panel for a certain model of mobile phone. After the "Aging Test" process, the automated optical inspection (AOI) equipment discovered a panel with a blurred, irregular bright spot-type Mura defect in the lower right corner of the display under a low grayscale (Gray Level 32) display.
[0117] The causes of Mura defects are extremely complex, potentially involving dozens of processes across multiple manufacturing stages, including the backlight module, thin-film transistor (TFT) array, organic light-emitting layer, and encapsulation layer. Traditional troubleshooting methods require multiple engineers from different departments to access and collaboratively analyze massive amounts of data, a process that can take hours or even days.
[0118] like Figure 10 As shown, the defect generation root cause tracing system provided in this application includes a data and knowledge layer, a core intelligent engine layer, and an application and interaction layer. The data and knowledge layer is responsible for accessing defect image streams, equipment sensor data, and MES / SCADA system process logs, and storing equipment failure mode knowledge bases and physical rule bases. The core intelligent engine layer includes three core modules: a defect generation inversion engine, a multi-source data association and causal reasoning engine. The application and interaction layer provides functions such as a root cause visualization interface, diagnostic report generation, and automatic maintenance work order creation.
[0119] The defect generation root cause tracing method and system workflow provided in this application are as follows: First, defect capture and process triggering are performed. The online AOI system captures the actual defect image (I_defect_real) of the defective panel and accurately marks the defective area in the image. After receiving the image, this system automatically obtains the panel's unique production serial number (e.g., Panel_ID: PXL-20240520-789) and automatically triggers the subsequent root cause tracing process.
[0120] Next, defect image inversion analysis is performed. Based on the Panel_ID, the system automatically retrieves the corresponding lower right corner image of a defect-free panel of the same model and batch from the standard image library, using it as the normal product template I_norm. The system synchronously inputs the real defect image I_defect_real and the normal product image template I_norm into the physically parameterized defect generation and inversion engine. The defect generation and inversion engine performs inverse optimization in the set physical parameter space, outputting as follows: Figure 11 The optimal combination of physical parameters (i.e., the first combination of physical parameters, best_θ) is shown. Simultaneously, the defect generation and inversion engine forward synthesizes a simulated Mura defect image, I_defect_gen_best, based on I_norm and best_θ. This simulated image highly matches the real defect in visual morphology and deep features, and will be retained for subsequent visualization and comparison verification on the terminal interface.
[0121] Then, multi-source data correlation is performed. The system first extracts the complete production history of the specific panel from the Manufacturing Execution System (MES) using the Panel_ID (e.g., the flow process is: cleaning, photolithography, etching, evaporation, and encapsulation). This history includes all the equipment and process chambers the panel passed through, as well as the precise absolute timestamps of entry and exit. Subsequently, based on a preset knowledge base, the system determines that this type of "low-frequency, blurred-boundary bright spot mura" is most likely caused by uniformity issues in the thin-film encapsulation TFE or organic light-emitting layer evaporation process. Accordingly, the system establishes a causal time window of [0, 2], locking it within a specific hourly interval before and after the process occurs, precisely focusing the analysis on these two core processes and the adjacent handling and cleaning steps.
[0122] Next, feature extraction is performed. Based on the aforementioned spatiotemporal alignment information, the system selectively extracts high-frequency sensor data and process logs from the factory's underlying data platform (such as the Statistical Process Control System (SPC) and Equipment Automation Program (EAP)) for the period when the panel passes through relevant equipment. The specific extraction scope includes: evaporation rate, crucible temperature, substrate temperature, QCM readings of the evaporator, and encapsulation material spraying pressure, nozzle movement path, UV curing intensity, cavity humidity, as well as positioning accuracy parameters and vibration sensor data of the relevant handling robots.
[0123] Next, causal reasoning and root cause matching are performed. The system queries the pre-built fault-defect causal knowledge base for the optimal combination of physical parameters, best_θ. The knowledge base rule response indicates that the features of "low frequency, blurry bright spots" have the highest correlation with "uneven vapor deposition film thickness" or "locally excessively thin encapsulation layer thickness" (prior matching degree reaches 0.85); the feature of "blurred boundary" has a corresponding correlation with "uneven encapsulation layer" (match degree 0.6). Then, evidence fusion and data inspection are performed. The vapor deposition machine data is checked: the system comparison found that when the panel is in the vapor deposition chamber, the temperature of the linear vapor deposition source nozzle responsible for the lower right corner of the panel showed an instantaneous fluctuation of ±3°C, which significantly exceeded the preset process control upper limit (±1.5°C). The underlying causal graph of the knowledge base shows that this physical anomaly will directly lead to an abnormal vapor deposition rate of organic materials in this area, thus forming uneven film thickness (intermediate phenomenon). Checking the packaging equipment data: Cross-validation of sensor data shows that the packaging spraying pressure and movement path are within the standard tolerance range, and the humidity inside the cavity is stable with no obvious abnormalities. Checking the handling data: No abnormal vibration frequency bands were found in the robot's time-series data, directly ruling out the possibility of material physical structure damage caused by vibration during handling. Historical pattern comparison: The system searched the historical data pool and found that within the past week, the same linear source nozzle of the same vapor deposition machine had experienced two similar temperature fluctuation alarms; and although the panels produced after the first two alarms did not report fatal defects in the final inspection, their photoelectric test parameters (such as brightness uniformity) showed a slight deterioration trend. Then, Bayesian probabilistic inference was performed: The system input the above evidence into the dynamic causal graph, combined the feature matching degree of the inverted physical parameters (weight 0.8), the evidence of abnormal intensity of real-time equipment data (weight 0.9), and the prior probability obtained based on historical retrieval (weight 0.7), and finally calculated the posterior probability of each hypothesis through the Bayesian update algorithm: Root cause hypothesis A: The temperature control of the third linear source nozzle of vapor deposition machine #07 is unstable, with a corresponding diagnostic confidence level of 94%. Root cause hypothesis B: Uneven localized coating of encapsulation material, with a diagnostic confidence level of 5%. Other causes, with a corresponding diagnostic confidence level of 1%.
[0124] like Figure 12 As shown, the final integrated output and decision support system automatically generates a structured diagnostic report containing a chain of evidence based on the reasoning results, and pushes it to the interactive terminals of on-site maintenance engineers and process engineers in real time. The terminal interface displays the real panel defect diagram, the inverted simulated defect diagram, and a visualized dynamic cause-effect diagram highlighting the transmission path of "abnormal temperature of the vapor deposition machine → uneven film thickness → Mura defect" on the same screen, and directly outputs the corresponding maintenance guidance work order, thereby guiding engineers directly to the vapor deposition machine #07 for nozzle maintenance, completing the intelligent decision-making closed loop from "defect perception" to "precise maintenance".
[0125] This embodiment demonstrates the powerful application value in the technology-intensive, lengthy, and defect-cost-high panel manufacturing industry. Through an automated process of "defect physical inversion → production history tracing → multi-source data intelligent correlation → causal probabilistic reasoning," the system accurately locates fuzzy visual defects (Mura) to specific production equipment (evaporation machine) and components (linear source nozzles) within minutes, providing clear repair guidance. This not only significantly shortens the mean time to repair and avoids potential quality risks for entire batches of panels, but also provides direct data-driven evidence for process optimization and predictive maintenance, achieving a leap from "perceiving defects" to "understanding root causes."
[0126] It should be understood that the technical solution of this application is not only applicable to the aforementioned high-precision display panel manufacturing industry, but can also be widely applied and transferred to any other intelligent manufacturing, discrete manufacturing, or long-process technology field with observable defect manifestations, quantifiable physical / mechanical generation mechanisms, and multi-source heterogeneous production line data characteristics. Typical application areas and scenarios of the solution of this application include, but are not limited to, semiconductor manufacturing and advanced packaging, new energy lithium battery manufacturing, automobile manufacturing and heavy industry stamping or welding, high-precision optical components and glass substrate manufacturing, and steel metallurgy and continuous rolling.
[0127] Furthermore, based on the defect generation root cause tracing method proposed in the embodiments of this application, the embodiments of this application also propose a defect generation root cause tracing system for implementing the defect generation root cause tracing method.
[0128] Figure 13 This is a schematic diagram of a defect generation root cause tracing system 200 proposed in an embodiment of this application.
[0129] refer to Figure 13 As shown, the defect generation root cause tracing system may include: a defect generation inversion module 210 and a causal reasoning module 220.
[0130] The specific functions of each module can be found in the previous text, and will not be repeated here.
[0131] Furthermore, this application also proposes a defect generation root cause tracing device, which includes a processor and a memory connected together. The memory is used to store program code, and the processor is used to call the program code to execute any of the defect generation root cause tracing methods proposed in this application.
[0132] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.
[0133] This application also provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.
[0134] Although this application has been described with reference to preferred embodiments, various modifications can be made thereto and components or steps can be replaced with equivalents without departing from the scope of this application. In particular, the technical features mentioned in the various embodiments can be combined in any manner, provided there is no conflict in structure or method steps. This application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
[0135] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for tracing the root causes of defects, characterized in that, The method includes: A defect generation model is obtained, which is used to synthesize a simulated defect image based on a normal product image and physical parameters in a set physical parameter space, so as to construct a mapping relationship between the simulated defect image and the physical parameters; Acquire the real defect image to be diagnosed, and use the defect generation inversion engine to invert the first physical parameter combination from the physical parameter space based on the real defect image and the mapping relationship; Using a causal reasoning engine, a joint reasoning is performed based on the associated equipment data corresponding to the real defect image and the first physical parameter combination to output a root cause diagnosis report. The root cause diagnosis report includes the faulty production equipment and fault mode that caused the defect.
2. The method according to claim 1, characterized in that, The defect generation model is constructed based on a diffusion model, and includes a physical parameter encoder, a conditional fusion module, and a multi-level conditional injection network, wherein: The physical parameter encoder is used to convert the physical parameters into physical embedding vectors; The conditional fusion module is used to obtain text feature embeddings, fuse the text feature embeddings with the physical embedding vectors, and generate fused conditional embeddings. The multi-level conditional injection network is used to perform conditional injection based on the fusion conditional embedding of the normal product image to generate the simulated defect image.
3. The method according to claim 2, characterized in that, The conditional fusion module is used to fuse the physical embedding vector with the text feature embedding to generate a fused conditional embedding, specifically including: Through the first attention stream, attention is calculated using the text feature embedding as the query vector and the physical embedding vector as the key vector and value vector to extract the constraint features of physical conditions on the text position. Through the second attention stream, the text feature embedding is used simultaneously as a query vector, key vector and value vector to perform self-attention calculation, and the semantic dependency features between text features are extracted. After the constraint features and semantic dependency features are fused by gating weight calculation, they are then combined with the text feature embedding through residual connection and layer normalization to output the fused conditional embedding.
4. The method according to claim 2 or 3, characterized in that, The multi-level conditional injection network includes a cross-attention layer and a feature affine modulation layer. The multi-level conditional injection network is used to perform conditional injection based on the fused conditional embedding of the normal product image, specifically including: During the denoising process, the latent representation of the normal product image is used as the basis. High-level semantic guidance is performed through the cross-attention layer using the fusion conditional embedding, and spatial feature modulation is performed through the feature affine modulation layer using the fusion conditional embedding to complete the conditional injection.
5. The method according to any one of claims 1-4, characterized in that, The defect generation model is trained using a training loss function, which includes: Pixel-level reconstruction loss used to constrain the pixel differences between the synthesized simulated defect images and the training ground values; Perceptual loss used to constrain the distance between the simulated defect image and the training ground value in the high-level feature space; and Physical constraint loss is used to penalize combinations of physical parameters that do not conform to the laws of physics.
6. The method according to any one of claims 1-5, characterized in that, Using a defect generation and inversion engine, the first combination of physical parameters is inverted from the physical parameter space based on the real defect image and the mapping relationship. Specifically, this includes: Obtain a normal product image corresponding to the actual defect image; The defect generation model is used to perform inverse optimization search of the first combination of physical parameters in the physical parameter space, so that the simulated defect image generated based on the normal product image and the first combination of physical parameters is the image most similar to the real defect image.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Based on the production identifier and timestamp of the defective product corresponding to the real defect image, the production path of the defective product is extracted from the manufacturing execution system to determine the causal time window. Obtain the associated equipment data corresponding to the real defect image. The associated equipment data includes the time sequence data and process logs of the relevant equipment in the production path within the causal time window.
8. The method according to any one of claims 1-7, characterized in that, The step of using a causal reasoning engine to perform joint reasoning based on the associated device data corresponding to the real defect image and the first physical parameter combination specifically includes: A dynamic cause-effect graph is constructed, which includes a physical cause-effect layer, an intermediate phenomenon layer, and a defect manifestation layer. The physical cause-effect layer contains equipment state nodes and / or process parameter nodes that characterize observable or controllable entities on the production line. The intermediate phenomenon layer contains phenomenon nodes that characterize invisible physical or chemical changes that occur in the product during processing. The defect manifestation layer contains defect nodes that characterize the final observable defect features of the product. The associated device data is input to the node corresponding to the physical causal layer, and the first physical parameter combination is input to the defect node corresponding to the defect manifestation layer; a probabilistic inference algorithm is run on the dynamic causal graph to solve the posterior probability of the faulty production equipment and the fault mode and output the root cause diagnosis report.
9. The method according to claim 8, characterized in that, The method further includes: Production data from the production line is collected in real time, and the conditional probability table of each node in the dynamic causal graph is dynamically updated using a Bayesian update algorithm. The causal discovery algorithm is run using production data accumulated within a preset historical time period to automatically add or delete nodes in the dynamic causal graph and / or adjust the causal relationships between nodes, thereby achieving dynamic updates of the dynamic causal graph.
10. The method according to any one of claims 1-9, characterized in that, The output root cause diagnosis report specifically includes: The interactive interface outputs the real defect image, the simulated defect image obtained by the defect generation model based on the first combination of physical parameters, a dynamic causal graph to characterize the decision evidence chain of the causal inference engine, and maintenance suggestions.
11. A defect generation root cause tracing system, characterized in that, Includes modules or units for performing the method as described in any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method as described in any one of claims 1 to 10.
13. A computer program product, characterized in that, It includes instructions that, when executed by a processor, cause the method as described in any one of claims 1 to 10 to be performed.