Zero sample industrial anomaly detection method and system based on Gaussian mixture expert geometric constraint
By employing a mixture of Gaussian expert-provided geometric constraints method, the problem of insufficient generalization in industrial anomaly detection is solved, achieving stable and accurate anomaly detection under complex backgrounds, with particularly excellent performance in the detection of internal cavity surfaces of automotive valve bodies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA UNIV OF TECH
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies face the problem of insufficient generalization in industrial anomaly detection scenarios with scarce categories and uncertain anomaly types. In particular, when faced with the joint expression of discrete state semantics and continuous context semantics and nonlinear appearance changes, the detection accuracy is limited.
The method of geometric constraints of Gaussian mixture expert prompts is adopted. By constructing a Gaussian mixture expert prompt distribution model, semantic geometric constraints and fusion sampling are implemented to generate the final text prompt. Feature vectors are extracted using text encoder and image encoder, and cosine similarity is calculated for anomaly detection.
It improves the stability and generalization ability of multi-type anomaly detection, enhances the stability of model detection and localization in complex backgrounds, and improves detection accuracy and cross-domain and cross-class robustness.
Smart Images

Figure CN122048908A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nondestructive testing technology, and in particular to a zero-sample industrial anomaly detection method and system that combines Gaussian mixture expert-hint geometric constraints, which is especially suitable for anomaly detection of precision industrial components such as the inner cavity surface of automotive valve bodies. Background Technology
[0002] Industrial anomaly detection is a crucial aspect of quality control, widely applied in parts processing, service life monitoring, and product assembly inspection. It is particularly important in consumer electronics manufacturing and precision component quality control, where precise detection and location of anomalies are essential. However, industrial anomalies often face challenges such as scarce categories and uncertain anomaly types. Classification- or reconstruction-based detection methods frequently suffer from insufficient generalization when dealing with novel or rare anomaly categories.
[0003] In recent years, large-scale language models have demonstrated strong capabilities in language understanding and generation, driving the development of vision-language models such as CLIP, BLIP-2, and DINOv2. These models achieve zero-shot industrial anomaly detection in complex environments by calculating the similarity between images and text. Although various vision-language model methods have made progress, most still rely on deterministic textual cues, resulting in insufficient expressive power and limiting the generalization ability to unseen categories. In industrial scenarios, anomaly detection faces two main challenges: first, the joint representation of discrete state semantics and continuous contextual semantics; and second, nonlinear appearance changes (such as local deformation and texture defects) leading to geometric inconsistencies between text and visual features, thus affecting the stability and generalization of detection.
[0004] To address this, we study a zero-sample industrial anomaly detection method that combines Gaussian mixture expert suggestion geometric constraints to improve the stability and generalization ability of detecting multiple types of anomalies. Summary of the Invention
[0005] To address the aforementioned technical problems, the present invention aims to provide a zero-sample industrial anomaly detection method and system that combines Gaussian mixture expert hint geometric constraints, thereby solving the technical problems of insufficient generalization and limited detection accuracy faced by existing industrial anomaly detection methods in scenarios with scarce categories and uncertain anomaly types.
[0006] The objective of this invention is achieved through the following technical solution: A zero-sample industrial anomaly detection method with Gaussian expert geometric constraints includes the following steps: Step S10. Construct a Gaussian mixture expert cue distribution model to generate a Gaussian mixture expert distribution cue set adapted to image semantics; Step S20. Implement semantic geometric constraints and fusion sampling for prompts, adjust the geometric position of the semantic center of expert prompts, fuse the weights of each semantic expert through expert gating, and generate the final text prompt based on the reparameterized sampling strategy; Step S30. Use a text encoder and an image encoder to extract the feature vectors of the final text prompt and the input image, respectively; Step S40. Calculate the cosine similarity between the text feature vector and the image feature vector, and use it as the scoring basis for anomaly detection to achieve zero-sample industrial anomaly detection.
[0007] Further, in step S10: prompting via the given input text. Modeling the Gaussian mixture expert suggestion distribution specifically includes: Construct a set of context words , This is used to represent the local semantic differences in background, material, texture, and lighting of the internal cavity surface image of an automotive valve body; Construct a set of state words , This is used to represent the normal state semantics and abnormal state semantics of an image of the internal cavity surface of a vehicle valve body; where the normal state semantics are... Abnormal state semantics are And satisfy: ; Constructing category words , used to characterize the category of the detected object; The text prompt Includes normal text prompts and abnormal text prompts , Text prompt Represented as: (1).
[0008] Furthermore, text prompts are treated as random variables that change with the semantic context and state of the image. Construct Gaussian priors for context words and state words respectively; where: Contextual Gaussian prior construction includes: input image Gaussian function Contextual semantic center Contextual word diagonal covariance matrix Then the context words Gaussian prior The calculation formula is: (2); The Gaussian prior construction for state words includes: for each state word label Let the center of normal state semantics and abnormal state semantics be . The diagonal covariance matrix of the state words is ; by state word cue vector Contextual word suggestion vectors The fusion constitutes a set of mixed Gaussian expert distribution hints. Then the state word Gaussian prior The calculation formula is: (3); Gaussian mixture expert distribution hint set The calculation formula is as follows: (4); Among them, the status label is recorded. Status word prompt vector Contextual word suggestion vector The Gaussian mixture expert distributed cue set provides a distributed basis for cue semantic geometric constraints and fusion sampling.
[0009] Furthermore, step S20 indicates that semantic geometric constraints and fusion sampling include: nonlinear geometric adjustment and expert-gated fusion; The nonlinear geometric adjustment includes direction adjustment and amplitude correction, specifically: Distribute the Gaussian mixture expert cue set according to semantic source. Divide into multiple semantic experts, and provide suggestions from each expert in the suggestion semantic space. semantic center Diagonal covariance matrix Combining into semantic prototypes : (5); The direction after nonlinear geometric adjustment With nonlinear geometric adjustment amplitude The calculation formula is: (6); in, These are the learning weight transformation matrices, For sigmoid activation function, For hyperbolic tangent function, For direction offset term, For amplitude offset; The nonlinear geometrically adjusted transform amplitude modulation factor for: (7); The It is a constant; based on the direction adjusted by the nonlinear geometry. Nonlinear geometric adjustment amplitude With conversion amplitude modulation factor Update the semantic centers of each expert suggestion, then the geometrically adjusted semantic centers of each expert suggestion will be obtained. The calculation formula is: (8); The expert gating fusion constructs a Softmax-based gating logic function and dynamically calculates the gating weights of each expert: After obtaining the adjusted semantic centers from each expert, a gating method is used to assign weights to each semantic expert. Then expert gating weight The calculation formula is: (9); in, For gated logic functions, The text is a CLS token; the gating weights are adaptively adjusted according to changes in image semantic information: when the image semantic classification result tends to be normal, the weight of the normal state expert increases; when it tends to be abnormal, the weight of the abnormal state expert increases; the weight of the context word expert reflects the degree of influence of the environmental background.
[0010] Furthermore, to enable the model to adjust the role of each expert according to different semantic changes, a diagonal covariance matrix for each expert's prompts is generated. Geometrically adjusted expert-suggested semantic centers By incorporating expert weights, we obtain the semantic center after incorporating expert weights. With the diagonal covariance matrix ; The semantic center after expert weighting The calculation formula is as follows: (10); The diagonal covariance matrix after incorporating expert weights The calculation formula is as follows: (11); Based on the semantic center after fusion expert weights With the diagonal covariance matrix Gaussian priors with fused expert weights The calculation formula is: (12); Furthermore, the random sampling process is explicitly expressed as a semantic center after incorporating expert weights through reparameterization. Diagonal covariance matrix after fusion of expert weights The final cue vector is obtained by summing the standard Gaussian noise. ; Final hint vector The calculation formula is as follows: (13); in, It is Gaussian noise. For the number of samples, ; Will Mapping to context words Abnormal status words Normal state words Three steps to obtain the final text prompt. , ,but: (14); In the formula, For the set of context words, This is used to represent the local semantic differences in background, material, texture, and lighting of the internal cavity surface image of an automotive valve body; For a set of state words, , used to represent the normal state semantics and abnormal state semantics of the internal cavity surface image of an automotive valve body; For normal state semantics, It is an abnormal state semantic and satisfies: .
[0011] Furthermore, in step S30: The final processed text prompt Input text encoder Extracting text features The calculation formula is: (15); Input the image of the inner cavity surface of the automotive valve body. Input Image Encoder Extracting image features The calculation formula is: (16).
[0012] Further, in step S40, the text features Image features Cosine similarity The calculation formula is: (17); Obtain surface anomaly detection results.
[0013] Furthermore, for anomaly detection in the surface image of the internal cavity of the automotive valve body, the generation of the Gaussian distribution cue includes: Construct normal experts and abnormal experts using the status terms "normal" and "abnormal" respectively; Construct a context expert using contextual terms such as "metallic luster", "cast texture", "internal cavity surface", "local deformation", "scratches", and "porosity"; By combining the image features of the inner cavity surface of the automotive valve body, the probability distribution of local deformation and texture defects is accurately modeled through the hybrid Gaussian expert prompt distribution model, generating Gaussian distributed prompts adapted to the automotive valve body inner cavity surface inspection task.
[0014] Furthermore, the method also includes: A multi-scale feature fusion mechanism is constructed to extract multi-scale image features at different levels of the image encoder and calculate the similarity with the corresponding scale text prompt features. The final anomaly detection score is obtained by weighted fusion of similarity scores at various scales, where shallow feature weights are used to detect subtle texture anomalies and deep feature weights are used to detect semantic-level anomalies.
[0015] A zero-sample industrial anomaly detection system with mixture Gaussian expert geometric constraints includes: The Gaussian mixture expert prompt generation module is used to construct a Gaussian mixture expert prompt distribution model to generate Gaussian distribution prompts that are adapted to the semantics of the image, including normal state expert units, abnormal state expert units and context word expert units; The geometric constraint and fusion sampling module is used to implement the geometric constraint and fusion sampling of the prompt semantics. It includes a geometric adjustment direction calculation unit, a geometric adjustment magnitude calculation unit, and an expert gating fusion unit, which are used to adjust the geometric position of the expert prompt semantic center and fuse the weights of each expert. The feature extraction module includes a pre-trained text encoder and an image encoder, which are used to extract feature vectors from the final text prompt and the input image, respectively. The anomaly scoring module is used to calculate the cosine similarity between text features and image features based on a reparameterized sampling strategy, and output anomaly detection scores.
[0016] Furthermore, the geometric adjustment direction calculation unit includes: Learnable weight transformation matrix, dimension ,in For feature dimensions; The Sigmoid activation function module is used to map the linear transformation result to the (0,1) interval and calculate the nonlinear geometric adjustment direction. The geometric adjustment amplitude calculation unit includes a hyperbolic tangent function module and a direction offset term configuration module, which are used to adapt to the nonlinear appearance change pattern of the internal cavity surface image of the automotive valve body.
[0017] Furthermore, the expert gate fusion unit includes: The weight calculation submodule dynamically calculates the weights of each expert based on the Softmax function; The adaptive adjustment submodule dynamically adjusts the weight allocation of normal experts, abnormal experts, and contextual word experts based on the image semantic classification results. The fusion output submodule merges the semantic centers of various experts using a weighted summation method to generate the final text prompt.
[0018] Furthermore, the system also includes: A dedicated interface module for inspecting the internal cavity surface of automotive valve bodies is used to receive image input of the internal cavity surface of automotive valve bodies and to perform optimized detection for local deformation, texture defects, scratches, and abnormal pore types. The real-time detection module is configured with reparameterized sampling times to balance calculation speed and anomaly scoring accuracy, meeting the real-time detection needs of industrial production lines.
[0019] Compared with the prior art, one or more embodiments of the present invention may have the following advantages: This invention proposes a zero-shot industrial anomaly detection method combining Gaussian mixture expert hints with geometric constraints. It aims to address the shortcomings of current zero-shot industrial anomaly detection methods, such as the inability of deterministic text hints to simultaneously address discrete-state semantics and continuous context semantics, as well as the lack of geometric constraints on hint semantics. This method introduces a distributed hint paradigm, upgrading the "single hint vector" to a learnable Gaussian mixture expert hint distribution. This allows the hints to cover not only discrete-state semantics but also adapt to continuous context semantic changes, improving the model's transferability and generalization ability from a semantic perspective. Furthermore, the geometric consistency constraint mechanism, by introducing manifold geometric constraints, shapes an "alignable geometry" for the hint semantic space, effectively enhancing cross-domain and cross-class robustness. Through expert-gated fusion and sampling, this method suppresses semantic drift and contextual perturbations at the geometric structure level, enhancing the model's anomaly detection and localization stability in the complex context of automotive valve body cavities. Finally, the distributed hints and geometric constraints are integrated into a unified module, achieving end-to-end optimization and demonstrating strong engineering application value. Attached Figure Description
[0020] Figure 1 This is a flowchart of zero-sample industrial anomaly detection based on hybrid Gaussian expert geometric constraints provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the Gaussian mixture expert hint distribution model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the principle of semantic geometric constraints and fusion sampling provided in this embodiment of the invention; Figure 4 This is a schematic diagram illustrating an application scenario for the detection of the inner cavity surface of an automotive valve body, provided by an embodiment of the present invention. Figure 5 This is a structural block diagram of the zero-sample industrial anomaly detection system provided in an embodiment of the present invention; Figure 6 This is a flowchart of the multi-scale feature fusion mechanism provided in the embodiments of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in further detail below with reference to the embodiments and accompanying drawings. Example 1
[0022] like Figure 1 As shown, this embodiment provides a zero-sample industrial anomaly detection method based on hybrid Gaussian expert geometric constraints, including the following steps: Step S10. Construct a Gaussian mixture expert cue distribution model to generate a Gaussian mixture expert distribution cue set adapted to image semantics; like Figure 2 As shown, a Gaussian mixture model is constructed with K experts, where expert 1 is the normal state expert, expert 2 is the abnormal state expert, and experts 3 to K are context word experts.
[0023] Define the set of state words The following text is a collection of words. and category word; the set of state words is used to represent the normal state semantics and abnormal state semantics of the surface image of the internal cavity of the automotive valve body; wherein the normal state semantics are: Abnormal state semantics are And satisfy: ; The aforementioned set of terms is used to represent the local semantic differences in background, material, texture, and lighting of the internal cavity surface image of the automotive valve body; Constructing category words , used to characterize the category of the detected object; The text prompt Includes normal text prompts and abnormal text prompts , Text prompt Represented as: (1) Treat text prompts as random variables that change with the semantic context and state of the image. Construct Gaussian priors for context words and state words respectively; where: Contextual Gaussian prior construction includes: input image Gaussian function Contextual semantic center Contextual word diagonal covariance matrix Then the context words Gaussian prior The calculation formula is: (2); The Gaussian prior construction for state words includes: for each state word label Let the center of normal state semantics and abnormal state semantics be . The diagonal covariance matrix of the state words is ; by state word cue vector Contextual word suggestion vectors The fusion constitutes a set of mixed Gaussian expert distribution hints. Then the state word Gaussian prior The calculation formula is: (3); Gaussian mixture expert distribution hint set The calculation formula is as follows: (4); Among them, the status label is recorded. Status word prompt vector Contextual word suggestion vector The Gaussian mixture expert distributed cue set provides a distributed basis for cue semantic geometric constraints and fusion sampling.
[0024] In this example, for the task of detecting anomalies on the surface of an automotive valve body cavity, the Gaussian distribution-based prompt generation strategy utilizes a close integration of state words, context words, and image features to accurately depict the statistical characteristics of local deformation and texture defects. Specifically, by modeling with a mixture of Gaussian expert prompt distribution, the state semantics (normal or abnormal) and context semantics (such as material, background, and lighting conditions) in the input text are transformed into a probability distribution. This not only covers deterministic semantics but also adapts to changes in continuous context, thereby enhancing the model's transferability and generalization ability at the semantic level. Subsequently, by introducing a nonlinear geometric adjustment mechanism to impose geometric consistency constraints on the expert prompt distribution, the geometric mismatch between the prompt semantic space and the image feature space is effectively resolved, improving the robustness of cross-domain and cross-class detection.
[0025] For the input file image First, initial features are extracted through an image encoder. The semantic centers of each expert are initialized as corresponding word vectors and dynamically adjusted based on image features during training.
[0026] Step S20. Implement semantic geometric constraints and fusion sampling for prompts, adjust the geometric position of the semantic center of expert prompts, fuse the weights of each semantic expert through expert gating, and generate the final text prompt based on the reparameterized sampling strategy; like Figure 3 As shown, step S20 includes: nonlinear geometric adjustment and expert gating fusion; The nonlinear geometric adjustment includes direction adjustment and amplitude correction, specifically: Distribute the Gaussian mixture expert cue set according to semantic source. Divide into multiple semantic experts, and provide suggestions from each expert in the suggestion semantic space. semantic center Diagonal covariance matrix Combining into semantic prototypes : (5); The direction after nonlinear geometric adjustment With nonlinear geometric adjustment amplitude The calculation formula is: (6); in, These are the learning weight transformation matrices, For sigmoid activation function, For hyperbolic tangent function, For direction offset term, For amplitude offset; The nonlinear geometrically adjusted transform amplitude modulation factor for: (7) The It is a constant; based on the direction adjusted by the nonlinear geometry. Nonlinear geometric adjustment amplitude With conversion amplitude modulation factor Update the semantic centers of each expert suggestion, then the geometrically adjusted semantic centers of each expert suggestion will be obtained. The calculation formula is: (8); The expert gating fusion constructs a Softmax-based gating logic function and dynamically calculates the gating weights of each expert: After obtaining the adjusted semantic centers from each expert, a gating method is used to assign weights to each semantic expert. Then expert gating weight The calculation formula is: (9); in, For gated logic functions, The text CLS token is used; the gating weights are adaptively adjusted according to changes in image semantic information: when the image semantic classification result tends towards the normal class, the weight of the normal state expert increases; when it tends towards the abnormal class, the weight of the abnormal state expert increases; the weight of the context word expert reflects the degree of influence of the environmental background. This dynamic weight allocation strategy can effectively identify and respond to key features in the image, ensuring the accuracy and flexibility of the model in handling different types of anomalies. Through the dynamic adjustment of expert weights, the model can more accurately capture state changes in the image and reduce misjudgments caused by contextual interference.
[0027] In this embodiment, the Gaussian mixture expert cue distribution model, through the combination of a learnable weight transformation matrix and a sigmoid activation function, achieves nonlinear geometric adjustment of the expert cue semantic center calculation direction. This adjustment directly reflects changes in image context and state semantics. Specifically, the weight transformation matrix can flexibly adjust the contribution level of each expert, while the sigmoid activation function ensures that the adjustment direction conforms to the probability distribution law, thereby effectively controlling the degree to which the cue semantics adapt to changes in image content. This mechanism enables the model to better cope with various types of anomalies in complex industrial scenarios, especially in the detection of anomalies on the surface of automotive valve bodies, accurately capturing subtle changes caused by contextual information, such as local deformation or texture defects, thereby improving the robustness and generalization of detection.
[0028] A hyperbolic tangent function and a specific directional bias term were selected to adapt to the nonlinear appearance variations of the automotive valve body cavity surface image. This setup allows the method to more accurately adjust the geometry of expert suggestions when dealing with complex situations such as local deformation and texture defects, thereby enhancing the robustness and accuracy of the detection process. Through the nonlinear transformation of the function, the expert suggestion distribution can better fit the changing trends of the actual semantics, which is particularly crucial when facing highly complex and subtle industrial anomaly detection scenarios. The appropriate selection of the directional bias term ensures more precise adjustment of the suggestion vector direction, enabling rapid responses to different anomaly types and contexts, effectively improving detection efficiency and recognition rate; it significantly enhances the performance of zero-shot industrial anomaly detection methods on automotive valve body cavity surface detection tasks, achieving more stable and broader anomaly category generalization capabilities.
[0029] To enable the model to adjust the role of each expert according to different semantic changes, a diagonal covariance matrix is generated for each expert's suggestions. Geometrically adjusted expert-suggested semantic centers By incorporating expert weights, we obtain the semantic center after incorporating expert weights. With the diagonal covariance matrix ; The semantic center after expert weighting The calculation formula is as follows: (10); Diagonal covariance matrix after incorporating expert weights The calculation formula is as follows: (11); Based on the semantic center after fusion expert weights With the diagonal covariance matrix Gaussian priors with fused expert weights The calculation formula is: (12); By reparameterizing, the random sampling process is explicitly expressed as a semantic center after incorporating expert weights. Diagonal covariance matrix after fusion of expert weights The final cue vector is obtained by summing the standard Gaussian noise. ; Final hint vector The calculation formula is as follows: (13); in, It is Gaussian noise. For the number of samples, ; Will Mapping to context words Abnormal status words Normal state words Three steps to obtain the final text prompt. , ,but: (14); In the formula, , This is a set of context words used to represent the local semantic differences in background, material, texture, and lighting of the internal cavity surface image of an automotive valve body; , This is a set of state words used to represent the normal and abnormal state semantics of images of the internal cavity surface of a car valve body; For normal state semantics, It is an abnormal state semantic and satisfies: .
[0030] Step S30. Use a text encoder and an image encoder to extract the feature vectors of the final text prompt and the input image, respectively; Specifically: The final processed text prompt Input text encoder Extracting text features The calculation formula is: (15); Input the image of the inner cavity surface of the automotive valve body. Input Image Encoder Extracting image features The calculation formula is: (16).
[0031] Step S40. Calculate the cosine similarity between the text feature vector and the image feature vector based on the medium parameterized sampling strategy, and use it as the scoring basis for anomaly detection to achieve zero-sample industrial anomaly detection.
[0032] The text features Image features Cosine similarity The calculation formula is as follows: (17).
[0033] The cosine similarity calculation method employs a reparameterization sampling strategy to form the sum of the semantic center and standard Gaussian noise. This strategy allows for the statistical estimation of the distance between text and image features. The reparameterization sampling mechanism dynamically adjusts the position of the expert prompt's semantic center, ensuring it closely follows changes in context words, abnormal state words, and normal state words. Furthermore, it introduces additional uncertainty to better simulate the semantic distribution in real-world industrial scenarios. Even when faced with nonlinear appearance changes or complex background interference, the model can effectively align the feature spaces of the text and image by adjusting the scale and direction of the Gaussian noise, thereby statistically capturing the similarity between text prompts and image features.
[0034] In this embodiment, the text encoder and image encoder employ pre-trained deep learning models, specifically the ViT architecture of the CLIP model. This architecture significantly enhances feature extraction capabilities, enabling more accurate capture of the semantic relationships between text and images. The ViT architecture divides the input into fixed-size patches and processes them as sequences, utilizing a self-attention mechanism to capture global features. This not only strengthens the model's understanding of input information but also ensures effective understanding and matching of previously unseen surface anomaly features, even in zero-shot industrial anomaly detection scenarios. Therefore, the solution in this embodiment can more accurately detect and identify surface anomalies in automotive valve bodies, improving detection accuracy and generalization ability. Of course, the ViT architecture is not the only option. In other embodiments, different deep learning models, such as CNNs or other Transformer variants, can be considered, as long as they provide sufficient feature abstraction capabilities and cross-modal matching efficiency.
[0035] In the above examples, the text prompt generation is not only based on the input text but also fully utilizes the image features extracted by the image encoder. This strategy, through the interaction between the semantic centers of context words and image features, enables the generated text prompts to more accurately reflect the contextual information of the actual image, significantly improving the relevance and effectiveness of the prompts. Specifically, by allowing the semantic centers of context words to dynamically respond to changes in image features, it is possible to capture details such as the background, material, and texture unique to the image, thereby generating text prompts that are more closely aligned with the current image scene. This dynamic adjustment mechanism not only enhances the model's understanding of the input image but also effectively overcomes the limitations of fixed text prompts in dealing with complex and ever-changing industrial scenarios. Especially in the detection of anomalies on the surface of automotive valve bodies, it can more robustly cope with various background interferences and uncertainties in anomaly types. Finally, through the adaptive text prompt generation method, the model can more accurately identify abnormal regions when calculating the cosine similarity between text features and image features, improving detection accuracy and stability.
[0036] Example 2: Application of automotive valve body cavity surface inspection like Figure 4 As shown, this embodiment addresses a specific application scenario for detecting abnormalities on the inner surface of an automotive valve body.
[0037] The internal cavity surface of automotive valve bodies is characterized by high reflectivity, complex curved surfaces, and casting textures. Anomalies include localized deformation, scratches, porosity, and textural abnormalities. Traditional methods are ill-suited to adapting to the diverse surface conditions and unseen anomaly types.
[0038] Specific implementation configuration: Normal expert, semantic center corresponds to "normal automotive valve body cavity surface"; abnormal expert, semantic center corresponds to "abnormal automotive valve body cavity surface"; context expert 1, corresponds to "metallic luster surface"; context expert 2, corresponds to "cast texture background"; context expert 3, corresponds to "local deformation features"; context expert 4, corresponds to "abnormal scratches"; context expert 5, corresponds to "abnormal porosity".
[0039] Geometric constraint parameter settings: The directional bias term is optimized to address the nonlinear variation characteristics of reflective metal surfaces, enhancing sensitivity to minor scratches; the weight transformation matrix is fine-tuned based on the valve body surface image after pre-training.
[0040] During detection: acquire images of the valve body cavity surface; calculate anomaly scores through steps S10-S40; when the score is below the threshold, output anomaly location and type suggestions.
[0041] Experimental validation: The method of this invention was validated on a test set containing 500 normal samples and 300 abnormal samples (covering scratches, pores, deformations, etc.). The area under the receiver operating characteristic curve of the present invention reached 0.94, which is significantly better than the 0.82 of the traditional fixed cue method, demonstrating its superior generalization ability under zero-sample conditions.
[0042] Example 3: Multi-scale feature fusion like Figure 5 As shown, to further enhance the detection capability of subtle anomalies, this embodiment introduces a multi-scale feature fusion mechanism, extracting features from multiple layers of the image encoder. For example, the shallow features of the 4th layer capture local texture and edge details, used to detect subtle scratches and dents; the mid-layer features of the 8th layer capture component structure and geometry, used to detect deformation and assembly errors; and the deep features of the 12th layer capture global semantic context, used to identify complex anomaly patterns, corresponding to receptive fields of different scales.
[0043] A Gaussian mixture of expert cue distributions is constructed for each scale, and a similarity score is calculated. The scores from each scale are then fused using learnable weights. Among these, a higher weight in the 4th layer of shallow features is beneficial for detecting subtle texture anomalies, while a higher weight in the 12th layer of deep features is beneficial for detecting semantic-level anomalies.
[0044] Multi-scale feature fusion employs the following approaches: a top-down path, which transmits deep semantic information to shallower layers to enhance the semantic consistency of detailed regions; lateral connections, which achieve alignment and complementarity of features at different levels; feature alignment, which addresses the spatial resolution differences of features at different scales; and cross-scale fusion, which generates a unified high-dimensional feature representation through Concat+Conv operations.
[0045] Hybrid Gaussian expert hint constraint integration: The fused visual features are combined with the core Gaussian distributed hints (contextual word prior, state word prior, geometric constraint sampling) for cosine similarity calculation, and finally an anomaly score map is output to achieve accurate localization.
[0046] Example 4: System Implementation like Figure 6 As shown, this embodiment also provides a zero-sample industrial anomaly detection system with hybrid Gaussian expert geometric constraints, including: The Gaussian mixture expert prompt generation module is used to construct a Gaussian mixture expert prompt distribution model to generate Gaussian distribution prompts that are adapted to the semantics of the image. It includes normal state expert units, abnormal state expert units, and context word expert units, and is used to maintain the learnable parameters of K experts.
[0047] The geometric constraint and fusion sampling module is used to implement semantic geometric constraints and fusion sampling for prompts; it includes: The geometric adjustment direction calculation unit realizes the composite calculation of the weight transformation matrix and the Sigmoid activation function; The geometric adjustment amplitude calculation unit realizes the composite calculation of the hyperbolic tangent function and the direction offset term; The expert-gated fusion unit uses the Softmax function to achieve dynamic weight calculation and fusion.
[0048] The feature extraction module includes a pre-trained text encoder and an image encoder, which are used to extract feature vectors from the final text prompt and the input image, respectively. The anomaly scoring module is used to calculate the cosine similarity between text features and image features based on a reparameterized sampling strategy, and output anomaly detection scores.
[0049] The geometric adjustment direction calculation unit includes: Learnable weight transformation matrix, dimension ,in For feature dimensions; The Sigmoid activation function module is used to map the linear transformation result to the (0,1) interval and calculate the nonlinear geometric adjustment direction. The geometric adjustment amplitude calculation unit includes a hyperbolic tangent function module and a direction offset term configuration module, which are used to adapt to the nonlinear appearance change pattern of the internal cavity surface image of the automotive valve body.
[0050] The expert gate fusion unit includes: a weight calculation submodule, which dynamically calculates the weights of each expert based on the Softmax function; an adaptive adjustment submodule, which dynamically adjusts the weight allocation of normal experts, abnormal experts, and context word experts according to the image semantic classification results; and a fusion output submodule, which fuses the semantic centers of each expert based on a weighted summation method to generate the final text prompt.
[0051] The dedicated interface module for inspecting the internal cavity surface of automotive valve bodies features an input / output interface optimized for specific application scenarios. It receives images of the internal cavity surface of automotive valve bodies and performs optimized detection for local deformation, texture defects, scratches, and porosity defects, outputting detection results and anomaly location heatmaps in real time.
[0052] The real-time detection module is configured with reparameterized sampling times to balance calculation speed and anomaly scoring accuracy, meeting the real-time detection needs of industrial production lines.
[0053] While the embodiments disclosed in this invention are as described above, the content is merely for the purpose of facilitating understanding of the invention and is not intended to limit the invention. Any person skilled in the art to which this invention pertains may make any modifications and variations in form and detail of the implementation without departing from the spirit and scope disclosed herein; however, the scope of patent protection for this invention shall still be determined by the scope defined in the appended claims.
Claims
1. A zero-sample industrial anomaly detection method with Gaussian expert geometric constraints, characterized in that, Includes the following steps: Step S10. Construct a Gaussian mixture expert cue distribution model to generate a Gaussian mixture expert distribution cue set adapted to image semantics; Step S20. Implement semantic geometric constraints and fusion sampling for prompts, adjust the geometric position of the semantic center of expert prompts, fuse the weights of each semantic expert through expert gating, and generate the final text prompt based on the reparameterized sampling strategy; Step S30. Use a text encoder and an image encoder to extract the feature vectors of the final text prompt and the input image, respectively; Step S40. Calculate the cosine similarity between the text feature vector and the image feature vector, and use it as the scoring basis for anomaly detection to achieve zero-sample industrial anomaly detection.
2. The zero-sample industrial anomaly detection method with mixture Gaussian expert geometric constraints according to claim 1, characterized in that, In step S10: prompts are given via input text. Modeling the Gaussian mixture expert suggestion distribution specifically includes: Construct a set of context words , This is used to represent the local semantic differences in background, material, texture, and lighting of the internal cavity surface image of an automotive valve body; Construct a set of state words , This is used to represent the normal state semantics and abnormal state semantics of an image of the internal cavity surface of a vehicle valve body; where the normal state semantics are... Abnormal state semantics are And satisfy: ; Constructing category words , used to characterize the category of the detected object; The text prompt Includes normal text prompts and abnormal text prompts , Text prompt Represented as: (1)。 3. The zero-sample industrial anomaly detection method with mixture Gaussian expert geometric constraints according to claim 2, characterized in that, Treat text prompts as random variables that change with the semantic context and state of the image. Construct Gaussian priors for context words and state words respectively; where: Contextual Gaussian prior construction includes: input image Gaussian function Contextual semantic center Diagonal covariance matrix of context words Then the context words Gaussian prior The calculation formula is: (2); The Gaussian prior construction for state words includes: for each state word label Let the center of normal state semantics and abnormal state semantics be . The diagonal covariance matrix of the state words is ; by state word cue vector Contextual word suggestion vectors The fusion constitutes a set of mixed Gaussian expert distribution hints. Then the state word Gaussian prior The calculation formula is: (3); Gaussian mixture expert distribution hint set The calculation formula is as follows: (4); Among them, the status label is recorded. Status word prompt vector Contextual word suggestion vector The Gaussian mixture expert distributed cue set provides a distributed basis for cue semantic geometric constraints and fusion sampling.
4. The zero-sample industrial anomaly detection method based on mixture Gaussian expert geometric constraints according to claim 3, characterized in that, The semantic geometric constraints and fusion sampling mentioned in step S20 include: nonlinear geometric adjustment and expert gating fusion. The nonlinear geometric adjustment includes direction adjustment and amplitude correction, specifically: Distribute the Gaussian mixture expert cue set according to semantic source. Divide into multiple semantic experts, and provide suggestions from each expert in the suggestion semantic space. semantic center Diagonal covariance matrix Combining into semantic prototypes : (5); The direction after nonlinear geometric adjustment With nonlinear geometric adjustment amplitude The calculation formula is: (6); in, These are the learning weight transformation matrices, For sigmoid activation function, For hyperbolic tangent function, For direction offset term, For amplitude offset; The nonlinear geometrically adjusted transform amplitude modulation factor for: (7); The It is a constant; based on the direction adjusted by the nonlinear geometry. Nonlinear geometric adjustment amplitude With conversion amplitude modulation factor Update the semantic centers of each expert suggestion, then the geometrically adjusted semantic centers of each expert suggestion will be obtained. The calculation formula is: (8); The expert gating fusion constructs a Softmax-based gating logic function and dynamically calculates the gating weights of each expert: After obtaining the adjusted semantic centers from each expert, a gating method is used to assign weights to each semantic expert. Then expert gating weight The calculation formula is: (9); in, For gated logic functions, The text is a CLS token; the gating weights are adaptively adjusted according to changes in image semantic information: when the image semantic classification result tends to be normal, the weight of the normal state expert increases; when it tends to be abnormal, the weight of the abnormal state expert increases; the weight of the context word expert reflects the degree of influence of the environmental background.
5. The zero-sample industrial anomaly detection method with hybrid Gaussian expert geometric constraints according to claim 4, characterized in that, To enable the model to adjust the role of each expert according to different semantic changes, a diagonal covariance matrix is generated for each expert's suggestions. Geometrically adjusted expert-suggested semantic centers By incorporating expert weights, we obtain the semantic center after incorporating expert weights. With the diagonal covariance matrix ; The semantic center after expert weighting The calculation formula is as follows: (10); The diagonal covariance matrix after incorporating expert weights The calculation formula is as follows: (11); Based on the semantic center after fusion expert weights With the diagonal covariance matrix Gaussian priors with fused expert weights The calculation formula is: (12); Furthermore, the random sampling process is explicitly expressed as a semantic center after incorporating expert weights through reparameterization. Diagonal covariance matrix after fusion of expert weights The final cue vector is obtained by summing the standard Gaussian noise. ; Final hint vector The calculation formula is as follows: (13); in, It is Gaussian noise. For the number of samples, ; Will Mapping to context words Abnormal status words Normal state words Three steps to obtain the final text prompt. , ,but: (14); In the formula, For the set of context words, This is used to represent the local semantic differences in background, material, texture, and lighting of the internal cavity surface image of an automotive valve body; For a set of state words, , used to represent the normal state semantics and abnormal state semantics of the internal cavity surface image of an automotive valve body; For normal state semantics, It is an abnormal state semantic and satisfies: .
6. The zero-sample industrial anomaly detection method with mixture Gaussian expert geometric constraints according to claim 1, characterized in that, In step S30: The final processed text prompt Input text encoder Extracting text features The calculation formula is: (15); Input the image of the inner cavity surface of the automotive valve body. Input Image Encoder Extracting image features The calculation formula is: (16)。 7. The zero-sample industrial anomaly detection method with mixture Gaussian expert geometric constraints according to claim 6, characterized in that, In step S40, the text features Image features Cosine similarity The calculation formula is: (17); Obtain surface anomaly detection results.
8. The zero-sample industrial anomaly detection method with mixture Gaussian expert geometric constraints according to claim 1, characterized in that, For anomaly detection in the surface image of the internal cavity of an automotive valve body, the generation of the Gaussian distribution cue includes: Normal experts and abnormal experts are constructed using the status terms "normal" and "abnormal" respectively; Contextual experts are constructed using contextual terms such as "metallic luster", "cast texture", "internal cavity surface", "local deformation", "scratches", and "porosity". By combining the image features of the inner cavity surface of the automotive valve body, the probability distribution of local deformation and texture defects is accurately modeled through the hybrid Gaussian expert prompt distribution model, generating Gaussian distributed prompts adapted to the automotive valve body inner cavity surface inspection task.
9. The zero-sample industrial anomaly detection method with hybrid Gaussian expert geometric constraints according to claim 1, characterized in that, The method further includes: A multi-scale feature fusion mechanism is constructed to extract multi-scale image features at different levels of the image encoder and calculate the similarity with the corresponding scale text prompt features. The final anomaly detection score is obtained by weighted fusion of similarity scores at various scales, where shallow feature weights are used to detect subtle texture anomalies and deep feature weights are used to detect semantic-level anomalies.
10. A zero-sample industrial anomaly detection system with mixture Gaussian expert geometric constraints, characterized in that, include: The Gaussian mixture expert prompt generation module is used to construct a Gaussian mixture expert prompt distribution model to generate Gaussian distribution prompts that are adapted to the semantics of the image, including normal state expert units, abnormal state expert units and context word expert units; The geometric constraint and fusion sampling module is used to implement the geometric constraint and fusion sampling of the prompt semantics. It includes a geometric adjustment direction calculation unit, a geometric adjustment magnitude calculation unit, and an expert gating fusion unit, which are used to adjust the geometric position of the expert prompt semantic center and fuse the weights of each expert. The feature extraction module includes a pre-trained text encoder and an image encoder, which are used to extract feature vectors from the final text prompt and the input image, respectively. The anomaly scoring module is used to calculate the cosine similarity between text features and image features based on a reparameterized sampling strategy, and output anomaly detection scores.
11. The zero-sample industrial anomaly detection system with hybrid Gaussian expert geometric constraints according to claim 10, characterized in that, The geometric adjustment direction calculation unit includes: Learnable weight transformation matrix, dimension ,in For feature dimensions; The Sigmoid activation function module is used to map the linear transformation result to the (0,1) interval and calculate the nonlinear geometric adjustment direction. The geometric adjustment amplitude calculation unit includes a hyperbolic tangent function module and a direction offset term configuration module, which are used to adapt to the nonlinear appearance change pattern of the internal cavity surface image of the automotive valve body.
12. The zero-sample industrial anomaly detection system with hybrid Gaussian expert geometric constraints according to claim 10, characterized in that, The expert gate fusion unit includes: The weight calculation submodule dynamically calculates the weights of each expert based on the Softmax function; The adaptive adjustment submodule dynamically adjusts the weight allocation of normal experts, abnormal experts, and contextual word experts based on the image semantic classification results. The fusion output submodule merges the semantic centers of various experts using a weighted summation method to generate the final text prompt.
13. The zero-sample industrial anomaly detection system with hybrid Gaussian expert geometric constraints according to claim 10, characterized in that, The system also includes: A dedicated interface module for inspecting the internal cavity surface of automotive valve bodies is used to receive image input of the internal cavity surface of automotive valve bodies and to perform optimized detection for local deformation, texture defects, scratches, and abnormal pore types. The real-time detection module is configured with reparameterized sampling times to balance calculation speed and anomaly scoring accuracy, meeting the real-time detection needs of industrial production lines.