Two-stage intelligent analysis method and system for small sample of casting defects

CN122175979BActive Publication Date: 2026-08-28NANCHANG HANGKONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610645485.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-28
Estimated Expiration
2046-05-12

AI Technical Summary

Technical Problem

传统方法通常采用常规图像处理算法或人工设计特征加分类器的方式,但这些方法在面对复杂的工业检测场景时,存在检测效率低、普适性差、难以应对挑战性的问题

Benefits of technology

[0008]与现有技术相比,本发明的有益效果是:通过构建类条件参考锚点并执行两阶段语义对齐训练,有效解决了小样本条件下同类缺陷样本不足、相似缺陷易混淆的问题,第一阶段引入混淆类间隔约束与全局回退原型,提升训练稳定性与样本利用效率,第二阶段结合低秩适配与封闭词表约束,增强语义融合与类别区分能力,推理阶段通过证据检索与一致性仲裁,输出缺陷结论、形貌及工艺证据,显著提高了检测精度、可解释性与工程可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122175979B_ABST
    Figure CN122175979B_ABST
Patent Text Reader

Abstract

The application provides a two-stage intelligent analysis method and system for small samples of casting defects, which comprises generating a topography text anchor point according to image visual features and generating a process cause text anchor point according to historical process data; mapping the class condition reference anchor point to a language embedding space; extracting real visual features, unfreezing low-rank adaptive parameters in a large language model, and obtaining an optimal detection model; based on the optimal detection model, performing similarity retrieval on the to-be-detected visual features and knowledge items to obtain a retrieval result, and calculating consistency, if the consistency is consistent, outputting a structured result of a defect category, evidence explanation and confidence. The application effectively solves the problems of insufficient similar defect samples and similar defects being easily confused under the condition of small samples.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A two-stage intelligent analysis method for small samples of casting defects, characterized in that, The method includes: Extract image visual features from the pre-training dataset, generate morphology text anchors based on the image visual features, and generate process cause text anchors based on historical process data. Based on the visual features of the images, the morphological text anchors, and the process cause text anchors of similar real images, a global rollback prototype is performed, and weighted fusion normalization is performed to obtain class conditional reference anchors, and the class conditional reference anchors are mapped to the language embedding space. Freeze the main parameters of the visual encoder and the large language model, and update the projection layer parameters. This step specifically includes: Initialize the parameters of the frozen visual encoder and the large language model, and freeze the main parameters of the frozen visual encoder and the large language model; Update the projection layer parameters from vision to language embedding space, and map the class conditional reference anchor point to the language embedding space through the projection layer; Calculate the category margin constraint of the pre-defined confusion defect pair, calculate the modeling loss of the large language model, and perform gradient descent operation. The expression for calculating the pre-defined confusion is: ; This indicates a confusion-type suppression defect. This indicates the index of the currently correct defect category. This indicates a presupposed set of easily confused defects. Represents cosine similarity. Indicates the first Embedded representation of class defect reference anchor points after projection layer mapping. Indicates the first Correct defect category text representation, Indicates the preset interval; The calculation expression for the modeling loss of the large language model is as follows: ; In the formula, This represents the language modeling loss in the first stage. This refers to language modeling. Indicates the total length of the target sequence. Indicates the time index in the target sequence. This represents the conditional probability distribution given by a large language model. Represents parameters of a large language model. Indicates time The target token to be output Indicates at time The previously generated token sequence; The real casting defect image is input into the frozen visual encoder to extract real visual features. The real visual features are then mapped to the language embedding space through the projection layer. The low-rank adaptation parameters in the large language model are unfrozen, and the target loss is generated to obtain the optimal detection model. This step specifically includes: The projection layer and the low-rank adaptation parameters are jointly updated; Calculate the loss of the generated target under the constraint of a closed defect vocabulary, and perform gradient descent operation; Extract the visual features to be detected from the image of the casting to be detected, and perform similarity retrieval between the visual features to be detected and knowledge entries based on the optimal detection model to obtain retrieval results. Calculate the consistency between the retrieval results, the defect category, and the category to which the selected evidence entry belongs. If the retrieval results, the defect category, and the category to which the selected evidence entry belongs are consistent, output the structured results of the defect category, evidence description, and confidence level.

2. The two-stage intelligent analysis method for small samples of casting defects according to claim 1, characterized in that, The steps of extracting image visual features from the pre-training dataset, generating morphology text anchors based on the image visual features, and generating process cause text anchors based on historical process data include: A predetermined number of sample images are extracted from the pre-trained dataset, and the visual features of the sample images are extracted using a visual encoder. Based on the visual features of the image, shape text anchors are generated, and based on process experience data and historical data, process cause text anchors are generated.

3. The two-stage intelligent analysis method for small samples of casting defects according to claim 1, characterized in that, The normalization expression is: ; In the formula, Indicates the first Class condition reference anchor point for class defects This represents the normalization operator. , , These represent three different non-negative weight coefficients. Indicates the first Visual prototypes of class defects Indicates the first Embedding of text anchor points for class defects Indicates the first Text anchor embedding of process causes of defects Indicates the first A regression reference prototype for class defects.

4. The two-stage intelligent analysis method for small samples of casting defects according to claim 1, characterized in that, After the step of calculating the consistency of the search results, defect categories, and the category to which the selected evidence item belongs, the method further includes: If the search results, the defect category, and the category of the selected evidence item are inconsistent, arbitration output will be triggered based on the difference in similarity, a preset threshold, or a conflict rule. Output the priority review results, along with a structured result including defect conclusions, evidence descriptions, confidence levels, and review flags.

5. A two-stage intelligent analysis system for small samples of casting defects, characterized in that, The system includes: The extraction and generation module is used to extract image visual features from the pre-training dataset, generate morphology text anchors based on the image visual features, and generate process cause text anchors based on historical process data. The fusion mapping module is used to perform global backtracking prototype based on the image visual features, the morphological text anchor points and the process cause text anchor points of the same real image, and to perform weighted fusion normalization to obtain class conditional reference anchor points, and to map the class conditional reference anchor points to the language embedding space. The freeze-update module is used to freeze the main parameters of the visual encoder and the large language model, and update the projection layer parameters. The freeze update module includes: An initialization unit is used to initialize the parameters of the frozen visual encoder and the large language model, and to freeze the main parameters of the frozen visual encoder and the large language model. The update mapping unit is used to update the projection layer parameters from vision to language embedding space, and to map the class conditional reference anchor point to the language embedding space through the projection layer; The computational unit is used to calculate the category margin constraint of the preset confusion defect pair, calculate the modeling loss of the large language model, and perform gradient descent operation, wherein the calculation expression for the preset confusion is: ; This indicates a confusion-type suppression defect. Indicates the current correct category index. This represents the set of defects that are easily confused with type c. This indicates a presupposed set of easily confused defects. Represents cosine similarity. Indicates the first Embedded representation of class defect reference anchor points after projection layer mapping. Indicates the first Class standard category text representation, Indicates the first Class standard category text representation, Indicates the preset interval; The calculation expression for the modeling loss of the large language model is as follows: ; In the formula, This represents the language modeling loss in the first stage. This refers to language modeling. Indicates the total length of the target sequence. Indicates the time index in the target sequence. This represents the conditional probability distribution given by a large language model. Represents parameters of a large language model. Indicates the first One target token, Indicates at time The previously generated token sequence; The input extraction module is used to input real casting defect images into the frozen visual encoder to extract real visual features, map the real visual features to the language embedding space through the projection layer, unfreeze the low-rank adaptation parameters in the large language model, and generate target loss to obtain the optimal detection model. The input extraction module includes: A joint update unit is used to jointly update the projection layer and the low-rank adaptation parameters; The computational execution unit is used to calculate the generation target loss under the constraint of a closed defect vocabulary and to perform gradient descent operations. The retrieval and calculation module is used to extract the visual features to be detected from the image of the casting to be detected, perform similarity retrieval between the visual features to be detected and knowledge entries based on the optimal detection model to obtain retrieval results, and calculate the consistency of the retrieval results, defect categories and the categories to which the selected evidence entries belong. If the retrieval results, defect categories and the categories to which the selected evidence entries belong are consistent, the structured results of defect category, evidence description and confidence level are output.

6. The two-stage intelligent analysis system for small samples of casting defects according to claim 5, characterized in that, The extraction and generation module includes: An extraction unit is used to extract a preset number of sample images from the pre-trained dataset and extract the image visual features from the sample images through a visual encoder. The generation unit is used to generate morphological text anchors based on the visual features of the image, and to generate process cause text anchors based on process experience data and historical data.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the two-stage intelligent analysis method for small samples of casting defects as described in any one of claims 1 to 4.

8. A storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the two-stage intelligent analysis method for small samples of casting defects as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Apparent defect detection method and system based on multi-modal large model

    CN119762485A

  • Few-sample industrial processing anomaly detection method based on pre-training model CLIP

    CN121883878A