Glass defect detection method and device, computer equipment and storage medium

The defect detection model trained using multimodal methods solves the problem of low detection efficiency when glass models change, and achieves efficient defect detection across different glass models, thus improving detection efficiency.

CN121998964APending Publication Date: 2026-05-08SHENZHEN SMARTMORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN SMARTMORE TECH CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing glass defect detection methods are inefficient when the glass type changes, requiring data to be collected and models to be trained again, resulting in low detection efficiency.

Method used

A multimodal approach is used to train the defect detection model. By utilizing text prompts, material model information, and defect annotation information, the model generates a structural feature map by acquiring the training image data of the first glass to be inspected and binding it with the defect annotation image data, thereby training a target defect detection model that can be reused across different models.

Benefits of technology

This technology enables defect detection of different glass models without the need to re-collect and label defect data or train a new model, improving detection efficiency and avoiding the time-consuming process of model retraining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998964A_ABST
    Figure CN121998964A_ABST
Patent Text Reader

Abstract

The invention relates to a glass defect detection method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring training image data corresponding to first to-be-detected glass; training a preset defect detection model by using the training image data to obtain a target defect detection model; the target defect detection model is obtained by training in a multi-modal mode based on the character prompt information, the material model information and the defect labeling information; acquiring to-be-detected image data corresponding to second to-be-detected glass; the second to-be-detected glass is different from the first to-be-detected glass in model; and performing defect detection on the to-be-detected image data by using the target defect detection model to obtain a defect detection result corresponding to the second to-be-detected glass. According to the invention, the detection efficiency of glass defects can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of defect detection technology, and in particular to a glass defect detection method, apparatus, computer equipment, and storage medium. Background Technology

[0002] As a core component of display devices, the surface quality of glass directly determines the display effect. During the production process, glass surfaces are prone to defects such as scratches and stains. These defects not only reduce display quality but can also render the entire display device unusable in severe cases. Therefore, accurate detection of glass defects is crucial.

[0003] Currently, glass defect detection methods typically involve: labeling defects for a single glass model, training a defect detection model using the labeled data, and then using the trained model to detect defects in that glass model. While this method is simple to operate and can handle large amounts of data, it becomes problematic when the glass model changes. The structural differences between the old and new models mean the existing model cannot recognize the structural features of the new model, easily misjudging structural differences as defects and leading to over-detection. To solve this problem, it is necessary to re-collect defect data for the new glass model, manually label it, and train a new model, resulting in low detection efficiency.

[0004] Therefore, improving the efficiency of glass defect detection has become an urgent problem to be solved. Summary of the Invention

[0005] Therefore, it is necessary to provide a glass defect detection method, apparatus, computer equipment, and storage medium to address the aforementioned technical problems and improve the detection efficiency of glass defects.

[0006] In a first aspect, this application provides a method for detecting glass defects, including: Obtain the training image data corresponding to the first glass to be detected; The target defect detection model is obtained by training the pre-set defect detection model using training image data. The target defect detection model is trained in a multimodal manner based on text prompts, material model information, and defect annotation information. Obtain the image data to be tested corresponding to the second glass to be tested; the second glass to be tested is of a different model than the first glass to be tested; The target defect detection model is used to detect defects in the image data to be inspected, and the defect detection results corresponding to the second glass to be inspected are obtained.

[0007] Secondly, this application provides a glass defect detection device, comprising: The first acquisition module is used to acquire training image data corresponding to the first glass to be detected; The model training module is used to train a preset defect detection model using training image data to obtain a target defect detection model. The target defect detection model is trained using a multimodal approach based on text prompts, material model information, and defect annotation information. The second acquisition module is used to acquire the image data to be detected corresponding to the second glass to be detected; the second glass to be detected is of a different model than the first glass to be detected. The defect detection module is used to perform defect detection on the image data to be inspected using the target defect detection model, and obtain the defect detection result corresponding to the second glass to be inspected.

[0008] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the method described above.

[0009] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.

[0010] Fifthly, this application provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described above.

[0011] The aforementioned glass defect detection method, apparatus, computer equipment, and storage medium acquire training image data of the first glass to be tested and train the defect detection model to obtain a target defect detection model that can be reused across different models. This model can directly detect defects in the image data of the second glass to be tested of different models without having to re-collect and label defect data and train a new model when changing models. This avoids the time-consuming model retraining in traditional detection methods, thereby improving the efficiency of glass defect detection. Attached Figure Description

[0012] Figure 1 An application environment diagram of a glass defect detection method provided in this application embodiment; Figure 2 A schematic flowchart of a glass defect detection method provided in an embodiment of this application; Figure 3 A flowchart illustrating a method for obtaining a target defect detection model provided in an embodiment of this application; Figure 4 A schematic diagram illustrating the training process of a target defect detection model provided in an embodiment of this application; Figure 5 A structural block diagram of a glass defect detection device provided in an embodiment of this application; Figure 6 An internal structural diagram of a computer device provided in an embodiment of this application; Figure 7 An internal structural diagram of another computer device provided in an embodiment of this application; Figure 8 This is an internal structural diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0014] First, let me explain some technical terms or phrases used in this application: Glass defect detection: refers to the process of identifying, locating and judging defects on or inside the glass surface through image recognition, feature analysis and other technical means. The aforementioned defects include, but are not limited to, scratches, particles, stains, bubbles and other abnormal areas that affect the performance or appearance quality of the glass.

[0015] Glass type: refers to the product type classified based on structural parameters such as glass size, thickness, number and distribution of holes, and curved edge design. Different glass types have different structural features and imaging characteristics.

[0016] Contrastive Language-Image Pre-training (CLIP model): This refers to a cross-modal model that is pre-trained on a large amount of language text and image data through contrastive learning. It has the ability to map textual and image information to the same feature space and can be used to extract semantic features of textual information and establish their association with image features.

[0017] Vision Transformer (VIT) model: refers to an image feature extraction model based on the Transformer architecture. It can capture global and local detail features of an image through a self-attention mechanism and can be used to extract structural region marking information in glass structure feature maps.

[0018] Transformer architecture: A deep learning model architecture based on the self-attention mechanism. Its core is to capture and model the global features of the data by calculating the correlation weights between elements in the input data. It can process sequence or image data without relying on loop or convolution operations.

[0019] A standard OK image for glass refers to an image of qualified glass selected for a specific model, free of any quality defects and with an imaging effect that meets production standards. This image must fully represent the inherent structural features of the corresponding glass model (e.g., the location and distribution of holes, the shape and extent of curved edges, etc.). In some embodiments, the imaging parameters of the standard OK image (e.g., light intensity, shooting angle, magnification, etc.) can be consistent with the imaging parameters used in the defect detection process.

[0020] Cross-self-attention mechanism: This is a feature association learning method that combines self-attention and cross-attention, widely used in feature fusion scenarios for multimodal or multi-sequence data. The model first learns the association weights of features within the same modality (or sequence) through self-attention (i.e., "self-attention"); then, it learns the association weights of features between different modalities (or sequences) through cross-attention (i.e., "cross-attention"). Through cross-self-attention, the model can simultaneously capture both the internal feature dependencies of a single modality and the cross-modal feature dependencies between multiple modalities, thus more comprehensively extracting effective information from the data.

[0021] Attention-weighted fusion is a multi-source feature fusion strategy based on "importance weights." Its core principle is to assign dynamically changing weight coefficients to features from different sources, and then achieve feature fusion through weighted summation. The core logic is as follows: first, an attention mechanism is used to calculate the weight of each feature (the larger the weight value, the higher the contribution of the feature to the task); then, the multi-source features are linearly or non-linearly combined according to the weight ratio. Compared to traditional equal-weight fusion (e.g., direct concatenation, average fusion), this method can adaptively highlight key features, suppress redundant or noisy features, and improve the effectiveness of the fused features.

[0022] Binary cross-entropy loss function: This is a loss function used for binary classification tasks. It measures the difference between the model's predicted probability and the true label, and is one of the most commonly used loss functions in classification tasks.

[0023] The Dice loss function is a loss function based on the Dice similarity coefficient, mainly used for segmentation tasks, especially suitable for scenarios with extremely imbalanced samples. The Dice similarity coefficient ranges from 0 to 1, with values ​​closer to 1 indicating a higher degree of overlap between the predicted result and the true label. The Dice loss function transforms the optimization objective of the segmentation task into minimizing the loss value by subtracting the Dice similarity coefficient from 1, i.e., maximizing the overlap of the segmentation results.

[0024] Please see Figure 1 , Figure 1This diagram illustrates the application environment of a glass defect detection method provided in this embodiment. Terminal 102 communicates with server 104 via a communication network. A data storage system stores the data that server 104 needs to process. The data storage system can be integrated onto server 104 or hosted on a cloud or other network server. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0025] like Figure 2 As shown, this application provides a method for detecting glass defects, which can be applied to... Figure 1 The method will be illustrated using terminal 102 or server 104 as examples. It is understood that the computer device may include at least one of a terminal and a server. The method includes the following steps: S101. Obtain the training image data corresponding to the first glass to be detected.

[0026] Specifically, for the first glass to be inspected, images of various defects that occur during its production process can be collected to obtain multiple defect images; these multiple defect images are manually annotated (or automatically annotated by machine) to clarify the location, shape, type and other information of the defects in each image, and obtain the corresponding defect annotation image data, so that the model can learn the morphological features of the defects.

[0027] The types of defects in the glass may include at least one of the following: scratches, particles, stains, bubbles, etc., without limitation.

[0028] In addition, a standard OK image (i.e., an image of qualified glass without defects) corresponding to the first glass to be tested can be obtained. The inherent structural regions of the glass, such as the hole structure and the imaging difference area of ​​the arc edge, are marked on the standard OK image. Based on the above markings, a corresponding structural feature map (i.e., the first glass model feature map data) is generated. For example, the marked structural regions are assigned a value of 1, and the remaining normal regions are assigned a value of 0. The structural feature map is then bound to all defect-marked images in the defect-marked image data. That is, each defect-marked image corresponds to the same structural feature map, without the need for adjustment based on different defect images, thereby obtaining training image data.

[0029] As can be seen, by annotating the structural regions of only one standard OK image, generating a unified structural feature map, and binding it with all defect-annotated image data, this operation simplifies the workload of structural annotation from "annotating each defect image" to "annotating once and reusing globally," completely eliminating a large amount of repetitive structural annotation work, significantly reducing the annotation cost and time cost of training data, and improving the efficiency of training data acquisition.

[0030] Furthermore, by generating structural feature maps based on a standard OK map, the model learns the general logic during training that "regions assigned a value of 1 in the structural feature map represent the inherent structure of the glass, and the inherent structure of the glass does not require defect detection," rather than the specific structural location and distribution of the first glass to be inspected. When switching to a new type of glass, only the structural feature map corresponding to the new type needs to be replaced, and the model can quickly adapt to the detection requirements of the new type based on the learned general logic. This training method lays a solid foundation for the model's cross-type reuse capability and is a core prerequisite for achieving the goal of not needing to retrain the model when changing types.

[0031] S102. The preset defect detection model is trained using training image data to obtain the target defect detection model. The target defect detection model is trained using a multimodal approach based on text prompts, material model information, and defect annotation information.

[0032] Among them, the defect detection model can be preset or defaulted in advance. It refers to the initial model in which the core framework such as network structure, loss function, and optimizer are predefined manually before formal training, but the parameters are not updated iteratively with training image data.

[0033] In some embodiments, the defect detection model can be a deep learning model.

[0034] In some embodiments, the training image data includes feature map data of a first glass model and defect annotation image data, where the first glass model is the model of a first glass to be inspected; see also Figure 3 , Figure 3 This is a flowchart illustrating a method for obtaining a target defect detection model according to an embodiment of this application. The method involves training a preset defect detection model using training image data to obtain the target defect detection model, including: A1. Determine the prompt word data corresponding to the feature map data of the first glass model; A2. The preset defect detection model is trained using the first glass model feature map data, defect annotation image data, and prompt word data to obtain the target defect detection model.

[0035] Among them, the prompt word data refers to the standardized text prompt information generated based on the feature map data of the first glass model, which is a digital semantic expression of the inherent structural information of the glass and the defect detection rules.

[0036] Specifically, a preset image analysis algorithm can be used to statistically analyze the types and corresponding quantities of regions with a value of 1 in the feature map data of the first glass model: Identify and count all regions marked as "hole structures" to obtain the specific number of holes, denoted as X; Identify and count all regions marked as "arc-edge imaging difference regions" to obtain the specific number of imaging abnormal regions, denoted as Y.

[0037] According to the preset prompt template: "This glass contains xx holes and xx imaging abnormal areas. Defect detection will not be performed on these areas"; the statistically obtained X and Y values ​​are filled into the prompt template to generate the corresponding text prompt. This text prompt is the prompt data corresponding to the feature map data of the first glass model, and one feature map data corresponds to only one prompt data.

[0038] It should be explained that during the statistical process, the annotation rules used when generating the feature map data of the first glass model must be used to distinguish the difference between the hole structure and the arc edge imaging area, so as to ensure that the X and Y counts accurately correspond to the structural areas in the feature map.

[0039] The preset image analysis algorithm can be one of the following: connected component analysis algorithm, contour detection algorithm, etc.

[0040] Next, the preset defect detection model can be trained using the first glass model feature map data, defect annotation image data, and prompt word data to obtain the target defect detection model.

[0041] It is evident that by combining the feature map data of the first glass model, the defect annotation image data, and the prompt word data to train the defect detection model, the model can learn the morphological features of defects while also accurately grasping the location and quantity of the inherent structural regions of the glass (holes, arc edge imaging difference regions, etc.) and the rule of "not performing defect detection" through the visual structural information of the feature map and the semantic constraint information of the prompt words. Thus, the trained target defect detection model has the ability to distinguish between the inherent structure of the glass and defects, avoiding over-detection caused by structural changes when changing models. It can directly adapt to the detection of new glass models without retraining the model, greatly improving detection efficiency.

[0042] In some embodiments, a preset defect detection model is trained using first glass model feature map data, defect annotation image data, and prompt word data to obtain a target defect detection model, including: B1. The prompt word data is processed using the first feature extraction model to obtain the first feature data; B2. The first glass model feature map data is processed using the second feature extraction model to obtain the second feature data; B3. The first feature data and the second feature data are fused to obtain the first fused feature data; B4. The first fused feature data and the defect annotation image data are fused to obtain the second fused feature data; B5. The preset defect detection model is trained using the first fusion feature data and the second fusion feature data to obtain the target defect detection model.

[0043] The first feature extraction model and the second feature extraction model can both be preset or defaulted in advance.

[0044] Specifically, the prompt word data can be input into the first feature extraction model for feature extraction to obtain the first feature data; similarly, the first glass model feature map data can be input into the second feature extraction model for feature extraction to obtain the second feature data.

[0045] In some embodiments, the first feature extraction model includes a contrastive language-image pre-trained model (CLIP model); cue word data is input into the CLIP model, which encodes the semantics of the cue word data, transforming abstract semantic information into a digitized language feature vector (i.e., first feature data), which contains semantic information about the number of glass structures and detection rules. Additionally, the second feature extraction model includes a visual transformation feature extraction model (VIT model); first glass model feature map data is input into the VIT model, which uses a self-attention mechanism to capture the global distribution features of structural regions in the feature map, transforming visual structural information into a digitized visual feature vector (i.e., second feature data), which contains the location and distribution information of the inherent glass structure.

[0046] Then, the first feature data and the second feature data can be fused to obtain the first fused feature data; then, the first fused feature data and the defect-annotated image data can be fused to obtain the second fused feature data. Specifically, a preset fusion method can be used to fuse the first fused feature data and the defect-annotated image data together to obtain the second fused feature data; wherein, the preset fusion method can include one of the following: feature splicing method, feature element-wise multiplication method, attention weighted fusion method, etc.

[0047] To illustrate, assuming the preset fusion method is attention-weighted fusion, the dimensions of the first fusion feature data and the defect-annotated image data can be unified. If the two dimensions are different, they can be mapped to the same dimension through a linear transformation layer. Then, the weights of the two types of features can be calculated using the attention mechanism to obtain two attention weights. Based on these two attention weights, the first fusion feature data and the defect-annotated image data are weighted and fused to obtain the second fusion feature data.

[0048] Finally, the first and second fused feature data can be used to train the preset defect detection model to obtain the target defect detection model.

[0049] Please see Figure 4 , Figure 4 The diagram illustrates the training process of a target defect detection model provided in this application embodiment. The specific process is as follows: 1. Obtain input Text prompt (i.e., the prompt data mentioned above): Prompts containing information about the inherent structure of the glass model (e.g., "This glass has 3 holes") are used to provide semantic constraints.

[0050] Material model feature map (i.e. the first glass model feature map data mentioned above): the binary structure feature map corresponding to the glass model (the inherent structure area is assigned a value of 1, and other areas are 0), which is used to provide visual structure information.

[0051] Image + Defect Annotation (i.e., the defect-annotated image data mentioned above): Glass defect images with manual annotations (marking the location / type of the actual defects), used as supervised data for training.

[0052] 2. Feature Extraction CLIP model: Receives "text prompt", extracts its semantic features, and outputs the first feature data (language modality features).

[0053] VIT model: Receives "material model feature map", extracts its visual structural features, and outputs second feature data (visual modal features).

[0054] 3. Cross-modal feature fusion Cross-self-attention mechanism: acquire "first feature data" (linguistic features) and "second feature data" (visual features), allow the two to pay attention to each other and establish semantic-visual association, and output first fused feature data (cross-modal features that fuse structural semantics and visual distribution).

[0055] The second fusion feature data: feature data generated by combining the supervision information of "graph + defect annotation" and containing defect discrimination rules (for structural constraints for subsequent model training).

[0056] 4. Model Training Feature injection: Input both the "first fusion feature data" and the "second fusion feature data" into the defect detection model and perform forward inference to obtain the predicted defect results.

[0057] Loss assessment (LOSS≤T1): Calculate the model loss based on the actual defect detection results and determine whether the model loss is lower than the preset threshold T1. If “yes” (or the model has reached the preset training rounds): the training is considered complete, and the target defect detection model is obtained; If “No”: Update the model parameters through backpropagation, and repeat the process of “inference → calculate model loss → update parameters” until the model loss reaches the target.

[0058] In some embodiments, step 2, “feature extraction”, can be performed by a multimodal feature extraction and fusion module; the multimodal feature extraction and fusion module may include a CLIP model and a VIT model; step 3, “cross-modal feature fusion”, can be performed by a cross-self-attention module.

[0059] It is evident that by using phased feature fusion and dual-feature joint training, the model learns two independent and complementary features: "cross-modal association between structural semantics and visual features" and "defect discrimination rules under structural constraints." This enhances the model's learning depth and robustness, avoids the loss of feature information or weight imbalance caused by single-feature training, and further improves the model's accuracy in distinguishing between the inherent structure and defects of glass.

[0060] In some embodiments, the first feature data and the second feature data are fused to obtain the first fused feature data, including: The first feature data and the second feature data are fused based on the cross-self-attention mechanism to obtain the first fused feature data.

[0061] Specifically, the first feature data (denoted as L) is obtained by extracting the prompt word data from the first feature extraction model (i.e., the CLIP model), with dimensions [N, dk]; where N is the sequence length of the language features (e.g., the word embedding sequence length of the prompt word), and dk is the feature dimension.

[0062] The second feature data (denoted as V1) is obtained by extracting the feature map data of the first glass model from the second feature extraction model (i.e., the VIT model). The dimension is [M, dk], where M is the sequence length of the visual features (such as the patch sequence length of the feature map), and dk is the feature dimension (which must be consistent with the dimension of the first feature data. If they are inconsistent, they can be unified through a linear transformation layer).

[0063] The cross-self-attention mechanism allows linguistic features (i.e., the first feature data mentioned above) and visual features (i.e., the second feature data mentioned above) to pay attention to each other. That is, linguistic features pay attention to the regions in visual features that match their semantics, while visual features pay attention to the semantics in linguistic features that match their visual distribution, thereby achieving deep fusion of cross-modal features.

[0064] To illustrate, we can first generate a query (Q), key (K), and value (V) matrix. Specifically, the first feature data can be used as the query matrix Q, and the second feature data as the key matrix K and value matrix V, realizing the cross-focus of language features on visual features (bidirectional cross-focus is also possible, but we take unidirectional cross-focus as an example here, which is consistent with the application scenario of the method in this embodiment). A linear transformation is then performed on the first feature data L to generate the query matrix Q, as follows: Q=L*W Q ; Among them, W Q This represents the learnable query weight matrix with dimensions [dk, dQ]; dQ represents the feature dimension of the query matrix Q. A linear transformation is performed on the second feature data V1 to generate the key matrix K and the value matrix V, as follows: K=V1*W K ; V=V1*W V ; Among them, W K W represents the learnable key weight matrix with dimensions [dk, dK], where dK represents the feature dimension of the key matrix K; V Let V represent the learnable value weight matrix with dimensions [dk, dV], where dV represents the feature dimension of the value matrix V.

[0065] By calculating the dot product of the query matrix Q and the key matrix K, the association score between the linguistic and visual features is obtained. Then, through scaling and the Softmax activation function, the association score is transformed into an attention weight matrix between 0 and 1, representing the degree of attention of the linguistic features to each visual feature region. The attention weight matrix and the value matrix V are weighted and summed to obtain the cross-attention feature (i.e., the first fused feature data), which integrates the semantic constraints of the linguistic features and the structural distribution information of the visual features.

[0066] In some embodiments, to improve the training stability and feature learning ability of the model, residual connections and layer normalization operations are typically added after the cross-attention mechanism: Residual connection: The first feature data L is concatenated or added with the cross-attention feature (the dimensions must be consistent) to obtain the residual feature.

[0067] Layer normalization: The residual features are layer normalized to obtain the final fused feature data (i.e., the first fused feature data).

[0068] In this way, the cross-self-attention mechanism enables the first feature data (semantic cue words) of the language modality and the second feature data (visual information of feature maps) of the visual modality to pay attention to each other and establish a precise association. This provides unified and closely related cross-modal feature data for subsequent model training, enabling the model to understand the correspondence between the semantic constraints of the inherent structure of glass and the visual distribution, thereby effectively improving the model training effect.

[0069] In some embodiments, a preset defect detection model is trained using first fused feature data and second fused feature data to obtain a target defect detection model, including: C1. The first fusion feature data and the second fusion feature data are processed by the defect detection model to obtain the predicted defect detection results; C2. Obtain the actual defect detection results corresponding to the first glass to be tested; C3. Determine the model loss based on the preset loss function, the predicted defect detection results, and the actual defect detection results; C4. Update the parameters of the defect detection model based on the model loss to obtain the target defect detection model.

[0070] The preset loss function can be preset in advance or defaulted. Specifically, the preset loss function can include the following: binary cross-entropy loss function, Dice loss function, etc.

[0071] Specifically, the first fusion feature data is a cross-modal feature fused through a cross-self-attention mechanism, containing precise correlation information between the semantic constraints of the glass's inherent structure and its visual distribution (e.g., the specific hole regions in the semantic correspondence feature map for "X holes"). The second fusion feature data is a feature fused through attention weighting, containing morphological features of defects under structural constraints (e.g., scratches, bubbles, and other defect features in unstructured areas). Together, they provide the defect detection model with a complete basis for judging "which areas are inherent structures (no need for detection)" and "which areas may contain defects (need to be identified)." The first and second fusion feature data can be input into the defect detection model for processing to obtain predicted defect detection results. For example, the defect detection model will perform a comprehensive analysis of the input features (i.e., the first and second fusion feature data). In the inherent structural regions marked by the first fusion feature data, the defect detection model suppresses the output of defect detection to avoid misclassifying structures as defects. In non-inherent structural regions, the defect detection model determines the existence of defects based on the defect morphological features in the second fusion feature data, and further determines the location, type, size, and other information of the defects. After inference, the defect detection model outputs the final predicted defect detection result, which may include the following information: mask or classification label, the specific location of the defect, the defect type (e.g., scratches, particles, bubbles, etc.), and the confidence score of the defect, etc.

[0072] To illustrate, suppose the core task of a defect detection model is "pixel-level defect localization," then its output, the predicted defect detection result, typically includes: Mask: A binary image with the exact same size as the input glass image, where white pixels (value 1) represent defect areas and black pixels (value 0) represent normal areas. For example, when detecting bubble defects in a certain type of photovoltaic glass, the mask will accurately circle the shape and position of each bubble.

[0073] Defect type: The defect detection model determines the defect type as "bubble" by analyzing the features of the masked area.

[0074] Confidence score: The confidence score for the entire masked area. For example, 0.92 means that the defect detection model has a 92% confidence that the area is a bubble defect.

[0075] For example, assuming the core task of a defect detection model is to "locate and classify defects," its output, the predicted defect detection results, typically include: Category labels: For example, labels such as "scratches", "particles", and "bubbles" directly identify the type of defect.

[0076] The specific location of the defect is represented by the coordinates of the bounding box, in the format (x1, y1, x2, y2); where (x1, y1) are the coordinates of the top-left corner of the bounding box, and (x2, y2) are the coordinates of the bottom-right corner. For example, when detecting a scratch defect on the windshield of a certain model of car, the output location is (120, 250, 360, 265), indicating that the scratch is located within this area of ​​the image.

[0077] Defect type: "Scratches".

[0078] Confidence score: The confidence score for this bounding box. For example, 0.88 means that the defect detection model is 88% confident that there is a scratch defect in the area.

[0079] Then, the actual defect detection result corresponding to the first glass to be tested can be obtained. Specifically, a preset mapping relationship between the glass to be tested and the defect detection result can be stored in advance, and the actual defect detection result corresponding to the first glass to be tested can be determined based on the mapping relationship. Alternatively, the actual defect detection result can be obtained by having staff manually inspect the first glass to be tested.

[0080] Next, the model loss can be determined based on the preset loss function, the predicted defect detection results, and the actual defect detection results. Specifically, the corresponding parameters in the predicted defect detection results and the actual defect detection results can be substituted into the preset loss function for calculation to obtain the model loss. For example, assuming the preset loss function is the Dice loss function, this loss function measures the detection accuracy by calculating the overlap between the predicted defect region in the predicted defect detection results and the actual defect region in the actual defect detection results. The lower the overlap, the higher the loss value.

[0081] Finally, the parameters of the defect detection model can be updated based on the model loss to obtain the target defect detection model. Specifically, based on the calculated model loss, the backpropagation algorithm is executed. Starting from the output layer of the defect detection model, the gradient of the loss function with respect to each learnable parameter (e.g., the attention weight matrix) in the defect detection model is calculated layer by layer. Then, a gradient descent optimizer (e.g., the Adam optimizer) is used to update all learnable parameters in the defect detection model according to the calculated gradient direction. The process of "forward inference to obtain prediction results → calculating model loss → backpropagation to calculate gradient → parameter update" is repeated until the model loss of the defect detection model drops to a stable threshold (e.g., 0.01) or reaches the preset number of training rounds. At this point, training is stopped, and the optimized model obtained is the target defect detection model.

[0082] Thus, by using the first fusion feature data, which integrates structural semantics and visual information, and the second fusion feature data, which incorporates structural constraints, as input to the model, and combining supervised training and iterative parameter updates based on real defect detection results, the model can accurately learn the rules for distinguishing between the inherent structural regions and defect regions of the glass. This effectively avoids misjudging the inherent structure as a defect, while improving the detection accuracy of real defects. Ultimately, a target defect detection model with strong generalization ability and adaptability to the detection of different types of glass is obtained.

[0083] S103. Obtain the image data to be detected corresponding to the second glass to be detected; the second glass to be detected is of a different model than the first glass to be detected; Specifically, another piece of glass with a different model than the first glass to be tested (i.e., the second glass model) can be selected and used as the second glass to be tested. Image data of the second glass to be tested can be collected to obtain the image data to be tested.

[0084] It should be explained that the second glass to be tested can also be of the same model as the first glass to be tested. Since the target defect detection model is trained from the training image data corresponding to the first glass to be tested, the target defect detection model can identify defects in the image data of the second glass to be tested.

[0085] In some embodiments, obtaining the image data to be detected corresponding to the second glass to be detected includes: S31. Determine the reference image corresponding to the second glass to be tested; S32. Determine the reference feature map corresponding to the reference image; S33. Determine the image data to be detected based on the reference feature map.

[0086] Specifically, a standard OK image of the second glass to be tested can be obtained and used as a reference image for the second glass to be tested. Then, a reference feature map corresponding to the reference image can be determined. Specifically, the inherent structural regions can be marked on the reference image using a preset annotation tool. The inherent structural regions can include "hole structures" and "imaging difference regions" (e.g., arc edge imaging difference regions, etc., which are inherent structural features of this type of glass, not defects). After the annotation is completed, the corresponding feature map (i.e., binary structural feature map, where the marked holes and imaging difference regions are assigned a value of 1, and other normal regions are assigned a value of 0) is automatically generated based on the marked structural regions.

[0087] The preset annotation tools may include at least one of the following: LabelMe, customized industrial annotation plugins, etc.

[0088] Finally, the image data to be detected can be determined based on the reference feature map. Specifically, the reference feature map can be directly used as the image data to be detected.

[0089] In this way, by configuring reference images and reference feature maps separately for different models of the second glass to be tested, the inherent structural information of the glass model can be accurately extracted, and the input image data to be tested can be preprocessed or constrained in a targeted manner based on the structural information, so as to avoid the inherent structure caused by model differences being misjudged as a defect.

[0090] In addition, the target defect detection model can be quickly adapted to new glass models by updating the glass feature map information without modifying the parameters or retraining it. This greatly improves the detection efficiency and cross-model detection versatility when changing production lines.

[0091] S104. Use the target defect detection model to perform defect detection on the image data to be detected, and obtain the defect detection result corresponding to the second glass to be detected.

[0092] The image data to be detected includes a reference feature map.

[0093] Specifically, reference prompting data can be determined based on the reference feature map and the preset prompting word template. Specifically, the reference feature map can be analyzed using a preset image analysis algorithm to obtain the specific number of holes, denoted as X1, and the specific number of imaging abnormal areas, denoted as Y1. The specific values ​​of X1 and Y1 obtained from the analysis are filled into the corresponding positions of the prompting word template to obtain the reference prompting data. Then, the reference feature map and reference prompting data can be input into the target defect detection model for processing to obtain the defect detection result corresponding to the second glass to be detected.

[0094] In summary, the glass defect detection method provided in this application obtains a target defect detection model that can be reused across different models by acquiring training image data of the first glass to be tested and training the defect detection model. This model can directly detect defects in the image data of the second glass to be tested of different models without having to re-collect and label defect data and train a new model when changing models. This avoids the time-consuming model retraining in traditional detection methods, thereby improving the detection efficiency of glass defects.

[0095] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0096] Based on the same inventive concept, this application also provides a glass defect detection device. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more glass defect detection device embodiments provided below can be found in the limitations of the glass defect detection method above, and will not be repeated here.

[0097] Please see Figure 5 , Figure 5 This application provides a structural block diagram of a glass defect detection device 500, which includes: The first acquisition module 501 is used to acquire training image data corresponding to the first glass to be detected; The model training module 502 is used to train a preset defect detection model using training image data to obtain a target defect detection model. The target defect detection model is trained using a multimodal approach based on text prompts, material model information, and defect annotation information. The second acquisition module 503 is used to acquire the image data to be detected corresponding to the second glass to be detected; the second glass to be detected is of a different model than the first glass to be detected. The defect detection module 504 is used to perform defect detection on the image data to be inspected using the target defect detection model, and obtain the defect detection result corresponding to the second glass to be inspected.

[0098] In some embodiments, the training image data includes first glass model feature map data and defect annotation image data, wherein the first glass model is the model of the first glass to be detected; in terms of training a preset defect detection model using the training image data to obtain a target defect detection model, the model training module 502 is specifically used for: Determine the prompt word data corresponding to the feature map data of the first glass model; The target defect detection model is obtained by training the pre-set defect detection model using the first glass model feature map data, defect annotation image data, and prompt word data.

[0099] In some embodiments, in training a preset defect detection model using first glass model feature map data, defect annotation image data, and prompt word data to obtain a target defect detection model, the model training module 502 is specifically used for: The prompt word data is processed by the first feature extraction model to obtain the first feature data; The second feature data is obtained by processing the feature map data of the first glass model using the second feature extraction model. The first feature data and the second feature data are fused to obtain the first fused feature data; The first fused feature data and the defect-annotated image data are fused to obtain the second fused feature data; The preset defect detection model is trained using the first fusion feature data and the second fusion feature data to obtain the target defect detection model.

[0100] In some embodiments, in training a preset defect detection model using first fused feature data and second fused feature data to obtain a target defect detection model, the model training module 502 is specifically used for: The first and second fused feature data are processed by the defect detection model to obtain the predicted defect detection results; Obtain the actual defect detection results corresponding to the first glass to be tested; The model loss is determined based on the preset loss function, the predicted defect detection results, and the actual defect detection results; The parameters of the defect detection model are updated based on the model loss to obtain the target defect detection model.

[0101] In some embodiments, in fusing the first feature data and the second feature data to obtain the first fused feature data, the model training module 502 is specifically used for: The first feature data and the second feature data are fused based on the cross-self-attention mechanism to obtain the first fused feature data.

[0102] In some embodiments, the second acquisition module 503 is specifically used for: acquiring the image data to be detected corresponding to the second glass to be detected; Determine the reference image corresponding to the second glass to be tested; Determine the reference feature map corresponding to the reference image; The image data to be detected is determined based on the reference feature map.

[0103] In some embodiments, the first feature extraction model includes a contrastive language-image pre-trained model; the second feature extraction model includes a visual transformation feature extraction model.

[0104] Each module in the aforementioned glass defect detection device 500 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0105] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data related to glass defect detection. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the glass defect detection method described above.

[0106] In some embodiments, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the steps in the aforementioned glass defect detection method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen; the input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs or touchpads set on the casing of the computer device, or external keyboards, touchpads or mice, etc.

[0107] Those skilled in the art will understand that Figure 6 or Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0108] In some embodiments, a computer device is provided, the computer device including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.

[0109] In some embodiments, such as Figure 8 The diagram shows the internal structure of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the above-described method embodiments.

[0110] In some embodiments, a computer program product is provided, which includes a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0111] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0112] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0113] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0114] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for detecting glass defects, characterized in that, include: Obtain the training image data corresponding to the first glass to be detected; The preset defect detection model is trained using the training image data to obtain the target defect detection model; The target defect detection model is trained using a multimodal approach based on text prompts, material model information, and defect annotation information. Obtain the image data to be tested corresponding to the second glass to be tested; the second glass to be tested is of a different model than the first glass to be tested; The target defect detection model is used to perform defect detection on the image data to be detected, and the defect detection result corresponding to the second glass to be detected is obtained.

2. The method according to claim 1, characterized in that, The training image data includes first glass model feature map data and defect annotation image data, wherein the first glass model is the model of the first glass to be tested; The step of training a preset defect detection model using the training image data to obtain a target defect detection model includes: Determine the prompt word data corresponding to the first glass model feature map data; The target defect detection model is obtained by training the preset defect detection model using the first glass model feature map data, the defect annotation image data, and the prompt word data.

3. The method according to claim 2, characterized in that, The step of training a preset defect detection model using the first glass model feature map data, the defect annotation image data, and the prompt word data to obtain a target defect detection model includes: The prompt word data is processed by the first feature extraction model to obtain the first feature data; The second feature data is obtained by processing the first glass model feature map data using the second feature extraction model. The first feature data and the second feature data are fused to obtain the first fused feature data; The first fused feature data and the defect-annotated image data are fused to obtain the second fused feature data; The preset defect detection model is trained using the first fused feature data and the second fused feature data to obtain the target defect detection model.

4. The method according to claim 3, characterized in that, The step of training a preset defect detection model using the first fused feature data and the second fused feature data to obtain a target defect detection model includes: The first fused feature data and the second fused feature data are processed by a defect detection model to obtain the predicted defect detection result; Obtain the actual defect detection results corresponding to the first glass to be tested; The model loss is determined based on the preset loss function, the predicted defect detection results, and the actual defect detection results; The defect detection model is updated with parameters based on the model loss to obtain the target defect detection model.

5. The method according to claim 3 or 4, characterized in that, The process of fusing the first feature data and the second feature data to obtain the first fused feature data includes: The first feature data and the second feature data are fused based on the cross-self-attention mechanism to obtain the first fused feature data.

6. The method according to any one of claims 1-4, characterized in that, The step of obtaining the image data to be detected corresponding to the second glass to be detected includes: Determine the reference image corresponding to the second glass to be tested; Determine the reference feature map corresponding to the reference image; The image data to be detected is determined based on the reference feature map.

7. The method according to claim 3, characterized in that, The first feature extraction model includes a contrastive language-image pre-trained model; the second feature extraction model includes a visual transformation feature extraction model.

8. A glass defect detection device, characterized in that, include: The first acquisition module is used to acquire training image data corresponding to the first glass to be detected; The model training module is used to train a preset defect detection model using the training image data to obtain a target defect detection model; the target defect detection model is trained using a multimodal approach based on text prompts, material model information, and defect annotation information. The second acquisition module is used to acquire the image data to be detected corresponding to the second glass to be detected; the second glass to be detected is of a different model than the first glass to be detected. The defect detection module is used to perform defect detection on the image data to be detected using the target defect detection model, and obtain the defect detection result corresponding to the second glass to be detected.

9. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.