A real-time image semantic segmentation system of self-supervised learning
Through a real-time image semantic segmentation system with self-supervised learning, combined with energy distribution analysis and vision-language similarity matching to generate alarm masks, adaptive spectral clustering and scarcity indicators, accurate identification and positioning of unknown defects are achieved, solving the problems of missed detection and forgetting in automated inspection systems in dynamic industrial environments, and meeting the real-time and adaptability requirements of edge devices.
Patent Information
- Application Number
- CN202511125544.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing automated inspection systems suffer from problems of missed detection and catastrophic forgetting when faced with identifying new types of defects in dynamic industrial production environments, making it difficult to achieve real-time and accurate defect identification on resource-constrained edge devices.
A real-time image semantic segmentation system using self-supervised learning generates warning masks through energy distribution analysis and vision-language similarity matching, generates pseudo labels by combining adaptive spectral clustering and scarcity indicators, uses a very small amount of manual annotation for confirmation, freezes the encoder to fine-tune the terminal, and realizes rapid self-evolution of the model and long-term performance conservation through self-attention reconstruction branches and heterogeneous parallel scheduling.
It achieves accurate identification and positioning of unknown defects, reduces the workload of manual labeling, prevents catastrophic forgetting, ensures the high real-time performance and continuous learning ability of the model on edge devices, and adapts to the needs of complex industrial scenarios.
Smart Images

Figure CN120655927B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time image semantic segmentation, and in particular to a real-time image semantic segmentation system based on self-supervisory learning. Background Art
[0002] In modern industrial manufacturing, quality inspection is crucial for ensuring product compliance and maintaining a company's reputation. Imagine a busy factory floor, where conveyor belts continuously transport products awaiting inspection, such as automotive parts, electronic components, or precision machinery. Amid the hum of machinery, high-resolution cameras positioned above key production line nodes capture real-time images of each product's surface. These products undergo rigorous defect inspection, such as identifying scratches, cracks, bubbles, or assembly errors, to prevent defective products from entering the market. In the past, this task primarily relied on experienced human inspectors, who individually examined products with the naked eye or through magnifying glasses. However, with the expansion of production scale and the increasing demand for efficiency, manual inspection has become inadequate to meet the high-throughput demands of modern industry. In recent years, the rise of computer vision technology has enabled automated inspection systems to gradually replace manual labor. These systems leverage advanced image processing and machine learning algorithms to achieve faster and more consistent defect detection. These systems must not only maintain low latency in high-speed production environments but also adapt to diverse product types and complex industrial scenarios to ensure accurate and reliable inspection results.
[0003] However, existing automated inspection systems face significant technical challenges in practical applications, especially in the field of defect recognition based on semantic segmentation. Traditional semantic segmentation models rely on pre-defined defect category datasets for training, and their design assumes that the defect types in the inspection environment are fixed.
[0004] However, this assumption often breaks down in dynamic industrial production environments. For example, when new manufacturing processes are introduced, raw materials are changed, or equipment parameters are adjusted, new types of defects may be introduced, such as specific types of surface depressions or microcracks that were not present in the model's initial training data.
[0005] The emergence of these new defects stems from the complexity and unpredictability of the production process, such as batch variations between material suppliers, fluctuations in ambient temperature and humidity, or anomalies caused by mechanical wear. Furthermore, because traditional models lack the ability to identify unknown defects, they either misclassify new defects as known categories or completely ignore these anomalies, resulting in frequent missed detections. In demanding industrial scenarios, such as automotive manufacturing or aerospace, the consequences of missed detections can be extremely serious: defective products can cause safety incidents, such as brake component failure or structural weaknesses in the fuselage, and even lead to large-scale recalls, resulting in millions of dollars in economic losses and damage to brand reputation. Furthermore, even when retraining the model by collecting new defect data, traditional methods still face the risk of catastrophic forgetting, where the model's ability to recognize old defects significantly decreases when learning new defects. This problem prevents the system from operating continuously in dynamic environments, limiting its applicability for real-time quality inspection, especially on resource-constrained edge devices. Summary of the Invention
[0006] (1) Technical problems solved
[0007] To address the shortcomings of existing technologies, the present invention provides a self-supervised learning, real-time image semantic segmentation system. It generates warning masks through energy distribution analysis and vision-language similarity matching; adaptive spectral clustering generates pseudo-labels and scarcity indicators; minimal manual annotation verification; frozen encoder fine-tuning, elastic memory coefficient regularization and distillation; self-attention reconstruction branch provides pixel-level replay; and automatic mixed precision and heterogeneous parallel scheduling enable edge deployment. This invention addresses the challenges of learning-forgetting cycles and the online-offline split, enabling rapid model self-evolution and long-term performance conservation, making it suitable for continuous learning in open environments. It also addresses the technical issues discussed in the background.
[0008] (2) Technical solution
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: a real-time image semantic segmentation system with self-supervised learning, including an alarm mask generation module that simultaneously performs energy distribution analysis and visual-language similarity matching on the input image, generates and outputs an unknown defect alarm mask at the pixel level based on a dual-threshold criterion;
[0010] The adaptive spectral clustering module maps the alarm mask to a unified embedding space and then inputs it into the adaptive spectral clustering module, automatically determining the number of clusters, generating pseudo labels, and calculating the scarcity index of each cluster;
[0011] The manual labeling module selects three representative samples from each cluster based on their scarcity for manual rapid labeling, and writes the confirmed results back to the pseudo-label table and category dictionary simultaneously;
[0012] The fine-tuning module only fine-tunes the tuning terminal and new cluster adaptation layer under the condition of freezing the encoder parameters. In each round of training, the information elasticity index and cross-layer forgetting index are output through the adaptive stability mapping module to dynamically adjust the EWC regularization and distillation weights.
[0013] Self-attention reconstruction module, the self-attention reconstruction branch is enabled in parallel during training, providing pixel-level playback signals to the segmentation branch through the residual map and updating the weights of the two branches in the same forward process.
[0014] The automatic mixed-precision inference module uses automatic mixed-precision inference on the fine-tuned model, mapping half-precision tensor calculations to the GPU, integer tensor calculations to the NPU, and full-precision control flow to the CPU, combined with heterogeneous parallel scheduling, to perform online incremental training and inference tasks and deploy them on the edge.
[0015] Further, forward inference is performed on the input image to generate class prediction logits for each pixel, and an energy score of each pixel is calculated based on the class prediction logits, where the energy score is obtained by applying a negative logarithm and function to the class prediction logits;
[0016] An energy threshold is set and pixels with energy scores exceeding the energy threshold are marked as potential unknown defect areas; a pre-trained vision-language model is used to extract pixel-level visual features of the input image.
[0017] Furthermore, a text feature vector is generated for each known defect category, and the cosine similarity between the visual feature of each pixel and the text features of all known categories is calculated and the maximum similarity is taken;
[0018] Set a similarity threshold and identify pixels whose maximum similarity is lower than the similarity threshold as unknown defects;
[0019] A dual-threshold criterion is applied to each pixel. When the energy score exceeds the energy threshold and the maximum similarity is lower than the similarity threshold, the pixel is marked as an unknown defect. An alarm mask is generated, in which pixels marked as unknown defects are assigned a value of 1 and the remaining pixels are assigned a value of 0. The alarm mask is output.
[0020] Further, performing adaptive spectral clustering on the feature set, including constructing a similarity matrix of the feature set;
[0021] The degree matrix is calculated based on the similarity matrix, and the Laplace matrix is constructed by the similarity matrix and the degree matrix. After the Laplace matrix is eigendecomposed, the eigenvalue gap method is used to determine the number of clusters. The first several eigenvectors are extracted and K-means clustering is applied to generate cluster labels.
[0022] Furthermore, three pixels with the smallest Euclidean distance to the cluster feature mean are selected from each cluster as representative samples, and the image areas corresponding to the representative samples are submitted to the manual annotation process to assign a defect category label to each representative sample.
[0023] Furthermore, a majority vote is performed on the defect category labels of the three representative samples of each cluster to determine the final category label of the cluster. The pseudo labels of all pixels in the cluster in the pseudo label map are updated to the final category label of the cluster, and the number and name of the new defect category are recorded in the category dictionary. The updated pseudo label map and category dictionary are output.
[0024] Furthermore, the encoder parameters are frozen to maintain the stability of feature extraction, the decoding head is fine-tuned with a decaying learning rate, and the new cluster adaptation layer is set to a higher initial learning rate than the decoding head to accelerate the learning of new features;
[0025] The information elasticity index is determined by calculating the deviation between the old task features and the initial features and applying a Gaussian kernel function. The cross-layer forgetting index is calculated by weighted averaging the output deviations of each layer of the decoding head on the old task samples.
[0026] Furthermore, the adaptive stability mapping module receives the information elasticity index and the cross-layer forgetting index and generates the elastic memory coefficient using the Sigmoid function;
[0027] The total loss function integrates the cross entropy loss of the new category, the EWC regularization loss weighted by (1-elastic memory coefficient) to protect key parameters, and the distillation loss weighted by the elastic memory coefficient to maintain the consistency of the output distribution of the old task.
[0028] Furthermore, the self-attention reconstruction branch is enabled in parallel, accessing the feature map output by the encoder, and generating an attention weight matrix through the self-attention module. The attention weight matrix captures the global context information of the feature map to generate a global context feature map;
[0029] The global context feature map is input into the reconstruction decoder to generate a reconstructed image;
[0030] The pixel-level difference between the reconstructed image and the original input image is calculated to generate a residual map, which is normalized and upsampled to the same resolution as the segmentation decoder feature map.
[0031] Furthermore, the upsampled residual map and the segmentation feature map are concatenated along the channel dimension to generate a fused feature map, which is used for subsequent calculations of the segmentation decoder to generate segmentation prediction results.
[0032] Define the mean square error loss of the reconstruction task and the cross entropy loss of the segmentation task, construct the total loss function as the weighted sum of the mean square error loss and the cross entropy loss, calculate the reconstructed image and segmentation prediction results simultaneously in the forward propagation, and update the weights of the reconstruction decoder and segmentation decoder through backpropagation.
[0033] Furthermore, the fine-tuned model will be automatically inferred using mixed precision, including:
[0034] Convert the weights and input tensors of the convolutional and fully connected layers from full precision to half precision, convert the input and output tensors of the activation and pooling layers from full precision to integer, and keep the input and output tensors of the Softmax layer and loss function in full precision.
[0035] Furthermore, through heterogeneous parallel scheduling, half-precision tensor calculations are allocated to the GPU, integer tensor calculations are allocated to the NPU, and full-precision control flow is allocated to the CPU;
[0036] Online incremental training and real-time inference are performed concurrently. The training task uses new pseudo-labels and fine-tuning strategies to update model parameters in the background, while the inference task uses the current model to perform real-time defect segmentation on the input image in the foreground.
[0037] After the model is deployed on the edge device, online incremental training is performed regularly, and the performance is tested using the validation set after the update. If the performance drops by more than the preset threshold, it is rolled back to the previous stable version.
[0038] (3) Beneficial effects
[0039] The present invention provides a real-time image semantic segmentation system with self-supervised learning, which has the following beneficial effects:
[0040] Through the dual-threshold criteria of energy distribution analysis and visual-language similarity matching, pixel-level alarm masks for unknown defects are accurately generated. Combined with uncertainty quantification and semantic verification, new defect areas that are difficult to detect with traditional models are accurately identified, providing a reliable abnormal area location basis for subsequent anomaly clustering and model updates.
[0041] By mapping the alarm mask to a unified embedding space, adaptive spectral clustering technology is used to automatically determine the number of clusters and generate pseudo labels, while also calculating a scarcity index. The application of adaptive spectral clustering can automatically discover the category structure of new defects, avoiding the difficulty of presetting the number of clusters in traditional clustering methods. The introduction of a scarcity index provides priority guidance for manual labeling, making the labeling process more targeted and significantly improving labeling efficiency.
[0042] Based on scarcity, only three representative samples are selected from each cluster for manual labeling, and the confirmation results are synchronized to the pseudo-label table and category dictionary. With minimal manual intervention, new defect categories can be accurately confirmed, significantly reducing the workload of manual labeling while ensuring the high reliability of the pseudo-label table.
[0043] While freezing the encoder parameters, only the tuning dock and new cluster adaptation layers are fine-tuned. The EWC regularization and distillation weights are dynamically adjusted using the elastic memory coefficient output by the adaptive stability mapping module, using the information elasticity index and the cross-layer forgetting index. This achieves a dynamic balance between new and old tasks, effectively preventing catastrophic forgetting while accelerating the learning of new tasks.
[0044] The self-attention reconstruction branch is enabled in parallel, providing pixel-level playback signals to the segmentation branch through the residual map, and updating the weights of both branches in the same forward process. By using the reconstruction task to assist the segmentation task, the model's ability to identify complex defects is significantly improved, while the accuracy of locating abnormal areas is enhanced, allowing the model to more accurately capture defect details, thereby improving overall segmentation performance.
[0045] By adopting automatic mixed-precision inference and heterogeneous parallel scheduling, the model is deployed at the edge, making full use of the limited resources of edge devices, ensuring high real-time performance and continuous learning capabilities, meeting the requirements of industrial quality inspection scenarios for low latency and high adaptability, and realizing rapid self-evolution of the model and long-term performance conservation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the structure of the real-time image semantic segmentation system for self-supervised learning of the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] See also Figure 1 The present invention provides a real-time image semantic segmentation system for self-supervised learning, comprising:
[0049] Step 1: Perform forward inference on the input image to generate class prediction logits for each pixel; calculate the energy score of each pixel based on the class prediction logits, where the energy score is obtained by applying the negative logarithm and function to the class prediction logits; set an energy threshold and mark pixels with energy scores exceeding the energy threshold as potential unknown defect areas; use a pre-trained vision-language model to extract pixel-level visual features of the input image;
[0050] Generate text feature vectors for known defect categories, calculate the cosine similarity between the visual features of each pixel and the text features of all known categories, and take the maximum similarity; set a similarity threshold and confirm that pixels with a maximum similarity lower than the similarity threshold are unknown defects; apply a dual-threshold criterion to each pixel, and mark the pixel as an unknown defect when the energy score exceeds the energy threshold and the maximum similarity is lower than the similarity threshold; generate an alarm mask, assigning a value of 1 to pixels marked as unknown defects and 0 to the remaining pixels; and output the alarm mask to the downstream process.
[0051] The step 1 includes the following:
[0052] Step 101: Energy distribution analysis
[0053] In energy distribution analysis, forward inference is first performed on the input image, and a deep learning model is used to generate a category prediction for each pixel, that is, a probability distribution of all known categories corresponding to each pixel. Next, an energy score is calculated for each pixel based on its probability distribution.
[0054] The specific calculation method is to take the natural logarithm of the predicted class probability for each pixel, then sum the logarithmic values for all categories and take the inverse to obtain the energy score for that pixel. Pixels with higher energy scores indicate that the model has less confidence in its judgment that the pixel belongs to a known class. A pre-set energy score threshold is then used to compare the energy score of each pixel against this threshold. If a pixel's energy score exceeds the preset energy score threshold, it is preliminarily marked as a potential unknown defect region; otherwise, it is considered more likely to belong to a known class or a normal region.
[0055] The energy score is calculated based on the logarithmic sum of probability distributions and can quantify the uncertainty of the model's predictions. When the model is fully trained for known categories, its probability distribution is typically highly concentrated. However, for unknown defects, due to a lack of training data, the probability distribution is more dispersed, resulting in a higher energy score. Therefore, as an uncertainty measure, the energy score can effectively identify pixels with low confidence in the model's predictions, which are more likely to correspond to unknown defects. In addition, by presetting the energy score threshold, the labeling sensitivity can be adjusted according to the specific application scenario. For example, the threshold can be increased in industrial inspections with high precision requirements, or lowered in scenarios with high recall requirements, thereby flexibly adapting to different needs.
[0056] Step 102: Visual-linguistic similarity matching
[0057] In vision-language similarity matching, a pre-trained vision-language model is first used to extract features from the input image, generating a visual feature vector for each pixel. Simultaneously, for known defect categories, a textual description of each category is predefined, such as "surface scratches" or "internal cracks," and the corresponding textual feature vector is generated using the vision-language model's text encoder. Next, for each pixel, the similarity between its visual feature vector and the textual feature vectors of all known categories is calculated using the cosine similarity between vectors, which is the dot product of two vectors divided by the product of their moduli.
[0058] The maximum similarity between each pixel and all known text feature vectors is taken as the maximum similarity. A similarity threshold is pre-set and the maximum similarity of each pixel is compared with this threshold. If the maximum similarity of a pixel is lower than the preset similarity threshold, it is considered to have a weak semantic association with the known categories, further supporting the judgment that it is an unknown defect; otherwise, it is considered to belong to a known category.
[0059] The vision-language model learns the semantic correspondence between images and text through pre-training. It can extract pixel-level visual features and match them with text descriptions, providing semantic verification for the identification of unknown defects. Cosine similarity, as a similarity calculation method, can effectively measure the directional consistency between feature vectors and reflect the closeness of the pixel's visual features to the semantics of known categories. When the maximum similarity of a pixel is below the threshold, it indicates that it does not match the semantics of all known categories, increasing its credibility as an unknown defect. This method can eliminate pseudo-anomaly regions that have similar visual features to known categories but higher energy scores through semantic verification, thereby improving the accuracy of unknown defect labeling.
[0060] Step 103: Generate an alarm mask using the dual threshold criterion
[0061] In the dual-threshold criterion for generating an alarm mask, the results of the aforementioned energy distribution analysis and visual-language similarity matching are combined to make a joint judgment on each pixel. The specific judgment method is: for each pixel, check whether its energy score exceeds the preset energy score threshold, and at the same time check whether its maximum similarity is lower than the preset similarity threshold. If a pixel meets both the conditions of the energy score exceeding the energy score threshold and the maximum similarity being lower than the similarity threshold, it is marked as an unknown defect area; otherwise, it is considered a known category or a normal area. Based on this judgment, an alarm mask is generated, in which pixels marked as unknown defects are assigned a value of 1 and other pixels are assigned a value of 0, forming a binary image mask used to represent the distribution of unknown defect areas in the input image.
[0062] The dual-threshold criterion combines energy scores and maximum similarity judgments, integrating information from model prediction uncertainty analysis and semantic similarity verification. The energy score reflects the confidence of the model, while maximum similarity provides semantic constraints. Combining these two can reduce the potential for misjudgment caused by a single metric. For example, relying solely on energy scores can mislabel some known but complex areas as unknown defects. Incorporating similarity judgments can effectively eliminate such false anomalies. The alarm masks generated by this method are more accurate and reliable, providing pixel-level precise labeling for subsequent anomaly analysis and processing, facilitating further system processing.
[0063] Step 104: Output to downstream
[0064] In the step of outputting to the downstream, the alarm mask generated by the dual-threshold criterion is output as the result and passed to the subsequent adaptive spectral clustering and pseudo-label generation process.
[0065] Specifically, the alarm mask is saved as an image, where pixel values of 1 represent unknown defects and values of 0 represent areas with known defects. The output process ensures that the resolution of the alarm mask matches the input image, allowing subsequent processes to directly use the mask for pixel-level analysis and processing.
[0066] The output of the alarm mask provides structured input data for subsequent adaptive spectral clustering and pseudo-label generation, ensuring the continuity of the entire process. Regions with pixel values of 1 clearly indicate the location of unknown defects, facilitating subsequent clustering algorithms to group and analyze these regions or further optimize the model through pseudo-label generation. The resolution of the alarm mask remains consistent with the input image, ensuring complete information transfer and enabling the system to achieve precise anomaly detection and model updates at the pixel level, thereby improving overall adaptive capabilities.
[0067] Step 2: Perform adaptive spectral clustering on the feature set, specifically including constructing a similarity matrix of the feature set, calculating a degree matrix based on the similarity matrix, constructing a Laplace matrix from the similarity matrix and the degree matrix, performing eigendecomposition on the Laplace matrix, determining the number of clusters using the eigenvalue gap method, extracting the first several eigenvectors and applying K-means clustering to generate cluster labels.
[0068] The second step includes the following:
[0069] Step 201: Map the alarm mask to the unified embedding space
[0070] In the process of mapping the warning mask to the unified embedding space, the input image is firstly subjected to feature extraction, and a pre-trained visual model is used to generate a feature map of the image, which contains a high-dimensional feature vector at each pixel position.
[0071] Next, the alarm mask generated in step S1 is adjusted to the same spatial resolution as the feature map. Pixels with a value of 1 in the alarm mask represent unknown defect areas, while pixels with a value of 0 represent non-defective or known category areas. After this adjustment, the two are spatially aligned. Feature vectors corresponding to pixels with a value of 1 in the alarm mask are then extracted from the feature map to form a feature set. This feature set contains the feature vectors of all unknown defect pixels.
[0072] The reason for using pre-trained visual models to extract features is that pre-trained models can capture deep semantic information in images and provide high-quality feature representations for easy subsequent processing. Figure 1 This ensures spatial alignment and avoids information loss or matching errors caused by resolution differences. Extracting a feature set of unknown defect pixels focuses on the features of the unknown defect area, reducing interference from irrelevant data. This improves computational efficiency and provides targeted input data for subsequent analysis.
[0073] Step 202: Adaptive spectral clustering
[0074] In adaptive spectral clustering, a similarity matrix is first constructed. Each element of the similarity matrix represents the similarity between two pixel feature vectors. The similarity is calculated using a Gaussian kernel function based on the negative exponential form of the Euclidean distance between feature vectors.
[0075] Next, we calculate the degree matrix, which is a diagonal matrix whose diagonal elements are the sum of the elements in the corresponding rows of the similarity matrix, representing the sum of the similarities between each pixel feature and all other features. Then, we construct the Laplacian matrix, which is obtained by subtracting the similarity matrix from the degree matrix.
[0076] After that, the Laplace matrix is eigendecomposed to obtain the first several eigenvalues and their corresponding eigenvectors sorted from small to large, and the optimal number of clusters is determined by the eigenvalue gap method. The eigenvalue gap method calculates the difference between adjacent eigenvalues and selects the position with the largest difference as the number of clusters.
[0077] Finally, the first several eigenvectors are extracted to form a new feature matrix, and the K-means clustering algorithm is applied to the new feature matrix to generate a cluster label for each pixel.
[0078] The Gaussian kernel function is used to construct the similarity matrix. This function reflects the local structure in the feature space, enhances the connections between similar features, and improves clustering accuracy. The degree and Laplacian matrices are constructed because they are standard methods for spectral clustering. They transform the clustering problem into a graph partitioning problem, facilitating dimensionality reduction and grouping through feature decomposition. The number of clusters is automatically selected based on the intrinsic structure of the data, improving the objectivity of the clustering results. The reduced feature space makes it easier to separate cluster structures and generate accurate cluster labels.
[0079] Step 203: Generate pseudo labels
[0080] To generate pseudo-labels, each cluster is first assigned a unique pseudo-class label that is independent of and unique to the labels of known defect classes. Next, the cluster labels generated by clustering are mapped back to the spatial resolution of the original image to generate a pseudo-label map. In this pseudo-label map, pixels with a value of 1 in the alarm mask are assigned a pseudo-class label based on their cluster label. Pixels with a value of 0 in the alarm mask are assigned a label of 0, indicating non-defective areas or known class regions. Finally, the pseudo-label map is output as input data for subsequent steps.
[0081] Each cluster is assigned a unique pseudo-category label, converting clustering results into semantic annotations to facilitate system identification of unknown defect categories and support incremental model learning. Cluster labels are mapped to the original image resolution, ensuring that the pseudo-label map is consistent with the input image space, facilitating subsequent training and validation. Pixels with alarm mask values of 1 and 0 are distinguished and assigned values separately, preserving the difference between known and unknown regions. This generates a structured pseudo-label map, supporting self-supervised learning and system optimization.
[0082] Step 204: Calculate the scarcity index of each cluster
[0083] When calculating the scarcity index for each cluster, we first calculate the feature mean of each cluster. The feature mean is the average of all pixel feature vectors within the cluster. Next, we calculate the internal compactness of the cluster. The internal compactness is the average of the squared Euclidean distances between each pixel feature vector within the cluster and its feature mean, which represents the degree of dispersion of features within the cluster.
[0084] Next, we calculate the inter-cluster separation, which is the minimum squared Euclidean distance between the feature mean of the current cluster and the feature means of all other clusters, and represents the degree of distinction between the cluster and other clusters. Finally, we define a scarcity index as the ratio of inter-cluster separation to internal compactness. A larger scarcity index value indicates that the cluster is more scarce and more likely to represent a new defect type.
[0085] Calculating feature means provides a representation of the cluster's center, facilitating subsequent measurement and simplifying the description of cluster features. Internal compactness measures similarity within a cluster, assessing the concentration of features within the cluster and reflecting cluster cohesion. Inter-cluster separation measures inter-cluster differences, quantifying the uniqueness of a cluster and identifying its distinction from other clusters. The scarcity index combines separation and compactness to comprehensively assess cluster scarcity, providing a priority basis for manual annotation, improving annotation efficiency and targeted model updates.
[0086] Step 3: Select the three pixels with the smallest Euclidean distance to the cluster feature mean from each cluster as representative samples. Submit the image area corresponding to the representative samples to the manual annotation process, assign a defect category label to each representative sample, perform a majority vote on the defect category labels of the three representative samples of each cluster to determine the final category label of the cluster, update the pseudo labels of all pixels in the cluster in the pseudo label map to the final category label of the cluster, and record the number and name of the new defect category in the category dictionary. Output the updated pseudo label map and category dictionary;
[0087] The step three includes the following:
[0088] Step 301: Select representative samples based on scarcity
[0089] When selecting representative samples based on scarcity, we first obtain the scarcity index for each cluster from step 2. This index measures the scarcity of a cluster and reflects its representativeness within the dataset. Next, within each cluster, we calculate the Euclidean distance between each pixel's feature vector and the cluster's mean. The Euclidean distance represents the straight-line distance between a feature vector and the cluster's mean in high-dimensional space.
[0090] Then, all pixels within the cluster are sorted in ascending order based on their Euclidean distance to the cluster feature mean. The three pixels with the smallest distance are selected. These pixels are considered closest to the cluster center and represent the cluster's typical features because of their closest distance to the cluster feature mean. Finally, the image regions corresponding to these three pixels are extracted from the input image. Image regions are fixed-size image patches centered around the pixels and are submitted to the manual annotation process.
[0091] The three pixels closest to the cluster feature mean are selected as representative samples. These samples best reflect the overall feature distribution of the cluster, effectively reducing the number of labeled samples while ensuring representativeness, improving labeling efficiency and the accuracy of subsequent processing. Euclidean distance is used as a metric. Euclidean distance accurately reflects the similarity between feature vectors in high-dimensional space, providing an objective and universal basis for sample selection. Image regions, rather than individual pixels, are submitted. Image regions contain contextual information surrounding the pixels, making it easier for human annotators to understand the sample content and accurately determine the defect category.
[0092] Step 302: Manual rapid labeling
[0093] During the manual rapid labeling process, the image areas corresponding to the three representative samples of each cluster are first displayed to the manual labeler for category judgment. Next, the manual labeler assigns a true defect category label to each representative sample based on the content of the image area. The defect category label may be a known defect category or a newly discovered defect category. Then, if a representative sample is confirmed to be a new defect category, a new number is assigned to the new defect category, and the number and category name are recorded in the category dictionary, which is a dynamic list used to store all defect categories. If the representative sample belongs to a known defect category, it is directly labeled using the number already in the category dictionary. Finally, a majority vote is performed on the defect category labels of the three representative samples of each cluster to determine the final category label of the cluster. The majority voting rule is to select the category that appears the most times among the three labels as the final category label. If the three labels are different, one is randomly selected from the three labels as the final category label.
[0094] The final cluster category label is determined by majority voting. Comprehensive judgment based on the labels of multiple representative samples can reduce the impact of individual sample labeling errors and improve the reliability and consistency of cluster category determination. Recording new defect categories in the category dictionary and dynamically expanding the defect category set can adapt to new defect types that may arise during the production process, enhancing the system's flexibility and applicability. Manually labeling three representative samples allows sufficient information to be obtained from a small number of sample labels, significantly reducing the manual labeling workload while ensuring accuracy.
[0095] Step 303: Synchronously write back the pseudo-label table and category dictionary
[0096] When writing back the pseudo-label table and category dictionary simultaneously, the pseudo-label map is first updated based on the final category label of each cluster, replacing the pseudo-labels of all pixels in the cluster with the final category label of the cluster to ensure the consistency of pixel categories within the cluster. Next, if the final category label of the cluster belongs to a completely new defect category, the number and name of the new defect category are added to the category dictionary. The category dictionary serves as a dynamic list for maintaining all known defect categories and newly added defect categories. If the final category label belongs to a known defect category, there is no need to change the category dictionary; it is only necessary to verify that the categories in the pseudo-label map are consistent with the existing categories in the category dictionary. Finally, the updated pseudo-label map and category dictionary are output for use in subsequent processing steps to support model training and optimization.
[0097] Updating the pseudo-label map and applying the manually annotated accurate category information to all pixels within the cluster corrects errors in the initial pseudo-labels, improves the overall accuracy of the pseudo-label map, and provides reliable data for subsequent model training. Dynamically updating the category dictionary, by recording and managing newly added defect categories, ensures the ability to identify unknown defects and enhances the model's adaptability and scalability. Outputting the updated pseudo-label map and category dictionary ensures that subsequent processing uses the latest category information and annotated data, maintaining the consistency and effectiveness of the entire self-supervised learning process.
[0098] Step 4: The encoder parameters are frozen to maintain the stability of feature extraction, and the decoding head is fine-tuned using a decaying learning rate. The new cluster adaptation layer is set higher than the initial learning rate of the decoding head to accelerate the learning of new features.
[0099] The information elasticity index is determined by calculating the deviation between the old task features and the initial features and applying a Gaussian kernel function. The cross-layer forgetting index is calculated by weighted averaging the output deviations of each layer of the decoding head on the old task samples. The adaptive stability mapping module receives the information elasticity index and the cross-layer forgetting index and uses the Sigmoid function to generate the elastic memory coefficient. The total loss function integrates the cross entropy loss of the new category, the EWC regularization loss weighted by (1-elastic memory coefficient) to protect key parameters, and the distillation loss weighted by the elastic memory coefficient to maintain the consistency of the old task output distribution.
[0100] The step 4 includes the following contents:
[0101] Step 401: Freeze encoder parameters
[0102] During the encoder parameter freezing process, the semantic segmentation model is first divided into an encoder and a decoder head. The encoder is responsible for extracting features from the input image, while the decoder head is responsible for mapping the extracted features into class predictions. Next, during the incremental learning phase, all encoder parameters are fixed, and gradient updates to the encoder parameters are not performed during backpropagation, ensuring that the features extracted by the encoder remain stable. Training is then focused on the decoder head and the new cluster adaptation layer, adjusting only the components directly related to class prediction without adjusting the encoder.
[0103] The encoder has been maturely trained in the old tasks, and the features it extracts are highly stable and reliable for the old categories. Fixed parameters can prevent incremental learning from interfering with the old task knowledge. By maintaining the stability of the feature extraction process, the model can effectively avoid forgetting the feature representation of the old categories, while providing a consistent input basis for subsequent decoding and adaptation to new categories.
[0104] Step 402: Fine-tune the dock and the new cluster adaptation layer
[0105] When fine-tuning the decoder head and the new cluster adaptation layer, a learning rate decay strategy is first implemented for the decoder head. A predetermined initial learning rate is set, and the learning rate is gradually reduced with each training round to limit the adjustment of the decoder head parameters and avoid excessive parameter changes. Next, a higher initial learning rate is set for the new cluster adaptation layer, ensuring that its initial learning rate is higher than that of the decoder head, thereby accelerating the learning of the new class features by the new cluster adaptation layer. Finally, supervised training is performed on samples from the new class using the cross-entropy loss function. The parameters of the decoder head and new cluster adaptation layer are adjusted by calculating the difference between the model's predicted probabilities and the true class labels.
[0106] Adopting a learning rate decay strategy, gradually reducing the learning rate of the decoder head can control the step size of parameter adjustment, preventing parameter oscillation caused by excessively large step sizes in the later stages of training, enhancing the model's convergence stability and improving its generalization ability on existing data. Setting a higher initial learning rate for the new cluster adaptation layer is crucial. The new cluster adaptation layer needs to quickly learn the feature representations of new categories, and a higher learning rate can accelerate this process. This improves the model's adaptability and learning efficiency to new tasks. The cross-entropy loss function is a commonly used loss calculation method in classification tasks. It effectively measures the gap between predicted results and true labels, providing accurate optimization guidance and ensuring that the model's predictive ability for new categories gradually improves.
[0107] Step 403: Calculate the information elasticity index and the cross-layer forgetting index
[0108] In the process of calculating the information elasticity index and the cross-layer forgetting index, the information elasticity index is calculated first. The information elasticity index is used to quantify the model's ability to retain old task knowledge when learning new categories.
[0109] Before the start of incremental training, the encoder's feature representation of the old task samples is recorded. Then, during the training process, the difference between the current feature representation and the pre-training feature representation is calculated. The difference is measured by the Euclidean distance of the feature vector, and then the distance difference is converted into an exponential form using a Gaussian kernel function. Finally, the exponential values of all old task samples are averaged to obtain the information elasticity index.
[0110] Next, we calculate the cross-layer forgetting index, which measures the degree to which the model has forgotten previous task knowledge at different levels. We select the shallow, mid, and deep layers of the decoder head and calculate the difference between the output of each layer on the previous task sample and the pre-training output. This difference is measured using the Euclidean distance of the output vectors. The differences are averaged across all previous task samples, and then the weighted sum of the average differences across each layer is taken to obtain the cross-layer forgetting index.
[0111] By quantifying the degree of change in the encoder's feature representations, we can objectively assess the model's ability to retain knowledge from previous tasks, providing a quantifiable metric to guide subsequent adjustments to training strategies. The Gaussian kernel function can smoothly convert distance differences into an exponential form, enhancing sensitivity to small changes; improving the accuracy of the assessment and reducing the impact of extreme values; and calculating a cross-layer forgetting index. The output changes at different layers of the decoder head reflect the model's forgetting at different levels of abstraction. Comprehensively monitoring these changes helps understand the distribution of forgetting. Providing multi-level forgetting information facilitates optimization of the model's overall performance. By weighting the summation of layer differences, different layers have varying degrees of influence on model predictions, and weighting can highlight the role of key layers.
[0112] Step 404: The adaptive stability mapping module outputs the elastic memory coefficient EMC
[0113] In the process of the adaptive stability mapping module outputting the elastic memory coefficient EMC, the calculated information elasticity index and cross-layer forgetting index are first input into the adaptive stability mapping module.
[0114] Next, the adaptive stability mapping module generates an elastic memory coefficient through a nonlinear mapping function and uses the Sigmoid function to process the input. The input of the Sigmoid function is the weighted value of the information elasticity index minus the cross-layer forgetting index, where the weight is determined by the preset hyperparameters. The output of the Sigmoid function is the elastic memory coefficient, and the value range of the elastic memory coefficient is between 0 and 1.
[0115] The Sigmoid function is used as the mapping function. It can smoothly map input values to the range of 0 to 1, making it suitable for representing the value characteristics of the coefficients. This generates a bounded and continuous output, facilitating subsequent adjustment of the loss function weights. The information elasticity index and the cross-layer forgetting index, which reflect feature stability and output consistency respectively, provide a comprehensive assessment of the current state of the model. This allows the elastic memory coefficient to more accurately reflect the model's balance between new and existing tasks. By setting hyperparameters to control weights, different application scenarios may require adjusting the emphasis on the information elasticity index and the cross-layer forgetting index. This enhances the system's flexibility and adapts it to diverse task requirements.
[0116] Step 405: Dynamically adjust EWC regularization and distillation weights
[0117] In the process of dynamically adjusting EWC regularization and distillation weights, the EWC regularization loss is first defined. The EWC regularization loss is achieved by performing a weighted penalty on the difference between the model parameters and the optimal parameters of the old task. The weight is determined by the Fisher information matrix, which is used to measure the importance of each parameter to the old task.
[0118] Then, a distillation loss is defined, which measures the difference in output distribution by computing the KL divergence between the current output of the model and the pre-training output on the old task samples. Then, the total loss function is defined as the weighted sum of the cross-entropy loss, the EWC regularization loss and the distillation loss, where the weight of the EWC regularization loss is 1 minus the elastic memory coefficient, and the weight of the distillation loss directly uses the elastic memory coefficient.
[0119] Using the EWC regularization loss, by imposing protective constraints on important parameters, the damage to old task knowledge in incremental learning can be effectively reduced. Maintain the predictive ability of the model on the old task, avoid catastrophic forgetting. Using the distillation loss, by keeping the consistency of the output distribution, the knowledge of the old task can be preserved while learning the new task. Smooth migration of new and old task knowledge, improve the overall stability of the model. Dynamically adjust the weights of EWC regularization loss and distillation loss, adaptively balance the learning needs of new and old tasks according to the steady state and forgetting degree of the model. Improve the efficiency of incremental learning and ensure the balanced performance of the model on new and old tasks. Use the elastic memory coefficient to adjust the weight, which can reflect the current state of the model in real time, realize the dynamic allocation of the weight, and make the training process more in line with the actual demand.
[0120] Step five, parallel enable self-attention reconstruction branch, access the feature map output by the encoder, generate attention weight matrix through the self-attention module, the attention weight matrix captures the global context information of the feature map to generate the global context feature map; the global context feature map is input into the reconstruction decoder to generate the reconstructed image; calculate the pixel-level difference between the reconstructed image and the original input image to generate the residual image, and the residual image is normalized and upsampled to the same resolution as the segmentation decoder feature map;
[0121] The upsampled residual image and the segmentation feature map are concatenated along the channel dimension to generate a fused feature map, which is used for subsequent calculations of the segmentation decoder to generate segmentation prediction results; define the mean square error loss of the reconstruction task and the cross-entropy loss of the segmentation task, and construct the total loss function as the weighted sum of the mean square error loss and the cross-entropy loss. Calculate the reconstructed image and the segmentation prediction result simultaneously in the forward propagation, and update the weights of the reconstruction decoder and the segmentation decoder through back propagation.
[0122] The step five includes the following contents:
[0123] Step 501, parallel enable self-attention reconstruction branch
[0124] In the process of parallel enabling the self-attention reconstruction branch, first, a self-attention module is connected after the feature map output by the encoder. The self-attention module generates an attention weight matrix by calculating the similarity between each position in the feature map, and the attention weight matrix is used to capture the global context information of the feature map.
[0125] The features of each position in the feature map are compared with the features of all other positions, similarity scores are calculated, and these scores are normalized to form an attention weight matrix. Then, the feature map is weighted using the attention weight matrix to generate a global context feature map that integrates semantic information from different regions of the image. By weighted summing the features of each position in the feature map through the attention weight matrix, a new feature representation is obtained. Then, the global context feature map is input into the reconstruction decoder to generate a reconstructed image, which is an approximate reconstruction of the original input image. The reconstruction decoder progressively upsamples and transforms the global context feature map to output pixel value representations of the same resolution as the original image.
[0126] The access self-attention module can effectively capture long-range dependencies in the image and improve the model's understanding of global semantics. The accuracy of the reconstruction task is enhanced, making the reconstructed image closer to the original image. By weighting the feature map using the attention weight matrix, the features of different positions are fused through weighted integration, which can highlight the regions that contribute to the reconstruction task. The quality of the reconstructed image is improved and the interference of irrelevant information is reduced. By supervising the model to learn the complete representation of the image through the reconstruction task, the model's ability to distinguish between normal and abnormal regions is improved. The segmentation task is better assisted to identify the defect region.
[0127] Step 502, generate residual map
[0128] In the process of generating the residual map, first calculate the pixel-level difference between the reconstructed image and the original input image.
[0129] For each pixel position, calculate the Euclidean distance between the RGB values of the reconstructed image and the original image at that position. The Euclidean distance is obtained by calculating the sum of the squares of the differences in the RGB three channels respectively and then taking the square root. A residual map is generated, which reflects the difference between the reconstructed image and the original image. Then, the residual map is normalized. The normalization operation scales the values in the residual map to the range of 0 to 1. Specifically, find the maximum and minimum values in the residual map, and linearly map each pixel value to this range to obtain a normalized residual map. The normalized residual map is convenient for subsequent fusion with the segmentation feature map.
[0130] In use, by comparing the reconstructed image with the original image, the abnormal region in the image can be highlighted, because the reconstruction error of the normal region is smaller and the reconstruction error of the abnormal region is larger; a residual map that highlights the defect region is generated, providing auxiliary information for the segmentation task. Normalization processing ensures that the numerical range of the residual map matches the segmentation feature map, avoiding the influence of too large numerical differences on the feature fusion effect. The stability and consistency of feature fusion are improved.
[0131] Step 503: Provide pixel-level playback signal
[0132] In the process of providing pixel-level playback signals, the normalized residual image is first upsampled to the same resolution as the segmentation decoder feature map. The upsampling method uses bilinear interpolation. According to the target resolution, the new pixel value is calculated by weighted average of the surrounding pixel values to ensure that the residual image and the feature map are spatially aligned. Next, in the middle layer of the segmentation decoder, the upsampled residual image and the segmentation feature map are spliced along the channel dimension to generate a fused feature map. The fused feature map contains the segmentation features and residual information. The splicing operation is to directly add the number of channels of the two to form a new feature representation. The fused feature map is then used to continue the subsequent calculations of the segmentation decoder to generate the final segmentation prediction result. Specifically, the decoder gradually upsamples and transforms the fused feature map to output the pixel-level classification result.
[0133] Upsampling the residual image ensures that it and the segmentation feature map have the same spatial resolution, facilitating feature fusion. This achieves spatial alignment of information and avoids errors caused by inconsistent resolution. Joining the residual image with the segmentation feature map, by introducing residual information, provides additional clues about abnormal regions for the segmentation task, enhancing the model's focus on defect areas. This improves the segmentation task's accuracy for identifying new defect types. Subsequent calculations using the fused feature map contain richer information, helping the model make more accurate predictions and improving the overall quality of the segmentation results.
[0134] Step 504: Update the weights of the two branches in the same forward process
[0135] In the process of updating the weights of the two branches in the same forward pass, we first define the loss function for the reconstruction task. We use mean squared error to monitor the consistency of the reconstructed image with the original image. The mean squared error measures the reconstruction quality by calculating the sum of the squares of the differences in each pixel value between the reconstructed image and the original image, and then taking the average. Next, we define the loss function for the segmentation task. We use cross-entropy loss to monitor the consistency of segmentation predictions and pseudo-labels. Cross-entropy loss optimizes the segmentation model by calculating the difference between the predicted probability distribution and the true label distribution.
[0136] Next, the total loss function is constructed as the weighted sum of the reconstruction loss and the segmentation loss. These two losses are weighted and added together using preset hyperparameters to control the relative importance of the two tasks. Finally, the reconstructed image and segmentation prediction are simultaneously calculated in the forward propagation. Backpropagation then simultaneously updates the weights of the reconstruction decoder and the segmentation decoder. Specifically, the network parameters are derived and adjusted based on the total loss, achieving coordinated optimization of the two branches.
[0137] When used, the mean square error can directly reflect the pixel-level difference between the reconstructed image and the original image, making it easier to optimize the reconstruction task, improve the accuracy of the reconstructed image, and thus improve the quality of the residual image. Using cross entropy loss as the segmentation loss is a commonly used loss function in classification tasks and can effectively optimize the accuracy of category predictions.
[0138] Improving the segmentation model's prediction accuracy and constructing a weighted sum as the total loss function, by balancing the losses of the reconstruction and segmentation tasks, enables collaborative training of the two tasks, enhancing knowledge transfer and performance balance between the new and old tasks. Updating weights in the same forward pass improves training efficiency and ensures simultaneous optimization of the reconstruction and segmentation tasks, reducing training time and improving overall model performance.
[0139] Step 6: Automatic mixed-precision inference is performed on the fine-tuned model. Specifically, the weights and input tensors of the convolutional and fully connected layers are converted from full precision to half precision, the input and output tensors of the activation and pooling layers are converted from full precision to integer, and the input and output tensors of the Softmax layer and loss function are kept at full precision.
[0140] Through heterogeneous parallel scheduling, half-precision tensor calculations are allocated to the GPU, integer tensor calculations are allocated to the NPU, and full-precision control flow is allocated to the CPU; online incremental training and real-time inference are executed concurrently, where the training task uses new pseudo-labels and fine-tuning strategies to update model parameters in the background, and the inference task uses the current model in the foreground to perform real-time defect segmentation on the input image; after the model is deployed on the edge device, online incremental training is performed regularly, and the performance is tested using the validation set after the update. If the performance deteriorates by more than the preset threshold, it is rolled back to the previous stable version.
[0141] The step six includes the following contents:
[0142] Step 601: Automatic mixed precision reasoning
[0143] During automatic mixed-precision inference, the tensors in the model are first classified into half-precision, integer, and full-precision tensors based on their computational requirements. Half-precision tensors are used for the weights and input tensors of convolutional and fully connected layers, integer tensors are used for the input and output tensors of activation functions and pooling layers, and full-precision tensors are used for the input and output tensors of the Softmax layer and loss function. Next, during inference, the model weights and activation values are automatically converted to the corresponding precision. The weights and input tensors of the convolutional and fully connected layers are converted from full precision to half precision, the input and output tensors of the activation functions are converted from full precision to integer, and the input and output tensors of the Softmax layer and loss function are maintained at full precision. Tensor cores in edge devices are then used to accelerate half-precision computations, and inference speed is improved through parallel matrix multiplication.
[0144] By assigning different types of tensors to appropriate precision, computational complexity and memory usage can be significantly reduced while maintaining accuracy. This improves model inference speed on edge devices, meeting real-time requirements. Tensor Cores are used to accelerate half-precision calculations. Tensor Cores are specifically designed for matrix operations and can efficiently process half-precision data. This further shortens inference time and meets the time-sensitive demands of industrial quality inspection.
[0145] Step 602: Heterogeneous parallel scheduling
[0146] During heterogeneous parallel scheduling, computing tasks are first mapped to different hardware based on tensor type: half-precision tensor computations are assigned to the GPU, integer tensor computations to the NPU, and full-precision control flow to the CPU. Next, task queues and priority scheduling algorithms are used to dynamically allocate tasks based on their real-time requirements and hardware load. Inference tasks are prioritized over training tasks, and data and model parallelism strategies are used to distribute tasks to the GPU and NPU.
[0147] Then, online incremental training and real-time inference are performed concurrently. The training task uses new pseudo labels and fine-tuning strategies to update model parameters in the background, and the inference task uses the current model to perform real-time defect segmentation on the input image in the foreground.
[0148] Heterogeneous parallel scheduling leverages the computing strengths of GPUs, NPUs, and CPUs to achieve efficient task coordination and meet the high performance demands of resource-constrained edge devices. A priority scheduling algorithm ensures that real-time-critical inference tasks are prioritized. Concurrent training and inference enable seamless integration of online model evolution and real-time applications, enhancing the system's continuous learning capabilities while maintaining normal production line operations.
[0149] Step 603: Edge deployment
[0150] During edge deployment, the fine-tuned model is first converted to a format supported by the edge device, converting the model to ONNX or TensorRT format. Next, the model is deployed on the edge device, combining automatic mixed precision and heterogeneous parallel scheduling to perform real-time inference, using the optimized model to perform millisecond-level semantic segmentation on input images. Online incremental training is then regularly performed using newly generated pseudo-labels and verified annotated data, updating model parameters within minutes. Finally, after each model update, performance is automatically tested using a validation set. If performance metrics degrade by more than a preset threshold, the model is rolled back to the previous stable version.
[0151] Converting models to a format supported by edge devices ensures they can run efficiently on resource-constrained edge devices, improving model compatibility and execution efficiency. Real-time inference meets the timeliness requirements of defect detection in industrial quality inspection scenarios and reduces production line downtime. Regular online incremental training continuously updates models to adapt to new defect types, while performance verification and rollback mechanisms prevent performance degradation caused by model updates, ensuring model reliability in open environments.
[0152] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0153] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0154] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only for some logical functions. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0155] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0156] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A real-time image semantic segmentation system with self-supervised learning, characterized by: include, The alarm mask generation module performs energy distribution analysis and visual-language similarity matching on the input image simultaneously, and generates and outputs unknown defect alarm masks at the pixel level based on the dual-threshold criterion; The adaptive spectral clustering module maps the alarm mask to a unified embedding space and then inputs it into the adaptive spectral clustering module to automatically determine the number of clusters, generate pseudo labels, and calculate the scarcity index of each cluster; The manual labeling module selects three representative samples from each cluster based on their scarcity for manual rapid labeling, and writes the confirmed results back to the pseudo-label table and category dictionary simultaneously; The fine-tuning module only fine-tunes the tuning terminal and new cluster adaptation layer under the condition of freezing the encoder parameters. In each round of training, the information elasticity index and cross-layer forgetting index are output through the adaptive stability mapping module to dynamically adjust the EWC regularization and distillation weights. Self-attention reconstruction module, which enables the self-attention reconstruction branch in parallel during training, provides pixel-level playback signals to the segmentation branch through the residual map and updates the weights of the two branches in the same forward process; The automatic mixed-precision inference module uses automatic mixed-precision inference on the fine-tuned model. It performs online incremental training and inference tasks, combining heterogeneous parallel scheduling with half-precision tensor computation on the GPU, integer tensor computation on the NPU, and full-precision control flow on the CPU. The module is deployed on the edge. Perform forward inference on the input image to generate class prediction logits for each pixel, and calculate an energy score for each pixel based on the class prediction logits, where the energy score is obtained by applying a negative logarithm and function to the class prediction logits; Set an energy threshold and mark pixels whose energy scores exceed the energy threshold as potential unknown defect areas. Use the pre-trained vision-language model to extract pixel-level visual features of the input image. Generate text feature vectors for known defect categories, calculate the cosine similarity between the visual features of each pixel and the text features of all known categories, and take the maximum similarity; Set a similarity threshold and identify pixels whose maximum similarity is lower than the similarity threshold as unknown defects; A dual-threshold criterion is applied to each pixel. When the energy score exceeds the energy threshold and the maximum similarity is lower than the similarity threshold, the pixel is marked as an unknown defect. An alarm mask is generated, and pixels marked as unknown defects are assigned a value of 1, and the remaining pixels are assigned a value of 0. The alarm mask is output.
2. The self-supervised learning real-time image semantic segmentation system according to claim 1, characterized in that: Perform adaptive spectral clustering on a set of features, Including, constructing a similarity matrix of feature sets; The degree matrix is calculated based on the similarity matrix, and the Laplace matrix is constructed by the similarity matrix and the degree matrix. After the Laplace matrix is eigendecomposed, the eigenvalue gap method is used to determine the number of clusters. The first several eigenvectors are extracted and K-means clustering is applied to generate cluster labels.
3. The self-supervised learning real-time image semantic segmentation system according to claim 2, characterized in that: From each cluster, the three pixels with the smallest Euclidean distance to the cluster feature mean are selected as representative samples. The image areas corresponding to the representative samples are submitted to the manual annotation process, and a defect category label is assigned to each representative sample.
4. The self-supervised learning real-time image semantic segmentation system according to claim 3, characterized in that: A majority vote is performed on the defect category labels of the three representative samples of each cluster to determine the final category label of the cluster. The pseudo labels of all pixels in the cluster in the pseudo label map are updated to the final category label of the cluster, and the number and name of the new defect category are recorded in the category dictionary. The updated pseudo label map and category dictionary are output.
5. The self-supervised learning real-time image semantic segmentation system according to claim 4, characterized in that: The encoder parameters are frozen to maintain the stability of feature extraction, the decoding head is fine-tuned with a decaying learning rate, and the new cluster adaptation layer is set to a higher initial learning rate than the decoding head to accelerate the learning of new features; The information elasticity index is determined by calculating the deviation between the old task features and the initial features and applying a Gaussian kernel function. The cross-layer forgetting index is calculated by weighted averaging the output deviations of each layer of the decoding head on the old task samples.
6. The self-supervised learning real-time image semantic segmentation system according to claim 5, characterized in that: The adaptive stability mapping module receives the information elasticity index and the cross-layer forgetting index and uses the Sigmoid function to generate the elastic memory coefficient; The total loss function integrates the cross entropy loss of the new category, the EWC regularization loss weighted by 1 minus the elastic memory coefficient to protect key parameters, and the distillation loss weighted by the elastic memory coefficient to maintain the consistency of the output distribution of the old task.
7. The self-supervised learning real-time image semantic segmentation system according to claim 6, characterized in that: In parallel, the self-attention reconstruction branch is enabled, which is connected to the feature map output by the encoder. The attention weight matrix is generated through the self-attention module. The attention weight matrix captures the global context information of the feature map to generate a global context feature map. The global context feature map is input into the reconstruction decoder to generate a reconstructed image; The pixel-level difference between the reconstructed image and the original input image is calculated to generate a residual map, which is normalized and upsampled to the same resolution as the segmentation decoder feature map.
8. The self-supervised learning real-time image semantic segmentation system according to claim 7, characterized in that: The upsampled residual map and the segmentation feature map are concatenated along the channel dimension to generate a fused feature map. The fused feature map is used for subsequent calculations of the segmentation decoder to generate segmentation prediction results. Define the mean square error loss of the reconstruction task and the cross entropy loss of the segmentation task, construct the total loss function as the weighted sum of the mean square error loss and the cross entropy loss, calculate the reconstructed image and segmentation prediction results simultaneously in the forward propagation, and update the weights of the reconstruction decoder and segmentation decoder through backpropagation.
9. The self-supervised learning real-time image semantic segmentation system according to claim 8, characterized in that: Automatic mixed-precision inference is used on the fine-tuned model, including: Convert the weights and input tensors of the convolutional and fully connected layers from full precision to half precision, convert the input and output tensors of the activation and pooling layers from full precision to integer, and keep the input and output tensors of the Softmax layer and loss function in full precision.
10. The self-supervised learning real-time image semantic segmentation system according to claim 9, characterized in that: Through heterogeneous parallel scheduling, half-precision tensor calculations are allocated to the GPU, integer tensor calculations are allocated to the NPU, and full-precision control flow is allocated to the CPU; Online incremental training and real-time inference are performed concurrently. The training task uses new pseudo-labels and fine-tuning strategies to update model parameters in the background, while the inference task uses the current model to perform real-time defect segmentation on the input image in the foreground. After the model is deployed on the edge device, online incremental training is performed regularly, and the performance is tested using the validation set after the update. If the performance drops by more than the preset threshold, it is rolled back to the previous stable version.
Citation Information
Patent Citations
Medical image classification method based on prior knowledge enhanced mask and alignment modeling
CN119399523A
Multi-task hybrid supervised medical image segmentation method and system based on federated learning
JP7386370B1