Graph-guided data labeling governance method for curved shell defect detection

By employing a graph-guided data annotation governance method, combined with mask-guided data augmentation and a dual-network nested framework, the annotation anomalies and consistency issues of industrial curved shell defect datasets were resolved. This improved the data annotation quality and model generalization ability, enabling the model to adapt to complex industrial scenarios and ensure production safety.

CN120726425BActive Publication Date: 2025-11-07SOUTHWEAT UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511216639.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-07
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as annotation anomalies, inconsistent annotation quality, high computational resource consumption, and long-tail distribution when processing industrial curved shell defect datasets. These issues make it difficult to generate high-quality data annotation sets, affecting model generalization ability and production safety.

Method used

A graph-guided data annotation governance approach is adopted, which combines mask-guided data augmentation and a dual-network nested framework, including a tail-aware decoupled contrastive learning model and a multi-scale self-representation learning model. Through soft label generation and active learning strategies, the data annotation process is optimized.

Benefits of technology

It significantly improves the data annotation quality and model generalization ability of defect detection in curved shells, adapts to complex industrial scenarios, improves annotation governance efficiency and accuracy, and ensures product quality and production safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726425B_ABST
    Figure CN120726425B_ABST
Patent Text Reader

Abstract

The application discloses a graph-guided data annotation management method for curved shell defect detection, and belongs to the technical field of data annotation management, and comprises the following steps: S1, collecting curved shell defect data and annotating the curved shell defect data; S2, based on the annotated curved shell defect data, performing data enhancement on the defect area through mask guidance to generate a data annotation set; S3, inputting the data annotation set into a double-network nested framework to generate a prediction result of a soft label; and S4, based on the prediction result of the soft label, adopting a graph-guided active learning strategy to iteratively optimize the data annotation set and the double-network nested framework to complete data annotation management. Through the data enhancement strategy of mask guidance and in combination with the double-network nested framework, the system improves the data annotation quality of high-resolution curved shell images with complex backgrounds such as curved distortion, high light interference and long-tail distribution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data labeling governance, and particularly relates to a graph-guided data labeling governance method for curved shell defect detection. BACKGROUND

[0002] At present, the labeling anomaly of industrial datasets has evolved into an inherent systematic problem, with diverse sources and complex manifestations. On the one hand, due to the high cost of expert labeling, many enterprises tend to use low-cost non-professional manual labeling, which leads to the influence of subjective factors on labeling quality and makes it difficult to ensure consistency and accuracy. On the other hand, the rotation scanning imaging mechanism in the data collection process of curved shell brings a unique labeling bias. For example, motion blur caused by high-speed rotation can cover up or fake defects, resulting in errors in the labeling results. At the same time, the high dynamic nature of the industrial production environment, such as changes in production line speed, material differences, and changes in the glossiness of the curved shell surface, can cause the same defect to be assigned different labels, further exacerbating the labeling consistency problem.

[0003] In addition, the structural and complex nature of industrial curved shell defect data further exacerbates the difficulty of data labeling governance. First, the high-resolution image of the unfolded curved shell has high-dimensional complex features and a large data size, significantly increasing the computational resource consumption and model training difficulty. Second, the long-tail distribution problem commonly found in industrial defect data is exacerbated by the curvature effect and the variability of production processes. During the generation process, the edge distortion and axial stretching of the curved surface result in a scarcity of some subtle defect samples, which further enhances the bias towards common classes and weakens the representation ability of tail classes during model training. Finally, the curved shell defects account for a very small proportion of the overall unfolded image, and their material is single and their color changes are limited, resulting in a high degree of mixing of surface texture and background information, greatly increasing the difficulty of distinguishing foreground defects from background interference.

[0004] To address the above problems, the Chinese invention patent with publication number CN115564960A proposes a network image label denoising method combining sample selection and label correction. The method includes the following steps: step S1, selecting clean samples through the cosine similarity of samples and class centers; step S2, selecting reusable samples from the remaining samples and correcting them through sample uncertainty dynamics; step S3, updating the network using clean samples and corrected reusable samples. This invention uses cosine similarity for sample selection and corrects labels through clean samples, improving sample utilization and fine-grained classification performance.

[0005] Another Chinese invention patent with publication number CN120030368A proposes a label correction crowd-sourcing result aggregation method and system based on deep clustering. The method first extracts task features through a pre-trained model as input data for the model. Second, a worker loss term is introduced, combined with the reconstruction loss and KL divergence in the variational deep embedding model, to define the comprehensive loss function of the model, and the model is trained through gradient descent method. During the training process, the latent variables are clustered using a Gaussian mixture model to obtain clustering clusters, and specific class labels are mapped to each clustering cluster by combining the worker's annotation information for the task. Finally, the contour coefficient of each task is calculated to identify tasks with low clustering accuracy, which are adjusted and corrected by worker labels to further improve the accuracy of the clustering results. This invention effectively reduces the interference of label noise on true value inference by clustering similar tasks and correcting tasks with low clustering quality using worker labels, improving the performance of the result aggregation model.

[0006] Although the existing two inventions have realized the correction of abnormal labels, there are still obvious deficiencies in dealing with the complexity and annotation abnormalities of industrial curved shell defect datasets. First, the semi-supervised sample selection method based on loss distribution or confidence to generate pseudo-labels for abnormal label correction is highly dependent on the accuracy of the initial model. In scenarios with a high proportion of annotation abnormalities or significant long-tail distribution, the initial classifier is easily affected by annotation abnormalities, leading to sample classification errors and further amplifying annotation governance errors. Second, the deep clustering method based on latent space to mine the semantic association between samples to optimize labels faces a high-dimensional bottleneck on high-resolution datasets, resulting in high computational complexity and difficulty in adapting to large-scale industrial data. At the same time, relying on single-scale feature representation, traditional deep clustering methods fail to fully utilize multi-scale feature information to depict the inherent prevalent structure of semi-supervised industrial datasets, limiting the ability of label correction.

[0007] In summary, the annotation governance of industrial high-resolution datasets urgently needs a comprehensive solution based on a dual-network nested framework to address the unique challenges of industrial curved shell defect datasets. This method not only significantly improves the quality and efficiency of annotation governance, generating high-quality data annotation sets and enhancing the generalization ability of models in complex industrial scenarios, but also adapts to the diverse needs of curved shell defect detection, ultimately ensuring product quality and industrial production safety. SUMMARY

[0008] To address the above deficiencies in the prior art, the graph-guided data annotation governance method for curved shell defect detection provided by the present invention fuses multiple advanced algorithms and technologies, combines a dual-network nested governance framework of contrastive learning and self-representation learning, and realizes the optimization and governance of high-dimensional complex industrial curved shell defect datasets, solving the problems of abnormalities and consistency commonly existing in industrial curved shell defect data annotation.

[0009] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a graph-guided data annotation and management method for defect detection in curved shells, comprising the following steps:

[0010] S1. Collect defect data of curved shell and annotate the defect data of curved shell;

[0011] S2. Based on the labeled curved shell defect data, data augmentation of the defect area is performed using masking guidance to generate a data label set;

[0012] S3. Input the data label set into the double-network nested framework to generate the prediction results of soft labels;

[0013] S4. Based on the prediction results of soft labels, a graph-guided active learning strategy is used to iteratively optimize the data label set and the double-network nested framework to complete the data label governance.

[0014] Furthermore: In S1, the curved shell defect data specifically refers to the curved shell defect image. After acquiring the curved shell defect image, it is preprocessed. The preprocessing method is as follows:

[0015] During the acquisition process, the rotation speed of the curved shell is controlled. Simultaneously, a surface acquisition consistency score is used to ensure a high degree of matching between the shell's rotation speed and the acquisition frequency of the line scan camera. This allows for the complete and continuous acquisition of images of defects in the curved shell. The surface acquisition consistency score is then used to further refine this process. The specific expression is:

[0016]

[0017] In the formula, N The number of image frames acquired. For the first m The shell rotation speed corresponding to the frame image. The target acquisition frequency for the linear scan camera. Maximum rotational speed is used for normalization. Indicates the first m Frame image and the first The region overlap rate of a frame image is used to measure the continuity of image acquisition.

[0018] In an image acquisition system, the contrast and visibility of defect features are improved by adjusting the light source angle and brightness. Curved background suppression is performed based on the line scan direction, and the pixel intensity after curved background suppression is... The specific expression is:

[0019]

[0020] In the formula, is the original pixel intensity, is the scan direction weight, is the suppression coefficient, is the local illumination estimate, is the global illumination mean, is the global illumination standard deviation;

[0021] The expression of the pixel intensity after the scan direction texture smoothing processing on the collected image is specifically:

[0022]

[0023] In the formula, is the original intensity of the neighborhood pixel, q is the scan direction neighborhood representing a one-dimensional window pixel set along the line scanning camera axis, is an adjustment factor based on the scan direction distance and the illumination intensity;

[0024]

[0025] In the formula, is the axial pixel distance, is the intensity difference, is the distance control parameter, is the intensity control parameter;

[0026] In S1, the method for labeling the curved shell defect data is specifically: a consensus-driven labeling mechanism is used to vote for the image, and the labeling result of the first m frame image is specifically:

[0027]

[0028] In the formula, is the total number of votes, is the expert model j for the labeling category of the first m frame image, is the candidate defect category, is the indicator function, is the judgment threshold.

[0029] Further, S2 includes the following sub-steps:

[0030] S21, based on the images in the labeled curved shell defect data, a label mask is generated through labeling or a bounding box, and the label mask is used to separate the defect region and the non-defect region ; ​​

[0031]

[0032]

[0033] wherein, I 1 is an image in the labeled curved shell defect data, M 2 is a labeled mask, is an element-wise multiplication;

[0034] S22, performing random data augmentation on the defect region, including applying at least one of geometric augmentation, illumination and color augmentation, and texture augmentation, the geometric augmentation including scaling, rotation, and translation, the illumination and color augmentation including brightness, contrast, hue, and saturation adjustment, and the texture augmentation including simulating imaging noise and applying Gaussian blur;

[0035] S23, performing image blending on the data-augmented defect region and the non-defect region through a smoothing mask to generate an image of the data-augmented curved shell defect data to obtain a data annotation set;

[0036]

[0037] wherein, is a smoothing mask, is the data-augmented defect region;

[0038]

[0039] wherein, is a Gaussian blur function, is a Gaussian blur standard deviation.

[0040] Further, in S3, the dual-network nested framework includes a contrast learning model based on tail-aware decoupling and a multi-scale self-representation learning model connected to each other, the contrast learning model includes a shared encoder and a projection head connected to each other, and the shared encoder is connected to the multi-scale self-representation learning model, wherein the shared encoder is a ConvNeXt-tiny network, and the workflow of the dual-network nested framework is specifically:

[0041] S31, inputting the data annotation set into the contrast learning model based on tail-aware decoupling, collecting negative samples using a class-aware negative sampling strategy, and generating anchor samples and other samples within a batch;

[0042] S32, input the anchor point sample and other samples into the ConvNeXt-tiny network to generate a multi-scale feature map, input the features after uniform channel into a multi-scale self-representation learning model to generate a consensus self-representation matrix, obtain an affinity matrix through the consensus self-representation matrix, and generate a prediction result of a soft label by using self-representation spectral probability propagation;

[0043] In S32, after the multi-scale feature map is generated, it is mapped to a low-dimensional feature space through a projection head, and a sample pair is constructed by randomly adding random disturbance in the feature space, which is used to calculate a reweighted decoupling loss function. By applying a dynamic weight to the tail class samples, the optimization process of the positive and negative sample pairs is decoupled.

[0044] Further, in S31, the method for collecting negative samples by using a class-aware negative sampling strategy is as follows:

[0045] S311, arrange the classes of the data annotation set in descending order according to the number of samples, and take the first 20% of the classes as the head class set H , take the last 20% of the classes as the tail class set T , and take the remaining classes as the intermediate class set M , calculate the frequency weight of each class;

[0046]

[0047] In the formula, w c is the inverse frequency weight of the class c , and n c is the number of samples of the class c .

[0048] S312, calculate the class-aware sampling probability:

[0049] In response to the fact that the class of the anchor point sample belongs to the tail class set, the negative sample is sampled from the head class, the intermediate class and the tail class, and the sampling probability is as follows:

[0050]

[0051] In the formula, Z i is the anchor point sample, c i is the class of the anchor point sample, p m is the probability of sampling from the intermediate class, c k is the class of the negative sample, is the weight of the negative sample class, p h is the probability of sampling from the head class, is a conditional formula.

[0052]

[0053] In response to the anchor sample's category belonging to either the head class set or the intermediate class set, negative samples are sampled from all classes, with their sampling probabilities... for:

[0054]

[0055] In the formula, C For all categories;

[0056] S313. Introduce semantic similarity guidance during the negative sample collection process, and combine it with category-aware sampling probability to obtain the final negative sample sampling probability. ;

[0057]

[0058] In the formula, For hyperparameters, For category-aware sampling probability, Normalized cosine similarity;

[0059]

[0060] In the formula, Features of other samples within the batch z j Cosine similarity with anchor samples, Temperature parameter used to control similarity sharpness;

[0061]

[0062] In the formula, It is an L2 norm;

[0063] S314. During the collection of negative samples, the set of categories of the negative samples is constrained to satisfy the minimum category coverage and diversity constraints. The expression for satisfying the minimum category coverage is as follows:

[0064]

[0065] In the formula, This is the floor function. This is the diversity control coefficient. n This represents the number of negative samples sampled for each anchor point sample.

[0066] The expression for the diversity constraint is as follows:

[0067] .

[0068] Further, in S32, the ConvNeXt-tiny network includes two-dimensional convolutions, normalization layers, initial levels, intermediate levels, deep levels, and final levels connected in sequence, each level includes several ConvNeXt residual blocks, and the multi-scale feature mapping includes first, second, third, and fourth feature mappings, and the workflow of the ConvNeXt-tiny network is specifically as follows:

[0069] A1, input image based on the ConvNeXt-tiny network I 2, two-dimensional convolution with an input kernel of 4x4, a stride of 2, and a channel number of 96, the output of the two-dimensional convolution is input into the normalization layer to obtain a feature map X 0;

[0070] A2, input the feature map X 0 into the initial level, the initial level includes three ConvNeXt residual blocks, and the first feature mapping X 1 is obtained, wherein the workflow of the ConvNeXt residual block is specifically as follows:

[0071] The input feature of the ConvNeXt residual block is subjected to feature space extraction using a depth convolution block with a kernel of 4x4, a stride of 2, and a channel number of 96, and is normalized by a LayerNorm operation to generate a normalized feature mapping; a 1x1 convolution is used to fuse the inter-channel information of the normalized feature mapping, expand the channel number from 96 to 384, and introduce a smooth nonlinearity using a GELU activation function; a 1x1 convolution is used to restore the channel number from 384 to 96, and the restored feature is combined with the input feature through a residual connection to obtain the output feature of the ConvNeXt residual block;

[0072] A3, input the first feature mapping X 1 into the intermediate level, adjust the channel number to 192 through a down-sampling layer, input the down-sampled feature into three ConvNeXt residual blocks, and obtain the second feature mapping X 2;

[0073] A4, input the second feature mapping X 2 into the deep level, adjust the channel number to 384 through a down-sampling layer, input the down-sampled feature into nine ConvNeXt residual blocks, and obtain the third feature mapping X 3;

[0074] A5, input the third feature mapping X 3 into the final level, adjust the channel number to 768 through a down-sampling layer, input the down-sampled feature into three ConvNeXt residual blocks, and obtain the fourth feature mapping X 4.

[0075] Furthermore: In S32, the workflow of the multi-scale self-representation learning model is as follows:

[0076] B1. The multi-scale feature mapping is unified to 64 channels through 1×1 convolution, and average pooling is performed to obtain the unified channel features.

[0077] B2. Input the unified channel features into the corresponding scale autoencoder. The autoencoder consists of three layers: encoder and decoder. An unbiased fully connected self-representation learning layer is embedded between the encoder and decoder to learn the self-representation coefficient matrix for each scale.

[0078] Specifically, a reconstruction loss is constructed based on the encoder's input and the decoder's output. Minimizing this reconstruction loss prevents model collapse. The specific expression is:

[0079]

[0080] In the formula, for l The input of the layer encoder, for l The output of the layer decoder, s As a scale;

[0081] Between the three-layer encoder and decoder, an unbiased fully connected self-representation learning layer is embedded into the refined features at each scale to learn the self-representation coefficient matrix. This unbiased fully connected self-representation learning layer uses an optimized self-representation loss to ensure that the features of each sample are represented by a linear combination of other samples at the same scale, while applying regularization constraints to avoid trivial solutions. The specific expression is:

[0082]

[0083] In the formula, As a refined feature, For balancing parameters, It is a self-representation coefficient matrix;

[0084] B3. The consensus self-representation matrix is ​​generated by fusing the self-representation coefficient matrices at various scales. The specific process is as follows:

[0085] Statistical features of the self-representation matrix at each scale are extracted using global average pooling, generating a global descriptor for the self-representation coefficient matrix. G ;

[0086]

[0087] In the formula, To stack the self-representation coefficient matrices into a third-order tensor along the channel dimension, is global average pooling;

[0088] The channel weight is obtained by measuring the contribution degree of each channel using the global description factor U ;

[0089]

[0090] wherein, is an unbiased fully connected network, is a ReLu activation function, is a Sigmoid activation function;

[0091] The self-representation coefficient matrix is weighted by the channel weight, and a consensus self-representation matrix is generated using a 3x3 convolution ;

[0092]

[0093] wherein, is a 3x3 convolution;

[0094] In S32, the method for obtaining the affinity matrix is specifically: an affinity matrix W is generated by taking the absolute value average of and the transpose thereof;

[0095]

[0096] wherein, is a transpose symbol;

[0097] The expression of the loss function of the optimized multi-scale self-representation learning model is specifically:

[0098]

[0099] wherein, and are weight coefficients, is a structure preserving loss;

[0100]

[0101] wherein, is the rank of the calculation matrix, is a Laplacian matrix, , is a diagonal matrix degree matrix, whose diagonal elements , is an element in the affinity matrix ;

[0102] In S32, the method for generating a soft label using self-representation spectral probability propagation is specifically:

[0103] C1, Affinity Matrix Perform feature decomposition, before extraction K The feature vectors constitute the embedded feature matrix. ;

[0104]

[0105] In the formula, For the front K 1 eigenvector;

[0106] C2. Use K-means clustering to calculate the Euclidean distance from the centroid to the samples in the curved shell defect dataset, and assign the samples to... K Each cluster generates a hard tag;

[0107] C3. Based on the affinity matrix and hard labels, calculate the sum of similarities between samples and each cluster, and calculate the sample similarity. i and belonging to category k The sum of the similarities of the samples The specific expression is:

[0108]

[0109] In the formula, For hard label k The sample set;

[0110] C4. Generate samples using Softmax transformation based on the sum of similarities between sample pairs and each cluster. i Prediction results of soft tags P i , , P ik For the sample i Category k The probability of a cluster satisfies ;

[0111]

[0112] In the formula, For scaling ;

[0113]

[0114] In the formula, This is a scaling factor to prevent Softmax overflow.

[0115] Furthermore: In S32, positive sample pairs are used to calculate the reweighted decoupling loss function. The specific expression is:

[0116]

[0117] wherein, N is the number of samples in the batch, is the negative sample loss weight for balancing positive and negative loss contribution, is the anchor sample feature projected by the head, is the positive sample pair feature, is the negative sample feature, is the positive sample loss, is the negative sample loss;

[0118]

[0119]

[0120] wherein, is the cosine similarity for measuring the similarity between features, is the temperature parameter for controlling the sharpness of similarity;

[0121]

[0122] wherein, N k is the negative sample set.

[0123] Further, in S3, the training method of the dual-network nested framework is specifically:

[0124] Training the contrast learning model:

[0125] The training set is input into the contrast learning model, and a reweighted decoupled loss function is applied for optimization. The optimizer is configured as SGD, the initial learning rate is 1e-3, and the weight decay is 5e-4. In order to avoid early shock, the training process is matched with a cosine annealing scheduler and a 5-epoch linear warm-up acceleration model to converge. The model is trained for 100 epochs, and in response to the fact that the validation loss fluctuation is less than 0.1% for 10 consecutive epochs, the network parameters are frozen, the ConvNeXt-tiny network is fixed as a multi-scale feature extractor, and multi-scale feature mapping is output.

[0126] Training the multi-scale self-representation learning model:

[0127] The multi-scale feature mapping is input, and the self-representation matrix is optimized through the loss function of the multi-scale self-representation learning model, wherein the reconstruction loss constrains the consistency of the input image and the decoding output, the self-representation loss optimizes the adjacency relationship, and the structure preservation loss forces the semantic structure of the multi-scale features to be uniform; the training process optimizer adopts AdamW, the initial learning rate is 5e-4, the model is trained for 100 epochs, the termination condition is set to be that the loss reduction rate is less than 0.01% for 15 consecutive epochs, and FP16 mixed precision is used in the training process and the early stopping strategy is adopted to ensure the convergence of the model.

[0128] Further, in S4, the data labeling set and the double network nested framework are iteratively optimized by using a graph-guided active learning strategy, and the method is specifically as follows:

[0129] In S41, based on the prediction result and uncertainty of the soft label of the sample, the sample membership and sample entropy value are calculated, the soft label of the sample is labeled and optimized, the labeled corrected label and pseudo label are generated, and the labeling optimization method is specifically as follows:

[0130] In response to the sample membership being greater than or equal to 0.9 or the sample entropy value being less than or equal to 0.1, the soft label of the sample is taken as a pseudo label.

[0131] In response to the sample membership being less than 0.9 and greater than or equal to 0.7 or the sample entropy value being less than or equal to 0.3 and greater than 0.1, it is judged whether the soft label of the sample is consistent with the original label, if yes, the soft label of the sample is taken as a labeled corrected label, and if no, the labeled corrected label is generated through an auditing link.

[0132] In response to the sample membership being less than 0.7 or the sample entropy value being greater than 0.3, the labeled corrected label is generated through an auditing link.

[0133] The method for generating the labeled corrected label through the auditing link is specifically as follows:

[0134] The soft label of the sample is voted by an expert model, and the soft label with an effective vote number exceeding a preset threshold is taken as a labeled corrected label.

[0135] In S42, the data labeling set is updated according to the sample after labeling optimization, and the double network nested framework is retrained according to the data labeling set, wherein the optimizer, momentum parameter and weight decay coefficient remain unchanged during the training process.

[0136] In S43, the data labeling set and the double network nested framework are iteratively optimized repeatedly, and after each round of iterative optimization, the comprehensive performance index of the contrast learning model on the validation set and the data labeling set is verified and compared, and in response to the performance improvement amplitude of two consecutive iterations being less than 0.5%, the iterative optimization of the data labeling set and the double network nested framework is completed.

[0137] The beneficial effects of the present application are:

[0138] (1) The present application provides a graph-guided data labeling management method for curved shell defect detection. Through a mask-guided data augmentation strategy, a dual-network nested framework combining a ConvNeXt-tiny network with tail-aware contrastive learning and a multi-scale self-representation learning model is used to improve the data labeling quality of high-resolution curved shell images with complex backgrounds such as curved distortion, high light interference, and long-tail distribution. In the initial stage, mask guidance is used to achieve targeted data augmentation, which helps the model focus on the essential features of defects and enhances its ability to adapt to diverse defect patterns. In the feature extraction stage, class-aware negative sampling and a reweighted decoupling loss function are introduced to effectively enhance the representation of curved shell defects and alleviate the labeling imbalance caused by long-tail distribution. In the soft label generation stage, a high-quality consensus self-representation matrix is generated by fusing multi-scale specific self-representation matrices to depict the internal structure of the data, and a discriminative soft label is generated using a self-representation spectral probability propagation mechanism. The entire labeling management process is based on information entropy constraints and graph-guided soft labels for active learning. Through iterative optimization, the quality of abnormal labeling is effectively corrected and improved, achieving efficient and accurate data labeling management. The labeling management method of the present application significantly improves the data labeling accuracy and model generalization ability of curved shell defect detection, effectively addressing the data labeling reliability challenges caused by curved characteristics, production process noise, and line acquisition differences.

[0139] (2) The present application uses a dual-network nested framework based on tail-aware decoupled contrastive learning model and multi-scale self-representation learning model for industrial curved shell defect data labeling management, significantly improving the efficiency and quality of labeling management. On the one hand, the contrastive learning model fully utilizes the powerful feature expression ability of convolutional neural networks to extract multi-scale spatial features of high-resolution images, effectively addressing the high-dimensional data characteristics caused by industrial curved shell line scanning and the background information redundancy caused by single material. On the other hand, self-representation learning is used to mine the semantic relationships between samples, capture the internal popular structure of the data, form semantic clusters, and improve the labeling consistency and correction accuracy. This method not only shares multi-scale features to enhance the representation ability of rare samples and alleviate the long-tail distribution problem of curved shell defect datasets, but also dynamically corrects abnormal labeling through graph-guided popular structure analysis, achieving efficient adaptation to complex noise patterns. Therefore, the graph-guided dual-network nested labeling management method proposed in the present application can adapt to the complexity of industrial curved shell defect datasets, achieve efficient feature representation and depiction of data latent distribution, and effectively improve the overall quality of data labeling, providing a solid data foundation for intelligent manufacturing.

[0140] (3) The method can cope with the challenges specific to the defect data set of the industrial curved shell, not only significantly improving the quality and efficiency of the annotation governance, generating high-level data annotation sets, and enhancing the generalization ability of the model in complex industrial scenarios, but also adapting to the diversified needs of curved shell defect detection, ultimately ensuring product quality and industrial production safety. BRIEF DESCRIPTION OF DRAWINGS

[0141] Figure 1 The figure-guided data annotation governance method for curved shell defect detection of the present application is a flowchart.

[0142] Figure 2 The figure is a schematic diagram of a double-network nested framework.

[0143] Figure 3 The figure is a flowchart of the iterative optimization method for data annotation sets and contrastive learning models. DETAILED DESCRIPTION

[0144] The specific embodiments of the present application are described below to facilitate understanding of the present application by those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and that any variations that are obvious to those skilled in the art within the spirit and scope of the present application as defined and determined by the appended claims are within the scope of protection.

[0145] As shown in Figure 1 In one embodiment of the present application, the figure-guided data annotation governance method for curved shell defect detection includes the following steps:

[0146] S1, collect curved shell defect data and annotate the curved shell defect data;

[0147] S2, based on the annotated curved shell defect data, perform data augmentation on the defect area through mask guidance to generate a data annotation set;

[0148] S3, input the data annotation set into a double-network nested framework to generate a prediction result of soft labels;

[0149] S4, based on the prediction result of soft labels, use a figure-guided active learning strategy to iteratively optimize the data annotation set and the double-network nested framework to complete data annotation governance.

[0150] In the data preparation stage, a simple curved shell defect data collection device was built to simulate the industrial production environment, which included an image acquisition system and a curved shell rotation driving system.

[0151] In S1, the curved shell defect data specifically refers to the curved shell defect image. After acquiring the curved shell defect image, it is preprocessed. The preprocessing method is as follows:

[0152] The image acquisition system utilizes a high-precision line scan camera to perform line scan image acquisition on a curved shell on a rotating roller platform. During the acquisition process, the curved shell rotation drive system controls the rotation speed of the curved shell. Simultaneously, a surface acquisition consistency score ensures a high degree of matching between the shell's rotation speed and the line scan camera's acquisition frequency, thereby acquiring complete and continuous images of defects in the curved shell, eliminating inter-frame redundancy and coverage blind spots, ensuring spatial sampling completeness, and achieving a surface acquisition consistency score. The specific expression is:

[0153]

[0154] In the formula, N The number of image frames acquired. For the first m The shell rotation speed corresponding to the frame image. The target acquisition frequency for the linear scan camera. Maximum rotational speed is used for normalization. Indicates the first m Frame image and the first The region overlap rate of a frame image is used to measure the continuity of image acquisition.

[0155] In the image acquisition system, the contrast and visibility of defect features of the image are improved by adjusting the angle and brightness of the light source, and the curved background is suppressed based on the line scan direction. Using this method, the image is further standardized, and problems such as uneven light reflection and shadows on curved surfaces caused by the line scan camera acquisition process are reduced.

[0156] Among them, the pixel intensity after curved background suppression The specific expression is:

[0157]

[0158] In the formula, Original pixel intensity For the scanning direction weights, The suppression coefficient, For local illumination estimation, the average intensity of the local window based on the scanning direction reflects the pixel. P The distribution of light in the surrounding area, This represents the global average illumination. The global illumination standard deviation is used to quantify the illumination level and variation of the entire image.

[0159] The acquired images undergo scan-direction texture smoothing to suppress periodic textures such as polishing streaks and scan band artifacts in images of curved shell defects acquired by a line scan camera. The pixel intensity after scan-direction texture smoothing is [not specified]. The specific expression is:

[0160]

[0161] In the formula, For neighboring pixels q The original strength, The neighborhood representation along the scanning direction is a one-dimensional window of pixels along the axis of the line scan camera, adapted to the line-by-line acquisition characteristics of line scan. It is an adjustment factor based on the scanning direction distance and illumination intensity, used to dynamically adjust the smoothing intensity, enhance smoothing in periodic texture areas, and preserve details in areas with drastic illumination changes;

[0162]

[0163] In the formula, The axial pixel distance Due to differences in strength, For distance control parameters, These are the control parameters for intensity;

[0164] In this embodiment, using the method described above, the curved shell defect data acquisition device collected 10,000 original images containing various surface defects of curved shells. These images comprehensively reflect the defect morphologies that may appear on the surface under diverse external interference conditions, providing diverse data support for model training.

[0165] In S1, the method for annotating the defect data of the curved shell is as follows: a consensus-driven annotation mechanism is used to perform voting-style annotation on the image, the first... m Annotation results of frame images The specific expression is:

[0166]

[0167] In the formula, The total number of votes. For expert models j For the first m Annotation categories for frame images, Candidate defect categories, For indicator functions, To determine the threshold.

[0168] In this embodiment, in order to effectively construct the labeled data set, the labeling expert team performs systematic preliminary labeling work on the images through the structured vector labeling platform, or establishes an expert model for labeling. These professionally labeled data sets will serve as the data labeling set in the subsequent active learning process of the labeling management algorithm, ensuring the high accuracy and reliability of the algorithm in the labeling optimization and management process.

[0169] In S2, in order to enhance the generalization ability of the model, a large number of redundant single backgrounds exist in the non-defect area of the surface image of the curved shell, which leads to that the global enhancement cannot provide effective information gain. In this embodiment, the random data enhancement is performed on the defect area only through the mask guidance in the image preprocessing stage. Compared with the traditional global enhancement, the mask-guided method can more specifically highlight the defect features. This enhancement method retains the original pixels of the non-defect area while realizing the seamless fusion of the enhanced area and the background in combination with the smoothing mask, so as to make the model focus on the essential features of the defects and enhance the adaptability of the model to different defect morphologies. S2 includes the following steps:

[0170] S21, based on the images in the labeled curved shell defect data, generating a labeling mask through labeling or a bounding box, and separating the defect area and the non-defect area using the labeling mask ;

[0171]

[0172]

[0173] In the formula, I 1 is an image in the labeled curved shell defect data, , H and W represent the height and width of the image, represents a matrix, M 2 is a labeling mask, represents a position is a defect image pixel coordinate, represents a non-defect image, is an element-wise multiplication;

[0174] S22, performing random data enhancement on the defect area, including applying at least one of geometric enhancement, illumination and color enhancement, and texture enhancement. The geometric enhancement includes scaling, rotation, and translation. The illumination and color enhancement includes brightness, contrast, hue, and saturation adjustment. The texture enhancement includes simulating imaging noise and applying Gaussian blur.

[0175] ​where, scaling simulates the change of defect shape, size and position, such as micro-crack magnification or reduction. The scaling process calculates the target pixel value by bilinear interpolation, scaling the pixels within the defect region bounding box:

[0176]

[0177] where, is the data-augmented defect region, and are random scaling factors for coordinates, is the bilinear interpolation function;

[0178] Rotation rotates the pixel coordinates around the center of the defect region, simulating the change of defect direction, such as scratches at different angles. The rotated pixel coordinates are calculated by applying an affine transformation of the rotation matrix to the defect region:

[0179]

[0180] where, is the random rotation angle, is the center coordinate of the defect region;

[0181] Translation simulates the change of defect position, such as defect formation at different positions, by moving the position of the defect region. The translation operation randomly moves the pixels within the region by the offset and updates the mask position:

[0182]

[0183] where, and are random translation factors for coordinates, , the translated defect is limited within the image boundary;

[0184] Illumination and color adjustment simulates different lighting conditions, material backgrounds and color shifts, etc. by modifying the brightness and color characteristics of the defect region. Among them, brightness adjustment is based on the maximum pixel value of the defect region, adding a uniform random brightness offset to each channel of the defect region, simulating the change of light intensity:

[0185]

[0186] where, is the random brightness offset, is the color channel, the pixel value of the defect region after brightness adjustment is clipped to the range [0, 1].

[0187] Contrast adjustment enhances or weakens the distinction between defects and background by scaling the deviation of defect region pixel values relative to their mean:

[0188]

[0189] In the formula, For random contrast factor, N The total number of pixels in the defect area is denoted as , and the adjusted pixel values ​​of the defect area are cropped to the range of [0,1].

[0190] Color dithering independently adjusts hue in the HSV color space. and saturation Simulates color shifts caused by changes in material or lighting:

[0191]

[0192]

[0193] In the formula, , For color space conversion functions, For random hue shift, This is a random saturation factor.

[0194] Texture modification simulates subtle texture variations in defects, such as scratches, graininess, and blurriness. Specifically, independent Gaussian random noise is added to each channel of the defect area to simulate imaging noise.

[0195]

[0196] In the formula, n For random noise values ​​that follow a Gaussian distribution, The noise standard deviation is calculated using the maximum pixel value in the defect area.

[0197] By applying Gaussian blur and performing a weighted average on the pixels in the defect area, the effect of low-resolution imaging or blurred defects can be simulated:

[0198]

[0199] In the formula, This represents the Gaussian kernel function, used for blurring operations. Relative to the current pixel Horizontal and vertical offset coordinates, The standard deviation is the Gaussian blur.

[0200] S23. By using a smoothing mask to blend the defective and non-defective regions after data augmentation, an image of the data-augmented curved shell defect data is generated. This yields the data annotation set;

[0201]

[0202] In the formula, is a smooth mask, is a defect region after data enhancement; in this embodiment, in order to avoid the boundary mutation of the defect region after data enhancement and the original non-defect region, a Gaussian blur is used to generate a smooth mask for image blending, and .

[0203]

[0204] In the formula, is a Gaussian blur function, is a Gaussian blur standard deviation.

[0205] In this embodiment, based on the mask-guided data enhancement method, the selective enhancement of the defect region effectively avoids unnecessary background interference caused by the high consistency of the background and the limited useful information in the industrial curved shell data set. This image enhancement strategy focuses on highlighting and optimizing defect features, providing a robust and high-quality data basis for subsequent self-representation-based graph-guided data annotation management and active learning processes, thereby effectively managing and optimizing the annotation anomalies.

[0206] In S3, the double-network nested framework includes a contrast learning model based on tail-aware decoupling and a multi-scale self-representation learning model connected to each other, and the double-network nested framework is as shown in Figure 2 The contrast learning model includes a shared encoder and a projection head connected to each other, and the shared encoder is connected to the multi-scale self-representation learning model, wherein the shared encoder is a ConvNeXt-tiny network. In order to effectively deal with the data annotation anomalies and consistency problems caused by the uneven illumination reflection of the curved shell due to the curved geometric structure, the contrast learning model first extracts features of the data set through the peripheral ConvNeXt-tiny network to obtain high-quality and multi-scale feature expressions; then, the above features are input into the multi-scale self-representation learning model to further mine and depict the internal structural information of the data, and realize in-depth modeling of the essential attributes of the labeled data; finally, a high-discrimination soft label is generated through the self-representation coefficient matrix to provide high-quality data annotation management and optimization support for downstream tasks.

[0207] The working process of the double-network nested framework is as follows:

[0208] In S31, the data annotation set is input into the contrast learning model based on tail-aware decoupling, and a class-aware negative sampling strategy is used to collect negative samples to generate anchor samples and other samples within a batch;

[0209] S32, input the anchor point sample and other samples into the ConvNeXt-tiny network to generate a multi-scale feature map, input the feature after uniform channel into a multi-scale self-representation learning model to generate a consensus self-representation matrix, obtain an affinity matrix through the consensus self-representation matrix, and generate a prediction result of a soft label by using self-representation spectral probability propagation;

[0210] In S32, after the multi-scale feature map is generated, it is mapped to a low-dimensional feature space through a projection head, and a sample pair is constructed by randomly adding random disturbance in the feature space, which is used to calculate a reweighted decoupling loss function, and a dynamic weight is applied to the tail class sample to decouple the optimization process of the positive and negative sample pairs.

[0211] In the embodiment, the dual-network nested framework adopts a contrast learning model based on tail-aware decoupling, which can not only effectively deal with the inherent problems such as high-dimensional complexity and serious long-tail distribution in the industrial curved shell defect data by virtue of the powerful feature extraction capability of contrast learning, but also can model the internal structure of the data through the self-representation matrix to deeply mine the essential relationship between the data. Therefore, the design of the dual-network cooperation can significantly improve the governance and optimization capability of the industrial curved shell defect data set labeling anomalies, and ensure that the labeling optimization result has high quality and high reliability.

[0212] The contrast learning model based on tail-aware decoupling aims to make full use of the semantic difference between the head class and the tail class through tail class awareness and contrast learning, optimize the representation quality of the tail class in the long-tail distribution data, and improve the discrimination ability of the model to the tail class samples. After training, the model freezes the parameters only as a feature extractor, outputs multi-scale features, and applies them to subsequent self-representation matrix graph construction to achieve efficient feature representation.

[0213] The contrast learning model introduces a class-aware negative sampling strategy and a reweighted decoupling loss function to optimize the representation quality of the tail class in the long-tail data distribution. Traditional contrast learning methods usually use uniform negative sampling and standard loss functions, which are difficult to effectively deal with the sparsity and feature expression deficiency of the tail class samples in the long-tail distribution. The present application dynamically adjusts the negative sample distribution and applies a dynamic weight to the tail class samples to decouple the optimization process of the positive and negative sample pairs, further improving the performance of the model in the long-tail distribution scenario.

[0214] In S31, the method of collecting negative samples by using the class-aware negative sampling strategy is as follows:

[0215] S311, arrange the classes of the data labeling set in descending order according to the number of samples, and take the first 20% of the classes as the head class set H , take the last 20% of the classes as the tail class set T , and take the remaining classes as the intermediate class set M , calculate the frequency weight of each class;

[0216]

[0217] wherein, w c is the inverse frequency weight of the category c , n c is the number of samples of the category c ;

[0218] S312, calculate the category-aware sampling probability:

[0219] In response to the category of the anchor sample belonging to the tail class set, the negative sample is sampled from the head class, the intermediate class and the tail class, and the sampling probability of the negative sample is :

[0220]

[0221] wherein, Z i is the anchor sample, c i is the category of the anchor sample, p m is the probability of sampling from the intermediate class, c k is the category of the negative sample, is the weight of the negative sample category, p h is the probability of sampling from the head class, is the conditional formula, in the above formula, when the anchor sample belongs to the tail class set, the probability of selecting the sample as the negative sample is

[0222]

[0223] In response to the category of the anchor sample belonging to the head class set or the intermediate class set, the negative sample is sampled from all categories, and the sampling probability of the negative sample is :

[0224]

[0225] wherein, C is all categories;

[0226] The above-mentioned category-aware sampling strategy enhances the negative sample selection of the tail class anchor point, and improves the semantic differentiation ability of the tail class and the head class.

[0227] S313, introduce semantic similarity guidance in the process of collecting negative samples, and obtain the final negative sample sampling probability combined with the category-aware sampling probability;

[0228]

[0229] wherein, is a hyper-parameter, is a class-aware sampling probability, is a normalized cosine similarity;

[0230] In the embodiment, the semantic similarity guidance ensures that the negative samples have a certain similarity with the anchor points in the feature space, forcing the contrast learning to further enhance the distinguishing ability of positive and negative samples.

[0231]

[0232] wherein, is the feature of other samples in the batch z j cosine similarity with the anchor sample, is a temperature parameter for controlling the sharpness of similarity;

[0233]

[0234] wherein, is an L2 norm;

[0235] In the embodiment, to further optimize the quality of negative samples, the application introduces semantic similarity guidance in the dynamic negative sampling process, selecting negative samples that are semantically similar to the anchor samples in the feature space.

[0236] S314, the class set of the negative sample is constrained during the collection of the negative sample, so as to satisfy the minimum class coverage and diversity constraint, and the expression satisfying the minimum class coverage is specifically:

[0237]

[0238] wherein, is a floor function, is a diversity control coefficient, n is the number of negative samples sampled for each anchor sample; in the embodiment, to avoid the negative sample set being too concentrated in certain classes, the application uses an intra-batch diversity constraint method in the above sampling process, to ensure that the negative sample set covers a diversified class distribution.

[0239] The expression of the diversity constraint is specifically:

[0240]

[0241] The diversity constraint maximizes the number of classes in the negative sample set, ensures maximum diversity, contains sufficient class information, and enhances the generalization ability of the tail class anchor in the contrast learning.

[0242] In S32, the ConvNeXt-tiny network includes two-dimensional convolutions, normalization layers, an initial level, an intermediate level, a deep level, and a final level connected in sequence, each level includes a plurality of ConvNeXt residual blocks, and the multi-scale feature mapping includes a first feature mapping, a second feature mapping, a third feature mapping, and a fourth feature mapping. The workflow of the ConvNeXt-tiny network is as follows:

[0243] A1, input image based on the ConvNeXt-tiny network I 2, , the input kernel is a two-dimensional convolution with a size of 4x4, a stride of 2, and a channel number of 96, the output of the two-dimensional convolution is input into the normalization layer, and a feature map X 0, ;

[0244] A2, input the feature map X 0 into the initial level, the initial level includes three ConvNeXt residual blocks, and a first feature mapping X 1 is obtained, which is used to capture low-level semantic information such as image edges and textures, ; wherein the workflow of the ConvNeXt residual block is as follows:

[0245] The input feature of the ConvNeXt residual block is subjected to feature space extraction using a depth convolution block with a kernel size of 4x4, a stride of 2, and a channel number of 96, and is normalized by a LayerNorm operation to generate a normalized feature mapping; a 1x1 convolution is used to fuse the inter-channel information of the normalized feature mapping, expand the channel number from 96 to 384, and introduce a smooth nonlinearity using a GELU activation function; a 1x1 convolution is used to restore the channel number from 384 to 96, and the restored feature is combined with the input feature through a residual connection to obtain the output feature of the ConvNeXt residual block; in this embodiment, the working principles of the ConvNeXt residual blocks are the same, so the workflows of the remaining ConvNeXt residual blocks are not described.

[0246] A3, input the first feature mapping X 1 into the intermediate level, adjust the channel number to 192 through a downsampling layer, input the downsampled feature into three ConvNeXt residual blocks, and obtain a second feature mapping X 2, which is used to obtain middle-level semantic information such as image shapes and patterns, ;

[0247] A4, input the second feature mapping X2Input deep level, adjust the number of channels to 384 through the downsampling layer, input the down-sampled features into 9 ConvNeXt residual blocks, third feature map X 3, for obtaining image high-level semantic information, ;

[0248] A5, the third feature map X 3Input final level, adjust the number of channels to 768 through the downsampling layer, input the down-sampled features into 3 ConvNeXt residual blocks, fourth feature map X 4, .

[0249] After the model obtains the feature map through the ConvNeXt-tiny network, the feature map is mapped to a low-dimensional feature space through the projection head and randomly selected to add a feature space random disturbance to construct a positive sample pair for subsequent sample contrast loss calculation.

[0250] In this embodiment, based on the multi-scale self-representation learning model of external feature extraction, through self-representation learning and multi-scale information fusion, the multi-level semantic information in the curved shell defect data set can be fully mined, and the subspace representation ability for industrial high-dimensional complex curved shell defect data can be significantly improved. The model design includes a multi-scale autoencoder, a multi-scale self-representation learning layer, and a self-adaptive attention fusion mechanism, which provides a solid support for subsequent data labeling management and optimization combined with active learning.

[0251] In S32, the workflow of the multi-scale self-representation learning model is as follows:

[0252] B1, unify the channels of the multi-scale feature map to 64 through 1x1 convolution, and perform average pooling to obtain the feature after unifying the channels;

[0253] B2, input the feature after unifying the channels into the corresponding scale of the autoencoder, the autoencoder includes three layers of encoders and decoders, and an unbiased fully connected self-representation learning layer with a size of is embedded between the encoder and the decoder to learn the self-representation coefficient matrix of each scale;

[0254] Wherein, the reconstruction loss is constructed according to the input of the encoder and the output of the decoder, and the model collapse is prevented by minimizing the reconstruction loss, and the expression of the reconstruction loss is as follows:

[0255]

[0256] In the formula, is the input of the l layer encoder, is the output of the l layer decoder,s is the scale;

[0257] An unbiased fully connected self-representation learning layer is embedded in the three-layer encoder and decoder for each scale of refined features to learn the self-representation coefficient matrix, which makes the feature of each sample represented by the linear combination of other samples in the same scale by optimizing the self-representation loss, while applying a regularization constraint to avoid trivial solutions. The expression of the self-representation loss is specifically as follows:

[0258]

[0259] In the formula, is the refined feature, is the balance parameter, is the self-representation coefficient matrix, is the Frobenius norm, and the first term ensures that the linear relationship between samples is captured, and the representation relationship between each data sample is learned. The second term is the Frobenius norm regularization, which forces the self-representation coefficient matrix to have a clear block diagonal effect. The last term avoids trivial solutions by constraining all samples are divided into different class clusters.

[0260] The self-representation matrix of each scale captures the subspace relationship at a specific scale, but the matrices of different scales may contain complementary or redundant information. To generate a unified subspace representation, these matrices are fused through an adaptive attention mechanism to generate a consensus self-representation matrix , which provides a basis for subsequent spectral clustering.

[0261] B3, the self-representation coefficient matrices of each scale are fused to generate a consensus self-representation matrix, and the process is specifically as follows:

[0262] The statistical features of each scale self-representation matrix are extracted through global average pooling to generate global description factors of the self-representation coefficient matrix G .

[0263]

[0264] In the formula, is the self-representation coefficient matrix stacked into a three-order tensor along the channel dimension, is the global average pooling;

[0265] The channel weight is obtained by measuring the contribution degree of each channel using the global description factor U .

[0266]

[0267] wherein, is an unbiased fully connected network to capture the relationship between different scale auto-encoding coefficient matrices, is a ReLu activation function, is a Sigmoid activation function;

[0268] The auto-encoding coefficient matrix is weighted by the channel weight, and a consensus auto-encoding matrix is generated using a 3x3 convolution ;

[0269]

[0270] wherein, is a 3x3 convolution;

[0271] In S32, the method for obtaining the affinity matrix is specifically: generating the affinity matrix W by taking the absolute value average of and the transpose thereof, to ensure the positive definiteness and stability of the spectral decomposition;

[0272]

[0273] wherein, is a transpose symbol;

[0274] The affinity matrix W integrates the subspace information of multi-scale auto-encoding and provides a global view of the relationship between samples. After obtaining the affinity matrix W, feature decomposition is performed thereon to extract the eigenvectors corresponding to the first K eigenvalues, K is the cluster number determined by the preset task. Feature decomposition can reduce the high-dimensional affinity relationship, filter out noise while preserving the main subspace structure, and enhance the robustness to complex patterns in the surface shell defect data. Finally, the eigenvector matrix is obtained by using the eigenvectors obtained by the feature decomposition, the K-means algorithm is applied to each row vector to minimize the intra-cluster variance, optimize the distance from the sample to the cluster center, and assign N samples to K clusters to realize clustering.

[0275] Based on this, the loss function of the multi-scale auto-encoding learning model is divided into three parts: the reconstruction loss ensures that the features generated by the three-layer encoder are reconstructed close to the original input through the decoder, retaining the data information; the auto-encoding loss optimizes the scale auto-encoding matrix and the consensus matrix to promote the block diagonal characteristics of the subspace structure; the structure preserving loss uses the Laplacian matrix to constrain the manifold structure of the features, enhancing the clustering consistency. The model optimizes these losses jointly to ensure the collaborative optimization of multi-scale feature representation and subspace clustering, and the expression of the loss function of the multi-scale auto-encoding learning model is specifically:

[0276]

[0277] wherein, and is a weight coefficient, is a structure preserving loss, in order to make high-dimensional nonlinear multi-scale feature data preserve the structure characteristics inherent in the data, a nonlinear manifold learning algorithm is used to construct the structure preserving loss;

[0278]

[0279] wherein, is the rank of the calculation matrix, is a Laplacian matrix, , is a diagonal matrix degree matrix, the diagonal elements , is an element in the affinity matrix , representing the similarity between samples i and j ;

[0280] In this embodiment, the multi-scale self-representation learning model optimizes multi-level data representation for the complex subspace structure of the industrial curved shell defect data set, fully utilizes multi-scale information from local details to global semantics, and enhances the robustness of the model to noise labels and subspace discrimination ability. The model realizes the deep integration of multi-scale feature collaborative modeling and clustering by adjusting the encoder reconstruction, self-representation learning and manifold structure learning, enhances the correction ability to noise labels, and ensures the accuracy and overall clustering consistency of sample subspace expression.

[0281] In the soft label generation stage, the sample membership probability weight is obtained by using the Softmax probability normalization combined with the similarity between samples in the affinity matrix, the class association of the sample in the multi-scale subspace is decoupled, so as to obtain accurate soft labels. In S32, the method for generating soft labels by using self-representation spectral probability propagation is specifically:

[0282] C1, perform feature decomposition on the affinity matrix , extract the first K feature vectors to form an embedding feature matrix V, representing the projection of the sample in the low-dimensional subspace, and normalize the feature embedding matrix to ensure scale consistency;

[0283]

[0284] wherein, is the first K feature vector;

[0285] C2. Use K-means clustering to calculate the Euclidean distance from the centroid to the samples in the curved shell defect dataset, and assign the samples to... K Each cluster generates a hard tag;

[0286] After obtaining the feature vector matrix, K-means clustering is used to calculate the Euclidean distance from the sample to the centroid, and the samples are assigned to K clusters to generate hard labels. Hard labels provide the basis for category division in the generation of soft labels, ensuring that samples are initially distinguished within the subspace.

[0287]

[0288] In the formula, For the first k The centroid of a cluster, v i For the sample i Embedded vector, For indicator functions, if The value is 1 if it is 1, otherwise it is 0.

[0289] C3. Based on the affinity matrix and hard labels, calculate the sum of similarities between samples and each cluster, and calculate the sample similarity. i and belonging to category k The sum of the similarity of the samples The specific expression is:

[0290]

[0291] In the formula, For hard label k The sample set;

[0292] C4. Generate samples using Softmax transformation based on the sum of similarities between sample pairs and each cluster. i Prediction results of soft tags P i The non-linear characteristics of Softmax amplify the probability of dominant clusters and compress the probability of weakly associated clusters, ensuring that soft labels accurately capture structural information. , P ik For the sample i Category k The probability of a cluster satisfies ;

[0293]

[0294] In the formula, For scaling ;

[0295]

[0296] wherein, is a scaling factor to prevent Softmax overflow.

[0297] The self-representation spectrum probability propagation soft label generation method in the embodiment aims to deeply mine multi-scale structural information in the consensus self-representation matrix, convert the implicit class relationship into a high-distinguishable soft label, and provide a reliable reference basis for data labeling management and optimization.

[0298] On the basis of the category-aware negative sampling strategy, the application designs a re-weighted decoupled loss function based on the loss function of contrastive learning, and applies a dynamic weight to tail class samples. Specifically, by decoupling the contributions of positive and negative samples, the loss function is divided into two parts of positive sample attraction and negative sample repulsion, avoiding the complex influence of the number of negative samples on the gradient. By assigning a higher weight to the tail class samples, the loss contribution of the positive and negative samples of the tail class is amplified, and the feature representation in the long-tail distribution scenario is optimized.

[0299] In S32, the positive samples are used to calculate the re-weighted decoupled loss function, and the expression of the re-weighted decoupled loss function is specifically as follows:

[0300]

[0301] wherein, N is the number of samples in the batch, is a negative sample loss weight used to balance the positive and negative loss contributions, is a positive sample feature, is a negative sample feature, is a positive sample loss, is a negative sample loss;

[0302] The contrastive learning model based on tail-aware decoupling in the embodiment optimizes the tail class feature representation for long-tail distribution data, fully utilizes the class distribution and semantic similarity information, improves the discrimination ability and generalization performance of the model, and thus enhances the performance of the model in the unbalanced data scenario. Therefore, the model relies on the category-aware negative sample selection mechanism and the joint loss function of decoupled weighting in the training stage to cooperatively optimize the tail class learning efficiency and overall prediction accuracy.

[0303]

[0304]

[0305] wherein, is a cosine similarity, which measures the similarity between features, is a temperature parameter used to control the sharpness of the similarity;

[0306] In the embodiment, the positive sample loss in the reweighted decoupled loss function encourages the anchor sample and its augmented view (positive sample pair) to be close in the feature space, and the similarity of the positive sample pair is maximized by maximizing the cosine similarity. By applying a higher weight to the tail class anchor, the attractive effect of the tail class positive sample is amplified, and the cohesion of the tail class feature is enhanced.

[0307]

[0308] In the formula, N k The negative sample set is.

[0309] In the embodiment, the negative sample loss pushes the anchor sample away from the negative sample feature, and the distinction between the anchor and the negative sample is enhanced by calculating the softplus loss for each negative sample. By applying a higher weight to the tail class anchor, the repulsive effect of the tail class negative sample is amplified, and the discrimination ability of the tail class feature is optimized.

[0310] In S3, the training method of the double network nested framework is specifically:

[0311] Training the contrast learning model:

[0312] The training set is input into the contrast learning model, and the reweighted decoupled loss is applied for optimization. The optimizer is configured as SGD, the initial learning rate is 1e-3, and the weight decay is 5e-4. In order to avoid early shock, the training process is combined with a cosine annealing scheduler and a 5epoch linear warm-up to accelerate model convergence. The model is trained for 100 epochs, and in response to the fact that the validation loss fluctuation is less than 0.1% for 10 consecutive epochs, the network parameters are frozen, the ConvNeXt-tiny network is fixed as a multi-scale feature extractor, and a multi-scale feature mapping is output.

[0313] Training the multi-scale self-representation learning model:

[0314] The multi-scale feature mapping is input, and the self-representation matrix is optimized by the loss function of the multi-scale self-representation learning model. The reconstruction loss constrains the consistency between the input image and the decoding output, the self-representation loss optimizes the adjacency relationship, and the structure preservation loss forces the semantic structure of the multi-scale feature to remain uniform. The optimizer used in the training process is AdamW, the initial learning rate is 5e-4, the model is trained for 100 epochs, and the termination condition is set to be that the loss decrease rate is less than 0.01% for 15 consecutive epochs. FP16 mixed precision is used in the training process, and an early stopping strategy is used to ensure model convergence.

[0315] As shown in Figure 3 In the embodiment, in S4, the data annotation set and the double network nested framework iterative optimization method are specifically:

[0316] S41, based on the prediction result and uncertainty of the soft label of the sample, the sample membership degree and sample entropy value are calculated, the soft label of the sample is labeled and optimized, the labeled corrected label and pseudo label are generated, and the labeling optimization method is specifically as follows:

[0317] In response to the sample membership degree being greater than or equal to 0.9 or the sample entropy value being less than or equal to 0.1, the soft label of the sample is taken as the pseudo label;

[0318] In response to the sample membership degree being less than 0.9 and greater than or equal to 0.7 or the sample entropy value being less than or equal to 0.3 and greater than 0.1, it is judged whether the soft label of the sample is consistent with the original label, if yes, the soft label of the sample is taken as the labeled corrected label, and if no, the labeled corrected label is generated through the auditing link;

[0319] In response to the sample membership degree being less than 0.7 or the sample entropy value being greater than 0.3, the labeled corrected label is generated through the auditing link;

[0320] The method for generating the labeled corrected label through the auditing link is specifically as follows:

[0321] The soft label of the sample is voted by the expert model, and the soft label with the effective vote number exceeding the preset threshold is taken as the labeled corrected label;

[0322] In the embodiment, the expert model will combine the image content such as defect texture, shape and P i probability distribution information to comprehensively judge the category and generate the labeled corrected label, so as to improve the labeling quality and consistency of the data set. The auditing link follows the multiple voting strategy during the original data labeling, and the labels of the potential abnormal labeling and high uncertainty samples are independently voted and audited by multiple expert models. The expert models independently give the judgment opinions based on the flexible (soft) labels generated by the model training. After collecting the voting results, it is judged whether the effective vote number exceeds the preset half threshold: if yes, it indicates that the defect category judgment reaches the expert consensus, and the obtained label is taken as the final labeled corrected label; if no, the voting is reorganized until enough effective vote numbers are obtained to ensure the accuracy and consistency of the labeling result.

[0323] S42, updating the data labeling set according to the sample after the labeling optimization, and retraining the double network nested framework according to the data labeling set, wherein the optimizer, momentum parameter and weight decay coefficient remain unchanged during the training process;

[0324] S43, repeating S41 to S42 to iteratively optimize the data labeling set and the double network nested framework, after each round of iterative optimization is completed, the comprehensive performance indicators of the contrast learning model on the verification set and the data labeling set are verified and compared, and in response to the performance improvement amplitude of the continuous two iterations being less than 0.5%, the iterative optimization of the data labeling set and the double network nested framework is completed.

[0325] After the annotation correction is completed, the network architecture, parameters and regularization strategies preset in the initial training stage are strictly reused to restart the model training process using the managed updated data set. The optimizer, momentum parameters and weight decay coefficient remain unchanged during the training process. After each round of iterative training is completed, the comprehensive performance indicators of the model on the clean validation set and the managed optimized data set need to be verified simultaneously. When the performance improvement amplitude of two consecutive iterations is less than 0.5%, it is determined that the model reaches the optimal state and the iteration is terminated. At this time, the annotation abnormality proportion of the data set of the double-network nested framework has converged to a stable state.

[0326] In the description of the present application, it should be understood that the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features. Therefore, the features limited by "first", "second", "third" can explicitly or implicitly include one or more of the features.

Claims

1. A method for graph-guided data labeling governance for curved shell defect detection, characterized in that, The method comprises the following steps: S1, collecting curved shell defect data, and labeling the curved shell defect data; S2, based on the labeled curved shell defect data, data augmentation is performed on the defect area through mask guidance to generate a data annotation set; S3, inputting the data annotation set into a double-network nested framework to generate a prediction result of a soft label; S4, based on the prediction result of the soft label, a graph-guided active learning strategy is used to iteratively optimize the data annotation set and the double-network nested framework to complete data annotation management; In S3, the double-network nested framework comprises a contrast learning model based on tail perception decoupling and a multi-scale self-representation learning model connected with each other, the contrast learning model comprises a shared encoder and a projection head connected with each other, and the shared encoder is connected with the multi-scale self-representation learning model, wherein the shared encoder is a ConvNeXt-tiny network, and a working process of the double-network nested framework is specifically as follows: In S31, the data annotation set is inputted into the contrast learning model based on tail perception decoupling, a class-aware negative sampling strategy is used to collect negative samples, and anchor samples and other samples in a batch are generated; In S32, the anchor samples and the other samples are inputted into the ConvNeXt-tiny network to generate multi-scale feature mappings, the features after being unified in a channel are inputted into the multi-scale self-representation learning model to generate a consensus self-representation matrix, an affinity matrix is obtained through the consensus self-representation matrix, and a prediction result of a soft label is generated through self-representation spectral probability propagation; In S32, after the multi-scale feature mappings are generated, they are mapped to a low-dimensional feature space through the projection head, and sample pairs are constructed by randomly adding random disturbances in the feature space, which are used to calculate a reweighted decoupling loss function, and a dynamic weight is applied to the tail class samples to decouple the optimization process of the positive and negative sample pairs.

2. The graph-guided data annotation governance method for curved shell defect detection according to claim 1, wherein, In S1, the curved shell defect data is specifically a curved shell defect image, and after the curved shell defect image is collected, it is preprocessed, and a specific method of the preprocessing is as follows: The rotation speed of the curved shell is controlled during the collection process, and the curved shell rotation speed is highly matched with the collection frequency of the linear array camera through the curved surface collection consistency score, so that the curved surface shell defect image is collected completely and continuously, and the curved surface collection consistency score The expression of the curved surface collection consistency score is specifically: In the formula, N The number of image frames acquired. For the first m The shell rotation speed corresponding to the frame image. The target acquisition frequency for the linear scan camera. Maximum rotational speed is used for normalization. Indicates the first m Frame image and the first The region overlap rate of a frame image is used to measure the continuity of image acquisition. In the image acquisition system, the contrast of the image and the visibility of the defect features are improved by adjusting the angle and brightness of the light source, and the curved surface background suppression is performed based on the line scanning direction, and the expression of the pixel intensity after the curved surface background suppression is specifically as follows: The expression is specifically as follows: wherein, is the original pixel intensity, is the scan direction weight, is the suppression coefficient, is the local illumination estimate, is the global illumination mean, is the global illumination standard deviation; The scanned image is subjected to scanning direction texture smoothing processing, and the pixel intensity after scanning direction texture smoothing processing is The expression is specifically: wherein, is the original intensity of the neighborhood pixel q , is a scan direction neighborhood representing a one-dimensional window of pixel sets along the camera axis of the line scan, is an adjustment factor based on the scan direction distance and the intensity of the illumination; wherein is an axial pixel distance, is an intensity difference, is a control parameter for the distance, is a control parameter for the intensity; In S1, the method for labeling the curved shell defect data is specifically as follows: a consensus-driven labeling mechanism is adopted to vote for the image, and the first m The labeling result of the frame image The expression of the formula is specifically as follows: In the formula, is the total number of votes, is the expert model j to the first m labeled class of the frame image, is the candidate defect class, is an indicator function, is a judgment threshold.

3. The graph-guided data annotation governing method for curved shell defect detection according to claim 1, wherein, S2 comprises the following steps: S21, based on the image in the labeled curved shell defect data, generate a labeled mask through labeling or a bounding box, and separate the defect area using the labeled mask and the non-defect area ; wherein I 1 is the image in the annotated curved shell defect data, M 2 is the annotation mask, is an element-wise multiplication; In S22, random data augmentation is performed on the defect area, including at least one of geometric enhancement, illumination and color enhancement, and texture enhancement, the geometric enhancement includes scaling, rotation and translation, the illumination and color enhancement includes brightness, contrast, hue and saturation adjustment, and the texture enhancement includes simulating imaging noise and applying Gaussian blur; S23, image blending is performed on the defect area and the non-defect area after data enhancement through a smoothing mask to generate an image of the data-enhanced curved shell defect data , to obtain a data annotation set; In the formula, is a smoothing mask, is a defect region after data enhancement; wherein is a Gaussian blur function, is a Gaussian blur standard deviation.

4. The graph-guided data annotation governance method for curved shell defect detection according to claim 1, wherein, In S31, a specific method of collecting negative samples by using the class-aware negative sampling strategy is as follows: S311. Sort the categories of the data label set in descending order by the number of samples, and take the top 20% of categories as the head class set. H The last 20% of categories are used as the tail category set. T The remaining categories serve as an intermediate class set. M Calculate the frequency weight for each category; wherein w c is the inverse frequency weight of the class c , n c is the number of samples of the class c . In S312, a class-aware sampling probability is calculated: In response to the category of the anchor sample belonging to the tail class set, the negative sample is sampled from the head class, the middle class and the tail class, and the sampling probability of the tail class is is: where Z i is an anchor sample, c i is a class of anchor samples, p m is a probability of sampling from an intermediate class, c k is a class of negative samples, is a weight of a negative sample class, p h is a probability of sampling from a head class, is a conditional formula; In response to the category of the anchor sample belonging to the head class set or the middle class set, the negative sample is sampled from all categories, and the sampling probability of each category is is: In the formulae, C for all classes; S313, introducing semantic similarity guidance in the process of collecting negative samples, combining category-aware sampling probability to obtain the final negative sample sampling probability ; wherein is a hyperparameter, is a class-aware sampling probability, is a normalized cosine similarity; wherein is the feature of other samples in the batch z j cosine similarity to anchor samples, is a temperature parameter used to control the sharpness of the similarity wherein is the L2 norm; In S314, the class set of the negative samples is constrained during the collection of the negative samples, so that the minimum class coverage and diversity constraints are met, and a specific expression of the minimum class coverage is as follows: wherein is a floor function, is a diversity control coefficient, n is the number of sampled negative samples for each anchor sample; A specific expression of the diversity constraint is as follows: 。 5. The graph-guided data annotation governance method for curved shell defect detection according to claim 1, wherein, In S32, the ConvNeXt-tiny network comprises a two-dimensional convolution, a normalization layer, an initial level, an intermediate level, a deep level and a final level connected in sequence, each level comprises a plurality of ConvNeXt residual blocks, the multi-scale feature mappings comprise a first feature mapping, a second feature mapping, a third feature mapping and a fourth feature mapping, and a working process of the ConvNeXt-tiny network is specifically as follows: A1, input image based on a ConvNeXt-tiny network I 2, a two-dimensional convolution with an input kernel of 4x4, a stride of 2, and a channel number of 96, the output of the two-dimensional convolution being input to a normalization layer to obtain a feature map X 0; A2, the feature map X 0input initial level, the initial level includes 3 ConvNeXt residual blocks, to obtain the first feature mapping X 1, wherein the workflow of the ConvNeXt residual block is specifically: The input features of the ConvNeXt residual block are subjected to feature space extraction using a depth convolution block with a kernel of 4x4, a stride of 2, and a channel number of 96, and are normalized by a LayerNorm operation to generate normalized feature maps; a 1x1 convolution is used to fuse the inter-channel information of the normalized feature maps, expand the channel number from 96 to 384, and introduce a smooth nonlinearity using a GELU activation function; a 1x1 convolution is used to restore the channel number from 384 to 96, and the restored features are combined with the input features through a residual connection to obtain the output features of the ConvNeXt residual block; A3, the first feature mapping X 1 input intermediate level, adjust the number of channels to 192 through the down-sampling layer, the down-sampled features input 3 ConvNeXt residual blocks, the second feature mapping X 2; A4, the second feature map X 2 input deep level, adjust the number of channels to 384 through down-sampling layer, the down-sampled feature input 9 ConvNeXt residual blocks, the third feature map X 3; A5, the third feature map X 3. The input final layer adjusts the number of channels to 768 through the down-sampling layer, and the down-sampled features are input into three ConvNeXt residual blocks, and the fourth feature map X 4.

6. The graph-guided data labeling governance method for curved shell defect detection according to claim 5, wherein, In S32, the workflow of the multi-scale self-representation learning model is as follows: B1, the multi-scale feature maps are subjected to 1x1 convolution to unify the channels of the multi-scale features to 64, and average pooling is performed to obtain the features after channel unification; B2, the features after channel unification are input into the corresponding scale autoencoder, which includes three layers of encoder and decoder, and an unbiased fully connected self-representation learning layer is embedded between the encoder and the decoder to learn the self-representation coefficient matrix of each scale; Wherein, the reconstruction loss is constructed according to the input of the encoder and the output of the decoder, the model collapse is prevented by minimizing the reconstruction loss, and the expression of the reconstruction loss is specifically: The expression of the reconstruction loss is specifically: wherein is l input to the layer encoder, is l output from the layer decoder, s is a scale; An unbiased fully connected self-representation learning layer is embedded for the refined features of each scale between the three-layer encoder and decoder, and a self-representation coefficient matrix is learned, the unbiased fully connected self-representation learning layer makes the features of each sample represented by the linear combination of other samples of the same scale through an optimized self-representation loss, and a regularization constraint is applied to avoid trivial solutions, the optimized self-representation loss The expression is specifically: wherein is a refining feature, is a balancing parameter, is a self-representation coefficient matrix; B3, the self-representation coefficient matrices of each scale are fused to generate a consensus self-representation matrix, and the process is as follows: The statistical features of the self-representation matrix of each scale are extracted by global average pooling to generate a global description factor of the self-representation coefficient matrix G ; wherein is the self-representation coefficient matrix stacked into a third-order tensor along the channel dimension, is the global average pooling. The channel weight is obtained by measuring the contribution degree of each channel by using a global description factor U ; wherein is a bias-free fully connected network, is a ReLu activation function, is a Sigmoid activation function; The self-representation coefficient matrix is weighted by each channel weight, and a consensus self-representation matrix is generated using a 3x3 convolution ; In the formulae, is a 3x3 convolution; In S32, the method for obtaining the affinity matrix is specifically: generating the affinity matrix W by taking the absolute value average of the affinity matrix and its transpose. and the transpose thereof. In the formulae, is a transposed symbol; The expression of the loss function of the multi-scale self-representation learning model is optimized as follows: wherein and are weight coefficients, is a structural preservation loss; wherein is the rank of the matrix, is the Laplacian matrix, , is a diagonal matrix with diagonal elements , is the affinity matrix in the set In S32, the method of generating soft labels using self-representation spectral probability propagation is as follows: C1, to the affinity matrix characteristic decomposition, extract the front K characteristic vector, constitute the embedding feature matrix ; In the formula, is the previous K feature vector; C2, using K-means clustering to calculate the Euclidean distance of samples in the curved shell defect data set to the centroid, assign the samples to K clusters, and generate hard labels; C3, based on the affinity matrix and the hard labels, calculate the sum of similarities of the sample to each cluster, calculate the sum of similarities of the sample to the samples belonging to the category i k The expression of the sum of similarities of the sample to the samples belonging to the category​​ In the formula, For hard label k The sample set; C4. The Softmax transformation is applied to the sum of similarities of each cluster to the sample to generate a soft label for the sample i P i , , P ik i k ;​​​​ In the formula, scaled by ; In the formula, is a scaling factor to prevent Softmax overflow.

7. The graph-guided data labeling governance method for curved shell defect detection according to claim 6, wherein, In S32, the positive sample pair is used to calculate a reweighted decoupling loss function, and the expression of the reweighted decoupling loss function is specifically as follows: The expression of the reweighted decoupling loss function is specifically as follows: wherein N is the number of samples in the batch, is the negative sample loss weight used to balance the positive and negative loss contributions, is the anchor sample feature output by the projection head, is the positive sample pair feature, is the negative sample feature, is the positive sample loss, is the negative sample loss; In the formula, To calculate the cosine similarity, measure the similarity between features, Temperature parameter for controlling the sharpness of the similarity. In the formula, N k is the negative sample set.

8. The method of claim 7, wherein the method further comprises: In S3, the training method of the double network nested framework is as follows: Train the contrast learning model: The training set is input into the contrast learning model, and a reweighted decoupled loss function is applied for optimization. The optimizer is configured as SGD, the initial learning rate is 1e-3, and the weight decay is 5e-4. To avoid early oscillation, the training process is combined with a cosine annealing scheduler and a 5epoch linear warm-up to accelerate model convergence. The model is trained for 100 epochs, and in response to a continuous 10epoch validation loss fluctuation of less than 0.1%, the network parameters are frozen, the ConvNeXt-tiny network is fixed as a multi-scale feature extractor, and the multi-scale feature maps are output. Train the multi-scale self-representation learning model: The multi-scale feature maps are input into the multi-scale self-representation learning model, and the loss function of the model is used to optimize the self-representation matrix. The reconstruction loss constrains the consistency of the input image and the decoding output, the self-representation loss optimizes the adjacency relationship, and the structure preservation loss forces the semantic structure of the multi-scale features to be consistent. The optimizer used in the training process is AdamW, the initial learning rate is 5e-4, the model is trained for 100 epochs, and the termination condition is set to a continuous 15epoch loss decrease rate of less than 0.01%. FP16 mixed precision is used in the training process, and an early stopping strategy is adopted to ensure model convergence.

9. The graph-guided data labeling governance method for curved shell defect detection according to claim 8, wherein, In S4, the graph-guided active learning strategy and the iterative optimization method of the double network nested framework are as follows: In S41, based on the prediction results and uncertainty of the soft labels of the samples, the sample membership and sample entropy values are calculated, the soft labels of the samples are labeled and optimized, the labeled corrected labels and pseudo labels are generated, and the labeling optimization method is as follows: In response to the sample membership being greater than or equal to 0.9 or the sample entropy value being less than or equal to 0.1, the soft label of the sample is taken as a pseudo label; In response to the sample membership being less than 0.9 and greater than or equal to 0.7 or the sample entropy value being less than or equal to 0.3 and greater than 0.1, it is judged whether the soft label of the sample is consistent with the original label, if yes, the soft label of the sample is taken as a label correction label, and if no, a label correction label is generated through an auditing link; In response to the sample membership being less than 0.7 or the sample entropy value being greater than 0.3, a label correction label is generated through the auditing link; The method for generating the label correction label through the auditing link is specifically as follows: The soft labels of the sample are voted by the expert model, and the soft label with an effective vote number exceeding a preset threshold is taken as a label correction label; S42, updating the data labeling set according to the sample after the labeling optimization, and retraining the double-network nested framework according to the data labeling set, wherein the optimizer, momentum parameter and weight decay coefficient remain unchanged during the training process; S43, iteratively optimizing the data labeling set and the double-network nested framework repeatedly, and after each round of iterative optimization is completed, the comprehensive performance index of the contrast learning model on the validation set and the data labeling set is verified and compared, and in response to the performance improvement amplitude of two consecutive iterations being less than 0.5%, the iterative optimization of the data labeling set and the double-network nested framework is completed.

Citation Information

Patent Citations

  • Network image label denoising method combining sample selection and label correction

    CN115564960A

  • Label correction crowdsourcing result convergence method and system based on deep clustering

    CN120030368A

  • Rolling metal surface defect automatic labeling method based on multi-task self-adaptive model

    CN119444759A

  • Target detection method for patrol inspection of broken strand conductor of power transmission line by unmanned aerial vehicle

    CN120544084A