Training method and system of image pre-training model and storage medium

By generating multi-resolution feature volumes and constructing positive and negative contrast images of flawed and qualified images, this method addresses the shortcomings of existing image pre-training methods in terms of accuracy and robustness in complex image processing, achieving efficient image classification and recognition, and improving the recognition accuracy of flawed images and the adaptability of the model.

CN120913029AActive Publication Date: 2025-11-07SHENZHEN WRITER INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510865699.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-07
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing image pre-training methods lack accuracy and robustness when processing complex images, and lack effective judgment of image texture clarity, leading to misjudgment and omission in the screening process of defective and qualified images. The models are poorly adaptable to subtle changes, making it difficult to achieve efficient image classification and recognition.

Method used

By acquiring original training images, extracting semantic information and generating high-dimensional feature data, constructing multi-resolution feature volumes, fusing texture clarity for image segmentation, constructing positive and negative contrast images of defective and qualified images for comparative training, and combining defective and qualified images for joint training to generate an integrated image pre-trained model.

Benefits of technology

It improves the representational ability of image features, enhances the recognition accuracy of flawed images and the generalization ability of the model, strengthens the model's ability to learn flawed features, optimizes the quality of the training dataset, ensures the reliability and robustness of the model, and enables more accurate identification and classification of different types of image features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913029A_ABST
    Figure CN120913029A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image pre-training, in particular to a training method and system of an image pre-training model and a storage medium. The method comprises the following steps: collecting an original training image, carrying out high-dimensional feature mapping and pyramid transformation to generate a multi-resolution feature body, carrying out texture definition judgment on the image based on the feature body, screening out a defective training image and a qualified training image, and for the defective training image, carrying out texture definition judgment on the qualified training image. A traceability correction technology is adopted to correct a defect correction training image, positive and negative contrast images are constructed to train a defect image pre-training model, meanwhile, weak disturbance transformation is performed on a qualified training image, a positive example pair before and after disturbance is generated to train a qualified image pre-training model, the two pre-training models are integrated into a collaborative sample stream, and the collaborative sample stream is extracted. And a high-performance integrated image pre-training model is obtained through integrated pre-training. According to the method, image quality grading and differential modeling are realized, the defect identification precision is improved, and the model generalization ability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image pre-training technology, and in particular to a training method, system and storage medium for an image pre-training model. Background Technology

[0002] Existing image pre-training methods often rely on simple feature extraction and classification algorithms, resulting in insufficient accuracy and robustness of the models when processing complex images. Current technologies lack a systematic approach to training data selection and processing, making it difficult to guarantee the diversity and representativeness of training samples, further impacting the model's generalization ability. This is particularly true in the identification and processing of flawed images, where insufficient data and inaccurate feature extraction are common problems. Furthermore, existing models lack effective judgment of image texture clarity, leading to misclassification and missed classification during the screening process between flawed and qualified images, affecting the quality of subsequent training. Especially in complex environments, the models exhibit poor adaptability to subtle changes, making efficient image classification and recognition difficult, further increasing the challenges in practical applications, particularly in industrial inspection and automated processing. Summary of the Invention

[0003] Therefore, it is necessary to provide a training method, system, and storage medium for an image pre-training model to solve at least one of the aforementioned technical problems.

[0004] To achieve the above objectives, a training method for an image pre-training model includes the following steps:

[0005] Step S1: After acquiring the original training images, extract the semantic information from the images and perform feature processing to generate structured high-dimensional feature data; reduce the high-dimensional feature data in three different proportions to finally generate a multi-resolution feature body;

[0006] Step S2: Analyze the texture sharpness of the image by fusing feature volumes at various resolutions, and divide the original image into flawed training images and qualified training images based on texture sharpness.

[0007] Step S3: Tracing and correcting the defective training image to become a defect-corrected training image, and constructing positive and negative contrast images based on the defect-corrected training image and the defective training image; performing contrast training by combining the positive and negative contrast images to construct a pre-trained model of the defective image.

[0008] Step S4: Perform weak perturbation transformation on qualified training images, and pair the images before and after the transformation one by one to construct positive pairs before and after perturbation; find regions with consistent features from the positive pairs before and after perturbation, and use these region images to train the pre-trained model of qualified images.

[0009] Step S5: the feature data sets of the defect image pre-training model and the qualified image pre-training model are integrated into a collaborative sample flow, and joint training is performed through the sample flow, so as to obtain an integrated image pre-training model.

[0010] The present application improves the representation ability of image features through high-dimensional feature mapping and pyramid transformation, so that the multi-resolution feature body better captures subtle texture changes, and the accuracy of texture clarity determination is significantly improved. The effective screening of defect training images and qualified training images optimizes the quality of the training data set, ensuring the reliability of subsequent model training. The traceability correction technology of the defect image ensures the accuracy and consistency of the training data, providing a real and representative sample for the model. The positive and negative contrast images provide rich contrast information for the defect image pre-training model, enhancing the model's learning ability of defect features. The weak disturbance transformation technology makes the qualified training images more diverse, improving the model's generalization ability and robustness. The construction of positive pairs further improves the model's adaptability to image changes. The integration of the collaborative sample flow effectively combines the defect images and the qualified images, and the final integrated image pre-training model performs better in handling complex image tasks, can more accurately recognize and classify different types of image features, realizes image quality grading and difference modeling, improves defect recognition accuracy, and enhances the model's generalization ability.

[0011] The present application also provides a training system for an image pre-training model, which is used to execute the training method of the image pre-training model as described above, and the training system for the image pre-training model comprises:

[0012] An image acquisition module is configured to extract semantic information from the image and perform feature processing after acquiring the original training image, to generate structured high-dimensional feature data; and the high-dimensional feature data is sequentially reduced by three different ratios, and finally a multi-resolution feature body is generated.

[0013] A texture screening module is configured to analyze the texture clarity of the image by fusing the feature bodies of different resolutions, and to divide the original image into defect training images and qualified training images according to the texture clarity.

[0014] A defect modeling module is configured to trace and correct the defect training images to defect correction training images, and to construct positive and negative contrast images based on the defect correction training images and the defect training images; and to construct a pre-training model of the defect image by contrast training combined with the positive and negative contrast images.

[0015] A positive example modeling module is configured to perform weak disturbance transformation based on the qualified training images, and to pair the images before and after transformation one by one to construct positive pairs before and after disturbance; to find out the areas with consistent features from the positive pairs before and after disturbance, and to use these area images to train a pre-training model of the qualified image.

[0016] The integrated training module is used for integrating the feature data of the defect image pre-training model and the qualified image pre-training model into a cooperative sample flow, and jointly training through the sample flow, so as to obtain an integrated image pre-training model.

[0017] The application can comprehensively obtain original training images, ensure the diversity and richness of data, improve the extraction accuracy of image features through feature dimension mapping and pyramid transformation, effectively capture subtle texture information in images through the generated multi-resolution feature body, accurately distinguish defect training images and qualified training images based on the clarity judgment of the texture screening module, optimize the quality of subsequent training data sets, improve the reliability of model training, ensure the authenticity of training data through the traceability correction technology of the defect modeling module, provide representative samples for the model, enhance the learning ability of the defect image pre-training model through the construction of positive and negative contrast images, make it more effectively identify and process defect features, the positive example modeling module expands the sample diversity of qualified training images through weak disturbance transformation, improves the generalization ability and robustness of the model, the integrated training module effectively combines the pre-training models of defect images and qualified images, realizes cooperative training, and finally obtains an integrated image pre-training model that exhibits superior performance in processing complex image tasks, can accurately identify and classify various types of image features, and greatly enhances the practicality and adaptability of the system.

[0018] The application also provides a computer readable storage medium storing a computer program, which, when executed, implements the training method of the image pre-training model according to any one of the preceding embodiments.

[0019] The application can efficiently store and manage the training program of the image pre-training model through the application of the computer readable storage medium, ensure the execution efficiency and accuracy of the program, provide stable data access and processing capability, support large-scale image data processing and analysis, promote the rapid integration of diversified training samples, optimize the image feature extraction and processing process, improve the flexibility and scalability of model training, ensure the integrity and consistency of training data, enhance the adaptability and robustness of the model, and finally realize a high-performance image pre-training model to meet the needs of complex image analysis tasks. BRIEF DESCRIPTION OF DRAWINGS

[0020] Fig. 1 It is a step flow schematic diagram of the training method of the image pre-training model;

[0021] Fig. 2 It is a detailed implementation step flow schematic diagram of step S2;

[0022] The objectives, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0023] The technical method of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0024] In addition, the accompanying drawings are only schematic illustrations of the present application, and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0025] It should be understood that although the terms "first", "second" and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element can be referred to as a second element, and similarly a second element can be referred to as a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0026] To achieve the above-mentioned purpose, please refer to Figs. 1-2 A training method of an image pre-training model, comprising the following steps:

[0027] Step S1: After collecting the original training images, the semantic information in the images is extracted and feature processed to generate structured high-dimensional feature data; the high-dimensional feature data is sequentially reduced by three different ratios, and finally a multi-resolution feature body is generated;

[0028] Step S2: The texture clarity of the image is analyzed by fusing the feature bodies of each resolution, and the original image is divided into flaw training images and qualified training images according to the texture clarity;

[0029] Step S3: The flaw training images are traced and corrected to flaw correction training images, and positive-negative contrast images are constructed based on the flaw correction training images and the flaw training images; comparison training is carried out combined with the positive-negative contrast images, so as to construct a pre-training model of flaw images;

[0030] Step S4: weak disturbance transformation is performed based on qualified training images, and the images before and after transformation are paired one by one to construct positive example pairs before and after disturbance; regions with consistent features are found from the positive example pairs before and after disturbance, and the images of these regions are used to train the pre-training model of qualified images;

[0031] Step S5: the feature data sets of the defect image pre-training model and the qualified image pre-training model are integrated into a cooperative sample flow, and joint training is performed through the sample flow, so as to obtain an integrated image pre-training model.

[0032] The present application improves the representation ability of image features through high-dimensional feature mapping and pyramid transformation, so that the multi-resolution feature body better captures subtle texture changes, and the accuracy of texture clarity determination is significantly improved. The effective screening of defect training images and qualified training images optimizes the quality of the training data set and ensures the reliability of subsequent model training. The traceability correction technology of defect images ensures the accuracy and consistency of the training data, provides real and representative samples for the model, and provides rich contrast information for the defect image pre-training model through the constructed positive and negative contrast images, enhances the learning ability of the model to defect features, and the weak disturbance transformation technology makes the qualified training images have higher diversity, improves the generalization ability and robustness of the model, the construction of positive example pairs further improves the adaptability of the model to image changes, and the integration of cooperative sample flow realizes the effective combination of defect images and qualified images. The integrated image pre-training model finally obtained has stronger performance in processing complex image tasks, can more accurately recognize and classify different types of image features, realizes image quality grading and difference modeling, improves defect recognition precision, and enhances model generalization ability.

[0033] In the embodiment of the present application, the training method of the image pre-training model comprises the following steps:

[0034] Step S1: after collecting original training images, extract semantic information in the images and perform feature processing to generate structured high-dimensional feature data; the high-dimensional feature data is sequentially reduced by three different ratios, and finally a multi-resolution feature body is generated;

[0035] In this embodiment, the target workpiece images under standard lighting conditions are collected by an industrial camera, each image has a resolution of 2048x2048 pixels, the shooting process maintains a fixed exposure value, a fixed shooting distance, and no background clutter interference, after the collection is completed, the images are uniformly formatted and converted into RGB three-channel data format, and the images are input into a feature mapping network based on a multi-layer residual structure. The network structure adopted is 4 layers of residual convolution blocks, each layer has a convolution kernel size of 3x3, a step of 1, uses a ReLU activation function, and outputs a high-dimensional feature mapping tensor with a size of 2048x2048x64. The feature tensor is input into a multi-scale pyramid construction module as input, and the construction module sequentially performs a pooling layer and a variable size convolution operation on the feature map. The pooling layer uses a maximum pooling method, and the scales are set to the original resolution, 1 / 2, 1 / 4, and 1 / 8, a total of four layers. After performing a multi-channel convolution operation with a convolution kernel size of 3x3 on each layer, a multi-resolution feature body is generated. The final multi-resolution feature body contains multi-scale texture information data tensors at the original image level and three compressed scales, and is stored as a multi-layer tensor structure for subsequent processing.

[0036] Step S2: Analyze the texture definition of the image by fusing the feature bodies of each resolution, and divide the original image into defect training images and qualified training images according to the texture definition;

[0037] In this embodiment, a texture definition judgment module is used to evaluate the texture quality of each scale of each original training image. The module is based on a frequency domain definition evaluation method and performs a fast Fourier transform (FFT) on the feature body at each scale to extract the frequency domain high-frequency component response value. The frequency energy distribution density threshold is defined as 0.25. Images with high-frequency energy above the threshold are marked as clear texture images, and images below the threshold are marked as fuzzy texture images. The four-scale results are combined to determine the final texture definition label of the image. The weight coefficient is determined by the original resolution ratio of the scale, and the proportion is 4:2:1:1. After the judgment is completed, all images are labeled as "qualified training images" or "defect training images".

[0038] Step S3: Trace and correct the defect training images to defect correction training images, and construct positive and negative contrast images based on the defect correction training images and the defect training images; perform contrast training combined with the positive and negative contrast images to construct a pre-training model of the defect image;

[0039] In this embodiment, the image region tracing reconstruction processing is performed on the defect training image, the reconstruction processing is based on the region-guided inverse mapping CNN (Region-Guided Inverse Mapping CNN), the clear image edge structure is taken as the guide, the low-frequency texture smoothing filling and structure defect recovery are performed on the defect region, the defect correction training image is output, the image retains part of the original defect trace, while the texture continuity is recovered, then each group of defect training images and the corresponding defect correction images form positive and negative contrast image pairs, and are input into the feature contrast extraction network, the network adopts the Siamese structure, two branches share the weight structure, the L2 distance difference is calculated after the image features are extracted, the difference contrast feature tensor is obtained, the difference tensor is input into the salient region positioning module, the salient mask map is generated based on the attention mechanism, the mask map has a value range of 0-1, which represents the salient degree of each pixel point, the mask map is spatially aligned with the defect image and the correction image, then the channel dimension is spliced, the splicing order is the original image, the enhanced image and the mask image, a three-channel ternary image is constructed, and is input into the triplet contrast training network, the network adopts the Triplet structure, the training loss function is TripletLoss, and finally the defect image pre-training model is obtained through back propagation, the model output is an image distinguishability feature embedding vector.

[0040] Step S4: based on the qualified training image, a weak disturbance transformation is performed, and the images before and after the transformation are paired one by one to construct positive example pairs before and after the disturbance; the regions with consistent features are found from the positive example pairs before and after the disturbance, and the pre-training model of the qualified image is trained by using the region images;

[0041] In this embodiment, for each qualified training image, a perturbation operation based on multi-scale block is performed, the block method adopts a square window-based block method, each image is divided into 64 sub-blocks of 32x32 pixels, then a local texture data extraction operation is performed on each sub-block, the extraction method uses a local binary pattern (LBP) to obtain a texture response vector of the sub-block, and then the texture response vector is direction gradient deconstructed, a Sobel operator is used to extract gradients in x and y directions respectively to generate a direction response vector of each sub-block, a dominant direction set is generated by comparing the maximum direction of the response strength, and then a local texture distribution data is generated according to the distribution of the dominant direction of each sub-block in the whole image, the data structure is a combination structure of the dominant direction and the intensity value of each sub-block, finally, the local texture distribution data is used to perform a weak perturbation transformation operation on each qualified image, the perturbation method is a Gaussian blur kernel convolution, the blur kernel size is set to 3x3, the standard deviation σ is 0.6, and the perturbation operation is only performed in the area where the texture dominant direction is consistent to ensure that the overall structure of the image is not damaged, and a perturbation qualified training image is generated, then the pre-perturbation image and the post-perturbation image are paired to form a pre-perturbation and post-perturbation positive example pair, which is input into a positive example comparison learning network, the network structure is a symmetric residual network, Cosine Similarity is used as a loss measurement standard, and a qualified image pre-training model is trained and output.

[0042] Step S5: The feature data of the flaw image pre-training model and the qualified image pre-training model are integrated into a collaborative sample stream, and joint training is performed through the sample stream, so as to obtain an integrated image pre-training model.

[0043] In this embodiment, the trained flaw image pre-training model and the qualified image pre-training model are loaded, the intermediate embedding layer features output by the two models are extracted using a consistent image input format, a sample feature stream is constructed, the two models are applied to the same batch of training images respectively, and the embedding features of the two models are output synchronously, the two embedding feature vectors are fused using a channel splicing method to obtain a collaborative sample stream, the feature stream is input into an integrated pre-training model in a subsequent training stage, the integrated model adopts an attention-guided feature weighting mechanism, different attention weights are applied to different parts of the input feature stream, the final image embedding feature is output through a fully connected layer, and the integrated model is trained and optimized in a contrast training strategy, so as to obtain an integrated image pre-training model which can be used for pre-training of multiple types of images, and the output model is saved in a standard ONNX format for subsequent inference.

[0044] Preferably, step S1 comprises the following steps:

[0045] Step S11: Cross-domain sampling is performed on the original image data to obtain a heterogeneous source data set; a semantic segmentation mapping is performed on the heterogeneous source data set to generate a semantic guidance tensor;

[0046] Step S12: high-dimensional feature mapping is performed on the original training image based on the semantic guidance tensor, wherein nested convolution operation is performed by using a channel dimension expansion and semantic alignment module, the convolution kernel size is set to 3*3, and the expansion rate range is [1, 2, 4], so as to obtain feature dimension mapping data;

[0047] Step S13: distribution consistency calibration is performed on the feature dimension mapping data to obtain normalized feature space data;

[0048] Step S14: multi-scale pyramid transformation is performed on the normalized feature space data, wherein the pyramid scale layer number is set to 3 layers, the down-sampling rate of each layer is 1, 0.5 and 0.25 in turn, and the channel compression rate is set to 0.5, so as to obtain a multi-resolution feature body.

[0049] In this embodiment, a heterogeneous source dataset is constructed by cross-domain sampling. This operation extracts images from public industrial image datasets and custom working condition image sets according to a proportion of 1000 images per category. The public dataset includes MVTec AD and DAGM2007, and the custom image set is collected by an industrial camera during actual pipeline production. The sampling standard is to collect once every five minutes, and the fixed angle 45-degree oblique light and overhead light combination method is used for each sample collection. The resolution of each image is fixed at 2048x2048 pixels. After sampling, the class label index is established by label matching rules, and then the sampled images are input into the pre-trained DeepLabV3+ semantic segmentation network. The ResNet-101 structure is selected for the backbone network, and the output layer dimension is fixed at 1 / 4 of the original image size. After interpolation up-sampling to restore the original image size and registration with the original image, the output is stacked according to the channel number of 3 to generate a semantic guidance tensor with a size of 2048x2048x3. This tensor is used to guide the high-dimensional feature mapping operation downstream. In the data structure, the semantic guidance tensor is stored in uint8 format, and each pixel value represents its semantic class index number, which is obtained by pre-training semantic network inference. The number of semantic categories is set to 13. By constructing a nested convolution structure and introducing a channel dimension expansion and semantic alignment module, the channel dimension expansion operation uses a Depthwise Separable Convolution (Deep Separable Convolution) structure to expand the feature channel number of the original image from 3 to 64. The convolution kernel size is uniform at 3x3, the step is set to 1, and the padding mode is SAME, which ensures that the output size remains unchanged. During the convolution process, three different inflation rates are used, namely 1, 2 and 4, for parallel convolution. After the output is feature spliced, a feature mapping tensor with a size of 2048x2048x192 is formed. Then, a cross-channel alignment operation is performed through the semantic alignment module. The method is to perform channel broadcasting on the semantic guidance tensor and element-wise multiplication with the feature mapping tensor to form a semantic-enhanced feature dimension mapping data. This operation relies on GPU parallel computing, and TensorFlow tf.nn.The depthwise_conv2d function implements all nested convolution operations, the convolution weight is initialized using the He Normal strategy, all weight parameters are updated step by step through the back propagation mechanism, and the distribution consistency calibration processing is performed on the feature dimension mapping data. Specifically, the distribution normalization method is used for feature space mapping standardization. First, the Batch Normalization operation is performed on the channel dimension, and the mean μ and standard deviation σ of each channel are calculated. Then, each pixel value is normalized according to the method of (x-μ) / σ, where μ and σ are values calculated by channel in the entire training batch. BatchSize is set to 16, and the tensor type processing with float32 floating-point precision is used during calculation. After normalization, all data values are limited within the range of approximately standard normal distribution. This processing process is realized by the torch.nn.BatchNorm2d module in the PyTorch framework. The normalized tensor is called normalized feature space data, and the size of the tensor is still 2048x2048x192. On this basis, the next step of pyramid structure construction is performed. By setting the pyramid three-layer scale structure, each layer corresponds to a down-sampling ratio of 1, 0.5 and 0.25 respectively. The first layer maintains the original resolution, the second layer is down-sampled to 1024x1024x192 through bilinear interpolation, and the third layer is down-sampled to 512x512x192. Then, the channel number of each layer is compressed. The channel compression rate is set to 0.5. The 1x1 convolution kernel is used for channel compression operation. The channel number of the first layer is compressed from 192 to 96, the channel number of the second layer is compressed from 192 to 96, and the channel number of the third layer is also processed to 96. Finally, three groups of feature tensors with different resolutions but consistent channels are obtained, each group of tensors corresponds to a scale layer of the pyramid. After the three groups of tensors are packaged in a unified format, a multi-resolution feature body is formed. The feature body retains spatial detail information and scale change information, which is used for subsequent texture judgment and training image classification processing. The construction of the entire pyramid structure is realized by using the tf.image.resize function of TensorFlow for down-sampling, and the tf.nn.conv2d function for 1x1 convolution channel compression.

[0050] Preferably, step S2 comprises the following steps:

[0051] Step S21: performing cross-scale aggregation on the multi-resolution feature body to obtain fused texture features; and performing local texture response mapping on the fused texture features to obtain local texture response distribution data.

[0052] Step S22: calculating local structure contrast based on the local texture response distribution data, wherein the texture window size is 8x8 pixels and the sliding step is 4 pixels; and inferring image texture sharpness based on the local structure contrast.

[0053] Step S23: performing threshold segmentation on the image texture definition based on a preset texture definition threshold to generate a definition segmentation label;

[0054] Step S24: screening the original training images into flaw training images and qualified training images through the definition segmentation label.

[0055] In this embodiment, the multi-resolution feature bodies extracted from the three-layer pyramid structure are respectively subjected to bilinear interpolation upsampling operation to unify the spatial resolution. In this operation, the image matrix interpolation method is adopted to scale the low-resolution feature map to be consistent with the highest resolution layer, and then the feature maps of the three scale layers are connected in the channel dimension to form the fused feature map after channel expansion. Next, a local texture response mapping module is introduced to perform local filtering processing on each spatial position in the fused feature map. The filter kernel used is a response template map. Through image local response operation, the gray response intensity of each pixel in the receptive field is analyzed. In actual operation, the local texture response tensor is constructed by using image local variance calculation and combining the response intensity changes. Finally, the response maximum value of the tensor in the spatial dimension is extracted to generate the local texture response distribution data. The whole operation process does not involve weight training and gradient propagation, but only pure forward mapping and response feature value output. The local texture response distribution data is divided into several continuous tiled non-overlapping regions, and the fixed texture window size is set as 8x8 pixels. The image is scanned in turn, and the moving interval of the sliding window is set as 4 pixels in both horizontal and vertical directions. In each 8x8 pixel region, the average gray response value and standard deviation value are first calculated. Then, the response difference value operation is performed between the current region and the adjacent eight-neighbor regions one by one, so as to obtain the response gradient difference between the current region and the adjacent regions. Then, the local contrast value of the region is represented by the ratio of the response gradient difference to the standard deviation. Further, the local structure contrast map is constructed by clustering and counting the distribution range of the contrast values in all local windows. Finally, the contrast map is aggregated and normalized in the channel direction within the full image range to generate a texture sharpness map in the continuous value domain. The value of each pixel represents the response degree of the texture sharpness in the local structure. The texture sharpness map is normalized to ensure that the numerical interval is concentrated between zero and one. Then, a unified static threshold reference line is set as the sharpness dividing point and mapped to the normalized image. The threshold value is derived from the response distribution median integral calculation result of a large number of qualified sample images on the local texture contrast map. When performing specific segmentation, the texture sharpness value of each pixel is compared with the pre-set threshold value. If it is greater than the threshold value, it is marked as a clear region pixel, otherwise it is marked as a fuzzy region pixel. Finally, the clear and fuzzy binary labels of all pixel points are used to generate a binary mask image. The mask image is the sharpness segmentation label image. The label image and the original image maintain the same spatial resolution and coordinate system. The sharpness segmentation label image corresponding to each image is obtained, and the ratio of the number of clear labels to the number of fuzzy labels in the label image is calculated. According to the ratio result, the division judgment standard is set. If the clear area ratio in an image is lower than the unified set proportion threshold, the image is classified into the defect training image set, otherwise it is classified into the qualified training image set.The determination of the proportion threshold is derived from the empirical dividing line calculated from 1000 manually annotated images in the manual screening experiment. To ensure the accuracy of the grouping results, after the classification is completed, the training images and the corresponding label images are packaged and stored in different training directory paths according to the grouping situation. The defect training image set is used for subsequent texture enhancement pre-training task, and the qualified training image set is used for standard feature extraction training path to improve the robustness and texture understanding ability of the overall image pre-training model.

[0056] Preferably, the defect training image in step S3 is traced and corrected to a defect correction training image, and the positive and negative contrast images are constructed based on the defect correction training image and the defect training image, including:

[0057] Extracting the defect features in the defect training image, and cutting the defect region image in the defect training image based on the defect features;

[0058] Performing defect tracing processing on the defect region image through the defect features, thereby obtaining image defect source data;

[0059] Performing source defect avoidance simulation on the image defect source data to generate simulated defect source avoidance data;

[0060] Performing defect correction processing on the defect region image through the simulated defect source avoidance data, thereby obtaining the defect correction training image;

[0061] Projecting the defect correction training image and the defect training image into the same space, and performing image comparison processing to obtain the positive and negative contrast images.

[0062] In this embodiment, the training image set marked as defect image by the definition segmentation label is loaded, and the gray scale information in the image channel is combined with the texture response map for feature cross coding. The specific operation is to unify the image gray scale map, definition mask map and local contrast map into a three-channel composite feature map in the form of channel stacking, and to call a differential feature extraction module based on an edge enhancement convolution kernel in the composite map. The module includes a 5x5 convolution kernel group for amplifying edge mutation areas and suppressing continuous smooth area response. After sliding across the entire image through the convolution kernel group to obtain a feature enhancement response map, a region growing method is further used to lock the region boundary in the feature set. After the boundary is closed, the region is extracted from the original image in the form of a mask to obtain a defect region image in each image. A fixed context expansion strategy is adopted during the cutting operation to increase a 6-pixel boundary around the cut region to retain the surrounding semantic context information. The output size is uniformly adjusted to 64x64 pixel resolution for subsequent defect tracking operations. Each defect region image obtained by cutting is input into a pre-constructed traceability network structure based on an expanded receptive field convolution stacking structure. The structure includes two series of encoding and decoding modules and an explicit location attention module. When inputting the defect image, it is first extracted through three convolution layers with sizes of 3x3, 5x5 and 7x7 in turn to obtain feature maps at different scales. Then, the explicit location attention module is used to weight process the feature response weight of each spatial position. The module calculates the attention response factor corresponding to each pixel by combining two-dimensional coordinate embedding and channel attention, and performs point multiplication mapping operation on the original feature map to generate a location marked feature map. Finally, the context semantic features of the original image are combined for deconvolution reconstruction. The reconstruction output is a spatial coordinate image data containing the starting point of defect generation, edge expansion trajectory and explicit defect texture change path. This data constitutes the image defect source data. All coordinates are uniformly mapped and output based on the original image space. The coordinate information and structure path in the traceability result are read and input into a defect avoidance generator based on guided edge morphological deformation repair. The generator consists of two sub-networks, namely the edge curvature reconstruction module and the morphological consistency constraint module. The former generates a defect boundary shape line graph based on the input path data, fits all discontinuous path segments using a Bezier curve, and constructs a continuous and derivable boundary curve. The latter uses the line graph as a mask to reconstruct the original image. In the morphological consistency constraint module, a specific regularization loss is used to constrain the gray scale continuity, edge smoothness and texture homogeneity of the reconstructed region. The reconstruction result maintains edge alignment and texture continuity with the original background region. After the simulated avoidance image is verified by the adversarial sample detection module, unstable images are removed. Only the avoidance data with texture boundary consistency greater than the threshold is output as the final simulated defect source avoidance data. The texture boundary consistency is calculated by the structural similarity index SSIM (Structural Similarity Index) and the minimum threshold is set to 0.82, The fusion interpolation reconstruction operation is performed based on the image inpainting inference network. First, the defect region positioning map in the original image is taken as an input mask to guide the input into the reconstruction network. Meanwhile, the source avoidance data and the original image gray feature map are concatenated in the channel dimension as feature input. The entire image inpainting inference network is based on the U-Net structure, and the jump connection features are fused to retain the image context structure information. The complete texture region is recovered through multi-scale decoding. In this process, rough semantic filling is performed in the low-resolution stage, and texture details are supplemented in the high-resolution stage to ensure that the repaired region and the surrounding image structure are seamlessly transitioned. Finally, the output is a corrected image consistent with the size of the original image. The image is the defect correction training image, which is saved in the same order as the original image and labeled as a cleaned image sample. An image embedding network based on a shared feature encoding structure is used. First, the embedding vectors of the two images are extracted through the weight-shared encoder module. The encoder is a residual structure convolutional network containing four convolution blocks. Each convolution block is composed of two 3x3 convolution layers and one 2x2 max pooling layer. The embedding vector is a 512-dimensional global feature vector output at the end of the encoder for each image. Then, the two sets of embedding vectors are input into the Euclidean distance module to calculate the distance matrix, while retaining the image coordinate level alignment features for pixel-level difference analysis. All image pairs are labeled with the same image number, and are constructed as image pair data structures with the defect image as the negative sample and the corrected image as the positive sample. Finally, the two images are spliced into a single-channel comparison image, with the defect training image on the left, the corrected image on the right, and the pixel-level difference heat map displayed in color-coded form in the middle. The positive and negative comparison images are constructed and saved under the comparison image training set path.

[0063] Preferably, the step S3 of performing contrast training on the positive and negative comparison images to construct the pre-training model of the defect image comprises:

[0064] extracting difference contrast features of the positive and negative comparison images;

[0065] locating the salient region of the image based on the difference contrast features, and determining a contrast salient mask based on the salient region of the image;

[0066] spatially aligning the contrast salient mask, the defect training image, and the defect correction training image, and performing channel splicing processing to generate an original-enhanced-salient mask ternary image;

[0067] extracting ternary image feature data in the original-enhanced-salient mask ternary image;

[0068] inputting the ternary image feature data into a contrast learning network for ternary contrast training to generate discriminative contrast features;

[0069] training the defect image pre-training model through the discriminative contrast features.

[0070] In this embodiment, the defect correction training image and the corresponding defect training image are respectively subjected to size normalization processing, and the images are adjusted to a unified resolution of 256x256 pixels by using a bilinear interpolation method. Subsequently, the two types of images are compared pixel by pixel by using a method based on a structural similarity index SSIM (Structural Similarity Index Measure). Under the condition that the size of a sliding window is set to 11x11 and the window step is set to 1, the brightness, contrast, and structural difference of a local region are calculated, thereby obtaining a difference matrix. Further, the pixel region with a difference value greater than 0.25 in the difference matrix is extracted as a difference contrast feature region. Finally, the output difference contrast feature is expressed in the form of a binary graph, wherein a pixel value of 1 represents a significant pixel region with a structural difference, and a pixel value of 0 represents a region without a difference. The binary graph generated by the foregoing is used as an initial mask and input into a saliency detection network based on a full convolutional network FCN (Fully Convolutional Network). The network structure comprises five convolutional layers and three up-sampling layers. In the convolutional layers, a ReLU activation function and a BatchNorm normalization operation are used to accelerate the convergence speed and suppress overfitting. After the input image is processed by the network, a saliency heat map with a dimension of 256x256 is output. Subsequently, a threshold segmentation method is used to extract a salient region, wherein the saliency threshold is set to 0.3. When the pixel value in the heat map is greater than 0.3 corresponds to the significant area, and finally the threshold segmentation result is morphologically closed to repair the edge breakage, and then an opening operation is performed to remove isolated noise points. After processing, the obtained significant area mask is the contrast significant mask. The area with a mask pixel value of 1 is used for subsequent spatial alignment and channel splicing. Affine transformation is performed on three types of images to unify the spatial reference system. The affine matrix is obtained by calculating the center point of the image and the geometric center point of the significant area. The affine transformation matrix is generated by using the getAffineTransform function in OpenCV, and then the warpAffine operation is performed on the image to obtain the image data in the unified space. After ensuring that the images are aligned, the three types of images are stacked according to the channel dimension. The defect training image is input as the first channel, the defect correction training image is input as the second channel, and the contrast significant mask is input as the third channel. Finally, a three-channel image data with a size of 256x256x3 is generated, which is called the original-enhancement-significant mask ternary image. The pixel data of the ternary image is stored in float32 format and normalized to the [0, 1] interval to adapt to the input requirements of the neural network. A set of shallow convolutional neural network models are constructed to extract feature information. The network structure includes three convolutional layers, each with a kernel size of 3x3, channel numbers of 32, 64 and 128 respectively, a step size of 1, and a same padding method. After convolution, BatchNorm normalization and ReLU activation functions are connected. After each convolution, a max pooling operation is added to compress spatial information. The pooling window is set to 2x2. Finally, the convolution output is flattened into a one-dimensional vector through the Flatten operation, and two fully connected layers are used for further feature projection processing. The output dimension of the first fully connected layer is 512, and the output dimension of the second fully connected layer is 128. The 128-dimensional vector is the ternary image feature data. The feature vector is normalized by L2 regularization to stabilize the subsequent contrast learning training. A contrast learning network with a Triplet Loss structure is constructed, where the Anchor input is the ternary image feature corresponding to the defect training image, the Positive input is the ternary image feature corresponding to the defect correction training image, and the Negative input is the unrelated image ternary feature randomly sampled from different image batches. A triplet loss function with a margin of 0.5 is used for training. The Loss function is set to max(D(anchor, positive)-D(anchor, negative)+margin, 0) form, where D is the Euclidean distance calculation function. The model is optimized, and the initial learning rate is set to 0.001, the batch size is set to 64, the training round is set to 120, and the Anchor and Negative features are mined in the batch during the training process, that is, the Negative sample with the smallest distance from the Anchor is selected from the current batch to form a difficult negative pair to enhance the model discrimination ability, and finally the contrast learning embedding model containing 128-dimensional contrast feature representation is output. Using the aforementioned trained contrast learning embedding model as the basis, the extracted discriminative contrast features are input into the backbone network as a supervision signal. The backbone network structure uses ResNet18, removes the final classification layer based on the original ResNet18 structure, and adds a fully connected layer with an output of 128 dimensions. The mean square error between the discriminative contrast features and the backbone network output features is used as the supervision loss function for training. During the training process, the parameters of the contrast learning embedding model are frozen, and only the ResNet18 network parameters are updated. The AdamW optimizer is used, the initial learning rate is 0.0005, the regularization parameter is set to 1e-4, and the training data uses the aforementioned generated triple image data. The batch size is set to 32, and the training round is set to 80 rounds. Finally, the output model is used for pre-training representation construction of the defect image features. All model weights are stored in float32 format. After training, the model parameters are exported as an ONNX format file for subsequent model deployment.

[0071] Especially important is that the difference contrast feature is used to locate the image salient region, and the contrast salient mask is determined based on the image salient region, including:

[0072] According to the response amplitude distribution of each channel in the difference contrast feature, channel focusing processing is performed to construct a focused difference feature map;

[0073] The focused difference feature map is subjected to spatial saliency projection to obtain an initial salient response map;

[0074] Based on the initial salient response map, region expansion fusion is performed to generate a salient candidate region map;

[0075] The salient candidate region map is subjected to boundary refinement processing to obtain an image salient region;

[0076] The contrast salient mask is constructed through the image salient region.

[0077] In this embodiment, the channel weighting module is used to analyze the input difference contrast feature tensor channel by channel. The shape of the difference contrast feature tensor is set as [C, H, W], where C is the number of channels, H and W are the height and width of the feature map, respectively. First, the response value of each channel is mapped to a one-dimensional vector by the global average pooling method. Then, the one-dimensional response vectors of all channels are normalized by the maximum and minimum normalization method, that is, the response value of each channel is subtracted by the minimum value of the channel and then divided by the difference between the maximum and minimum values. The normalized response is used to generate the channel attention weight map. Then, the channel attention weight map and the original difference contrast feature are multiplied channel by channel to construct a focused difference feature map. This map focuses more on the channel component response that contains obvious difference regions in terms of semantic expression ability. A 1x1 convolution kernel in the two-dimensional convolutional neural network (Convolutional Neural Network) is used to perform channel fusion processing on the focused difference feature map, reducing the channel dimension to 1 to form a single-channel spatial map. The convolution operation weight is generated by the He initialization method and fixed to the pre-training state to maintain the stability of channel fusion. Then, the sliding window maximum pooling operation is used on the reduced image in the spatial dimension. The window size is set to 5x5, the step is 1, and the padding value is 0. This processing is used to enhance the response value of the focused region and improve the spatial contrast. The final generated initial salient response map has a size of [1, H, W], which retains the high response region of spatial saliency. The pixel value of the initial salient response map is normalized to between 0 and 1. The threshold segmentation method is used to determine the starting point of the candidate region. The threshold is set to 0.6. All pixel points greater than the threshold are determined as seed point regions. Then, the seed point region is expanded based on the morphological dilation operation. The structure element is set to a 3x3 cross kernel, and the iteration number is 2. After each expansion, a region average fusion operation is performed, that is, the pixel values in each connected region are averaged and assigned to the entire region to smooth the response difference. After expansion and fusion, a salient candidate region map is generated, which is a binary image with the same size as the original image. The salient candidate region map is subjected to Gaussian blur processing. The convolution kernel size is set to 5x5, and the standard deviation is 1.2, for removing noise edge, using Canny operator for edge detection, set low threshold value to 50, high threshold value to 150, through the Canny edge detection result extraction region contour boundary line, then through the edge point fitting method generates continuous closed region, fitting method adopts least square polygon approximation, the boundary after contour fitting is used for cutting candidate region graph, the obtained image salient region is the region mask in the closed contour, the mask only contains the connected response structure in the focus area range, convert the image salient region graph into mask format, define the pixels with response value greater than 0 in the mask as 1, and the rest as 0, form a single channel binary mask graph, then the mask graph is subjected to size normalization processing, keep it with the same size as the input image and the enhanced image, the corresponding operation is bilinear interpolation (Bilinear Interpolation) resampling method, then the mask is aligned with the input image channel, the mask is copied to 3 channels to form a 3 channel mask graph, which is used for subsequent input with the original graph and the enhanced graph, the salient mask graph is the contrast salient mask, which has spatial precision and semantic contrast feature consistency, which can be directly used as an input auxiliary channel of the saliency contrast module in the training stage, the mask value is only 0 or 1, and the gray distribution area is not included, so as to ensure the binary input standard.

[0078] Preferably, in step S4, weak perturbation transformation is performed based on the qualified training images, and the images before and after transformation are paired one by one to construct the positive example pairs before and after perturbation, including:

[0079] The qualified training images are subjected to multi-scale block processing, and local texture data in each sub-block is extracted;

[0080] Local texture distribution data of the local texture data is calculated;

[0081] Weak perturbation transformation is performed on the qualified training images based on the local texture distribution data, so as to generate perturbed qualified training images;

[0082] The perturbed qualified training images and the qualified training images are paired one by one, so as to generate positive example pairs before and after perturbation.

[0083] In this embodiment, the size of the qualified training image is normalized to a uniform scale of 512x512 pixels. The original image is then size-standardized using linear interpolation, and then the image is divided into blocks according to the multi-scale setting. The block size is set to three levels: 32x32, 64x64, and 128x128. At each scale, a non-overlapping sliding window is used to scan the image region and extract the sub-block content. Each scale corresponds to a different fine-grained structure extraction purpose. For each sub-block, the Gray Level Co-occurrence Matrix method is used to extract local texture data. During the Gray Level Co-occurrence Matrix extraction process, the gray level is set to 256 levels. The energy, contrast, entropy, and correlation of each sub-block in the 0-degree, 45-degree, 90-degree, and 135-degree directions are calculated. The texture data of all sub-blocks form the texture feature set of the image at different scales, which represents the texture composition and spatial distribution state of the local region in the image. The local texture parameter matrix extracted at each scale is independently normalized. The maximum value normalization method is used, i.e., each dimension of the texture feature parameter is divided by its maximum value in the image sub-block set to obtain a normalized texture feature tensor with a range of 0 to 1. Then, a texture probability distribution model is constructed at each scale. This model is realized by histogram estimation of each normalized texture parameter. The number of histogram bins is set to 32. Each type of texture parameter probability distribution interval is accurately divided. The number of normalized texture values in each bin is counted and the proportion is calculated to generate a distribution vector of the texture feature. The distribution vectors at all scales are combined to form a multi-scale texture distribution vector, which serves as a statistical basis for subsequent disturbance function construction. The dimension of this vector is equal to the product of the texture feature dimension, the scale number, and the bin number. Based on the multi-scale texture distribution vector constructed in the previous step, a distribution consistency constraint is introduced in the disturbance transformation. First, the original image is enhanced based on Fourier frequency domain disturbance. The image is transformed to the frequency domain using discrete Fourier transform (DFT). The disturbance frequency band is set to the low frequency region, and the frequency range is the spectrum block within 16 pixels from the center of the image frequency domain. Random disturbance noise is added to this region. The noise in the amplitude direction follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.03. The phase direction remains unchanged. Then, the disturbed image in the frequency domain is restored to the spatial domain using inverse discrete Fourier transform (IDFT). The low-frequency texture in the disturbed image changes slightly but the structure remains intact. Then, the local texture distribution vector of the disturbed image is calculated and compared with the original distribution vector using cosine similarity. The similarity threshold is set to 0.92, if not satisfied, re-generate the perturbation noise and repeat the above frequency domain perturbation process until the texture distribution consistency is satisfied, the finally generated perturbation image is the perturbation qualified training image, an index numbering system is constructed for all qualified training images, each image is given a unique number, and the corresponding perturbation image is kept consistent with the original image number, in the pairing process, a mapping dictionary is constructed with the number as the key value, the original image and the perturbation image form a key-value pair structure, and the one-to-one correspondence between the original image and its perturbation image is realized by calling mode, the paired images are saved in tensor mode, each pair of image combination forms a four-dimensional tensor with a size of [2, C, H, W], wherein C is the channel number and is set to 3, H and W are the image height and width, respectively, both are 512, the tensor is input into the contrast learning module in the training network as a positive example pair, and in the data loading stage, it is assembled into a batch tensor through batch reading operation, the batch size is 16, that is, 16 pairs of perturbation before and after images are input each time, through the index mechanism, the strict pairing relationship between the perturbation before and after images can be maintained without affecting the data structure of the image itself, and the texture distribution and structural feature correlation between the positive example pairs are fully preserved.

[0084] Especially important is that the local texture distribution data of the local texture data comprises:

[0085] The local texture data is directionally gradient-decomposed, thereby obtaining a direction response vector;

[0086] The maximum direction of response intensity in the direction response vector is counted;

[0087] Based on the maximum direction of response intensity, spatial aggregation is performed and direction distribution frequency data is analyzed;

[0088] The dominant texture direction distribution data is determined through the direction distribution frequency data and the maximum direction of response intensity;

[0089] According to the dominant texture direction distribution data, the local texture data is texture intensity depicted, thereby obtaining a texture response intensity matrix;

[0090] Based on the texture response intensity matrix, stability weighted reconstruction is performed, thereby generating the local texture distribution data.

[0091] In this embodiment, the original image is divided according to a fixed scale, and the image block size is set to 64x64 pixels. Each sub-block is processed independently as a processing unit. The Sobel operator is used to convolve the pixel gray scale values in each sub-block. The horizontal Sobel kernel [-1, 0, 1; -2, 0, 2; -1, 0, 1] and the vertical Sobel kernel [-1, -2, -1; 0, 0, 0; 1, 2, 1] are used to act on the image gray scale data, respectively. The gradient amplitude in the x and y directions for each pixel point is calculated. Then, the arctangent function is used to calculate the direction angle θ, where θ is equal to the arctangent of the y direction gradient value divided by the x direction gradient value. The range of the direction angle θ is mapped to the closed interval of 0 to 180 degrees by adjustment. The direction angle corresponding to each pixel point is recorded as the direction response value of the point. The direction response values of all pixels in the sub-block are counted to form a response vector containing 4096 angle values. This vector is the direction response vector of the image sub-block. The gradient intensity corresponding to each pixel direction angle θ in the direction response vector is paired and bound. The intensity value is the square root of the sum of the squares of the Sobel convolution calculated amplitude. The intensity value of each pixel point is the square root of the sum of the squares of the x and y direction gradient values. A histogram is constructed for each direction angle θ. The angle division granularity is set to 10 degrees, and a total of 18 direction intervals are set. The intensity values corresponding to the pixels falling into each direction interval are accumulated. The intensity sum vector corresponding to all intervals is vectorized to obtain a direction intensity vector with a length of 18. The index of the element with the maximum intensity value in the vector is searched. The median value in the direction interval corresponding to the index is defined as the main direction angle of the sub-block. The value of the main direction angle belongs to the set {5 degrees, 15 degrees, 25 degrees, …, 175 degrees}. A direction consistency aggregation kernel is constructed in each image sub-block with the main direction angle as the center. The kernel size is set to 7x7 pixels. The angle difference between each pixel point and the main direction angle is calculated. The allowed difference range is set to ±15 degrees. The pixels that meet the angle difference condition are aggregated. The frequency of each direction angle in the aggregation region is calculated by dividing the number of pixels falling into the corresponding direction angle interval by the total number of pixels in the aggregation region. The statistical result forms an 18-dimensional direction frequency vector. The direction frequency vector represents the local direction consistency feature of the sub-block under the guidance of the main direction angle. The value of each dimension in the direction frequency vector is between 0 and 1, and the total sum is 1. The vector provides a basis for subsequent dominant texture direction recognition. The direction frequency vector obtained in the previous step and the main direction angle are combined to construct a dominant direction distribution description vector. The vector has a dimension of 18. The construction method is to set the interval containing the main direction angle as the weight center and amplify the frequency values in the center interval and its adjacent left and right two direction intervals by weighting. The weighting coefficients are set as follows: the center interval weight is 1.5, the left and right two intervals weights are 1.2, and the weights of the remaining direction dimensions are 1.0, multiply all dimensions by the corresponding weight and then re-normalize, so that the sum of the vector values is 1, and the direction frequency weighted vector after normalization is the dominant texture direction distribution data of the image sub-block, which describes the direction consistency and strengthens the dominant direction angle feature, for each pixel, according to the frequency value of the direction interval where the direction angle is located in the dominant texture direction distribution data, the specific process is to read the direction angle θ of the pixel, and then it is classified into the nearest direction interval, and the value of the direction interval in the dominant texture direction distribution vector is obtained as the direction response weight of the pixel, then the weight is multiplied by the gradient intensity value to generate a new weighted response value, and the weighted response values of all pixels in the image sub-block are combined to form a matrix, the dimension of the matrix is the same as the size of the image sub-block, that is, 64*64, and the matrix is called a texture response intensity matrix, which is used to describe the direction consistency response intensity distribution of each pixel in the sub-block, and the texture response intensity matrix is edge filtered, bilateral filtering (Bilateral Filtering) is used to denoise and smooth the texture response intensity matrix, the spatial domain Gaussian kernel parameter is set to σs=3, and the intensity domain Gaussian kernel parameter is set to σr=0.1, and the smoothed texture response matrix is output after processing, then the matrix is analyzed for regional stability, and the structure tensor (Structure Tensor) method is used to calculate the local stability metric, the structure tensor is constructed for each pixel and the eigenvalues and eigenvectors are calculated, if the maximum eigenvalue is significantly greater than the minimum eigenvalue, the point is determined as a stable texture point, the average response value of all stable texture points in each sub-block is calculated, and the average value is used to construct a weighted mask, and the mask is applied to the smoothed response matrix to retain the stable region response value and suppress the unstable region response value, finally the local texture distribution data of the image sub-block is output, and all the sub-blocks are spliced to form the local texture distribution map of the whole image.

[0092] Preferably, the step S4 of finding the regions with consistent features from the positive example pairs before and after the disturbance and training the pre-trained model of the qualified image using the region images of the regions comprises:

[0093] Extracting feature data of each layer of the positive example pairs in the positive example pairs before and after the disturbance;

[0094] Respectively evaluating the consistency of the feature data of each layer of the positive example pairs, thereby generating feature consistency distribution data;

[0095] Filtering high-consistency regions based on a preset consistency distribution threshold and the feature consistency distribution data;

[0096] Disassembling the corresponding region images of the qualified training image through the high-consistency regions;

[0097] Training the pre-trained model of the qualified image through the corresponding region images.

[0098] In this embodiment, the forward propagation is performed on the two positive example images before and after perturbation by the convolutional neural network (CNN) structure of each layer of the pre-trained image model. The specific process is as follows: the size of the input image is uniformly adjusted to 224x224 pixels, the pixel value is standardized to between 0 and 1, and the feature maps of different levels are extracted by sequentially passing through the multiple convolution layers, batch normalization layers and activation function layers (such as ReLU, linear rectifier unit) of the model. The extracted feature maps include shallow edge information, middle texture features and deep semantic features. The size of the shallow feature map is 56x56x64, the size of the middle feature map is 28x28x128, and the size of the deep feature map is 14x14x256. The above feature maps are obtained from the corresponding levels of the positive example images before and after perturbation, and are saved in the form of multi-dimensional tensor data for subsequent consistency evaluation processing. Cosine similarity is used as the consistency evaluation index. The feature vectors of the same position in the same layer before and after perturbation are unfolded into one-dimensional vectors along the channel dimension, the dot product of the two one-dimensional vectors is calculated, and the cosine similarity value of the position is obtained by dividing the product of the Euclidean norms of the respective vectors. The value range is limited to -1 to 1. To avoid the influence of negative values, negative values are uniformly truncated to 0. The calculation result forms a two-dimensional consistency distribution matrix with the size of the corresponding feature map. For example, a 28x28 consistency matrix is output for a 28x28 size feature map, and each element in the matrix corresponds to the similarity score of the feature vectors before and after perturbation at that position. This operation is repeated for all levels, and a set of multi-level consistency distribution matrices is finally generated and stored as floating-point tensor data for screening high-consistency regions. The consistency distribution threshold is set to 0.85, and each layer consistency distribution matrix is traversed. The elements in the matrix that are greater than or equal to 0.The element position of 85 is marked as a high-consistency region, a binary operation is used to assign a value of 1 to positions higher than a threshold value and a value of 0 to positions lower than the threshold value, then spatial clustering is performed on the binary matrix, a continuous connected region is selected, the area size of each connected region is calculated, noise regions with an area smaller than 16 pixels are filtered out, and the remaining connected regions are taken as valid blocks of the high-consistency region of the layer. For feature maps of different layers, the high-consistency region is mapped back to the original input image coordinate system through upsampling method combined with the corresponding image spatial scale to form a complete high-consistency region mask. The mask data is used for corresponding region image disassembly operation. According to the mapped high-consistency region mask, image segmentation operation is performed on the qualified training image, the image size is unified to 224*224 pixels, and the pixel region with a mask value of 1 is taken as the region of interest for extraction. A mask-based cropping method is applied to extract the minimum rectangular bounding box region containing the complete connected region. The bounding box coordinates are calculated according to the mask pixel distribution. The image content of the cropped region is taken as the corresponding region image after disassembly. If the size of the bounding box is insufficient for 112*112 pixels, 0-value pixels are filled around the bounding box to reach the size, ensuring the consistency of the input size in the subsequent training batch. The cropped region image is saved in the standard RGB three-channel format, and the storage format is lossless PNG file. All disassembled region images constitute a training sample set. The region image sample set obtained by disassembly is subjected to data enhancement processing, including random horizontal flipping, random rotation by ±15 degrees, and random color jitter. The enhanced data batch is input to the pre-trained model in units of 64 images. The model structure is based on the ResNet50 deep residual network, and the cross-entropy loss function is selected as the loss function. Batch normalization layers and Dropout layers with a dropout rate of 0.3 are set during the training process. The accuracy and loss are calculated using the validation set after the end of each epoch during the training process. The number of training iterations is set to 100 rounds. After the training is completed, the model weight file and network parameters are saved.

[0099] Preferably, step S5 comprises the following steps:

[0100] Step S51: extracting output feature vectors of the flaw image pre-trained model and the qualified image pre-trained model, and performing complementary feature mapping based on the output feature vectors to obtain dual-modal feature fusion data;

[0101] Step S52: performing class distribution re-estimation based on the dual-modal feature fusion data to obtain class attribution distribution data; filtering distribution uncertain data in the dual-modal feature fusion data based on the class attribution distribution data, and constructing distribution uncertain label data;

[0102] Step S53: performing self-supervised pseudo-label expansion simulation on the distribution uncertain label data to generate pseudo-label expansion data; performing multi-task standard distribution on the pseudo-label expansion data to obtain multi-task training distribution data;

[0103] Step S54: Task cooperative training scheduling is performed according to the multi-task training distribution data, so as to obtain a cooperative sample stream; integrated model joint training is performed based on the cooperative sample stream, so as to obtain an integrated image pre-training model.

[0104] In this embodiment, the defect image and the qualified image are input into the respective corresponding pre-trained model, the model structure adopts ResNet50 architecture, the input image size is unified to 224x224 pixels, after the processing of each layer of convolution, pooling and residual module, the output feature vector is finally extracted from the fully connected layer or the global average pooling layer of the model, the feature vector dimension is set to 2048, two feature vectors with the same dimension are obtained for the same image pair, then the two feature vectors are complementarily mapped by using a feature mapping function, the mapping process adopts a bilinear pooling operation, specifically, the outer product of the two vectors is calculated to obtain a high-dimensional feature matrix, and then the low-rank decomposition is used to reduce the dimension to a 1024-dimensional fusion vector, the fusion vector is normalized to make each dimension value between 0 and 1, forming a double-modal feature fusion data, a multi-class Softmax classifier is used to calculate the forward propagation of the double-modal feature fusion data, and the prediction probability distribution of each class is output, the probability distribution forms the class attribution distribution data, the number of classes is set to 10 according to the task, the probability value is in the range of 0 to 1 and the sum is 1, the entropy value is calculated based on the distribution as the uncertainty measure, the entropy value formula adopts the standard discrete entropy formula, the higher the entropy value represents the more uncertain the distribution, the fusion feature data with the entropy value greater than 0.7 is marked as distribution uncertain data, then the corresponding prediction class label and the entropy value data are constructed as distribution uncertain label data, the label format adopts one-hot encoding method, and the entropy value weight is stored in the tensor data structure, for the distribution uncertain label data, first, a data enhancement strategy is introduced to diversify the corresponding image, including random cropping, rotation ± 10 degrees, color transformation and Gaussian noise superposition, the enhanced image is input into the model to predict the class probability again, the original label and the enhanced prediction label are combined to generate pseudo labels by consistency regularization method, the pseudo label adopts probability threshold screening method, the probability is highest and more than 0.The category of 8 is taken as a pseudo-label category, the pseudo-label data is merged with the original label data to form an expanded data set, then the training tasks are divided according to the task target, including a classification task, an adversarial training task and a feature reconstruction task, the expanded data set is distributed to different task sub-networks according to the label attribute and the feature distribution, and the multi-task learning framework is used to form multi-task training distribution data, the data format includes input features, label information and task identification, so that the data can be correctly routed to the corresponding sub-task module in the training process, a task scheduler is constructed, the multi-task training distribution data is input, the task scheduler dynamically adjusts the scheduling order of the training samples according to the task priority and the resource occupation, the scheduling strategy adopts a polling weighting method, so that the task samples are uniformly distributed in the training cycle, and the training sample flow is input into the corresponding task sub-network after scheduling, each sub-network structure includes a shared backbone network and a special task head, the gradient accumulation technology is used to optimize the multi-task training process, the classification loss, the adversarial loss and the reconstruction loss are weighted and summed through a joint loss function in the training process, the weight parameters are respectively set as 0.5, 0.3 and 0.2, the optimizer adopts AdamW, the learning rate is set as 0.00005, the training iteration number is 150 rounds, the weights of the task sub-networks are integrated after the training is completed, the integration method adopts a weighted average method, the weights are distributed according to the task verification accuracy, and finally the integrated image pre-training model with the multi-modal feature fusion and multi-task cooperation capability is generated, the weight parameter file and the model structure configuration are stored in a special format file, and are called for subsequent image recognition and classification tasks.

[0105] The application further provides a training system of an image pre-training model, which is used for executing the training method of the image pre-training model.

[0106] The image acquisition module is used for acquiring original training images, extracting semantic information in the images and performing feature processing to generate structured high-dimensional feature data; and the high-dimensional feature data is sequentially reduced at three different ratios to finally generate multi-resolution feature bodies.

[0107] The texture screening module is used for analyzing the texture definition of the images by fusing the feature bodies of different resolutions, and dividing the original images into flaw training images and qualified training images according to the texture definition.

[0108] The flaw modeling module is used for tracing and correcting the flaw training images into flaw correction training images, and constructing positive-negative contrast images based on the flaw correction training images and the flaw training images; and performing contrast training combined with the positive-negative contrast images, so as to construct a pre-training model of the flaw images.

[0109] The positive example modeling module is used to perform weak perturbation transformation on qualified training images and pair the images before and after the transformation to construct positive example pairs before and after perturbation; it identifies regions with consistent features from the positive example pairs before and after perturbation and uses these region images to train the pre-trained model of qualified images.

[0110] The ensemble training module is used to combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and then perform joint training through this sample stream to obtain the ensemble image pre-trained model.

[0111] This invention, through the application of an image acquisition module, can comprehensively acquire original training images, ensuring the diversity and richness of the data. Feature dimension mapping and pyramid transformation improve the accuracy of image feature extraction. The generated multi-resolution feature volume effectively captures subtle texture information in the image. The texture screening module, based on sharpness judgment, accurately distinguishes between flawed training images and qualified training images, optimizing the quality of subsequent training datasets and improving the reliability of model training. The source correction technology of the flaw modeling module ensures the authenticity of the training data and provides representative samples for the model. By constructing positive and negative contrast images, the learning ability of the flawed image pre-trained model is enhanced, enabling it to more effectively identify and process flawed features. The positive example modeling module expands the sample diversity of qualified training images through weak perturbation transformation, improving the model's generalization ability and robustness. The integrated training module effectively combines the pre-trained models of flawed images and qualified images, realizing collaborative training. The final integrated image pre-trained model exhibits superior performance when handling complex image tasks, accurately identifying and classifying various types of image features, greatly enhancing the system's practicality and adaptability.

[0112] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed, implements the training method of the image pre-training model as described in any of the above claims.

[0113] This invention utilizes computer-readable storage media to efficiently store and manage the training program of an image pre-trained model, ensuring the program's execution efficiency and accuracy, providing stable data access and processing capabilities, supporting the processing and analysis of large-scale image data, promoting the rapid integration of diverse training samples, optimizing the image feature extraction and processing flow, enhancing the flexibility and scalability of model training, ensuring the integrity and consistency of training data, strengthening the model's adaptability and robustness, and ultimately achieving a high-performance image pre-trained model to meet the needs of complex image analysis tasks.

[0114] Therefore, the embodiments should be regarded, at any point, as being exemplary and not limiting, the scope of the application being defined by the appended claims and not by the above description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.

[0115] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training an image pre-training model, characterized in that, The method comprises the following steps: Step S1: After collecting the original training image, the semantic information in the image is extracted and feature processing is performed to generate structured high-dimensional feature data; the high-dimensional feature data is sequentially reduced by three different ratios to finally generate a multi-resolution feature body; Step S2: The texture clarity of the image is analyzed by fusing the feature bodies of each resolution, and the original image is divided into defect training images and qualified training images according to the texture clarity; Step S3: The defect training image is traced and corrected to a defect correction training image, and a positive-negative contrast image is constructed based on the defect correction training image and the defect training image; combined with the positive-negative contrast image, contrast training is performed to construct a pre-training model of the defect image; Step S4: Based on the qualified training image, a weak disturbance transformation is performed, and the images before and after transformation are paired one by one to construct a positive example pair before and after disturbance; the regions with consistent features are found from the positive example pair before and after disturbance, and the images of these regions are used to train the pre-training model of the qualified image; Step S5: The feature data of the defect image pre-training model and the qualified image pre-training model is integrated into a cooperative sample stream, and joint training is performed through the sample stream to obtain an integrated image pre-training model.

2. The training method of the image pre-training model according to claim 1, characterized in that, Step S1 comprises the following steps: Step S11: Cross-domain sampling is performed on the original image data to obtain a heterogeneous source data set; semantic segmentation mapping is performed on the heterogeneous source data set to generate a semantic guide tensor; Step S12: High-dimensional feature mapping is performed on the original training image based on the semantic guide tensor, wherein nested convolution operation is performed by using a channel dimension expansion and semantic alignment module, the convolution kernel size is set to 3*3, and the expansion rate range is [1, 2, 4], so as to obtain feature dimension mapping data; Step S13: The feature dimension mapping data is subjected to distribution consistency calibration to obtain normalized feature space data; Step S14: Multi-scale pyramid transformation is performed on the normalized feature space data, wherein the pyramid scale layer number is set to 3, the down-sampling rate of each layer is 1, 0.5 and 0.25 in turn, and the channel compression rate is set to 0.5, so as to obtain a multi-resolution feature body.

3. The method of Claim 1, wherein, Step S2 comprises the following steps: Step S21: The multi-resolution feature body is subjected to cross-scale aggregation to obtain fused texture features; local texture response mapping is performed on the fused texture features to obtain local texture response distribution data; Step S22: Local structure contrast is calculated according to the local texture response distribution data, wherein the texture window size is 8*8 pixels, the sliding step length is 4 pixels, and the image texture clarity is inferred based on the local structure contrast; Step S23: The image texture clarity is subjected to threshold segmentation based on a preset texture clarity threshold to generate a clarity segmentation label; Step S24: The original training image is classified into defect training images and qualified training images through the clarity segmentation label.

4. The training method of the image pre-training model according to claim 1, characterized in that, In step S3, the defect training image is traced and corrected to a defect correction training image, and a positive-negative contrast image is constructed based on the defect correction training image and the defect training image, comprising: The defect features in the defect training image are extracted, and the defect region image in the defect training image is cut based on the defect features; The image defect source data is obtained by defect tracing processing on the defect area image through the defect feature; The simulated defect source avoidance data is generated by source defect avoidance simulation on the image defect source data; The defect correction training image is obtained by defect correction processing on the defect area image through the simulated defect source avoidance data; The defect correction training image and the defect training image are projected into the same space, and image comparison processing is performed to obtain positive-negative comparison images.

5. The method of Claim 1, wherein, The pre-training model of the defect image is constructed by comparison training in step S3 combined with the positive-negative comparison images, including: difference comparison features of the positive-negative comparison images are extracted; the image salient region is located based on the difference comparison features, and the comparison salient mask is determined based on the image salient region; the comparison salient mask, the defect training image and the defect correction training image are spatially aligned, and channel splicing processing is performed to generate an original-enhanced-salient mask ternary image; ternary image feature data in the original-enhanced-salient mask ternary image is extracted; the ternary image feature data is input into the comparison learning network for ternary comparison training to generate discriminative comparison features; the defect image pre-training model is trained through the discriminative comparison features.

6. The method of Claim 1, wherein, In step S4, the weak perturbation transformation is performed based on the qualified training image, and the images before and after the transformation are paired one by one to construct the positive example pairs before and after the perturbation, including: the qualified training image is processed by multi-scale block processing, and local texture data in each sub-block is extracted; local texture distribution data of the local texture data is calculated; the weak perturbation transformation is performed on the qualified training image based on the local texture distribution data, thereby generating the perturbed qualified training image; the perturbed qualified training image and the qualified training image are paired one by one to generate the positive example pairs before and after the perturbation.

7. The method of Claim 1, wherein, In step S4, the regions with consistent features are found from the positive example pairs before and after the perturbation, and the pre-training model of the qualified image is trained using these region images, including: extracting feature data of each layer of the positive example pairs before and after the perturbation; consistency evaluation is performed on each layer of the positive example pair feature data to generate feature consistency distribution data; high consistency regions are screened out based on a preset consistency distribution threshold and the feature consistency distribution data; the corresponding region image of the qualified training image is disassembled through the high consistency region; the pre-training model of the qualified image is trained through the corresponding region image.

8. The method of Claim 1, wherein, Step S5 includes the following steps: Step S51: extract the output feature vector of the defect image pre-training model and the qualified image pre-training model, and perform complementary feature mapping based on the output feature vector to obtain double-modal feature fusion data; Step S52: re-estimate the class distribution based on the double-modal feature fusion data to obtain class attribution distribution data; screen the distribution uncertain data in the double-modal feature fusion data based on the class attribution distribution data, and construct distribution uncertain label data; Step S53: generate pseudo-label expansion data by self-supervised pseudo-label expansion simulation on the distribution uncertain label data; perform multi-task standard allocation through the pseudo-label expansion data to obtain multi-task training allocation data; Step S54: task coordination training scheduling is performed according to the multi-task training distribution data, so as to obtain a coordination sample stream; and integrated model joint training is performed based on the coordination sample stream, so as to obtain an integrated image pre-training model.

9. A training system of an image pre-training model, characterized in that, The training system of the image pre-training model comprises: An image acquisition module is configured to acquire original training images, extract semantic information in the images, and perform feature processing to generate structured high-dimensional feature data; and the high-dimensional feature data is sequentially reduced by three different scales to finally generate multi-resolution feature bodies. A texture screening module is configured to analyze texture definition of the images by fusing the feature bodies of different resolutions, and divide the original images into flaw training images and qualified training images according to the texture definition. A flaw modeling module is configured to trace and correct the flaw training images to flaw correction training images, and construct positive-negative contrast images based on the flaw correction training images and the flaw training images; and perform contrast training by combining the positive-negative contrast images, so as to construct a pre-training model of the flaw images. A positive example modeling module is configured to perform weak disturbance transformation based on the qualified training images, and pair the images before and after the transformation one by one to construct positive example pairs before and after the disturbance; find out regions with consistent features from the positive example pairs before and after the disturbance, and use the region images to train a pre-training model of the qualified images. An integrated training module is configured to integrate feature data of the pre-training model of the flaw images and the pre-training model of the qualified images into a coordination sample stream, and perform joint training through the sample stream to obtain an integrated image pre-training model.

10. A computer readable storage medium storing a computer program, characterized in that, The computer program, when executed, implements the training method of the image pre-training model according to any one of claims 1 to 8. The computer program, when executed, implements the training method of the image pre-training model according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Auto-encoder anomaly detection method based on comparative learning

    CN114724043A

  • Abnormity detection method based on twin auto-encoders and bidirectional information depth supervision

    CN116645369A

  • Track defect detection method and device based on pre-training large model

    CN117455902A

  • Image processing model training method and device, electronic equipment and storage medium

    CN118071652A

  • Self-supervised image anomaly detection method based on patch aggregation and discrimination

    CN119131475A

Cited By

  • Method and device for detecting morphology of roller surface

    CN121353813A

  • Battery pack box body welding seam surface defect detection method based on machine vision

    CN121810678A

  • A battery pack box weld surface defect detection method based on machine vision

    CN121810678B

  • Image alignment method and device and storage medium

    CN122223079A