Training method, system and storage medium of image pre-training model

By constructing a training method using multi-resolution feature volumes and positive-negative contrast images, this study addresses the shortcomings of existing image pre-training methods in terms of accuracy and robustness in complex image processing. It achieves efficient defect image recognition and classification, thereby improving the model's adaptability and robustness.

CN120913029BActive Publication Date: 2026-04-28SHENZHEN WRITER INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN WRITER INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-06-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing image pre-training methods lack accuracy and robustness when processing complex images, and lack effective judgment of image texture clarity, leading to misjudgment and omission in the screening process of defective and qualified images. The models are poorly adaptable to subtle changes, making it difficult to achieve efficient image classification and recognition.

Method used

By acquiring original training images, extracting semantic information and generating high-dimensional feature data, constructing multi-resolution feature volumes, fusing texture clarity for image segmentation, constructing positive and negative contrast images of defective and qualified images for comparative training, and combining defective and qualified images for joint training to generate an integrated image pre-trained model.

Benefits of technology

It improves the representational ability of image features, enhances the recognition accuracy of flawed images and the generalization ability of the model, strengthens the model's ability to learn flawed features, optimizes the quality and consistency of training data, and improves the adaptability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913029B_ABST
    Figure CN120913029B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image pre-training, and particularly relates to a training method and system of an image pre-training model and a storage medium. The method comprises the following steps: collecting original training images, performing high-dimensional feature mapping and pyramid transformation to generate multi-resolution feature bodies, determining the texture definition of the images based on the feature bodies, screening out defective training images and qualified training images, for the defective training images, correcting them into defective correction training images by using a trace correction technology, and constructing positive and negative contrast images to train a defective image pre-training model, at the same time, performing weak disturbance transformation on the qualified training images to generate positive example pairs before and after disturbance to train a qualified image pre-training model, integrating the two pre-training models into a cooperative sample stream, and obtaining a high-performance integrated image pre-training model through integrated pre-training. The present application realizes image quality grading and difference modeling, improves the defect recognition precision, and enhances the model generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image pre-training technology, and in particular to a training method, system and storage medium for an image pre-training model. Background Technology

[0002] Existing image pre-training methods often rely on simple feature extraction and classification algorithms, resulting in insufficient accuracy and robustness of the models when processing complex images. Current technologies lack a systematic approach to training data selection and processing, making it difficult to guarantee the diversity and representativeness of training samples, further impacting the model's generalization ability. This is particularly true in the identification and processing of flawed images, where insufficient data and inaccurate feature extraction are common problems. Furthermore, existing models lack effective judgment of image texture clarity, leading to misclassification and missed classification during the screening process between flawed and qualified images, affecting the quality of subsequent training. Especially in complex environments, the models exhibit poor adaptability to subtle changes, making efficient image classification and recognition difficult, further increasing the challenges in practical applications, particularly in industrial inspection and automated processing. Summary of the Invention

[0003] Therefore, it is necessary to provide a training method, system, and storage medium for an image pre-training model to solve at least one of the aforementioned technical problems.

[0004] To achieve the above objectives, a training method for an image pre-training model includes the following steps:

[0005] Step S1: After acquiring the original training images, extract the semantic information from the images and perform feature processing to generate structured high-dimensional feature data; reduce the high-dimensional feature data in three different proportions to finally generate a multi-resolution feature body;

[0006] Step S2: Analyze the texture sharpness of the image by fusing feature volumes at various resolutions, and divide the original image into flawed training images and qualified training images based on texture sharpness.

[0007] Step S3: Tracing and correcting the defective training image to become a defect-corrected training image, and constructing positive and negative contrast images based on the defect-corrected training image and the defective training image; performing contrast training by combining the positive and negative contrast images to construct a pre-trained model of the defective image.

[0008] Step S4: Perform weak perturbation transformation on qualified training images, and pair the images before and after the transformation one by one to construct positive pairs before and after perturbation; find regions with consistent features from the positive pairs before and after perturbation, and use these region images to train the pre-trained model of qualified images.

[0009] Step S5: Combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and perform joint training through this sample stream to obtain the integrated image pre-trained model.

[0010] This invention enhances the representational power of image features through high-dimensional feature mapping and pyramid transformation, enabling multi-resolution feature volumes to better capture subtle texture changes and significantly improving the accuracy of texture clarity determination. Effective screening of defective and qualified training images optimizes the quality of the training dataset, ensuring the reliability of subsequent model training. The defect image source correction technique ensures the accuracy and consistency of the training data, providing the model with authentic and representative samples. The constructed positive and negative contrast images provide rich comparative information for the defective image pre-training model, enhancing the model's ability to learn defective features. The weak perturbation transformation technique gives qualified training images greater diversity, improving the model's generalization ability and robustness. The construction of positive pairs further improves the model's adaptability to image changes. The integration of collaborative sample streams achieves an effective combination of defective and qualified images. The resulting integrated image pre-training model exhibits stronger performance when handling complex image tasks, more accurately identifying and classifying different types of image features, realizing image quality grading and difference modeling, improving defect recognition accuracy, and enhancing model generalization ability.

[0011] The present invention also provides a training system for an image pre-training model, for performing the training method of the image pre-training model as described above, the training system for the image pre-training model comprising:

[0012] The image acquisition module is used to acquire the original training images, extract the semantic information from the images, perform feature processing, and generate structured high-dimensional feature data; the high-dimensional feature data is then scaled down in three different proportions to finally generate a multi-resolution feature body;

[0013] The texture screening module is used to analyze the texture sharpness of an image by fusing feature volumes at various resolutions, and to classify the original image into flawed training images and qualified training images based on the texture sharpness.

[0014] The defect modeling module is used to trace and correct defective training images into defect-corrected training images, and to construct positive and negative contrast images based on the defect-corrected training images and the defective training images; and to perform comparative training by combining the positive and negative contrast images to construct a pre-trained model of defective images.

[0015] The positive example modeling module is used to perform weak perturbation transformation on qualified training images and pair the images before and after the transformation to construct positive example pairs before and after perturbation; it identifies regions with consistent features from the positive example pairs before and after perturbation and uses these region images to train the pre-trained model of qualified images.

[0016] The ensemble training module is used to combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and then perform joint training through this sample stream to obtain the ensemble image pre-trained model.

[0017] This invention, through the application of an image acquisition module, can comprehensively acquire original training images, ensuring the diversity and richness of the data. Feature dimension mapping and pyramid transformation improve the accuracy of image feature extraction. The generated multi-resolution feature volume effectively captures subtle texture information in the image. The texture screening module, based on sharpness judgment, accurately distinguishes between flawed training images and qualified training images, optimizing the quality of subsequent training datasets and improving the reliability of model training. The source correction technology of the flaw modeling module ensures the authenticity of the training data and provides representative samples for the model. By constructing positive and negative contrast images, the learning ability of the flawed image pre-trained model is enhanced, enabling it to more effectively identify and process flawed features. The positive example modeling module expands the sample diversity of qualified training images through weak perturbation transformation, improving the model's generalization ability and robustness. The integrated training module effectively combines the pre-trained models of flawed images and qualified images, realizing collaborative training. The final integrated image pre-trained model exhibits superior performance when handling complex image tasks, accurately identifying and classifying various types of image features, greatly enhancing the system's practicality and adaptability.

[0018] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed, implements the training method of the image pre-training model as described in any of the above claims.

[0019] This invention utilizes computer-readable storage media to efficiently store and manage the training program of an image pre-trained model, ensuring the program's execution efficiency and accuracy, providing stable data access and processing capabilities, supporting the processing and analysis of large-scale image data, promoting the rapid integration of diverse training samples, optimizing the image feature extraction and processing flow, enhancing the flexibility and scalability of model training, ensuring the integrity and consistency of training data, strengthening the model's adaptability and robustness, and ultimately achieving a high-performance image pre-trained model to meet the needs of complex image analysis tasks. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the steps of a training method for an image pre-training model.

[0021] Figure 2 This is a detailed flowchart illustrating the implementation steps of step S2;

[0022] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0024] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0025] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0026] To achieve the above objectives, please refer to Figures 1 to 2 A training method for an image pre-training model includes the following steps:

[0027] Step S1: After acquiring the original training images, extract the semantic information from the images and perform feature processing to generate structured high-dimensional feature data; reduce the high-dimensional feature data in three different proportions to finally generate a multi-resolution feature body;

[0028] Step S2: Analyze the texture sharpness of the image by fusing feature volumes at various resolutions, and divide the original image into flawed training images and qualified training images based on texture sharpness.

[0029] Step S3: Tracing and correcting the defective training image to become a defect-corrected training image, and constructing positive and negative contrast images based on the defect-corrected training image and the defective training image; performing contrast training by combining the positive and negative contrast images to construct a pre-trained model of the defective image.

[0030] Step S4: Perform weak perturbation transformation on qualified training images, and pair the images before and after the transformation one by one to construct positive pairs before and after perturbation; find regions with consistent features from the positive pairs before and after perturbation, and use these region images to train the pre-trained model of qualified images.

[0031] Step S5: Combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and perform joint training through this sample stream to obtain the integrated image pre-trained model.

[0032] This invention enhances the representational power of image features through high-dimensional feature mapping and pyramid transformation, enabling multi-resolution feature volumes to better capture subtle texture changes and significantly improving the accuracy of texture clarity determination. Effective screening of defective and qualified training images optimizes the quality of the training dataset, ensuring the reliability of subsequent model training. The defect image source correction technique ensures the accuracy and consistency of the training data, providing the model with authentic and representative samples. The constructed positive and negative contrast images provide rich comparative information for the defective image pre-training model, enhancing the model's ability to learn defective features. The weak perturbation transformation technique gives qualified training images greater diversity, improving the model's generalization ability and robustness. The construction of positive pairs further improves the model's adaptability to image changes. The integration of collaborative sample streams achieves an effective combination of defective and qualified images. The resulting integrated image pre-training model exhibits stronger performance when handling complex image tasks, more accurately identifying and classifying different types of image features, realizing image quality grading and difference modeling, improving defect recognition accuracy, and enhancing model generalization ability.

[0033] In this embodiment of the invention, the training method of the image pre-training model includes the following steps:

[0034] Step S1: After acquiring the original training images, extract the semantic information from the images and perform feature processing to generate structured high-dimensional feature data; reduce the high-dimensional feature data in three different proportions to finally generate a multi-resolution feature body;

[0035] In this embodiment, images of the target workpiece under standard lighting conditions are acquired using an industrial camera. Each image has a resolution of 2048×2048 pixels. The shooting process maintains a fixed exposure value, a fixed shooting distance, and is free of background clutter. After acquisition, the images are uniformly transcoded into RGB three-channel data format and input into a feature mapping network constructed based on a multi-layer residual structure. The network structure consists of four residual convolutional blocks, each with a kernel size of 3×3 and a stride of 1, using the ReLU activation function. The output is a 2048×2048 pixel image. An 8×64 high-dimensional feature mapping tensor is used as input to the multi-scale pyramid construction module. The construction module sequentially performs pooling layers and variable-size convolution operations on the feature map. The pooling layers adopt max pooling, and the scales are set to the original resolution, 1 / 2, 1 / 4, and 1 / 8, for a total of four layers. After performing a multi-channel convolution operation with a kernel size of 3×3 on each layer, a multi-resolution feature volume is generated. The final multi-resolution feature volume contains multi-scale texture information data tensors at the original image level and at the three compressed scales, which are stored as a multi-layer tensor structure for subsequent processing.

[0036] Step S2: Analyze the texture sharpness of the image by fusing feature volumes at various resolutions, and divide the original image into flawed training images and qualified training images based on texture sharpness.

[0037] In this embodiment, a texture sharpness determination module is used to evaluate the texture quality of each original training image at each scale. This module is built based on the frequency domain sharpness evaluation method. It performs Fast Fourier Transform (FFT) on the feature volume at each scale to extract the response values ​​of high-frequency components in the frequency domain. The frequency domain energy distribution density threshold is defined as 0.25. Images with a high-frequency energy ratio higher than the threshold are marked as sharp texture images, and images with a high-frequency energy ratio lower than the threshold are marked as blurry texture images. The final texture sharpness label of the image is determined by a weighted voting method based on the results of the four scales. The weight coefficients are determined by the original resolution ratio of the scales, with a ratio of 4:2:1:1. After the determination is completed, all images are labeled as "qualified training images" or "flawed training images".

[0038] Step S3: Tracing and correcting the defective training image to become a defect-corrected training image, and constructing positive and negative contrast images based on the defect-corrected training image and the defective training image; performing contrast training by combining the positive and negative contrast images to construct a pre-trained model of the defective image.

[0039] In this embodiment, image region reconstruction processing is performed on the defective training images. The reconstruction processing is based on a region-guided inverse mapping (GRM) convolutional network. A CNN (Neural Network Array) is used to guide the processing of flawed regions by performing low-frequency texture smoothing and structural restoration, outputting a flaw-corrected training image that retains some original flaw traces while restoring texture continuity. Each flawed training image is then paired with its corresponding flaw-corrected image to form a positive-negative contrast image pair, which is then input into a feature contrast extraction network. This network uses a Siamese architecture where two branches share a weight structure. After extracting image features, L2 distance difference is calculated to obtain a difference contrast feature tensor. This difference tensor is input into a salient region localization module, which generates a salient mask image based on an attention mechanism. The mask image has a value range of 0 to 1, representing the saliency of each pixel. After spatial alignment with the flawed and corrected images, the mask image is concatenated along the channel dimension in the order of original image, enhanced image, and mask image, forming a three-channel ternary image. This image is then input into a triplet contrast training network, which uses a Triplet architecture with a TripletLoss loss function. Finally, a pre-trained flawed image model is obtained through backpropagation, and the model outputs an image discriminative feature embedding vector.

[0040] Step S4: Perform weak perturbation transformation on qualified training images, and pair the images before and after the transformation one by one to construct positive pairs before and after perturbation; find regions with consistent features from the positive pairs before and after perturbation, and use these region images to train the pre-trained model of qualified images.

[0041] In this embodiment, a multi-scale block-based perturbation operation is performed on each qualified training image. The block-based method adopts a square window-based block division, dividing each image into 64 32×32 pixel sub-blocks. Subsequently, a local texture data extraction operation is performed on each sub-block using Local Binary Pattern (LBP). The texture response vector of each sub-block is obtained using the pattern method. Then, the texture response vector is deconstructed by directional gradient. The Sobel operator is used to extract gradients in the x and y directions to generate the directional response vector of each sub-block. By comparing the directions with the maximum response intensity, a set of dominant directions is generated. Then, local texture distribution data is generated based on the distribution of the dominant directions of each sub-block in the whole image. This data structure is a combination structure of the dominant direction and intensity value of each sub-block. Finally, the local texture distribution data is used to perform a weak perturbation transformation operation on each qualified image. The perturbation method is Gaussian blur kernel convolution, the blur kernel size is set to 3×3, the standard deviation σ is 0.6, and the perturbation operation is only performed in the region where the dominant texture directions are consistent to ensure that the overall structure of the image is not destroyed. Perturbed qualified training images are generated. Then, the images before and after perturbation are paired one by one to form positive pairs before and after perturbation. These pairs are input into the positive example contrast learning network. The network structure is a symmetric residual network. Cosine Similarity is used as the loss metric to train the pre-trained model of the output qualified images.

[0042] Step S5: Combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and perform joint training through this sample stream to obtain the integrated image pre-trained model.

[0043] In this embodiment, the pre-trained models for flawed images and qualified images are loaded and trained. Using a consistent image input format, the intermediate embedding layer features output by the two models are extracted to construct a sample feature stream. The two models are then applied to the same batch of training images and their respective embedding features are output synchronously. The two embedding feature vectors are fused using channel concatenation to obtain a collaborative sample stream. This feature stream is input into the ensemble pre-trained model in the subsequent training stage. The ensemble model adopts an attention-guided feature weighting mechanism, applying different attention weights to each part of the input feature stream. The final image embedding features are output through a fully connected layer and trained and optimized using a contrastive training strategy. Finally, an ensemble image pre-trained model that can be used for pre-training of multiple types of images is obtained. The output model is saved in standard ONNX format for subsequent inference.

[0044] Preferably, step S1 includes the following steps:

[0045] Step S11: Perform cross-domain sampling on the original image data to obtain a heterogeneous source dataset; perform semantic segmentation mapping on the heterogeneous source dataset to generate a semantic guidance tensor;

[0046] Step S12: Perform high-dimensional feature mapping on the original training image based on the semantically guided tensor. The channel dimension expansion and semantic alignment modules are used to perform nested convolution operations. The convolution kernel size is set to 3×3 and the dilation rate ranges from [1,2,4] to obtain the feature dimension mapping data.

[0047] Step S13: Perform distribution consistency calibration on the feature dimension mapping data to obtain normalized feature space data;

[0048] Step S14: Perform multi-scale pyramid transformation on the normalized feature space data, wherein the number of pyramid scale layers is set to 3, the downsampling factor of each layer is 1, 0.5 and 0.25 respectively, and the channel compression rate is set to 0.5 to obtain multi-resolution feature volume.

[0049] In this embodiment, a heterogeneous source dataset is constructed through cross-domain sampling. This operation involves extracting images from a public industrial image dataset and a custom working condition image set, with 1000 images per class. The public dataset includes MVTec AD and DAGM2007, while the custom image set is collected using industrial cameras in actual production lines. The sampling standard is once every five minutes. For each class of samples, a combination of fixed-angle 45-degree oblique lighting and top lighting is used. The resolution of each image is fixed at 2048×2048 pixels. After sampling, a category label index is established using label matching rules. The sampled images are then input into a pre-trained DeepLabV3+ semantic segmentation network, where the backbone network uses a ResNet-101 structure, and the output layer dimension is fixed at 1 / 4 of the original image size. After being restored to the original image size through interpolation and upsampling, the image is registered with the original image and then stacked with 3 channels to generate a semantic guidance tensor of size 2048×2048×3. This tensor is used to guide downstream high-dimensional feature mapping operations. In the data structure, the semantic guidance tensor is stored in uint8 format, and each pixel value represents its semantic category index number, which is obtained from the inference output of the pre-trained semantic network. The number of semantic categories is set to 13. By constructing a nested convolutional structure and introducing channel dimension expansion and semantic alignment modules, the channel dimension expansion operation adopts Depthwise. The Separable Convolution (SAME) architecture expands the number of feature channels in the original image from the initial 3 channels to 64 channels. The kernel size is uniformly 3×3, the stride is set to 1, and the padding mode is SAME mode to ensure the output size remains unchanged. Three different dilation rates (1, 2, and 4) are used in parallel convolutions. The output is concatenated to form a feature mapping tensor of size 2048×2048×192. Then, a semantic alignment module performs cross-channel alignment by broadcasting the semantic guidance tensor to each channel and performing element-wise multiplication with the feature mapping tensor, thus forming semantically enhanced feature dimension mapping data. This operation relies on GPU parallel computing, using tf.nn in TensorFlow.The `depthwise_conv2d` function implements all nested convolution operations. Convolution weights are initialized using the He Normal strategy, and all weight parameters are updated incrementally through backpropagation. Distribution consistency calibration is performed on the feature dimension mapping data, specifically using distribution normalization to standardize the feature space mapping. First, batch processing is performed on the channel dimension. The normalization operation calculates the mean μ and standard deviation σ for each channel, and then normalizes each pixel value according to (x-μ) / σ, where μ and σ are the channel-wise statistical values ​​across the entire training batch. The batch size is set to 16, and the calculation uses a floating-point tensor with float32 precision. After normalization, all data values ​​are constrained to a range approximating a standard normal distribution. This process is implemented by the torch.nn.BatchNorm2d module in the PyTorch framework. The normalized tensor is called the normalized feature space data, and its size remains 2048×2048×192. Based on this, the next step is to construct the pyramid structure. A three-layer pyramid scale structure is set, with each layer corresponding to a downsampling factor of 1, 0.5, and 0.25, respectively. The first layer maintains the original resolution, and the second layer uses double lines... The first layer is downsampled to 1024×1024×192 using interpolation, and the third layer is downsampled again to 512×512×192. Then, the number of channels in each layer is compressed with a channel compression ratio of 0.5 using a 1×1 convolution kernel. The number of channels in the first layer is compressed from 192 to 96, the second layer from 192 to 96, and the third layer is similarly compressed to 96. This results in three sets of feature tensors with different resolutions but consistent channel counts. Each set of tensors corresponds to a scale layer of the pyramid. These three sets of tensors are encapsulated in a unified format to form a multi-resolution feature volume. This feature volume retains spatial detail and scale variation information, which is used for subsequent texture determination and training image classification. In practice, the entire pyramid structure is constructed using TensorFlow's `tf.image.resize` function for downsampling and the `tf.nn.conv2d` function for 1×1 convolutional channel compression.

[0050] Preferably, step S2 includes the following steps:

[0051] Step S21: Aggregate the multi-resolution feature volumes across scales to obtain fused texture features; perform local texture response mapping on the fused texture features to obtain local texture response distribution data;

[0052] Step S22: Calculate the local structure contrast based on the local texture response distribution data, where the texture window size is 8×8 pixels and the sliding step size is 4 pixels, and infer the image texture clarity based on the local structure contrast;

[0053] Step S23: Perform threshold segmentation on the image texture sharpness based on a preset texture sharpness threshold to generate sharpness segmentation labels;

[0054] Step S24: Filter the original training images into flawed training images and qualified training images by using sharpness segmentation labels.

[0055] In this embodiment, bilinear interpolation upsampling is performed on the multi-resolution feature maps extracted from the three-layer pyramid structure to unify the spatial resolution. In this operation, image matrix interpolation is used to scale the low-resolution feature maps to match the highest resolution layer. Then, the feature maps from the three scale layers are concatenated along the channel dimension to form a channel-expanded fused feature map. Next, a local texture response mapping module is introduced to perform local filtering on each spatial location in the fused feature map. The filtering kernel used is a response template map. Gray-level response intensity analysis is performed on each pixel in the receptive field through image local response computation. In practice, local texture is constructed by calculating the local variance of the image and combining it with changes in response intensity. The response tensor is used to extract the maximum response value of the tensor in the spatial dimension to generate local texture response distribution data. The entire operation does not involve weight training or gradient propagation; it only performs pure forward mapping and response feature value output. The local texture response distribution data is divided into several consecutive, non-overlapping tiled regions. The entire image is scanned sequentially using a fixed texture window size of 8×8 pixels. The sliding window's movement interval is set to 4 pixels horizontally and vertically. In each 8×8 pixel region, the average grayscale response value and standard deviation are first calculated. Then, the response difference between the current region and its eight neighboring regions is calculated one by one to obtain the response gradient difference between the current region and its neighboring regions. Finally, the response gradient is used as the basis for further calculations. The ratio of the difference to the standard deviation represents the local contrast value of the region. Further, by clustering and statistically analyzing the distribution range of contrast values ​​across all local windows, a complete local structure contrast map is constructed. Finally, this contrast map is aggregated and normalized across the entire image according to the channel direction, generating a texture sharpness map in a continuous value domain. The value corresponding to each pixel represents its response to the texture sharpness in the local structure. The texture sharpness map is normalized and scaled to ensure that the numerical range is concentrated between zero and one. Subsequently, a unified static threshold reference line is set as the sharpness boundary point and mapped onto the normalized image. This threshold is derived from the rounded calculation result of the median of the response distribution of a large number of qualified sample images on the local texture contrast map. During specific segmentation, each... The texture sharpness value corresponding to each pixel is compared with a preset threshold. If it is greater than the threshold, the pixel is marked as a sharp region; otherwise, it is marked as a blurry region. Finally, a binary mask is generated based on the sharpness and blurry binary labels of all pixels. This mask is the sharpness segmentation label map. This label map maintains the same spatial resolution as the original image and uses the same coordinate system for alignment. The sharpness segmentation label map for each image is obtained, and the ratio of the number of sharp labels to the number of blurry labels for all pixels in the label map is counted. Based on this ratio, a classification judgment standard is set. If the proportion of sharp regions in an image is lower than a uniformly set threshold, the image is assigned to the defective training image set; otherwise, it is assigned to the qualified training image set.The threshold for this proportion was determined based on an empirical dividing line calculated from 1000 manually labeled images during a manual screening experiment. To ensure the accuracy of the grouping results, after classification, the training images and their corresponding label images were packaged separately according to the grouping and stored in different training directory paths. The set of flawed training images was used for subsequent texture enhancement pre-training tasks, while the set of qualified training images was used for the standard feature extraction training path to improve the robustness and texture understanding ability of the overall image pre-training model.

[0056] Preferably, step S3, which involves correcting the defective training image to a defect-corrected training image and constructing a positive-negative contrast image based on the defect-corrected training image and the defective training image, includes:

[0057] Extract defect features from defective training images and crop defective regions from the training images based on these features;

[0058] By using defect features, defect source processing is performed on defect area images to obtain image defect source data;

[0059] Simulate source defect avoidance for image defect source data to generate simulated defect source avoidance data;

[0060] By simulating defect source avoidance data, defect area images are processed to correct defects, resulting in defect correction training images.

[0061] The defect correction training image and the defect training image are projected into the same space and image comparison processing is performed to obtain positive and negative contrast images.

[0062] In this embodiment, a training image set labeled as defective images by sharpness segmentation is loaded, and feature cross-coding is performed by combining grayscale information and texture response maps in the image channels. Specifically, the image grayscale map, sharpness mask map, and local contrast map are uniformly constructed into a three-channel composite feature map in a channel stacking manner. A differential feature extraction module based on edge enhancement convolution kernels is called in this composite map. This module contains a 5×5 convolution kernel group to amplify edge abrupt regions and suppress the response of continuous smooth regions. After obtaining the feature enhancement response map by sliding across the entire image through this convolution kernel group, a region growing method is further used to lock the boundaries of the regions in the feature set. After the boundaries are closed, the region is cropped and extracted from the original image in the form of a mask to obtain each The image of the defective region in the image is processed by a fixed context expansion strategy during the cropping operation. A 6-pixel boundary is added around the cropped region to preserve the surrounding semantic context information. The output size is uniformly adjusted to 64×64 pixel resolution for subsequent defect tracking operations. Each cropped defective region image is input into a pre-constructed source tracing network structure based on an expanded receptive field convolutional stacking structure. This structure contains two cascaded encoder-decoder modules and an explicit positional attention module. When the defective image is input, it first passes through three convolutional layers of sizes 3×3, 5×5, and 7×7 to extract feature maps at different scales. Then, the explicit positional attention module weights the feature response weights at each spatial location. This module uses a two-dimensional... The attention response factor for each pixel is calculated by combining coordinate embedding and channel attention. A dot product mapping operation is then performed on the original feature map to generate a positional calibration feature map. Finally, deconvolution reconstruction is performed using the contextual semantic features of the original image. The reconstruction output is spatial coordinate image data containing the defect generation starting point, edge expansion trajectory, and explicit defect texture change path. This data constitutes the source data of image defects. All coordinates are uniformly mapped and output based on the original image space. The coordinate information and structural path from the source tracing results are read and input into a defect avoidance generator based on guided edge deformation repair. This generator consists of two sub-networks: an edge curvature reconstruction module and a morphological consistency constraint module. The former constructs a fitting function based on the input path data. A wireframe diagram of the defect boundary morphology is generated. Bézier curves are used to fit all discontinuous path segments and construct a continuous differentiable boundary curve. The latter uses the wireframe diagram as a mask to reconstruct and constrain the original image. In the morphological consistency constraint module, a specific regularization loss is used to impose triple constraints on the gray-level continuity, edge smoothness, and texture homogeneity of the reconstructed region. The reconstruction result maintains edge alignment and texture coherence with the original background region. All simulated avoidance images are verified by the adversarial example detection module and then discarded to generate unstable images. Only avoidance data with texture boundary consistency greater than a threshold are retained as the final simulated defect source avoidance data output. Texture boundary consistency is calculated by the image structure similarity index SSIM (Structural Similarity Index) and the minimum threshold is set to 0.82. An image inpainting inference network is used to perform fusion interpolation reconstruction. First, the defect region localization map in the original image is used as an input mask to guide the input into the reconstruction network. At the same time, the source avoidance data and the grayscale feature map of the original image are concatenated in the channel dimension as feature input. The entire image inpainting inference network is based on the U-Net structure, fusing skip connection features to preserve the image context structure information. The complete texture region is restored through multi-scale decoding. In this process, coarse semantic filling is performed at the low resolution stage, and texture detail supplementation is performed at the high resolution stage to ensure a seamless transition between the repaired area and the surrounding image structure. The final output is a corrected image with the same size as the original image. This image is the defect correction training image. During the storage process, the numbering is kept consistent with the original image and it is labeled as a cleaned image sample. Image embedding based on a shared feature encoding structure is used. The network first extracts embedding vectors from two images using a weight-shared encoder module. This encoder is a residual convolutional network containing four convolutional blocks, each consisting of two 3×3 convolutional layers and one 2×2 max-pooling layer. The embedding vector is the 512-dimensional global feature vector output by the encoder for each image. Then, the two sets of embedding vectors are input into a Euclidean distance module to calculate the distance matrix. Simultaneously, image coordinate-level alignment features are preserved for pixel-level difference analysis. All image pairs are labeled with the same image number, constructing an image pair data structure where the defective image is the negative sample and the corrected image is the positive sample. Finally, the two images are stitched together to form a single-channel contrast image. The left side is the defective training image, the right side is the corrected image, and the middle displays a pixel-level difference heatmap in color-coded form, forming a positive and negative contrast image, which is uniformly stored in the contrast image training set path.

[0063] Preferably, step S3, which combines positive and negative contrast images for comparative training to construct a pre-trained model of the defective image, includes:

[0064] Extract the difference contrast features from positive and negative contrast images;

[0065] The salient regions of the image are located based on the difference contrast features, and a contrast salient mask is determined based on the salient regions of the image.

[0066] Spatially align the contrast saliency mask, the defect training image, and the defect correction training image, and then perform channel stitching to generate the original-enhanced-saliency mask ternary image.

[0067] Extract ternary image feature data from the original-enhanced-salient mask ternary image;

[0068] The triple image feature data is input into the contrast learning network to perform triple contrast training, thereby generating discriminative contrast features;

[0069] A pre-trained model of flawed images is trained using discriminative contrast features.

[0070] In this embodiment, the defect correction training image and the corresponding defect training image are both normalized in size. Bilinear interpolation is used to adjust the images to a uniform resolution of 256×256 pixels. Then, a pixel-by-pixel comparison is performed on the two types of images using the Structural Similarity Index Measure (SSIM). With a sliding window size of 11×11 and a window stride of 1, the brightness, contrast, and structural differences of local regions are calculated to obtain a difference matrix. Pixel regions with a difference value greater than 0.25 in the difference matrix are further extracted as difference contrast feature regions. The final output difference contrast features are expressed as a binary image, where a pixel value of 1 indicates a significant pixel region with structural differences, and a pixel value of 0 indicates a region with no differences. The generated binary image is used as the initial mask input to a fully convolutional network (FCN). In the saliency detection network of the Convolutional Network (CNN), the network structure consists of five convolutional layers and three upsampling layers. ReLU activation and BatchNorm normalization are used in the convolutional layers to accelerate convergence and suppress overfitting. After processing, the input image outputs a 256×256 saliency heatmap. A threshold segmentation method is then used to extract salient regions, with a saliency threshold set to 0.3. A saliency threshold is set when the pixel value in the heatmap is greater than 0.At point 3, the salient region is defined. The threshold segmentation result is then subjected to morphological closing to repair edge breaks, followed by opening to remove isolated noise points. The resulting salient region mask is the contrast saliency mask. Regions with a pixel value of 1 in the mask are used for subsequent spatial alignment and channel stitching. Affine transformations are performed on the three types of images to unify the spatial reference frame. The affine matrix is ​​obtained by calculating the image center point and the geometric center point of the salient region. The affine transformation matrix is ​​generated using the `getAffineTransform` function in OpenCV, and then the `warpAffine` operation is performed on the image to obtain image data in a unified space. After ensuring image alignment, the three types of images are stacked according to the channel dimension. The defect training image is used as the first channel input, the defect correction training image as the second channel input, and the contrast saliency mask as the third channel input. Finally, a three-channel image data of size 256×256×3 is generated, called the original-enhanced-salient mask ternary map. For example, the pixel data of the ternary image is stored in float32 format and normalized to the [0,1] interval to adapt to the input requirements of the neural network. A set of shallow convolutional neural network models is constructed to extract feature information. This network structure contains three convolutional layers, each with a kernel size of 3×3, 32, 64, and 128 channels respectively, a stride of 1, and same padding. After convolution, BatchNorm normalization and ReLU activation function are applied. Max pooling is added after each convolution to compress spatial information, with a pooling window of 2×2. Finally, the convolution output is flattened into a one-dimensional vector through the Flatten operation, and further feature projection processing is performed using two fully connected layers. The first fully connected layer has an output dimension of 512, and the second fully connected layer has an output dimension of 128. This 128-dimensional vector is the ternary image feature data. The feature vector is normalized using L2 regularization to stabilize subsequent contrastive learning training. A set of Triplet layers is constructed. A contrastive learning network with a Loss (triple loss) structure is used. The Anchor input consists of triple image features corresponding to the flawed training image, the Positive input consists of triple image features corresponding to the flaw-corrected training image, and the Negative input consists of triple image features randomly sampled from different image batches. Training is performed using a triple loss function with a margin of 0.5. The Loss function is set to the form max(D(anchor,positive)-D(anchor,negative)+margin,0), where D is the Euclidean distance calculation function. Model optimization is performed with an initial learning rate set to 0.001, batch size set to 64, training epochs set to 120. During training, intra-batch hard example mining is performed on Anchor and Negative features, that is, selecting the Negative sample with the smallest distance to the Anchor from the current batch to form a hard negative pair to enhance the model's discriminative ability. After training, the output is a contrastive learning embedding model containing 128-dimensional contrastive feature representations. Using the aforementioned trained contrastive learning embedding model as a foundation, the extracted discriminative contrastive features are used as supervision signals input to the backbone network. The backbone network structure adopts ResNet18, and the final classification layer is removed from the original ResNet18 structure and a layer with an output of 12 is added. An 8-dimensional fully connected layer is used for training. The mean squared error between discriminative contrastive features and the backbone network output features is used as the supervised loss function. During training, the contrastive learning embedding model parameters are frozen, and only the ResNet18 network parameters are updated. The AdamW optimizer is used with an initial learning rate of 0.0005 and a regularization parameter set to 1e-4. The training data uses the previously generated ternary image data, with a batch size of 32 and 80 training epochs. The final output model is used for pre-training representation construction of flawed image features. All model weights are stored in float32 format. After training, the model parameters are exported as an ONNX format file for subsequent model deployment.

[0071] Of particular importance is the localization of salient regions in an image based on contrast features, and the determination of a contrast saliency mask based on these salient regions, including:

[0072] Channel focusing processing is performed based on the response amplitude distribution of each channel in the difference comparison features to construct a focused difference feature map;

[0073] Spatially saliency projection is performed on the focused difference feature map to obtain the initial saliency response map;

[0074] Regional expansion and fusion are performed based on the initial salient response map to generate a salient candidate region map;

[0075] The boundaries of the salient candidate region map are refined to obtain the salient regions of the image;

[0076] Contrast-salience masks are constructed using salient regions of the image.

[0077] In this embodiment, a channel-weighted module is used to perform channel-by-channel response analysis on the input difference contrast feature tensor. The shape of the difference contrast feature tensor is set to [C, H, W], where C is the number of channels, and H and W are the height and width of the feature map, respectively. First, the response value of each channel is mapped to a one-dimensional vector using the Global Average Pooling method. Then, the one-dimensional response vectors of all channels are normalized using the minus-max normalization method, i.e., subtracting the minimum value of each channel from the response value and then dividing by the difference between the maximum and minimum values. The normalized response is used to generate a channel attention weight map. Then, the channel attention weight map is multiplied channel-by-channel with the original difference contrast features to construct a focused difference feature map. This map focuses more on the channel component responses containing obvious difference regions in terms of semantic expressive power. A two-dimensional convolutional neural network is used. The 1×1 convolutional kernel in the network performs channel fusion processing on the focal difference feature map, reducing the channel dimension to 1 to form a single-channel spatial map. The weights of this convolution operation are generated using the He initialization method and fixed to the pre-training state to maintain the stability of channel fusion. Then, a sliding window max pooling operation is applied to the spatial dimension of the reduced image. The window size is set to 5×5, the stride is 1, and the padding value is 0. This processing is used to enhance the response value of the focal region and improve the spatial contrast. The final generated initial salient response map has a size of [1, H, W], which retains the high response region with spatial saliency. The pixel values ​​of the initial salient response map are normalized to 0 to 1. In this process, a threshold segmentation method is used to determine the starting point of the candidate region. The threshold is set to 0.6, and all pixels with positions greater than this threshold are identified as seed point regions. Then, the seed point regions are subjected to region expansion processing based on morphological dilation operation. The structuring element is set to a 3×3 cross-shaped kernel, and the number of iterations is 2. After each expansion, a region averaging and fusion operation is performed, that is, the pixel values ​​in each connected region are averaged and assigned to the entire region to smooth the response differences. After expansion and fusion, a salient candidate region map is generated, and the output is a binary image of the same size as the original image. The salient candidate region map is subjected to Gaussian blur processing, and the convolution kernel size is set to 5×5 with a standard deviation of 1.2. After removing noisy edges, the Canny operator is used for edge detection, with a low threshold of 50 and a high threshold of 150. The Canny edge detection results are used to extract the region contour boundaries. Then, continuous closed regions are generated through edge point fitting using least-squares polygon approximation. The bounding edges after contour fitting are used to crop the candidate region map. The resulting salient region is a mask within the closed contour. This mask only contains the connected response structure within the focal region. The salient region map is converted to a mask format, defining pixels with response values ​​greater than 0 as 1 and others as 0, forming a single-channel binary mask image. The mask image is then normalized to maintain the same size as the input and enhanced images; the corresponding operation is bilinear interpolation. The interpolation resampling method is followed by a mask alignment operation with the input image channels, copying the mask to three channels to form a 3-channel mask image. This mask image is used for subsequent concatenation with the original and enhanced images. This saliency mask image is the contrast saliency mask, possessing spatial accuracy and semantic contrast feature consistency. During the training phase, it can be directly used as the auxiliary input channel for the saliency contrast module. The mask values ​​are only 0 or 1, excluding grayscale distribution areas, ensuring binary input standardization.

[0078] Preferably, step S4 involves performing a weak perturbation transformation based on the qualified training images and pairing the images before and after the transformation one by one to construct positive pairs before and after the perturbation, including:

[0079] The qualified training images are processed by multi-scale block segmentation, and local texture data in each sub-block is extracted;

[0080] Calculate the local texture distribution data of the local texture data;

[0081] Weak perturbation transformation is applied to qualified training images based on local texture distribution data to generate perturbation qualified training images;

[0082] The perturbation-compliant training images are paired one by one with the compliant training images to generate positive pairs before and after the perturbation.

[0083] In this embodiment, the size of the qualified training images is normalized to a uniform scale of 512×512 pixels. The original images are standardized using linear interpolation. Then, the images are divided into blocks according to a multi-scale setting, with block scales of 32×32, 64×64, and 128×128. At each scale, a non-overlapping sliding window method is used to sequentially scan the image region and extract sub-block content. Each scale corresponds to a different fine-grained structure extraction objective. A gray-level co-occurrence matrix (GLCM) is applied to each sub-block. The Matrix (GLCM) method is used to extract local texture data. During the GLCM extraction process, the gray levels are set to 256, and four texture parameters—energy, contrast, entropy, and correlation—are calculated for each sub-block at 0°, 45°, 90°, and 135°. The texture data of all sub-blocks are combined to form the complete texture feature set of the image at different scales. This complete set is used to characterize the texture composition and spatial distribution of local regions in the image. The local texture parameter matrix extracted at each scale is independently normalized using maximum value normalization, i.e., dividing each dimension's texture feature parameter by its maximum value in the image sub-block set to obtain a normalized texture feature tensor ranging from 0 to 1. Then, a texture probability distribution model is constructed at each scale. This model is achieved by histogram estimation of each normalized texture parameter. The histogram bins are set to 32, precisely dividing the probability distribution interval of each type of texture parameter. The number of normalized texture values ​​within each bin is counted, and their proportion is calculated, thereby generating a texture feature distribution vector. The distribution vectors at all scales are concatenated to form a multi-scale texture distribution vector, which serves as the statistical basis for constructing the subsequent perturbation function. The dimension of this vector is equal to the texture feature dimension multiplied by the number of scales and the number of bins. Based on the multi-scale texture distribution vector constructed in the previous step, a distribution consistency constraint is introduced in the perturbation transformation. First, the original image is enhanced by perturbation in the Fourier frequency domain. The image is transformed to the frequency domain representation using Discrete Fourier Transform (DFT). The perturbation frequency band is set as the low-frequency region, and the frequency range is the spectral block within a 16-pixel range extending outward from the center of the image in the frequency domain. Perturbation noise is randomly added to this region. This noise follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.03 in the amplitude direction, while maintaining the original value in the phase direction. Then, the inverse discrete Fourier transform (IDFT) is used to restore the perturbated image in the frequency domain to the spatial domain. The low-frequency texture in the perturbed image undergoes slight changes, but the structure remains intact. Next, the local texture distribution vector of the perturbated image is calculated, and it is compared with the original distribution vector using cosine similarity. The similarity threshold is set to 0.92. If not satisfied, the perturbation noise is regenerated and the above frequency domain perturbation process is repeated until the texture distribution consistency is satisfied. The finally generated perturbation image is the qualified perturbation training image. An index numbering system is constructed for all qualified training images, each image is assigned a unique number, and the corresponding perturbation image is kept consistent with the original image number. In the pairing process, a mapping dictionary is constructed with the number as the key value. The original image and the perturbation image form a key-value pair structure. The one-to-one correspondence between the original image and its perturbation image is realized by calling. The paired images are stored in tensor form. Each pair of images is combined to form a size of [2]. A four-dimensional tensor [C, H, W] is used, where C is the number of channels (set to 3), H and W are the image height and width (both 512), respectively. This tensor is input into the contrastive learning module of the training network as positive pairs. During the data loading phase, it is assembled into batch tensors through batch read operations. The batch size is 16, meaning that 16 pairs of images before and after perturbation are input for each training iteration. Through an indexing mechanism, the strict pairing relationship between the images before and after perturbation is maintained without affecting the image data structure itself, ensuring that the texture distribution and structural feature correlation between the positive pairs are fully preserved.

[0084] Of particular importance is that the local texture distribution data used to calculate local texture data includes:

[0085] The directional gradient is deconstructed from the local texture data to obtain the directional response vector;

[0086] The direction of maximum response intensity in the statistical directional response vector;

[0087] Spatial aggregation is performed based on the direction of maximum response intensity, and the frequency data of directional distribution are analyzed;

[0088] The dominant texture orientation distribution data is determined by the orientation distribution frequency data and the direction of maximum response intensity.

[0089] The local texture data is characterized by texture intensity based on the distribution data of the dominant texture direction, thereby obtaining the texture response intensity matrix;

[0090] Stability-weighted reconstruction is performed based on the texture response intensity matrix to generate local texture distribution data.

[0091] In this embodiment, the original image is divided according to a fixed scale, with the image block size set to 64×64 pixels. Each sub-block is processed independently as a processing unit. Within each sub-block, the pixel grayscale values ​​are convolved using the Sobel operator. A horizontal Sobel kernel [-1,0,1; -2,0,2; -1,0,1] and a vertical Sobel kernel [-1,-2,-1; 0,0,0; 1,2,1] are applied to the image grayscale data to calculate the gradient magnitude of each pixel in the x and y directions. Then, the arctangent function is used to calculate the direction angle θ, where θ is equal to the arctangent of the y-direction gradient value divided by the x-direction gradient value. The range of the direction angle θ is adjusted to 0 to 180 degrees. The mapping is completed within the closed interval. The orientation angle corresponding to each pixel is recorded as the orientation response value of that point. All orientation response values ​​are statistically analyzed along the sub-block dimension, forming a response vector containing 4096 angle values. This vector is the orientation response vector of the image sub-block. The gradient intensity corresponding to each pixel orientation angle θ in the orientation response vector is paired and bound. The intensity value is the square root of the sum of squared magnitudes calculated by Sobel convolution, i.e., the intensity value of each pixel is the square root of the sum of the squared gradient values ​​in the x and y directions. A histogram is constructed for each orientation angle θ, with an angle division granularity of 10 degrees, resulting in a total of 18 orientation intervals. The intensity values ​​corresponding to pixels falling within each orientation interval are accumulated. The intensity sum corresponding to each interval is vectorized to obtain a direction intensity vector of length 18. The index of the element with the largest intensity value in this vector is searched, and the midpoint of the direction interval corresponding to this index is defined as the principal direction angle of that sub-block. The principal direction angle takes the value of {5 degrees, 15 degrees, 25 degrees, ..., 175 degrees}. A direction consistency aggregation kernel is constructed in each image sub-block centered on the principal direction angle, with a kernel size of 7×7 pixels. The angle difference between each pixel and the principal direction angle is calculated, with an allowable difference range of ±15 degrees. Pixels that meet the angle difference condition are aggregated, and the frequency of occurrence of each direction angle is counted within the aggregation region. The frequency is calculated by dividing the number of pixels falling within the corresponding direction angle interval by... The total number of pixels in the aggregated region is statistically analyzed to form an 18-dimensional orientation frequency vector. This vector characterizes the local orientation consistency of the sub-block under the guidance of the main orientation angle. Each dimension of the orientation frequency vector has a value between 0 and 1, and the sum of the dimensions is 1. This vector provides the foundation for subsequent dominant texture orientation recognition. Combined with the orientation frequency vector obtained in the previous step and the main orientation angle, a dominant orientation distribution description vector is constructed. This vector has 18 dimensions and is constructed by setting the interval containing the main orientation angle as the weight center. The frequency values ​​in this interval and its two adjacent left and right orientation intervals are weighted and amplified. The weighting coefficients are set as follows: the weight of the center interval is 1.5, the weight of each of the two left and right intervals is 1.2, and the weight of the remaining orientation dimensions is set to 1.0. After multiplying all dimensions by their corresponding weights and renormalizing, the sum of the vector values ​​is kept to 1. This normalized directional frequency-weighted vector is the dominant texture direction distribution data of the image sub-block. This data describes directional consistency while strengthening the dominant feature of the main direction angle. Each pixel is mapped and weighted according to the frequency value of its directional angle interval in the dominant texture direction distribution data. Specifically, the directional angle θ of the pixel is read, it is assigned to the nearest directional interval, and the value of this directional interval in the dominant texture direction distribution vector is obtained as the directional response weight of the pixel. Then, this weight is multiplied by its gradient intensity value to generate a new weighted response value. The weighted response values ​​of all pixels in the image sub-block are combined into a matrix. The dimension of this matrix is ​​the same as the size of the image sub-block, i.e., 64×64. This matrix is ​​called the texture response intensity matrix. The texture response intensity matrix is ​​used to characterize the directional consistency response intensity distribution of each pixel in the sub-block. Edge filtering is performed on the texture response intensity matrix using bilateral filtering. The texture response intensity matrix is ​​denoised and smoothed using a filtering method. The Gaussian kernel parameter in the spatial domain is set to σs = 3, and the Gaussian kernel parameter in the intensity domain is set to σr = 0.1. The resulting smooth texture response matrix is ​​then output. Next, regional stability analysis is performed on this matrix. The structure tensor method is used to calculate the local stability metric. A structure tensor is constructed for each pixel, and its eigenvalues ​​and eigenvectors are calculated. If the largest eigenvalue is significantly greater than the smallest eigenvalue, the point is considered a stable texture point. Within each sub-block, all stable texture points are counted, and their average response value is calculated. This average value is used to construct a weighted mask, which is applied to the smoothed response matrix to retain the response values ​​of stable regions and suppress those of unstable regions. Finally, the local texture distribution data of the image sub-block is output. The outputs of all sub-blocks are stitched together to form the local texture distribution map of the entire image.

[0092] Preferably, step S4, which involves identifying regions with consistent features from the positive example pairs before and after the perturbation and using these region images to train a pre-trained model for qualified images, includes:

[0093] Extract the feature data of each layer of positive example pairs in the positive example pairs before and after the perturbation;

[0094] Consistency assessment of feature data is performed on positive examples at each layer to generate feature consistency distribution data;

[0095] Highly consistent regions are selected based on a preset consistency distribution threshold and characteristic consistency distribution data.

[0096] The corresponding region images of qualified training images are decomposed by high consistency regions;

[0097] A qualified image pre-training model is trained using images of the corresponding regions.

[0098] In this embodiment, the pre-trained image model uses a convolutional neural network (CNN) structure to perform forward propagation on the two positive example images before and after perturbation. Specifically, the input image size is uniformly adjusted to 224×224 pixels, and pixel values ​​are standardized to between 0 and 1. Feature maps at different levels are extracted sequentially through multiple convolutional layers, batch normalization layers, and activation function layers (such as ReLU). The extracted feature maps include shallow edge information, mid-level texture features, and deep semantic features. The shallow feature map size is 56×56×64, the mid-level feature map size is 28×28×128, and the deep feature map size is 14×14×256. These feature maps are obtained from the corresponding levels of the positive example images before and after perturbation, and saved as multidimensional tensor data for subsequent consistency evaluation. Cosine similarity is used. Similarity is used as a consistency evaluation metric. First, the feature vectors of corresponding positions in the same layer before and after the perturbation are expanded into one-dimensional vectors along the channel dimension. The dot product of the two one-dimensional vectors is calculated and divided by the product of their respective Euclidean norms to obtain the cosine similarity value at that position. The value range is limited to -1 to 1. To avoid the influence of negative values, negative values ​​are uniformly truncated to 0. The calculation result forms a two-dimensional consistency distribution matrix corresponding to the feature map size. For example, for a 28×28 feature map, a 28×28 consistency matrix is ​​output. Each element in the matrix corresponds to the similarity score of the feature vector at that position before and after the perturbation. This operation is repeated for all levels, and finally a multi-level consistency distribution matrix set is generated. This set is stored as floating-point tensor data to filter high consistency regions. The consistency distribution threshold is set to 0.85. Each layer of the consistency distribution matrix is ​​traversed, and elements in the matrix greater than or equal to 0 are selected.The element positions of 85 are marked as high-consistency regions. Binarization is used to assign a value of 1 to positions above the threshold and a value of 0 to positions below the threshold. Then, spatial clustering is performed on the binarized matrix to select continuous connected regions. The area of ​​each connected region is calculated, and noise regions with an area less than 16 pixels are filtered out. The remaining connected regions are considered valid blocks of the high-consistency region for that layer. For the feature maps of different layers, combined with the corresponding image spatial scale, an upsampling method is used to map the high-consistency regions back to the original input image coordinate system, forming a complete high-consistency region mask. The mask data is used for image decomposition operations in the corresponding regions. Based on the mapped high-consistency region mask, image segmentation is performed on qualified training images. The image size is uniformly 224×224 pixels. Pixel regions with a value of 1 in the mask are extracted as regions of interest. A mask-based cropping method is applied to extract the smallest rectangular bounding box region containing the complete connected regions. The bounding box coordinates are calculated based on the mask pixel distribution. The image content of the domain is used as the corresponding region image after decomposition. If the bounding box size is less than 112×112 pixels, zero-value pixels are padded around the bounding box to achieve that size, ensuring consistent input size for subsequent training batches. The cropped region images are saved in standard RGB three-channel format as lossless PNG files. All decomposed region images constitute the training sample set. Data augmentation processing is performed on the decomposed region image sample set, including random horizontal flipping, random rotation ±15 degrees, and random color jitter. Augmented data batches are input into the pre-trained model in units of 64 images. The model structure is based on a ResNet50 deep residual network, and the cross-entropy loss function is used. During training, batch normalization layers and dropout layers with a dropout rate of 0.3 are set. After each epoch, the accuracy and loss are calculated using the validation set. The number of training iterations is set to 100 epochs. After training, the model weight file and network parameters are saved.

[0099] Preferably, step S5 includes the following steps:

[0100] Step S51: Extract the output feature vectors of the defective image pre-trained model and the qualified image pre-trained model, and perform complementary feature mapping based on the output feature vectors to obtain dual-modal feature fusion data;

[0101] Step S52: Re-estimate the category distribution based on the dual-modal feature fusion data to obtain the category attribution distribution data; filter the distributionally uncertain data in the dual-modal feature fusion data based on the category attribution distribution data, and construct the distributionally uncertain label data;

[0102] Step S53: Perform self-supervised pseudo-label augmentation simulation on the uncertain label data to generate pseudo-label augmented data; perform multi-task standard allocation using the pseudo-label augmented data to obtain multi-task training allocation data;

[0103] Step S54: Perform task collaborative training scheduling based on multi-task training allocation data to obtain a collaborative sample stream; perform joint training of the ensemble model based on the collaborative sample stream to obtain an ensemble image pre-trained model.

[0104] In this embodiment, defective and qualified images are input into their respective pre-trained models. Both models use the ResNet50 architecture, and the input image size is uniformly 224×224 pixels. After processing through convolutional, pooling, and residual modules, output feature vectors are extracted from either the fully connected layer or the global average pooling layer. The feature vector dimension is set to 2048. For the same image pair, two feature vectors with the same dimension are obtained. Then, a feature mapping function is used to perform complementary mapping between the two feature vectors. The mapping process employs bilinear pooling. The pooling operation involves performing an outer product of two vectors to obtain a high-dimensional feature matrix. This matrix is ​​then reduced to a 1024-dimensional fusion vector using low-rank decomposition. The fusion vector is normalized so that each dimension's value is distributed between 0 and 1, forming bimodal feature fusion data. A multi-class Softmax classifier is used for forward propagation calculation on this bimodal feature fusion data, outputting the predicted probability distribution for each class. This probability distribution forms the class attribution distribution data. The number of classes is set to 10 according to the task, with probability values ​​ranging from 0 to 1 and summing to 1. Based on this distribution, an entropy index is calculated as a measure of uncertainty. The entropy formula uses the standard discrete entropy formula; a higher entropy value indicates a higher uncertainty. The more uncertain the distribution, the more likely the fused feature data with an entropy value greater than 0.7 is to be labeled as data with uncertain distribution. Subsequently, the corresponding predicted category labels are constructed together with the entropy data to form data with uncertain distribution. The label format adopts one-hot encoding and is stored in a tensor data structure with entropy weights. For data with uncertain distribution, a data augmentation strategy is first introduced to perform various transformations on the corresponding image, including random cropping, rotation ±10 degrees, color transformation, and Gaussian noise superposition. After augmentation, the image is input into the model to re-predict the category probability. Pseudo-labels are generated by combining the original label and the augmented predicted label through a consistency regularization method. The pseudo-labels are selected using a probability thresholding method, with the highest probability exceeding 0.The category of 8 is used as the pseudo-label category. The pseudo-label data is merged with the original label data to form an expanded dataset. Then, training tasks are divided according to the task objective, including classification tasks, adversarial training tasks, and feature reconstruction tasks. Based on label attributes and feature distribution, tasks are assigned to the expanded dataset. A multi-task learning framework is used to distribute the data to different task sub-networks, forming multi-task training allocation data. The data format includes input features, label information, and task identifiers to ensure that the data can be correctly routed to the corresponding sub-task module during training. A task scheduler is built, and the multi-task training allocation data is input. The task scheduler dynamically adjusts the scheduling order of training samples according to task priority and resource consumption. The scheduling strategy adopts a round-robin weighted method to ensure that the samples of each task are evenly distributed within the training cycle. The training samples flow through... After scheduling, the data are input into the corresponding task sub-networks. Each sub-network structure includes a shared backbone network and a dedicated task head. Gradient accumulation technology is used to optimize the multi-task training process. During training, the classification loss, adversarial loss, and reconstruction loss are weighted and summed using a joint loss function with weight parameters set to 0.5, 0.3, and 0.2, respectively. The optimizer used is AdamW, with a learning rate of 0.00005 and 150 training iterations. After training, the weights of each task sub-network are integrated into the model using a weighted average method. The weights are allocated based on the task validation accuracy. The final generated integrated image pre-trained model has multi-modal feature fusion and multi-task collaborative capabilities. The weight parameter file and model structure configuration are stored in a dedicated format file for subsequent image recognition and classification tasks.

[0105] The present invention also provides a training system for an image pre-training model, for performing the training method of the image pre-training model as described above, the training system for the image pre-training model comprising:

[0106] The image acquisition module is used to acquire the original training images, extract the semantic information from the images, perform feature processing, and generate structured high-dimensional feature data; the high-dimensional feature data is then scaled down in three different proportions to finally generate a multi-resolution feature body;

[0107] The texture screening module is used to analyze the texture sharpness of an image by fusing feature volumes at various resolutions, and to classify the original image into flawed training images and qualified training images based on the texture sharpness.

[0108] The defect modeling module is used to trace and correct defective training images into defect-corrected training images, and to construct positive and negative contrast images based on the defect-corrected training images and the defective training images; and to perform comparative training by combining the positive and negative contrast images to construct a pre-trained model of defective images.

[0109] The positive example modeling module is used to perform weak perturbation transformation on qualified training images and pair the images before and after the transformation to construct positive example pairs before and after perturbation; it identifies regions with consistent features from the positive example pairs before and after perturbation and uses these region images to train the pre-trained model of qualified images.

[0110] The ensemble training module is used to combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and then perform joint training through this sample stream to obtain the ensemble image pre-trained model.

[0111] This invention, through the application of an image acquisition module, can comprehensively acquire original training images, ensuring the diversity and richness of the data. Feature dimension mapping and pyramid transformation improve the accuracy of image feature extraction. The generated multi-resolution feature volume effectively captures subtle texture information in the image. The texture screening module, based on sharpness judgment, accurately distinguishes between flawed training images and qualified training images, optimizing the quality of subsequent training datasets and improving the reliability of model training. The source correction technology of the flaw modeling module ensures the authenticity of the training data and provides representative samples for the model. By constructing positive and negative contrast images, the learning ability of the flawed image pre-trained model is enhanced, enabling it to more effectively identify and process flawed features. The positive example modeling module expands the sample diversity of qualified training images through weak perturbation transformation, improving the model's generalization ability and robustness. The integrated training module effectively combines the pre-trained models of flawed images and qualified images, realizing collaborative training. The final integrated image pre-trained model exhibits superior performance when handling complex image tasks, accurately identifying and classifying various types of image features, greatly enhancing the system's practicality and adaptability.

[0112] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed, implements the training method of the image pre-training model as described in any of the above claims.

[0113] This invention utilizes computer-readable storage media to efficiently store and manage the training program of an image pre-trained model, ensuring the program's execution efficiency and accuracy, providing stable data access and processing capabilities, supporting the processing and analysis of large-scale image data, promoting the rapid integration of diverse training samples, optimizing the image feature extraction and processing flow, enhancing the flexibility and scalability of model training, ensuring the integrity and consistency of training data, strengthening the model's adaptability and robustness, and ultimately achieving a high-performance image pre-trained model to meet the needs of complex image analysis tasks.

[0114] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of the equivalents of the application be incorporated into the invention.

[0115] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A training method for an image pre-training model, characterized in that, Includes the following steps: Step S1: After acquiring the original training images, extract the semantic information from the images and perform feature processing to generate structured high-dimensional feature data; reduce the high-dimensional feature data in three different proportions to finally generate a multi-resolution feature body; Step S2: Analyze the texture sharpness of the image by fusing feature volumes at various resolutions, and divide the original image into flawed training images and qualified training images based on texture sharpness. Step S3: Source-correct the defective training image to create a defect-corrected training image, and construct positive and negative contrast images based on the defect-corrected training image and the defective training image; perform comparative training using the positive and negative contrast images to construct a pre-trained model for the defective image. Step S3, source-correcting the defective training image to create a defect-corrected training image, and constructing positive and negative contrast images based on the defect-corrected training image and the defective training image, includes: Extract defect features from defective training images and crop defective regions from the training images based on these features; By analyzing defect features, defect source processing is performed on defective areas in the image to obtain the source data of image defects. Specifically: The defective region image is input into the pre-constructed source tracing network structure. When the defective image is input, the convolutional layer extracts feature maps at different scales, performs weighted processing on the feature response weights of each spatial location, calculates the attention response factor corresponding to each pixel, and performs a dot product mapping operation on the original feature map to generate a location calibration feature map. Combined with the contextual semantic features of the original image, deconvolution reconstruction is performed. The reconstruction output is spatial coordinate image data containing the defect generation starting point, edge expansion trajectory, and explicit defect texture change path. This data constitutes the image defect source data. The source data of image defects are used to simulate defect avoidance, in order to generate simulated defect source avoidance data, specifically: The coordinate information and structural path in the source tracing results are read and input into the defect avoidance generator. The generator includes a shape consistency constraint module. In the shape consistency constraint module, a specific regularization loss is used to impose triple constraints on the gray-level continuity, edge smoothness and texture homogeneity of the reconstructed area. After all simulated avoidance images are verified, unstable images are removed and only avoidance data with texture boundary consistency greater than the threshold are retained as simulated defect source avoidance data output. By simulating defect source avoidance data, defect area images are processed to correct defects, resulting in defect correction training images. The defect correction training image and the defect training image are projected into the same space and image comparison processing is performed to obtain positive and negative contrast images. Step S4: Perform weak perturbation transformation on qualified training images, and pair the images before and after the transformation one by one to construct positive pairs before and after perturbation; find regions with consistent features from the positive pairs before and after perturbation, and use these region images to train the pre-trained model of qualified images. Step S5: Combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and perform joint training through this sample stream to obtain the integrated image pre-trained model.

2. The training method for the image pre-training model according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Perform cross-domain sampling on the original image data to obtain a heterogeneous source dataset; perform semantic segmentation mapping on the heterogeneous source dataset to generate a semantic guidance tensor; Step S12: Perform high-dimensional feature mapping on the original training image based on the semantically guided tensor. The channel dimension expansion and semantic alignment modules are used to perform nested convolution operations. The convolution kernel size is set to 3×3 and the dilation rate ranges from [1,2,4] to obtain the feature dimension mapping data. Step S13: Perform distribution consistency calibration on the feature dimension mapping data to obtain normalized feature space data; Step S14: Perform multi-scale pyramid transformation on the normalized feature space data, wherein the number of pyramid scale layers is set to 3, the downsampling factor of each layer is 1, 0.5 and 0.25 respectively, and the channel compression rate is set to 0.5 to obtain multi-resolution feature volume.

3. The training method for the image pre-training model according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Aggregate the multi-resolution feature volumes across scales to obtain fused texture features; perform local texture response mapping on the fused texture features to obtain local texture response distribution data; Step S22: Calculate the local structure contrast based on the local texture response distribution data, where the texture window size is 8×8 pixels and the sliding step size is 4 pixels, and infer the image texture clarity based on the local structure contrast; Step S23: Perform threshold segmentation on the image texture sharpness based on a preset texture sharpness threshold to generate sharpness segmentation labels; Step S24: Filter the original training images into flawed training images and qualified training images by using sharpness segmentation labels.

4. The training method for the image pre-training model according to claim 1, characterized in that, Step S3, which combines positive and negative contrast images for comparative training to construct a pre-trained model of the flawed image, includes: Extract the difference contrast features from positive and negative contrast images; The salient regions of the image are located based on the difference contrast features, and a contrast salient mask is determined based on the salient regions of the image. Spatially align the contrast saliency mask, the defect training image, and the defect correction training image, and then perform channel stitching to generate the original-enhanced-saliency mask ternary image. Extract ternary image feature data from the original-enhanced-salient mask ternary image; The triple image feature data is input into the contrast learning network to perform triple contrast training, thereby generating discriminative contrast features; A pre-trained model of flawed images is trained using discriminative contrast features.

5. The training method for the image pre-training model according to claim 1, characterized in that, Step S4 involves performing a weak perturbation transformation based on the qualified training images and pairing the images before and after the transformation one by one to construct positive pairs before and after the perturbation, including: The qualified training images are processed by multi-scale block segmentation, and local texture data in each sub-block is extracted; Calculate the local texture distribution data of the local texture data; Weak perturbation transformation is applied to qualified training images based on local texture distribution data to generate perturbation qualified training images; The perturbation-compliant training images are paired one by one with the compliant training images to generate positive pairs before and after the perturbation.

6. The training method for the image pre-training model according to claim 1, characterized in that, Step S4 involves identifying regions with consistent features from the positive pairs before and after the perturbation, and using these region images to train a pre-trained model for qualified images. Extract the feature data of each layer of positive example pairs in the positive example pairs before and after the perturbation; Consistency assessment of feature data is performed on positive examples at each layer to generate feature consistency distribution data; Highly consistent regions are selected based on a preset consistency distribution threshold and characteristic consistency distribution data. The corresponding region images of qualified training images are decomposed by high consistency regions; A qualified image pre-training model is trained using images of the corresponding regions.

7. The training method for the image pre-training model according to claim 1, characterized in that, Step S5 includes the following steps: Step S51: Extract the output feature vectors of the defective image pre-trained model and the qualified image pre-trained model, and perform complementary feature mapping based on the output feature vectors to obtain dual-modal feature fusion data; Step S52: Re-estimate the category distribution based on the dual-modal feature fusion data to obtain the category attribution distribution data; filter the distributionally uncertain data in the dual-modal feature fusion data based on the category attribution distribution data, and construct the distributionally uncertain label data; Step S53: Perform self-supervised pseudo-label augmentation simulation on the uncertain label data to generate pseudo-label augmented data; use the pseudo-label augmented data for multi-task allocation to obtain multi-task training allocation data; Step S54: Perform task collaborative training scheduling based on multi-task training allocation data to obtain a collaborative sample stream; perform joint training of the ensemble model based on the collaborative sample stream to obtain an ensemble image pre-trained model.

8. A training system for an image pre-training model, characterized in that, The training system for performing the training method of the image pre-training model as described in claim 1 includes: The image acquisition module is used to acquire the original training images, extract the semantic information from the images, perform feature processing, and generate structured high-dimensional feature data; the high-dimensional feature data is then scaled down in three different proportions to finally generate a multi-resolution feature body; The texture screening module is used to analyze the texture sharpness of an image by fusing feature volumes at various resolutions, and to classify the original image into flawed training images and qualified training images based on the texture sharpness. The defect modeling module is used to trace and correct defective training images into defect-corrected training images, and to construct positive and negative contrast images based on the defect-corrected training images and the defective training images; and to perform comparative training by combining the positive and negative contrast images to construct a pre-trained model of defective images. The positive example modeling module is used to perform weak perturbation transformation on qualified training images and pair the images before and after the transformation to construct positive example pairs before and after perturbation; it identifies regions with consistent features from the positive example pairs before and after perturbation and uses these region images to train the pre-trained model of qualified images. The ensemble training module is used to combine the feature datasets of the flawed image pre-trained model and the qualified image pre-trained model into a collaborative sample stream, and then perform joint training through this sample stream to obtain the ensemble image pre-trained model.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed, it implements the training method of the image pre-training model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing model training method and device, electronic equipment and storage medium

    CN118071652A

  • Pumped storage power station construction anomaly detection method and system based on unmanned aerial vehicle image analysis

    CN119888507A