Method and device for generating surface defect image of steel wire rope
By combining the defect image generation method of the Stable Diffusion model and the Controlnet module, the problem of generating high-quality wire rope defect images is solved, the training effect and robustness of the detection model are improved, and complex defect features can be better dealt with.
Patent Information
- Application Number
- CN202510516156.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The prior art is difficult to generate high-quality wire rope defect images, resulting in limited improvement in machine vision detection effects and inability to effectively deal with complex defect characteristics.
The defect image generation method based on the Stable Diffusion model and Controlnet module is adopted to generate high-quality wire rope target defect images through basic feature extraction networks, defect feature generation networks and fusion decoding networks, and combine defect reference images with defective reference images, defect feature text and defect reference images.
The quality and efficiency of the generation of defect images on the surface of the wire rope is improved, the training effect and robustness of the detection model are enhanced, and complex defect features can be better simulated in real scenes.
Smart Images

Figure CN120563646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image generation method and device, in particular to a method and device for generating a surface defect image of a steel wire rope. Background Art
[0002] Steel wire rope is widely used in many industries such as construction, mining, transportation, and petroleum due to its excellent strength, toughness, and adaptability. It not only excels in load-bearing capacity, but can also work for a long time in harsh environments. It is a core component of a variety of heavy equipment and machinery.
[0003] Although wire ropes offer excellent performance, they can also develop various damage issues over time. Common types of damage include broken wires, wear, deformation, and corrosion. Understandably, if problems with wire ropes are not discovered promptly, the damage will continue to increase over time. If no action is taken when damage reaches a certain level, safety accidents can occur.
[0004] When detecting defects in wire ropes, traditional detection methods include magnetic induction testing, ultrasonic testing, and eddy current testing. Among them, magnetic induction testing detects broken wires and cracks through changes in magnetic flux, which is suitable for detecting internal defects in wire ropes but has high equipment costs. Ultrasonic testing uses the reflection and attenuation characteristics of waves to detect changes in the internal structure of wire ropes. It has high accuracy but complex signal processing. Eddy current testing quickly detects surface and shallow defects through changes in conductivity. It is suitable for online detection, but is not sensitive to deep defects in wire ropes.
[0005] With the development of machine learning, researchers have proposed visual inspection methods for wire rope surface defects. Visual inspection methods include manual visual inspection and machine vision inspection. Manual visual inspection relies on the experience of the inspector and is low-cost, but inefficient and highly subjective. Machine vision inspection uses industrial cameras and algorithms (such as YOLO and Mask R-CNN) to automatically detect defects such as broken wires and wear in wire ropes. It offers high accuracy and speed, but requires a high initial equipment investment. Overall, machine vision offers advantages in automation and accuracy, making it the mainstream method for modern wire rope appearance inspection.
[0006] However, in real-world scenarios, researchers face the challenge of data scarcity due to the scarcity and difficulty in labeling wire rope defect data. In order to meet the needs of machine vision inspection, data enhancement can be used on wire rope defect data. Traditional data enhancement methods mainly include geometric transformations (such as rotation, flipping, scaling, cropping, and translation), color transformations (such as brightness, contrast, and saturation adjustment), and noise addition. These methods expand the dataset through simple image transformations and improve the robustness of the detection model to common deformations. However, their limitations are: the generated samples lack diversity and cannot simulate the complex defect characteristics in real scenes (such as random deformation and combinations of multiple defects). The coverage of data distribution is limited, which may lead to limited improvement in the actual defect detection effect.
[0007] In order to solve this problem, generating high-quality wire rope defect images has become an important research task. By generating realistic defect images, the training effect of the detection model can be greatly improved, and the accuracy and robustness of defect detection can be improved.
[0008] Currently, methods used to generate defect images are primarily divided into traditional defect image generation methods and deep learning-based methods, each with its own characteristics and limitations. In traditional machine learning, generative models are a concept distinct from discriminative models. Generative models aim to learn the essential characteristics of a data distribution and generate new data that conforms to that distribution.
[0009] In the research of generative models, Wang Xinyan proposed a diffusion model solution for defect sample simulation generation to address the problem of ineffective training of deep learning models due to the difficulty in obtaining defect samples. He improved the diffusion model and incorporated the attention mechanism and DDIM accelerated sampling technology into it, achieving rapid image generation. To address the problems of small number of X-ray image samples and class imbalance, Peng Yi introduced the diffusion model to the field of weld defect image enhancement and proposed an improved diffusion model algorithm to expand samples to improve the detection capability of rare categories of defects. To improve the skip-layer connection of the noise estimation network in the IDDPM model, he proposed the IDDPM_CCT model.
[0010] The defect image generation method based on deep learning and generative models shows great potential in improving data diversity and quality. However, due to the particularity of wire ropes, how to effectively generate high-quality wire rope defect images is a technical problem that urgently needs to be solved. Summary of the Invention
[0011] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and device for generating a wire rope surface defect image, which can achieve the generation of high-quality wire rope surface defect images and improve the efficiency and reliability of wire rope surface defect image generation.
[0012] According to the technical solution provided by the present invention, a method for generating a surface defect image of a wire rope comprises:
[0013] Providing defect image generation conditions, and loading the provided defect image generation conditions into the constructed defect image generation model, wherein,
[0014] The defect image generation model includes a basic feature extraction network, a defect feature generation network and a fusion decoding network.
[0015] The defect image generation conditions include a complete reference image, defect feature text, and a defect reference image, wherein the type of the steel wire rope in the complete reference image is consistent with the type of the steel wire rope in the defect reference image;
[0016] Loading the complete reference image into the basic feature extraction network to obtain a complete reference latent vector extracted by the basic feature extraction network, and loading the extracted complete reference latent vector into the fusion decoding network;
[0017] Loading the defect feature text and the defect reference image into the defect feature generation network to generate a text-guided defect latent vector through the defect feature generation network, and loading the generated text-guided defect latent vector into the fusion decoding network;
[0018] A fusion decoding network is used to perform fusion decoding processing on the flawless reference latent vector and the text-guided defect latent vector, so as to generate a wire rope target defect image after the fusion decoding processing, wherein the defect state of the wire rope target defect image is consistent with the defect state described by the defect feature text.
[0019] The basic feature extraction network is constructed based on at least a feature extraction Stable Diffusion model and an IP-Adapter module, wherein:
[0020] For the loaded intact reference image, use the variational autoencoder in the feature extraction Stable Diffusion model to perform variational autoencoding on the intact reference image to generate a flawless encoding latent vector;
[0021] The IP-Adapter module uses a decoupled cross-attention mechanism to convert the complete encoding latent vector into the complete reference cross-attention feature, and injects the complete reference cross-attention feature into the U-Net module of the feature extraction Stable Diffusion model to generate the complete reference latent vector after denoising by the U-Net module.
[0022] The defect feature generation network is constructed based on at least the defect generation Stable Diffusion model and the Controlnet module, wherein:
[0023] Loading the defect reference image into the Controlnet module to generate defect constraints via the Controlnet module, and loading the defect constraints into the defect generation Stable Diffusion model;
[0024] The defect reference image and the defect feature text are simultaneously loaded into the defect generation stable diffusion model, so that the defect generation stable diffusion model generates a text-guided defect latent vector under the defect constraint condition.
[0025] The fusion decoding network includes a latent vector fusion unit and a latent vector decoder, wherein,
[0026] A latent vector fusion unit is used to perform latent vector fusion processing on the intact reference latent vector and the text-guided defect latent vector, so as to generate a wire rope fusion latent vector after the latent vector fusion processing;
[0027] The latent vector decoder decodes the wire rope fusion latent vector to generate a wire rope target image after decoding.
[0028] When performing latent vector fusion processing, it includes:
[0029] aligning the perfect reference latent vector with the text-guided defect latent vector in the latent space to generate a perfect aligned latent vector and a text-guided defect aligned latent vector after the latent space alignment;
[0030] Performing frequency domain feature extraction on the above-mentioned perfect alignment latent vector and text-guided defect alignment latent vector to obtain perfect alignment frequency domain features and text-guided defect alignment frequency domain features, respectively, wherein the perfect alignment frequency domain features include perfect alignment low-frequency features and perfect alignment high-frequency features, and the text-guided defect alignment frequency domain features include defect alignment low-frequency features and defect alignment high-frequency features;
[0031] Perform low-frequency fusion on the intact aligned low-frequency features and the defect aligned low-frequency features to generate aligned fused low-frequency features;
[0032] Perform high-frequency fusion of the intact alignment high-frequency features and the defect alignment high-frequency features to generate aligned fusion high-frequency features;
[0033] Based on the aligned fusion low-frequency features and the aligned fusion high-frequency features, the aligned combined frequency domain features are combined to generate aligned combined frequency domain features after the combination, and the generated aligned combined frequency domain features are transformed into the time domain to generate the wire rope fusion latent vector after the time domain transformation.
[0034] When the perfect reference latent vector is aligned with the text-guided defective latent vector in the latent space, we have:
[0035]
[0036] Among them, h IP is the perfect alignment latent vector, h CN Latent vector for text-guided defect alignment, W IP is the mapping matrix of the missing reference latent vector, W CN is the mapping matrix of the text-guided defect latent vector, b IP is the bias term of the missing reference latent vector, b CN is the bias term of the text-guided defect latent vector, z IP is the complete reference latent vector, z CN Bootstrapping defect latent vectors for text.
[0037] When generating aligned combined frequency domain features, we have:
[0038]
[0039] Among them, Z fuse To align the combined frequency domain features, To align and fuse low-frequency features, To align and fuse high-frequency features, α is the background fusion control coefficient, β is the defect enhancement coefficient, To perfectly align the low frequency features, To perfectly align high frequency features, Align low-frequency features for defects, Align high-frequency features for defects.
[0040] When providing a defect reference image, the providing method includes:
[0041] Provide defect reference source images,
[0042] Convert the defect reference source image into the HSV color space and the Lab color space respectively, and then perform binarization operations in the HSV color space and the Lab color space respectively, so as to generate an HSV color space binarized image and a Lab color space binarized image respectively after the binarization operation;
[0043] Performing a union operation on the HSV color space binary image and the Lab color space binary image to form a spatially fused binary image;
[0044] Based on the HSV color space of the defect reference source image and the spatial fusion binary image, the hue change state in the HSV color space is calculated and determined;
[0045] Based on the Lab color space of the defect reference source image and the spatial fusion binary image, calculate and determine the brightness change state in the Lab color space;
[0046] Based on the calculated hue change state and brightness change state,
[0047] When the hue changes significantly, hue separation is performed in the HSV color space of the defect reference source image, and a defect reference image is generated after separation;
[0048] When the brightness change is significant, brightness enhancement is performed in the Lab color space of the defect reference source image, and a defect reference image is generated after the brightness enhancement.
[0049] When building a defect image generation model, the construction method includes:
[0050] Constructing a defect image generation basic model, wherein the defect image generation basic model includes a preliminarily trained basic feature extraction network, a preliminarily trained defect feature generation network, and an untrained fusion decoding network;
[0051] Construct a basic model training dataset for training the basic model for defect image generation, where:
[0052] The basic model training data set includes a non-defective training sample set and a defective training sample set.
[0053] The defect-free training sample set includes a plurality of defect-free training samples, each of which includes a defect-free first-category steel wire rope training image;
[0054] The defect training sample set includes a plurality of defect training samples, each defect training sample includes a second type of steel wire rope training image with defects and a training image text for characterizing the defect state of the second type of steel wire rope training image;
[0055] The defect image generation basic model is trained based on the configured basic model training conditions, where:
[0056] During model training, a defect-free training sample is loaded into the basic feature extraction network, and a defect training sample is loaded into the defect feature generation network. After that, a wire rope training defect image is generated through the fusion decoding network.
[0057] Calculating a loss function for model training based on the wire rope defect training image and the second type of wire rope training image in the current defect training sample, and adjusting network parameters of the defect image generation basic model based on the calculated loss function until the defect image generation basic model is trained to a target state;
[0058] The defect image generation base model that reaches the target state based on model training generates the required defect image generation model.
[0059] A device for generating a surface defect image of a wire rope comprises a defect image generating device and a defect image generating model deployed in the defect image generating device, wherein:
[0060] The defect image generation conditions of the wire rope are obtained, and the defect image generation device uses the above-mentioned method to perform defect image generation processing, and generates a target defect image of the wire rope after the defect image generation processing.
[0061] The advantages of the present invention are as follows: a flawless reference image is loaded into a basic feature extraction network to obtain a flawless reference latent vector through extraction by the basic feature extraction network; a defect feature text and a defect reference image are loaded into a defect feature generation network to generate a text-guided defect latent vector through the defect feature generation network; thereafter, a fusion decoding network is used to perform fusion decoding processing on the flawless reference latent vector and the text-guided defect latent vector to generate a target defect image of the wire rope after the fusion decoding processing.
[0062] The basic feature extraction network is constructed based on the feature extraction Stable Diffusion model and the IP-Adapter module. At the same time, the defect feature generation network is constructed based on the defect generation Stable Diffusion model and the Controlnet module. This can effectively resolve the conflicts caused by combining with a Stable Diffusion model, while better realizing the generation of wire rope surface defect images and improving the generation quality, efficiency and reliability of wire rope target defect images. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 A schematic diagram of an embodiment of generating a surface defect image of a steel wire rope according to the present invention.
[0064] Figure 2 This is a structural block diagram of an embodiment of the defect image generation model of the present invention.
[0065] Figure 3 FIG. 1 is a schematic diagram of an embodiment of a complete reference image according to the present invention.
[0066] Figure 4 This is a schematic diagram of a first embodiment of a defect reference image according to the present invention.
[0067] Figure 5 This is a schematic diagram of a first embodiment of generating a target defect image of a wire rope according to the present invention. DETAILED DESCRIPTION
[0068] The present invention will be further described below with reference to specific drawings and embodiments.
[0069] In order to achieve the generation of high-quality wire rope surface defect images and improve the efficiency and reliability of wire rope surface defect image generation, the present invention provides a method for generating wire rope surface defect images. Specifically, the generation method includes:
[0070] Providing defect image generation conditions, and loading the provided defect image generation conditions into the constructed defect image generation model, wherein,
[0071] The defect image generation model includes a basic feature extraction network, a defect feature generation network and a fusion decoding network.
[0072] The defect image generation conditions include a complete reference image, defect feature text, and a defect reference image, wherein the type of the steel wire rope in the complete reference image is consistent with the type of the steel wire rope in the defect reference image;
[0073] Loading the complete reference image into the basic feature extraction network to obtain a complete reference latent vector extracted by the basic feature extraction network, and loading the extracted complete reference latent vector into the fusion decoding network;
[0074] Loading the defect feature text and the defect reference image into the defect feature generation network to generate a text-guided defect latent vector through the defect feature generation network, and loading the generated text-guided defect latent vector into the fusion decoding network;
[0075] A fusion decoding network is used to perform fusion decoding processing on the flawless reference latent vector and the text-guided defect latent vector, so as to generate a wire rope target defect image after the fusion decoding processing, wherein the defect state of the wire rope target defect image is consistent with the defect state described by the defect feature text.
[0076] As can be seen from the above description, the wire rope surface defect image specifically refers to the presence of defects in the wire rope image, and the defects should be distributed on the surface of the wire rope in the wire rope image. Among them, the types of wire rope defects can generally be broken wires, wear, rust and / or deformation. In order to improve the quality of the generated wire rope surface defect image, Figure 1 An embodiment of the present invention for generating a surface defect image of a wire rope is shown. As can be seen from the figure, when generating a surface defect image of a wire rope, defect image generation conditions should be provided and loaded into a defect image generation model. The defect image generation conditions and the defect image generation model are explained in detail below.
[0077] In specific implementation, the defect image generation conditions include a flawless reference image, a defect feature text, and a defect reference image, wherein the flawless reference image specifically means that there are no defects on the surface of the wire rope in the image, such as Figure 3 As shown in the figure; the defect reference image specifically refers to the surface image of the wire rope with defects, such as Figure 4 , Figure 5 This is an embodiment of generating a target defect image of a wire rope. Generally, the size of the intact reference image and the defective reference image are the same.
[0078] It should be noted that defects within the defect reference image generally include one or more of the aforementioned broken wires, wear, corrosion, and deformation. Generally, the type of wire rope within the intact reference image is consistent with the type of wire rope within the defect reference image, meaning that wire ropes of the same type should be used as generation conditions. Defect feature text is primarily used to guide the generation of defect distribution within the surface defect image. The defect feature text should include the defect type, the number of defects corresponding to the defect type, and / or the distribution of defects within the wire rope surface defect image.
[0079] Figure 2 An embodiment of the defect image generation model of the present invention is shown in FIG. As can be seen from the figure, the defect image generation model should at least include a basic feature extraction network, a defect feature generation network, and a fusion decoding network, wherein the basic feature extraction network and the defect feature generation network should both be connected to the fusion decoding network. It should be noted that when the defect image generation conditions are loaded into the defect image generation model, it specifically refers to loading the intact reference image into the basic feature extraction network, and loading the defect feature text and the defect reference image into the defect feature generation network at the same time. At this time, the fusion decoding network serves as the output layer of the defect image generation model. Therefore, a wire rope surface defect image can be generated through the fusion decoding network. Specifically, the wire rope surface defect image generated by the fusion decoding network is the target defect image of the wire rope.
[0080] It should be understood that the defect state of the wire rope target defect image is consistent with the defect state described in the defect feature text. The defect state described in the defect feature text can refer to the above-mentioned corresponding explanation of the defect feature text. Therefore, the defect state contained in the wire rope target defect image may generally include the type of defect, the number and distribution of the corresponding defect type, etc.
[0081] After loading the intact reference image into the basic feature extraction network, the basic feature extraction network can extract the intact reference latent vector; after loading the defect feature text and the defect reference image into the defect feature generation network, the defect feature generation network can generate the text-guided defect latent vector. Thereafter, the fusion decoding network performs fusion decoding processing on the intact reference latent vector and the text-guided defect latent vector, and finally generates the wire rope target defect image. The specific method and process of generating the wire rope target defect image will be described in detail below.
[0082] In one embodiment of the present invention, the basic feature extraction network is constructed based on at least a feature extraction StableDiffusion model and an IP-Adapter module, wherein:
[0083] For the loaded intact reference image, use the variational autoencoder in the feature extraction Stable Diffusion model to perform variational autoencoding on the intact reference image to generate a flawless encoding latent vector;
[0084] The IP-Adapter module uses a decoupled cross-attention mechanism to convert the complete encoding latent vector into the complete reference cross-attention feature, and injects the complete reference cross-attention feature into the U-Net module of the feature extraction Stable Diffusion model to generate the complete reference latent vector after denoising by the U-Net module.
[0085] In order to extract the complete reference latent vector of the complete reference image, the basic feature extraction network can be constructed based on the feature extraction Stable Diffusion model and the IP-Adapter module, wherein the feature extraction Stable Diffusion model can adopt the existing commonly used Stable Diffusion model. It should be noted that when the existing commonly used Stable Diffusion model is adopted, the feature extraction Stable Diffusion model may include a variational autoencoder (VAE) for image encoding, a text encoder and a U-Net module; the IP-Adapter (Image Projection Adapter) module is an extension module of the Stable Diffusion model, which aims to more accurately control the image generation process by combining image and text prompts (Prompt), and can enhance the model's response ability to specific input prompts (prompt words or other auxiliary information).
[0086] In specific implementation, when a basic feature extraction network is constructed based on the feature extraction Stable Diffusion model and the IP-Adapter module, the connection and coordination between the IP-Adapter module and the feature extraction Stable Diffusion model can be consistent with the existing technology. For example, the IP-Adapter module can be connected to the variational autoencoder and U-Net module in the Stable Diffusion model. As can be seen from the above description, when using the basic feature extraction network to generate a complete reference latent vector, only a complete reference image is required.
[0087] To meet the requirement of generating a complete reference latent vector, a decoupled cross-attention mechanism should be adopted to separate the text encoder and the variational autoencoder within the feature extraction Stable Diffusion model. This allows the basic feature extraction network formed when text is input into the text encoder to process the complete reference image and generate a complete reference latent vector. Specifically, the method and process of separating the text encoder and the variational autoencoder using the decoupled cross-attention mechanism are consistent with the existing technology, and the details of the decoupled cross-attention mechanism will not be repeated here.
[0088] During operation, the variational autoencoder in the feature extraction Stable Diffusion model receives a flawless reference image and performs variational autoencoding on the received flawless reference image to generate a flawless encoding latent vector after variational autoencoding. The generated flawless encoding latent vector can generate a flawless reference cross-attention feature through the IP-Adapter module, and the flawless reference cross-attention feature is injected into the U-Net module of the feature extraction Stable Diffusion model to generate a flawless reference latent vector after denoising processing through the U-Net module. Therefore, the basic feature extraction network of the present invention can extract and enhance the visual features of the flawless reference image and generate a clear, real, and flawless flawless reference latent vector.
[0089] It should be understood that when the basic feature extraction network adopts the above-mentioned structural form, the generated complete reference latent vector is the representation of the complete reference image in the low-dimensional abstract space. The latent vectors of the present invention all express the same meaning. For details, please refer to the description here.
[0090] It's understood that the methods and processes for the variational autoencoder to perform variational autoencoding on the perfect reference image, the IP-Adapter module to generate perfect reference cross-attention features, and the U-Net module to generate the perfect reference latent vector through denoising are all consistent with existing technologies and will not be further described here. After generating the perfect reference latent vector, since self-decoding is not required, the self-decoder in the Stable Diffusion model can be omitted.
[0091] In one embodiment of the present invention, the defect feature generation network is constructed based on at least a defect generation StableDiffusion model and a Controlnet module, wherein:
[0092] Loading the defect reference image into the Controlnet module to generate defect constraints via the Controlnet module, and loading the defect constraints into the defect generation Stable Diffusion model;
[0093] The defect reference image and the defect feature text are simultaneously loaded into the defect generation stable diffusion model, so that the defect generation stable diffusion model generates a text-guided defect latent vector under the defect constraint condition.
[0094] In order to generate text-guided defect latent vectors, the defect feature extraction network may include a defect generation Stable Diffusion model and a Controlnet module, wherein the defect generation Stable Diffusion model may adopt an existing commonly used Stable Diffusion model. Therefore, the situation of the defect generation Stable Diffusion model can refer to the corresponding description above and will not be repeated here. The Controlnet module is a neural network-based architecture, which is mainly used to control the defect generation Stable Diffusion model through additional input to achieve precise control of the generated content. The Controlnet module in the defect feature generation network can also adopt an existing commonly used form. Therefore, when the ControlNet module is used to control the defect generation Stable Diffusion model, the connection and coordination form between the Controlnet module and the defect generation Stable Diffusion model can be consistent with the existing technology, and the specific content of the generation of the defect generation Stable Diffusion model can be controlled.
[0095] When generating a text-guided defect latent vector, the defect reference image should be loaded into the Controlnet module, and the defect reference image and the defect feature text should be loaded into the defect generation Stable Diffusion model at the same time. Among them, the Controlnet module processes the defect reference image to generate defect constraints. The specific method and process of generating defect constraints can be consistent with the existing technology and will not be repeated here. When the defect reference image and the defect feature text are loaded into the defect generation Stable Diffusion model at the same time, the defect reference image is subjected to variational self-encoding by the variational autoencoder in the defect generation Stable Diffusion model to generate a defect encoding latent vector, and the defect feature text is subjected to text encoding by the text encoder to generate a defect feature text encoding latent vector. The constraint condition is integrated with the defect encoding latent vector through the attention mechanism. Thereafter, the defect feature text encoding latent vector is combined to control the denoising process of the U-Net module in the defect generation Stable Diffusion model, and a text-guided defect latent vector can be generated.
[0096] The above description only illustrates one embodiment of the process of generating text-guided defect latent vectors using a defect feature generation network. It is understood that when a defect feature generation network is constructed using a Stable Diffusion model and a Controlnet module, the detailed process of generating text-guided defect latent vectors can be consistent with the prior art and will not be further described here. In specific implementations, when generating text-guided defect latent vectors, since direct decoding of the text-guided defect latent vectors is not required, the self-decoder of the Stable Diffusion model should not be included in the defect feature generation network.
[0097] As can be seen from the above description, structural control (such as edge detection, posture control, depth map, etc.) can be achieved through the Controlnet module, while the IP-Adapter module is mainly used to process the style or features of the reference image. Therefore, if the Controlnet module and the IP-Adapter module are combined with a Stable Diffusion model at the same time, it may affect the U-Net module and cause conflicts during feature extraction. For example, the Controlnet module forces shape constraints, but the IP-Adapter module may want to adjust the shape to better match the reference style. At this time, the U-Net module in the Stable Diffusion model may cause conflicts during feature extraction. In one embodiment of the present invention, a basic feature extraction network is constructed based on the feature extraction Stable Diffusion model and the IP-Adapter module. At the same time, a defect feature generation network is constructed based on the defect generation Stable Diffusion model and the Controlnet module. This can effectively resolve the conflicts caused by combining with a Stable Diffusion model at the same time, but can better achieve the generation of wire rope surface defect images.
[0098] In one embodiment of the present invention, the fusion decoding network includes a latent vector fusion unit and a latent vector decoder, wherein:
[0099] A latent vector fusion unit is used to perform latent vector fusion processing on the intact reference latent vector and the text-guided defect latent vector, so as to generate a wire rope fusion latent vector after the latent vector fusion processing;
[0100] The latent vector decoder decodes the wire rope fusion latent vector to generate a wire rope target image after decoding.
[0101] It should be understood that after generating the intact reference latent vector and the text-guided defect latent vector, they need to be fused and decoded. To achieve this fusion decoding process, the fusion decoding network should include a latent vector fusion unit and a latent vector decoder connected to the latent vector fusion unit. The latent vector fusion unit is connected to the basic feature extraction network and the defect feature generation network so that the latent vector fusion unit can be used to perform latent vector fusion processing on the intact reference latent vector and the text-guided defect latent vector to generate a fused wire rope latent vector. Thereafter, the latent vector decoder is used to decode the fused wire rope latent vector and generate a wire rope target image after decoding.
[0102] In specific implementations, the latent vector decoder can employ the autodecoder commonly used in stable diffusion models. Therefore, the method and process for decoding the wire rope fusion latent vector using the latent vector decoder are consistent with existing techniques. The following describes the method and process for performing latent vector fusion processing in the latent vector fusion unit.
[0103] In order to achieve fusion decoding, the corresponding output ends of the U-Net module in the feature extraction Stable Diffusion model and the U-Net module in the defect generation Stable Diffusion model should be connected to the latent vector fusion unit so that the latent vector fusion unit can perform latent vector fusion processing on the complete reference latent vector and the text-guided defect latent vector.
[0104] In one embodiment of the present invention, performing latent vector fusion processing includes:
[0105] aligning the perfect reference latent vector with the text-guided defect latent vector in the latent space to generate a perfect aligned latent vector and a text-guided defect aligned latent vector after the latent space alignment;
[0106] Performing frequency domain feature extraction on the above-mentioned perfect alignment latent vector and text-guided defect alignment latent vector to obtain perfect alignment frequency domain features and text-guided defect alignment frequency domain features, respectively, wherein the perfect alignment frequency domain features include perfect alignment low-frequency features and perfect alignment high-frequency features, and the text-guided defect alignment frequency domain features include defect alignment low-frequency features and defect alignment high-frequency features;
[0107] Perform low-frequency fusion on the intact aligned low-frequency features and the defect aligned low-frequency features to generate aligned fused low-frequency features;
[0108] Perform high-frequency fusion of the intact alignment high-frequency features and the defect alignment high-frequency features to generate aligned fusion high-frequency features;
[0109] Based on the aligned fusion low-frequency features and the aligned fusion high-frequency features, the aligned combined frequency domain features are combined to generate aligned combined frequency domain features after the combination, and the generated aligned combined frequency domain features are transformed into the time domain to generate the wire rope fusion latent vector after the time domain transformation.
[0110] It should be understood that due to the difference in the latent space distribution of the perfect reference latent vector and the text-guided defect latent vector, directly fusing the perfect reference latent vector with the text-guided defect latent vector in an additive manner may lead to feature mismatch, which in turn will cause the generated wire rope target image decoded by the latent vector decoder to fail to meet the generation requirements. Therefore, it is necessary to align the perfect reference latent vector and the text-guided defect latent vector in the latent space (Latent Space Alignment, LSA).
[0111] In a specific implementation, the perfect reference latent vector and the text-guided defect latent vector can be linearly mapped using a linear mapping based on statistical analysis, and then aligned after the linear mapping. Specifically, when the perfect reference latent vector and the text-guided defect latent vector are aligned in the latent space, we have:
[0112]
[0113] Among them, h IP is the perfect alignment latent vector, h CN Latent vector for text-guided defect alignment, W IP is the mapping matrix of the missing reference latent vector, W CN is the mapping matrix of the text-guided defect latent vector, b IP is the bias term of the missing reference latent vector, b CN is the bias term of the text-guided defect latent vector, z IP is the complete reference latent vector, z CN Bootstrapping defect latent vectors for text.
[0114] It should be noted that when using linear mapping based on statistical analysis for potential space alignment, the above-mentioned mapping matrix W should be calculated IP , mapping matrix W CN , bias term b IP and the bias term b CN , the following is an example to illustrate the specific calculation process, specifically:
[0115]
[0116] Among them, IP,i is the i-th position feature vector in the perfect alignment latent vector, N is the number of position feature vectors in the perfect alignment latent vector and the text-guided defect latent vector, μ IP is the mean of the N position feature vectors in the perfect alignment latent vector, Ψ CN,i is the i-th position feature vector in the text-guided defect latent vector, μ CN is the mean of the N position feature vectors in the text-guided defect latent vector, Σ IP is the covariance matrix of the perfect aligned latent vector, Σ CN is the covariance matrix of the text-guided defect latent vector.
[0117] It should be understood that by using the same U-Net backbone structure for the basic feature extraction network and the defect generation network, the number of positional feature vectors within the flawless alignment latent vector and the text-guided defect latent vector should be the same. Furthermore, after obtaining the flawless alignment latent vector and the text-guided defect latent vector, the corresponding positional feature vectors can be obtained using commonly used techniques in the art.
[0118] In specific implementation, based on the above calculation results, we have:
[0119]
[0120] It can be understood that the calculation of the above mapping matrix W IP , mapping matrix W CN , bias term b IP and the bias term b CN After that, the above calculation method can be used to align the perfect reference latent vector and the text-guided defective latent vector in the latent space, and the perfect aligned latent vector h can be obtained respectively. IP , the text-guided defect alignment latent vector h CN .
[0121] When extracting frequency domain features from the above-mentioned perfect alignment latent vectors and text-guided defect alignment latent vectors, Fourier transform can be used to convert the perfect alignment latent vectors and text-guided defect alignment latent vectors into the frequency domain for frequency domain feature extraction. Specifically, the following steps are performed:
[0122]
[0123] Among them, Z IP For the perfect alignment latent vector h IP Frequency domain characteristics of Z CN is the frequency domain feature of the latent vector for text-guided defect alignment, and F is the Fourier transform.
[0124] It should be noted that for the perfect alignment latent vector h IP Perform Fourier transform and obtain the frequency domain feature Z IP The method and process of Fourier transform are consistent with the existing technology, and the method and process of Fourier transform are not repeated here. IP And the frequency domain feature Z CN Finally, the frequency features can be decomposed into low-frequency features representing global background information and high-frequency features representing local details and defects according to the frequency range. Specifically:
[0125]
[0126] Among them, Z is the frequency domain feature Z IP Or frequency domain feature Z CN , when Z is the frequency domain feature Z IP When LowPass(Z) is used, the low-frequency features can be fully aligned. After that, the corresponding perfect alignment high-frequency features can be calculated When Z is the frequency domain feature Z CNWhen LowPass(Z) is used, the defect alignment low-frequency feature can be obtained. Afterwards, the corresponding defect alignment high-frequency features can be calculated
[0127] Specifically, LowPass(Z) represents a low-pass filtering process, wherein when performing the low-pass filtering process, a low-frequency cutoff frequency and a high-frequency starting frequency should be determined. For the low-frequency cutoff frequency and the high-frequency starting frequency, an embodiment may be:
[0128]
[0129] Where: f max Indicates the maximum frequency, maximum frequency f max It is related to the image size. One feasible calculation method is:
[0130]
[0131] Where M is a parameter related to the size of the intact reference image or the defective reference image. If the size of the intact reference image or the defective reference image is M×M, the corresponding maximum frequency f can be calculated. max .
[0132] In one embodiment of the present invention, when generating the aligned combined frequency domain features, there are:
[0133]
[0134] Among them, Z fuse To align the combined frequency domain features, To align and fuse low-frequency features, To align and fuse high-frequency features, α is the background fusion control coefficient, β is the defect enhancement coefficient, To perfectly align the low frequency features, To perfectly align high frequency features, Align low-frequency features for defects, Align high-frequency features for defects.
[0135] Specifically, when the low-frequency features of the perfect alignment and the low-frequency features of the defect alignment are fused at low frequencies, we have: Among them, α∈[0,1]; when the high-frequency features of the perfect alignment and the high-frequency features of the defect alignment are fused at high frequency, we have: The defect enhancement coefficient β also ranges from [0, 1]. It can be understood that the background fusion control coefficient α can be used to control the low-frequency fusion state, while the defect enhancement coefficient β can be used to control the high-frequency fusion state. In other words, the weighted fusion state of low-frequency features and high-frequency features can be controlled by the values of the background fusion control coefficient α and the defect enhancement coefficient β to ensure a balanced expression of background and defects.
[0136] After obtaining the aligned combined frequency domain features in the above manner, the generated aligned combined frequency domain features can be transformed into the time domain to generate a wire rope fusion latent vector after the time domain transformation. It can be understood that the generated wire rope fusion latent vector can comprehensively represent the normal background and defect features, that is, it can improve the quality of the subsequent decoding and generation of the wire rope target defect image using the latent vector decoder.
[0137] In one embodiment of the present invention, when providing a defect reference image, the providing method includes:
[0138] Provide defect reference source images,
[0139] Convert the defect reference source image into the HSV color space and the Lab color space respectively, and then perform binarization operations in the HSV color space and the Lab color space respectively, so as to generate an HSV color space binarized image and a Lab color space binarized image respectively after the binarization operation;
[0140] Performing a union operation on the HSV color space binary image and the Lab color space binary image to form a spatially fused binary image;
[0141] Based on the HSV color space of the defect reference source image and the spatial fusion binary image, the hue change state in the HSV color space is calculated and determined;
[0142] Based on the Lab color space of the defect reference source image and the spatial fusion binary image, calculate and determine the brightness change state in the Lab color space;
[0143] Based on the calculated hue change state and brightness change state,
[0144] When the hue changes significantly, hue separation is performed in the HSV color space of the defect reference source image, and a defect reference image is generated after separation;
[0145] When the brightness change is significant, brightness enhancement is performed in the Lab color space of the defect reference source image, and a defect reference image is generated after the brightness enhancement.
[0146] It should be noted that the defect reference source image generally refers to the image generated by capturing the wire rope in the working scene. The working scene of the wire rope generally includes an environment such as underground coal mines. The wire rope in the working scene is generally oily and muddy. Therefore, when the wire rope in the working scene is image captured to obtain the defect reference source image, and the defect reference image is directly generated using the defect reference source image, the quality of the defect reference image is low, which in turn leads to a low quality of the generated wire rope target defect image.
[0147] The color space of the defect reference source image is generally an RGB color space. The defect reference source image can be converted into the HSV color space and the Lab color space respectively through the technical means commonly used in this technical field. The specific conversion methods to the HSV color space and the Lab color space can be consistent with the existing technology and will not be repeated here. According to the characteristics of the wire rope, in the HSV color space, the hue (H) can be used to distinguish between oil stains (yellowish, brown) and mud stains (gray-brown), the saturation (S) can reflect the degree of pollution, and the brightness (V) helps to identify low-brightness pollution areas. The brightness (L) channel in the Lab color space can enhance the structural details of the wire rope, and the a / b channel can be used to distinguish between oil stains (a channel is reddish, b channel is yellowish) and mud stains (a / b channels are close to neutral gray). Therefore, after converting to the HSV color space and the Lab color space, the defect reference source image can be effectively processed and the quality of the generated defect reference image can be improved.
[0148] After conversion to the HSV color space, the HSV color space can be binarized to generate an HSV color space binary image. Similarly, the Lab color space can be binarized to generate a Lab color space binary image. The following example illustrates the binarization method and process based on the characteristics of oil and mud on the wire rope.
[0149] Specifically, when performing binarization in the HSV color space, an HSV color space mask should be constructed. Thereafter, the HSV color space mask is used to determine the binarized value of each pixel in the HSV color space. For the HSV color space mask, there is:
[0150]
[0151] Where M HSV is the HSV color space mask, M oil,H is the oil mask in HSV color space, M mud,H is the HSV color space dirt mask, M H is the HSV color space hue mask, S low is the first saturation threshold, S mid is the second saturation threshold, V low is the first brightness threshold, V high is the second brightness threshold, H low is the first hue threshold, H high is the second hue threshold, ∪ is the union operation, and ∩ is the intersection operation.
[0152] In specific implementation, the first saturation threshold S low Can be set to 20, the second saturation threshold S midCan be set to 40, the first brightness threshold V low Can be set to 80, the second brightness threshold V high Can be set to 200, the first threshold of hue H low Can be set to 10, the first threshold of hue H high It can be set to 40; of course, it can also be set to other values, which can be determined according to the working scenario of the defect reference source image.
[0153] It can be understood that when the HSV color space mask is used for binarization, the corresponding channel values of the hue H channel, saturation S channel, and brightness V channel corresponding to each pixel are determined. Thereafter, the corresponding channel values are used to determine the HSV color space oil mask M corresponding to the current pixel. oil,H , HSV color space dirt mask M mud,H And HSV color space hue mask M H The corresponding value can then be obtained to obtain the HSV color space mask M corresponding to the current pixel. HSV Specifically, when the HSV color space oil mask M oil,H , HSV color space dirt mask M mud,H And HSV color space hue mask M H When there is a value of 1 in the HSV color space mask M corresponding to the current pixel HSV The value should be 1 if and only if the HSV color space oil mask M oil,H , HSV color space dirt mask M mud,H And HSV color space hue mask M H When both are 0, the HSV color space mask M corresponding to the current pixel HSV The value should be 0.
[0154] In specific implementation, for each pixel in the HSV color space, the HSV color space mask M corresponding to each pixel is determined. HSV Value, based on the HSV color space mask M of all pixels HSV The value is to realize the binarization processing of the HSV color space and obtain the HSV color space binary image.
[0155] When binarizing the Lab color space, the method can be similar to the binarization of the HSV color space mentioned above. One feasible method includes:
[0156] Construct a Lab color space mask, which is:
[0157]
[0158] Among them, M Lab is the Lab color space mask, Moil,L is the Lab color space oil mask, M mud,L is the Lab color space dirt mask, L low is the first threshold of L channel, L mid is the second threshold of L channel, a thresh is the a channel threshold, b low is the first threshold of channel b, thresh is the second threshold of channel b, |a| is the absolute value of channel a, and |b| is the absolute value of channel b.
[0159] In specific implementation, the first threshold value L of the L channel low Can be set to 40, L channel second threshold L mid Can be set to 60, a channel threshold a thresh Can be set to 10, b channel first threshold b low Can be set to 20, b channel second threshold b thresh Can be set to 10.
[0160] After constructing the Lab color space mask, you can refer to the method for determining the HSV color space mask value mentioned above. That is, the method and process for generating the Lab color space binary image can refer to the corresponding description of the HSV color space binary image mentioned above, which will not be repeated here.
[0161] From the above description, it can be seen that the Lab color space binary image and the HSV color space binary image are both binary images. When the HSV color space binary image and the Lab color space binary image are subjected to a union operation, a spatially fused binary image can be formed. It can be seen that the spatially fused binary image includes all pixels with a binary value of 1 in the Lab color space binary image and the HSV color space binary image.
[0162] After obtaining the spatially fused binary image, the hue change state in the HSV color space can be calculated based on the defect reference source image and the spatially fused binary image. Simultaneously, the brightness change state in the Lab color space can be calculated based on the defect reference source image and the spatially fused binary image. The following examples illustrate how to calculate the hue and brightness transformation states.
[0163] When calculating the hue change state in the HSV color space, a feasible calculation method is:
[0164]
[0165] Among them, ΔH is the hue change state, is the mean hue of the polluted area, is the mean hue of the background area, Munion (i, j) is the value of the pixel at the coordinate position (i, j) in the spatial fusion binary image, and H(i, j) is the value of the hue H channel of the pixel at the coordinate position (i, j) in the HSV color space.
[0166] When calculating the brightness change state in the Lab color space, a feasible calculation method is:
[0167]
[0168] Among them, ΔL is the brightness change state, is the average brightness of the polluted area, is the mean brightness of the background area, and L(i, j) is the value of the L channel of the pixel with coordinate position (i, j) in the Lab color space.
[0169] After the hue change state and brightness change state are determined by the above method, further judgment is required. Specifically:
[0170] When the hue change is significant, hue separation is performed in the HSV color space of the defective reference source image, and a defective reference image is generated after separation. When the hue change state ΔH>30°, the hue change can be considered significant. Thereafter, hue separation can be performed using technical means commonly used in this technical field, that is, the method and process of hue separation can be consistent with the existing technology and will not be repeated here.
[0171] When the brightness change is significant, brightness enhancement is performed in the Lab color space of the defect reference source image, and a defect reference image is generated after the brightness enhancement. When the brightness change state ΔL is greater than 20, the brightness change can be considered significant. Thereafter, existing commonly used technical means can be used for brightness enhancement.
[0172] During specific implementation, the hue change threshold and the brightness change threshold can be determined according to the situation of the working scene where the wire rope is located. For example, the image of the wire rope in the working scene can be captured, and the hue change threshold and the brightness change threshold can be statistically obtained. After that, the calculated hue change state can be compared with the determined hue change threshold to determine whether the hue change is a significant state. Similarly, it can be determined whether the brightness change state is in a significant state.
[0173] When it is determined that the hue change is significant, the defect reference source image needs to be hue separated in the HSV color space, and after hue separation, it can be transformed into the RGB color space to generate a defect reference image. The hue separation method can adopt the commonly used method in the existing technology, which will not be repeated here.
[0174] When a significant brightness change is determined, brightness enhancement is performed in the Lab color space of the defect reference source image. After brightness enhancement, the image is converted to the RGB color space to generate a defect reference image. It is understood that the method for performing brightness enhancement in the Lab color space is consistent with existing techniques and will not be further described here.
[0175] The above shows an embodiment of processing a defect reference source image to generate a defect reference image. Of course, other methods can also be used to process the defect reference source image, and the specific method is based on whether a high-quality defect reference image can be obtained. Examples are not given here one by one.
[0176] In one embodiment of the present invention, when constructing a defect image generation model, the construction method includes:
[0177] Constructing a defect image generation basic model, wherein the defect image generation basic model includes a preliminarily trained basic feature extraction network, a preliminarily trained defect feature generation network, and an untrained fusion decoding network;
[0178] Construct a basic model training dataset for training the basic model for defect image generation, where:
[0179] The basic model training data set includes a non-defective training sample set and a defective training sample set.
[0180] The defect-free training sample set includes a plurality of defect-free training samples, each of which includes a defect-free first-category steel wire rope training image;
[0181] The defect training sample set includes a plurality of defect training samples, each defect training sample includes a second type of steel wire rope training image with defects and a training image text for characterizing the defect state of the second type of steel wire rope training image;
[0182] The defect image generation basic model is trained based on the configured basic model training conditions, where:
[0183] During model training, a defect-free training sample is loaded into the basic feature extraction network, and a defect training sample is loaded into the defect feature generation network. After that, a wire rope training defect image is generated through the fusion decoding network.
[0184] Calculating a loss function for model training based on the wire rope defect training image and the second type of wire rope training image in the current defect training sample, and adjusting network parameters of the defect image generation basic model based on the calculated loss function until the defect image generation basic model is trained to a target state;
[0185] The defect image generation base model that reaches the target state based on model training generates the required defect image generation model.
[0186] It should be understood that the constructed defect image generation basic model should be consistent with the above-mentioned defect image generation model. For details, please refer to the above description. After that, the defect image generation basic model can be trained, and the defect image generation model can be obtained after model training. In specific implementation, when constructing the defect image generation basic model, it should include a basic feature extraction network that has undergone preliminary training, a defect feature generation network that has undergone preliminary training, and an untrained fusion decoding network. That is, when constructing the defect image generation basic model, the basic feature extraction network and the defect feature generation network should be trained first and obtained. The following is a specific description of the situation of pre-training the basic feature extraction network and the defect feature generation network.
[0187] It should be noted that in order to obtain the basic feature extraction network after preliminary training, the basic feature extraction basic network should be constructed first, and then the basic feature extraction basic network should be trained so that the basic feature extraction network after preliminary training can be obtained after training. The situation of the basic feature extraction network can refer to the corresponding description above. The difference is that the basic feature extraction basic network should include a self-decoder, such as the self-decoder of the feature extraction Stable Diffusion model; similarly, in order to obtain the defect feature generation network after preliminary training, the defect feature generation basic network should also be constructed first, and then the defect feature generation basic network should be trained so that the defect feature generation network after preliminary training can be obtained after training. It can be understood that the defect feature generation basic network also includes a self-decoder, such as the self-decoder of the defect generation Stable Diffusion model.
[0188] When training the basic feature extraction network, initial training should be performed using a defect-free training sample set. The defect-free training sample set includes first-category wire rope training images, and the type of wire rope in the first-category wire rope training images should be the same as the type of wire rope corresponding to the defect-free reference images and defective reference images. During specific training, at least necessary training conditions, including the extraction network loss function, should be configured. These necessary training conditions may include the number of training iterations and the optimization algorithm. These conditions can be selected based on specific needs, such as a batch size of 2 to 4, a number of training iterations of 150 to 200, a learning rate of 2e-4, and an AdamW optimizer with a weight decay value of 1e-6.
[0189] After each round of training of the basic feature extraction network using a complete training sample set, the loss value of the extraction network loss function should be calculated. The extraction network loss function can adopt the MSE (Mean Squared Error) loss function. When calculating the loss value, a feasible method is:
[0190] The defect-free training sample is loaded into the basic feature extraction basic network, and the defect-free feature extraction training image corresponding to the current defect-free training sample is obtained, wherein the defect-free feature extraction training image is the image output by the self-decoder of the feature extraction StableDiffusion model. Among all the generated defect-free feature extraction training images, one defect-free feature extraction training image is selected and the corresponding defect-free training sample is determined. After that, the loss value of the extraction network loss function can be calculated by comparing the selected defect-free feature extraction training image with the first type of wire rope training image in the corresponding defect-free training sample. Then, the loss value is:
[0191]
[0192] Among them, MSE1 is the loss value of the extraction network loss function, I(i,j) is the pixel value of the coordinate (i,j) in the first type of wire rope training image, and I o (i, j) is the pixel value at coordinate (i, j) in the defect-free feature extraction training image, m is the number of pixel rows corresponding to the defect-free feature extraction training image and the first type of wire rope training image, and n is the number of pixel columns corresponding to the defect-free feature extraction training image and the first type of wire rope training image.
[0193] In specific implementation, after calculating the loss value MSE1 of the extraction network loss function, the loss value MSE1 should be compared with the feature extraction training loss threshold. When the loss value MSE1 is greater than the feature extraction training loss threshold, the network parameters of the basic feature extraction network should be adjusted in reverse and trained again until the loss value MSE1 is no greater than the feature extraction training loss threshold.
[0194] Specifically, after the loss value MSE1 is no greater than the feature extraction training loss threshold, an operator can perform manual visual judgment and subjectively determine whether the image quality of the defect-free feature extraction training image meets the standard. If so, training is terminated, thereby obtaining a basic feature extraction network after preliminary training. If not, the feature extraction training loss threshold can be lowered, and the basic feature extraction network model is further trained until the loss value MSE1 is no greater than the feature extraction training loss threshold and the image quality of the defect-free feature extraction training image is visually determined to meet the standard. Generally, the feature extraction training loss threshold can be set to 0.1.
[0195] When training the basic network for defect feature generation, a defect training sample set should be used, wherein each defect training sample includes a second-class wire rope training image with defects and a training image text used to characterize the defect state of the second-class wire rope training image. The type of wire rope corresponding to the second-class wire rope training image belongs to the same type as the type of wire rope in the above-mentioned intact reference image. The training image text describes the defect state in the second-class wire rope training image. The defect state can be referred to the corresponding description above.
[0196] When constructing the defect training sample set, images of defective wire ropes are captured. These images are then processed to generate defect reference images from the defect reference source images. After this processing, a second-category wire rope training image is generated. Once the second-category wire rope training images are obtained, a prompt word is generated for each defect within the second-category wire rope training image. The prompt word describes the wire rope defect characteristics, resulting in characteristic information about the wire rope defects and, in turn, the training image text. This indicates that the defect characteristics within the training image text correspond one-to-one to the defects in the second-category wire rope training images.
[0197] When constructing a defect training sample set, the defect types within the defect training sample should cover typical defect types, such as broken wires, wear, corrosion, and deformation. Training image text can be stored in JSON files or tables (such as CSV).
[0198] In order to obtain defect training samples, each defect image should be carefully annotated. Specifically, when annotating, at least the number of defects in the image should be marked, such as "1 broken wire" or "2 rust spots"; the location of the defect in the image should be determined, such as "upper left corner" or "central area"; and it should be classified as broken wire, wear, rust, or deformation. The annotation information is stored in a structured format, such as containing the image path and annotation feature information. Then, defect prompt words (i.e., feature text information) are generated, including the number of defects, defect location, and defect type. For example, Example 1: "There is a broken wire in the upper left corner of the wire rope, with a length of approximately 5 mm." Example 2: "There are 3 rust spots distributed on the surface of the wire, evenly distributed on the right." Each image forms a one-to-one correspondence with the corresponding prompt word, which is convenient for storage and retrieval. Data is stored in the form of tables (such as) or files (such as JSON).
[0199] It should be understood that when training the basic network for defect feature generation, it is still necessary to configure necessary training conditions, including generating the network loss function. The necessary training conditions may include the number of training iterations and the optimizer, etc., which can be selected according to needs. For example, you can refer to the training condition settings used when training the basic network for basic feature extraction mentioned above.
[0200] During training, the defect training samples are loaded into the defect feature generation basic network, and the defect feature generation basic network generates corresponding defect generation training images, wherein the defect generation training images are images output by the self-decoder of the defect generation StableDiffusion model. After each round of training of the defect feature generation basic network using the defect training sample set, one defect generation training image is selected from all the generated defect generation training images, and the corresponding defect training sample is determined. Thereafter, the loss value of the generation network loss function can be calculated by comparing the selected defect generation training image with the second type of wire rope training image in the corresponding determined defect training sample. The generation network loss function can adopt the MSE (Mean Squared Error) loss function. When the MES loss function is adopted, the corresponding description of the extraction network loss function can be referred to above.
[0201] In specific implementation, after calculating the loss value of the generation network loss function, the loss value should be compared with the defect generation training loss threshold. When the loss value is greater than the defect generation training loss threshold, the network parameters of the defect feature generation network should be adjusted in reverse and training should be performed again until the loss value is no greater than the defect generation training loss threshold.
[0202] Specifically, when the loss value is no greater than the defect generation training loss threshold, an operator can perform manual visual judgment and subjectively determine whether the image quality of the defect generation training image meets the standard. If so, training is terminated, thereby obtaining a defect feature generation network after preliminary training. If not, the defect generation training loss threshold can be lowered, and the defect feature generation basic network model is trained again until the loss value is no greater than the defect generation training loss threshold, and the image quality of the defect generation training image is determined to meet the standard through observation. Generally, the defect generation training loss threshold can be set to 0.1.
[0203] From the above description, it can be seen that after obtaining the basic feature extraction network and defect feature generation network after preliminary training, the autoencoders used in the preliminary training can be removed respectively, and connected with the fusion decoding network according to the above description to construct the basic model for defect image generation.
[0204] In order to construct a basic model for defect image generation, the basic model training conditions should generally be configured. The configured basic model training conditions should include the basic model loss function, number of iterations, and optimizer, etc. The configured basic model training conditions can be selected according to needs. For example, you can refer to the training condition settings used when training the basic network for basic feature extraction mentioned above.
[0205] When training the defect image generation basic model using the basic model training dataset, a defect-free training sample is loaded into the basic feature extraction network, and a defect training sample is loaded into the defect feature generation network. A wire rope training defect image is then generated via the fusion decoding network. After a round of training of the defect image generation basic model using the basic model training dataset, a loss function for model training is calculated based on the wire rope training defect image and the second type of wire rope training image within the current defect training sample. The network parameters of the defect image generation basic model are adjusted based on the calculated loss function until the defect image generation basic model is trained to the target state. Specifically, the method for training the defect image generation basic model to the target state can be referred to the corresponding description above and will not be repeated here.
[0206] It should be noted that the defect image generation model required is generated based on the defect image generation base model that has reached the target state through model training, thereby achieving the construction of the defect image generation model. Of course, other methods can also be used to construct the defect image generation model. The specific method can be selected according to needs and will not be detailed here.
[0207] Furthermore, a device for generating a surface defect image of a wire rope can be obtained. In one embodiment of the present invention, the device comprises a defect image generating device, and a defect image generating model is deployed in the defect image generating device, wherein:
[0208] The defect image generation conditions of the wire rope are obtained, and the defect image generation device uses the above-mentioned method to perform defect image generation processing, and generates a target defect image of the wire rope after the defect image generation processing.
[0209] Specifically, the defect image generation device can utilize commonly available computer equipment. The type of defect image generation device can be selected based on needs, so as to meet the requirements for generating a target defect image of the wire rope. The defect image generation model can be deployed within the defect image generation device using methods commonly used in the art. Thereafter, the target defect image of the wire rope can be generated using the methods described above. For details, please refer to the above description and will not be repeated here.
Claims
1. A method for generating a surface defect image of a wire rope, characterized in that: The generation method comprises: Providing defect image generation conditions, and loading the provided defect image generation conditions into the constructed defect image generation model, wherein, The defect image generation model includes a basic feature extraction network, a defect feature generation network and a fusion decoding network. The defect image generation conditions include a complete reference image, defect feature text, and a defect reference image, wherein the type of the steel wire rope in the complete reference image is consistent with the type of the steel wire rope in the defect reference image; Loading the complete reference image into the basic feature extraction network to obtain a complete reference latent vector extracted by the basic feature extraction network, and loading the extracted complete reference latent vector into the fusion decoding network; Loading the defect feature text and the defect reference image into the defect feature generation network to generate a text-guided defect latent vector through the defect feature generation network, and loading the generated text-guided defect latent vector into the fusion decoding network; A fusion decoding network is used to perform fusion decoding processing on the flawless reference latent vector and the text-guided defect latent vector, so as to generate a wire rope target defect image after the fusion decoding processing, wherein the defect state of the wire rope target defect image is consistent with the defect state described by the defect feature text.
2. The method for generating a wire rope surface defect image according to claim 1, wherein: The basic feature extraction network is constructed based on at least a feature extraction Stable Diffusion model and an IP-Adapter module, wherein: For the loaded intact reference image, use the variational autoencoder in the feature extraction Stable Diffusion model to perform variational autoencoding on the intact reference image to generate a flawless encoding latent vector; The IP-Adapter module uses a decoupled cross-attention mechanism to convert the complete encoding latent vector into the complete reference cross-attention feature, and injects the complete reference cross-attention feature into the U-Net module of the feature extraction Stable Diffusion model to generate the complete reference latent vector after denoising by the U-Net module.
3. The method for generating a wire rope surface defect image according to claim 1, wherein: The defect feature generation network is constructed based on at least the defect generation Stable Diffusion model and the Controlnet module, wherein: Loading the defect reference image into the Controlnet module to generate defect constraints via the Controlnet module, and loading the defect constraints into the defect generation Stable Diffusion model; The defect reference image and the defect feature text are simultaneously loaded into the defect generation stable diffusion model, so that the defect generation stable diffusion model generates a text-guided defect latent vector under the defect constraint condition.
4. The method for generating a wire rope surface defect image according to claim 1, wherein: The fusion decoding network includes a latent vector fusion unit and a latent vector decoder, wherein, A latent vector fusion unit is used to perform latent vector fusion processing on the intact reference latent vector and the text-guided defect latent vector, so as to generate a wire rope fusion latent vector after the latent vector fusion processing; The latent vector decoder decodes the wire rope fusion latent vector to generate a wire rope target image after decoding.
5. The method for generating a surface defect image of a steel wire rope according to claim 4, wherein: When performing latent vector fusion processing, it includes: aligning the perfect reference latent vector with the text-guided defect latent vector in the latent space to generate a perfect aligned latent vector and a text-guided defect aligned latent vector after the latent space alignment; Performing frequency domain feature extraction on the above-mentioned perfect alignment latent vector and text-guided defect alignment latent vector to obtain perfect alignment frequency domain features and text-guided defect alignment frequency domain features, respectively, wherein the perfect alignment frequency domain features include perfect alignment low-frequency features and perfect alignment high-frequency features, and the text-guided defect alignment frequency domain features include defect alignment low-frequency features and defect alignment high-frequency features; Perform low-frequency fusion on the intact aligned low-frequency features and the defect aligned low-frequency features to generate aligned fused low-frequency features; Perform high-frequency fusion of the intact alignment high-frequency features and the defect alignment high-frequency features to generate aligned fusion high-frequency features; Based on the aligned fusion low-frequency features and the aligned fusion high-frequency features, the aligned combined frequency domain features are combined to generate aligned combined frequency domain features after the combination, and the generated aligned combined frequency domain features are transformed into the time domain to generate the wire rope fusion latent vector after the time domain transformation.
6. The method for generating a wire rope surface defect image according to claim 5, wherein: When the perfect reference latent vector is aligned with the text-guided defective latent vector in the latent space, we have: Among them, h IP is the perfect alignment latent vector, h CN Latent vector for text-guided defect alignment, W IP is the mapping matrix of the missing reference latent vector, W CN is the mapping matrix of the text-guided defect latent vector, b IP is the bias term of the missing reference latent vector, b CN is the bias term of the text-guided defect latent vector, z IP is the complete reference latent vector, z CN Bootstrapping defect latent vectors for text.
7. The method for generating a wire rope surface defect image according to claim 6, wherein: When generating aligned combined frequency domain features, we have: Among them, Z fuse To align the combined frequency domain features, To align and fuse low-frequency features, To align and fuse high-frequency features, α is the background fusion control coefficient, β is the defect enhancement coefficient, To perfectly align the low frequency features, To perfectly align high frequency features, Align low-frequency features for defects, Align high-frequency features for defects.
8. The method for generating a surface defect image of a steel wire rope according to any one of claims 1 to 7, wherein: When providing a defect reference image, the providing method includes: Provide defect reference source images, Convert the defect reference source image into the HSV color space and the Lab color space respectively, and then perform binarization operations in the HSV color space and the Lab color space respectively, so as to generate an HSV color space binarized image and a Lab color space binarized image respectively after the binarization operation; Performing a union operation on the HSV color space binary image and the Lab color space binary image to form a spatially fused binary image; Based on the HSV color space of the defect reference source image and the spatial fusion binary image, the hue change state in the HSV color space is calculated and determined; Based on the Lab color space of the defect reference source image and the spatial fusion binary image, calculate and determine the brightness change state in the Lab color space; Based on the calculated hue change state and brightness change state, When the hue changes significantly, hue separation is performed in the HSV color space of the defect reference source image, and a defect reference image is generated after separation; When the brightness change is significant, brightness enhancement is performed in the Lab color space of the defect reference source image, and a defect reference image is generated after the brightness enhancement.
9. The method for generating a surface defect image of a steel wire rope according to any one of claims 1 to 7, wherein: When building a defect image generation model, the construction method includes: Constructing a defect image generation basic model, wherein the defect image generation basic model includes a preliminarily trained basic feature extraction network, a preliminarily trained defect feature generation network, and an untrained fusion decoding network; Construct a basic model training dataset for training the basic model for defect image generation, where: The basic model training data set includes a non-defective training sample set and a defective training sample set. The defect-free training sample set includes a plurality of defect-free training samples, each of which includes a defect-free first-category steel wire rope training image; The defect training sample set includes a plurality of defect training samples, each defect training sample includes a second type of steel wire rope training image with defects and a training image text for characterizing the defect state of the second type of steel wire rope training image; The defect image generation basic model is trained based on the configured basic model training conditions, where: During model training, a defect-free training sample is loaded into the basic feature extraction network, and a defect training sample is loaded into the defect feature generation network. After that, a wire rope training defect image is generated through the fusion decoding network. Calculating a loss function for model training based on the wire rope defect training image and the second type of wire rope training image in the current defect training sample, and adjusting network parameters of the defect image generation basic model based on the calculated loss function until the defect image generation basic model is trained to a target state; The defect image generation base model that reaches the target state based on model training generates the required defect image generation model.
10. A device for generating a surface defect image of a steel wire rope, characterized in that: A defect image generation device is included, and a defect image generation model is deployed in the defect image generation device, wherein: The defect image generation conditions of the wire rope are obtained, and the defect image generation device adopts the method described in any one of claims 1 to 9 to perform defect image generation processing, and generates a target defect image of the wire rope after the defect image generation processing.
Citation Information
Patent Citations
Defect image generation method and device, electronic equipment and storage medium
CN115393231A
Steel wire rope surface defect identification method based on feature fusion
CN115565011A
Defect detection method, device and equipment based on few-sample learning and medium
CN119151928A
X-ray image defect detection method and device, equipment and storage medium
CN119251171A
Cited By
Defect image generation method and device, computer equipment and readable storage medium
CN121147200A
AI self-supervision compensation rope microdefect detection method
CN121391858A