Steel wire rope defect detection method and device, computer equipment and storage medium
By fusing image and text information using a multimodal large model, the problem of insufficient multimodal information fusion and low automation in wire rope defect detection in existing technologies is solved, enabling efficient and accurate diagnosis of the health status of wire ropes and improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202511423106.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-13
AI Technical Summary
Existing wire rope defect detection technologies suffer from several problems, including heavy reliance on single images or manual annotation, insufficient fusion of multimodal information, low accuracy in detecting internal and extremely small defects, and poor automation in online monitoring.
A multimodal large model is used for image feature encoding, text feature encoding, and modality fusion processing. EVA-CLIP-G and SkilNet-NLG are used to generate cross-modal representations. Data augmentation is performed by combining a controllable diffusion model. Defects are located and classified through cross-modal cross-attention mechanism and dynamic query.
It enables efficient, accurate, and robust diagnosis of the health status of wire ropes, improves detection efficiency and accuracy, reduces reliance on manual labeling, enhances the ability to identify subtle and internal defects, and supports the automation of online monitoring.
Smart Images

Figure CN121330359A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to large models, more specifically to a steel wire rope defect detection method and device, computer equipment and storage medium. BACKGROUND
[0002] Steel wire rope defect detection is crucial for ensuring the safety of various hoisting equipment and cableway systems, effectively preventing accidents caused by steel wire rope breakage. By timely discovering and addressing potential defects, the service life of steel wire ropes can be significantly extended, reducing maintenance costs.
[0003] Existing detection techniques include manual visual inspection, ultrasonic physical detection, and single-modal visual detection. Manual visual inspection is a method that relies on experienced technicians to identify steel wire rope defects. By using the naked eye, magnifying glass, or portable microscope, the steel wire rope is inspected section by section under different lighting conditions and observation angles to discover surface defects such as broken wires, cracks, local deformation, and corrosion. This method has the advantage of high flexibility, allowing the identification of some uncommon defects and the generation of detailed natural language descriptions as a reference for subsequent image labeling. However, its disadvantage is the high dependence on subjective judgment, making it difficult to implement large-scale online monitoring. Ultrasonic and other physical detection methods are mainly used to detect internal defects in steel wire ropes. By utilizing the propagation characteristics of high-frequency ultrasonic waves in materials, echoes are generated when cracks, cavities, or delaminations are encountered. By analyzing parameters such as arrival time, amplitude attenuation, and spectral features of these echoes, the location and depth of internal defects can be accurately determined. In addition, there are methods such as magnetic flux leakage detection and acoustic emission detection, which are used to detect surface or near-surface cracks and record stress wave signals generated during crack propagation, respectively. Although these methods can reveal "invisible to the naked eye" defects, they require professional time-series analysis and labeling processes. With the development of computer vision and deep learning technologies, single-modal visual-based steel wire rope defect detection technology has gradually matured. This technology uses high-resolution industrial cameras to capture steel wire rope surface images and optimizes image quality through a series of image processing techniques. Subsequently, advanced deep learning models are used for defect detection. Although this method performs well in real-time monitoring, it still faces challenges in identifying extremely small or deep defects.
[0004] Single-modal image data mainly includes visible light and infrared images, each of which has different advantages: RGB images provide high-resolution surface texture information, while infrared images can reveal crack tips or wear areas caused by thermal stress concentration. However, single-modal data is susceptible to environmental factors, limiting its application range. To overcome the limitations of single-modal images, researchers have proposed multi-modal image fusion strategies to combine RGB and infrared images to improve the robustness of detection through different fusion stages (early, middle, and late). However, this approach also presents additional challenges, such as sensor calibration, spatio-temporal alignment, and data imbalance. The labeling methods for steel wire rope defects include rectangular bounding boxes, polygons, and pixel-level segmentation masks, each with its own application scenarios. Although high-precision labeling helps improve model performance, it also increases labor costs and may result in biased results due to differences between annotators.
[0005] As can be seen, current steel wire rope defect detection technologies have their own advantages and disadvantages. Manual visual inspection is inefficient, physical detection methods such as ultrasonic waves are costly and complex to operate, single-modal visual detection is susceptible to environmental factors, and multi-modal image fusion faces more technical and cost challenges. Additionally, constructing a high-quality dataset is a significant challenge, mainly due to sample scarcity, single modality, and time-consuming and labor-intensive labeling processes.
[0006] Therefore, it is necessary to design a new method to overcome the deficiencies of existing technologies, such as heavy reliance on single images or manual annotation, insufficient multi-modal information fusion, low detection accuracy for internal and extremely small defects, and poor online monitoring automation, to achieve efficient, accurate, and robust diagnosis of steel wire rope health status. SUMMARY
[0007] The purpose of the present application is to overcome the deficiencies of the prior art and provide a steel wire rope defect detection method, device, computer equipment, and storage medium.
[0008] To achieve the above-mentioned purpose, the present application adopts the following technical solution: a steel wire rope defect detection method, comprising:
[0009] Obtaining a picture to be detected and a prompt text;
[0010] Inputting the picture to be detected and the prompt text into a multi-modal large model for image feature encoding, text feature encoding, modal fusion processing, and positioning and classifying different types of defects to obtain a detection result;
[0011] Outputting the detection result.
[0012] A further technical solution is that the multi-modal large model includes an image information encoder, a text information encoder, a modal fusion module, a detector module, and a prompt word module.
[0013] The image information encoder maps the to-be-detected picture to a high-dimensional hidden space shared with natural language description to form a visual feature;
[0014] The text information encoder extracts a semantic vector related to the defect in the prompt text to form a text feature;
[0015] The modal fusion module fuses the visual feature and the text feature to generate a cross-modal expression to obtain a fusion feature;
[0016] The detector module locates and classifies different types of defects based on the fusion feature;
[0017] The prompt word module automatically adjusts the attention weight according to the business knowledge to improve the positioning accuracy and the classification accuracy.
[0018] Further technical solutions thereof are as follows: the to-be-detected picture and the prompt text are input into a multi-modal large model for image feature encoding, text feature encoding, modal fusion processing, and locating and classifying different types of defects to obtain a detection result, including:
[0019] The EVA-CLIP-G is used to map the to-be-detected picture to a high-dimensional hidden space shared with natural language description to obtain a visual feature;
[0020] The SkilNet-NLG is used to process the prompt text containing specific defect information to extract a semantic vector related to the defect to obtain a text feature;
[0021] The cross-modal cross-attention mechanism is used to fuse the visual feature and the text feature to generate a consistent cross-modal expression to obtain a fusion feature;
[0022] Based on the fusion feature, a dynamic query and a multi-head cross-attention layer are used to locate and classify different types of defects to obtain a detection result.
[0023] Further technical solutions thereof are as follows: the training process of the multi-modal large model includes:
[0024] A multi-source data is acquired, and a basic data set is constructed based on the multi-source data;
[0025] A controllable diffusion model is used to perform data augmentation based on the basic data set to obtain a sample set;
[0026] An initial multi-modal large model is constructed;
[0027] The initial multi-modal large model is trained using the sample set to obtain a multi-modal large model.
[0028] The further technical solution is as follows: the acquisition of multi-source data and the construction of a basic dataset based on the multi-source data include:
[0029] Data on fault detection, missed detection, and normal operation are collected synchronously from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs to obtain multi-source data, which includes fault types and corresponding images.
[0030] The multi-source data is preprocessed and stored;
[0031] The stored multi-source data is labeled to obtain the basic dataset.
[0032] The further technical solution is as follows: the controllable diffusion model includes a semantic control layer, a physical constraint module, and a dynamic noise scheduling layer;
[0033] The semantic control layer encodes the defect type and location parameters into a conditional vector, and integrates it into the diffusion model through the Cross-Attention mechanism to guide the generation of images with specific defects.
[0034] The physical constraint module is used to embed prior knowledge of the physical structure of the steel wire rope into the UNet architecture and enhance it through the weights of the pre-trained segmentation model to ensure that the generated defect images meet the actual physical characteristics and standard requirements.
[0035] Dynamic noise scheduling is used to optimize the noise addition strategy during the diffusion process using adaptive scheduling.
[0036] The further technical solution is as follows: the data augmentation based on the basic dataset using a controllable diffusion model to obtain a sample set includes:
[0037] The basic dataset is combined with user input parameters, and a multimodal encoder is used to transform the semantic information of defect category, size, and location into feature vectors readable by the generative model. A controllable diffusion model is then used to automatically generate diverse defect images to obtain a sample set.
[0038] The present invention also provides a wire rope defect detection device, comprising:
[0039] The acquisition unit is used to acquire the image to be detected and the prompt text;
[0040] The detection unit is used to input the image to be detected and the prompt text into a multimodal large model to perform image feature encoding, text feature encoding, modal fusion processing, and to locate and classify different types of defects in order to obtain detection results;
[0041] The output unit is used to output the detection results.
[0042] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0043] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0044] The advantages of this invention compared to existing technologies are as follows: By integrating a multimodal large model and utilizing the image to be detected and the prompt text for image feature encoding, text feature encoding, and modal fusion processing, this invention achieves precise location and classification of different types of defects. This overcomes the problems of existing technologies, such as heavy reliance on single images or manual annotation, insufficient multimodal information fusion, low detection accuracy of internal and extremely small defects, and poor automation in online monitoring. Specifically, this method not only fully utilizes the complementary advantages of image and text information, improving the robustness and generalization ability of defect identification in complex environments, but also significantly improves detection efficiency and accuracy through automated data processing, achieving efficient and accurate diagnosis of the health status of wire ropes. This method greatly reduces reliance on manual annotation, enhances the model's ability to identify minute and internal defects, and supports the automation of online monitoring, providing strong technical support for the safety management of wire ropes.
[0045] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a schematic diagram illustrating an application scenario of the wire rope defect detection method provided in this embodiment of the invention.
[0048] Figure 2 A schematic flowchart of the wire rope defect detection method provided in an embodiment of the present invention;
[0049] Figure 3 A schematic block diagram of a wire rope defect detection device provided in an embodiment of the present invention;
[0050] Figure 4 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0053] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0054] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0055] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the wire rope defect detection method provided in this embodiment of the invention. Figure 2 This is a schematic flowchart illustrating the wire rope defect detection method provided in this embodiment of the invention. The method is applied in a server. The server interacts with the terminal, and by combining multimodal large-scale model processing technology, it effectively overcomes the shortcomings of existing technologies, such as heavy reliance on single images or manual annotation, insufficient multimodal information fusion, low accuracy in detecting internal and extremely small defects, and poor automation in online monitoring. Specifically, it first uses a multimodal large-scale model, including an image information encoder, a text information encoder, a modal fusion module, a detector module, and a prompt word module, to process the acquired image to be detected and its related prompt text, achieving efficient encoding and fusion of image features and text features, thereby accurately locating and classifying different types of wire rope defects. Furthermore, a controllable diffusion model is used for data augmentation, which not only enriches the diversity of training samples but also optimizes the quality of generated images through a semantic control layer, a physical constraint module, and dynamic noise scheduling, ensuring that they conform to actual physical characteristics. Ultimately, this series of techniques work together to achieve efficient, accurate, and robust diagnosis of the health status of wire ropes, greatly improving the automation level and accuracy of wire rope defect detection.
[0056] Figure 2 This is a schematic flowchart of the wire rope defect detection method provided in an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S130.
[0057] S110. Obtain the image to be detected and the prompt text.
[0058] In this embodiment, the images to be inspected refer to surface images of wire ropes collected through various methods (such as online monitoring systems, manual inspections, and on-site feedback). These high-resolution RGB images have a resolution of 4096×2160 and cover a variety of typical application scenarios, including wind power, mines, bridges, cable cars, cranes, and tower cranes. Each image undergoes preprocessing steps, including outlier filtering, missing field interpolation and deduplication, and is stored after uniformly formatting timestamps and geographic location information. In addition, these images also contain pixel-level masking and morphological parameter annotation information for wire rope defects (such as broken wires, wear, and corrosion).
[0059] The cue text is a natural language description provided for a specific defect or feature in the image to be detected. This description can specify the defect's location (e.g., "top left area," "bottom right corner"), type (e.g., "crack," "corrosion," "broken wire"), and size (e.g., "approximately 120×15 pixels"). The cue text not only enhances the model's understanding of the image but also guides multimodal large models to more accurately locate and classify different types of defects. To improve diversity and accuracy, the cue text is typically generated using templates and combined with a thesaurus to increase variation, for example: "A crack appears in the top left area of the image, approximately 120×15 pixels," or "A severe corrosion is visible in the top center, covering an area of approximately 80×20 pixels." These cue texts are ultimately linked with the corresponding image file and its annotation using a unique ID to form a high-quality "cue word-image" training pair, used for fine-tuning and optimization of the downstream visual-language localization model. In this way, the model can better understand and identify subtle defects in the image, improving overall detection accuracy and robustness.
[0060] S120. Input the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modal fusion processing, and location and classification of different types of defects to obtain the detection result.
[0061] In this embodiment, the detection result refers to the final output information after processing by the aforementioned multimodal large model, which includes the location (boundary box coordinates), type (such as broken wires, cracks, corrosion, etc.), and severity (such as size, damage level, etc.) of various defects on the wire rope surface. This information not only helps to accurately identify and classify various defects, but also provides a scientific basis for subsequent maintenance and repair work, thereby effectively improving the efficiency and accuracy of monitoring the health status of the wire rope.
[0062] Specifically, the multimodal large model includes an image information encoder, a text information encoder, a modality fusion module, a detector module, and a prompt word module;
[0063] Specifically, the image information encoder maps the image to be detected to a high-dimensional latent space shared with the natural language description, forming visual features; it also maps the image to be detected (high-resolution RGB steel wire rope image) to a high-dimensional latent space shared with the natural language description, forming visual features. Using EVA-CLIP-G as the core component, pre-trained weights are used to transform the image features into a representation that aligns with the natural language description, capturing texture, shape, and global semantic information.
[0064] The text information encoder extracts semantic vectors related to defects from the prompt text to form text features; it also extracts semantic vectors related to defects from the prompt text (a natural language description of a specific defect or feature) to form text features. Using SkilNet-NLG as the text encoder, it efficiently extracts semantic information directly related to the location and classification of wire rope defects from the text by activating only the sub-modules most relevant to the task through a sparse activation mechanism.
[0065] The modality fusion module fuses visual and textual features to generate cross-modal representations, resulting in fused features. It also fuses visual features generated by an image encoder and textual features generated by a text encoder to generate cross-modal representations, again resulting in fused features. This module, based on BERT's two-stream joint attention network architecture, achieves deep fusion across multiple cross-modal attention layers, ensuring full interaction between visual and linguistic features at both spatial and semantic levels.
[0066] The detector module locates and classifies different types of defects based on these fused features. First, the fused features are mapped to a spatial pyramid feature map (FPN) to accommodate defects at different scales. Then, a dynamic query vector is extracted from the [CLS] vector of the text prompt, carrying three types of prompt information: "location," "type," and "size," as prior guidance for localization. Through a multi-head cross-attention mechanism, the query vector interacts with the features at each layer, ultimately generating bounding box regression results and defect category scores. These are then processed using non-maximum suppression (NMS) or a learned ensemble strategy to output the final detection box.
[0067] The prompt word module automatically adjusts attention weights based on business knowledge, improving localization accuracy and classification accuracy. A soft prompting mechanism is employed, encoding the business knowledge required for wire rope detection into a set of continuously learnable vectors, which are then input into the model along with the original text and visual features. During fine-tuning, by minimizing downstream localization and classification losses, the prompt vectors gradually learn the representations that best activate the relevant skill modules, significantly improving the model's performance in real-world application scenarios.
[0068] Through this multi-level and multi-dimensional comprehensive processing approach, this embodiment can not only significantly improve the accuracy and robustness of wire rope defect detection, but also significantly reduce the cost of manual annotation, quickly adapt to new scenarios, and enhance the model's few-sample generalization ability.
[0069] In one embodiment, step S120 described above may include steps S121 to S124.
[0070] S121. The image to be detected is mapped to a high-dimensional latent space shared with the natural language description using EVA-CLIP-G to obtain visual features.
[0071] In this embodiment, visual features refer to the high-level semantic representation generated after encoding the input 4096×2160 resolution steel wire rope image using the EVA-CLIP-G model. This representation not only captures low-level visual features in the image (such as color and texture) but also includes high-level conceptual information (such as shape and global structure). Specifically, EVA-CLIP-G uses its pre-trained weights to map the steel wire rope surface image to a high-dimensional latent space shared with natural language descriptions, enabling defect concepts such as "broken wire," "crack," and "corrosion" to be directly compared with image features, thereby providing a more stable representation and higher defect recognition accuracy.
[0072] S122. The prompt text containing specific defect information is processed by SkilNet-NLG to extract semantic vectors related to the defects in order to obtain text features.
[0073] In this embodiment, text features refer to semantic information directly related to the location and classification of wire rope defects extracted from specific defect warning text (e.g., "A crack appears in the upper left area, approximately 120×15 pixels"). SkilNet-NLG is used as the text encoder, employing a Transformer architecture and implementing sparse activation by introducing skill modules. Based on the content of the warning text, the model only activates the skill modules most relevant to it (such as location description, defect type identification, and size quantization), thereby efficiently extracting key semantic vectors from the text. These vectors represent the specific information about the defect contained in the warning text and can be effectively fused with image features.
[0074] S123. The visual features and text features are fused through a cross-modal cross-attention mechanism to generate a consistent cross-modal representation, thereby obtaining fused features.
[0075] In this embodiment, fused features refer to semantic information directly related to the location and classification of wire rope defects extracted from specific defect warning text (e.g., "A crack appears in the upper left area, approximately 120×15 pixels"). SkilNet-NLG is used as the text encoder, employing a Transformer architecture and implementing sparse activation by introducing skill modules. Based on the content of the warning text, the model only activates the skill modules most relevant to it (such as location description, defect type identification, size quantization, etc.), thereby efficiently extracting key semantic vectors from the text. These vectors represent the specific information about the defect contained in the warning text and can be effectively fused with image features.
[0076] S124. Based on the fusion features, use dynamic query and multi-head cross-attention layer to locate and classify different types of defects to obtain detection results.
[0077] Extract K dynamic query vectors carrying three types of prompt information: "direction", "type", and "size" from the [CLS] vector of the text prompts and the output of each skill module.
[0078] Each dynamic query interacts sequentially with features from each layer of the pyramid through a multi-head cross-attention layer. After combining spatial and semantic information, it enters two parallel small feedforward networks for bounding box regression and defect category scoring. Finally, the decoder transforms the regression results to the image coordinate system and applies non-maximum suppression or a learned ensemble strategy based on confidence, outputting several high-confidence detection boxes.
[0079] After processing through the above steps, the output includes information on the location (boundary box coordinates), type (such as broken wires, cracks, corrosion, etc.), and severity (such as size, damage level, etc.) of various defects on the wire rope surface. This information not only helps to accurately identify and classify various defects, but also provides a scientific basis for subsequent maintenance and repair work, thereby effectively improving the efficiency and accuracy of monitoring the health status of the wire rope.
[0080] In this embodiment, as Figure 1 As shown, EVA-CLIP-G is used as the core component of the image feature encoder. EVA-CLIP-G is pre-trained on hundreds of millions of "image-text" data sets, enabling the mapping of wire rope surface images to a high-dimensional latent space shared with natural language descriptions. This allows for direct comparison between defect concepts such as "broken wire," "cracks," and "corrosion" and image features. Unlike traditional methods that only focus on low-level visual features (such as VQ-VAE), EVA-CLIP-G captures texture, shape, and global semantic information, is compatible with rich language cues, and provides a more stable representation in the detection of subtle cracks or microscopic damage. In practice, EVA-CLIP-G is first used to encode the input 4096×2160 resolution wire rope image, and its [CLS] embedding is taken as the global visual vector. This vector is then aligned with the text vector obtained by SkilNet-NLG encoding the defect cues through a linear mapping layer and fed into a Joint Attention Network for cross-modal attention fusion. This design not only improves the recall and accuracy of defect detection, but also enhances the model's generalization ability under conditions of a small number of real samples.
[0081] This paper employs SkilNet-NLG, a Transformer architecture, as the text feature encoder and introduces "skill modules" to achieve sparse activation. These skill modules include "location description," "defect type identification," "size quantization," and "severity assessment," ensuring that the model activates only relevant sub-modules when loading a task, avoiding the negative transfer problem encountered by traditional multi-task models when handling weakly related tasks. For example, for the prompt "a crack appears in the upper left area, approximately 120×15 pixels," the model activates the corresponding three skill modules, while others remain silent. Finally, the output vector at the [CLS] position serves as the global semantic representation of the prompt, which can be concatenated with image features or fused with cross-attention for subsequent defect localization and classification head training. This method improves the accuracy and efficiency of aligning text prompts with image features.
[0082] The multimodal fusion module employs a BERT-based dual-stream joint attention network architecture. Image and text information are extracted independently by their respective Transformer encoders and then deeply fused across multiple cross-modal attention layers. Specifically, the high-resolution steel wire rope surface image is first segmented into several candidate regions, and the EVA-CLIP-G model is used to generate visual feature vectors. Simultaneously, the defect cues corresponding to the image are fed into the BERT text encoder, generating context-aware semantic vectors through the Masked Language Modeling task. Then, bidirectional multi-head attention operations are performed in N layers of joint attention units to ensure sufficient interaction between visual and linguistic features at both spatial and semantic levels, forming a highly consistent cross-modal representation. During the pre-training phase, three tasks are optimized in parallel: region mask prediction, text mask prediction, and image-text matching. After pre-training, fine-tuning is performed using a small amount of labeled data, achieving accurate localization and classification of various steel wire rope defects.
[0083] The detector module generates the final defect detection box based on cross-modal fusion features. First, the visual-linguistic features fused by a joint attention network are mapped to a spatial pyramid feature map (FPN) to accommodate cracks and broken wires at different scales. Second, K dynamic query vectors are extracted from the [CLS] vector of the text prompts and the outputs of each skill module, carrying three types of prompt information—"location," "type," and "size"—as prior guidance for localization. In the detector backbone, each dynamic query interacts sequentially with the features of each layer of the pyramid through a multi-head cross-attention layer. After fusing spatial and semantic information, the query vector enters two parallel small feedforward networks: one for bounding box regression and the other for defect category scoring. All query results are transformed to the image coordinate system by the decoder and subjected to non-maximum suppression or a learned ensemble strategy based on confidence, ultimately outputting a high-confidence detection box.
[0084] To efficiently incorporate the business knowledge required for wire rope inspection (such as "strand location," "defect type," and "severity level") into a multimodal large-scale model, a "soft cue" mechanism was adopted. First, a set of basic cue word templates were defined and converted into initial text embeddings using SkilNet-NLG. During model fine-tuning, all parameters of the large-scale pre-trained model were frozen, and gradient updates were performed only on the cue embedding vectors to minimize downstream localization and classification losses. The visual features of each image and the corresponding soft cue vector interact in a joint attention layer, automatically adjusting the attention weights based on the business information in the cue, focusing more on the potential defect locations and shapes. Experiments show that this continuous vector-based cue design significantly improves localization accuracy and classification accuracy while reducing the cost of manual parameter tuning.
[0085] For the aforementioned multimodal large model, the training process includes:
[0086] Step 1: Obtain multi-source data and construct a basic dataset based on the multi-source data;
[0087] Step 2: Based on the aforementioned basic dataset, perform data augmentation using a controlled diffusion model to obtain a sample set;
[0088] Step 3: Construct the initial multimodal large model;
[0089] Step 4: Train the initial multimodal large model using the sample set to obtain the multimodal large model.
[0090] Step one may include:
[0091] Data on fault detection, missed detection, and normal operation are collected synchronously from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs to obtain multi-source data, which includes fault types and corresponding images.
[0092] The multi-source data is preprocessed and stored;
[0093] The stored multi-source data is labeled to obtain the basic dataset.
[0094] The controllable diffusion model includes a semantic control layer, a physical constraint module, and a dynamic noise scheduling layer;
[0095] The semantic control layer encodes the defect type and location parameters into a conditional vector, which is then fused into the diffusion model via a cross-attention mechanism to guide the generation of images with specific defects.
[0096] The physical constraint module is used to embed prior knowledge of the physical structure of the steel wire rope into the UNet architecture and enhance it through weights of the pre-trained segmentation model to ensure that the generated defect images meet the actual physical characteristics and standard requirements.
[0097] Dynamic noise scheduling is used to optimize the noise addition strategy during the diffusion process using adaptive scheduling.
[0098] In one embodiment, step two may include:
[0099] The basic dataset is combined with user input parameters, and a multimodal encoder is used to transform the semantic information of defect category, size, and location into feature vectors readable by the generative model. A controllable diffusion model is then used to automatically generate diverse defect images to obtain a sample set.
[0100] During training, data on misdetections, missed detections, and normal operation are collected periodically from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs. Detailed metadata, including collection time, location, equipment model, and fault type, is recorded. The data acquisition end filters outliers, interpolates missing fields, and removes duplicates from the raw data. Timestamps and geographic location information are then uniformly formatted and categorized for storage in a relational database or distributed file system, with scheduled backup and recovery strategies implemented. To ensure high accuracy and consistency of image labels, a rigorous annotation manual is developed. Combining pre-trained model suggestions and Canny operator candidate region extraction, detailed polygon / rectangular bounding box annotations are completed by annotation personnel, and quality is strictly controlled through sampling quality checks and a double-check mechanism.
[0101] In the model training phase, a subset is first used for pre-training. Based on key metrics such as accuracy, recall, and robustness, a high-quality training set is constructed by selecting baseline samples that meet performance requirements. Subsequently, cross-validation is performed on the full dataset to verify generalization ability. Finally, the data collection frequency, scenario coverage, and annotation standards are dynamically adjusted based on feedback from field applications, constructing a closed loop of "data collection-preprocessing-storage-annotation-data filtering-training-feedback" to continuously improve data quality and model performance.
[0102] The semantic control layer encodes defect types such as broken wires, wear, and corrosion, as well as information such as polar radius r, azimuth angle θ, and size parameters, into a 128-dimensional conditional vector, which is then injected into the diffusion process through the Cross-Attention mechanism.
[0103] The physical constraint module embeds prior knowledge such as the helical arrangement and core topology of the wire rope into the jump connections of the UNet and loads pre-trained segmentation model weights to ensure that the synthesized textures conform to the ASTM E1572 standard.
[0104] The dynamic noise scheduling uses an adaptive scheduling function, which can optimize high-frequency details in the defect area after 1000 diffusion steps, so that the peak signal-to-noise ratio (PSNR) of the generated image reaches more than 32dB.
[0105] To construct a training set covering multiple scenarios, over 500 real steel wire rope images with a resolution of 4096×2160 were first collected, covering six typical application scenarios: wind power, mining, bridges, cable cars, cranes, and tower cranes. Pixel-level masks and morphological parameter annotations were performed for 12 defect types. Simultaneously, synthetic defect images were generated using physical simulations with COMSOL and Ansys Fluent, constructing a pre-training set containing 2000 synthetic samples.
[0106] To achieve controllable separation of base texture and defect deformation, a VAE encoder is introduced to split each image into content latent code (base texture) and defect latent code (local deformation) to support independent control and diverse combinations of different defect parameters.
[0107] The first stage involves pre-training on a synthetic dataset (Batch Size = 32, Learning Rate = 5 × 10). -5 The second stage involves fine-tuning on a real dataset (50 rounds). (Batch Size = 16, Learning Rate = 1 × 10⁻⁶). -6 (30 rounds). Only the UNet ontology and noise scheduling parameters were unfrozen, while the weights of the semantic control layer and physical constraint module remained fixed. The optimization objective was a hybrid loss, including standard diffusion reconstruction loss, perceptual loss based on VGG16 feature maps, and consistency loss measured by comparing the generated and real defect histograms using Hough transform. Through this architecture and strategy, the model converged quickly with a small number of real samples, accurately reproducing various minute and deep defect features while maintaining high-fidelity synthesis. After the conditional diffusion process, the model generated 5000 high-fidelity synthesized images. Subsequently, the Canny operator was used for edge screening on all synthesized samples, and the images were manually reviewed by annotators to remove artifacts and images that did not conform to physical laws, ultimately forming a high-quality augmented dataset of 5000 images, providing sufficient and diverse sample support for subsequent detection model training.
[0108] Based on 5000 labeled images and their defect border information, a "prompt word-image" training pair is automatically generated. First, the coordinates (x, y, w, h) of each defect border are parsed and mapped to its relative position in the image (e.g., top left, center right, bottom center, etc.). Then, the border size (w, h) is converted to an "approximately w × h pixels" format and combined with the defect type (broken wire / crack / corrosion, etc.) to fill a predefined template, such as "A crack appears in the top left area of the image, approximately 120 × 15 pixels." To enhance diversity, 10 synonym templates (e.g., "A… is visible," "A… defect was detected") and a synonym library for type and location ("crack" / "fissure," "top right" / "upper right," etc.) are prepared, and these are randomly replaced to generate the final prompt.
[0109] All prompts and their corresponding image files and annotations are linked together by a unique ID. During the trial period, 200 pairs were randomly selected for manual review and error correction. The templates and mapping rules were adjusted, and finally, thousands of high-quality "prompt word-image" training pairs were output for fine-tuning of the downstream visual-language localization model.
[0110] The training data for existing technologies is scarce, especially defective samples with diversity and complexity.
[0111] This embodiment combines a wire rope defect type library with user input parameters, and utilizes a multimodal encoder to transform semantic information such as defect category, size, and location into feature vectors readable by the generated model. Based on this, a Controlled Diffusion Model (CDM) is used to automatically generate diverse defect images. After generation, the images undergo both physical plausibility post-processing and manual review, significantly increasing the quantity and diversity of defect images while ensuring that the new samples accurately reflect real-world working conditions, thus completely alleviating the performance constraints imposed by the scarcity of training data.
[0112] Existing technologies, relying on a single modality (such as images only), struggle to fully capture defect information, and their detection accuracy is low in complex environments (such as changes in lighting or occlusion).
[0113] In this embodiment, a high-resolution RGB steel wire rope image and natural language prompts are fed in parallel into a unified multimodal encoder to form a fused visual-linguistic feature representation. Based on this, a large-scale visual-linguistic model centered on VIT-NLG uses a cross-attention mechanism to accurately locate and classify complex defects such as micro-cracks and localized corrosion. This multimodal approach not only maximizes the guiding value of text prompts for image features but also significantly improves the model's robustness and generalization ability under varying lighting conditions, occlusion, and background interference.
[0114] S130, Output the detection results.
[0115] The test results are output to the terminal for display.
[0116] In this embodiment, by constructing a multimodal dataset driven by "image-text" soft prompts, and combining high-fidelity augmented samples generated by controlled diffusion with natural language prompts, the variety and semantic dimension of the training data are enriched, enabling the model to maintain high accuracy in complex scenarios.
[0117] By using a controllable diffusion model to automatically generate synthetic images and complete the annotation process through a semi-automated workflow, and then using templated regular expressions to generate text labels in batches, the workload of manually writing prompts is greatly reduced, achieving an efficiency improvement of several times compared to traditional methods.
[0118] By employing soft hints and sparse activation text encoder technology, only a few hint vectors and joint attention layer parameters need to be fine-tuned to quickly adapt to new scenarios and achieve real-time inference at the second level on edge devices, significantly reducing computing power and development costs.
[0119] By combining physical priors with large-scale "image-text" pre-training, the model can achieve performance close to that of full-scale training in the fine-tuning stage with a small number of real samples, and has stronger generalization ability across scenes and defect types.
[0120] The aforementioned wire rope defect detection method integrates a multimodal large model and utilizes image feature encoding, text feature encoding, and modal fusion processing based on the image to be detected and the prompt text. This enables precise location and classification of different types of defects, overcoming the problems of existing technologies such as heavy reliance on single images or manual annotation, insufficient multimodal information fusion, low detection accuracy of internal and extremely small defects, and poor automation of online monitoring. Specifically, this method not only fully utilizes the complementary advantages of image and text information, improving the robustness and generalization ability of defect recognition in complex environments, but also significantly improves detection efficiency and accuracy through automated data processing, achieving efficient and accurate diagnosis of the health status of wire ropes. This method greatly reduces reliance on manual annotation, enhances the model's ability to identify minute and internal defects, and supports the automation of online monitoring, providing strong technical support for the safety management of wire ropes.
[0121] Figure 3 This is a schematic block diagram of a wire rope defect detection device 300 provided in an embodiment of the present invention. Figure 3 As shown, corresponding to the above-described wire rope defect detection method, the present invention also provides a wire rope defect detection device 300. This wire rope defect detection device 300 includes a unit for performing the above-described wire rope defect detection method, and the device can be configured in a server. Specifically, please refer to... Figure 3 The wire rope defect detection device 300 includes an acquisition unit 301, a detection unit 302, and an output unit 303.
[0122] The acquisition unit 301 is used to acquire the image to be detected and the prompt text; the detection unit 302 is used to input the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modal fusion processing, and to locate and classify different types of defects to obtain the detection result; the output unit 303 is used to output the detection result.
[0123] In one embodiment, the detection unit 302 includes a visual feature extraction subunit, a text feature extraction subunit, a fusion feature subunit, and a result detection subunit.
[0124] The visual feature extraction subunit is used to map the image to be detected to a high-dimensional latent space shared with the natural language description using EVA-CLIP-G to obtain visual features; the text feature extraction subunit is used to process the prompt text containing specific defect information through SkilNet-NLG to extract semantic vectors related to the defects to obtain text features; the feature fusion subunit is used to fuse the visual features and the text features through a cross-modal cross-attention mechanism to generate a consistent cross-modal expression to obtain fused features; and the result detection subunit is used to locate and classify different types of defects based on the fused features using dynamic query and multi-head cross-attention layers to obtain detection results.
[0125] In one embodiment, the device further includes:
[0126] The model training unit is used to acquire multi-source data and construct a basic dataset based on the multi-source data; perform data augmentation using a controllable diffusion model based on the basic dataset to obtain a sample set; construct an initial multimodal large model; and train the initial multimodal large model using the sample set to obtain a multimodal large model.
[0127] In one embodiment, the model training unit is further configured to synchronously collect data on false detections, missed detections, and normal operation from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs to obtain multi-source data, wherein the multi-source data includes fault types and corresponding images; preprocess the multi-source data and store it; and label the stored multi-source data to obtain a basic dataset.
[0128] In one embodiment, the model training unit is used to combine the basic dataset with user input parameters, use a multimodal encoder to transform the semantic information of defect category, size, and location into feature vectors readable by the generated model, and use a controllable diffusion model to automatically generate diverse defect images to obtain a sample set.
[0129] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned wire rope defect detection device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0130] The aforementioned wire rope defect detection device 300 can be implemented as a computer program, which can perform tasks such as... Figure 4 It runs on the computer device shown.
[0131] Please see Figure 4 , Figure 4This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0132] See Figure 4 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0133] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a wire rope defect detection method.
[0134] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0135] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a wire rope defect detection method.
[0136] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0137] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:
[0138] Acquire the image to be detected and the prompt text; input the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modal fusion processing, and localization and classification of different types of defects to obtain the detection result; output the detection result.
[0139] The multimodal large model includes an image information encoder, a text information encoder, a modality fusion module, a detector module, and a prompt word module;
[0140] The image information encoder maps the image to be detected to a high-dimensional latent space shared with the natural language description to form visual features; the text information encoder extracts semantic vectors related to defects from the prompt text to form text features; the modality fusion module fuses visual features and text features to generate cross-modal expressions and obtain fused features; the detector module locates and classifies different types of defects based on these fused features; and the prompt word module automatically adjusts attention weights according to business knowledge to improve positioning accuracy and classification accuracy.
[0141] In one embodiment, when the processor 502 implements the step of inputting the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modality fusion processing, and locating and classifying different types of defects to obtain detection results, the specific implementation steps are as follows:
[0142] The image to be detected is mapped to a high-dimensional latent space shared with the natural language description using EVA-CLIP-G to obtain visual features. The prompt text containing specific defect information is processed by SkilNet-NLG to extract semantic vectors related to the defects, thus obtaining text features. The visual features and text features are fused through a cross-modal cross-attention mechanism to generate a consistent cross-modal representation, thus obtaining fused features. Based on the fused features, dynamic querying and multi-head cross-attention layers are used to locate and classify different types of defects to obtain detection results.
[0143] In one embodiment, when implementing the training steps of the multimodal large model, the processor 502 specifically implements the following steps:
[0144] Acquire multi-source data and construct a basic dataset based on the multi-source data; perform data augmentation using a controllable diffusion model based on the basic dataset to obtain a sample set; construct an initial multimodal large model; train the initial multimodal large model using the sample set to obtain a multimodal large model.
[0145] The controllable diffusion model includes a semantic control layer, a physical constraint module, and a dynamic noise scheduling layer;
[0146] The semantic control layer encodes the defect type and location parameters into a conditional vector, which is then fused into the diffusion model via a cross-attention mechanism to guide the generation of images with specific defects.
[0147] The physical constraint module is used to embed prior knowledge of the physical structure of the steel wire rope into the UNet architecture and enhance it through weights of the pre-trained segmentation model to ensure that the generated defect images meet the actual physical characteristics and standard requirements.
[0148] Dynamic noise scheduling is used to optimize the noise addition strategy during the diffusion process using adaptive scheduling.
[0149] In one embodiment, when implementing the step of acquiring multi-source data and constructing a basic dataset based on the multi-source data, the processor 502 specifically implements the following steps:
[0150] Data on fault detection, missed detection, and normal operation are collected synchronously from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs to obtain multi-source data, which includes fault types and corresponding images. The multi-source data is preprocessed and stored. The stored multi-source data is then labeled to obtain a basic dataset.
[0151] In one embodiment, when implementing the step of performing data augmentation using a controlled diffusion model based on the base dataset to obtain a sample set, the processor 502 specifically implements the following steps:
[0152] The basic dataset is combined with user input parameters, and a multimodal encoder is used to transform the semantic information of defect category, size, and location into feature vectors readable by the generative model. A controllable diffusion model is then used to automatically generate diverse defect images to obtain a sample set.
[0153] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0154] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0155] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:
[0156] Acquire the image to be detected and the prompt text; input the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modal fusion processing, and localization and classification of different types of defects to obtain the detection result; output the detection result.
[0157] The multimodal large model includes an image information encoder, a text information encoder, a modality fusion module, a detector module, and a prompt word module;
[0158] The image information encoder maps the image to be detected to a high-dimensional latent space shared with the natural language description to form visual features; the text information encoder extracts semantic vectors related to defects from the prompt text to form text features; the modality fusion module fuses visual features and text features to generate cross-modal expressions and obtain fused features; the detector module locates and classifies different types of defects based on these fused features; and the prompt word module automatically adjusts attention weights according to business knowledge to improve positioning accuracy and classification accuracy.
[0159] In one embodiment, when the processor executes the computer program to implement the steps of inputting the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modality fusion processing, and locating and classifying different types of defects to obtain detection results, the specific implementation is as follows:
[0160] The image to be detected is mapped to a high-dimensional latent space shared with the natural language description using EVA-CLIP-G to obtain visual features. The prompt text containing specific defect information is processed by SkilNet-NLG to extract semantic vectors related to the defects, thus obtaining text features. The visual features and text features are fused through a cross-modal cross-attention mechanism to generate a consistent cross-modal representation, thus obtaining fused features. Based on the fused features, dynamic querying and multi-head cross-attention layers are used to locate and classify different types of defects to obtain detection results.
[0161] In one embodiment, when the processor executes the computer program to implement the training steps of the multimodal large model, it specifically implements the following steps:
[0162] Acquire multi-source data and construct a basic dataset based on the multi-source data; perform data augmentation using a controllable diffusion model based on the basic dataset to obtain a sample set; construct an initial multimodal large model; train the initial multimodal large model using the sample set to obtain a multimodal large model.
[0163] The controllable diffusion model includes a semantic control layer, a physical constraint module, and a dynamic noise scheduling layer;
[0164] The semantic control layer encodes the defect type and location parameters into a conditional vector, which is then fused into the diffusion model via a cross-attention mechanism to guide the generation of images with specific defects.
[0165] The physical constraint module is used to embed prior knowledge of the physical structure of the steel wire rope into the UNet architecture and enhance it through weights of the pre-trained segmentation model to ensure that the generated defect images meet the actual physical characteristics and standard requirements.
[0166] Dynamic noise scheduling is used to optimize the noise addition strategy during the diffusion process using adaptive scheduling.
[0167] In one embodiment, when the processor executes the computer program to implement the step of acquiring multi-source data and constructing a basic dataset based on the multi-source data, it specifically implements the following steps:
[0168] Data on fault detection, missed detection, and normal operation are collected synchronously from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs to obtain multi-source data, which includes fault types and corresponding images. The multi-source data is preprocessed and stored. The stored multi-source data is then labeled to obtain a basic dataset.
[0169] In one embodiment, when the processor executes the computer program to implement the step of data augmentation using a controlled diffusion model based on the base dataset to obtain a sample set, it specifically implements the following steps:
[0170] The basic dataset is combined with user input parameters, and a multimodal encoder is used to transform the semantic information of defect category, size, and location into feature vectors readable by the generative model. A controllable diffusion model is then used to automatically generate diverse defect images to obtain a sample set.
[0171] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0172] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0173] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0174] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0175] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0176] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for detecting defects in steel wire ropes, characterized in that, include: Obtain the image to be detected and the prompt text; The image to be detected and the prompt text are input into a multimodal large model for image feature encoding, text feature encoding, modal fusion processing, and localization and classification of different types of defects to obtain the detection results; Output the detection results.
2. The method for detecting defects in steel wire ropes according to claim 1, characterized in that, The multimodal large model includes an image information encoder, a text information encoder, a modality fusion module, a detector module, and a prompt word module; The image information encoder maps the image to be detected to a high-dimensional latent space shared with the natural language description, forming visual features; The text information encoder extracts semantic vectors related to defects from the prompt text to form text features; The modality fusion module fuses visual features and textual features to generate cross-modal representations, thus obtaining fused features; The detector module locates and classifies different types of defects based on these fused features; The prompt word module automatically adjusts attention weights based on business knowledge, improving positioning accuracy and classification accuracy.
3. The method for detecting defects in steel wire ropes according to claim 2, characterized in that, The process of inputting the image to be detected and the prompt text into a multimodal large model for image feature encoding, text feature encoding, modal fusion processing, and locating and classifying different types of defects to obtain detection results includes: The image to be detected is mapped to a high-dimensional latent space shared with the natural language description using EVA-CLIP-G to obtain visual features; The prompt text containing specific defect information is processed by SkilNet-NLG to extract semantic vectors related to the defects in order to obtain text features; By fusing the visual features and the text features through a cross-modal cross-attention mechanism, a consistent cross-modal representation is generated to obtain the fused features; Based on the fusion features, dynamic querying and multi-head cross-attention layer localization and classification of different types of defects are used to obtain detection results.
4. The method for detecting defects in steel wire ropes according to claim 1, characterized in that, The training process of the multimodal large model includes: Acquire multi-source data and construct a basic dataset based on the multi-source data; Based on the aforementioned basic dataset, a controlled diffusion model is used to augment the data to obtain a sample set; Construct an initial large multimodal model; The initial multimodal large model is trained using the sample set to obtain the multimodal large model.
5. The method for detecting defects in wire ropes according to claim 4, characterized in that, The acquisition of multi-source data and the construction of a basic dataset based on the multi-source data include: Data on fault detection, missed detection, and normal operation are collected synchronously from online monitoring systems, manual inspection reports, on-site feedback, and equipment logs to obtain multi-source data, which includes fault types and corresponding images. The multi-source data is preprocessed and stored; The stored multi-source data is labeled to obtain the basic dataset.
6. The method for detecting defects in wire ropes according to claim 4, characterized in that, The controllable diffusion model includes a semantic control layer, a physical constraint module, and a dynamic noise scheduling layer; The semantic control layer encodes the defect type and location parameters into a conditional vector, and integrates it into the diffusion model through the Cross-Attention mechanism to guide the generation of images with specific defects. The physical constraint module is used to embed prior knowledge of the physical structure of the steel wire rope into the UNet architecture and enhance it through the weights of the pre-trained segmentation model to ensure that the generated defect images meet the actual physical characteristics and standard requirements. Dynamic noise scheduling is used to optimize the noise addition strategy during the diffusion process using adaptive scheduling.
7. The method for detecting defects in steel wire ropes according to claim 6, characterized in that, The process of augmenting the data using a controlled diffusion model based on the aforementioned base dataset to obtain a sample set includes: The basic dataset is combined with user input parameters, and a multimodal encoder is used to transform the semantic information of defect category, size, and location into feature vectors readable by the generative model. A controllable diffusion model is then used to automatically generate diverse defect images to obtain a sample set.
8. A wire rope defect detection device, characterized in that, include: The acquisition unit is used to acquire the image to be detected and the prompt text; The detection unit is used to input the image to be detected and the prompt text into a multimodal large model to perform image feature encoding, text feature encoding, modal fusion processing, and to locate and classify different types of defects in order to obtain detection results; The output unit is used to output the detection results.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Cited By
Part surface quality detection method, device, equipment and medium
CN121527086A