Defect detection method, device and equipment for oil-immersed equipment and storage medium
By constructing a defect detection model using CLIP and Mamba modules, the problems of low accuracy and efficiency in oil-immersed equipment detection were solved, achieving high-precision defect identification and ensuring stable equipment operation.
Patent Information
- Application Number
- CN202512042741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-20
AI Technical Summary
Existing defect detection methods for oil-immersed electrical equipment are inaccurate and inefficient, making it difficult to effectively identify minute defects. This can lead to serious malfunctions such as corona discharge, partial discharge, or even insulation breakdown during equipment operation.
A defect detection model based on the CLIP cross-modal vision-language pre-trained model and the Mamba state-space sequence modeling module is constructed. Combined with a visual context-driven adaptive text prompting mechanism, the coupling between image features and defect semantics is strengthened by dynamically generating scene-related text prompts. The Mamba module is introduced to jointly model local details and long-range dependencies, thereby improving the model's discrimination stability and generalization ability under complex conditions.
It enables high-precision, cross-domain robust detection of multiple types of defects in oil-immersed equipment with a small sample size, improving the accuracy and efficiency of defect detection and ensuring stable operation of equipment in complex environments.
Smart Images

Figure CN121708005A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power system technology, and in particular to a defect detection method, apparatus, equipment and storage medium for oil-immersed equipment. Background Technology
[0002] As core equipment in power transmission and distribution systems, the manufacturing quality of oil-immersed electrical equipment directly affects the long-term safe and stable operation of the power grid. During the production process, even minor defects in the oil-paper insulation system, porcelain bushing surface, and weld structure of oil-immersed electrical equipment can gradually evolve into serious faults such as corona discharge, partial discharge, or even insulation breakdown during operation due to electrothermal stress, mechanical vibration, or environmental factors, leading to equipment shutdown or large-scale power outages. Therefore, defect detection during the production process is crucial for ensuring the quality of oil-immersed equipment.
[0003] Traditional defect detection methods mainly rely on manual visual inspection, electrical and withstand voltage tests, and oil sample and leakage inspection. Manual visual inspection is usually carried out by inspectors who visually inspect and manually record each equipment component to detect obvious defects such as surface scratches and cracks. Electrical and withstand voltage tests verify the overall insulation level of the equipment by applying power frequency voltage, impulse voltage, or partial discharge tests, and are used to detect serious internal insulation defects or process defects.
[0004] However, traditional detection methods suffer from poor accuracy and low efficiency in defect detection due to incomplete detection. Summary of the Invention
[0005] This application provides a defect detection method, apparatus, equipment, and storage medium for oil-immersed equipment, which aims to improve the accuracy and efficiency of defect detection.
[0006] In a first aspect, embodiments of this application provide a defect detection method for an oil-immersed device, comprising:
[0007] Acquire the image to be processed from the oil immersion equipment;
[0008] The image to be processed is subjected to image standardization processing to obtain a standard image to be processed. The image standardization processing is used to unify the format of the image data.
[0009] The standard image to be processed is input into a preset defect detection model to obtain the defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image.
[0010] In one or more embodiments, before inputting the image to be processed into a preset defect detection model to obtain defect information of the oil-immersed equipment, the method includes:
[0011] Multiple defect images of the oil-immersed equipment are preprocessed to obtain multiple standard defect images. The preprocessing includes image normalization and data augmentation. The data augmentation is used to add noise and adjust the amplitude of the image data.
[0012] The initial detection model is trained based on the multiple standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model. The initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scanning space unit.
[0013] In one or more embodiments, training an initial detection model based on the plurality of standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model includes:
[0014] S1, For each standard defect image, the standard defect image is input into the adaptive context prompting unit to obtain semantic prior information;
[0015] S2, determine image feature data and text feature data based on the semantic prior information, the standard defect image, and the visual-language unit;
[0016] S3, input the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data;
[0017] S4, Based on the enhanced image feature data and the text feature data, determine the defect information corresponding to the standard defect image;
[0018] S5. Based on the defect information corresponding to the standard defect image and the labeled defect information corresponding to each defect image, update the parameters of the initial detection model, repeat steps S1-S5 until the preset loss converges to the first preset threshold, and obtain the defect detection model.
[0019] In one or more embodiments, the adaptive context prompting unit includes a first image encoder and a linear layer;
[0020] Accordingly, the step of inputting the standard defect image into the adaptive context prompting unit to obtain semantic prior information includes:
[0021] The standard defect image is input into the first image encoder to obtain category feature data, which carries contextual information of the standard defect image.
[0022] The category feature data is input into the linear layer to obtain the semantic prior information.
[0023] In one or more embodiments, the visual-language unit includes a second image encoder and a text encoder;
[0024] Accordingly, determining image feature data and text feature data based on the semantic prior information, the standard defect image, and the visual-language unit includes:
[0025] The semantic prior information and the standard defect image are input into the second image encoder to determine the image feature data;
[0026] The preset text prompt information is input into the text encoder to determine the text prompt data corresponding to the standard defect image;
[0027] The text prompt data and the semantic prior information are concatenated to determine the text feature data, wherein the data dimensions of the text prompt data and the semantic prior information are consistent.
[0028] In one or more embodiments, the hybrid scan space unit includes a hybrid scan encoder, a state space unit, and a hybrid scan decoder;
[0029] Accordingly, inputting the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data includes:
[0030] The image feature data is input into the hybrid scan encoder to obtain hybrid image feature data;
[0031] The hybrid image feature data is input into the state space unit to obtain hidden feature data;
[0032] The hybrid image feature data and the hidden feature data are input into the hybrid scan decoder to obtain enhanced image feature data.
[0033] In one or more embodiments, determining the defect information corresponding to the standard defect image based on the enhanced image feature data and the text feature data includes:
[0034] Determine the similarity value between the enhanced image feature data and the text feature data;
[0035] Based on the similarity value and the second preset threshold, the defect information corresponding to the standard defect image is determined, and the defect information includes text description information.
[0036] Secondly, embodiments of this application provide a defect detection device for an oil-immersed equipment, comprising:
[0037] The acquisition module is used to acquire the image to be processed from the oil immersion equipment;
[0038] The processing module is used to perform image standardization processing on the image to be processed to obtain a standard image to be processed. The image standardization processing is used to unify the format of the image data.
[0039] The detection module is used to input the standard image to be processed into a preset defect detection model to obtain the defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image.
[0040] In one or more embodiments, before inputting the image to be processed into a preset defect detection model to obtain the defect information of the oil-immersed equipment, the detection module is further configured to:
[0041] Multiple defect images of the oil-immersed equipment are preprocessed to obtain multiple standard defect images. The preprocessing includes image normalization and data augmentation. The data augmentation is used to add noise and adjust the amplitude of the image data.
[0042] The initial detection model is trained based on the multiple standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model. The initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scanning space unit.
[0043] In one or more embodiments, the detection module trains an initial detection model based on the plurality of standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model, specifically for:
[0044] S1, For each standard defect image, the standard defect image is input into the adaptive context prompting unit to obtain semantic prior information;
[0045] S2, determine image feature data and text feature data based on the semantic prior information, the standard defect image, and the visual-language unit;
[0046] S3, input the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data;
[0047] S4, Based on the enhanced image feature data and the text feature data, determine the defect information corresponding to the standard defect image;
[0048] S5. Based on the defect information corresponding to the standard defect image and the labeled defect information corresponding to each defect image, update the parameters of the initial detection model, repeat steps S1-S5 until the preset loss converges to the first preset threshold, and obtain the defect detection model.
[0049] In one or more embodiments, the adaptive context prompting unit includes a first image encoder and a linear layer;
[0050] Accordingly, the detection module inputs the standard defect image into the adaptive context prompting unit to obtain semantic prior information, specifically for:
[0051] The standard defect image is input into the first image encoder to obtain category feature data, which carries contextual information of the standard defect image.
[0052] The category feature data is input into the linear layer to obtain the semantic prior information.
[0053] In one or more embodiments, the visual-language unit includes a second image encoder and a text encoder;
[0054] Accordingly, the detection module, based on the semantic prior information, the standard defect image, and the visual-language unit, determines image feature data and text feature data, specifically for:
[0055] The semantic prior information and the standard defect image are input into the second image encoder to determine the image feature data;
[0056] The preset text prompt information is input into the text encoder to determine the text prompt data corresponding to the standard defect image;
[0057] The text prompt data and the semantic prior information are concatenated to determine the text feature data, wherein the data dimensions of the text prompt data and the semantic prior information are consistent.
[0058] In one or more embodiments, the hybrid scan space unit includes a hybrid scan encoder, a state space unit, and a hybrid scan decoder;
[0059] Accordingly, the detection module inputs the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data, specifically for:
[0060] The image feature data is input into the hybrid scan encoder to obtain hybrid image feature data;
[0061] The hybrid image feature data is input into the state space unit to obtain hidden feature data;
[0062] The hybrid image feature data and the hidden feature data are input into the hybrid scan decoder to obtain enhanced image feature data.
[0063] In one or more embodiments, the detection module determines the defect information corresponding to the standard defect image based on the enhanced image feature data and the text feature data, specifically for:
[0064] Determine the similarity value between the enhanced image feature data and the text feature data;
[0065] Based on the similarity value and the second preset threshold, the defect information corresponding to the standard defect image is determined, and the defect information includes text description information.
[0066] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0067] The memory stores computer-executed instructions;
[0068] The processor executes computer execution instructions stored in the memory, such that the processor, when executed, is used to implement the method described in the first aspect and any of the embodiments above.
[0069] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods described in the first aspect and any of the embodiments above.
[0070] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, is used to implement the defect detection method for an oil-immersed device as described in the first aspect and various possible implementations of the first aspect.
[0071] This application provides a method, apparatus, device, and storage medium for defect detection of oil-immersed equipment. The method first acquires an image of the oil-immersed equipment to be processed, then performs image standardization processing on the image to obtain a standard image to be processed. Image standardization processing is used to unify the format of the image data. The standard image to be processed is then input into a preset defect detection model to obtain defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the corresponding labeled defect information for each defect image. In the above method, by standardizing the image to be processed, differences between images from different sources and under different shooting conditions can be eliminated, ensuring that the images subsequently input into the model have a uniform format. The defect detection model constructed based on multiple defect images and their corresponding labeled defect information can effectively identify different defects in the oil-immersed equipment. Since the model is trained based on real data, its generalization ability is enhanced, enabling high detection accuracy in practical applications. The combination of image standardization processing and the defect detection model can effectively improve the defect recognition capability of the oil-immersed equipment, thereby improving the accuracy and efficiency of defect detection. Attached Figure Description
[0072] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0073] Figure 1 A flowchart illustrating the defect detection method for oil-immersed equipment provided in this application embodiment. Figure 1 ;
[0074] Figure 2 A flowchart illustrating the defect detection method for oil-immersed equipment provided in this application embodiment. Figure 2 ;
[0075] Figure 3 Schematic diagram of the structure of the initial detection model provided in the embodiments of this application Figure 1 ;
[0076] Figure 4 Schematic diagram of the structure of the initial detection model provided in the embodiments of this application Figure 2 ;
[0077] Figure 5 Schematic diagram of the structure of the initial detection model provided in the embodiments of this application Figure 3 ;
[0078] Figure 6 This is a schematic diagram of hybrid scanning provided for an embodiment of this application;
[0079] Figure 7 This is a schematic diagram of the structure of the hybrid scanning spatial unit provided in the embodiments of this application;
[0080] Figure 8 A flowchart illustrating the defect detection method for oil-immersed equipment provided in this application embodiment. Figure 3 ;
[0081] Figure 9 A schematic diagram of the defect detection device for an oil-immersed equipment provided in an embodiment of this application;
[0082] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0083] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0084] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0085] Before introducing the embodiments of this application, the application background of the embodiments of this application will be explained first:
[0086] As core equipment in power transmission and distribution systems, the manufacturing quality of oil-immersed electrical equipment directly affects the long-term safe and stable operation of the power grid. During the production process, even minor defects in the oil-paper insulation system, porcelain bushing surface, weld structure, and nameplate markings of oil-immersed electrical equipment can gradually evolve into serious faults such as corona discharge, partial discharge, or even insulation breakdown during operation due to electrothermal stress, mechanical vibration, or environmental factors, leading to equipment shutdowns or large-scale power outages. Therefore, defect detection during the production process is crucial for ensuring the quality of oil-immersed equipment.
[0087] Traditional testing methods primarily rely on manual visual inspection, electrical and withstand voltage tests, and oil sample and leakage checks. Manual visual inspection typically involves inspectors visually inspecting and manually recording the components such as oil tanks, porcelain bushings, flanges, conductive rods, lead terminals, nameplates, and outer casing welds on the production line or at the final inspection station. This is used to identify obvious surface scratches, cracks, rust, oil seepage, missing parts, or incorrect installation. Electrical and withstand voltage tests verify the overall insulation level of the equipment by applying power frequency voltage, impulse voltage, or partial discharge tests, and are used to detect serious internal insulation or manufacturing defects. Oil sample and leakage checks involve observing the tank and bushings under static conditions, applying leak detection agents, and extracting oil samples to check for leaks and to assess the quality of the insulating oil.
[0088] However, manual visual inspection relies on experience and is prone to fatigue and missed detections, electrical and pressure tests can only detect serious defects, and oil sample and leak inspections are time-consuming and difficult to quantify.
[0089] Existing defect detection methods mainly use industrial cameras to collect images of local areas of equipment, and combine algorithms such as grayscale threshold segmentation, edge detection, and morphological operations to identify defects such as oil stains, missing nameplates, and loose bolts. However, they are sensitive to changes in lighting, reflections, and differences in equipment color, and can only detect local areas.
[0090] In summary, existing detection methods suffer from poor defect detection accuracy and low detection efficiency.
[0091] This application provides a defect detection method for oil-immersed equipment, aiming to solve the aforementioned technical problems of existing technologies. The technical concept of this application is as follows: Existing defect detection methods for oil-immersed equipment suffer from poor robustness and insufficient generalization ability, resulting in low detection accuracy and efficiency. To address these issues, this application proposes constructing a defect detection model based on CLIP (a cross-modal visual-language pre-trained model) and Mamba (a state-space sequence modeling module), combined with a visual context-driven adaptive text prompt mechanism, to achieve low-sample, high-precision, and cross-domain robust detection of various types of defects in oil-immersed electrical equipment. The technical solution of this application uses CLIP's cross-modal semantic alignment capability as its core, dynamically generating scene-related text prompts to strengthen the coupling between image features and defect semantics. Simultaneously, the Mamba module is introduced to jointly model local details and long-range dependencies, compensating for the shortcomings of traditional models in identifying minute defects. Finally, through a structured optimization strategy, the model's discrimination stability and generalization ability under complex working conditions are improved, achieving accurate detection of defects in oil-immersed equipment.
[0092] The execution subject of this application embodiment is an electronic device, which can be a terminal device, such as a laptop, desktop computer, or tablet computer, or a server. In practical applications, whether the electronic device is a terminal device or a server can be determined according to the actual situation, and no specific limitation is imposed on it.
[0093] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0094] Figure 1 A flowchart illustrating the defect detection method for oil-immersed equipment provided in this application embodiment. Figure 1 .like Figure 1 As shown, the defect detection method for this oil-immersed equipment includes the following steps:
[0095] S110. Acquire the image to be processed from the oil-immersed equipment;
[0096] In this step, an image to be processed is acquired that clearly reflects the surface and key component condition of the oil-immersed equipment. This image must cover areas of the oil-immersed equipment prone to defects, providing high-quality raw data for subsequent processing.
[0097] In one possible implementation, an inspection robot equipped with a high-definition industrial camera is used to autonomously take pictures during the shutdown or operation intervals of oil-immersed equipment, focusing on key parts such as oil tank welds, sleeve flanges, and conductive rods; alternatively, a fixedly installed monitoring camera can capture images of the equipment surface in real time and automatically select images with sufficient clarity and no obvious reflection as images to be processed.
[0098] S120. Perform image standardization processing on the image to be processed to obtain a standard image to be processed;
[0099] Image standardization is used to unify the format of image data.
[0100] In this step, the images to be processed are standardized to unify the data format and eliminate interference from differences in shooting equipment and environment. Through format regularization and pixel adjustment, images suitable for subsequent model input are provided.
[0101] In one possible implementation, the image standardization process does not change the key information such as the location and shape of the defect. It only unifies the format of the image to be processed. For example, the acquired image to be processed is uniformly converted into RGB three-channel format, cropped or scaled to a fixed size, and then standardized using mean and standard deviation to map the pixel values to the range of [-1, 1]. Finally, slight noise interference in the image is removed to complete the format unification and feature calibration, and output a standard image to be processed.
[0102] S130. Input the standard image to be processed into the preset defect detection model to obtain the defect information of the oil-immersed equipment;
[0103] The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image.
[0104] In this step, the defect detection model, which is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image, can efficiently extract image features and match defect patterns, and output accurate defect judgment results. Therefore, the standard image to be processed is input into the preset defect detection model to directly obtain the defect information of the oil-immersed equipment.
[0105] In one possible implementation, the standard image input to be processed is fused with a contrastive learning-based visual-language pre-training (CLIP) model and a defect detection model from the Mamba module. This defect detection model includes two processing stages:
[0106] In the first stage, semantic prior information is extracted from the input image, scene-related text prompts are dynamically generated, and these prompts are injected into the visual-language CLIP model to determine the text feature data of the embedding context. Image feature data is also extracted through the visual encoder of the CLIP model.
[0107] In the second stage, the image feature data is processed by the Mamba module to enhance the association between local defects and global structure, capture long-range dependencies using state space structure, and improve the sensitivity to fine-grained local details through local enhancement mechanisms.
[0108] Finally, the data is matched with the text feature data embedded in the context to output the defect information of the oil-immersed equipment. The defect information includes whether a defect exists, the type of defect (such as cracks, oil seepage, scratches), the location coordinates of the defect, and the confidence level.
[0109] The defect detection method for oil-immersed equipment provided in this application first acquires an image of the oil-immersed equipment to be processed, then performs image standardization processing on the image to be processed to obtain a standard image to be processed. Image standardization processing is used to unify the format of image data. The standard image to be processed is then input into a preset defect detection model to obtain defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the corresponding labeled defect information for each defect image. In this embodiment, by standardizing the image to be processed, differences between images from different sources and under different shooting conditions can be eliminated, ensuring that the images subsequently input into the model have a uniform format. The defect detection model constructed based on multiple defect images and their corresponding labeled defect information can effectively identify different defects in the oil-immersed equipment. Since the model is trained based on real data, its generalization ability is enhanced, enabling it to achieve high detection accuracy in practical applications. The combination of image standardization processing and the defect detection model can effectively improve the defect recognition capability of the oil-immersed equipment, thereby improving the accuracy and efficiency of defect detection.
[0110] Based on the above embodiments, Figure 2 A flowchart illustrating the defect detection method for oil-immersed equipment provided in this application embodiment. Figure 2 Prior to step S130 above, the defect detection method for this oil-immersed equipment further includes the following steps:
[0111] S210. Preprocess multiple defect images of the oil-immersed equipment to obtain multiple standard defect images;
[0112] The preprocessing includes image normalization and data augmentation, with data augmentation used to add noise and adjust the amplitude of the image data.
[0113] In this step, multiple defect images of the oil-immersion equipment are standardized and data augmentation processing, including noise addition and amplitude adjustment, is performed. Standardization eliminates interference caused by differences in equipment and cameras, while data augmentation simulates changes in the real scene, allowing the model to adapt to complex detection environments.
[0114] In one possible implementation, multiple defect images of the oil-immersed equipment are first standardized, and all defect images are uniformly converted into RGB three-channel format, scaled to a fixed size, and then the pixel values are normalized to the range of [-1, 1].
[0115] Then, data augmentation is performed on the standardized image data, such as adding Gaussian noise or salt-and-pepper noise to simulate sensor interference. By adjusting the amplitude, brightness and contrast are changed to simulate lighting changes. At the same time, edge enhancement is performed on images with minor defects to highlight features such as cracks and oil seepage. Finally, a pre-processed standard defect image is output, thereby improving the robustness of subsequent models to the image.
[0116] S220. Train the initial detection model based on multiple standard defect images and the labeled defect information corresponding to each defect image to obtain a defect detection model;
[0117] The initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scan space unit.
[0118] In this step, the adaptive contextual cueing unit in the initial detection model enhances cross-modal alignment, the visual-language unit extracts core features, the hybrid scanning spatial unit captures details and long-range dependencies, and the parameters of the initial detection model are optimized based on multiple standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model.
[0119] In one possible implementation, multiple standard defect images and the labeled defect information corresponding to each defect image are divided according to a preset ratio to obtain a training set, a validation set, and a test set.
[0120] Figure 3 Schematic diagram of the structure of the initial detection model provided in the embodiments of this application Figure 1 .like Figure 3 As shown, the initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scan space unit.
[0121] An initial model is built based on the PyTorch framework. Standard defect images from the training set are input into the visual-language unit to extract image features. The adaptive context prompting unit dynamically generates and encodes text prompts based on the global features of the image. The hybrid scan space unit performs multi-directional scanning and state space modeling of the image features to enhance the features of minor defects.
[0122] The model parameters are iteratively trained and optimized using the labeled defect category and location as the true labels, with the defect recognition accuracy on the validation set stabilizing above 95%. Training is then stopped to obtain the defect detection model.
[0123] Finally, the defect detection model was tested using a test set to obtain the evaluation index data of the model.
[0124] In one possible implementation, step S220 described above may include the following steps:
[0125] S1, For each standard defect image, input the standard defect image into the adaptive context prompting unit to obtain semantic prior information;
[0126] For example, the adaptive context prompting unit parses the global features of a standard defect image, mines key information such as device type, defect location, and background environment, and forms targeted semantic prior information to avoid the problem of insufficient generalization of general text prompts.
[0127] In one possible implementation, the adaptive context prompting unit includes a first image encoder and a linear layer;
[0128] Accordingly, one possible implementation of step S1 above includes the following steps:
[0129] S11. Input the standard defect image into the first image encoder to obtain category feature data;
[0130] Among them, the category feature data carries contextual information of the standard defect image;
[0131] For example, the first image encoder uses a CLIP visual encoder with frozen parameters, which can accurately capture contextual information such as device type, defect location, and surface condition in the image without additional training by utilizing its large-scale pre-trained feature extraction capabilities, ensuring the generalization and correlation of category feature data.
[0132] In one possible implementation, a standardized image of a "weld crack in an oil-immersed transformer tank" defect is input into a first image encoder. This first image encoder then extracts global semantic features of the image through multi-layer convolution and attention mechanisms, and outputs category feature data. .
[0133] Category feature data It includes equipment and area information such as "oil-immersed transformer" and "tank weld", as well as scene context such as "crack morphology" and "surface reflectivity", providing comprehensive support for subsequent semantic prior generation.
[0134] S12. Input the category feature data into the linear layer to obtain semantic prior information.
[0135] For example, the linear layer performs dimensionality transformation and semantic extraction on the categorical feature data, mapping high-dimensional image features into semantic prior information that can be directly used for text prompt generation.
[0136] The core function of the linear layer is to align feature dimensions and enhance key contextual cues, ensuring that the semantic information output accurately matches the generation requirements of subsequent text prompts, thereby achieving deep coupling between images and text.
[0137] In one possible implementation, Figure 4 Schematic diagram of the structure of the initial detection model provided in the embodiments of this application Figure 2 .like Figure 4 As shown, the adaptive contextual cueing unit includes a first image encoder and a linear layer.
[0138] The standard defect image is input into the first image encoder (such as a CLIP visual encoder with frozen parameters) to obtain category feature data. This feature, after undergoing a linear transformation through a linear layer, is projected into a semantic vector embedding visual context information. Then, key features are filtered through activation functions, and finally semantic prior information is output.
[0139] S2, based on semantic prior information, standard defect images, and visual-language units, determine image feature data and text feature data;
[0140] For example, the visual-language unit extracts image features from a standard defect image to obtain image feature data, and combines this with semantic prior information to determine text feature data.
[0141] The visual-language unit encodes the image and the adaptively generated text prompts respectively, mapping them to the same feature space, providing a unified feature support for subsequent defect matching.
[0142] In one possible implementation, the visual-language unit includes a second image encoder and a text encoder;
[0143] Accordingly, one possible implementation of step S2 above includes the following steps:
[0144] S21. Input the semantic prior information and standard defect image into the second image encoder to determine the image feature data;
[0145] For example, the second image encoder uses a CLIP visual encoder (with finely adjustable parameters) to fuse semantic priors with original image information and extract more discriminative image features.
[0146] Injecting semantic prior information allows the encoder to focus on defect-related regions, avoid interference from irrelevant backgrounds, strengthen the association between defect features and device context, and improve the accuracy of feature representation.
[0147] In one possible implementation, semantic prior information (such as oil-immersed sleeves, flange areas, and linear cracks) is input into the CLIP visual encoder along with the standardized crack defect image. The semantic cues are fused with the image pixel information through an attention mechanism, and the features such as crack texture and edge morphology of the flange area are extracted in a key way. The output is high-dimensional image feature data, which also includes equipment context, defect location and morphology information.
[0148] S22. Input the preset text prompt information into the text encoder to determine the text prompt data corresponding to the standard defect image;
[0149] For example, the text encoder converts preset text prompts into vector data that can be matched with image features. The preset text prompts cover device type, defect status and category, providing a semantic basis for cross-modal alignment.
[0150] The text encoder uses the CLIP text encoder to ensure that the output text prompt data and image feature data are in the same feature space, which meets the requirements of subsequent fusion.
[0151] S23. Concatenate the text prompt data and semantic prior information to determine the text feature data;
[0152] The text prompt data and semantic prior information have the same data dimensions.
[0153] For example, text prompt data and semantic prior information with consistent data dimensions are concatenated to achieve deep fusion of text semantics and image context. After concatenation, complete semantic information can be retained, so that the text feature data has both preset semantics and image-derived contextual clues, thereby enhancing the accuracy of cross-modal matching.
[0154] In one possible implementation, the concatenation of text prompt data and semantic prior information is performed within the adaptive context prompt unit.
[0155] Category feature data The semantic prior information obtained by dimensionality matching projection through a linear layer is concatenated with the text prompt data output by the text encoder of the vision-language unit. This text prompt data includes the text prompt for "abnormal state". And the text prompt for "normal status" The concatenated multimodal representations are then subjected to linear mapping to obtain the final optimized text feature data. This provides more discriminative text features for subsequent cross-modal alignment and anomaly detection.
[0156] Category feature data It can be represented as:
[0157]
[0158] in, This represents a standard defect image.
[0159] Semantic vectors embedding visual context information It can be represented as:
[0160]
[0161] Text feature data It can be represented as:
[0162]
[0163]
[0164]
[0165] In one possible implementation, Figure 5 Schematic diagram of the structure of the initial detection model provided in the embodiments of this application Figure 3 .like Figure 5 As shown, the visual-language unit includes a second image encoder and a text encoder. Semantic prior information and a standard defect image are input into the second image encoder to determine image feature data. Simultaneously, preset text prompt information is input into the text encoder to determine the text prompt data corresponding to the standard defect image. Further, the image feature data is input into the hybrid scanning space unit.
[0166] S3, input the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data;
[0167] For example, the hybrid scanning spatial unit focuses on enhancing the detailed representation and long-range correlation of image feature data, making up for the problem that the original image features are insufficient in capturing small defects. Through multi-directional scanning and state space modeling, it simultaneously mines local defect details and global structural information, and outputs more recognizable enhanced image feature data.
[0168] In one possible implementation, the hybrid scan space unit includes a hybrid scan encoder, a state space unit, and a hybrid scan decoder;
[0169] Accordingly, one possible implementation of step S3 above includes the following steps:
[0170] S31. Input the image feature data into the hybrid scan encoder to obtain hybrid image feature data;
[0171] For example, the hybrid scanning encoder, which addresses the characteristics of discrete defect distribution and edge-prone faults in oil-immersed equipment, employs a multi-directional scanning strategy to rearrange image features, complementing the CLIP visual encoder's focus on the central region, thus ensuring that edge defects are not missed.
[0172] In one possible implementation, Figure 6 This is a schematic diagram of hybrid scanning provided for an embodiment of this application. For example... Figure 6 As shown, the image feature data is split into a sequence of feature blocks of fixed size, which are then input into a hybrid scanning encoder to rearrange and encode the feature blocks. The feature weights of edge areas such as oil tank welds and sleeve flanges are emphasized, and the final output is hybrid image feature data with consistent dimensions, highlighting defect association information in different directions.
[0173] S32. Input the hybrid image feature data into the state space unit to obtain the hidden feature data;
[0174] For example, the state space unit, as the core of the Mamba module, has the ability to process sequences with "memory" and can realize the correlation modeling between local defects and global structure by capturing the long-range dependencies of mixed image feature data.
[0175] State-space units, by explicitly modeling the dynamic process of "input → state → output", can capture long-range dependencies and global context information while maintaining computational efficiency, making them suitable for efficient modeling of long sequence data (such as speech, time series, image patch sequences, etc.).
[0176] In one possible implementation, the mixed image feature data is input into the state space unit in the scanning order, and then processed through state equations. Update the internal state by merging information from historical feature blocks and the current feature block.
[0177] S33. Input the hybrid image feature data and hidden feature data into the hybrid scan decoder to obtain enhanced image feature data.
[0178] For example, the hybrid scanning decoder restores the hidden feature data to the spatial layout corresponding to the original image features, and then fuses it with the hybrid image feature data, which not only preserves local details, but also incorporates global correlation information, thereby improving the recognition of defect features.
[0179] In one possible implementation, the hybrid scan decoder first reverse-engineers the hidden feature data along the original scan direction to obtain a spatially aligned feature map. This map is then element-wise added to the hybrid image feature data, retaining the original feature information through residual connections while simultaneously superimposing long-range dependent features. The final output is enhanced image feature data, which highlights details of local defects such as minute cracks and slight oil seepage, while also containing information about the overall structure of the equipment, providing strong discriminative feature support for subsequent defect determination.
[0180] In one possible implementation, Figure 7 This is a schematic diagram of the structure of the hybrid scanning spatial unit provided in an embodiment of this application. Figure 7 As shown, the hybrid scan spatial unit mainly includes a hybrid scan encoder, state space units (SSMs), and a hybrid scan decoder. It also includes two Layer Normalization (LN) layers, two LinearLayer layers, and one Depth-wise Convolution layer.
[0181] The enhanced image feature data output by the hybrid scan spatial cell is represented as follows:
[0182]
[0183] in, This represents the input image feature data.
[0184] S4. Determine the defect information corresponding to the standard defect image based on the enhanced image feature data and text feature data;
[0185] For example, by leveraging the semantic correlation between enhanced image features and text features, and calculating their similarity in a unified feature space, defect information including key information such as defect category and location can be output.
[0186] In one possible implementation, the similarity between the enhanced image feature data and the embeddings of "abnormal state" and "normal state" in the text feature data is calculated. If the similarity with the text feature of "oil seepage anomaly" is higher than 0.85 (reference value), an oil seepage defect is determined to exist. At the same time, the defect area is located based on the feature response location, and the defect category (oil seepage), location coordinates (flange area (x1, y1, x2, y2)) and confidence level (0.92) are output to form complete defect information.
[0187] In one possible implementation, step S4 above may include the following steps:
[0188] S41. Determine the similarity value between the enhanced image feature data and the text feature data;
[0189] For example, the similarity of cross-modal features between enhanced image feature data and text feature data is calculated to establish the association between image defects and text semantics.
[0190] Enhanced image feature data contains local defect details and global structural information, while text feature data integrates preset defect semantics and image context. Since the two are in the same feature space, the similarity value can directly reflect the degree of defect matching.
[0191] In one possible implementation, a cosine similarity algorithm is used to calculate the similarity score between the enhanced image feature data and the text feature data. If the image contains obvious cracks or defects, the output similarity score is 0.91; if it is a normal image, the similarity score with the abnormal text features is 0.35.
[0192] S42. Determine the defect information corresponding to the standard defect image based on the similarity value and the second preset threshold.
[0193] The defect information includes textual descriptions.
[0194] For example, when the similarity value exceeds the second preset threshold, it is determined that the standard defect image has a corresponding defect and the defect information, including the text description information, is output completely; if the similarity value does not exceed the second preset threshold, the defect information of the standard defect image is determined to be normal.
[0195] The second preset threshold is calibrated based on validation set data to balance the false negative rate and false positive rate. At the same time, the text description information is associated with process standards to improve the interpretability of the results.
[0196] In one possible implementation, a second preset threshold is set to 0.85, and a similarity value of 0.91 (higher than the second preset threshold) indicates the presence of a crack defect. The output defect information includes: defect category (crack), location (sleeve flange area (x1, y1, x2, y2)), confidence level (0.91), and text description information: "A linear crack exists in the oil-immersed sleeve flange area, which meets the judgment rule in the process standard that 'cracks on the equipment surface with a length exceeding 0.3mm must be marked'". If the similarity value is 0.35 (lower than the second preset threshold), it is judged as normal, and the text description information is: "No abnormal defects on the equipment surface, which meets the production quality standards".
[0197] S5. Based on the defect information corresponding to the standard defect image and the labeled defect information corresponding to each defect image, update the parameters of the initial detection model. Repeat steps S1-S5 until the preset loss converges to the first preset threshold to obtain the defect detection model.
[0198] For example, a loss function is used to measure the difference between the predicted defect information and the labeled information. The larger the difference, the larger the loss value; the smaller the difference, the smaller the loss value. The model parameters are adjusted through backpropagation, and the training is iterated until the loss converges, ensuring that the model can still stably identify multiple types of defects under conditions of few samples.
[0199] In one possible implementation, the focus loss function is used to calculate the loss value between the predicted defect information and the labeled defect information (category: oil seepage, location: flange area (x1, y1, x2, y2)), and the parameters of the adaptive context cue unit, the visual-language unit, and the hybrid scan space unit are updated through backpropagation.
[0200] Focus loss is an improved loss method based on cross-entropy loss, primarily addressing the problems of class imbalance and easy-to-classify samples dominating training. In scenarios with severe positive / negative sample imbalance (such as foreground / background in object detection, or normal / abnormal in defect detection), ordinary cross-entropy is easily dominated by a large number of "easy-to-classify samples"—these samples are already accurately predicted, but their sheer number crowds out the learning space for "difficult samples." Focus loss addresses this by multiplying the cross-entropy by a decay factor related to prediction confidence, automatically reducing the loss weight of samples that have been correctly and with high confidence classification, while maintaining a higher weight for inaccurately predicted or difficult-to-distinguish samples.
[0201] Focus loss between the defect information corresponding to the standard defect image and the labeled defect information corresponding to each defect image. It is expressed as follows:
[0202]
[0203] in, This represents the defect information corresponding to the standard defect image. This represents the labeled defect information corresponding to each defect image.
[0204] Further, repeat steps S1-S5, using a validation set to evaluate performance after each training round. When the focus loss value converges to the first preset threshold of 0.05 and does not decrease for three consecutive rounds, stop training to obtain the final defect detection model.
[0205] The defect detection method for oil-immersed equipment provided in this application first preprocesses multiple defect images of the oil-immersed equipment to obtain multiple standard defect images. The preprocessing includes image normalization and data augmentation. The data augmentation is used to add noise and adjust the amplitude of the image data. Then, an initial detection model is trained based on the multiple standard defect images and the labeled defect information corresponding to each defect image to obtain a defect detection model. The initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scanning space unit. In this embodiment, standardization processing eliminates the influence of different shooting conditions, equipment, and environmental factors, ensuring that all defect images are consistent in size, brightness, contrast, etc. This consistency is crucial for improving the training effect and detection accuracy of subsequent models. Data augmentation techniques such as noise addition and amplitude adjustment can effectively increase the diversity of training data, helping the model learn more defect features and thus improving the model's generalization ability. Training the initial detection model by combining multiple standard defect images and their annotation information ensures that the model fully learns the features of different types of defects, improving the defect recognition accuracy. The adaptive contextual cueing unit can dynamically adjust the detection strategy according to the contextual information of the input image, thereby improving the detection accuracy. The visual-language unit is used to combine image information with text descriptions, helping the model better understand the nature and background of defects. The hybrid scan space unit can promote feature extraction from different regions of the image, enhance the ability to recognize defects in complex backgrounds, and improve the overall performance of the model.
[0206] Based on the above embodiments, Figure 8 A flowchart illustrating the defect detection method for oil-immersed equipment provided in this application embodiment. Figure 3 Combining Figure 8 The specific process of the defect detection method for oil-immersed equipment provided in the embodiments of this application is described below:
[0207] S810: Obtain multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image;
[0208] For example, multiple standard defect images and the labeled defect information corresponding to each defect image are divided according to a preset ratio to obtain a training set, a validation set, and a test set.
[0209] S820. Preprocess multiple defect images in the training set to obtain multiple standard defect images;
[0210] S830. Train the initial detection model using multiple defect images in the training set and the labeled defect information corresponding to each defect image.
[0211] S840. Use the validation set to evaluate and validate the initial detection model after training;
[0212] S850. Determine whether the performance of the initial detection model after training meets the preset standard. If yes, obtain the defect detection model and execute step S860. Otherwise, adjust the parameters of the initial detection model and continue to execute step S830.
[0213] S860. Use the test set to test the defect detection model and obtain model evaluation index data.
[0214] Based on the above embodiments, the following are embodiments of the apparatus involved in this application:
[0215] Figure 9 This is a schematic diagram of the defect detection device for an oil-immersed equipment provided in an embodiment of this application. Figure 9 As shown, the defect detection device 900 for the oil-immersed equipment includes:
[0216] The acquisition module 910 is used to acquire the image to be processed from the oil immersion equipment.
[0217] The processing module 920 is used to perform image standardization processing on the image to be processed to obtain a standard image to be processed. The image standardization processing is used to unify the format of the image data.
[0218] The detection module 930 is used to input the standard image to be processed into the preset defect detection model to obtain the defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image.
[0219] In one or more embodiments, before inputting the image to be processed into a preset defect detection model to obtain defect information of the oil-immersed equipment, the detection module 930 is further configured to:
[0220] Multiple defect images of oil-immersed equipment are preprocessed to obtain multiple standard defect images. The preprocessing includes image normalization and data augmentation. Data augmentation is used to add noise and adjust the amplitude of the image data.
[0221] The initial detection model is trained based on multiple standard defect images and the labeled defect information corresponding to each defect image to obtain a defect detection model. The initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scan space unit.
[0222] In one or more embodiments, the detection module 930 trains an initial detection model based on multiple standard defect images and the labeled defect information corresponding to each defect image to obtain a defect detection model, specifically used for:
[0223] S1, For each standard defect image, input the standard defect image into the adaptive context prompting unit to obtain semantic prior information;
[0224] S2, based on semantic prior information, standard defect images, and visual-language units, determine image feature data and text feature data;
[0225] S3, input the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data;
[0226] S4. Determine the defect information corresponding to the standard defect image based on the enhanced image feature data and text feature data;
[0227] S5. Based on the defect information corresponding to the standard defect image and the labeled defect information corresponding to each defect image, update the parameters of the initial detection model. Repeat steps S1-S5 until the preset loss converges to the first preset threshold to obtain the defect detection model.
[0228] In one or more embodiments, the adaptive context prompting unit includes a first image encoder and a linear layer;
[0229] Correspondingly, the detection module 930 inputs the standard defect image into the adaptive context prompting unit to obtain semantic prior information, specifically used for:
[0230] A standard defect image is input into the first image encoder to obtain category feature data, wherein the category feature data carries contextual information of the standard defect image;
[0231] The categorical feature data is input into the linear layer to obtain semantic prior information.
[0232] In one or more embodiments, the visual-language unit includes a second image encoder and a text encoder;
[0233] Accordingly, the detection module 930, based on semantic prior information, standard defect images, and visual-linguistic units, determines image feature data and text feature data, specifically for:
[0234] The semantic prior information and standard defect images are input into the second image encoder to determine the image feature data;
[0235] Input the preset text prompt information into the text encoder to determine the text prompt data corresponding to the standard defect image;
[0236] Text prompt data and semantic prior information are concatenated to determine text feature data, where the data dimensions of text prompt data and semantic prior information are consistent.
[0237] In one or more embodiments, the hybrid scan space unit includes a hybrid scan encoder, a state space unit, and a hybrid scan decoder;
[0238] Correspondingly, the detection module 930 inputs the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data, specifically used for:
[0239] Image feature data is input into a hybrid scan encoder to obtain hybrid image feature data;
[0240] The mixed image feature data is input into the state space unit to obtain the hidden feature data;
[0241] The hybrid image feature data and hidden feature data are input into the hybrid scan decoder to obtain enhanced image feature data.
[0242] In one or more embodiments, the detection module 930 determines the defect information corresponding to the standard defect image based on the enhanced image feature data and text feature data, specifically for:
[0243] Determine the numerical similarity between the enhanced image feature data and the text feature data;
[0244] Based on the similarity value and the second preset threshold, the defect information corresponding to the standard defect image is determined, wherein the defect information includes text description information.
[0245] The defect detection device for oil-immersed equipment provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0246] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 10 As shown, the electronic device 1000 includes: a processor 1010, a memory 1020, and a bus 1030;
[0247] The memory 1020 is used to store computer-executed instructions from the processor 1010;
[0248] The processor 1010 is configured to execute the technical solutions of any of the foregoing method embodiments by executing computer execution instructions.
[0249] The specific implementation process of processor 1010 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0250] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0251] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0252] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0253] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0254] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0255] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0256] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0257] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0258] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0259] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0260] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory (RAM), magnetic disks, or optical disks.
[0261] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0262] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A defect detection method for oil-immersed equipment, characterized in that, include: Acquire the image to be processed from the oil immersion equipment; The image to be processed is subjected to image standardization processing to obtain a standard image to be processed. The image standardization processing is used to unify the format of the image data. The standard image to be processed is input into a preset defect detection model to obtain the defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image.
2. The method according to claim 1, characterized in that, Before inputting the image to be processed into a preset defect detection model to obtain the defect information of the oil-immersed equipment, the method includes: Multiple defect images of the oil-immersed equipment are preprocessed to obtain multiple standard defect images. The preprocessing includes image normalization and data augmentation. The data augmentation is used to add noise and adjust the amplitude of the image data. The initial detection model is trained based on the multiple standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model. The initial detection model includes an adaptive contextual cueing unit, a visual-language unit, and a hybrid scanning space unit.
3. The method according to claim 2, characterized in that, The step of training an initial detection model based on the plurality of standard defect images and the labeled defect information corresponding to each defect image to obtain the defect detection model includes: S1, For each standard defect image, the standard defect image is input into the adaptive context prompting unit to obtain semantic prior information; S2, determine image feature data and text feature data based on the semantic prior information, the standard defect image, and the visual-language unit; S3, input the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data; S4, Based on the enhanced image feature data and the text feature data, determine the defect information corresponding to the standard defect image; S5. Based on the defect information corresponding to the standard defect image and the labeled defect information corresponding to each defect image, update the parameters of the initial detection model, repeat steps S1-S5 until the preset loss converges to the first preset threshold, and obtain the defect detection model.
4. The method according to claim 3, characterized in that, The adaptive context prompting unit includes a first image encoder and a linear layer; Accordingly, the step of inputting the standard defect image into the adaptive context prompting unit to obtain semantic prior information includes: The standard defect image is input into the first image encoder to obtain category feature data, which carries contextual information of the standard defect image. The category feature data is input into the linear layer to obtain the semantic prior information.
5. The method according to claim 3 or 4, characterized in that, The visual-language unit includes a second image encoder and a text encoder; Accordingly, determining image feature data and text feature data based on the semantic prior information, the standard defect image, and the visual-language unit includes: The semantic prior information and the standard defect image are input into the second image encoder to determine the image feature data; The preset text prompt information is input into the text encoder to determine the text prompt data corresponding to the standard defect image; The text prompt data and the semantic prior information are concatenated to determine the text feature data, wherein the data dimensions of the text prompt data and the semantic prior information are consistent.
6. The method according to claim 3 or 4, characterized in that, The hybrid scan space unit includes a hybrid scan encoder, a state space unit, and a hybrid scan decoder; Accordingly, inputting the image feature data into the hybrid scanning spatial unit to obtain enhanced image feature data includes: The image feature data is input into the hybrid scan encoder to obtain hybrid image feature data; The hybrid image feature data is input into the state space unit to obtain hidden feature data; The hybrid image feature data and the hidden feature data are input into the hybrid scan decoder to obtain enhanced image feature data.
7. The method according to claim 3 or 4, characterized in that, The step of determining the defect information corresponding to the standard defect image based on the enhanced image feature data and the text feature data includes: Determine the similarity value between the enhanced image feature data and the text feature data; Based on the similarity value and the second preset threshold, the defect information corresponding to the standard defect image is determined, and the defect information includes text description information.
8. A defect detection device for oil-immersed equipment, characterized in that, include: The acquisition module is used to acquire the image to be processed from the oil immersion equipment; The processing module is used to perform image standardization processing on the image to be processed to obtain a standard image to be processed. The image standardization processing is used to unify the format of the image data. The detection module is used to input the standard image to be processed into a preset defect detection model to obtain the defect information of the oil-immersed equipment. The defect detection model is determined based on multiple defect images of the oil-immersed equipment and the labeled defect information corresponding to each defect image.
9. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.